Planet Python
Last update: September 03, 2021 10:40 AM UTC
September 03, 2021
Kushal Das
Default values, documentation and Ansible
While testing my qubes_ansible project on the upcoming Qubes OS 4.1 project, I noticed something really strange. But, before getting into that, this Ansible module and the connection plugin are for Qubes OS only, and based on the excellent Python modules provided by the Qubes team.
The error goes like this during the fact gathering steps (reformatted for the blog):
fatal: [debian-10]: UNREACHABLE! => {
"changed": false,
"msg": "Failed to create temporary directory.In some cases, you may have
been able to authenticate and did not have permissions on the target
directory. Consider changing the remote tmp path in ansible.cfg to a path
rooted in \"/tmp\", for more error information use -vvv. Failed command
was: ( umask 77 && mkdir -p \"` echo ~The *user* is the default user in
Qubes./.ansible/tmp `\"&& mkdir \"` echo ~The *user* is the default user in
Qubes./.ansible/tmp/ansible-tmp-1630548982.9355698-7707-90110802425258 `\"
&& echo ansible-tmp-1630548982.9355698-7707-90110802425258=\"` echo ~The
*user* is the default user in
Qubes./.ansible/tmp/ansible-tmp-1630548982.9355698-7707-90110802425258 `\"
), exited with result 1, stderr output: mkdir: cannot create directory
‘~The ~The *user* account as default in Qubes OS. ~The ~The *user* account
as default in Qubes OS. account as default in Qubes OS. ~The ~The *user*
account as default in Qubes OS. ~The ~The *user* account as default in
Qubes OS. account as default in Qubes OS. account as default in Qubes OS.
is the default user in Qubes.’: File name too long\n",
"unreachable": true
}
Most important part is the default user's home directory part, echo ~The
user is the default user in Qubes./.ansible/tmp. For a moment I totally
freaked out, as this looks like documentation. After reading the code more, I
can see it is coming from the DOCUMENTATION variable in my plugin. After
playing around a bit more and trying out different values I can see that the
default value mentioned in the documentation is becoming the default value in
the Python code.
After searching more I can see that the Ansible developers want the documentation string to be the gold standard and the code is parsing it find the default values. In my mind this is more confusing. I would expect the default value to be declared inside of the code.
Parsing the DOCUMENTATION and then finding the default values there in a
Python code still does not fit in my brain. Fixed the issue for now, let me see
what other surprises are waiting in the future.
September 03, 2021 02:21 AM UTC
Podcast.__init__
Monitor The Health Of Your Machine Learning Products In Production With Evidently
You've got a machine learning model trained and running in production, but that's only half of the battle. Are you certain that it is still serving the predictions that you tested? Are the inputs within the range of tolerance that you designed? Monitoring machine learning products is an essential step of the story so that you know when it needs to be retrained against new data, or parameters need to be adjusted. In this episode Emeli Dral shares the work that she and her team at Evidently are doing to build an open source system for tracking and alerting on the health of your ML products in production. She discusses the ways that model drift can occur, the types of metrics that you need to track, and what to do when the health of your system is suffering. This is an important and complex aspect of the machine learning lifecycle, so give it a listen and then try out Evidently for your own projects.
Summary
You’ve got a machine learning model trained and running in production, but that’s only half of the battle. Are you certain that it is still serving the predictions that you tested? Are the inputs within the range of tolerance that you designed? Monitoring machine learning products is an essential step of the story so that you know when it needs to be retrained against new data, or parameters need to be adjusted. In this episode Emeli Dral shares the work that she and her team at Evidently are doing to build an open source system for tracking and alerting on the health of your ML products in production. She discusses the ways that model drift can occur, the types of metrics that you need to track, and what to do when the health of your system is suffering. This is an important and complex aspect of the machine learning lifecycle, so give it a listen and then try out Evidently for your own projects.
Announcements
- Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
- When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
- Your host as usual is Tobias Macey and today I’m interviewing Emeli Dral about monitoring machine learning models in production with Evidently
Interview
- Introductions
- How did you get introduced to Python?
- Can you describe what Evidently is and the story behind it?
- What are the metrics that are useful for determining the performance and health of a machine learning model?
- What are the questions that you are trying to answer with those metrics?
- How does monitoring of machine learning models compare to monitoring of infrastructure or "traditional" software projects?
- What are the failure modes for a model?
- Can you describe the design and implementation of Evidently?
- How has the architecture changed or evolved since you started working on it?
- What categories of model is Evidently designed to work with?
- What are some strategies for making models conducive to monitoring?
- What is involved in monitoring a model on a continuous basis?
- What are some considerations when establishing useful thresholds for metrics to alert on?
- Once an alert has been triggered what is the process for resolving it?
- If the training process takes a long time, how can you mitigate the impact of a model failure until the new/updated version is deployed?
- What are the most interesting, innovative, or unexpected ways that you have seen Evidently used?
- What are the most interesting, unexpected, or challenging lessons that you have learned while working on Evidently?
- When is Evidently the wrong choice?
- What do you have planned for the future of Evidently?
Keep In Touch
- @EmeliDral on Twitter
- emeli-dral on GitHub
Picks
- Tobias
Links
The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA
September 03, 2021 01:03 AM UTC
September 02, 2021
Python for Beginners
Check If a List has Duplicate Elements
Lists are the most used data structures in Python. While programming, you may land into a situation where you will need a list containing only unique elements or you want to check if a list has duplicate elements. In this article, we will look at different ways to check if a list has duplicate elements in it.
Check if a list has duplicate Elements using Sets
We know that sets in Python contain only unique elements. We can use this property of sets to check if a list has duplicate elements or not.
For this, we will create a set from the elements of the list. After that, we will check the size of the list and the set. If the size of both the objects are equal, it will confirm that the list has no duplicate elements. If the size of the set is greater than the list, it will mean that the list contains duplicate elements. We can understand this from the following example.
def check_duplicate(l):
mySet = set(l)
if len(mySet) == len(l):
print("List has no duplicate elements.")
else:
print("The list contains duplicate elements")
list1 = [1, 2, 3, 4, 5, 6, 7]
print("List1 is:", list1)
check_duplicate(list1)
list2 = [1, 2, 1, 2, 4, 6, 7]
print("List2 is:", list2)
check_duplicate(list2)
Output:
List1 is: [1, 2, 3, 4, 5, 6, 7]
List has no duplicate elements.
List2 is: [1, 2, 1, 2, 4, 6, 7]
The list contains duplicate elements
In the above approach, we need to create a set from all the elements of the list. After that, we also check the size of the set and the list. These operations are very costly.
Instead of using this approach, we can search only for the first duplicate element. To do this, we will start from the first element of the list and will keep adding them to the set. Before adding the elements to the set, we will check if the element is already present in the set or not. If yes, the list contains duplicate elements. If we are able to add each element of the list to the set, the list does not contain any duplicate element. This can be understood from the following example.
def check_duplicate(l):
visited = set()
has_duplicate = False
for element in l:
if element in visited:
print("The list contains duplicate elements.")
has_duplicate = True
break
else:
visited.add(element)
if not has_duplicate:
print("List has no duplicate elements.")
list1 = [1, 2, 3, 4, 5, 6, 7]
print("List1 is:", list1)
check_duplicate(list1)
list2 = [1, 2, 1, 2, 4, 6, 7]
print("List2 is:", list2)
check_duplicate(list2)
Output:
List1 is: [1, 2, 3, 4, 5, 6, 7]
List has no duplicate elements.
List2 is: [1, 2, 1, 2, 4, 6, 7]
The list contains duplicate elements.
Check if a list has duplicate elements using the count() method
To check if a list has only unique elements, we can also count the occurrence of the different elements in the list. For this, we will use the count() method. The count() method, when invoked on a list, takes the element as input argument and returns the number of times the element is present in the list.
For checking if the list contains duplicate elements, we will count the frequency of each element. At the same time, we will also maintain a list of visited elements so that we don’t have to count the occurrences of the visited elements. Once the count of any element is found to be greater than one, it will prove that the list has duplicate elements. We can implement this as follows.
def check_duplicate(l):
visited = set()
has_duplicate = False
for element in l:
if element in visited:
pass
elif l.count(element) == 1:
visited.add(element)
elif l.count(element) > 1:
has_duplicate = True
print("The list contains duplicate elements.")
break
if not has_duplicate:
print("List has no duplicate elements.")
list1 = [1, 2, 3, 4, 5, 6, 7, 8]
print("List1 is:", list1)
check_duplicate(list1)
list2 = [1, 2, 1, 2, 4, 6, 7, 8]
print("List2 is:", list2)
check_duplicate(list2)
Output:
List1 is: [1, 2, 3, 4, 5, 6, 7, 8]
List has no duplicate elements.
List2 is: [1, 2, 1, 2, 4, 6, 7, 8]
The list contains duplicate elements.
Check if a list has duplicate elements using the counter() method
We can also use the counter() method to check if a list has only unique elements or not. The counter() method. The counter() method takes an iterable object as an input and returns a python dictionary in which the keys consist of the elements of the iterable object and the values associated with the keys are the frequency of the elements. After getting the frequency of each element of the list using the counter() method, we can check if the frequency of any element is greater than one or not. If yes, the list contains duplicate elements. Otherwise not.
from collections import Counter
def check_duplicate(l):
counter = Counter(l)
has_duplicate = False
frequencies = counter.values()
for i in frequencies:
if i > 1:
has_duplicate = True
print("The list contains duplicate elements.")
break
if not has_duplicate:
print("List has no duplicate elements.")
list1 = [1, 2, 3, 4, 5, 6, 7, 8]
print("List1 is:", list1)
check_duplicate(list1)
list2 = [1, 2, 1, 2, 4, 6, 7, 8]
print("List2 is:", list2)
check_duplicate(list2)
Output:
List1 is: [1, 2, 3, 4, 5, 6, 7, 8]
List has no duplicate elements.
List2 is: [1, 2, 1, 2, 4, 6, 7, 8]
The list contains duplicate elements.
Conclusion
In this article, we have discussed four ways to check if a list has only unique elements or not. We have used sets, count() and counter() methods to implement our approaches. To learn more about lists, you can read this article on list comprehension.
The post Check If a List has Duplicate Elements appeared first on PythonForBeginners.com.
September 02, 2021 03:19 PM UTC
Mike Driscoll
Creating a File Search GUI with wxPython
Have you ever needed to search for a file on your computer? Most operating systems have a way to do this. Windows Explorer has a search function and there’s also a search built-in to the Start Menu now. Other operating systems like Mac and Linux are similar. There are also applications that you can download that are sometimes faster at searching your hard drive than the built-in ones are.
In this article, you will be creating a simple file search utility using wxPython.
You will want to support the following tasks for the file search tool:
- Search by file type
- Case sensitive searches
- Search in sub-directories
You can download the source code from this article on GitHub.
Let’s get started!
Designing Your File Search Utility
It is always fun to try to recreate a tool that you use yourself. However in this case, you will just take the features mentioned above and create a straight-forward user interface. You can use a wx.SearchCtrl for searching for files and an ObjectListView for displaying the results. For this particular utility, a wx.CheckBox or two will work nicely for telling your application to search in sub-directories or if the search term is case-sensitive or not.
Here is a mockup of what the application will eventually look like:
File Search Mockup
Now that you have a goal in mind, let’s go ahead and start coding!
Creating the File Search Utility
Your search utility will need two modules. The first module will be called main and it will hold your user interface and most of the application’s logic. The second module is named search_threads and it will contain the logic needed to search your file system using Python’s threading module. You will use pubsub to update the main module as results are found.
The main script
The main module has the bulk of the code for your application. If you go on and enhance this application, the search portion of the code could end up having the majority of the code since that is where a lot of the refinement of your code should probably go.
Regardless, here is the beginning of main:
# main.py import os import sys import subprocess import time import wx from ObjectListView import ObjectListView, ColumnDefn from pubsub import pub from search_threads import SearchFolderThread, SearchSubdirectoriesThread
This time around, you will be using a few more built-in Python modules, such as os, sys, subprocess and time. The other imports are pretty normal, with the last one being a couple of classes that you will be creating based around Python’s Thread class from the threading module.
For now though, let’s just focus on the main module.
Here’s the first class you need to create:
class SearchResult:
def __init__(self, path, modified_time):
self.path = path
self.modified = time.strftime('%D %H:%M:%S',
time.gmtime(modified_time))
The SearchResult class is used for holding information about the results from your search. It is also used by the ObjectListView widget. Currently, you will use it to hold the full path to the search result as well as the file’s modified time. You could easily enhance this to also include file size, creation time, etc.
Now let’s create the MainPanel which houses most of UI code:
class MainPanel(wx.Panel):
def __init__(self, parent):
super().__init__(parent)
self.search_results = []
self.main_sizer = wx.BoxSizer(wx.VERTICAL)
self.create_ui()
self.SetSizer(self.main_sizer)
pub.subscribe(self.update_search_results, 'update')
The __init__() method gets everything set up. Here you create the main_sizer, an empty list of search_results and a listener or subscription using pubsub. You also call create_ui() to add the user interface widgets to the panel.
Let’s see what’s in create_ui() now:
def create_ui(self):
# Create the widgets for the search path
row_sizer = wx.BoxSizer()
lbl = wx.StaticText(self, label='Location:')
row_sizer.Add(lbl, 0, wx.ALL | wx.CENTER, 5)
self.directory = wx.TextCtrl(self, style=wx.TE_READONLY)
row_sizer.Add(self.directory, 1, wx.ALL | wx.EXPAND, 5)
open_dir_btn = wx.Button(self, label='Choose Folder')
open_dir_btn.Bind(wx.EVT_BUTTON, self.on_choose_folder)
row_sizer.Add(open_dir_btn, 0, wx.ALL, 5)
self.main_sizer.Add(row_sizer, 0, wx.EXPAND)
There are quite a few widgets to add to this user interface. To start off, you add a row of widgets that consists of a label, a text control and a button. This series of widgets allows the user to choose which directory they want to search using the button. The text control will hold their choice.
Now let’s add another row of widgets:
# Create search filter widgets row_sizer = wx.BoxSizer() lbl = wx.StaticText(self, label='Limit search to filetype:') row_sizer.Add(lbl, 0, wx.ALL|wx.CENTER, 5) self.file_type = wx.TextCtrl(self) row_sizer.Add(self.file_type, 0, wx.ALL, 5) self.sub_directories = wx.CheckBox(self, label='Sub-directories') row_sizer.Add(self.sub_directories, 0, wx.ALL | wx.CENTER, 5) self.case_sensitive = wx.CheckBox(self, label='Case-sensitive') row_sizer.Add(self.case_sensitive, 0, wx.ALL | wx.CENTER, 5) self.main_sizer.Add(row_sizer)
This row of widgets contains another label, a text control and two instances of wx.Checkbox. These are the filter widgets which control what you are searching for. You can filter based on any of the following:
- The file type
- Search sub-directories (when checked) or just the chosen directory
- The search term is case-sensitive
The latter two options are represented by using the wx.Checkbox widget.
Let’s add the search control next:
# Add search bar
self.search_ctrl = wx.SearchCtrl(
self, style=wx.TE_PROCESS_ENTER, size=(-1, 25))
self.search_ctrl.Bind(wx.EVT_SEARCHCTRL_SEARCH_BTN, self.on_search)
self.search_ctrl.Bind(wx.EVT_TEXT_ENTER, self.on_search)
self.main_sizer.Add(self.search_ctrl, 0, wx.ALL | wx.EXPAND, 5)
The wx.SearchCtrl is the widget to use for searching. You could quite easily use a wx.TextCtrl instead though. Regardless, in this case you bind to the press of the Enter key and to the mouse click of the magnifying class within the control. If you do either of these actions, you will call search().
Now let’s add the last two widgets and you will be done with the code for create_ui():
# Search results widget
self.search_results_olv = ObjectListView(
self, style=wx.LC_REPORT | wx.SUNKEN_BORDER)
self.search_results_olv.SetEmptyListMsg("No Results Found")
self.main_sizer.Add(self.search_results_olv, 1, wx.ALL | wx.EXPAND, 5)
self.update_ui()
show_result_btn = wx.Button(self, label='Open Containing Folder')
show_result_btn.Bind(wx.EVT_BUTTON, self.on_show_result)
self.main_sizer.Add(show_result_btn, 0, wx.ALL | wx.CENTER, 5)
The results of your search will appear in your ObjectListView widget. You also need to add a button that will attempt to show the result in the containing folder, kind of like how Mozilla Firefox has a right-click menu called “Open Containing Folder” for opening downloaded files.
The next method to create is on_choose_folder():
def on_choose_folder(self, event):
with wx.DirDialog(self, "Choose a directory:",
style=wx.DD_DEFAULT_STYLE,
) as dlg:
if dlg.ShowModal() == wx.ID_OK:
self.directory.SetValue(dlg.GetPath())
You need to allow the user to select a folder that you want to conduct a search in. You could let the user type in the path, but that is error-prone and you might need to add special error checking. Instead, you opt to use a wx.DirDialog, which prevents the user from entering a non-existent path. It is possible for the user to select the folder, then delete the folder before executing the search, but that would be an unlikely scenario.
Now you need a way to open a folder with Python:
def on_show_result(self, event):
"""
Attempt to open the folder that the result was found in
"""
result = self.search_results_olv.GetSelectedObject()
if result:
path = os.path.dirname(result.path)
try:
if sys.platform == 'darwin':
subprocess.check_call(['open', '--', path])
elif 'linux' in sys.platform:
subprocess.check_call(['xdg-open', path])
elif sys.platform == 'win32':
subprocess.check_call(['explorer', path])
except:
if sys.platform == 'win32':
# Ignore error on Windows as there seems to be
# a weird return code on Windows
return
message = f'Unable to open file manager to {path}'
with wx.MessageDialog(None, message=message,
caption='Error',
style= wx.ICON_ERROR) as dlg:
dlg.ShowModal()
The on_show_result() method will check what platform the code is running under and then attempt to launch that platform’s file manager. Windows uses Explorer while Linux uses xdg-open for example.
During testing, it was noticed that on Windows, Explorer returns a non-zero result even when it opens Explorer successfully, so in that case you just ignore the error. But on other platforms, you can show a message to the user that you were unable to open the folder.
The next bit of code you need to write is the on_search() event handler:
def on_search(self, event):
search_term = self.search_ctrl.GetValue()
file_type = self.file_type.GetValue()
file_type = file_type.lower()
if '.' not in file_type:
file_type = f'.{file_type}'
if not self.sub_directories.GetValue():
# Do not search sub-directories
self.search_current_folder_only(search_term, file_type)
else:
self.search(search_term, file_type)
When you click the “Search” button, you want it to do something useful. That is where the code above comes into play. Here you get the search_term and the file_type. To prevent issues, you put the file type in lower case and you will do the same thing during the search.
Next you check to see if the sub_directories check box is checked or not. If sub_directories is unchecked, then you call search_current_folder_only(); otherwise you call search().
Let’s see what goes into search() first:
def search(self, search_term, file_type):
"""
Search for the specified term in the directory and its
sub-directories
"""
folder = self.directory.GetValue()
if folder:
self.search_results = []
SearchSubdirectoriesThread(folder, search_term, file_type,
self.case_sensitive.GetValue())
Here you grab the folder that the user has selected. In the event that the user has not chosen a folder, the search button will not do anything. But if they have chosen something, then you call the SearchSubdirectoriesThread thread with the appropriate parameters. You will see what the code in that class is in a later section.
But first, you need to create the search_current_folder_only() method:
def search_current_folder_only(self, search_term, file_type):
"""
Search for the specified term in the directory only. Do
not search sub-directories
"""
folder = self.directory.GetValue()
if folder:
self.search_results = []
SearchFolderThread(folder, search_term, file_type,
self.case_sensitive.GetValue())
This code is pretty similar to the previous function. Its only difference is that it executesSearchFolderThread instead of SearchSubdirectoriesThread.
The next function to create is update_search_results():
def update_search_results(self, result):
"""
Called by pubsub from thread
"""
if result:
path, modified_time = result
self.search_results.append(SearchResult(path, modified_time))
self.update_ui()
When a search result is found, the thread will post that result back to the main application using a thread-safe method and pubsub. This method is what will get called assuming that the topic matches the subscription that you created in the __init__(). Once called, this method will append the result to search_results and then call update_ui().
Speaking of which, you can code that up now:
def update_ui(self):
self.search_results_olv.SetColumns([
ColumnDefn("File Path", "left", 300, "path"),
ColumnDefn("Modified Time", "left", 150, "modified")
])
self.search_results_olv.SetObjects(self.search_results)
The update_ui() method defines the columns that are shown in your ObjectListView widget. It also calls SetObjects() which will update the contents of the widget and show your search results to the user.
To wrap up the main module, you will need to write the Search class:
class Search(wx.Frame):
def __init__(self):
super().__init__(None, title='Search Utility',
size=(600, 600))
pub.subscribe(self.update_status, 'status')
panel = MainPanel(self)
self.statusbar = self.CreateStatusBar(1)
self.Show()
def update_status(self, search_time):
msg = f'Search finished in {search_time:5.4} seconds'
self.SetStatusText(msg)
if __name__ == '__main__':
app = wx.App(False)
frame = Search()
app.MainLoop()
This class creates the MainPanel which holds most of the widgets that the user will see and interact with. It also sets the initial size of the application along with its title. There is also a status bar that will be used to communicate to the user when a search has finished and how long it took for said search to complete.
Here is what the application will look like:

Now let’s move on and create the module that holds your search threads.
The search_threads Module
The search_threads module contains the two Thread classes that you will use for searching your file system. The thread classes are actually quite similar in their form and function.
Let’s get started:
# search_threads.py import os import time import wx from pubsub import pub from threading import Thread
These are the modules that you will need to make this code work. You will be using the os module to check paths, traverse the file system and get statistics from files. You will use pubsub to communicate with your application when your search returns results.
Here is the first class:
class SearchFolderThread(Thread):
def __init__(self, folder, search_term, file_type, case_sensitive):
super().__init__()
self.folder = folder
self.search_term = search_term
self.file_type = file_type
self.case_sensitive = case_sensitive
self.start()
This thread takes in the folder to search in, the search_term to look for, a file_type filter and whether or not the search term is case_sensitive. You take these in and assign them to instance variables of the same name. The point of this thread is only to search the contents of the folder that is passed-in, not its sub-directories.
You will also need to override the thread’s run() method:
def run(self):
start = time.time()
for entry in os.scandir(self.folder):
if entry.is_file():
if self.case_sensitive:
path = entry.name
else:
path = entry.name.lower()
if self.search_term in path:
_, ext = os.path.splitext(entry.path)
data = (entry.path, entry.stat().st_mtime)
wx.CallAfter(pub.sendMessage, 'update', result=data)
end = time.time()
# Always update at the end even if there were no results
wx.CallAfter(pub.sendMessage, 'update', result=[])
wx.CallAfter(pub.sendMessage, 'status', search_time=end-start)
Here you collect the start time of the thread. Then you use os.scandir() to loop over the contents of the folder. If the path is a file, you will check to see if the search_term is in the path and has the right file_type. Should both of those return True, then you get the requisite data and send it to your application using wx.CallAfter(), which is a thread-safe method.
Finally you grab the end_time and use that to calculate the total run time of the search and then send that back to the application. The application will then update the status bar with the search time.
Now let’s check out the other class:
class SearchSubdirectoriesThread(Thread):
def __init__(self, folder, search_term, file_type, case_sensitive):
super().__init__()
self.folder = folder
self.search_term = search_term
self.file_type = file_type
self.case_sensitive = case_sensitive
self.start()
The SearchSubdirectoriesThread thread is used for searching not only the passed-in folder but also its sub-directories. It accepts the same arguments as the previous class.
Here is what you will need to put in its run() method:
def run(self):
start = time.time()
for root, dirs, files in os.walk(self.folder):
for f in files:
full_path = os.path.join(root, f)
if not self.case_sensitive:
full_path = full_path.lower()
if self.search_term in full_path and os.path.exists(full_path):
_, ext = os.path.splitext(full_path)
data = (full_path, os.stat(full_path).st_mtime)
wx.CallAfter(pub.sendMessage, 'update', result=data)
end = time.time()
# Always update at the end even if there were no results
wx.CallAfter(pub.sendMessage, 'update', result=[])
wx.CallAfter(pub.sendMessage, 'status', search_time=end-start)
For this thread, you need to use os.walk() to search the passed in folder and its sub-directories. Besides that, the conditional statements are virtually the same as the previous class.
Wrapping Up
Creating search utilities is not particularly difficult, but it can be time-consuming. Figuring out the edge cases and how to account for them is usually what takes the longest when creating software. In this article, you learned how to create a utility to search for files on your computer.
Here are a few enhancements that you could add to this program:
- Add the ability to stop the search
- Prevent multiple searches from occurring at the same time
- Add other filters
Related Reading
Want to learn how to create more GUI applications with wxPython? Then check out these resources below:
-
Creating a GUI Application for NASA’s API with wxPython (article)
- Creating GUI Applications with wxPython (book) on Leanpub, Gumroad, or Amazon.
The post Creating a File Search GUI with wxPython appeared first on Mouse Vs Python.
September 02, 2021 12:30 PM UTC
Stack Abuse
Guide to Numpy's arange() Function
Intro
Numpy is the most popular mathematical computing Python library. It offers a great number of mathematical tools including but not limited to multi-dimensional arrays and matrices, mathematical functions, number generators, and a lot more.
One of the fundamental tools in NumPy is the ndarray - an N-dimensional array. Today, we're going to create ndarrays, generated in certain ranges using the NumPy.arange() function.
Parameters and Return
numpy.arange([start, ]stop, [step, ]dtype=None)
Returns evenly spaced values within a given interval where:
- start is a number (integer or real) from which the array starts from. It is optional.
- stop is a number (integer or real) which the array ends at and is not included in it.
- step is a number that sets the spacing between the consecutive values in the array. It is optional and is 1 by default.
- dtype is the type of output for array elements. It is None by default.
The method returns an ndarray of of evenly spaced values. If the array returns floating-point elements the array's length will be ceil((stop - start)/step).
np.arange() by Example
Importing NumPy
To start working with NumPy, we need to import it, as it's an external library:
import NumPy as np
If not installed, you can easily install it via pip:
$ pip install numpy
All-Argument np.arange()
Let's see how arange() works with all the arguments for the function. For instance, say we want a sequence to start at 0, stop at 10, with a step size of 3, while producing integers.
In a Python environement, or REPL, let's generate a sequence in a range:
>>> result_array = np.arange(start=0, stop=10, step=2, dtype=int)
The array is an ndarray containing the generated elements:
>>> result_array
array([0, 2, 4, 6, 8])
It's worth noting that the stop element isn't included, while the start element is included, hence we have a 0 but not a 10 even though the next element in the sequence should be a 10.
Note: As usual, you an provide positional arguments, without naming them or named arguments:
array = np.arange(start=0, stop=10, step=2, dtype=int)
# These two statements are the same
array = np.arange(0, 10, 2, int)
For the sake of brevity, the latter is oftentimes used, and the positions of these arguments must follow the sequence of start, stop, step and dtype.
np.arange() with stop
If only one argument is provided, it will be treated as the stop value. It will output all numbers up to but not including the stop number, with a default step of 1 and start of 0:
>>> result_array = np.arange(5)
>>> result_array
array([0, 1, 2, 3, 4])
np.arange() with start and stop
With two arguments, they default to start and stop, with a default step of 1 - so you can easily create a specific range without thinking about the step size:
>>> result_array = np.arange(5, 10)
>>> result_array
array([5, 6, 7, 8, 9])
Like with previous examples, you can also use floating point numbers here instead of integers. For example, we can start at 5.5:
>>> result_array = np.arange(5.5, 11.75)
The resulting array will be:
>>> result_array
array([ 5.5, 6.5, 7.5, 8.5, 9.5, 10.5, 11.5])
np.arange() with start, stop and step
The default dtype is None and in that case, ints are used so having an integer-based range is easy to create with a start, stop and step. For instance, let's generate a sequence of all the even numbers between 6 (inclusive) and 22 (exclusive):
>>> result_array = np.arange(6, 22, 2)
The result will be all even numbers between 6 up to but not including 22:
>>> result_array
array([ 6, 8, 10, 12, 14, 16, 18, 20])
np.arange() for Reversed Ranges
We can also pass in negative parameters into the np.arange() function to get a reversed array of numbers.
The start will be the larger number we want to start counting from, the stop will be the lower one, and the step will be a negative number:
result_array = np.arange(start=30,stop=14, step=-3)
The result will be an array of descending numbers with a negative step of 3:
>>> result_array
array([30, 27, 24, 21, 18, 15])
Creating Empty NDArrays with np.arange()
We can also create an empty arange as follows:
>>> result_array = np.arange(0)
The result will be an empty array:
>>> result_array
array([], dtype=int32)
This happens because 0 is the stop value we've set, and the start value is also 0 by default. So, the counting stops before starting.
Another case where the result will be an empty array is when the start value is higher than the stop value while the step is positive. For example:
>>> result_array = np.arange(start=30, stop=10, step=1)
The result will also be an empty array.
>>> result_array
array([], dtype=int32)
This can also happen the other way around. We can start with a small number, stop at a larger number, and have the step as a negative number. The output will be an empty array too:
>>> result_array = np.arange(start=10, stop=30, step=-1)
This also results in an empty ndarray:
>>> result_array
array([], dtype=int32)
Supported Data Types for np.arange()
The
dtypeargument, which defaults tointcan be any valid NumPy data type.
Note: This isn't to be confused with standard Python data types, though.
You can use the shorthand version for some of the more common datatypes, or the full name, prefixed with np.:
np.arange(..., dtype=int)
np.arange(..., dtype=np.int32)
np.arange(..., dtype=np.int64)
For some other data types, such as np.csignle, you'll prefix the type with np.:
>>> result_array = np.arange(start=10, stop=30, step=1, dtype=np.csingle)
>>> result_array
array([10.+0.j, 11.+0.j, 12.+0.j, 13.+0.j, 14.+0.j, 15.+0.j, 16.+0.j,
17.+0.j, 18.+0.j, 19.+0.j, 20.+0.j, 21.+0.j, 22.+0.j, 23.+0.j,
24.+0.j, 25.+0.j, 26.+0.j, 27.+0.j, 28.+0.j, 29.+0.j],
dtype=complex64)
A common short-hand data type is a float:
>>> result_array = np.arange(start=10, stop=30, step=1, dtype=float)
>>> result_array
array([10., 11., 12., 13., 14., 15., 16., 17., 18., 19., 20., 21., 22.,
23., 24., 25., 26., 27., 28., 29.])
For a list of all supported NumPy data types, take a look at the official documentation.
np.arange() vs np.linspace()
np.linspace() is similar to np.arange() in returning evenly spaced arrays. However, there are a couple of differences.
With np.linspace(), you specify the number of samples in a certain range instead of specifying the step. In addition, you can include endpoints in the returned array. Another difference is that np.linspace() can generate multiple arrays instead of returning only one array.
This is a simple example of np.linspace() with the endpoint included and 5 samples:
>>> result_array = np.linspace(0, 20, num=5, endpoint=True)
>>> result_array
array([ 0., 5., 10., 15., 20.])
Here, both the number of samples and the step size is 5, but that's coincidental:
>>> result_array = np.linspace(0, 20, num=2, endpoint=True)
>>> result_array
array([ 0., 20.])
Here, we make two points between 0 and 20, so they're naturally 20 steps apart. You can also the endpoint to False and np.linspace()will behave more likenp.arange()` in that it doesn't include the final element:
>>> result_array = np.linspace(0, 20, num=5, endpoint=False)
>>> result_array
array([ 0., 4., 8., 12., 16.])
np.arange() vs built-in range()
The Python's built-in range() function and np.arange() share a lot of similarities but have slight differences. In the following sections, we're going to highlight some of the similarities and differences between them.
Parameters and Returns
The main similarities are that they both have a start, stop, and step. Additionally, they are both start inclusive, and stop exclusive, with a default step of 1.
However:
- np.arange()
- Can handle multiple data types including floats and complex numbers
- returns a
ndarray - The array is fully created in memory
- range()
- Can handle only integers
- Returns a
rangeobject - Generates numbers on demand
Efficiency and Speed
There are some speed and efficiency differences between np.arange() and the built-in range() function. The range function generates the numbers on demand and doesn't create them in-memory, upfront.
This helps speed the process up if you know you'll break somewhere in that range: For example:
for i in range(100000000):
if i == some_number:
break
This will consume less memory since not all numbers are created in advance. This also makes ndarrays slower to initially construct.
However, if you still need the whole range of numbers in-memory, np.arange() is significantly faster than range() when the full range of numbers comes into play, after they've been constructed.
For instance, if we just iterate through them, the time it takes to create the arrays makes np.arange() perform slower due to the higher upfront cost:
$ python -m timeit "for i in range(100000): pass"
200 loops, best of 5: 1.13 msec per loop
$ python -m timeit "import numpy as np" "for i in np.arange(100000): pass"
100 loops, best of 5: 3.83 msec per loop
Conclusion
This guide aims to help you understand how the np.arange() function works and how to generate sequences of numbers.
Here's a quick recap of what we just covered.
np.arange()has 4 parameters:- start is a number (integer or real) from which the array starts from. It is optional.
- stop is a number (integer or real) which the array ends at and is not included in it.
- step is a number that sets the spacing between the consecutive values in the array. It is optional and is 1 by default.
- dtype is the type of output for array elements. It is
Noneby default.
- You can use multiple dtypes with arange including ints, floats, and complex numbers.
- You can generate reversed ranges by having the larger number as the start, the smaller number as the stop, and the step as a negative number.
np.linspace()is similar tonp.arange()in generating a range of numbers but differs in including the ability to include the endpoint and generating a number of samples instead of steps, which are computed based on the number of samples.np.arange()is more efficient than range when you need the whole array created. However, the range is better if you know you'll break somewhere when looping.
September 02, 2021 08:30 AM UTC
Python Bytes
#248 while True: stand up, sit down
<p><strong>Watch the live stream:</strong></p> <a href='/sitelet?url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DeIEGTZnsyCg' style='font-weight: bold;'>Watch on YouTube</a><br> <br> <p><strong>About the show</strong></p> <p>Sponsored by <strong>us:</strong></p> <ul> <li>Check out the <a href="/sitelet?url=https%3A%2F%2Ftraining.talkpython.fm%2Fcourses%2Fall"><strong>courses over at Talk Python</strong></a></li> <li>And <a href="/sitelet?url=https%3A%2F%2Fpythontest.com%2Fpytest-book%2F"><strong>Brian’s book too</strong></a>!</li> </ul> <p>Special guest: <strong>Paul Everitt</strong></p> <p><strong>Brain #1:</strong> <a href="/sitelet?url=https%3A%2F%2Fthreeofwands.com%2Fwhy-i-use-attrs-instead-of-pydantic%2F"><strong>Why I use attrs instead of pydantic</strong></a></p> <ul> <li><strong>Tin Tvrtković,</strong> <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2Ftintvrtkovic">@tintvrtkovic</a></li> <li>attrs vs dataclasses <ul> <li>Since dataclasses are a strict subset of attrs functionality. Recommend using attrs in most cases over dataclasses</li> <li>attrs is faster, has more features, releases more frequently, offers over a wider range of Python versions.</li> </ul></li> <li>attrs vs Pydantic <ul> <li>attrs is a library for generating the boring parts of writing classes; <ul> <li>Pydantic is that but also</li> <li>a complex validation library.</li> <li>a structuring/unstructuring library, ex converting to json and back</li> </ul></li> <li>attrs has opt-in validation that you have more control over</li> <li>cattrs can be used for structuring/unstructuring</li> <li>converters are opt-in for attrs, built into Pydantic, and can be wrong. <ul> <li>example using Pendulum that Pydantic mishandles</li> </ul></li> </ul></li> <li>Summary <ul> <li>attrs + cattrs + validators where necessary, converters where necessary</li> <li>will be faster</li> <li>you’ll have more control</li> <li>Kind of a “small, sharp, specialized tools” vs “swiss army knife” comparison.</li> </ul></li> </ul> <p><strong>Michael #2:</strong> <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2Fwhereismyjetpac%2Fstatus%2F1430694757320347648"><strong>mclfy</strong></a></p> <ul> <li>via __dann__</li> <li>Mcfly is an incredible Ctrl+r replacement</li> <li>McFly replaces your default <code>ctrl-r</code> shell history search with an intelligent search engine that takes into account your working directory and the context of recently executed commands. </li> <li>McFly's suggestions are prioritized in real time with a small neural network.</li> <li>Features <ul> <li>Rebinds <code>ctrl-r</code> to bring up a full-screen reverse history search prioritized with a small neural network.</li> <li>Augments your shell history to track command exit status, timestamp, and execution directory in a SQLite database.</li> <li>Maintains your normal shell history file as well so that you can stop using McFly whenever you want.</li> <li>Includes a simple action to scrub any history item from the McFly database and your shell history files.</li> <li>Designed to be extensible for other shells in the future.</li> <li>Written in Rust, so it's fast and safe.</li> </ul></li> </ul> <p><strong>Paul #3: Textual and</strong> <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2Fwillmcgugan%2Fstatus%2F1426267903733768193"><strong>boilerplate removal</strong></a></p> <ul> <li>In the race to make Textual the most talked-about package in Python Bytes history…</li> <li>I’d like to zoom in on a Twitter discussion he had about removing boilerplate</li> <li>I have traditionally been opposed to the convention-over-configuration approach that most successful Python projects have taken</li> <li>I dislike magic variable and file names, prefer explicit is better than implicit, actual <em>symbols</em></li> <li>Lately, because of…tooling</li> <li>But Will’s approach to “boilerplate removal” is compelling, as it remains mypy friendly</li> <li>Still, I find it flawed…code meant to be read 2 years from now…that stuff that is implied-away, worries me</li> <li>Will is great at working-in-the-open, being a gentle, encouraging public figure</li> </ul> <p><strong>Brian #4:</strong> <a href="/sitelet?url=https%3A%2F%2Fgithub.com%2FErotemic%2Fxdoctest"><strong>xdoctest</strong></a> </p> <ul> <li>“The <code>xdoctest</code> package is a re-write of Python's builtin <code>doctest</code> module. It replaces the old regex-based parser with a new abstract-syntax-tree based parser (using Python's <code>ast</code> module). The goal is to make doctests easier to write, simpler to configure, and encourage the pattern of test driven development.”</li> <li>“The main enhancements <code>xdoctest</code> offers over <code>doctest</code> are: <ol> <li>All lines in the doctest can now be prefixed with <code>>>></code>. Old-style doctests with <code>...</code> are still valid.</li> <li>Additionally, the multi-line strings don't require any prefix (but its ok if they do have either prefix).</li> <li>Tests are executed in blocks, rather than line-by-line, thus comment-based directives (e.g. <code># doctest: +SKIP</code>) are now applied to an entire block, rather than just a single line.</li> <li>Tests without a "want" statement will ignore any stdout / final evaluated value. This makes it easy to use simple assert statements to perform checks in code that might write to stdout.</li> <li>If your test has a "want" statement and ends with both a value and stdout, both are checked, and the test will pass if either matches.</li> <li>Output from multiple sequential print statements can now be checked by a single "got" statement. (new in 0.4.0).”</li> </ol></li> <li>Features I love <ul> <li>“The new got/want tester is very permissive by default; it ignores differences in whitespace”</li> <li>You can make doctest normalize whitespace, but why should you have to?</li> </ul></li> </ul> <p><strong>Michael #5:</strong> <a href="/sitelet?url=https%3A%2F%2Fmedium.com%2F%40davidkongfilm%2Fhow-i-hacked-my-standing-desk-with-a-raspberry-pi-a50ed14c7f6f"><strong>Automate the standing desk with python</strong></a></p> <ul> <li>via Joe Riedley, by David Kong</li> <li>“When I first started using it, I was very excited, but I quickly found myself sitting all day, in spite of the fancy desk.”</li> <li>I took off a few screws and … voila! A row of pins neatly exposed right in front.</li> <li>The pins in my control box, when connected correctly, simulate the pressing of the buttons on the front of the box.</li> <li><a href="/sitelet?url=https%3A%2F%2Fwww.raspberrypi.org%2Fproducts%2Fraspberry-pi-zero%2F"><strong>Raspberry Pi Zero</strong></a>, the simplest, most basic version. It doesn’t have all the bells and whistles, but it does everything I needed for this simple project, and it’s just $5(!).</li> <li>And the code</li> </ul> <pre><code> from gpiozero import LED # The LED library allows easy pin control from time import sleep import randomrelay = LED(17) # I connected the relay to pin 17 and groundwhile True: relay.on() sleep(1) relay.off() sleep(random.randint(45, 60) * 60) </code></pre> <p><strong>Paul #6:</strong> <a href="/sitelet?url=https%3A%2F%2Fcookiecutter-hypermodern-python.readthedocs.io%2Fen%2F2021.4.15%2F"><strong>Hypermodern Python Cookiecutter</strong></a></p> <ul> <li>I’ve been noodling with some code the last two years about bringing frontend DX to Python web dev</li> <li>Learning and talking more than adoption</li> <li>Running a modern Python project is a LOT of housekeeping</li> <li>Hypermodern Python Cookiecutter from Claudio Jolowicz teleported me to a state of the art I was looking for</li> <li>Poetry, Nox, GHA, pre-commit, flake8, PyPI uploads from CI, release drafter, Black, prettier, pytest, mypy, Sphinx and friends, GitHub labeler</li> <li>It’s NOT AT ALL just a cookiecutter</li> <li>The best part…it’s an enormously-detailed user guide, some blog posts with the “why”, it’s actively maintained</li> <li>The PR workflow is really well explained and wired up</li> <li>This could be…a course, a webinar</li> <li>Thanks Claudio</li> </ul> <p><strong>Extras</strong></p> <p>Michael:</p> <ul> <li><a href="/sitelet?url=https%3A%2F%2Fwww.surveymonkey.com%2Fr%2Fsecure-your-supply-chain"><strong>ActiveState's 2021 Software Supply Chain Security Survey</strong></a></li> <li><a href="/sitelet?url=https%3A%2F%2Fpythoninsider.blogspot.com%2F2021%2F08%2Fpython-397-and-3812-are-now-available.html"><strong>Python 3.9.7 and 3.8.12 are now available</strong></a></li> <li>From Shlomi Lanton, on your #2 Brian talked about having a history of all files to find the ones that were updated last, so I created <a href="/sitelet?url=https%3A%2F%2Fgithub.com%2FshlomiLan%2Fgrampa"><strong>granpa</strong></a> </li> <li>Also: <a href="/sitelet?url=https%3A%2F%2Fgithub.com%2Fnp-8%2Fwakepy"><strong>wakepy</strong></a> now works correctly on macOS</li> </ul> <p><strong>Joke:</strong> <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2Fismonkeyuser%2Fstatus%2F1430413027950612481%2F"><strong>Meaning</strong></a></p>
September 02, 2021 08:00 AM UTC
Read the Docs
Read the Docs newsletter - September 2021
Welcome to the latest edition of our monthly newsletter, where we share the most relevant updates around Read the Docs, offer a summary of new features we shipped during the previous month, and share what we’ll be focusing on in the near future.
Company highlights
- We have published the first release candidate of version 1.0.0 of our Sphinx theme, which adds support for recent versions of Sphinx and docutils among other things, and announced our future plans for it. Check out the linked blog post to know more.
- The first part of the new Read the Docs tutorial, which we wrote as part of our CZI grant, is online! We hope that it serves as a better starting point for people seeking to learn how to use our platform.
New features
- Commercial users can now share specific versions of their projects and log out from the flyout menu.
- We added the possibility for users to remove themselves from a project without having to ask an owner to remove them (as long as they are not the only owner of the project).
- We fixed a small usability issue with our search-as-you-type box, that ignored the first typed characters under certain circumstances.
- We have documented our unofficial support for Gitea, and made other minor documentation fixes.
Thanks to our external contributors Mozi, Maksudul Haque, Stefano Costa, and Christian Clauss.
You can always see the latest changes to our platforms in our Read the Docs Changelog.
Upcoming features
- Ana will continue researching tools for conducting automated testing on our Sphinx theme, review pull requests corresponding to the 1.1 milestone, and make some small styling improvements to our documentation.
- Anthony will work with Ana on testing our Sphinx theme, in addition to resuming work on our new user interface and doing some financial updates.
- Eric will keep pushing updates on our Commercial landing page and continue working on our sales processes. He’s also working to continue building EthicalAds as well.
- Juan Luis will continue working on our Read the Docs tutorial, improving our onboarding experience, and put our email marketing and on-site notifications to work.
- Manuel will finish the work on our new Docker images and build process, and keep designing our upcoming GitHub Application.
- Santos will implement a new Slack integration, work with Manuel on the GitHub Application, and build a user interface for audit tracking.
Possible issues
We have been making changes to how we store cookies to make our site more secure. This has caused some minor issues in certain web browsers or to users that were embedding private documentation pages inside an iframe.
Considering using Read the Docs for your next Sphinx or MkDocs project? Check out our documentation to get started!
September 02, 2021 12:00 AM UTC
September 01, 2021
Python for Beginners
Binary Search Tree in Python
You can use different data structures such as a python dictionary, a list, a tuple, or a set in programs. But these data structures are not sufficient for implementing hierarchical structures in the programs. In this article, we will study about binary search tree data structure and will implement them in python for better understanding.
What is a Binary Tree?
A binary tree is a tree data structure in which each node can have a maximum of 2 children. It means that each node in a binary tree can have either one, or two or no children. Each node in a binary tree contains data and references to its children. Both the children are named as left child and the right child according to their position. The structure of a node in a binary tree is shown in the following figure.
Node of a Binary Tree
We can implement a binary tree node in python as follows.
class BinaryTreeNode:
def __init__(self, data):
self.data = data
self.leftChild = None
self.rightChild=None
What is a Binary Search Tree?
A binary search tree is a binary tree data structure with the following properties.
- There are no duplicate elements in a binary search tree.
- The element at the left child of a node is always less than the element at the current node.
- The left subtree of a node has all elements less than the current node.
- The element at the right child of a node is always greater than the element at the current node.
- The right subtree of a node has all elements greater than the current node.
Following is an example of a binary search tree that satisfies all the properties discussed above.
Binary search tree
Now we will implement some of the basic operations on a binary search tree.
How to Insert an Element in a Binary Search Tree?
We will use the properties of binary search trees to insert elements into it. If we want to insert an element at a specific node, three conditions may arise.
- The current node can be an empty node i.e. None. In this case, we will create a new node with the element to be inserted and will assign the new node to the current node.
- The element to be inserted can be greater than the element at the current node. In this case, we will insert the new element in the right subtree of the current node as the right subtree of any node contains all the elements greater than the current node.
- The element to be inserted can be less than the element at the current node. In this case, we will insert the new element in the left subtree of the current node as the left subtree of any node contains all the elements lesser than the current node.
To insert an element, we will start from the root node and will insert the element into the binary search tree according to the above defined rules. The algorithm to insert elements in a binary search tree is implemented as in Python as follows.
class BinaryTreeNode:
def __init__(self, data):
self.data = data
self.leftChild = None
self.rightChild = None
def insert(root, newValue):
# if binary search tree is empty, create a new node and declare it as root
if root is None:
root = BinaryTreeNode(newValue)
return root
# if newValue is less than value of data in root, add it to left subtree and proceed recursively
if newValue < root.data:
root.leftChild = insert(root.leftChild, newValue)
else:
# if newValue is greater than value of data in root, add it to right subtree and proceed recursively
root.rightChild = insert(root.rightChild, newValue)
return root
root = insert(None, 50)
insert(root, 20)
insert(root, 53)
insert(root, 11)
insert(root, 22)
insert(root, 52)
insert(root, 78)
node1 = root
node2 = node1.leftChild
node3 = node1.rightChild
node4 = node2.leftChild
node5 = node2.rightChild
node6 = node3.leftChild
node7 = node3.rightChild
print("Root Node is:")
print(node1.data)
print("left child of the node is:")
print(node1.leftChild.data)
print("right child of the node is:")
print(node1.rightChild.data)
print("Node is:")
print(node2.data)
print("left child of the node is:")
print(node2.leftChild.data)
print("right child of the node is:")
print(node2.rightChild.data)
print("Node is:")
print(node3.data)
print("left child of the node is:")
print(node3.leftChild.data)
print("right child of the node is:")
print(node3.rightChild.data)
print("Node is:")
print(node4.data)
print("left child of the node is:")
print(node4.leftChild)
print("right child of the node is:")
print(node4.rightChild)
print("Node is:")
print(node5.data)
print("left child of the node is:")
print(node5.leftChild)
print("right child of the node is:")
print(node5.rightChild)
print("Node is:")
print(node6.data)
print("left child of the node is:")
print(node6.leftChild)
print("right child of the node is:")
print(node6.rightChild)
print("Node is:")
print(node7.data)
print("left child of the node is:")
print(node7.leftChild)
print("right child of the node is:")
print(node7.rightChild)
Output:
Root Node is:
50
left child of the node is:
20
right child of the node is:
53
Node is:
20
left child of the node is:
11
right child of the node is:
22
Node is:
53
left child of the node is:
52
right child of the node is:
78
Node is:
11
left child of the node is:
None
right child of the node is:
None
Node is:
22
left child of the node is:
None
right child of the node is:
None
Node is:
52
left child of the node is:
None
right child of the node is:
None
Node is:
78
left child of the node is:
None
right child of the node is:
None
How to search an element in a Binary search Tree?
As you know that a binary search tree cannot have duplicate elements, we can search any element in a binary search tree using the following rules that are based on the properties of the binary search trees. We will start from the root and follow these properties
- If the current node is empty, we will say that the element is not present in the binary search tree.
- If the element in the current node is greater than the element to be searched, we will search the element in its left subtree as the left subtree of any node contains all the elements lesser than the current node.
- If the element in the current node is less than the element to be searched, we will search the element in its right subtree as the right subtree of any node contains all the elements greater than the current node.
- If the element at the current node is equal to the element to be searched, we will return True.
The algorithm to search an element in a binary search tree based on the above properties is implemented in the following program.
class BinaryTreeNode:
def __init__(self, data):
self.data = data
self.leftChild = None
self.rightChild = None
def insert(root, newValue):
# if binary search tree is empty, create a new node and declare it as root
if root is None:
root = BinaryTreeNode(newValue)
return root
# if newValue is less than value of data in root, add it to left subtree and proceed recursively
if newValue < root.data:
root.leftChild = insert(root.leftChild, newValue)
else:
# if newValue is greater than value of data in root, add it to right subtree and proceed recursively
root.rightChild = insert(root.rightChild, newValue)
return root
def search(root, value):
# node is empty
if root is None:
return False
# if element is equal to the element to be searched
elif root.data == value:
return True
# element to be searched is less than the current node
elif root.data > value:
return search(root.leftChild, value)
# element to be searched is greater than the current node
else:
return search(root.rightChild, value)
root = insert(None, 50)
insert(root, 20)
insert(root, 53)
insert(root, 11)
insert(root, 22)
insert(root, 52)
insert(root, 78)
print("53 is present in the binary tree:", search(root, 53))
print("100 is present in the binary tree:", search(root, 100))
Output:
53 is present in the binary tree: True
100 is present in the binary tree: False
Conclusion
In this article, we have discussed binary search trees and their properties. We have also implemented the algorithms to insert elements into a binary search tree and to search elements in a binary search tree in Python. To learn more about data structures in Python, you can read this article on Linked list in python.
The post Binary Search Tree in Python appeared first on PythonForBeginners.com.
September 01, 2021 12:48 PM UTC
Mike Driscoll
Unit Conversion with Python and the Pint Package
Do you need to work measurements often? What about converting from one unit of measurement to another? There is a Python package called Pint that makes working with quantities easy to do. Pint allows you do arithmetic operations between a numerical value and a quantity as well. You can see the many different unit types included with Pint on their GitHub project.
Let’s get started by learning how to install Pint!
Installation
You can install Pint using pip like this:
python3 -m pip install pint
If you are a conda user, then you would want to use this command instead:
conda install -c conda-forge pint
Now that you have Pint installed, you are ready to learn how to use it!
Getting Started with Pint
One of the coolest features of Pint is that you can use it to convert from one unit type to another. For example, you might want to convert from some Imperial unit to a Metric unit.
A popular use case would be to convert from miles to kilometers. Open up your Python REPL (or IDLE) and try out the following code:
>>> from pint import UnitRegistry
>>> ureg = UnitRegistry()
>>> distance = 5 * ureg.mile
>>> distance
<Quantity(5, 'mile')>
>>> distance.to("kilometer")
<Quantity(8.04672, 'kilometer')>
Here you create a Quantity object named distance. You set its value to 5 miles. Then to convert it to kilometers, you call the distance’s to() method and pass in the new quantity name that you want. The result is that 5 miles is converted to 8.04672 kilometers.
You can convert the quantity to different unit types within the same system too. For example, you could convert kilometers to centimeters, if you wanted to:
>>> from pint import UnitRegistry
>>> ureg = UnitRegistry()
>>> distance_in_km = 5 * ureg.kilometer
>>> distance_in_km
<Quantity(5, 'kilometer')>
>>> distance_in_cm = distance_in_km.to("centimeter")
>>> distance_in_cm
<Quantity(500000.0, 'centimeter')>
Pint Parses Strings
One of Pint’s cool features is that you can specify quantities using strings. That means you can do stuff like this:
>>> my_quantity = ureg.Quantity >>> my_quantity(2.54, 'centimeter') <Quantity(2.54, 'centimeter')>
Or you can simplify it, even more, to simply:
>>> my_quantity = ureg.Quantity
>>> my_quantity('2.54in')
<Quantity(2.54, 'inch')>
Now that you have played around with converting between different unit types, you are ready to learn about string formatting with Pint.
Using String Formatting with Pint
Pint supports formatting using Python’s .format() and by using f-strings. Here is an example from the Pint tutorial:
>>> ureg = ureg.Quantity
>>> accel = 1.3 * ureg['meter/second**2']
>>> print(f'The str is {accel}')
The str is 1.3 meter / second ** 2
When the f-string is evaluated, the Quantity object is converted into a more human-readable format.
Pint goes farther than that by extending Python’s formatting capabilities. Here is an example of their custom “pretty print”:
>>> ureg = ureg.Quantity
>>> accel = 1.3 * ureg['meter/second**2']
>>> # Pretty print
>>> 'The pretty representation is {:P}'.format(accel)
'The pretty representation is 1.3 meter/second²'
Pint also supports custom printing of LaTeX and HTML for Jupyter Notebooks.
Wrapping Up
Pint is a really nice Python package. While this tutorial doesn’t cover it, Pint allows you to set your locale so that the unit names match your language. If you regularly work with quantities that need to be converted between unit types (like cm to mm or inches to centimeters), this may be just what you need to make your coding life easier.
More Neat Python Packages
Want to learn about other neat 3rd party Python packages? Check out the following articles:
The post Unit Conversion with Python and the Pint Package appeared first on Mouse Vs Python.
September 01, 2021 12:30 PM UTC
PyBites
Facial Recognition with Python
Identifying faces
I was asked by Bob to write a guest article for the PyBites blog, so whilst this isn’t my first blog article, it is my first ever guest blog article of which I’m immensely proud and very pleased to have written for Pybites.
In this article, I will detail how I used the face_recognition and Pillow modules to extract and then identify faces from a bunch of photographs. I want to credit Brad Traversy as this idea was originally taken from a Traversy Media video I had bookmarked a while ago and recently rewatched (link below this article). Using the same functionality I tweaked the scripts to allow input for a directory of photos and eventually, it will be incorporated into my second PDM Django application.
They are very rough and ready scripts for a proof of concept which I demonstrated to Bob during one of our weekly code check-in calls.
I created two scripts, one to extract faces from a bunch of photos and store them as jpeg files in a specified directory. The second script takes a bunch of known faces (I guess these are the control set) and compares them to photographs to identify the faces in random photos. On the whole, it’s very accurate and fast at identifying the known faces.
Extract the faces
The first script was fairly simple to implement. We have a directory of images in which we want to extract all images of faces, we’ll call this directory unknown. The script essentially scans through each photo, identifies the face and stores this face image as a new jpeg in another directory we’ll call ‘extract’. This file is created with the title of the original image, and the face location within the image.
import face_recognition
from PIL import Image
import os
unknown_faces = os.listdir("../frec/unknown/")
for image in unknown_faces:
image_of_people = face_recognition.load_image_file(f"../frec/unknown/{image}")
unknown_face_locations = face_recognition.face_locations(image_of_people)
for face_location in unknown_face_locations:
top, right, bottom, left = face_location
face_image = image_of_people[top:bottom, left:right]
pil_image = Image.fromarray(face_image)
pil_image.save(f"../frec/extract/{image}_{top}.jpg")So from an image like this:
You would get two separate images in the extract directory like these:
Having tested this on my own photos, the module is so good it identified a face in a poster within one of my photographs which you will be able to see if you download the code from the repository.
Facial Recognition Time
The second script written does all the clever stuff insofar as it will take the directory of images with unknown faces, compare them to the known faces images and then ‘draw’ on the original image a square around the face with the name of the individual identified, if indeed it identifies a face, otherwise it will draw a square around the face with ‘Unknown Person’ in place of the name.
import face_recognition
import os
from PIL import Image, ImageDraw
unknown_faces = os.listdir("../frec/unknown/")
Bob_image = face_recognition.load_image_file("../frec/known/Bob.jpeg")
Bob_encoding = face_recognition.face_encodings(Bob_image)[0]
Julian_image = face_recognition.load_image_file("../frec/known/Julian.jpeg")
Julian_encoding = face_recognition.face_encodings(Julian_image)[0]
known_face_encodings = [
Bob_encoding,
Julian_encoding,
]
known_face_names = [
"Bob",
"Julian",
]
for ukface in unknown_faces:
ukimage = face_recognition.load_image_file(f"../frec/unknown/{ukface}")
ukface_locations = face_recognition.face_locations(ukimage)
ukface_encodings = face_recognition.face_encodings(ukimage, ukface_locations)
# Convert to PIL format
pil_image = Image.fromarray(ukimage)
# Set up drawing on image
draw = ImageDraw.Draw(pil_image)
for (top, right, bottom, left), ukface_encoding in zip(
ukface_locations, ukface_encodings
):
matches = face_recognition.compare_faces(known_face_encodings, ukface_encoding)
name = "Unknown Person"
if True in matches:
first_match_index = matches.index(True)
name = known_face_names[first_match_index]
# Draw Box
draw.rectangle(
((left - 10, top - 10), (right + 10, bottom + 10)), outline=(227, 236, 75)
)
# Draw Label
text_width, text_height = draw.textsize(name)
draw.rectangle(
((left - 10, bottom - text_height + 2), (right + 10, bottom + 10)),
fill=(227, 236, 75),
outline=(227, 236, 75),
)
draw.text((left, bottom - text_height + 5), name, fill=(0, 0, 0, 0))
del draw
pil_image.save(f"../frec/identified/{ukface}_scanned.jpg")
#pil_image.show()
I’ve played around with the functionality of the ‘draw.rectangle’ function to try and capture as much of the face as possible inside of the square as originally the face was obscured by the square.
So again from the given image with unknown faces:
Using two different images of Bob and Julian to match against:
You would end up with:
The accuracy, as I said previously, is amazing and you can tweak the values of the face_recognition package to widen or narrow the match. It is also very fast at scanning a directory of images and displaying the results on the screen or saving them to another directory, which would be very easy with a small tweak to the script.
I believe you are also able to take advantage of GPU processing but I was unable to test this in my current environment.
I have expanded the scripts in the repo to contain more examples of known / unkown faces so feel free to download and have a play.
For those of you that are interested in the source material which I modified, please take a look at Brad’s video on YouTube, without that I wouldn’t have found out about the face_recognition module. Hopefully you can adapt this code to suit your own purpose as I did.
I also had to tweak my scripts to increase the chance of a match otherwise Todd Dewey from Ice Road Truckers was being identified as me (I’m not sure who should be more flattered by that!). That’s what the model and num_jitters options are for in the scripts found in the repo.
The docs for the face_recognition module can be found here Face Recognition Docs
The docs for Pillow can be found here Pillow Docs
My repository for these scripts can be found here GitHub
And Brad Traversy’s repo to accompany his video are here GitHub
September 01, 2021 07:57 AM UTC
Tryton News
Newsletter for September 2021
We hope that everybody had a nice Summer and enjoyed their holidays. The Tryton team continued working on the ERP and we are back with a resume of the latest improvements.
Changes for the User
We added a frame around the image widget. This makes the widget cleaner when empty.
More party identifiers and tax identifiers have been added for Austria, Ukraine and Vietnam.
The rule keywords from statement lines are now stored in such way that they can be used for future matching. This adds a form of learning behavior to the statement rules engine.
We added a new wizard to split accounting lines. This is useful to reschedule payable or receivable lines by applying a new maturity date to each new line. The wizard can also be used on dunning and invoice lines to reschedule them.
It is now possible to set accounts for taxes of type “None”. This is useful for taxes that are entered manually on the invoice because the account will be filled in automatically.
New Modules
The Stock Package Shipping Sendcloud Module allows package labels to be generated for shipments made by any of Sendcloud’s supported carriers.
The Account Budget Module provides the ability to set budgets for accounts over a defined period of time. These budgets can then be used to track the total amount from relevant transactions against the budgeted amount.
The Analytic Budget Module provides the ability to set budgets for analytic accounts over a defined period of time. These budgets can then be used to track the total amount from relevant transactions against the budgeted amount.
The Product Image Module adds images to each product and variant.
Changes for the System Administrator
We improved the error management in the script used to import postal codes.
Changes for the Developer
We moved and renamed the cost_warehouse from the product_cost_warehouse module to the warehouse in the stock module. By doing this it can now be used by any module that depends on the stock module.
The complete locale definition for the user’s language is now sent to the clients.
Proteus now also fills in the wizard actions attribute when the result is an empty list.
The number widgets’ width attribute is now also used as its default display width.
The currency module defines a new Monetary field. This is derived from the Numeric field by adding a currency attribute which contains the name of the field which stores the currency. The desktop clients render these fields using the monetary format and with the currency symbol by default.
It is now also possible to use a string as the digits value on number fields (instead of the usual pair of integers). The string must contain the name of a Many2One field which points to a Model that inherits from DigitsMixin and that provides a get_digits method.
This allowed the removal of all the Function fields that provided the currency and unit digits.
Another benefit is that clients cache the value for each DigitsMixin record for 1 day by default, so this change also reduces the load on the server.
We reduced the number of times we save the cost values when doing multiple moves.
We no longer try to read records that were deleted after being instantiated from a browse list.
The digits argument to the format_number method of Report is now optional. If it is not specified, or is set to None, it will display all the significant digits.
1 post - 1 participant
September 01, 2021 07:00 AM UTC
Django Weblog
Django bugfix release: 3.2.7
Today we've issued the 3.2.7 bugfix release.
The release package and checksums are available from our downloads page, as well as from the Python Package Index. The PGP key ID used for this release is Mariusz Felisiak: 2EF56372BA48CD1B.
September 01, 2021 05:54 AM UTC
August 31, 2021
TestDriven.io
Django REST Framework and Elasticsearch
This tutorial looks at how to integrate Django REST Framework with Elasticsearch.
Sandipan Dey
Probabilistic Deep Learning with Tensorflow
In this blog, we shall discuss on how to implement probabilistic deep learning models using Tensorflow. The problems to be discussed in this blog appeared in the exercises / projects in the coursera course “Probabilistic Deep Learning“, by Imperial College, London, as a part of TensorFlow 2 for Deep Learning Specialization. The problem statements / … Continue reading Probabilistic Deep Learning with Tensorflow
PyCoder’s Weekly
Issue #488 (Aug. 31, 2021)
#488 – AUGUST 31, 2021
View in Browser »
Python Ranks #1 in IEEE “Top Programming Languages”
“Python dominates as the de facto platform for new technologies” and “Learn Python. That’s the biggest takeaway we can give you from its continued dominance of IEEE Spectrum’s annual interactive rankings of the top programming languages. You don’t have to become a dyed-in-the-wool Pythonista, but learning the language well enough to use one of the vast number of libraries written for it is probably worth your time.”
IEEE.ORG
skybison: Instagram’s Experimental Performance Oriented Greenfield Implementation of Python
“Skybison is experimental performance-oriented greenfield implementation of Python 3.8. It contains a number of performance optimizations, including: small objects; a moving GC; hidden classes; bytecode inline caching; type-specialized bytecode; an experimental template JIT.”
GITHUB.COM/FACEBOOKEXPERIMENTAL
Start Your Free Scout APM Trial, No CC Needed, and Receive a $5 Donation to the OSS of Your Choice
Scout APM is leading-edge application performance and error monitoring designed to help devs find and fix observability issues before the customer ever sees them. You can connect your error reporting and APM data on one platform, with Scout’s new error monitoring feature add-on →
SCOUT APM sponsor
How to Use Optional Arguments When Defining Python Functions
In this tutorial, you’ll learn about optional arguments in Python and how to define functions with default values. You’ll also learn how to create functions that accept any number of arguments using args and kwargs.
REAL PYTHON
Python Project-Local Virtualenv Management
On UNIX-like operating systems you can have the Python equivalent of node_modules today, for every Python version, without changing your workflows.
HYNEK SCHLAWACK
Humble Software Bundle: Python Superpowers 2021
Pick up the awesome programming potential of Python with software like Mastering PyCharm (2021 Edition) & Object-Oriented Programming (OOP) in Python. Pay what you want & support charity!
HUMBLEBUNDLE.COM
Join the PyCon US 2022 Team!
The PyCon US organizers are looking for motivated volunteers who want to contribute their time and knowledge to make this year’s conference a great success.
PYCON US
Python 3.9.7 and 3.8.12 Are Now Available
More info in the full changelog.
CPYTHON DEV BLOG
Discussions
math.sqrt vs numpy.sqrt vs x ** 0.5 Performance Discussion
Andrej Karpathy (Director of AI at Tesla) shares an interesting performance observation on this Twitter thread that turns into a tale about accurate benchmarking. Calculating math.sqrt(1337.0) appears to be 10x faster than numpy.sqrt(1337.0). Python’s built-in square root (x ** 0.5) appears to be even faster. However, most of the performance differences seem to come from the benchmark setup, as Ishan Bhatt explains in this writeup.
TWITTER.COM/KARPATHY
Python Jobs
Data Engineer - Python & PostgreSQL (Newport Beach, CA, USA)
Sr. Backend Developer (Amsterdam, Netherlands)
Backend Software Engineer (Anywhere)
Articles & Tutorials
A Python Data Scientist’s Guide to the Apple Silicon Transition
A break down of what Apple Silicon means for Python users today, especially those doing scientific computing and data science: what works, what doesn’t, and where this might be going.
STANLEY SEIBERT
Write an SQL Query Builder in 150 Lines of Python
“This is the fourth article in a series about writing my own SQL query builder. Today, we’ll rewrite it from scratch, explore API design, learn when to be lazy, and look at worse and better ways of doing things – all in 150 lines of Python!”
ANDGRAVITY.COM
Rev APIs Solve All of Your Speech-to-Text Needs
Rev.ai is the most sophisticated automatic speech recognition in the world. Our speech-to-text APIs are more accurate, easier to use, and have less bias than competitors like Google, Amazon, and Microsoft. Try Rev.ai free for five hours right now →
REV.AI sponsor
Splitting Datasets With scikit-learn and train_test_split()
Learn why it’s important to split your dataset in supervised machine learning and how to do that with train_test_split() from the widely used scikit-learn package.
REAL PYTHON video
Building With CircuitPython & Constraints of Python for Microcontrollers
Can you make a version of Python that fits within the memory constraints of a microcontroller and have it still feel like Python? That is the intention behind CircuitPython. This week on the show, Scott Shawcroft, who is the project lead for CircuitPython.
REAL PYTHON podcast
Parsing in Python: Tools and Libraries You Can Use
“We present and compare all possible alternatives you can use to parse languages in Python. From libraries to parser generators, we present all options.”
GABRIELE TOMASSETTI
Low-Level Cache API in Django
Caching in Django can be implemented on different levels (or parts of the site). This article looks at how to use the low-level cache API in Django.
J-O ERIKSSON
SonarLint Free and Open Source IDE Extension for Python Devs
Working in VS Code, PyCharm, Visual Studio, or Eclipse? SonarLint helps you find & fix Code Quality and Code Security issues in your Python codebase!
SONARSOURCE sponsor
Python Behind the Scenes: How Async/Await Works in Python
“The async/await pattern can be explained in a simple manner if you start from the ground up. And that’s what we’re going to do today.”
VICTOR SKVORTSOV
Using libsqlite3 Directly From Python With ctypes
How to use ctypes to run SQLite queries without using the built-in sqlite3 Python package, and without compiling anything.
GITHUB.COM/MICHALC
Projects & Code
Events
Real Python Office Hours (Virtual)
September 1, 2021
REALPYTHON.COM
PyConline AU 2021
September 10 to September 13, 2021
PYCON.ORG.AU
Happy Pythoning!
This was PyCoder’s Weekly Issue #488.
View in Browser »
[ Subscribe to 🐍 PyCoder’s Weekly 💌 – Get the best Python news, articles, and tutorials delivered to your inbox once a week >> Click here to learn more ]
Will McGugan
Pretty printing JSON with Rich
If you work with JSON regularly (90% of Python developers I suspect) you might appreciate the print_json function just landed in Rich v10.9.0
If you call this function with a string, Rich will decode the string, reformat it, and print it to the console with nice syntax highlighting. Here's an example:
from rich import print_json
print_json('{"foo": [false, true, null]}')
Here's the output:
Note that the atomic values false, true, and null have their own color. I find this helpful when scanning a JSON blob.
If you call print_json with a data keyword argument it will encode that data and pretty print it in the same way.
data = {
"foo": [
3.1427,
(
"Paul Atreides",
"Vladimir Harkonnen",
"Thufir Hawat",
),
],
"atomic": (False, True, None),
}
from rich import print_json
print_json(data=data)
Here's the output:
Note that Rich will remove color if you pipe the output of your script to another program, so you can safely add syntax highlighting to your CLI tools.
You can also pretty print JSON files from the command line with the following:
python -m rich.json data.json
Here's an example of the output:
This is admittedly a small addition to Rich but I'm already finding it helpful.
Follow @willmcgugan on Twitter for Rich and Textual updates.
Quansight Labs Blog
CZI EOSS4 Grants at Quansight Labs
Here, at Quansight Labs, our goal is to work on sustaining the future of Open Source. We make sure we can live up to that goal by spending a significant amount of time working on impactful and critical infrastructure and projects within the Scientific Ecosystem.
As such, our goals align with those of the Chan Zuckerberg Initiative and, in particular, the Essential Open Source Software for Science (EOSS) program that supports tools essential to biomedical research via funds for software maintenance, growth, development, and community engagement.
CZI’s Essential Open Source Software for Science program supports software maintenance, growth, development, and community engagement for open source tools critical to science. And the Chan Zuckerberg Initiative was founded in 2015 to help solve some of society’s toughest challenges — from eradicating disease and improving education, to addressing the needs of our local communities. Their mission is to build a more inclusive, just, and healthy future for everyone.
Today, we are thrilled to announce that the team at Quansight Labs has been awarded five EOSS Cycle 4 grants to work on several projects within the PyData ecosystem. This post will introduce the successful grantees and their objectives for these two-year long grants.
Read more… (5 min remaining to read)
Python for Beginners
How to Extract a Date from a .txt File in Python
In this tutorial, we’ll examine the different ways you can extract a date from a .txt file using Python programming. Python is a versatile language—as you’ll discover—and there are many solutions for this problem.
First, we’ll look at using regular expression patterns to search text files for dates that fit a predefined format. We’ll learn about using the re library and creating our own regular expression searches.
We’ll also examine datetime objects and use them to convert strings into data models. Lastly, we’ll see how the datefinder module simplifies the process of searching a text file for dates that haven’t been formatted, like we might find in natural language content.
Extract a Date from a .txt File using Regular Expression
Dates are written in many different formats. Sometimes people write month/day/year. Other dates might include times of the day, or the day of the week (Wednesday July 8, 2021 8:00PM).
How dates are formatted is a factor to consider before we go about extracting them from text files.
For instance, if a date follows the month/date/year format, we can find it using a regular expression pattern. With regular expression, or regex for short, we can search a text by matching a string to a predefined pattern.
The beauty of regular expression is that we can use special characters to create powerful search patterns. For instance, we can craft a pattern that will find all the formatted dates in the following body of text.
minutes.txt
10/14/2021 – Meeting with the client.
07/01/2021 – Discussed marketing strategies.
12/23/2021 – Interviewed a new team lead.
01/28/2018 – Changed domain providers.
06/11/2017 – Discussed moving to a new office.
Example: Finding formatted dates with regex
import re
# open the text file and read the data
file = open("minutes.txt",'r')
text = file.read()
# match a regex pattern for formatted dates
matches = re.findall(r'(\d+/\d+/\d+)',text)
print(matches)
Output
[’10/14/2021′, ’07/01/2021′, ’12/23/2021′, ’01/28/2018′, ’06/11/2017′]
The regex pattern here uses special characters to define the strings we want to extract from the text file. The characters d and + tell regex we’re looking for multiple digits within the text.
We can also use regex to find dates that are formatted in different ways. By altering our regex pattern, we can find dates that use either a forward slash (\) or a dash (–) as the separator.
This works because regex allows for optional characters in the search pattern. We can specify that either character—a forward slash or dash—is an acceptable match.
apple2.txt
The first Apple II was sold on 07-10-1977. The last of the Apple II
models were discontinued on 10/15/1994.
Example: Matching dates with a regex pattern
import re
# open a text file
f = open("apple2.txt", 'r')
# extract the file's content
content = f.read()
# a regular expression pattern to match dates
pattern = "\d{2}[/-]\d{2}[/-]\d{4}"
# find all the strings that match the pattern
dates = re.findall(pattern, content)
for date in dates:
print(date)
f.close()
Output
07-10-1977
10/15/1994
Examining the full extent of regex’s potential is beyond the scope of this tutorial. Try experimenting with some of the following special characters to learn more about using regular expression patterns to extract a date—or other information—from a .txt file.
Special Characters in Regex
- \s – A space character
- \S – Any character except for a space character
- \d – Any digit from 0 to 9
- \D – And any character except for a digit
- \w – Any word of characters or digits [a-zA-Z0-9]
- \W – Any non-word characters
Extract a Datetime Object from a .txt File
In Python we can use the datetime library for manipulating dates and working with time. The datetime library comes pre-packed with Python, so there’s no need to install it.
By using datetime objects, we have more control over string data read from text files. For example, we can use a datetime object to get a copy of the current date and time of our computer.
import datetime
now = datetime.datetime.now()
print(now)
Output
2021-07-04 20:15:49.185380
In the following example, we’ll extract a date from a company .txt file that mentions a scheduled meeting. Our employer needs us to scan a group of such documents for dates. Later, we plan to add the information we gather to a SQLite database.
We’ll begin by defining a regex pattern that will match our date format. Once a match is found, we’ll use it to create a datetime object from the string data.
schedule.txt
schedule.txt
The project begins next month. Denise has scheduled a meeting in the conference room at the Embassy Suits on 10-7-2021.
Example: Creating datetime objects from file data
import re
from datetime import datetime
# open the data file
file = open("schedule.txt", 'r')
text = file.read()
match = re.search(r'\d+-\d+-\d{4}', text)
# create a new datetime object from the regex match
date = datetime.strptime(match.group(), '%d-%m-%Y').date()
print(f"The date of the meeting is on {date}.")
file.close()
Output
The date of the meeting is on 2021-07-10.
Extracting Dates from a Text File with the Datefinder Module
The Python datefinder module can locate dates in a body of text. Using the find_dates() method, it’s possible to search text data for many different types of dates. Datefinder will return any dates it finds in the form of a datetime object.
Unlike the other packages we’ve discussed in this guide, Python does not come with datefinder. The easiest way to install the datefinder module is to use pip from the command prompt.
pip install datefinder
With datefinder installed, we’re ready to open files and extract data. For this example, we’ll use a text document that introduces a fictitious company project. Using datefinder, we’ll extract each date from the .txt file, and print their datimeobject counterparts.
Feel free to save the file locally and follow along.
project_timeline.txt
PROJECT PEPPER
All team members must read the project summary by
January 4th, 2021.
The first meeting of PROJECT PEPPER begins on 01/15/2021
at 9:00am. Please find the time to read the following links by then.
created on 08-12-2021 at 05:00 PM
This project file has dates in many formats. Dates are written using dashes and forward slashes. What’s worse, the month January is written out. How can we find all these dates with Python?
Example: Using datefinder to extract dates from file data
import datefinder
# open the project schedule
file = open("project_timeline.txt",'r')
content = file.read()
# datefinder will find the dates for us
matches = list(datefinder.find_dates(content))
if len(matches) > 0:
for date in matches:
print(date)
else:
print("Found no dates.")
file.close()
Output
2021-01-04 00:00:00
2021-01-15 09:00:00
2021-08-12 17:00:00
As you can see from the output, datefinder is able to find a variety of date formats in the text. Not only is the package capable of recognizing the names of months, but it also recognizes the time of day if it’s included in the text.
In another example, we’ll use the datefinder package to extract a date from a .txt file that includes the dates for a popular singer’s upcoming tour.
tour_dates.txt
Saturday July 25, 2021 at 07:00 PM Inglewood, CA
Sunday July 26, 2021 at 7 PM Inglewood, CA
09/30/2021 7:30PM Foxbourough, MA
Example: Extract a tour date and times from a .txt file with datefinder
import datefinder
# open the project schedule
file = open("tour_dates.txt",'r')
content = file.read()
# datefinder will find the dates for us
matches = list(datefinder.find_dates(content))
if len(matches) > 0:
print("TOUR DATES AND TIMES")
print("--------------------")
for date in matches:
# use f string to format the text
print(f"{date.date()} {date.time()}")
else:
print("Found no dates.")
file.close()
Output
TOUR DATES AND TIMES
——————–
2021-07-25 19:00:00
2021-07-26 19:00:00
2021-09-30 19:30:00
As you can see from the examples, datefinder can find many different types of dates and times. This is useful if the dates you’re looking for don’t have a certain format, as will often be the case in natural language data.
Summary
In this post, we’ve covered several methods of how to extract a date or time from a .txt file. We’ve seen the power of regular expression to find matches in string data, and we’ve seen how to convert that data into a Python datetime object.
Finally, if the dates in your text files don’t have a specified format—as will be the case in most files with natural language content—try the datefinder module. With this Python package, it’s possible to extract dates and times from a text file that aren’t conveniently formatted ahead of time.
Related Posts
If you enjoyed this tutorial and are eager to learn more about Python—and we sincerely hope you are—follow these links for more great guides from Python for Beginners.
- How to use Python concatenation to join strings
- Using Python try catch to mitigate errors and prevent crashes
The post How to Extract a Date from a .txt File in Python appeared first on PythonForBeginners.com.
Real Python
Splitting Datasets With scikit-learn and train_test_split()
One of the key aspects of supervised machine learning is model evaluation and validation. When you evaluate the predictive performance of your model, it’s essential that the process be unbiased. Using train_test_split() from the data science library scikit-learn, you can split your dataset into subsets that minimize the potential for bias in your evaluation and validation process.
In this course, you’ll learn:
- Why you need to split your dataset in supervised machine learning
- Which subsets of the dataset you need for an unbiased evaluation of your model
- How to use
train_test_split()to split your data - How to combine
train_test_split()with prediction methods
In addition, you’ll get information on related tools from sklearn.model_selection.
[ Improve Your Python With 🐍 Python Tricks 💌 – Get a short & sweet Python Trick delivered to your inbox every couple of days. >> Click here to learn more and see examples ]
Stack Abuse
Random Projection: Theory and Implementation in Python with Scikit-Learn
Introduction
This guide is an in-depth introduction to an unsupervised dimensionality reduction technique called Random Projections. A Random Projection can be used to reduce the complexity and size of data, making the data easier to process and visualize. It is also a preprocessing technique for input preparation to a classifier or a regressor.
Random Projection is typically applied to highly-dimensional data, where other techniques such as Principal Component Analysis (PCA) can't do the data justice.
In this guide, we'll delve into the details of Johnson-Lindenstrauss lemma, which lays the mathematical foundation of Random Projections. We'll also show how to perform Random Projection using Python's Scikit-Learn library, and use it to transform input data to a lower-dimensional space.
Theory is theory, and practice is practice. As a practical illustration, we'll load the Reuters Corpus Volume I Dataset, and apply Gaussian Random Projection and Sparse Random Projection to it.
What is a Random Projection of a Dataset?
Put simply:
Random Projection is a method of dimensionality reduction and data visualization that simplifies the complexity of high-dimensional datasets.
The method generates a new dataset by taking the projection of each data point along a randomly chosen set of directions. The projection of a single data point onto a vector is mathematically equivalent to taking the dot product of the point with the vector.

Given a data matrix \(X\) of dimensions \(mxn\) and a \(dxn\) matrix \(R\) whose columns are the vectors representing random directions, the Random Projection of \(X\) is given by \(X_p\).
X p = X REach vector representing a random direction, has dimensionality \(n\), which is the same as all data points of \(X\). If we take \(d\) random directions, then we end up with a \(d\) dimensional transformed dataset. For the purpose of this tutorial, we'll fix a few notations:
m: Total example points/samples of input data.n: Total features/attributes of the input data. It is also the dimensionality of the original data.d: Dimensionality of the transformed data.
The idea of Random Projections is very similar to Principal Component Analysis (PCA), fundementally. However, in PCA, the projection matrix is computed via eigenvectors, which can be computationally expensive for large matrices.
When performing Random Projection, the vectors are chosen randomly making it very efficient. The name "projection" may be a little misleading as the vectors are chosen randomly, the transformed points are mathematically not true projections but close to being true projections.
The data with reduced dimensions is easier to work with. Not only can it be visualized but it can also be used in the pre-processing stage to reduce the size of the original data.
A Simple Example
Just to understand how the transformation works, let's take the following simple example.
Suppose our input matrix \(X\) is given by:
X = [ 1 3 2 0 0 1 2 1 1 3 0 0 ]And the projection matrix is given by:
R = 1 2 [ 1 − 1 1 1 1 − 1 1 1 ]The projection of X onto R is:
X p = X R = 1 2 [ 6 0 4 0 4 2 ]We started with three points in a four-dimensional space, and with clever matrix operations ended up with three transformed points in a two-dimensional space.
Note, some important attributes of the projection matrix \(R\). Each column is a unit matrix, i.e., the norm of each column is one. Also, the dot product of all columns taken pairwise (in this case only column 1 and column 2) is zero, indicating that both column vectors are orthogonal to each other.
This makes the matrix, an Orthonormal Matrix. However, in case of the Random Projection technique, the projection matrix does not have to be a true orthonormal matrix when very high-dimensional data is involved.
The success of Random Projection is based on an awesome mathematical finding known as Johnson-Lindenstrauss lemma, which is explained in detail in the following section!
The Johnson-Lindenstrauss lemma
The Johnson-Lindenstrauss lemma is the mathematical basis for Random Projection:
The Johnson-Lindenstrauss lemma states that if the data points lie in a very high-dimensional space, then projecting such points on simple random directions preserves their pairwise distances.
Preserving pairwise distances implies that the pairwise distances between points in the original space are the same or almost the same as the pairwise distance in the projected lower-dimensional space.
Thus, the structure of data and clusters within data are maintained in a lower-dimensional space, while the complexity and size of data are reduced substantially.
In this guide, we refer to the difference in the actual and projected pairwise distances as the "distortion" in data, which is introduced due to its projection in a new space.
Johnson-Lindenstrauss lemma also provides a "safe" measure of the number of dimensions to project the data points onto so that the error/distortion lies within a certain range, so finding the target number of dimensions is made easy.
Mathematically, given a pair of points \((x_1,x_2)\) and their corresponding projections \((x_1',x_2')\) defines an eps-embedding:
$$
(1 - \epsilon) |x_1 - x_2|^2 < |x_1' - x_2'|^2 < (1 + \epsilon) |x_1 - x_2|^2
$$
The Johnson-Lindenstrauss lemma specifies the minimum dimensions of the lower-dimensional space so that the above eps-embedding is maintained.
Determining the Random Directions of the Projection Matrix
Two well-known methods for determining the projection matrix are:
-
Gaussian Random Projection: The projection matrix is constructed by choosing elements randomly from a Gaussian distribution with mean zero.
-
Sparse Random Projection: This is a comparatively simpler method, where each vector component is a value from the set {-k,0,+k}, where k is a constant. One simple scheme for generating the elements of this matrix, also called the
Achlioptasmethod is to set \(k=\sqrt 3\):
The method above is equivalent to choosing the numbers from {+k,0,-k} based on the outcome of the roll of a dice. If the dice score is 1, then choose +k. If the dice score is in the range [2,5], choose 0, and choose -k for a dice score of 6.
A more general method uses a density parameter to choose the Random Projection matrix. Setting \(s=\frac{1}{\text{density}}\), the elements of the Random Projection matrix are chosen as:
The general recommendation is to set the density parameter to \(\frac{1}{\sqrt n}\).
As mentioned earlier, for both the Gaussian and sparse methods, the projection matrix is not a true orthonormal matrix. However, it has been shown that in high dimensional spaces, the randomly chosen matrix using either of the above two methods is close to an orthonormal matrix.
Random Projection Using Scikit-Learn
The Scikit-Learn library provides us with the random_projection module, that has three important classes/modules:
johnson_lindenstrauss_min_dim(): For determining the minimum number of dimensions of transformed data when given a sample sizem.GaussianRandomProjection: Performs Gaussian Random Projections.SparseRandomProjection: Performs Sparse Random Projections.
We'll demonstrate all the above three in the sections below, but first let's import the classes and functions we'll be using:
from sklearn.random_projection import SparseRandomProjection, johnson_lindenstrauss_min_dim
from sklearn.random_projection import GaussianRandomProjection
import numpy as np
from matplotlib import pyplot as plt
import sklearn.datasets as dt
from sklearn.metrics.pairwise import euclidean_distances
Determining the Minimum Number of Dimensions Via Johnson Lindenstrauss lemma
The johnson_lindenstrauss_min_dim() function determines the minimum number of dimensions d, which the input data can be mapped to when given the number of examples m, and the eps or \(\epsilon\) parameter.
The code below experiments with a different number of samples to determine the minimum size of the lower-dimensional space, which maintains a certain "safe" distortion of data.
Additionally, it plots log(d) against different values of eps for different sample sizes m.
An important thing to note is that the Johnson Lindenstrauss lemma determines the size of the lower-dimensional space \(d\) only based on the number of example points \(m\) in the input data. The number of attributes or features \(n\) of the original data is irrelevant:
eps = np.arange(0.001, 0.999, 0.01)
colors = ['b', 'g', 'm', 'c']
m = [1e1, 1e3, 1e7, 1e10]
for i in range(4):
min_dim = johnson_lindenstrauss_min_dim(n_samples=m[i], eps=eps)
label = 'Total samples = ' + str(m[i])
plt.plot(eps, np.log10(min_dim), c=colors[i], label=label)
plt.xlabel('eps')
plt.ylabel('log$_{10}$(d)')
plt.axhline(y=3.5, color='k', linestyle=':')
plt.legend()
plt.show()

From the plot above, we can see that for small values of eps, d is quite large but decreases as eps approaches one. The dimensionality is below 3500 (the dotted black line) for mid to large values of eps.
This shows that applying Random Projections only makes sense to high-dimensional data, of the order of thousands of features. In such cases, a high reduction in dimensionality can be achieved.
Random Projections are, therefore, very successful for text or image data, which involve a large number of input features, where Principal Component Analysis would
Data Transformation
Python includes the implementation of both Gaussian Random Projections and Sparse Random Projections in its sklearn library via the two classes GaussianRandomProjection and SparseRandomProjection respectively. Some important attributes for these classes are (the list is not exhaustive):
n_components: Number of dimensions of the transformed data. If it is set toauto, then the optimal dimensions are determined before projectioneps: The parameter of Johnson-Lindenstrauss lemma, which controls the number of dimensions so that the distortion in projected data is kept within a certain bound.density: Only applicable forSparseRandomProjection. The default value isauto, which sets \(s=\frac{1}{\sqrt n}\) for the selection of the projection matrix.
Like other dimensionality reduction classes of sklearn, both these classes include the standard fit() and fit_transform() methods. A notable set of attributes, which come in handy are:
n_components: The number of dimensions of the new space on which the data is projected.components_: The transformation or projection matrix.density_: Only applicable toSparseRandomProjection. It is the value ofdensitybased on which the elements of the projection matrix are computed.
Random Projection with GaussianRandomProjection
Let's start off with the GaussianRandomProjection class. The values of the projection matrix are plotted as a histogram and we can see that they follow a Gaussian distribution with mean zero. The size of the data matrix is reduced from 5000 to 3947:
X_rand = np.random.RandomState(0).rand(100, 5000)
proj_gauss = GaussianRandomProjection(random_state=0)
X_transformed = proj_gauss.fit_transform(X_rand)
# Print the size of the transformed data
print('Shape of transformed data: ' + str(X_transformed.shape))
# Generate a histogram of the elements of the transformation matrix
plt.hist(proj_gauss.components_.flatten())
plt.title('Histogram of the flattened transformation matrix')
plt.show()
This code results in:
Shape of transformed data: (100, 3947)

Random Projection with SparseRandomProjection
The code below demonstrates how data transformation can be made using a Sparse Random Projection. The entire transformation matrix is composed of three distinct values, whose frequency plot is also shown below.
Note that the transformation matrix is a SciPy sparse csr_matrix. The following code accesses the non-zero values of the csr_matrix and stores them in p. Next, it uses p to get the counts of the elements of the sparse projection matrix:
proj_sparse = SparseRandomProjection(random_state=0)
X_transformed = proj_sparse.fit_transform(X_rand)
# Print the size of the transformed data
print('Shape of transformed data: ' + str(X_transformed.shape))
# Get data of the transformation matrix and store in p.
# p consists of only 2 non-zero distinct values, i.e., pos and neg
# pos and neg are determined below
p = proj_sparse.components_.data
total_elements = proj_sparse.components_.shape[0] *\
proj_sparse.components_.shape[1]
pos = p[p>0][0]
neg = p[p<0][0]
print('Shape of transformation matrix: '+ str(proj_sparse.components_.shape))
counts = (sum(p==neg), total_elements - len(p), sum(p==pos))
# Histogram of the elements of the transformation matrix
plt.bar([neg, 0, pos], counts, width=0.1)
plt.xticks([neg, 0, pos])
plt.suptitle('Histogram of flattened transformation matrix, ' +
'density = ' +
'{:.2f}'.format(proj_sparse.density_))
plt.show()
This results in:
Shape of transformed data: (100, 3947)
Shape of transformation matrix: (3947, 5000)

The histogram is in agreement with the method of generating a sparse Random Projection matrix as discussed in the previous section. The zero is selected with probability (1-1/100 = 0.99), hence around 99% of values of this matrix are zero. Utilizing the data structures and routines for sparse matrices makes this transformation method very fast and efficient on large datasets.
Practical Random Projections With the Reuters Corpus Volume 1 Dataset
This section illustrates Random Projections on the Reuters Corpus Volume I Dataset. The dataset is freely accessible online, though for our purposes, it's easiest to looad via Scikit-Learn.
The sklearn.datasets module contains a fetch_rcv1() function that downloads and imports the dataset.
Note: The dataset may take a few minutes to download, if you've never imported it beforehand through this method. Since there's no progress bar, it may appear as if the script is hanging without progressing further. Give it a bit of time, when you run it initially.
The RCV1 dataset is a multilabel dataset, i.e., each data point can belong to multiple classes at the same time, and consists of 103 classes. Each data point has a dimensionality of a whopping 47,236, making it an ideal case for applying fast and cheap Random Projections.
To demonstrate the effectiveness of Random Projections, and to keep things simple, we'll select 500 data points that belong to at least one of the first three classes. The fetch_rcv1() function retrieves the dataset and returns an object with data and targets, both of which are sparse CSR matrices from SciPy.
Let's fetch the Reuters Corpus and prepare it for data transformation:
total_points = 500
# Fetch the dataset
dat = dt.fetch_rcv1()
# Select the sparse matrix's non-zero targets
target_nz = dat.target.nonzero()
# Select only indices of target_nz for data points that belong to
# either of class 1,2,3
ind_class_123 = np.asarray(np.where((target_nz[1]==0) |\
(target_nz[1]==1) |\
(target_nz[1] == 2))).flatten()
# Choose only 500 indices randomly
np.random.seed(0)
ind_class_123 = np.random.choice(ind_class_123, total_points,
replace=False)
# Retreive the row indices of data matrix and target matrix
row_ind = target_nz[0][ind_class_123]
X = dat.data[row_ind,:]
y = np.array(dat.target[row_ind,0:3].todense())
After data preparation, we need a function that creates a visualization of the projected data. To have an idea of the quality of transformation, we can compute the following three matrices:
dist_raw: Matrix of the pairwise Euclidean distances of the actual data points.dist_transform: Matrix of the pairwise Euclidean distances of the transformed data points.abs_diff: Matrix of the absolute difference ofdist_rawanddist_actual
The abs_diff_dist matrix is a good indicator of the quality of the data transformation. Close to zero or small values in this matrix indicate low distortion and a good transformation. We can directly display an image of this matrix or generate a histogram of its values to visually assess the transformation. We can also compute the average of all the values of this matrix to get a single quantitative measure for comparison.
The function create_visualization() creates three plots. The first graph is a scatter plot of projected points along the first two random directions. The second plot is an image of the absolute difference matrix and the third is the histogram of the values of the absolute difference matrix:
def create_visualization(X_transform, y, abs_diff):
fig,ax = plt.subplots(nrows=1, ncols=3, figsize=(20,7))
plt.subplot(131)
plt.scatter(X_transform[y[:,0]==1,0], X_transform[y[:,0]==1,1], c='r', alpha=0.4)
plt.scatter(X_transform[y[:,1]==1,0], X_transform[y[:,1]==1,1], c='b', alpha=0.4)
plt.scatter(X_transform[y[:,2]==1,0], X_transform[y[:,2]==1,1], c='g', alpha=0.4)
plt.legend(['Class 1', 'Class 2', 'Class 3'])
plt.title('Projected data along first two dimensions')
plt.subplot(132)
plt.imshow(abs_diff)
plt.colorbar()
plt.title('Visualization of absolute differences')
plt.subplot(133)
ax = plt.hist(abs_diff.flatten())
plt.title('Histogram of absolute differences')
fig.subplots_adjust(wspace=.3)
Reuters Dataset: Gaussian Random Projection
Let's apply Gaussian Random Projection to the Reuters dataset. The code below runs a for loop for different eps values. If the minimum safe dimensions returned by johnson_lindenstrauss_min_dim is less than the actual data dimensions, then it calls the fit_transform() method of GaussianRandomProjection. The create_visualization() function is then called to create a visualization for that value of eps.
At every iteration, the code also stores the mean absolute difference and the percentage reduction in dimensionality achieved by Gaussian Random Projection:
reduction_dim_gauss = []
eps_arr_gauss = []
mean_abs_diff_gauss = []
for eps in np.arange(0.1, 0.999, 0.2):
min_dim = johnson_lindenstrauss_min_dim(n_samples=total_points, eps=eps)
if min_dim > X.shape[1]:
continue
gauss_proj = GaussianRandomProjection(random_state=0, eps=eps)
X_transform = gauss_proj.fit_transform(X)
dist_raw = euclidean_distances(X)
dist_transform = euclidean_distances(X_transform)
abs_diff_gauss = abs(dist_raw - dist_transform)
create_visualization(X_transform, y, abs_diff_gauss)
plt.suptitle('eps = ' + '{:.2f}'.format(eps) + ', n_components = ' + str(X_transform.shape[1]))
reduction_dim_gauss.append(100-X_transform.shape[1]/X.shape[1]*100)
eps_arr_gauss.append(eps)
mean_abs_diff_gauss.append(np.mean(abs_diff_gauss.flatten()))





The images of the absolute difference matrix and its corresponding histogram indicate that most of the values are close to zero. Hence, a large majority of the pair of points maintain their actual distance in the low dimensional space, retaining the original structure of data.
To assess the quality of transformation, let's plot the mean absolute difference against eps. Also, the higher the value of eps, the greater the dimensionality reduction. Let's also plot the percentage reduction vs. eps in a second sub-plot:
fig,ax = plt.subplots(nrows=1, ncols=2, figsize=(10,5))
plt.subplot(121)
plt.plot(eps_arr_gauss, mean_abs_diff_gauss, marker='o', c='g')
plt.xlabel('eps')
plt.ylabel('Mean absolute difference')
plt.subplot(122)
plt.plot(eps_arr_gauss, reduction_dim_gauss, marker = 'o', c='m')
plt.xlabel('eps')
plt.ylabel('Percentage reduction in dimensionality')
fig.subplots_adjust(wspace=.4)
plt.suptitle('Assessing the Quality of Gaussian Random Projections')
plt.show()

We can see that using Gaussian Random Projection we can reduce the dimensionality of data to more than 99%! Though, this does come at the cost of a higher distortion of data.
Reuters Dataset: Sparse Random Projection
We can do a similar comparison with sparse Random Projection:
reduction_dim_sparse = []
eps_arr_sparse = []
mean_abs_diff_sparse = []
for eps in np.arange(0.1, 0.999, 0.2):
min_dim = johnson_lindenstrauss_min_dim(n_samples=total_points, eps=eps)
if min_dim > X.shape[1]:
continue
sparse_proj = SparseRandomProjection(random_state=0, eps=eps, dense_output=1)
X_transform = sparse_proj.fit_transform(X)
dist_raw = euclidean_distances(X)
dist_transform = euclidean_distances(X_transform)
abs_diff_sparse = abs(dist_raw - dist_transform)
create_visualization(X_transform, y, abs_diff_sparse)
plt.suptitle('eps = ' + '{:.2f}'.format(eps) + ', n_components = ' + str(X_transform.shape[1]))
reduction_dim_sparse.append(100-X_transform.shape[1]/X.shape[1]*100)
eps_arr_sparse.append(eps)
mean_abs_diff_sparse.append(np.mean(abs_diff_sparse.flatten()))





In the case of Random Projection, the absolute difference matrix appears similar to the one of Gaussian projection. The projected data on the first two dimensions, however, has a more interesting pattern, with many points mapped on the coordinate axis.
Let's also plot the mean absolute difference and percentage reduction in dimensionality for various values of the eps parameter:
fig,ax = plt.subplots(nrows=1, ncols=2, figsize=(10,5))
plt.subplot(121)
plt.plot(eps_arr_sparse, mean_abs_diff_sparse, marker='o', c='g')
plt.xlabel('eps')
plt.ylabel('Mean absolute difference')
plt.subplot(122)
plt.plot(eps_arr_sparse, reduction_dim_sparse, marker = 'o', c='m')
plt.xlabel('eps')
plt.ylabel('Percentage reduction in dimensionality')
fig.subplots_adjust(wspace=.4)
plt.suptitle('Assessing the Quality of Sparse Random Projections')
plt.show()

The trend of the two graphs is similar to that of a Gaussian Projection. However, the mean absolute difference for Gaussian Projection is lower than that of Random Projection.
Conclusions
In this guide, we discussed the details of two main types of Random Projections, i.e., Gaussian and sparse Random Projection.
We presented the details of the Johnson-Lindenstrauss lemma, the mathematical basis for these methods. We then showed how this method can be used to transform data using Python's sklearn library.
We also illustrated the two methods on a real-life Reuters Corpus Volume I Dataset.
We encourage the reader to try out this method in supervised classification or regression tasks at the pre-processing stage when dealing with very high-dimensional datasets.
Zero to Mastery
Python Monthly 💻🐍 August 2021
21st issue of Python Monthly! Read by 20,000+ Python developers every month. This monthly Python newsletter is focused on keeping you up to date with the industry and keeping your skills sharp, without wasting your valuable time.
Talk Python to Me
#332: Robust Python
Does it seem like your Python projects are getting bigger and bigger? Are you feeling the pain as your codebase expands and gets tougher to debug and maintain? Patrick Viafore is here to help us write more maintainable, longer-lived, and more enjoyable Python code.<br/> <br/> <strong>Links from the show</strong><br/> <br/> <div><b>Pat on Twitter</b>: <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2FPatViaforever" target="_blank" rel="noopener">@PatViaforever</a><br/> <b>Robust Python Book</b>: <a href="/sitelet?url=https%3A%2F%2Fwww.oreilly.com%2Flibrary%2Fview%2Frobust-python%2F9781098100650%2F" target="_blank" rel="noopener">oreilly.com</a><br/> <b>Typing in Python</b>: <a href="/sitelet?url=https%3A%2F%2Fdocs.python.org%2F3%2Flibrary%2Ftyping.html" target="_blank" rel="noopener">docs.python.org</a><br/> <b>mypy</b>: <a href="/sitelet?url=http%3A%2F%2Fmypy-lang.org%2F" target="_blank" rel="noopener">mypy-lang.org</a><br/> <b>SQLModel</b>: <a href="/sitelet?url=https%3A%2F%2Fsqlmodel.tiangolo.com%2F" target="_blank" rel="noopener">sqlmodel.tiangolo.com</a><br/> <b>CUPID principles @ relevant time</b>: <a href="/sitelet?url=https%3A%2F%2Fovercast.fm%2F%2BBYsRlGnE%2F19%3A06" target="_blank" rel="noopener">overcast.fm</a><br/> <b>Stevedore package</b>: <a href="/sitelet?url=https%3A%2F%2Fdocs.openstack.org%2Fstevedore%2Flatest%2F" target="_blank" rel="noopener">docs.openstack.org</a><br/> <b>Watch YouTube live stream edition</b>: <a href="/sitelet?url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DQU3JO4dwT-s" target="_blank" rel="noopener">youtube.com</a><br/> <b>Episode transcripts</b>: <a href="/sitelet?url=https%3A%2F%2Ftalkpython.fm%2Fepisodes%2Ftranscript%2F332%2Frobust-python" target="_blank" rel="noopener">talkpython.fm</a><br/> <br/> <b>Stay in touch with us</b><br/> <b>Subscribe on YouTube (for live streams)</b>: <a href="/sitelet?url=https%3A%2F%2Ftalkpython.fm%2Fyoutube" target="_blank" rel="noopener">youtube.com</a><br/> <b>Follow Talk Python on Twitter</b>: <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2Ftalkpython" target="_blank" rel="noopener">@talkpython</a><br/> <b>Follow Michael on Twitter</b>: <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2Fmkennedy" target="_blank" rel="noopener">@mkennedy</a><br/></div><br/> <strong>Sponsors</strong><br/> <a href='/sitelet?url=https%3A%2F%2Fclubhouse.io%2Ftalkpython'>Clubhouse</a><br> <a href='/sitelet?url=https%3A%2F%2Ftalkpython.fm%2Fmasterworks'>Masterworks.io</a><br> <a href='/sitelet?url=https%3A%2F%2Ftalkpython.fm%2Fassemblyai'>AssemblyAI</a><br> <a href='/sitelet?url=https%3A%2F%2Ftalkpython.fm%2Ftraining'>Talk Python Training</a>
PyBites
How to handle environment variables in Python
In this article I will share 3 libraries I often use to isolate my environment variables from production code.
Why is this important?
Separate config from code
As we can read in The Twelve-Factor App / III. Config:
Apps sometimes store config as constants in the code. This is a violation of twelve-factor, which requires strict separation of config from code.
https://12factor.net/config
Basically you want to be able to make config changes independently from code changes.
We also want to hide secret keys and API credentials! Notice that git is very persistent (PyCon talk: Oops, I committed my password to GitHub) so it’s important to get this right from the start.
First package: python-dotenv
These days I mostly use python-dotenv which makes this straightforward.
First install the library and add it to your requirements (or if you use Poetry it will automatically update your .toml file):
pip install python-dotenv
Secondly make an .env file with your environment variables in it.
It’s important that you ignore this file with git, otherwise you will end up committing sensitive data to your repo / project.
What I usually do is commit an empty .env-example (or .env-template) file so other developers know what they should set (see examples here and here).
So a new developer (or me checking out the repo on another machine) can do a cp .env-template .env and populate the variables. As the (checked out) .gitignore file contains .env, git won’t show it as a file to be staged for commit.
Then, to load in the variables from this file we use two lines of code:
from dotenv import load_dotenv
load_dotenv()
You can now access the environment variables using os.environ, for example:
BACKGROUND_IMG = os.environ["THUMB_BACKGROUND_IMAGE"]
FONT_FILE = os.environ["THUMB_FONT_TTF_FILE"]
To load the config without touching the environment, you can use dotenv_values(".env") which works the same as load_dotenv, except it doesn’t touch the environment, it just returns a dict with the values parsed from the .env file.
Check out the README for additional options.
Second package: python-decouple
Another library I have been using a lot with Django is python-decouple.
The process is pretty similar:
pip install python-decouple
Create an .env file with your config variables and “gitignore” it.
Then in your code you can use the config object. As per the example in the docs:
from decouple import config
SECRET_KEY = config('SECRET_KEY')
DEBUG = config('DEBUG', default=False, cast=bool)
EMAIL_HOST = config('EMAIL_HOST', default='localhost')
EMAIL_PORT = config('EMAIL_PORT', default=25, cast=int)
The casting and the ability to specify defaults are really convenient.
Another useful option is the Csv helper. For example having this in our .env file for our platform (a Django app):
ALLOWED_HOSTS=.localhost, .herokuapp.com
We can retrieve this variable in settings.py like this:
ALLOWED_HOSTS = config('ALLOWED_HOSTS', cast=Csv())
Third package: dj-database-url
And while we are here, there is one more package I want to show you: dj-database-url, which makes it easier to load in your database URL.
As per the docs:
The dj_database_url.config method returns a Django database connection dictionary, populated with all the data specified in your URL. There is also a conn_max_age argument to easily enable Django’s connection pool.
https://pypi.org/project/dj-database-url/
And here is how to use it:
import dj_database_url
DATABASES = {
'default': dj_database_url.config(
default=config('DATABASE_URL')
)
}
Nice and clean!
This is what I mostly use, for more options, check out python-decouple‘s README here.
Python Tips
As a recap, here is the python-decouple code in a concise tip you can easily paste into your project:
# pip install python-decouple dj-database-url
from decouple import config, Csv
import dj_database_url
SECRET_KEY = config('SECRET_KEY')
DEBUG = config('DEBUG', default=False, cast=bool)
ALLOWED_HOSTS = config('ALLOWED_HOSTS', cast=Csv())
DATABASES = {
'default': dj_database_url.config(
default=config('DATABASE_URL')
)
}
We love practical tips like these, to get our growing collection check out our book: PyBites Python Tips – 250 Bulletproof Python Tips That Will Instantly Make You A Better Developer
And with that we got a wrap. I hope this has been useful and will make it easier for you to separate config from code, which I wholeheartedly agree with The Twelve-Factor App, is important.
— Bob
Glyph Lefkowitz
Unproblematize
The essence of software engineering is solving problems.
The first impression of this insight will almost certainly be that it seems like a good thing. If you have a problem, then solving it is great!
But software engineers are more likely to have mental health problems1 than those who perform mechanical labor, and I think our problem-oriented world-view has something to do with that.
So, how could solving problems be a problem?
As an example, let’s consider the idea of a bug tracker.
For many years, in the field of software, any system used to track work has been commonly referred to as a “bug tracker”. In recent years, the labels have become more euphemistic and general, and we might now call them “issue trackers”. We have Sapir-Whorfed2 our way into the default assumption that any work that might need performing is a degenerate case of a problem.
We can contrast this with other fields. Any industry will need to track work that must be done. For example, in doing some light research for this post, I discovered that the relevant term of art in construction3 is typically “Project Management” or “Task Management” software. “Projects” and “Tasks” are no less hard work, but the terms do have a different valence than “Bugs” and “Issues”.
I don’t think we can start to fix this ... problem ... by attempting to change the terminology. Firstly, the domain inherently lends itself to this sort of language, which is why it emerged in the first place.
Secondly, Atlassian has desperately been trying to get everybody to call their bug tracker a “software development tool” where you write “stories” for years, and nobody does. It’s an issue tracker where you file bugs, and that’s what everyone calls it and describes what they do with it. Even they have to protest, perhaps a bit too much, that it’s “way more than a bug and issue tracker”4.
This pervasive orientation towards “problems” as the atom of work does extend to any knowledge work, and thereby to any “productivity system”. Any to-do list is, at its core, a list of problems. You wouldn’t put an item on the list if you were happy with the way the world was. Therefore every unfinished item in any to-do list is a little pebble of worry.
As of this writing, I have almost 1000 unfinished tasks on my personal to-do list.
This is to say nothing of any tasks I have to perform at work, not to mention the implicit א0 of additional unfinished tasks once one considers open source issue trackers for projects I work on.
It’s not really reasonable to opt out of this habit of problematizing everything. This monument to human folly that I’ve meticulously constructed out of the records of aspirations which exceed my capacity is, in fact, also an excellent prioritization tool. If you’re a good engineer, or even just good at making to-do lists, you’ll inevitably make huge lists of problems. On some level, this is what it means to set an intention to make the world — or at least your world — better.
On a different level though, this is how you set out to systematically give yourself anxiety, depression, or both. It’s clear from a wealth of neurological research that repeated experiences and thoughts change neural structures5. Thinking the same thought over and over literally re-wires your brain. Thinking the thought “here is another problem” over and over again forever is bound to cause some problems of its own.
The structure of to-do apps, bug trackers and the like is such that when an item is completed — when a problem is solved — it is subsequently removed from both physical view and our mind’s eye. What would be the point of simply lingering on a completed task? All the useful work is, after all, problems that haven’t been solved yet. Therefore the vast majority of our time is spent contemplating nothing but problems, prompting the continuous potentiation6 of neural pathways which lead to despair.
I don’t want to pretend that I have a cure for this self-inflicted ailment. I do, however, have a humble suggestion for one way to push back just a little bit against the relentless, unending tide of problems slowly eroding the shores of our souls: a positivity journal.
By “journal”, I do mean a private journal. Public expressions of positivity7 can help; indeed, some social and cultural support for expressing positivity is an important tool for maintaining a positive mind-set. However, it may not be the best starting point.
Unfortunately, any public expression becomes a discourse, and any discourse inevitably becomes a dialectic. Any expression of a view in public is seen by some as an invitation to express its opposite8. Therefore one either becomes invested in defending the boundaries of a positive community space — a psychically exhausting task in its own right — or one must constantly entertain the possibility that things are, in fact, bad, when one is trying to condition one’s brain to maintain the ability to recognize when things are actually good.
Thus my suggestion to write something for yourself, and only for yourself.
Personally, I use a template that I fill out every day, with four sections:
-
“Summary”. Summarize the day in one sentence that encapsulates its positive vibes. Honestly I put this in there because the Notes app (which is what I’m using to maintain this) shows a little summary of the contents of the note, and I was getting annoyed by just seeing “Proud:” as the sole content of that summary. But once I did so, I found that it helps to try to synthesize a positive narrative, as your brain may be constantly trying to assemble a negative one. It can help to write this last, even if it’s up at the top of your note, once you’ve already filled out some of the following sections.
-
“I’m proud of:”. First, focus on what you personally have achieved through your skill and hard work. This can be very difficult, if you are someone who has a habit of putting yourself down. Force yourself to acknowledge that you did something useful, even if you didn’t finish anything, you almost certainly made progress and that progress deserves celebration.
-
“I’m grateful to:”. Who are you grateful to? Why? What did they do for you? Once you’ve made the habit of allowing yourself to acknowledge your own accomplishments, it’s easy to see those; pay attention to the ways in which others support and help you. Thank them by name.
-
“I’m lucky because:”. Particularly in post-2020 hell-world it’s easy to feel like every random happenstance is an aggravating tragedy. But good things happen randomly all the time, and it’s easy to fail to notice them. Take a moment to notice things that went well for no good reason, because you’re definitely going to feel attacked by the universe when bad things happen for no good reason; and they will.
Although such a journal is private, it’s helpful to actually write out the answers, to focus on them, to force yourself to get really specific.
I hope this tool is useful to someone out there. It’s not going to solve any problems, but perhaps it will make the world seem just a little brighter.
-
“Maintaining Mental health on Software Development Teams”, Lena Kozar and Vova Vovk, in InfoQ ↩
-
“Construction Task and Project Tracking”, from Raptor Project Management Software ↩
-
Jira Features List, Atlassian Software ↩
-
“Culture Wires the Brain: A Cognitive Neuroscience Perspective”, Denise C. Park and Chih-Mao Huang, Perspect Psychol Sci. 2010 Jul 1; 5(4): 391–400. ↩
-
Long-term potentiation and learning, J L Martinez Jr, B E Derrick ↩
-
The #PositivePython hashtag on Twitter was a lovely experiment and despite my cautions here about public solutions to this problem, it’s generally pleasant to participate in. ↩
Python Piedmont Triad User Group
Lunch and learn series
PYPTUG Lunch and Learn
In order to help those starting out with python, we are starting a lunch and learn series. You can see upcoming lunch and learns on our meetup page:
https://www.meetup.com/PYthon-Piedmont-Triad-User-Group-PYPTUG/

