Sitelet https://web.archive.org/web/20210903131430/https://planetpython.org/

skip to navigation
skip to content

Planet Python

Last update: September 03, 2021 10:40 AM UTC

September 03, 2021


Kushal Das

Default values, documentation and Ansible

While testing my qubes_ansible project on the upcoming Qubes OS 4.1 project, I noticed something really strange. But, before getting into that, this Ansible module and the connection plugin are for Qubes OS only, and based on the excellent Python modules provided by the Qubes team.

The error goes like this during the fact gathering steps (reformatted for the blog):

fatal: [debian-10]: UNREACHABLE! => {
    "changed": false,
    "msg": "Failed to create temporary directory.In some cases, you may have
    been able to authenticate and did not have permissions on the target
    directory. Consider changing the remote tmp path in ansible.cfg to a path
    rooted in \"/tmp\", for more error information use -vvv. Failed command
    was: ( umask 77 && mkdir -p \"` echo ~The *user* is the default user in
    Qubes./.ansible/tmp `\"&& mkdir \"` echo ~The *user* is the default user in
    Qubes./.ansible/tmp/ansible-tmp-1630548982.9355698-7707-90110802425258 `\"
    && echo ansible-tmp-1630548982.9355698-7707-90110802425258=\"` echo ~The
    *user* is the default user in
    Qubes./.ansible/tmp/ansible-tmp-1630548982.9355698-7707-90110802425258 `\"
    ), exited with result 1, stderr output: mkdir: cannot create directory
    ‘~The ~The *user* account as default in Qubes OS. ~The ~The *user* account
    as default in Qubes OS. account as default in Qubes OS. ~The ~The *user*
    account as default in Qubes OS. ~The ~The *user* account as default in
    Qubes OS. account as default in Qubes OS. account as default in Qubes OS.
    is the default user in Qubes.’: File name too long\n",
    "unreachable": true
}

Most important part is the default user's home directory part, echo ~The user is the default user in Qubes./.ansible/tmp. For a moment I totally freaked out, as this looks like documentation. After reading the code more, I can see it is coming from the DOCUMENTATION variable in my plugin. After playing around a bit more and trying out different values I can see that the default value mentioned in the documentation is becoming the default value in the Python code.

After searching more I can see that the Ansible developers want the documentation string to be the gold standard and the code is parsing it find the default values. In my mind this is more confusing. I would expect the default value to be declared inside of the code.

Parsing the DOCUMENTATION and then finding the default values there in a Python code still does not fit in my brain. Fixed the issue for now, let me see what other surprises are waiting in the future.

September 03, 2021 02:21 AM UTC


Podcast.__init__

Monitor The Health Of Your Machine Learning Products In Production With Evidently

You've got a machine learning model trained and running in production, but that's only half of the battle. Are you certain that it is still serving the predictions that you tested? Are the inputs within the range of tolerance that you designed? Monitoring machine learning products is an essential step of the story so that you know when it needs to be retrained against new data, or parameters need to be adjusted. In this episode Emeli Dral shares the work that she and her team at Evidently are doing to build an open source system for tracking and alerting on the health of your ML products in production. She discusses the ways that model drift can occur, the types of metrics that you need to track, and what to do when the health of your system is suffering. This is an important and complex aspect of the machine learning lifecycle, so give it a listen and then try out Evidently for your own projects.

Summary

You’ve got a machine learning model trained and running in production, but that’s only half of the battle. Are you certain that it is still serving the predictions that you tested? Are the inputs within the range of tolerance that you designed? Monitoring machine learning products is an essential step of the story so that you know when it needs to be retrained against new data, or parameters need to be adjusted. In this episode Emeli Dral shares the work that she and her team at Evidently are doing to build an open source system for tracking and alerting on the health of your ML products in production. She discusses the ways that model drift can occur, the types of metrics that you need to track, and what to do when the health of your system is suffering. This is an important and complex aspect of the machine learning lifecycle, so give it a listen and then try out Evidently for your own projects.

Announcements

  • Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
  • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
  • Your host as usual is Tobias Macey and today I’m interviewing Emeli Dral about monitoring machine learning models in production with Evidently

Interview

  • Introductions
  • How did you get introduced to Python?
  • Can you describe what Evidently is and the story behind it?
  • What are the metrics that are useful for determining the performance and health of a machine learning model?
    • What are the questions that you are trying to answer with those metrics?
  • How does monitoring of machine learning models compare to monitoring of infrastructure or "traditional" software projects?
  • What are the failure modes for a model?
  • Can you describe the design and implementation of Evidently?
    • How has the architecture changed or evolved since you started working on it?
  • What categories of model is Evidently designed to work with?
    • What are some strategies for making models conducive to monitoring?
  • What is involved in monitoring a model on a continuous basis?
  • What are some considerations when establishing useful thresholds for metrics to alert on?
    • Once an alert has been triggered what is the process for resolving it?
    • If the training process takes a long time, how can you mitigate the impact of a model failure until the new/updated version is deployed?
  • What are the most interesting, innovative, or unexpected ways that you have seen Evidently used?
  • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Evidently?
  • When is Evidently the wrong choice?
  • What do you have planned for the future of Evidently?

Keep In Touch

Picks

Links

The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA

September 03, 2021 01:03 AM UTC

September 02, 2021


Python for Beginners

Check If a List has Duplicate Elements

Lists are the most used data structures in Python. While programming, you may land into a situation where you will need a list containing only unique elements or you want to check if a list has duplicate elements. In this article, we will look at different ways to check if a list has duplicate elements in it.

Check if a list has duplicate Elements using Sets

We know that sets in Python contain only unique elements. We can use this property of sets to check if a list has duplicate elements or not. 

For this, we will create a set from the elements of the list. After that, we will check the size of the list and the set. If the size of both the objects are equal, it will confirm that the list has no duplicate elements. If the size of the set is greater than the list, it will mean that the list contains duplicate elements.  We can understand this from the following example.

def check_duplicate(l):
    mySet = set(l)
    if len(mySet) == len(l):
        print("List has no duplicate elements.")
    else:
        print("The list contains duplicate elements")


list1 = [1, 2, 3, 4, 5, 6, 7]
print("List1 is:", list1)
check_duplicate(list1)
list2 = [1, 2, 1, 2, 4, 6, 7]
print("List2 is:", list2)
check_duplicate(list2)

Output:

List1 is: [1, 2, 3, 4, 5, 6, 7]
List has no duplicate elements.
List2 is: [1, 2, 1, 2, 4, 6, 7]
The list contains duplicate elements

In the above approach, we need to create a set from all the elements of the list. After that, we also check the size of the set and the list. These operations are very costly. 

Instead of using this approach, we can search only for the first duplicate element. To do this, we will start from the first element of the list and will keep adding them to the set. Before adding the elements to the set, we will check if the element is already present in the set or not. If yes, the list contains duplicate elements. If we are able to add each element of the list to the set, the list does not contain any duplicate element. This can be understood from the following example.

def check_duplicate(l):
    visited = set()
    has_duplicate = False
    for element in l:
        if element in visited:
            print("The list contains duplicate elements.")
            has_duplicate = True
            break
        else:
            visited.add(element)
    if not has_duplicate:
        print("List has no duplicate elements.")


list1 = [1, 2, 3, 4, 5, 6, 7]
print("List1 is:", list1)
check_duplicate(list1)
list2 = [1, 2, 1, 2, 4, 6, 7]
print("List2 is:", list2)
check_duplicate(list2)

Output:

List1 is: [1, 2, 3, 4, 5, 6, 7]
List has no duplicate elements.
List2 is: [1, 2, 1, 2, 4, 6, 7]
The list contains duplicate elements.

Check if a list has duplicate elements using the count() method

To check if a list has only unique elements, we can also count the occurrence of the different elements in the list. For this, we will use the count() method. The count() method, when invoked on a list, takes the element as input argument and returns the number of times the element is present in the list.  

For checking if the list contains duplicate elements, we will count the frequency of each element. At the same time, we will also maintain a list of visited elements so that we don’t have to count the occurrences of the visited elements. Once the count of any element is found to be greater than one, it will prove that the list has duplicate elements. We can implement this as follows.

def check_duplicate(l):
    visited = set()
    has_duplicate = False
    for element in l:
        if element in visited:
            pass
        elif l.count(element) == 1:
            visited.add(element)
        elif l.count(element) > 1:
            has_duplicate = True
            print("The list contains duplicate elements.")
            break
    if not has_duplicate:
        print("List has no duplicate elements.")


list1 = [1, 2, 3, 4, 5, 6, 7, 8]
print("List1 is:", list1)
check_duplicate(list1)
list2 = [1, 2, 1, 2, 4, 6, 7, 8]
print("List2 is:", list2)
check_duplicate(list2)

Output:

List1 is: [1, 2, 3, 4, 5, 6, 7, 8]
List has no duplicate elements.
List2 is: [1, 2, 1, 2, 4, 6, 7, 8]
The list contains duplicate elements.

 Check if a list has duplicate elements using the counter() method

We can also use the counter() method to check if a list has only unique elements or not. The counter() method. The counter() method takes an iterable object as an input and returns a python dictionary in which the keys consist of the elements of the iterable object and the values associated with the keys are the frequency of the elements. After getting the frequency of each element of the list using the counter() method, we can check if the frequency of any element is greater than one or not. If yes, the list contains duplicate elements. Otherwise not.

from collections import Counter


def check_duplicate(l):
    counter = Counter(l)
    has_duplicate = False
    frequencies = counter.values()
    for i in frequencies:
        if i > 1:
            has_duplicate = True
            print("The list contains duplicate elements.")
            break
    if not has_duplicate:
        print("List has no duplicate elements.")


list1 = [1, 2, 3, 4, 5, 6, 7, 8]
print("List1 is:", list1)
check_duplicate(list1)
list2 = [1, 2, 1, 2, 4, 6, 7, 8]
print("List2 is:", list2)
check_duplicate(list2)

Output:

List1 is: [1, 2, 3, 4, 5, 6, 7, 8]
List has no duplicate elements.
List2 is: [1, 2, 1, 2, 4, 6, 7, 8]
The list contains duplicate elements.

Conclusion

In this article, we have discussed four ways to check if a list has only unique elements or not. We have used sets, count() and counter() methods to implement our approaches. To learn more about lists, you can read this article on list comprehension.

The post Check If a List has Duplicate Elements appeared first on PythonForBeginners.com.

September 02, 2021 03:19 PM UTC


Mike Driscoll

Creating a File Search GUI with wxPython

Have you ever needed to search for a file on your computer? Most operating systems have a way to do this. Windows Explorer has a search function and there’s also a search built-in to the Start Menu now. Other operating systems like Mac and Linux are similar. There are also applications that you can download that are sometimes faster at searching your hard drive than the built-in ones are.

In this article, you will be creating a simple file search utility using wxPython.

You will want to support the following tasks for the file search tool:

You can download the source code from this article on GitHub.

Let’s get started!

Designing Your File Search Utility

It is always fun to try to recreate a tool that you use yourself. However in this case, you will just take the features mentioned above and create a straight-forward user interface. You can use a wx.SearchCtrl for searching for files and an ObjectListView for displaying the results. For this particular utility, a wx.CheckBox or two will work nicely for telling your application to search in sub-directories or if the search term is case-sensitive or not.

Here is a mockup of what the application will eventually look like:

File Search MockupFile Search Mockup

Now that you have a goal in mind, let’s go ahead and start coding!

Creating the File Search Utility

Your search utility will need two modules. The first module will be called main and it will hold your user interface and most of the application’s logic. The second module is named search_threads and it will contain the logic needed to search your file system using Python’s threading module. You will use pubsub to update the main module as results are found.

The main script

The main module has the bulk of the code for your application. If you go on and enhance this application, the search portion of the code could end up having the majority of the code since that is where a lot of the refinement of your code should probably go.

Regardless, here is the beginning of main:

# main.py

import os
import sys
import subprocess
import time
import wx

from ObjectListView import ObjectListView, ColumnDefn
from pubsub import pub
from search_threads import SearchFolderThread, SearchSubdirectoriesThread

This time around, you will be using a few more built-in Python modules, such as os, sys, subprocess and time. The other imports are pretty normal, with the last one being a couple of classes that you will be creating based around Python’s Thread class from the threading module.

For now though, let’s just focus on the main module.

Here’s the first class you need to create:

class SearchResult:

    def __init__(self, path, modified_time):
        self.path = path
        self.modified = time.strftime('%D %H:%M:%S',
                                      time.gmtime(modified_time))

The SearchResult class is used for holding information about the results from your search. It is also used by the ObjectListView widget. Currently, you will use it to hold the full path to the search result as well as the file’s modified time. You could easily enhance this to also include file size, creation time, etc.

Now let’s create the MainPanel which houses most of UI code:

class MainPanel(wx.Panel):

    def __init__(self, parent):
        super().__init__(parent)
        self.search_results = []
        self.main_sizer = wx.BoxSizer(wx.VERTICAL)
        self.create_ui()
        self.SetSizer(self.main_sizer)
        pub.subscribe(self.update_search_results, 'update')

The __init__() method gets everything set up. Here you create the main_sizer, an empty list of search_results and a listener or subscription using pubsub. You also call create_ui() to add the user interface widgets to the panel.

Let’s see what’s in create_ui() now:

def create_ui(self):
    # Create the widgets for the search path
    row_sizer = wx.BoxSizer()
    lbl = wx.StaticText(self, label='Location:')
    row_sizer.Add(lbl, 0, wx.ALL | wx.CENTER, 5)
    self.directory = wx.TextCtrl(self, style=wx.TE_READONLY)
    row_sizer.Add(self.directory, 1, wx.ALL | wx.EXPAND, 5)
    open_dir_btn = wx.Button(self, label='Choose Folder')
    open_dir_btn.Bind(wx.EVT_BUTTON, self.on_choose_folder)
    row_sizer.Add(open_dir_btn, 0, wx.ALL, 5)
    self.main_sizer.Add(row_sizer, 0, wx.EXPAND)

There are quite a few widgets to add to this user interface. To start off, you add a row of widgets that consists of a label, a text control and a button. This series of widgets allows the user to choose which directory they want to search using the button. The text control will hold their choice.

Now let’s add another row of widgets:

# Create search filter widgets
row_sizer = wx.BoxSizer()
lbl = wx.StaticText(self, label='Limit search to filetype:')
row_sizer.Add(lbl, 0, wx.ALL|wx.CENTER, 5)

self.file_type = wx.TextCtrl(self)
row_sizer.Add(self.file_type, 0, wx.ALL, 5)

self.sub_directories = wx.CheckBox(self, label='Sub-directories')
row_sizer.Add(self.sub_directories, 0, wx.ALL | wx.CENTER, 5)

self.case_sensitive = wx.CheckBox(self, label='Case-sensitive')
row_sizer.Add(self.case_sensitive, 0, wx.ALL | wx.CENTER, 5)
self.main_sizer.Add(row_sizer)

This row of widgets contains another label, a text control and two instances of wx.Checkbox. These are the filter widgets which control what you are searching for. You can filter based on any of the following:

The latter two options are represented by using the wx.Checkbox widget.

Let’s add the search control next:

# Add search bar
self.search_ctrl = wx.SearchCtrl(
    self, style=wx.TE_PROCESS_ENTER, size=(-1, 25))
self.search_ctrl.Bind(wx.EVT_SEARCHCTRL_SEARCH_BTN, self.on_search)
self.search_ctrl.Bind(wx.EVT_TEXT_ENTER, self.on_search)
self.main_sizer.Add(self.search_ctrl, 0, wx.ALL | wx.EXPAND, 5)

The wx.SearchCtrl is the widget to use for searching. You could quite easily use a wx.TextCtrl instead though. Regardless, in this case you bind to the press of the Enter key and to the mouse click of the magnifying class within the control. If you do either of these actions, you will call search().

Now let’s add the last two widgets and you will be done with the code for create_ui():

# Search results widget
self.search_results_olv = ObjectListView(
    self, style=wx.LC_REPORT | wx.SUNKEN_BORDER)
self.search_results_olv.SetEmptyListMsg("No Results Found")
self.main_sizer.Add(self.search_results_olv, 1, wx.ALL | wx.EXPAND, 5)
self.update_ui()

show_result_btn = wx.Button(self, label='Open Containing Folder')
show_result_btn.Bind(wx.EVT_BUTTON, self.on_show_result)
self.main_sizer.Add(show_result_btn, 0, wx.ALL | wx.CENTER, 5)

The results of your search will appear in your ObjectListView widget. You also need to add a button that will attempt to show the result in the containing folder, kind of like how Mozilla Firefox has a right-click menu called “Open Containing Folder” for opening downloaded files.

The next method to create is on_choose_folder():

def on_choose_folder(self, event):
    with wx.DirDialog(self, "Choose a directory:",
                      style=wx.DD_DEFAULT_STYLE,
                      ) as dlg:
        if dlg.ShowModal() == wx.ID_OK:
            self.directory.SetValue(dlg.GetPath())

You need to allow the user to select a folder that you want to conduct a search in. You could let the user type in the path, but that is error-prone and you might need to add special error checking. Instead, you opt to use a wx.DirDialog, which prevents the user from entering a non-existent path. It is possible for the user to select the folder, then delete the folder before executing the search, but that would be an unlikely scenario.

Now you need a way to open a folder with Python:

def on_show_result(self, event):
    """
    Attempt to open the folder that the result was found in
    """
    result = self.search_results_olv.GetSelectedObject()
    if result:
        path = os.path.dirname(result.path)
        try:
            if sys.platform == 'darwin':
                subprocess.check_call(['open', '--', path])
            elif 'linux' in sys.platform:
                subprocess.check_call(['xdg-open', path])
            elif sys.platform == 'win32':
                subprocess.check_call(['explorer', path])
        except:
            if sys.platform == 'win32':
                # Ignore error on Windows as there seems to be
                # a weird return code on Windows
                return

            message = f'Unable to open file manager to {path}'
            with wx.MessageDialog(None, message=message,
                                  caption='Error',
                                  style= wx.ICON_ERROR) as dlg:
                dlg.ShowModal()

The on_show_result() method will check what platform the code is running under and then attempt to launch that platform’s file manager. Windows uses Explorer while Linux uses xdg-open for example.

During testing, it was noticed that on Windows, Explorer returns a non-zero result even when it opens Explorer successfully, so in that case you just ignore the error. But on other platforms, you can show a message to the user that you were unable to open the folder.

The next bit of code you need to write is the on_search() event handler:

def on_search(self, event):
    search_term = self.search_ctrl.GetValue()
    file_type = self.file_type.GetValue()
    file_type = file_type.lower()
    if '.' not in file_type:
        file_type = f'.{file_type}'

    if not self.sub_directories.GetValue():
        # Do not search sub-directories
        self.search_current_folder_only(search_term, file_type)
    else:
        self.search(search_term, file_type)

When you click the “Search” button, you want it to do something useful. That is where the code above comes into play. Here you get the search_term and the file_type. To prevent issues, you put the file type in lower case and you will do the same thing during the search.

Next you check to see if the sub_directories check box is checked or not. If sub_directories is unchecked, then you call search_current_folder_only(); otherwise you call search().

Let’s see what goes into search() first:

def search(self, search_term, file_type):
    """
    Search for the specified term in the directory and its
    sub-directories
    """
    folder = self.directory.GetValue()
    if folder:
        self.search_results = []
        SearchSubdirectoriesThread(folder, search_term, file_type,
                                   self.case_sensitive.GetValue())

Here you grab the folder that the user has selected. In the event that the user has not chosen a folder, the search button will not do anything. But if they have chosen something, then you call the SearchSubdirectoriesThread thread with the appropriate parameters. You will see what the code in that class is in a later section.

But first, you need to create the search_current_folder_only() method:

def search_current_folder_only(self, search_term, file_type):
    """
    Search for the specified term in the directory only. Do
    not search sub-directories
    """
    folder = self.directory.GetValue()
    if folder:
        self.search_results = []
        SearchFolderThread(folder, search_term, file_type,
                           self.case_sensitive.GetValue())

This code is pretty similar to the previous function. Its only difference is that it executes
SearchFolderThread instead of SearchSubdirectoriesThread.

The next function to create is update_search_results():

def update_search_results(self, result):
    """
    Called by pubsub from thread
    """
    if result:
        path, modified_time = result
        self.search_results.append(SearchResult(path, modified_time))
    self.update_ui()

When a search result is found, the thread will post that result back to the main application using a thread-safe method and pubsub. This method is what will get called assuming that the topic matches the subscription that you created in the __init__(). Once called, this method will append the result to search_results and then call update_ui().

Speaking of which, you can code that up now:

def update_ui(self):
    self.search_results_olv.SetColumns([
        ColumnDefn("File Path", "left", 300, "path"),
        ColumnDefn("Modified Time", "left", 150, "modified")
    ])
    self.search_results_olv.SetObjects(self.search_results)

The update_ui() method defines the columns that are shown in your ObjectListView widget. It also calls SetObjects() which will update the contents of the widget and show your search results to the user.

To wrap up the main module, you will need to write the Search class:

class Search(wx.Frame):

    def __init__(self):
        super().__init__(None, title='Search Utility',
                         size=(600, 600))
        pub.subscribe(self.update_status, 'status')
        panel = MainPanel(self)
        self.statusbar = self.CreateStatusBar(1)
        self.Show()

    def update_status(self, search_time):
        msg = f'Search finished in {search_time:5.4} seconds'
        self.SetStatusText(msg)

if __name__ == '__main__':
    app = wx.App(False)
    frame = Search()
    app.MainLoop()

This class creates the MainPanel which holds most of the widgets that the user will see and interact with. It also sets the initial size of the application along with its title. There is also a status bar that will be used to communicate to the user when a search has finished and how long it took for said search to complete.

Here is what the application will look like:

A Search Utility

Now let’s move on and create the module that holds your search threads.

The search_threads Module

The search_threads module contains the two Thread classes that you will use for searching your file system. The thread classes are actually quite similar in their form and function.

Let’s get started:

# search_threads.py

import os
import time
import wx

from pubsub import pub
from threading import Thread

These are the modules that you will need to make this code work. You will be using the os module to check paths, traverse the file system and get statistics from files. You will use pubsub to communicate with your application when your search returns results.

Here is the first class:

class SearchFolderThread(Thread):

    def __init__(self, folder, search_term, file_type, case_sensitive):
        super().__init__()
        self.folder = folder
        self.search_term = search_term
        self.file_type = file_type
        self.case_sensitive = case_sensitive
        self.start()

This thread takes in the folder to search in, the search_term to look for, a file_type filter and whether or not the search term is case_sensitive. You take these in and assign them to instance variables of the same name. The point of this thread is only to search the contents of the folder that is passed-in, not its sub-directories.

You will also need to override the thread’s run() method:

def run(self):
    start = time.time()
    for entry in os.scandir(self.folder):
        if entry.is_file():
            if self.case_sensitive:
                path = entry.name
            else:
                path = entry.name.lower()

            if self.search_term in path:
                _, ext = os.path.splitext(entry.path)
                data = (entry.path, entry.stat().st_mtime)
                wx.CallAfter(pub.sendMessage, 'update', result=data)
    end = time.time()
    # Always update at the end even if there were no results
    wx.CallAfter(pub.sendMessage, 'update', result=[])
    wx.CallAfter(pub.sendMessage, 'status', search_time=end-start)

Here you collect the start time of the thread. Then you use os.scandir() to loop over the contents of the folder. If the path is a file, you will check to see if the search_term is in the path and has the right file_type. Should both of those return True, then you get the requisite data and send it to your application using wx.CallAfter(), which is a thread-safe method.

Finally you grab the end_time and use that to calculate the total run time of the search and then send that back to the application. The application will then update the status bar with the search time.

Now let’s check out the other class:

class SearchSubdirectoriesThread(Thread):

    def __init__(self, folder, search_term, file_type, case_sensitive):
        super().__init__()
        self.folder = folder
        self.search_term = search_term
        self.file_type = file_type
        self.case_sensitive = case_sensitive
        self.start()

The SearchSubdirectoriesThread thread is used for searching not only the passed-in folder but also its sub-directories. It accepts the same arguments as the previous class.

Here is what you will need to put in its run() method:

def run(self):
    start = time.time()
    for root, dirs, files in os.walk(self.folder):
        for f in files:
            full_path = os.path.join(root, f)
            if not self.case_sensitive:
                full_path = full_path.lower()

            if self.search_term in full_path and os.path.exists(full_path):
                _, ext = os.path.splitext(full_path)
                data = (full_path, os.stat(full_path).st_mtime)
                wx.CallAfter(pub.sendMessage, 'update', result=data)

    end = time.time()
    # Always update at the end even if there were no results
    wx.CallAfter(pub.sendMessage, 'update', result=[])
    wx.CallAfter(pub.sendMessage, 'status', search_time=end-start)

For this thread, you need to use os.walk() to search the passed in folder and its sub-directories. Besides that, the conditional statements are virtually the same as the previous class.

Wrapping Up

Creating search utilities is not particularly difficult, but it can be time-consuming. Figuring out the edge cases and how to account for them is usually what takes the longest when creating software. In this article, you learned how to create a utility to search for files on your computer.

Here are a few enhancements that you could add to this program:

Related Reading

Want to learn how to create more GUI applications with wxPython? Then check out these resources below:

The post Creating a File Search GUI with wxPython appeared first on Mouse Vs Python.

September 02, 2021 12:30 PM UTC


Stack Abuse

Guide to Numpy's arange() Function

Intro

Numpy is the most popular mathematical computing Python library. It offers a great number of mathematical tools including but not limited to multi-dimensional arrays and matrices, mathematical functions, number generators, and a lot more.

One of the fundamental tools in NumPy is the ndarray - an N-dimensional array. Today, we're going to create ndarrays, generated in certain ranges using the NumPy.arange() function.

Parameters and Return

numpy.arange([start, ]stop, [step, ]dtype=None)

Returns evenly spaced values within a given interval where:

The method returns an ndarray of of evenly spaced values. If the array returns floating-point elements the array's length will be ceil((stop - start)/step).

np.arange() by Example

Importing NumPy

To start working with NumPy, we need to import it, as it's an external library:

import NumPy as np

If not installed, you can easily install it via pip:

$ pip install numpy

All-Argument np.arange()

Let's see how arange() works with all the arguments for the function. For instance, say we want a sequence to start at 0, stop at 10, with a step size of 3, while producing integers.

In a Python environement, or REPL, let's generate a sequence in a range:

>>> result_array = np.arange(start=0, stop=10, step=2, dtype=int)

The array is an ndarray containing the generated elements:

>>> result_array
array([0, 2, 4, 6, 8])

It's worth noting that the stop element isn't included, while the start element is included, hence we have a 0 but not a 10 even though the next element in the sequence should be a 10.

Note: As usual, you an provide positional arguments, without naming them or named arguments:

array = np.arange(start=0, stop=10, step=2, dtype=int)
# These two statements are the same
array = np.arange(0, 10, 2, int)

For the sake of brevity, the latter is oftentimes used, and the positions of these arguments must follow the sequence of start, stop, step and dtype.

np.arange() with stop

If only one argument is provided, it will be treated as the stop value. It will output all numbers up to but not including the stop number, with a default step of 1 and start of 0:

>>> result_array = np.arange(5)
>>> result_array
array([0, 1, 2, 3, 4])

np.arange() with start and stop

With two arguments, they default to start and stop, with a default step of 1 - so you can easily create a specific range without thinking about the step size:

>>> result_array = np.arange(5, 10)
>>> result_array
array([5, 6, 7, 8, 9])

Like with previous examples, you can also use floating point numbers here instead of integers. For example, we can start at 5.5:

>>> result_array = np.arange(5.5, 11.75)

The resulting array will be:

>>> result_array
array([ 5.5,  6.5,  7.5,  8.5,  9.5, 10.5, 11.5])

np.arange() with start, stop and step

The default dtype is None and in that case, ints are used so having an integer-based range is easy to create with a start, stop and step. For instance, let's generate a sequence of all the even numbers between 6 (inclusive) and 22 (exclusive):

>>> result_array = np.arange(6, 22, 2)

The result will be all even numbers between 6 up to but not including 22:

>>> result_array
array([ 6,  8, 10, 12, 14, 16, 18, 20])

np.arange() for Reversed Ranges

We can also pass in negative parameters into the np.arange() function to get a reversed array of numbers.

The start will be the larger number we want to start counting from, the stop will be the lower one, and the step will be a negative number:

result_array = np.arange(start=30,stop=14, step=-3)

The result will be an array of descending numbers with a negative step of 3:

>>> result_array
array([30, 27, 24, 21, 18, 15])

Creating Empty NDArrays with np.arange()

We can also create an empty arange as follows:

>>> result_array = np.arange(0)

The result will be an empty array:

>>> result_array
array([], dtype=int32)

This happens because 0 is the stop value we've set, and the start value is also 0 by default. So, the counting stops before starting.

Another case where the result will be an empty array is when the start value is higher than the stop value while the step is positive. For example:

>>> result_array = np.arange(start=30, stop=10, step=1)

The result will also be an empty array.

>>> result_array
array([], dtype=int32)

This can also happen the other way around. We can start with a small number, stop at a larger number, and have the step as a negative number. The output will be an empty array too:

>>> result_array = np.arange(start=10, stop=30, step=-1)

This also results in an empty ndarray:

>>> result_array
array([], dtype=int32)

Supported Data Types for np.arange()

The dtype argument, which defaults to int can be any valid NumPy data type.

Note: This isn't to be confused with standard Python data types, though.

You can use the shorthand version for some of the more common datatypes, or the full name, prefixed with np.:

np.arange(..., dtype=int)
np.arange(..., dtype=np.int32)
np.arange(..., dtype=np.int64)

For some other data types, such as np.csignle, you'll prefix the type with np.:

>>> result_array = np.arange(start=10, stop=30, step=1, dtype=np.csingle)
>>> result_array
array([10.+0.j, 11.+0.j, 12.+0.j, 13.+0.j, 14.+0.j, 15.+0.j, 16.+0.j,
       17.+0.j, 18.+0.j, 19.+0.j, 20.+0.j, 21.+0.j, 22.+0.j, 23.+0.j,
       24.+0.j, 25.+0.j, 26.+0.j, 27.+0.j, 28.+0.j, 29.+0.j],
      dtype=complex64)

A common short-hand data type is a float:

>>> result_array = np.arange(start=10, stop=30, step=1, dtype=float)
>>> result_array
array([10., 11., 12., 13., 14., 15., 16., 17., 18., 19., 20., 21., 22.,
       23., 24., 25., 26., 27., 28., 29.])

For a list of all supported NumPy data types, take a look at the official documentation.

np.arange() vs np.linspace()

np.linspace() is similar to np.arange() in returning evenly spaced arrays. However, there are a couple of differences.

With np.linspace(), you specify the number of samples in a certain range instead of specifying the step. In addition, you can include endpoints in the returned array. Another difference is that np.linspace() can generate multiple arrays instead of returning only one array.

This is a simple example of np.linspace() with the endpoint included and 5 samples:

>>> result_array = np.linspace(0, 20, num=5, endpoint=True)
>>> result_array
array([ 0.,  5., 10., 15., 20.])

Here, both the number of samples and the step size is 5, but that's coincidental:

>>> result_array = np.linspace(0, 20, num=2, endpoint=True)
>>> result_array
array([ 0., 20.])

Here, we make two points between 0 and 20, so they're naturally 20 steps apart. You can also the endpoint to False and np.linspace()will behave more likenp.arange()` in that it doesn't include the final element:

>>> result_array = np.linspace(0, 20, num=5, endpoint=False)
>>> result_array
array([ 0.,  4.,  8., 12., 16.])

np.arange() vs built-in range()

The Python's built-in range() function and np.arange() share a lot of similarities but have slight differences. In the following sections, we're going to highlight some of the similarities and differences between them.

Parameters and Returns

The main similarities are that they both have a start, stop, and step. Additionally, they are both start inclusive, and stop exclusive, with a default step of 1.

However:

  1. Can handle multiple data types including floats and complex numbers
  2. returns a ndarray
  3. The array is fully created in memory
  1. Can handle only integers
  2. Returns a range object
  3. Generates numbers on demand

Efficiency and Speed

There are some speed and efficiency differences between np.arange() and the built-in range() function. The range function generates the numbers on demand and doesn't create them in-memory, upfront.

This helps speed the process up if you know you'll break somewhere in that range: For example:

for i in range(100000000):
    if i == some_number:
        break

This will consume less memory since not all numbers are created in advance. This also makes ndarrays slower to initially construct.

However, if you still need the whole range of numbers in-memory, np.arange() is significantly faster than range() when the full range of numbers comes into play, after they've been constructed.

For instance, if we just iterate through them, the time it takes to create the arrays makes np.arange() perform slower due to the higher upfront cost:

$ python -m timeit "for i in range(100000): pass"
200 loops, best of 5: 1.13 msec per loop

$ python -m timeit "import numpy as np" "for i in np.arange(100000): pass"
100 loops, best of 5: 3.83 msec per loop

Conclusion

This guide aims to help you understand how the np.arange() function works and how to generate sequences of numbers.

Here's a quick recap of what we just covered.

  1. np.arange() has 4 parameters:
    • start is a number (integer or real) from which the array starts from. It is optional.
    • stop is a number (integer or real) which the array ends at and is not included in it.
    • step is a number that sets the spacing between the consecutive values in the array. It is optional and is 1 by default.
    • dtype is the type of output for array elements. It is None by default.
  2. You can use multiple dtypes with arange including ints, floats, and complex numbers.
  3. You can generate reversed ranges by having the larger number as the start, the smaller number as the stop, and the step as a negative number.
  4. np.linspace() is similar to np.arange() in generating a range of numbers but differs in including the ability to include the endpoint and generating a number of samples instead of steps, which are computed based on the number of samples.
  5. np.arange() is more efficient than range when you need the whole array created. However, the range is better if you know you'll break somewhere when looping.

September 02, 2021 08:30 AM UTC


Python Bytes

#248 while True: stand up, sit down

<p><strong>Watch the live stream:</strong></p> <a href='/sitelet?url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DeIEGTZnsyCg' style='font-weight: bold;'>Watch on YouTube</a><br> <br> <p><strong>About the show</strong></p> <p>Sponsored by <strong>us:</strong></p> <ul> <li>Check out the <a href="/sitelet?url=https%3A%2F%2Ftraining.talkpython.fm%2Fcourses%2Fall"><strong>courses over at Talk Python</strong></a></li> <li>And <a href="/sitelet?url=https%3A%2F%2Fpythontest.com%2Fpytest-book%2F"><strong>Brian’s book too</strong></a>!</li> </ul> <p>Special guest: <strong>Paul Everitt</strong></p> <p><strong>Brain #1:</strong> <a href="/sitelet?url=https%3A%2F%2Fthreeofwands.com%2Fwhy-i-use-attrs-instead-of-pydantic%2F"><strong>Why I use attrs instead of pydantic</strong></a></p> <ul> <li><strong>Tin Tvrtković,</strong> <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2Ftintvrtkovic">@tintvrtkovic</a></li> <li>attrs vs dataclasses <ul> <li>Since dataclasses are a strict subset of attrs functionality. Recommend using attrs in most cases over dataclasses</li> <li>attrs is faster, has more features, releases more frequently, offers over a wider range of Python versions.</li> </ul></li> <li>attrs vs Pydantic <ul> <li>attrs is a library for generating the boring parts of writing classes; <ul> <li>Pydantic is that but also</li> <li>a complex validation library.</li> <li>a structuring/unstructuring library, ex converting to json and back</li> </ul></li> <li>attrs has opt-in validation that you have more control over</li> <li>cattrs can be used for structuring/unstructuring</li> <li>converters are opt-in for attrs, built into Pydantic, and can be wrong. <ul> <li>example using Pendulum that Pydantic mishandles</li> </ul></li> </ul></li> <li>Summary <ul> <li>attrs + cattrs + validators where necessary, converters where necessary</li> <li>will be faster</li> <li>you’ll have more control</li> <li>Kind of a “small, sharp, specialized tools” vs “swiss army knife” comparison.</li> </ul></li> </ul> <p><strong>Michael #2:</strong> <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2Fwhereismyjetpac%2Fstatus%2F1430694757320347648"><strong>mclfy</strong></a></p> <ul> <li>via __dann__</li> <li>Mcfly is an incredible Ctrl+r replacement</li> <li>McFly replaces your default <code>ctrl-r</code> shell history search with an intelligent search engine that takes into account your working directory and the context of recently executed commands. </li> <li>McFly's suggestions are prioritized in real time with a small neural network.</li> <li>Features <ul> <li>Rebinds <code>ctrl-r</code> to bring up a full-screen reverse history search prioritized with a small neural network.</li> <li>Augments your shell history to track command exit status, timestamp, and execution directory in a SQLite database.</li> <li>Maintains your normal shell history file as well so that you can stop using McFly whenever you want.</li> <li>Includes a simple action to scrub any history item from the McFly database and your shell history files.</li> <li>Designed to be extensible for other shells in the future.</li> <li>Written in Rust, so it's fast and safe.</li> </ul></li> </ul> <p><strong>Paul #3: Textual and</strong> <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2Fwillmcgugan%2Fstatus%2F1426267903733768193"><strong>boilerplate removal</strong></a></p> <ul> <li>In the race to make Textual the most talked-about package in Python Bytes history…</li> <li>I’d like to zoom in on a Twitter discussion he had about removing boilerplate</li> <li>I have traditionally been opposed to the convention-over-configuration approach that most successful Python projects have taken</li> <li>I dislike magic variable and file names, prefer explicit is better than implicit, actual <em>symbols</em></li> <li>Lately, because of…tooling</li> <li>But Will’s approach to “boilerplate removal” is compelling, as it remains mypy friendly</li> <li>Still, I find it flawed…code meant to be read 2 years from now…that stuff that is implied-away, worries me</li> <li>Will is great at working-in-the-open, being a gentle, encouraging public figure</li> </ul> <p><strong>Brian #4:</strong> <a href="/sitelet?url=https%3A%2F%2Fgithub.com%2FErotemic%2Fxdoctest"><strong>xdoctest</strong></a> </p> <ul> <li>“The <code>xdoctest</code> package is a re-write of Python's builtin <code>doctest</code> module. It replaces the old regex-based parser with a new abstract-syntax-tree based parser (using Python's <code>ast</code> module). The goal is to make doctests easier to write, simpler to configure, and encourage the pattern of test driven development.”</li> <li>“The main enhancements <code>xdoctest</code> offers over <code>doctest</code> are: <ol> <li>All lines in the doctest can now be prefixed with <code>&gt;&gt;&gt;</code>. Old-style doctests with <code>...</code> are still valid.</li> <li>Additionally, the multi-line strings don't require any prefix (but its ok if they do have either prefix).</li> <li>Tests are executed in blocks, rather than line-by-line, thus comment-based directives (e.g. <code># doctest: +SKIP</code>) are now applied to an entire block, rather than just a single line.</li> <li>Tests without a "want" statement will ignore any stdout / final evaluated value. This makes it easy to use simple assert statements to perform checks in code that might write to stdout.</li> <li>If your test has a "want" statement and ends with both a value and stdout, both are checked, and the test will pass if either matches.</li> <li>Output from multiple sequential print statements can now be checked by a single "got" statement. (new in 0.4.0).”</li> </ol></li> <li>Features I love <ul> <li>“The new got/want tester is very permissive by default; it ignores differences in whitespace”</li> <li>You can make doctest normalize whitespace, but why should you have to?</li> </ul></li> </ul> <p><strong>Michael #5:</strong> <a href="/sitelet?url=https%3A%2F%2Fmedium.com%2F%40davidkongfilm%2Fhow-i-hacked-my-standing-desk-with-a-raspberry-pi-a50ed14c7f6f"><strong>Automate the standing desk with python</strong></a></p> <ul> <li>via Joe Riedley, by David Kong</li> <li>“When I first started using it, I was very excited, but I quickly found myself sitting all day, in spite of the fancy desk.”</li> <li>I took off a few screws and … voila! A row of pins neatly exposed right in front.</li> <li>The pins in my control box, when connected correctly, simulate the pressing of the buttons on the front of the box.</li> <li><a href="/sitelet?url=https%3A%2F%2Fwww.raspberrypi.org%2Fproducts%2Fraspberry-pi-zero%2F"><strong>Raspberry Pi Zero</strong></a>, the simplest, most basic version. It doesn’t have all the bells and whistles, but it does everything I needed for this simple project, and it’s just $5(!).</li> <li>And the code</li> </ul> <pre><code> from gpiozero import LED # The LED library allows easy pin control from time import sleep import randomrelay = LED(17) # I connected the relay to pin 17 and groundwhile True: relay.on() sleep(1) relay.off() sleep(random.randint(45, 60) * 60) </code></pre> <p><strong>Paul #6:</strong> <a href="/sitelet?url=https%3A%2F%2Fcookiecutter-hypermodern-python.readthedocs.io%2Fen%2F2021.4.15%2F"><strong>Hypermodern Python Cookiecutter</strong></a></p> <ul> <li>I’ve been noodling with some code the last two years about bringing frontend DX to Python web dev</li> <li>Learning and talking more than adoption</li> <li>Running a modern Python project is a LOT of housekeeping</li> <li>Hypermodern Python Cookiecutter from Claudio Jolowicz teleported me to a state of the art I was looking for</li> <li>Poetry, Nox, GHA, pre-commit, flake8, PyPI uploads from CI, release drafter, Black, prettier, pytest, mypy, Sphinx and friends, GitHub labeler</li> <li>It’s NOT AT ALL just a cookiecutter</li> <li>The best part…it’s an enormously-detailed user guide, some blog posts with the “why”, it’s actively maintained</li> <li>The PR workflow is really well explained and wired up</li> <li>This could be…a course, a webinar</li> <li>Thanks Claudio</li> </ul> <p><strong>Extras</strong></p> <p>Michael:</p> <ul> <li><a href="/sitelet?url=https%3A%2F%2Fwww.surveymonkey.com%2Fr%2Fsecure-your-supply-chain"><strong>ActiveState's 2021 Software Supply Chain Security Survey</strong></a></li> <li><a href="/sitelet?url=https%3A%2F%2Fpythoninsider.blogspot.com%2F2021%2F08%2Fpython-397-and-3812-are-now-available.html"><strong>Python 3.9.7 and 3.8.12 are now available</strong></a></li> <li>From Shlomi Lanton, on your #2 Brian talked about having a history of all files to find the ones that were updated last, so I created <a href="/sitelet?url=https%3A%2F%2Fgithub.com%2FshlomiLan%2Fgrampa"><strong>granpa</strong></a> </li> <li>Also: <a href="/sitelet?url=https%3A%2F%2Fgithub.com%2Fnp-8%2Fwakepy"><strong>wakepy</strong></a> now works correctly on macOS</li> </ul> <p><strong>Joke:</strong> <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2Fismonkeyuser%2Fstatus%2F1430413027950612481%2F"><strong>Meaning</strong></a></p>

September 02, 2021 08:00 AM UTC


Read the Docs

Read the Docs newsletter - September 2021

Welcome to the latest edition of our monthly newsletter, where we share the most relevant updates around Read the Docs, offer a summary of new features we shipped during the previous month, and share what we’ll be focusing on in the near future.

Company highlights

New features

Thanks to our external contributors Mozi, Maksudul Haque, Stefano Costa, and Christian Clauss.

You can always see the latest changes to our platforms in our Read the Docs Changelog.

Upcoming features

  • Ana will continue researching tools for conducting automated testing on our Sphinx theme, review pull requests corresponding to the 1.1 milestone, and make some small styling improvements to our documentation.
  • Anthony will work with Ana on testing our Sphinx theme, in addition to resuming work on our new user interface and doing some financial updates.
  • Eric will keep pushing updates on our Commercial landing page and continue working on our sales processes. He’s also working to continue building EthicalAds as well.
  • Juan Luis will continue working on our Read the Docs tutorial, improving our onboarding experience, and put our email marketing and on-site notifications to work.
  • Manuel will finish the work on our new Docker images and build process, and keep designing our upcoming GitHub Application.
  • Santos will implement a new Slack integration, work with Manuel on the GitHub Application, and build a user interface for audit tracking.

Possible issues

We have been making changes to how we store cookies to make our site more secure. This has caused some minor issues in certain web browsers or to users that were embedding private documentation pages inside an iframe.


Considering using Read the Docs for your next Sphinx or MkDocs project? Check out our documentation to get started!

September 02, 2021 12:00 AM UTC

September 01, 2021


Python for Beginners

Binary Search Tree in Python

You can use different data structures such as a python dictionary, a list, a tuple, or a set in programs. But these data structures are not sufficient for implementing hierarchical structures in the programs. In this article, we will study about binary search tree data structure and will implement them in python for better understanding. 

What is a Binary Tree?

A binary tree is a tree data structure in which each node can have a maximum of 2 children.  It means that each node in a binary tree can have either one, or two or no children. Each node in a binary tree contains data and references to its children. Both the children are named as left child and the right child according to their position. The structure of a node in a binary tree is shown in the following figure.

Binary tree nodeNode of a Binary Tree

We can implement a binary tree node in python as follows.

class BinaryTreeNode:
  def __init__(self, data):
    self.data = data
    self.leftChild = None
    self.rightChild=None

What is a Binary Search Tree?

A binary search tree is a binary tree data structure with the following properties.

Following is an example of a binary search tree that satisfies all the properties discussed above.

Binary search tree in PythonBinary search tree

Now we will implement some of the basic operations on a binary search tree.

How to Insert an Element in a Binary Search Tree?

We will use the properties of binary search trees to insert elements into it. If we want to insert an element at a specific node, three conditions may arise.

  1. The current node can be an empty node i.e. None. In this case, we will create a new node with the element to be inserted and will assign the new node to the current node.
  2. The element to be inserted can be greater than the element at the current node. In this case, we will insert the new element in the right subtree of the current node as the right subtree of any node contains all the elements greater than the current node.
  3. The element to be inserted can be less than the element at the current node. In this case, we will insert the new element in the left subtree of the current node as the left subtree of any node contains all the elements lesser than the current node.

To insert an element, we will start from the root node and will insert the element into the binary search tree according to the above defined rules. The algorithm to insert elements in a binary search tree is implemented as in Python as follows.

class BinaryTreeNode:
    def __init__(self, data):
        self.data = data
        self.leftChild = None
        self.rightChild = None


def insert(root, newValue):
    # if binary search tree is empty, create a new node and declare it as root
    if root is None:
        root = BinaryTreeNode(newValue)
        return root
    # if newValue is less than value of data in root, add it to left subtree and proceed recursively
    if newValue < root.data:
        root.leftChild = insert(root.leftChild, newValue)
    else:
        # if newValue is greater than value of data in root, add it to right subtree and proceed recursively
        root.rightChild = insert(root.rightChild, newValue)
    return root


root = insert(None, 50)
insert(root, 20)
insert(root, 53)
insert(root, 11)
insert(root, 22)
insert(root, 52)
insert(root, 78)
node1 = root
node2 = node1.leftChild
node3 = node1.rightChild
node4 = node2.leftChild
node5 = node2.rightChild
node6 = node3.leftChild
node7 = node3.rightChild
print("Root Node is:")
print(node1.data)

print("left child of the node is:")
print(node1.leftChild.data)

print("right child of the node is:")
print(node1.rightChild.data)

print("Node is:")
print(node2.data)

print("left child of the node is:")
print(node2.leftChild.data)

print("right child of the node is:")
print(node2.rightChild.data)

print("Node is:")
print(node3.data)

print("left child of the node is:")
print(node3.leftChild.data)

print("right child of the node is:")
print(node3.rightChild.data)

print("Node is:")
print(node4.data)

print("left child of the node is:")
print(node4.leftChild)

print("right child of the node is:")
print(node4.rightChild)

print("Node is:")
print(node5.data)

print("left child of the node is:")
print(node5.leftChild)

print("right child of the node is:")
print(node5.rightChild)

print("Node is:")
print(node6.data)

print("left child of the node is:")
print(node6.leftChild)

print("right child of the node is:")
print(node6.rightChild)

print("Node is:")
print(node7.data)

print("left child of the node is:")
print(node7.leftChild)

print("right child of the node is:")
print(node7.rightChild)

Output:

Root Node is:
50
left child of the node is:
20
right child of the node is:
53
Node is:
20
left child of the node is:
11
right child of the node is:
22
Node is:
53
left child of the node is:
52
right child of the node is:
78
Node is:
11
left child of the node is:
None
right child of the node is:
None
Node is:
22
left child of the node is:
None
right child of the node is:
None
Node is:
52
left child of the node is:
None
right child of the node is:
None
Node is:
78
left child of the node is:
None
right child of the node is:
None

How to search an element in a Binary search Tree?

As you know that a binary search tree cannot have duplicate elements, we can search any element in a binary search tree using the following rules that are based on the properties of the binary search trees. We will start from the root and follow these properties

  1. If the current node is empty, we will say that the element is not present in the binary search tree.
  2. If the element in the current node is greater than the element to be searched, we will search the element in its left subtree as the left subtree of any node contains all the elements lesser than the current node.
  3. If the element in the current node is less than the element to be searched, we will search the element in its right subtree as the right subtree of any node contains all the elements greater  than the current node.
  4. If the element at the current node is equal to the element to be searched, we will return True.

The algorithm to search an element in a binary search tree based on the above properties is implemented in the following program. 

class BinaryTreeNode:
    def __init__(self, data):
        self.data = data
        self.leftChild = None
        self.rightChild = None


def insert(root, newValue):
    # if binary search tree is empty, create a new node and declare it as root
    if root is None:
        root = BinaryTreeNode(newValue)
        return root
    # if newValue is less than value of data in root, add it to left subtree and proceed recursively
    if newValue < root.data:
        root.leftChild = insert(root.leftChild, newValue)
    else:
        # if newValue is greater than value of data in root, add it to right subtree and proceed recursively
        root.rightChild = insert(root.rightChild, newValue)
    return root


def search(root, value):
    # node is empty
    if root is None:
        return False
    # if element is equal to the element to be searched
    elif root.data == value:
        return True
    # element to be searched is less than the current node
    elif root.data > value:
        return search(root.leftChild, value)
    # element to be searched is greater than the current node
    else:
        return search(root.rightChild, value)


root = insert(None, 50)
insert(root, 20)
insert(root, 53)
insert(root, 11)
insert(root, 22)
insert(root, 52)
insert(root, 78)
print("53 is present in the binary tree:", search(root, 53))
print("100 is present in the binary tree:", search(root, 100))

Output:

53 is present in the binary tree: True
100 is present in the binary tree: False

Conclusion

In this article, we have discussed binary search trees and their properties. We have also implemented the algorithms to insert elements into a binary search tree and to search elements in a binary search tree in Python. To learn more about data structures in Python, you can read this article on Linked list in python.

The post Binary Search Tree in Python appeared first on PythonForBeginners.com.

September 01, 2021 12:48 PM UTC


Mike Driscoll

Unit Conversion with Python and the Pint Package

Do you need to work measurements often? What about converting from one unit of measurement to another? There is a Python package called Pint that makes working with quantities easy to do. Pint allows you do arithmetic operations between a numerical value and a quantity as well. You can see the many different unit types included with Pint on their GitHub project.

Let’s get started by learning how to install Pint!

Installation

You can install Pint using pip like this:

python3 -m pip install pint

If you are a conda user, then you would want to use this command instead:

conda install -c conda-forge pint

Now that you have Pint installed, you are ready to learn how to use it!

Getting Started with Pint

One of the coolest features of Pint is that you can use it to convert from one unit type to another. For example, you might want to convert from some Imperial unit to a Metric unit.

A popular use case would be to convert from miles to kilometers. Open up your Python REPL (or IDLE) and try out the following code:

>>> from pint import UnitRegistry
>>> ureg = UnitRegistry()
>>> distance = 5 * ureg.mile
>>> distance
<Quantity(5, 'mile')>
>>> distance.to("kilometer")
<Quantity(8.04672, 'kilometer')>

Here you create a Quantity object named distance. You set its value to 5 miles. Then to convert it to kilometers, you call the distance’s to() method and pass in the new quantity name that you want. The result is that 5 miles is converted to 8.04672 kilometers.

You can convert the quantity to different unit types within the same system too. For example, you could convert kilometers to centimeters, if you wanted to:

>>> from pint import UnitRegistry 
>>> ureg = UnitRegistry()
>>> distance_in_km = 5 * ureg.kilometer
>>> distance_in_km
<Quantity(5, 'kilometer')>
>>> distance_in_cm = distance_in_km.to("centimeter")
>>> distance_in_cm
<Quantity(500000.0, 'centimeter')>

Pint Parses Strings

One of Pint’s cool features is that you can specify quantities using strings. That means you can do stuff like this:

>>> my_quantity = ureg.Quantity
>>> my_quantity(2.54, 'centimeter')
<Quantity(2.54, 'centimeter')>

Or you can simplify it, even more, to simply:

>>> my_quantity = ureg.Quantity
>>> my_quantity('2.54in')
<Quantity(2.54, 'inch')>

Now that you have played around with converting between different unit types, you are ready to learn about string formatting with Pint.

Using String Formatting with Pint

Pint supports formatting using Python’s .format() and by using f-strings. Here is an example from the Pint tutorial:

>>> ureg = ureg.Quantity
>>> accel = 1.3 * ureg['meter/second**2']
>>> print(f'The str is {accel}')
The str is 1.3 meter / second ** 2

When the f-string is evaluated, the Quantity object is converted into a more human-readable format.

Pint goes farther than that by extending Python’s formatting capabilities. Here is an example of their custom “pretty print”:

>>> ureg = ureg.Quantity
>>> accel = 1.3 * ureg['meter/second**2']
>>> # Pretty print
>>> 'The pretty representation is {:P}'.format(accel)
'The pretty representation is 1.3 meter/second²'

Pint also supports custom printing of LaTeX and HTML for Jupyter Notebooks.

Wrapping Up

Pint is a really nice Python package. While this tutorial doesn’t cover it, Pint allows you to set your locale so that the unit names match your language. If you regularly work with quantities that need to be converted between unit types (like cm to mm or inches to centimeters), this may be just what you need to make your coding life easier.

More Neat Python Packages

Want to learn about other neat 3rd party Python packages? Check out the following articles:

The post Unit Conversion with Python and the Pint Package appeared first on Mouse Vs Python.

September 01, 2021 12:30 PM UTC


PyBites

Facial Recognition with Python

Identifying faces

I was asked by Bob to write a guest article for the PyBites blog, so whilst this isn’t my first blog article, it is my first ever guest blog article of which I’m immensely proud and very pleased to have written for Pybites.

In this article, I will detail how I used the face_recognition and Pillow modules to extract and then identify faces from a bunch of photographs. I want to credit Brad Traversy as this idea was originally taken from a Traversy Media video I had bookmarked a while ago and recently rewatched (link below this article). Using the same functionality I tweaked the scripts to allow input for a directory of photos and eventually, it will be incorporated into my second PDM Django application.

They are very rough and ready scripts for a proof of concept which I demonstrated to Bob during one of our weekly code check-in calls.

I created two scripts, one to extract faces from a bunch of photos and store them as jpeg files in a specified directory. The second script takes a bunch of known faces (I guess these are the control set) and compares them to photographs to identify the faces in random photos. On the whole, it’s very accurate and fast at identifying the known faces.

Extract the faces

The first script was fairly simple to implement. We have a directory of images in which we want to extract all images of faces, we’ll call this directory unknown. The script essentially scans through each photo, identifies the face and stores this face image as a new jpeg in another directory we’ll call ‘extract’. This file is created with the title of the original image, and the face location within the image.

import face_recognition
from PIL import Image
import os

unknown_faces = os.listdir("../frec/unknown/")

for image in unknown_faces:
    image_of_people = face_recognition.load_image_file(f"../frec/unknown/{image}")
    unknown_face_locations = face_recognition.face_locations(image_of_people)

    for face_location in unknown_face_locations:
        top, right, bottom, left = face_location

        face_image = image_of_people[top:bottom, left:right]
        pil_image = Image.fromarray(face_image)
        pil_image.save(f"../frec/extract/{image}_{top}.jpg")

So from an image like this:

IMG1 bob and julian small 1

You would get two separate images in the extract directory like these:

IMG2 bob and julian small.jpeg 32 IMG3 bob and julian small.jpeg 72

Having tested this on my own photos, the module is so good it identified a face in a poster within one of my photographs which you will be able to see if you download the code from the repository.

Facial Recognition Time

The second script written does all the clever stuff insofar as it will take the directory of images with unknown faces, compare them to the known faces images and then ‘draw’ on the original image a square around the face with the name of the individual identified, if indeed it identifies a face, otherwise it will draw a square around the face with ‘Unknown Person’ in place of the name.

import face_recognition
import os
from PIL import Image, ImageDraw

unknown_faces = os.listdir("../frec/unknown/")
Bob_image = face_recognition.load_image_file("../frec/known/Bob.jpeg")
Bob_encoding = face_recognition.face_encodings(Bob_image)[0]
Julian_image = face_recognition.load_image_file("../frec/known/Julian.jpeg")
Julian_encoding = face_recognition.face_encodings(Julian_image)[0]

known_face_encodings = [
    Bob_encoding,
    Julian_encoding,
]

known_face_names = [
    "Bob",
    "Julian",
]

for ukface in unknown_faces:
    ukimage = face_recognition.load_image_file(f"../frec/unknown/{ukface}")
    ukface_locations = face_recognition.face_locations(ukimage)
    ukface_encodings = face_recognition.face_encodings(ukimage, ukface_locations)

    # Convert to PIL format
    pil_image = Image.fromarray(ukimage)

    # Set up drawing on image
    draw = ImageDraw.Draw(pil_image)
    for (top, right, bottom, left), ukface_encoding in zip(
        ukface_locations, ukface_encodings
    ):
        matches = face_recognition.compare_faces(known_face_encodings, ukface_encoding)

        name = "Unknown Person"

        if True in matches:
            first_match_index = matches.index(True)
            name = known_face_names[first_match_index]

        # Draw Box
        draw.rectangle(
            ((left - 10, top - 10), (right + 10, bottom + 10)), outline=(227, 236, 75)
        )

        # Draw Label
        text_width, text_height = draw.textsize(name)
        draw.rectangle(
            ((left - 10, bottom - text_height + 2), (right + 10, bottom + 10)),
            fill=(227, 236, 75),
            outline=(227, 236, 75),
        )
        draw.text((left, bottom - text_height + 5), name, fill=(0, 0, 0, 0))

    del draw
    pil_image.save(f"../frec/identified/{ukface}_scanned.jpg")
    #pil_image.show()

I’ve played around with the functionality of the ‘draw.rectangle’ function to try and capture as much of the face as possible inside of the square as originally the face was obscured by the square.

So again from the given image with unknown faces:

IMG1 bob and julian small

Using two different images of Bob and Julian to match against:

IMG4 Bob IMG5 Julian

You would end up with:

IMG6 bob and julian small.jpeg scanned

The accuracy, as I said previously, is amazing and you can tweak the values of the face_recognition package to widen or narrow the match. It is also very fast at scanning a directory of images and displaying the results on the screen or saving them to another directory, which would be very easy with a small tweak to the script.

I believe you are also able to take advantage of GPU processing but I was unable to test this in my current environment.

I have expanded the scripts in the repo to contain more examples of known / unkown faces so feel free to download and have a play.

For those of you that are interested in the source material which I modified, please take a look at Brad’s video on YouTube, without that I wouldn’t have found out about the face_recognition module. Hopefully you can adapt this code to suit your own purpose as I did.

I also had to tweak my scripts to increase the chance of a match otherwise Todd Dewey from Ice Road Truckers was being identified as me (I’m not sure who should be more flattered by that!). That’s what the model and num_jitters options are for in the scripts found in the repo.

The docs for the face_recognition module can be found here Face Recognition Docs

The docs for Pillow can be found here Pillow Docs

My repository for these scripts can be found here GitHub

And Brad Traversy’s repo to accompany his video are here GitHub

September 01, 2021 07:57 AM UTC


Tryton News

Newsletter for September 2021

We hope that everybody had a nice Summer and enjoyed their holidays. The Tryton team continued working on the ERP and we are back with a resume of the latest improvements.

Changes for the User

We added a frame around the image widget. This makes the widget cleaner when empty.

More party identifiers and tax identifiers have been added for Austria, Ukraine and Vietnam.

The rule keywords from statement lines are now stored in such way that they can be used for future matching. This adds a form of learning behavior to the statement rules engine.

We added a new wizard to split accounting lines. This is useful to reschedule payable or receivable lines by applying a new maturity date to each new line. The wizard can also be used on dunning and invoice lines to reschedule them.

It is now possible to set accounts for taxes of type “None”. This is useful for taxes that are entered manually on the invoice because the account will be filled in automatically.

New Modules

The Stock Package Shipping Sendcloud Module allows package labels to be generated for shipments made by any of Sendcloud’s supported carriers.

The Account Budget Module provides the ability to set budgets for accounts over a defined period of time. These budgets can then be used to track the total amount from relevant transactions against the budgeted amount.

The Analytic Budget Module provides the ability to set budgets for analytic accounts over a defined period of time. These budgets can then be used to track the total amount from relevant transactions against the budgeted amount.

The Product Image Module adds images to each product and variant.

Changes for the System Administrator

We improved the error management in the script used to import postal codes.

Changes for the Developer

We moved and renamed the cost_warehouse from the product_cost_warehouse module to the warehouse in the stock module. By doing this it can now be used by any module that depends on the stock module.

The complete locale definition for the user’s language is now sent to the clients.

Proteus now also fills in the wizard actions attribute when the result is an empty list.

The number widgets’ width attribute is now also used as its default display width.

The currency module defines a new Monetary field. This is derived from the Numeric field by adding a currency attribute which contains the name of the field which stores the currency. The desktop clients render these fields using the monetary format and with the currency symbol by default.

It is now also possible to use a string as the digits value on number fields (instead of the usual pair of integers). The string must contain the name of a Many2One field which points to a Model that inherits from DigitsMixin and that provides a get_digits method.
This allowed the removal of all the Function fields that provided the currency and unit digits.
Another benefit is that clients cache the value for each DigitsMixin record for 1 day by default, so this change also reduces the load on the server.

We reduced the number of times we save the cost values when doing multiple moves.

We no longer try to read records that were deleted after being instantiated from a browse list.

The digits argument to the format_number method of Report is now optional. If it is not specified, or is set to None, it will display all the significant digits.

1 post - 1 participant

Read full topic

September 01, 2021 07:00 AM UTC


Django Weblog

Django bugfix release: 3.2.7

Today we've issued the 3.2.7 bugfix release.

The release package and checksums are available from our downloads page, as well as from the Python Package Index. The PGP key ID used for this release is Mariusz Felisiak: 2EF56372BA48CD1B.

September 01, 2021 05:54 AM UTC

August 31, 2021


TestDriven.io

Django REST Framework and Elasticsearch

This tutorial looks at how to integrate Django REST Framework with Elasticsearch.

August 31, 2021 10:28 PM UTC


Sandipan Dey

Probabilistic Deep Learning with Tensorflow

In this blog, we shall discuss on how to implement probabilistic deep learning models using Tensorflow. The problems to be discussed in this blog appeared in the exercises / projects in the coursera course “Probabilistic Deep Learning“, by Imperial College, London, as a part of TensorFlow 2 for Deep Learning Specialization. The problem statements / … Continue reading Probabilistic Deep Learning with Tensorflow

August 31, 2021 09:40 PM UTC


PyCoder’s Weekly

Issue #488 (Aug. 31, 2021)

#488 – AUGUST 31, 2021
View in Browser »

The PyCoder’s Weekly Logo


Python Ranks #1 in IEEE “Top Programming Languages”

“Python dominates as the de facto platform for new technologies” and “Learn Python. That’s the biggest takeaway we can give you from its continued dominance of IEEE Spectrum’s annual interactive rankings of the top programming languages. You don’t have to become a dyed-in-the-wool Pythonista, but learning the language well enough to use one of the vast number of libraries written for it is probably worth your time.”
IEEE.ORG

skybison: Instagram’s Experimental Performance Oriented Greenfield Implementation of Python

“Skybison is experimental performance-oriented greenfield implementation of Python 3.8. It contains a number of performance optimizations, including: small objects; a moving GC; hidden classes; bytecode inline caching; type-specialized bytecode; an experimental template JIT.”
GITHUB.COM/FACEBOOKEXPERIMENTAL

Start Your Free Scout APM Trial, No CC Needed, and Receive a $5 Donation to the OSS of Your Choice

alt

Scout APM is leading-edge application performance and error monitoring designed to help devs find and fix observability issues before the customer ever sees them. You can connect your error reporting and APM data on one platform, with Scout’s new error monitoring feature add-on →
SCOUT APM sponsor

How to Use Optional Arguments When Defining Python Functions

In this tutorial, you’ll learn about optional arguments in Python and how to define functions with default values. You’ll also learn how to create functions that accept any number of arguments using args and kwargs.
REAL PYTHON

Python Project-Local Virtualenv Management

On UNIX-like operating systems you can have the Python equivalent of node_modules today, for every Python version, without changing your workflows.
HYNEK SCHLAWACK

Humble Software Bundle: Python Superpowers 2021

Pick up the awesome programming potential of Python with software like Mastering PyCharm (2021 Edition) & Object-Oriented Programming (OOP) in Python. Pay what you want & support charity!
HUMBLEBUNDLE.COM

Join the PyCon US 2022 Team!

The PyCon US organizers are looking for motivated volunteers who want to contribute their time and knowledge to make this year’s conference a great success.
PYCON US

Python 3.9.7 and 3.8.12 Are Now Available

More info in the full changelog.
CPYTHON DEV BLOG

Discussions

math.sqrt vs numpy.sqrt vs x ** 0.5 Performance Discussion

Andrej Karpathy (Director of AI at Tesla) shares an interesting performance observation on this Twitter thread that turns into a tale about accurate benchmarking. Calculating math.sqrt(1337.0) appears to be 10x faster than numpy.sqrt(1337.0). Python’s built-in square root (x ** 0.5) appears to be even faster. However, most of the performance differences seem to come from the benchmark setup, as Ishan Bhatt explains in this writeup.
TWITTER.COM/KARPATHY

Python Jobs

Data Engineer - Python & PostgreSQL (Newport Beach, CA, USA)

Research Affiliates

Sr. Backend Developer (Amsterdam, Netherlands)

GUTS Tickets

Backend Software Engineer (Anywhere)

Catalpa International

More Python Jobs >>>

Articles & Tutorials

A Python Data Scientist’s Guide to the Apple Silicon Transition

A break down of what Apple Silicon means for Python users today, especially those doing scientific computing and data science: what works, what doesn’t, and where this might be going.
STANLEY SEIBERT

Write an SQL Query Builder in 150 Lines of Python

“This is the fourth article in a series about writing my own SQL query builder. Today, we’ll rewrite it from scratch, explore API design, learn when to be lazy, and look at worse and better ways of doing things – all in 150 lines of Python!”
ANDGRAVITY.COM

Rev APIs Solve All of Your Speech-to-Text Needs

alt

Rev.ai is the most sophisticated automatic speech recognition in the world. Our speech-to-text APIs are more accurate, easier to use, and have less bias than competitors like Google, Amazon, and Microsoft. Try Rev.ai free for five hours right now →
REV.AI sponsor

Splitting Datasets With scikit-learn and train_test_split()

Learn why it’s important to split your dataset in supervised machine learning and how to do that with train_test_split() from the widely used scikit-learn package.
REAL PYTHON video

Building With CircuitPython & Constraints of Python for Microcontrollers

Can you make a version of Python that fits within the memory constraints of a microcontroller and have it still feel like Python? That is the intention behind CircuitPython. This week on the show, Scott Shawcroft, who is the project lead for CircuitPython.
REAL PYTHON podcast

Parsing in Python: Tools and Libraries You Can Use

“We present and compare all possible alternatives you can use to parse languages in Python. From libraries to parser generators, we present all options.”
GABRIELE TOMASSETTI

Low-Level Cache API in Django

Caching in Django can be implemented on different levels (or parts of the site). This article looks at how to use the low-level cache API in Django.
J-O ERIKSSON

SonarLint Free and Open Source IDE Extension for Python Devs

Working in VS Code, PyCharm, Visual Studio, or Eclipse? SonarLint helps you find & fix Code Quality and Code Security issues in your Python codebase!
SONARSOURCE sponsor

Python Behind the Scenes: How Async/Await Works in Python

“The async/await pattern can be explained in a simple manner if you start from the ground up. And that’s what we’re going to do today.”
VICTOR SKVORTSOV

Using libsqlite3 Directly From Python With ctypes

How to use ctypes to run SQLite queries without using the built-in sqlite3 Python package, and without compiling anything.
GITHUB.COM/MICHALC

Handling Environment Variables in Python

BOB BELDERBOS

Simulating a Direct Digital Frequency Synthesizer in Python

NASH REILLEY

Projects & Code

kmk_firmware: Clackety Keyboards Powered by Python

GITHUB.COM/KMKFW

ordered: Entropy-Controlled Contexts in Python

GITHUB.COM/HYPERC-AI

gopygo: Pure Python Go Parser, AST and Unparser Library

GITHUB.COM/UP9INC

cattrs: Complex Custom Class Converters for attrs

GITHUB.COM/TINCHE

msl-apollo-entry-guidance: Python Implementation of the Apollo Entry Guidance Algorithm Used by NASA’s MSL Spacecraft

GITHUB.COM/THOMASANTONY

pylectronics: Reproduce Digital Electronics in Python

GITHUB.COM/FGARCI03

sqlmodel: SQL Databases in Python, Designed for Simplicity, Compatibility, and Robustness

GITHUB.COM/TIANGOLO

Events

Real Python Office Hours (Virtual)

September 1, 2021
REALPYTHON.COM

PyConline AU 2021

September 10 to September 13, 2021
PYCON.ORG.AU


Happy Pythoning!
This was PyCoder’s Weekly Issue #488.
View in Browser »

alt

[ Subscribe to 🐍 PyCoder’s Weekly 💌 – Get the best Python news, articles, and tutorials delivered to your inbox once a week >> Click here to learn more ]

August 31, 2021 07:30 PM UTC


Will McGugan

Pretty printing JSON with Rich

If you work with JSON regularly (90% of Python developers I suspect) you might appreciate the print_json function just landed in Rich v10.9.0

If you call this function with a string, Rich will decode the string, reformat it, and print it to the console with nice syntax highlighting. Here's an example:

from rich import print_json
print_json('{"foo": [false, true, null]}')

Here's the output:

© 2021 Will McGugan

Calling print_json with a string will decode the JSON and pretty print it.

Note that the atomic values false, true, and null have their own color. I find this helpful when scanning a JSON blob.

If you call print_json with a data keyword argument it will encode that data and pretty print it in the same way.

data = {
    "foo": [
        3.1427,
        (
            "Paul Atreides",
            "Vladimir Harkonnen",
            "Thufir Hawat",
        ),
    ],
    "atomic": (False, True, None),
}
from rich import print_json
print_json(data=data)

Here's the output:

© 2021 Will McGugan

Calling the print_json function with a data keyword argument.

Note that Rich will remove color if you pipe the output of your script to another program, so you can safely add syntax highlighting to your CLI tools.

You can also pretty print JSON files from the command line with the following:

python -m rich.json data.json

Here's an example of the output:

© 2021 Will McGugan

Pretty printing a JSON file from the command line

This is admittedly a small addition to Rich but I'm already finding it helpful.

Follow @willmcgugan on Twitter for Rich and Textual updates.

August 31, 2021 05:43 PM UTC


Quansight Labs Blog

CZI EOSS4 Grants at Quansight Labs

Here, at Quansight Labs, our goal is to work on sustaining the future of Open Source. We make sure we can live up to that goal by spending a significant amount of time working on impactful and critical infrastructure and projects within the Scientific Ecosystem.

As such, our goals align with those of the Chan Zuckerberg Initiative and, in particular, the Essential Open Source Software for Science (EOSS) program that supports tools essential to biomedical research via funds for software maintenance, growth, development, and community engagement.

CZI’s Essential Open Source Software for Science program supports software maintenance, growth, development, and community engagement for open source tools critical to science. And the Chan Zuckerberg Initiative was founded in 2015 to help solve some of society’s toughest challenges — from eradicating disease and improving education, to addressing the needs of our local communities. Their mission is to build a more inclusive, just, and healthy future for everyone.

Today, we are thrilled to announce that the team at Quansight Labs has been awarded five EOSS Cycle 4 grants to work on several projects within the PyData ecosystem. This post will introduce the successful grantees and their objectives for these two-year long grants.

Read more… (5 min remaining to read)

August 31, 2021 05:01 PM UTC


Python for Beginners

How to Extract a Date from a .txt File in Python

In this tutorial, we’ll examine the different ways you can extract a date from a .txt file using Python programming. Python is a versatile language—as you’ll discover—and there are many solutions for this problem.

First, we’ll look at using regular expression patterns to search text files for dates that fit a predefined format. We’ll learn about using the re library and creating our own regular expression searches.

We’ll also examine datetime objects and use them to convert strings into data models. Lastly, we’ll see how the datefinder module simplifies the process of searching a text file for dates that haven’t been formatted, like we might find in natural language content.

Extract a Date from a .txt File using Regular Expression

Dates are written in many different formats. Sometimes people write month/day/year. Other dates might include times of the day, or the day of the week (Wednesday July 8, 2021 8:00PM).

How dates are formatted is a factor to consider before we go about extracting them from text files. 

For instance, if a date follows the month/date/year format, we can find it using a regular expression pattern. With regular expression, or regex for short, we can search a text by matching a string to a predefined pattern. 

The beauty of regular expression is that we can use special characters to create powerful search patterns. For instance, we can craft a pattern that will find all the formatted dates in the following body of text.

minutes.txt
10/14/2021 – Meeting with the client.
07/01/2021 – Discussed marketing strategies.
12/23/2021 – Interviewed a new team lead.
01/28/2018 – Changed domain providers.
06/11/2017 – Discussed moving to a new office.

Example: Finding formatted dates with regex

import re

# open the text file and read the data
file = open("minutes.txt",'r')

text = file.read()
# match a regex pattern for formatted dates
matches = re.findall(r'(\d+/\d+/\d+)',text)

print(matches)

Output

[’10/14/2021′, ’07/01/2021′, ’12/23/2021′, ’01/28/2018′, ’06/11/2017′]

The regex pattern here uses special characters to define the strings we want to extract from the text file. The characters d and + tell regex we’re looking for multiple digits within the text.

We can also use regex to find dates that are formatted in different ways. By altering our regex pattern, we can find dates that use either a forward slash (\) or a dash (–) as the separator.

This works because regex allows for optional characters in the search pattern. We can specify that either character—a forward slash or dash—is an acceptable match.

apple2.txt
The first Apple II was sold on 07-10-1977. The last of the Apple II
models were discontinued on 10/15/1994.

Example: Matching dates with a regex pattern

import re

# open a text file
f = open("apple2.txt", 'r')

# extract the file's content
content = f.read()

# a regular expression pattern to match dates
pattern = "\d{2}[/-]\d{2}[/-]\d{4}"

# find all the strings that match the pattern
dates = re.findall(pattern, content)

for date in dates:
    print(date)

f.close()

Output

07-10-1977
10/15/1994

Examining the full extent of regex’s potential is beyond the scope of this tutorial. Try experimenting with some of the following special characters to learn more about using regular expression patterns to extract a date—or other information—from a .txt file.

Special Characters in Regex

Extract a Datetime Object from a .txt File

In Python we can use the datetime library for manipulating dates and working with time. The datetime library comes pre-packed with Python, so there’s no need to install it.

By using datetime objects, we have more control over string data read from text files. For example, we can use a datetime object to get a copy of the current date and time of our computer.

import datetime

now = datetime.datetime.now()
print(now)

Output

2021-07-04 20:15:49.185380

In the following example, we’ll extract a date from a company .txt file that mentions a scheduled meeting. Our employer needs us to scan a group of such documents for dates. Later, we plan to add the information we gather to a SQLite database.

We’ll begin by defining a regex pattern that will match our date format. Once a match is found, we’ll use it to create a datetime object from the string data.

schedule.txt

schedule.txt
The project begins next month. Denise has scheduled a meeting in the conference room at the Embassy Suits on 10-7-2021.

Example: Creating datetime objects from file data

import re
from datetime import datetime

# open the data file
file = open("schedule.txt", 'r')
text = file.read()

match = re.search(r'\d+-\d+-\d{4}', text)
# create a new datetime object from the regex match
date = datetime.strptime(match.group(), '%d-%m-%Y').date()
print(f"The date of the meeting is on {date}.")
file.close()

Output

The date of the meeting is on 2021-07-10.

Extracting Dates from a Text File with the Datefinder Module

The Python datefinder module can locate dates in a body of text. Using the find_dates() method, it’s possible to search text data for many different types of dates. Datefinder will return any dates it finds in the form of a datetime object.

Unlike the other packages we’ve discussed in this guide, Python does not come with datefinder. The easiest way to install the datefinder module is to use pip from the command prompt.

pip install datefinder

With datefinder installed, we’re ready to open files and extract data. For this example, we’ll use a text document that introduces a fictitious company project. Using datefinder, we’ll extract each date from the .txt file, and print their datimeobject counterparts.

Feel free to save the file locally and follow along.

project_timeline.txt
PROJECT PEPPER

All team members must read the project summary by
January 4th, 2021.

The first meeting of PROJECT PEPPER begins on 01/15/2021

at 9:00am. Please find the time to read the following links by then.
created on 08-12-2021 at 05:00 PM

This project file has dates in many formats. Dates are written using dashes and forward slashes. What’s worse, the month January is written out. How can we find all these dates with Python?

Example: Using datefinder to extract dates from file data

import datefinder

# open the project schedule
file = open("project_timeline.txt",'r')

content = file.read()

# datefinder will find the dates for us
matches = list(datefinder.find_dates(content))

if len(matches) > 0:
    for date in matches:
        print(date)
else:
    print("Found no dates.")

file.close()

Output
2021-01-04 00:00:00
2021-01-15 09:00:00
2021-08-12 17:00:00

As you can see from the output, datefinder is able to find a variety of date formats in the text. Not only is the package capable of recognizing the names of months, but it also recognizes the time of day if it’s included in the text.

In another example, we’ll use the datefinder package to extract a date from a .txt file that includes the dates for a popular singer’s upcoming tour.

tour_dates.txt
Saturday July 25, 2021 at 07:00 PM     Inglewood, CA
Sunday July 26, 2021 at 7 PM     Inglewood, CA
09/30/2021 7:30PM  Foxbourough, MA

Example: Extract a tour date and times from a .txt file with datefinder

import datefinder

# open the project schedule
file = open("tour_dates.txt",'r')

content = file.read()

# datefinder will find the dates for us
matches = list(datefinder.find_dates(content))

if len(matches) > 0:
    print("TOUR DATES AND TIMES")
    print("--------------------")
    for date in matches:
        # use f string to format the text
        print(f"{date.date()}     {date.time()}")
else:
    print("Found no dates.")
file.close()

Output

TOUR DATES AND TIMES
——————–
2021-07-25     19:00:00
2021-07-26     19:00:00
2021-09-30     19:30:00

As you can see from the examples, datefinder can find many different types of dates and times. This is useful if the dates you’re looking for don’t have a certain format, as will often be the case in natural language data.

Summary

In this post, we’ve covered several methods of how to extract a date or time from a .txt file. We’ve seen the power of regular expression to find matches in string data, and we’ve seen how to convert that data into a Python datetime object.

Finally, if the dates in your text files don’t have a specified format—as will be the case in most files with natural language content—try the datefinder module. With this Python package, it’s possible to extract dates and times from a text file that aren’t conveniently formatted ahead of time.

Related Posts

If you enjoyed this tutorial and are eager to learn more about Python—and we sincerely hope you are—follow these links for more great guides from Python for Beginners.

The post How to Extract a Date from a .txt File in Python appeared first on PythonForBeginners.com.

August 31, 2021 02:52 PM UTC


Real Python

Splitting Datasets With scikit-learn and train_test_split()

One of the key aspects of supervised machine learning is model evaluation and validation. When you evaluate the predictive performance of your model, it’s essential that the process be unbiased. Using train_test_split() from the data science library scikit-learn, you can split your dataset into subsets that minimize the potential for bias in your evaluation and validation process.

In this course, you’ll learn:

In addition, you’ll get information on related tools from sklearn.model_selection.


[ Improve Your Python With 🐍 Python Tricks 💌 – Get a short & sweet Python Trick delivered to your inbox every couple of days. >> Click here to learn more and see examples ]

August 31, 2021 02:00 PM UTC


Stack Abuse

Random Projection: Theory and Implementation in Python with Scikit-Learn

Introduction

This guide is an in-depth introduction to an unsupervised dimensionality reduction technique called Random Projections. A Random Projection can be used to reduce the complexity and size of data, making the data easier to process and visualize. It is also a preprocessing technique for input preparation to a classifier or a regressor.

Random Projection is typically applied to highly-dimensional data, where other techniques such as Principal Component Analysis (PCA) can't do the data justice.

In this guide, we'll delve into the details of Johnson-Lindenstrauss lemma, which lays the mathematical foundation of Random Projections. We'll also show how to perform Random Projection using Python's Scikit-Learn library, and use it to transform input data to a lower-dimensional space.

Theory is theory, and practice is practice. As a practical illustration, we'll load the Reuters Corpus Volume I Dataset, and apply Gaussian Random Projection and Sparse Random Projection to it.

What is a Random Projection of a Dataset?

Put simply:

Random Projection is a method of dimensionality reduction and data visualization that simplifies the complexity of high-dimensional datasets.

The method generates a new dataset by taking the projection of each data point along a randomly chosen set of directions. The projection of a single data point onto a vector is mathematically equivalent to taking the dot product of the point with the vector.

random projections illustration

Given a data matrix \(X\) of dimensions \(mxn\) and a \(dxn\) matrix \(R\) whose columns are the vectors representing random directions, the Random Projection of \(X\) is given by \(X_p\).

X p = X R

Each vector representing a random direction, has dimensionality \(n\), which is the same as all data points of \(X\). If we take \(d\) random directions, then we end up with a \(d\) dimensional transformed dataset. For the purpose of this tutorial, we'll fix a few notations:

The idea of Random Projections is very similar to Principal Component Analysis (PCA), fundementally. However, in PCA, the projection matrix is computed via eigenvectors, which can be computationally expensive for large matrices.

When performing Random Projection, the vectors are chosen randomly making it very efficient. The name "projection" may be a little misleading as the vectors are chosen randomly, the transformed points are mathematically not true projections but close to being true projections.

The data with reduced dimensions is easier to work with. Not only can it be visualized but it can also be used in the pre-processing stage to reduce the size of the original data.

A Simple Example

Just to understand how the transformation works, let's take the following simple example.

Suppose our input matrix \(X\) is given by:

X = [ 1 3 2 0 0 1 2 1 1 3 0 0 ]

And the projection matrix is given by:

R = 1 2 [ 1 − 1 1 1 1 − 1 1 1 ]

The projection of X onto R is:

X p = X R = 1 2 [ 6 0 4 0 4 2 ]

We started with three points in a four-dimensional space, and with clever matrix operations ended up with three transformed points in a two-dimensional space.

Note, some important attributes of the projection matrix \(R\). Each column is a unit matrix, i.e., the norm of each column is one. Also, the dot product of all columns taken pairwise (in this case only column 1 and column 2) is zero, indicating that both column vectors are orthogonal to each other.

This makes the matrix, an Orthonormal Matrix. However, in case of the Random Projection technique, the projection matrix does not have to be a true orthonormal matrix when very high-dimensional data is involved.

The success of Random Projection is based on an awesome mathematical finding known as Johnson-Lindenstrauss lemma, which is explained in detail in the following section!

The Johnson-Lindenstrauss lemma

The Johnson-Lindenstrauss lemma is the mathematical basis for Random Projection:

The Johnson-Lindenstrauss lemma states that if the data points lie in a very high-dimensional space, then projecting such points on simple random directions preserves their pairwise distances.

Preserving pairwise distances implies that the pairwise distances between points in the original space are the same or almost the same as the pairwise distance in the projected lower-dimensional space.

Thus, the structure of data and clusters within data are maintained in a lower-dimensional space, while the complexity and size of data are reduced substantially.

In this guide, we refer to the difference in the actual and projected pairwise distances as the "distortion" in data, which is introduced due to its projection in a new space.

Johnson-Lindenstrauss lemma also provides a "safe" measure of the number of dimensions to project the data points onto so that the error/distortion lies within a certain range, so finding the target number of dimensions is made easy.

Mathematically, given a pair of points \((x_1,x_2)\) and their corresponding projections \((x_1',x_2')\) defines an eps-embedding:

$$
(1 - \epsilon) |x_1 - x_2|^2 < |x_1' - x_2'|^2 < (1 + \epsilon) |x_1 - x_2|^2
$$

The Johnson-Lindenstrauss lemma specifies the minimum dimensions of the lower-dimensional space so that the above eps-embedding is maintained.

Determining the Random Directions of the Projection Matrix

Two well-known methods for determining the projection matrix are:

R i j = 3 { + 1  with probability  1 6 0  with probability  2 3 − 1  with probability  1 6

The method above is equivalent to choosing the numbers from {+k,0,-k} based on the outcome of the roll of a dice. If the dice score is 1, then choose +k. If the dice score is in the range [2,5], choose 0, and choose -k for a dice score of 6.

A more general method uses a density parameter to choose the Random Projection matrix. Setting \(s=\frac{1}{\text{density}}\), the elements of the Random Projection matrix are chosen as:

R i j = { + s d  with probability  1 2 s 0  with probability  1 − 1 s − s d  with probability  1 2 s

The general recommendation is to set the density parameter to \(\frac{1}{\sqrt n}\).

As mentioned earlier, for both the Gaussian and sparse methods, the projection matrix is not a true orthonormal matrix. However, it has been shown that in high dimensional spaces, the randomly chosen matrix using either of the above two methods is close to an orthonormal matrix.

Random Projection Using Scikit-Learn

The Scikit-Learn library provides us with the random_projection module, that has three important classes/modules:

We'll demonstrate all the above three in the sections below, but first let's import the classes and functions we'll be using:

from sklearn.random_projection import SparseRandomProjection, johnson_lindenstrauss_min_dim
from sklearn.random_projection import GaussianRandomProjection
import numpy as np
from matplotlib import pyplot as plt
import sklearn.datasets as dt
from sklearn.metrics.pairwise import euclidean_distances

Determining the Minimum Number of Dimensions Via Johnson Lindenstrauss lemma

The johnson_lindenstrauss_min_dim() function determines the minimum number of dimensions d, which the input data can be mapped to when given the number of examples m, and the eps or \(\epsilon\) parameter.

The code below experiments with a different number of samples to determine the minimum size of the lower-dimensional space, which maintains a certain "safe" distortion of data.

Additionally, it plots log(d) against different values of eps for different sample sizes m.

An important thing to note is that the Johnson Lindenstrauss lemma determines the size of the lower-dimensional space \(d\) only based on the number of example points \(m\) in the input data. The number of attributes or features \(n\) of the original data is irrelevant:

eps = np.arange(0.001, 0.999, 0.01)
colors = ['b', 'g', 'm', 'c']
m = [1e1, 1e3, 1e7, 1e10]
for i in range(4):
    min_dim = johnson_lindenstrauss_min_dim(n_samples=m[i], eps=eps)
    label = 'Total samples = ' + str(m[i])
    plt.plot(eps, np.log10(min_dim), c=colors[i], label=label)
    
plt.xlabel('eps')
plt.ylabel('log$_{10}$(d)')
plt.axhline(y=3.5, color='k', linestyle=':')
plt.legend()
plt.show()

how to determine size of lower dimensional space for random projections

From the plot above, we can see that for small values of eps, d is quite large but decreases as eps approaches one. The dimensionality is below 3500 (the dotted black line) for mid to large values of eps.

This shows that applying Random Projections only makes sense to high-dimensional data, of the order of thousands of features. In such cases, a high reduction in dimensionality can be achieved.

Random Projections are, therefore, very successful for text or image data, which involve a large number of input features, where Principal Component Analysis would

Data Transformation

Python includes the implementation of both Gaussian Random Projections and Sparse Random Projections in its sklearn library via the two classes GaussianRandomProjection and SparseRandomProjection respectively. Some important attributes for these classes are (the list is not exhaustive):

Like other dimensionality reduction classes of sklearn, both these classes include the standard fit() and fit_transform() methods. A notable set of attributes, which come in handy are:

Random Projection with GaussianRandomProjection

Let's start off with the GaussianRandomProjection class. The values of the projection matrix are plotted as a histogram and we can see that they follow a Gaussian distribution with mean zero. The size of the data matrix is reduced from 5000 to 3947:

X_rand = np.random.RandomState(0).rand(100, 5000)
proj_gauss = GaussianRandomProjection(random_state=0)
X_transformed = proj_gauss.fit_transform(X_rand)

# Print the size of the transformed data
print('Shape of transformed data: ' + str(X_transformed.shape))

# Generate a histogram of the elements of the transformation matrix
plt.hist(proj_gauss.components_.flatten())
plt.title('Histogram of the flattened transformation matrix')
plt.show()

This code results in:

Shape of transformed data: (100, 3947)

gaussian random projection scikit learn

Random Projection with SparseRandomProjection

The code below demonstrates how data transformation can be made using a Sparse Random Projection. The entire transformation matrix is composed of three distinct values, whose frequency plot is also shown below.

Note that the transformation matrix is a SciPy sparse csr_matrix. The following code accesses the non-zero values of the csr_matrix and stores them in p. Next, it uses p to get the counts of the elements of the sparse projection matrix:

proj_sparse = SparseRandomProjection(random_state=0)
X_transformed = proj_sparse.fit_transform(X_rand)

# Print the size of the transformed data
print('Shape of transformed data: ' + str(X_transformed.shape))

# Get data of the transformation matrix and store in p. 
# p consists of only 2 non-zero distinct values, i.e., pos and neg
# pos and neg are determined below
p = proj_sparse.components_.data
total_elements = proj_sparse.components_.shape[0] *\
                  proj_sparse.components_.shape[1]
pos = p[p>0][0]
neg = p[p<0][0]
print('Shape of transformation matrix: '+ str(proj_sparse.components_.shape))
counts = (sum(p==neg), total_elements - len(p), sum(p==pos))
# Histogram of the elements of the transformation matrix
plt.bar([neg, 0, pos], counts, width=0.1)
plt.xticks([neg, 0, pos])
plt.suptitle('Histogram of flattened transformation matrix, ' + 
             'density = ' +
             '{:.2f}'.format(proj_sparse.density_))
plt.show()

This results in:

Shape of transformed data: (100, 3947)
Shape of transformation matrix: (3947, 5000)

sparse random projections scikit learn

The histogram is in agreement with the method of generating a sparse Random Projection matrix as discussed in the previous section. The zero is selected with probability (1-1/100 = 0.99), hence around 99% of values of this matrix are zero. Utilizing the data structures and routines for sparse matrices makes this transformation method very fast and efficient on large datasets.

Practical Random Projections With the Reuters Corpus Volume 1 Dataset

This section illustrates Random Projections on the Reuters Corpus Volume I Dataset. The dataset is freely accessible online, though for our purposes, it's easiest to looad via Scikit-Learn.

The sklearn.datasets module contains a fetch_rcv1() function that downloads and imports the dataset.

Note: The dataset may take a few minutes to download, if you've never imported it beforehand through this method. Since there's no progress bar, it may appear as if the script is hanging without progressing further. Give it a bit of time, when you run it initially.

The RCV1 dataset is a multilabel dataset, i.e., each data point can belong to multiple classes at the same time, and consists of 103 classes. Each data point has a dimensionality of a whopping 47,236, making it an ideal case for applying fast and cheap Random Projections.

To demonstrate the effectiveness of Random Projections, and to keep things simple, we'll select 500 data points that belong to at least one of the first three classes. The fetch_rcv1() function retrieves the dataset and returns an object with data and targets, both of which are sparse CSR matrices from SciPy.

Let's fetch the Reuters Corpus and prepare it for data transformation:

total_points = 500
# Fetch the dataset
dat = dt.fetch_rcv1()
# Select the sparse matrix's non-zero targets
target_nz = dat.target.nonzero()
# Select only indices of target_nz for data points that belong to 
# either of class 1,2,3
ind_class_123 = np.asarray(np.where((target_nz[1]==0) |\
                                    (target_nz[1]==1) |\
                                    (target_nz[1] == 2))).flatten()
# Choose only 500 indices randomly
np.random.seed(0)
ind_class_123 = np.random.choice(ind_class_123, total_points, 
                                 replace=False)

# Retreive the row indices of data matrix and target matrix
row_ind = target_nz[0][ind_class_123]
X = dat.data[row_ind,:]
y = np.array(dat.target[row_ind,0:3].todense())

After data preparation, we need a function that creates a visualization of the projected data. To have an idea of the quality of transformation, we can compute the following three matrices:

The abs_diff_dist matrix is a good indicator of the quality of the data transformation. Close to zero or small values in this matrix indicate low distortion and a good transformation. We can directly display an image of this matrix or generate a histogram of its values to visually assess the transformation. We can also compute the average of all the values of this matrix to get a single quantitative measure for comparison.

The function create_visualization() creates three plots. The first graph is a scatter plot of projected points along the first two random directions. The second plot is an image of the absolute difference matrix and the third is the histogram of the values of the absolute difference matrix:

def create_visualization(X_transform, y, abs_diff):
    fig,ax = plt.subplots(nrows=1, ncols=3, figsize=(20,7))

    plt.subplot(131)
    plt.scatter(X_transform[y[:,0]==1,0], X_transform[y[:,0]==1,1], c='r', alpha=0.4)
    plt.scatter(X_transform[y[:,1]==1,0], X_transform[y[:,1]==1,1], c='b', alpha=0.4)
    plt.scatter(X_transform[y[:,2]==1,0], X_transform[y[:,2]==1,1], c='g', alpha=0.4)
    plt.legend(['Class 1', 'Class 2', 'Class 3'])
    plt.title('Projected data along first two dimensions')

    plt.subplot(132)
    plt.imshow(abs_diff)
    plt.colorbar()
    plt.title('Visualization of absolute differences')

    plt.subplot(133)
    ax = plt.hist(abs_diff.flatten())
    plt.title('Histogram of absolute differences')

    fig.subplots_adjust(wspace=.3) 

Reuters Dataset: Gaussian Random Projection

Let's apply Gaussian Random Projection to the Reuters dataset. The code below runs a for loop for different eps values. If the minimum safe dimensions returned by johnson_lindenstrauss_min_dim is less than the actual data dimensions, then it calls the fit_transform() method of GaussianRandomProjection. The create_visualization() function is then called to create a visualization for that value of eps.

At every iteration, the code also stores the mean absolute difference and the percentage reduction in dimensionality achieved by Gaussian Random Projection:

reduction_dim_gauss = []
eps_arr_gauss = []
mean_abs_diff_gauss = []
for eps in np.arange(0.1, 0.999, 0.2):

    min_dim = johnson_lindenstrauss_min_dim(n_samples=total_points, eps=eps)
    if min_dim > X.shape[1]:
        continue
    gauss_proj = GaussianRandomProjection(random_state=0, eps=eps)
    X_transform = gauss_proj.fit_transform(X)
    dist_raw = euclidean_distances(X)
    dist_transform = euclidean_distances(X_transform)
    abs_diff_gauss = abs(dist_raw - dist_transform) 

    create_visualization(X_transform, y, abs_diff_gauss)
    plt.suptitle('eps = ' + '{:.2f}'.format(eps) + ', n_components = ' + str(X_transform.shape[1]))
    
    reduction_dim_gauss.append(100-X_transform.shape[1]/X.shape[1]*100)
    eps_arr_gauss.append(eps)
    mean_abs_diff_gauss.append(np.mean(abs_diff_gauss.flatten()))

RCV1 dataset gaussian random projections

RCV1 dataset gaussian random projections

RCV1 dataset gaussian random projections

RCV1 dataset gaussian random projections

RCV1 dataset gaussian random projections

The images of the absolute difference matrix and its corresponding histogram indicate that most of the values are close to zero. Hence, a large majority of the pair of points maintain their actual distance in the low dimensional space, retaining the original structure of data.

To assess the quality of transformation, let's plot the mean absolute difference against eps. Also, the higher the value of eps, the greater the dimensionality reduction. Let's also plot the percentage reduction vs. eps in a second sub-plot:

fig,ax = plt.subplots(nrows=1, ncols=2, figsize=(10,5))
plt.subplot(121)
plt.plot(eps_arr_gauss, mean_abs_diff_gauss, marker='o', c='g')
plt.xlabel('eps')
plt.ylabel('Mean absolute difference')

plt.subplot(122)
plt.plot(eps_arr_gauss, reduction_dim_gauss, marker = 'o', c='m')
plt.xlabel('eps')
plt.ylabel('Percentage reduction in dimensionality')

fig.subplots_adjust(wspace=.4) 
plt.suptitle('Assessing the Quality of Gaussian Random Projections')
plt.show()

RCV1 random projections reduction quality

We can see that using Gaussian Random Projection we can reduce the dimensionality of data to more than 99%! Though, this does come at the cost of a higher distortion of data.

Reuters Dataset: Sparse Random Projection

We can do a similar comparison with sparse Random Projection:

reduction_dim_sparse = []
eps_arr_sparse = []
mean_abs_diff_sparse = []
for eps in np.arange(0.1, 0.999, 0.2):

    min_dim = johnson_lindenstrauss_min_dim(n_samples=total_points, eps=eps)
    if min_dim > X.shape[1]:
        continue
    sparse_proj = SparseRandomProjection(random_state=0, eps=eps, dense_output=1)
    X_transform = sparse_proj.fit_transform(X)
    dist_raw = euclidean_distances(X)
    dist_transform = euclidean_distances(X_transform)
    abs_diff_sparse = abs(dist_raw - dist_transform) 

    create_visualization(X_transform, y, abs_diff_sparse)
    plt.suptitle('eps = ' + '{:.2f}'.format(eps) + ', n_components = ' + str(X_transform.shape[1]))
    
    reduction_dim_sparse.append(100-X_transform.shape[1]/X.shape[1]*100)
    eps_arr_sparse.append(eps)
    mean_abs_diff_sparse.append(np.mean(abs_diff_sparse.flatten()))

RCV1 dataset sparse random projections

RCV1 dataset sparse random projections

RCV1 dataset sparse random projections

RCV1 dataset sparse random projections

RCV1 dataset sparse random projections

In the case of Random Projection, the absolute difference matrix appears similar to the one of Gaussian projection. The projected data on the first two dimensions, however, has a more interesting pattern, with many points mapped on the coordinate axis.

Let's also plot the mean absolute difference and percentage reduction in dimensionality for various values of the eps parameter:

fig,ax = plt.subplots(nrows=1, ncols=2, figsize=(10,5))
plt.subplot(121)
plt.plot(eps_arr_sparse, mean_abs_diff_sparse, marker='o', c='g')
plt.xlabel('eps')
plt.ylabel('Mean absolute difference')

plt.subplot(122)
plt.plot(eps_arr_sparse, reduction_dim_sparse, marker = 'o', c='m')
plt.xlabel('eps')
plt.ylabel('Percentage reduction in dimensionality')

fig.subplots_adjust(wspace=.4) 
plt.suptitle('Assessing the Quality of Sparse Random Projections')
plt.show()

sparse random projections reduction quality

The trend of the two graphs is similar to that of a Gaussian Projection. However, the mean absolute difference for Gaussian Projection is lower than that of Random Projection.

Conclusions

In this guide, we discussed the details of two main types of Random Projections, i.e., Gaussian and sparse Random Projection.

We presented the details of the Johnson-Lindenstrauss lemma, the mathematical basis for these methods. We then showed how this method can be used to transform data using Python's sklearn library.

We also illustrated the two methods on a real-life Reuters Corpus Volume I Dataset.

We encourage the reader to try out this method in supervised classification or regression tasks at the pre-processing stage when dealing with very high-dimensional datasets.

August 31, 2021 10:30 AM UTC


Zero to Mastery

Python Monthly 💻🐍 August 2021

21st issue of Python Monthly! Read by 20,000+ Python developers every month. This monthly Python newsletter is focused on keeping you up to date with the industry and keeping your skills sharp, without wasting your valuable time.

August 31, 2021 10:00 AM UTC


Talk Python to Me

#332: Robust Python

Does it seem like your Python projects are getting bigger and bigger? Are you feeling the pain as your codebase expands and gets tougher to debug and maintain? Patrick Viafore is here to help us write more maintainable, longer-lived, and more enjoyable Python code.<br/> <br/> <strong>Links from the show</strong><br/> <br/> <div><b>Pat on Twitter</b>: <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2FPatViaforever" target="_blank" rel="noopener">@PatViaforever</a><br/> <b>Robust Python Book</b>: <a href="/sitelet?url=https%3A%2F%2Fwww.oreilly.com%2Flibrary%2Fview%2Frobust-python%2F9781098100650%2F" target="_blank" rel="noopener">oreilly.com</a><br/> <b>Typing in Python</b>: <a href="/sitelet?url=https%3A%2F%2Fdocs.python.org%2F3%2Flibrary%2Ftyping.html" target="_blank" rel="noopener">docs.python.org</a><br/> <b>mypy</b>: <a href="/sitelet?url=http%3A%2F%2Fmypy-lang.org%2F" target="_blank" rel="noopener">mypy-lang.org</a><br/> <b>SQLModel</b>: <a href="/sitelet?url=https%3A%2F%2Fsqlmodel.tiangolo.com%2F" target="_blank" rel="noopener">sqlmodel.tiangolo.com</a><br/> <b>CUPID principles @ relevant time</b>: <a href="/sitelet?url=https%3A%2F%2Fovercast.fm%2F%2BBYsRlGnE%2F19%3A06" target="_blank" rel="noopener">overcast.fm</a><br/> <b>Stevedore package</b>: <a href="/sitelet?url=https%3A%2F%2Fdocs.openstack.org%2Fstevedore%2Flatest%2F" target="_blank" rel="noopener">docs.openstack.org</a><br/> <b>Watch YouTube live stream edition</b>: <a href="/sitelet?url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DQU3JO4dwT-s" target="_blank" rel="noopener">youtube.com</a><br/> <b>Episode transcripts</b>: <a href="/sitelet?url=https%3A%2F%2Ftalkpython.fm%2Fepisodes%2Ftranscript%2F332%2Frobust-python" target="_blank" rel="noopener">talkpython.fm</a><br/> <br/> <b>Stay in touch with us</b><br/> <b>Subscribe on YouTube (for live streams)</b>: <a href="/sitelet?url=https%3A%2F%2Ftalkpython.fm%2Fyoutube" target="_blank" rel="noopener">youtube.com</a><br/> <b>Follow Talk Python on Twitter</b>: <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2Ftalkpython" target="_blank" rel="noopener">@talkpython</a><br/> <b>Follow Michael on Twitter</b>: <a href="/sitelet?url=https%3A%2F%2Ftwitter.com%2Fmkennedy" target="_blank" rel="noopener">@mkennedy</a><br/></div><br/> <strong>Sponsors</strong><br/> <a href='/sitelet?url=https%3A%2F%2Fclubhouse.io%2Ftalkpython'>Clubhouse</a><br> <a href='/sitelet?url=https%3A%2F%2Ftalkpython.fm%2Fmasterworks'>Masterworks.io</a><br> <a href='/sitelet?url=https%3A%2F%2Ftalkpython.fm%2Fassemblyai'>AssemblyAI</a><br> <a href='/sitelet?url=https%3A%2F%2Ftalkpython.fm%2Ftraining'>Talk Python Training</a>

August 31, 2021 08:00 AM UTC


PyBites

How to handle environment variables in Python

In this article I will share 3 libraries I often use to isolate my environment variables from production code.

Why is this important?

Separate config from code

As we can read in The Twelve-Factor App / III. Config:

Apps sometimes store config as constants in the code. This is a violation of twelve-factor, which requires strict separation of config from code.

https://12factor.net/config

Basically you want to be able to make config changes independently from code changes.

We also want to hide secret keys and API credentials! Notice that git is very persistent (PyCon talk: Oops, I committed my password to GitHub) so it’s important to get this right from the start.

First package: python-dotenv

These days I mostly use python-dotenv which makes this straightforward.

First install the library and add it to your requirements (or if you use Poetry it will automatically update your .toml file):

pip install python-dotenv

Secondly make an .env file with your environment variables in it.

It’s important that you ignore this file with git, otherwise you will end up committing sensitive data to your repo / project.

What I usually do is commit an empty .env-example (or .env-template) file so other developers know what they should set (see examples here and here).

So a new developer (or me checking out the repo on another machine) can do a cp .env-template .env and populate the variables. As the (checked out) .gitignore file contains .env, git won’t show it as a file to be staged for commit.

Then, to load in the variables from this file we use two lines of code:

from dotenv import load_dotenv

load_dotenv()

You can now access the environment variables using os.environ, for example:

BACKGROUND_IMG = os.environ["THUMB_BACKGROUND_IMAGE"]
FONT_FILE = os.environ["THUMB_FONT_TTF_FILE"]

To load the config without touching the environment, you can use dotenv_values(".env") which works the same as load_dotenv, except it doesn’t touch the environment, it just returns a dict with the values parsed from the .env file.

Check out the README for additional options.

Second package: python-decouple

Another library I have been using a lot with Django is python-decouple.

The process is pretty similar:

pip install python-decouple

Create an .env file with your config variables and “gitignore” it.

Then in your code you can use the config object. As per the example in the docs:

from decouple import config

SECRET_KEY = config('SECRET_KEY')
DEBUG = config('DEBUG', default=False, cast=bool)
EMAIL_HOST = config('EMAIL_HOST', default='localhost')
EMAIL_PORT = config('EMAIL_PORT', default=25, cast=int)

The casting and the ability to specify defaults are really convenient.

Another useful option is the Csv helper. For example having this in our .env file for our platform (a Django app):

ALLOWED_HOSTS=.localhost, .herokuapp.com

We can retrieve this variable in settings.py like this:

ALLOWED_HOSTS = config('ALLOWED_HOSTS', cast=Csv())

Third package: dj-database-url

And while we are here, there is one more package I want to show you: dj-database-url, which makes it easier to load in your database URL.

As per the docs:

The dj_database_url.config method returns a Django database connection dictionary, populated with all the data specified in your URL. There is also a conn_max_age argument to easily enable Django’s connection pool.

https://pypi.org/project/dj-database-url/

And here is how to use it:

import dj_database_url

DATABASES = {
    'default': dj_database_url.config(
        default=config('DATABASE_URL')
    )
}

Nice and clean!

This is what I mostly use, for more options, check out python-decouple‘s README here.


Python Tips

As a recap, here is the python-decouple code in a concise tip you can easily paste into your project:

# pip install python-decouple dj-database-url

from decouple import config, Csv
import dj_database_url

SECRET_KEY = config('SECRET_KEY')
DEBUG = config('DEBUG', default=False, cast=bool)
ALLOWED_HOSTS = config('ALLOWED_HOSTS', cast=Csv())

DATABASES = {
    'default': dj_database_url.config(
        default=config('DATABASE_URL')
    )
}

We love practical tips like these, to get our growing collection check out our book: PyBites Python Tips – 250 Bulletproof Python Tips That Will Instantly Make You A Better Developer

And with that we got a wrap. I hope this has been useful and will make it easier for you to separate config from code, which I wholeheartedly agree with The Twelve-Factor App, is important.

— Bob

August 31, 2021 07:50 AM UTC


Glyph Lefkowitz

Unproblematize

The essence of software engineering is solving problems.

The first impression of this insight will almost certainly be that it seems like a good thing. If you have a problem, then solving it is great!

But software engineers are more likely to have mental health problems1 than those who perform mechanical labor, and I think our problem-oriented world-view has something to do with that.

So, how could solving problems be a problem?


As an example, let’s consider the idea of a bug tracker.

For many years, in the field of software, any system used to track work has been commonly referred to as a “bug tracker”. In recent years, the labels have become more euphemistic and general, and we might now call them “issue trackers”. We have Sapir-Whorfed2 our way into the default assumption that any work that might need performing is a degenerate case of a problem.

We can contrast this with other fields. Any industry will need to track work that must be done. For example, in doing some light research for this post, I discovered that the relevant term of art in construction3 is typically “Project Management” or “Task Management” software. “Projects” and “Tasks” are no less hard work, but the terms do have a different valence than “Bugs” and “Issues”.

I don’t think we can start to fix this ... problem ... by attempting to change the terminology. Firstly, the domain inherently lends itself to this sort of language, which is why it emerged in the first place.

Secondly, Atlassian has desperately been trying to get everybody to call their bug tracker a “software development tool” where you write “stories” for years, and nobody does. It’s an issue tracker where you file bugs, and that’s what everyone calls it and describes what they do with it. Even they have to protest, perhaps a bit too much, that it’s “way more than a bug and issue tracker”4.


This pervasive orientation towards “problems” as the atom of work does extend to any knowledge work, and thereby to any “productivity system”. Any to-do list is, at its core, a list of problems. You wouldn’t put an item on the list if you were happy with the way the world was. Therefore every unfinished item in any to-do list is a little pebble of worry.

As of this writing, I have almost 1000 unfinished tasks on my personal to-do list.

This is to say nothing of any tasks I have to perform at work, not to mention the implicit א‎0 of additional unfinished tasks once one considers open source issue trackers for projects I work on.

It’s not really reasonable to opt out of this habit of problematizing everything. This monument to human folly that I’ve meticulously constructed out of the records of aspirations which exceed my capacity is, in fact, also an excellent prioritization tool. If you’re a good engineer, or even just good at making to-do lists, you’ll inevitably make huge lists of problems. On some level, this is what it means to set an intention to make the world — or at least your world — better.

On a different level though, this is how you set out to systematically give yourself anxiety, depression, or both. It’s clear from a wealth of neurological research that repeated experiences and thoughts change neural structures5. Thinking the same thought over and over literally re-wires your brain. Thinking the thought “here is another problem” over and over again forever is bound to cause some problems of its own.

The structure of to-do apps, bug trackers and the like is such that when an item is completed — when a problem is solved — it is subsequently removed from both physical view and our mind’s eye. What would be the point of simply lingering on a completed task? All the useful work is, after all, problems that haven’t been solved yet. Therefore the vast majority of our time is spent contemplating nothing but problems, prompting the continuous potentiation6 of neural pathways which lead to despair.


I don’t want to pretend that I have a cure for this self-inflicted ailment. I do, however, have a humble suggestion for one way to push back just a little bit against the relentless, unending tide of problems slowly eroding the shores of our souls: a positivity journal.

By “journal”, I do mean a private journal. Public expressions of positivity7 can help; indeed, some social and cultural support for expressing positivity is an important tool for maintaining a positive mind-set. However, it may not be the best starting point.

Unfortunately, any public expression becomes a discourse, and any discourse inevitably becomes a dialectic. Any expression of a view in public is seen by some as an invitation to express its opposite8. Therefore one either becomes invested in defending the boundaries of a positive community space — a psychically exhausting task in its own right — or one must constantly entertain the possibility that things are, in fact, bad, when one is trying to condition one’s brain to maintain the ability to recognize when things are actually good.

Thus my suggestion to write something for yourself, and only for yourself.

Personally, I use a template that I fill out every day, with four sections:

Although such a journal is private, it’s helpful to actually write out the answers, to focus on them, to force yourself to get really specific.

I hope this tool is useful to someone out there. It’s not going to solve any problems, but perhaps it will make the world seem just a little brighter.


  1. “Maintaining Mental health on Software Development Teams”, Lena Kozar and Vova Vovk, in InfoQ ↩

  2. Wikipedia page for “Linguistic Relativity” ↩

  3. “Construction Task and Project Tracking”, from Raptor Project Management Software ↩

  4. Jira Features List, Atlassian Software ↩

  5. “Culture Wires the Brain: A Cognitive Neuroscience Perspective”, Denise C. Park and Chih-Mao Huang, Perspect Psychol Sci. 2010 Jul 1; 5(4): 391–400. ↩

  6. Long-term potentiation and learning, J L Martinez Jr, B E Derrick ↩

  7. The #PositivePython hashtag on Twitter was a lovely experiment and despite my cautions here about public solutions to this problem, it’s generally pleasant to participate in. ↩

  8. As we well know. ↩

August 31, 2021 07:03 AM UTC


Python Piedmont Triad User Group

Lunch and learn series

PYPTUG Lunch and Learn


In order to help those starting out with python, we are starting a lunch and learn series. You can see upcoming lunch and learns on our meetup page:


https://www.meetup.com/PYthon-Piedmont-Triad-User-Group-PYPTUG/



August 31, 2021 05:41 AM UTC