Sitelet https://web.archive.org/web/20200827094200/https://planetpython.org/

skip to navigation
skip to content

Planet Python

Last update: August 27, 2020 07:47 AM UTC

August 27, 2020


Kushal Das

PrivChat with Tor: 2020-08-28

Tomorrow at 17:00UTC, Tor Project is hosting the next session of PrivChat, titled "The Good, the Bad, and the Ugly of Censorship Circumvention". You can watch it live on Youtube.

PrivChat tomorrow 17:00UTC on youtube

This 2nd edition of PrivChat is about the Good, the Bad and the Ugly that is happening in the front lines of censorship circumvention. Cory Doctorow will be the host for the evening, and the following people will be participating:

Don't miss your chance to listen to them. You can ask questions via the Youtube chat.

August 27, 2020 03:43 AM UTC


PSF GSoC students blogs

Weekly Check-in #13

<meta name="uuid" content="uuidCXied8VStJ11"><meta charset="utf-8">

What did I do this week?

I added test cases to check whether the orchestrator can run operations in parallel. I also started working on the input network. Currently, I have added all the input network and an orchestrator instance to the main node.

​

What's next?

Did I get stuck somewhere?

Yes. I got stuck in a deadlock. It was resolved with the help from mentors.

August 27, 2020 02:56 AM UTC


Sebastian Witowski

Find Item in a List

Find a number

If you want to find the first number that matches some criteria, what do you do? The easiest way is to write a loop that checks numbers one by one and returns when it finds the correct one.

Let’s say we want to get the first number divided by 42 and 43 (that’s 1806). If we don’t have a predefined set of elements (in this case, we want to check all the numbers starting from 1), we might use a “while loop”.

# find_item.py

def while_loop():
    item = 1
    # You don't need to use parentheses, but they improve readability
    while True:
        if (item % 42 == 0) and (item % 43 == 0):
            return item
        item += 1

It’s pretty straightforward:

Find a number in a list

If we have a list of items that we want to check, we will use a “for loop” instead. I know that the number I’m looking for is smaller than 10 000, so let’s use that as the upper limit:

# find_item.py

def for_loop():
    for item in range(1, 10000):
        if (item % 42 == 0) and (item % 43 == 0):
            return item

Let’s compare both solutions (benchmarks are done with Python 3.8 - I describe the whole setup in the Introduction article):

$ python -m timeit -s "from find_item import while_loop" "while_loop()"
2000 loops, best of 5: 134 usec per loop

$ python -m timeit -s "from find_item import for_loop" "for_loop()"
2000 loops, best of 5: 103 usec per loop

“While loop” is around 30% slower than the “for loop” (134/103≈1.301).

Loops are optimized to iterate over a collection of elements. Trying to manually do the iteration (for example, by referencing elements in a list through an index variable) will be a slower and often over-engineered solution.

Python 2 flashback

In Python 3, the range() function is lazy. It won't initialize an array of 10 000 elements, but it will generate them as needed. It doesn't matter if we say range(1, 10000) or range(1, 1000000) - there will be no difference in speed. But it was not the case in Python 2!

In Python 2, functions like range, filter, or zip were eager, so they would always create the whole collection when initialized. All those elements would be loaded to the memory, increasing the execution time of your code and its memory usage. To avoid this behavior, you had to use their lazy equivalents like xrange, ifilter, or izip.

Out of curiosity, let's see how slow is the for_loop() function if we run it with Python 2.7.18 (the latest and last version of Python 2):

$ pyenv shell 2.7.18
$ python -m timeit -s "from find_item import for_loop" "for_loop()"
10000 loops, best of 3: 151 usec per loop
That's almost 50% slower than running the same function in Python 3 (151/103≈1.4660). Updating Python version is one of the easiest performance wins you can get!

If you are wondering what's pyenv and how to use it to quickly switch Python versions, check out this section of my PyCon 2020 workshop on Python tools.

Let’s go back to our “while loop” vs. “for loop” comparison. Does it matter if the element we are looking for is at the beginning or at the end of the list?

def while_loop2():
    item = 1
    while True:
        if (item % 98 == 0) and (item % 99 == 0):
            return item
        item += 1

def for_loop2():
    for item in range(1, 10000):
        if (item % 98 == 0) and (item % 99 == 0):
            return item

This time, we are looking for number 9702, which is at the very end of our list. Let’s measure the performance:

$ python -m timeit -s "from find_item import while_loop2" "while_loop2()"
500 loops, best of 5: 710 usec per loop

$ python -m timeit -s "from find_item import for_loop2" "for_loop2()"
500 loops, best of 5: 578 usec per loop

There is almost no difference. “While loop” is around 22% slower this time (710/578≈1.223). I performed a few more tests (up to a number close to 100 000 000), and the difference was always similar (in the range of 20-30% slower).

Find a number in an infinite list

So far, the collection of items we wanted to iterate over was limited to the first 10 000 numbers. But what if we don’t know the upper limit? In this case, we can use the count function from the itertools library.

from itertools import count

def count_numbers():
    for item in count(1):
        if (item % 42 == 0) and (item % 43 == 0):
            return item

count(start=0, step=1) will start counting numbers from the start parameter, adding the step in each iteration. In my case, I need to change the start parameter to 1, so it works the same as the previous examples.

count works almost the same as the “while loop” that we made at the beginning. How about the speed?

$ python -m timeit -s "from find_item import count_numbers" "count_numbers()"
2000 loops, best of 5: 109 usec per loop

It’s almost the same as the “for loop” version. So count is a good replacement if you need an infinite counter.

What about a list comprehension?

A typical solution for iterating over a list of items is to use a list comprehension. But we want to exit the iteration as soon as we find our number, and that’s not easy to do with a list comprehension. It’s a great tool to go over the whole collection, but not in this case.

Let’s see how bad it is:

def list_comprehension():
    return [item for item in range(1, 10000) if (item % 42 == 0) and (item % 43 == 0)][0]
$ python -m timeit -s "from find_item import list_comprehension" "list_comprehension()"
500 loops, best of 5: 625 usec per loop

That’s really bad - it’s a few times slower than other solutions! It takes the same amount of time, no matter if we search for the first or last element. And we can’t use count here.

But using a list comprehension points us in the right direction - we need something that returns the first element it finds and then stops iterating. And that thing is a generator! We can use a generator expression to grab the first element matching our criteria.

Find item with a generator expression

def generator():
    return next(item for item in count(1) if (item % 42 == 0) and (item % 43 == 0))

The whole code looks very similar to a list comprehension, but we can actually use count. Generator expression will execute only enough code to return the next element. Each time you call next(), it will resume work in the same place where it stopped the last time, grab the next item, return it, and stop again.

$ python -m timeit -s "from find_item import generator" "generator()"
2000 loops, best of 5: 110 usec per loop

It takes almost the same amount of time as the best solution we have found so far. And I find this syntax much easier to read - as long as we don’t put too many ifs there!

Generators have the additional benefit of being able to “suspend” and “resume” counting. We can call next() multiple times, and each time we get the next element matching our criteria. If we want to get the first three numbers that can be divided by 42 and 43 - here is how easily we can do this with a generator expression:

def generator_3_items():
    gen = (item for item in count(1) if (item % 42 == 0) and (item % 43 == 0))
    return [next(gen), next(gen), next(gen)]

Compare it with the “for loop” version:

def for_loop_3_items():
    items = []
    for item in count(1):
        if (item % 42 == 0) and (item % 43 == 0):
            items.append(item)
            if len(items) == 3:
                return items

Let’s benchmark both versions:

$ python -m timeit -s "from find_item import for_loop_3_items" "for_loop_3_items()"
1000 loops, best of 5: 342 usec per loop

$ python -m timeit -s "from find_item import generator_3_items" "generator_3_items()"
1000 loops, best of 5: 349 usec per loop

Performance-wise, both functions are almost identical. So when would you use one over the other? “For loop” lets you write more complex code. You can’t put nested “if” statements or multiline code with side effects inside a generator expression. But if you only do simple filtering, generators can be much easier to read.

Be careful with nested ifs!

Nesting too many "if" statements makes code difficult to follow and reason about. And it's easy to make mistakes.

In the last example, if we don't nest the second if, it will be checked in each iteration. But we only need to check it when we modify the items list. It might be tempting to write the following code:

def for_loop_flat():
    items = []
    for item in count(1):
        if (item % 42 == 0) and (item % 43 == 0):
            items.append(item)
        if len(items) == 3:
            return items
This version is easier to follow, but it's also much slower!

$ python -m timeit -s "from find_item import for_loop_3_items" "for_loop_3_items()"
1000 loops, best of 5: 323 usec per loop

$ python -m timeit -s "from find_item import for_loop_flat" "for_loop_flat()"
500 loops, best of 5: 613 usec per loop
If you forget to nest ifs, your code will be 90% slower (613/323≈1.898).

Conclusions

Generator expression combined with next() is a great way to grab one or more elements based on specific criteria. It’s memory-efficient, fast, and easy to read - as long as you keep it simple. When the number of “if statements” in the generator expression grows, it becomes much harder to read (and write).

With complex filtering criteria or many ifs, “for loop” is a more suitable choice that doesn’t sacrifice the performance.

August 27, 2020 12:00 AM UTC

August 26, 2020


Mike Driscoll

Blackberry Released an Anti-Malware Tool Written in Python

In case you missed it earlier this month, Blackberry released a tool of theirs that they use for reverse engineering malware. That tool is called PE Tree and is open-source and written in Python.

Blackberry used the popular PyQt5 GUI toolkit to write that displays a tree view of portable executables, which makes it easier dump and reconstruct malware that is in memory.

The PR Tree tool works on Windows, Mac and Linux. It can run as a standalone application or as a plugin for IDAPython, which itself is a plugin for a disassembler.

This sounds like a really neat tool. If nothing else, it will be a good application to use for learning how to create a real-world GUI with Python.

The post Blackberry Released an Anti-Malware Tool Written in Python appeared first on The Mouse Vs. The Python.

August 26, 2020 07:05 PM UTC


Codementor

Top 5 Python Courses You can Join today for FREE

5 best free Python online courses for beginners to learn Python 3 in depth.

August 26, 2020 03:33 PM UTC


PyPy Development

PyPy 7.3.1 released

The PyPy team is proud to release the version 7.3.1 of PyPy, which includes two different interpreters:
  • PyPy2.7, which is an interpreter supporting the syntax and the features of Python 2.7 including the stdlib for CPython 2.7.13
  • PyPy3.6: which is an interpreter supporting the syntax and the features of Python 3.6, including the stdlib for CPython 3.6.9.
The interpreters are based on much the same codebase, thus the multiple release. This is a micro release, no APIs have changed since the 7.3.0 release in December, but read on to find out what is new.

Conda Forge now supports PyPy as a Python interpreter. The support right now is being built out. After this release, many more c-extension-based packages can be successfully built and uploaded. This is the result of a lot of hard work and good will on the part of the Conda Forge team. A big shout out to them for taking this on.

We have worked with the Python packaging group to support tooling around building third party packages for Python, so this release updates the pip and setuptools installed when executing pypy -mensurepip to pip>=20. This completes the work done to update the PEP 425 python tag from pp373 to mean “PyPy 7.3 running python3” to pp36 meaning “PyPy running Python 3.6” (the format is recommended in the PEP). The tag itself was changed in 7.3.0, but older pip versions build their own tag without querying PyPy. This means that wheels built for the previous tag format will not be discovered by pip from this version, so library authors should update their PyPy-specific wheels on PyPI.

Development of PyPy is transitioning to https://foss.heptapod.net/pypy/pypy. This move was covered more extensively in the blog post from last month.

The CFFI backend has been updated to version 14.0. We recommend using CFFI rather than c-extensions to interact with C, and using cppyy for performant wrapping of C++ code for Python. The cppyy backend has been enabled experimentally for win32, try it out and let use know how it works.

Enabling cppyy requires a more modern C compiler, so win32 is now built with MSVC160 (Visual Studio 2019). This is true for PyPy 3.6 as well as for 2.7.

We have improved warmup time by up to 20%, performance of io.StringIO to match if not be faster than CPython, and improved JIT code generation for generators (and generator expressions in particular) when passing them to functions like sum, map, and map that consume them. Performance of closures has also be improved in certain situations.

As always, this release fixed several issues and bugs raised by the growing community of PyPy users. We strongly recommend updating. Many of the fixes are the direct result of end-user bug reports, so please continue reporting issues as they crop up.
You can find links to download the v7.3.1 releases here:
http://pypy.org/download.html
We would like to thank our donors for the continued support of the PyPy project. If PyPy is not quite good enough for your needs, we are available for direct consulting work.

We would also like to thank our contributors and encourage new people to join the project. PyPy has many layers and we need help with all of them: PyPy and RPython documentation improvements, tweaking popular modules to run on PyPy, or general help with making RPython’s JIT even better. Since the previous release, we have accepted contributions from 13 new contributors, thanks for pitching in.

If you are a Python library maintainer and use c-extensions, please consider making a cffi / cppyy version of your library that would be performant on PyPy. In any case both cibuildwheel and the multibuild system support building wheels for PyPy wheels.

 

What is PyPy?

PyPy is a very compliant Python interpreter, almost a drop-in replacement for CPython 2.7, 3.6, and soon 3.7. It’s fast (PyPy and CPython 2.7.x performance comparison) due to its integrated tracing JIT compiler.

We also welcome developers of other dynamic languages to see what RPython can do for them.

This PyPy release supports:
  • x86 machines on most common operating systems (Linux 32/64 bits, Mac OS X 64 bits, Windows 32 bits, OpenBSD, FreeBSD)
  • big- and little-endian variants of PPC64 running Linux,
  • s390x running Linux
  • 64-bit ARM machines running Linux.

What else is new?

For more information about the 7.3.1 release, see the full changelog.

Please update, and continue to help us make PyPy better.

Cheers,
The PyPy team



The PyPy Team 

August 26, 2020 02:28 PM UTC


Real Python

Common Python Data Structures (Guide)

Data structures are the fundamental constructs around which you build your programs. Each data structure provides a particular way of organizing data so it can be accessed efficiently, depending on your use case. Python ships with an extensive set of data structures in its standard library.

However, Python’s naming convention doesn’t provide the same level of clarity that you’ll find in other languages. In Java, a list isn’t just a list—it’s either a LinkedList or an ArrayList. Not so in Python. Even experienced Python developers sometimes wonder whether the built-in list type is implemented as a linked list or a dynamic array.

In this tutorial, you’ll learn:

  • Which common abstract data types are built into the Python standard library
  • How the most common abstract data types map to Python’s naming scheme
  • How to put abstract data types to practical use in various algorithms

Note: This tutorial is adapted from the chapter “Common Data Structures in Python” in Python Tricks: The Book. If you enjoy what you read below, then be sure to check out the rest of the book.

Free Download: Get a sample chapter from Python Tricks: The Book that shows you Python's best practices with simple examples you can apply instantly to write more beautiful + Pythonic code.

Dictionaries, Maps, and Hash Tables

In Python, dictionaries (or dicts for short) are a central data structure. Dicts store an arbitrary number of objects, each identified by a unique dictionary key.

Dictionaries are also often called maps, hashmaps, lookup tables, or associative arrays. They allow for the efficient lookup, insertion, and deletion of any object associated with a given key.

Phone books make a decent real-world analog for dictionary objects. They allow you to quickly retrieve the information (phone number) associated with a given key (a person’s name). Instead of having to read a phone book front to back to find someone’s number, you can jump more or less directly to a name and look up the associated information.

This analogy breaks down somewhat when it comes to how the information is organized to allow for fast lookups. But the fundamental performance characteristics hold. Dictionaries allow you to quickly find the information associated with a given key.

Dictionaries are one of the most important and frequently used data structures in computer science. So, how does Python handle dictionaries? Let’s take a tour of the dictionary implementations available in core Python and the Python standard library.

dict: Your Go-To Dictionary

Because dictionaries are so important, Python features a robust dictionary implementation that’s built directly into the core language: the dict data type.

Python also provides some useful syntactic sugar for working with dictionaries in your programs. For example, the curly-brace ({ }) dictionary expression syntax and dictionary comprehensions allow you to conveniently define new dictionary objects:

>>>
>>> phonebook = {
...     "bob": 7387,
...     "alice": 3719,
...     "jack": 7052,
... }

>>> squares = {x: x * x for x in range(6)}

>>> phonebook["alice"]
3719

>>> squares
{0: 0, 1: 1, 2: 4, 3: 9, 4: 16, 5: 25}

There are some restrictions on which objects can be used as valid keys.

Python’s dictionaries are indexed by keys that can be of any hashable type. A hashable object has a hash value that never changes during its lifetime (see __hash__), and it can be compared to other objects (see __eq__). Hashable objects that compare as equal must have the same hash value.

Immutable types like strings and numbers are hashable and work well as dictionary keys. You can also use tuple objects as dictionary keys as long as they contain only hashable types themselves.

For most use cases, Python’s built-in dictionary implementation will do everything you need. Dictionaries are highly optimized and underlie many parts of the language. For example, class attributes and variables in a stack frame are both stored internally in dictionaries.

Python dictionaries are based on a well-tested and finely tuned hash table implementation that provides the performance characteristics you’d expect: O(1) time complexity for lookup, insert, update, and delete operations in the average case.

There’s little reason not to use the standard dict implementation included with Python. However, specialized third-party dictionary implementations exist, such as skip lists or B-tree–based dictionaries.

Besides plain dict objects, Python’s standard library also includes a number of specialized dictionary implementations. These specialized dictionaries are all based on the built-in dictionary class (and share its performance characteristics) but also include some additional convenience features.

Let’s take a look at them.

collections.OrderedDict: Remember the Insertion Order of Keys

Read the full article at https://realpython.com/python-data-structures/ »


[ Improve Your Python With 🐍 Python Tricks 💌 – Get a short & sweet Python Trick delivered to your inbox every couple of days. >> Click here to learn more and see examples ]

August 26, 2020 02:00 PM UTC


François Dion

Jupyter: JUlia PYThon and R

it's "ggplot2", not "ggplot", but it is ggplot()

 

Did you know that @projectJupyter's Jupyter Notebook (and JupyterLab) name came from combining 3 programming languages: JUlia, PYThon and R.

Readers of my blog do not need an introduction to Python. But what about the other 2?  

Today we will talk about R. Actually, R and Python, on the Raspberry Pi.

R Origin

R traces its origins to the S statistical programming language, developed in the 1970s at Bell Labs by John M. Chambers. He is also the author of books such as Computational Methods for Data Analysis (1977) and Graphical Methods for Data Analysis (1983). R is an open source implementation of that statistical language. It is compatible with S but also has enhancements over the original.

 

A quick getting started guide is available here: https://support.rstudio.com/hc/en-us/sections/200271437-Getting-Started

 



Installing Python

As a recap, in case you don't have Python 3 and a few basic modules, the installation goes as follow (open a terminal window first):


pi@raspberrypi: $ sudo apt install python3 python3-dev build-essential

pi@raspberrypi: $ sudo pip3 install jedi pandas numpy


Installing R

Installing R is equally easy:

 

pi@raspberrypi: $ sudo apt install r-recommended

 

We also need to install a few development packages:


pi@raspberrypi: $ sudo apt install libffi-dev libcurl4-openssl-dev libxml2-dev


This will allow us to install many packages in R. Now that R is installed, we can start it:


pi@raspberrypi: $ R

Installing packages

Once inside R, we can install packages using install.packages('name') where name is the name of the package. For example, to install ggplot2 (to install tidyverse, simply replace ggplot2 with tidyverse):

> install.packages('ggplot2')


To load it:


> library(ggplot2)

And we can now use it. We will use the mpg dataset and plot displacement vs highway miles per gallon and set the color to:

>ggplot(mpg, aes(displ, hwy, colour=class))+

 geom_point()



Combining R and Python

We can go at this 2 ways, from Python call R, or from R call Python. Here, from R we will call Python.

First, we need to install reticulate (the package that interfaces with Python):

> install.packages('reticulate')

And load it:

> library(reticulate)

We can verify which python binary that reticulate is using:

> py_config()

 Then we can use it to execute some python code. For example, to import the os module and use os.listdir(), from R we do ($ works a bit in a similar fashion to Python's .):

> os <- import("os")
> os$listdir(".")

Or even enter a Python REPL:

> repl_python()
>>> import pandas as pd

>>>


Type exit to leave the Python REPL.

One more trick: Radian

we will now exit R (quit()) and install radian, a command line REPL for R that is fully aware of the reticulate and Python integration:

pi@raspberrypi: $ sudo pip3 install radian


pi@raspberrypi: $ radian

This is just like the R REPL, only better. And you can switch to python very quickly by typing ~:

r$> ~

As soon as the ~ is typed, radian enters the python mode by itself:

r$> reticulate::repl_python()

>>> 

Hitting backspace at the beginning of the line switches back to the R REPL:

r$> 


I'll cover more functionality in a future post.


Francois Dion
@f_dion

August 26, 2020 01:28 PM UTC


Codementor

Mistakes Most Students Make While Learning Programming

We grew up hearing that every individual is different, but surprisingly, I have seen hundreds of students repeating the same mistakes when they start coding. And I have learned that such blunders need to be corrected as soon as possible. Indeed, you have to work hard to learn programming. But hard work without knowing what mistakes to avoid is like trying to beat the wind.

August 26, 2020 12:57 PM UTC


IslandT

Python dictionary example – return the word pattern for a given word in the form of the decimal number

In this article we will solve a python question on codewars by using the Python dictionary. Our strategy here is to use the python dictionary to keep those words that have already appeared before so we will not increase the number when we see the same word again.

Below is the entire problem we need to solve.

In cryptanalysis, words patterns can be a useful tool in cracking simple ciphers.

A word pattern is a description of the patterns of letters occurring in a word, where each letter is given an integer code in order of appearance. So the first letter is given the code 0, and second is then assigned 1 if it is different to the first letter or 0 otherwise, and so on.

As an example, the word “hello” would become “0.1.2.2.3”. For this task case-sensitivity is ignored, so “hello”, “helLo” and “heLlo” will all return the same word pattern.

Your task is to return the word pattern for a given word. All words provided will be non-empty strings of alphabetic characters only, i.e. matching the regex “[a-zA-Z]+”.

Here is the solution to the above problem.

def word_pattern(word):
    
    word_list = list(word)
    first = False
    crypto_str = ""
    increment = 0
    
    word_dict = dict() # create an empty dictionary to keep increment number
    
    for char in word_list:
        
        char = char.lower() # change all characters to lower case first before make the comparison
        
        if not first:

            first = True
            crypto_str += str(increment)
            word_dict[char] = increment
            
        else:

            if char in word_dict:
                crypto_str += "." + str(word_dict[char])
            else:
                increment += 1
                crypto_str += "." + str(increment)
                word_dict[char] = increment

    return crypto_str

As you can see from above, we will assign 0 to the first character in the string and increase the decimal number each time we see a different word, but if the word has already appeared before then we will use back the same decimal number. All these are made possible with help from the Python dictionary object.

Do you have a better solution? Leave your own solution in the comment box below this post.

August 26, 2020 12:51 PM UTC


Codementor

How to Learn Programming with Zero Stress

So today I’m gonna share on how to eliminate this kind of stress, use the best ways to learn programming, and enjoy coding happily ever after.

August 26, 2020 12:02 PM UTC


PSF GSoC students blogs

GSoC Week 12: Final Week

What I did this week?

I worked on smaller bug fixes this week.

What is coming up next?

I am having exams this week so I am taking a break.

Have I got stuck anywhere?

There are no blocking issues for me at this moment.

August 26, 2020 06:25 AM UTC

Weekly Check-In | GSoc | #13

Greetings, People of the world!

The last week of coding phase was easy, kinda fun but emotional since its like something really amazing is coming to an end. Though will always stay a part of this community but the GSOC experience is something full of thrill and excitement at every moment.

 

1. What did I do this week?

Built a few animated icons and they are now available on https://icons.eosdesignsystem.com/ To anyone who might go and check them out, Do try editing them too according to how you need them. wink

 

2. What is coming up next?

Wrapping things up and looking forward to add more animated icons to the animated icon set. And any final changes or updates that my mentors will suggest.

 

3. Did you get stuck anywhere?

Nope! It was a simple easy week. Was fun. Ahh! gonna miss this experience!!!

 

August 26, 2020 04:11 AM UTC


Matt Layman

Administer All The Things

In the previous Understand Django article, we used models to see how Django stores data in a relational database. We covered all the tools to bring your data to life in your application. In this article, we will focus on the built-in tools that Django provides to help us manage that data. What Is The Django Admin? When you run an application, you’ll find data that needs special attention. Maybe you’re creating a blog and need to create and edit tags or categories.

August 26, 2020 12:00 AM UTC

August 25, 2020


PyCoder’s Weekly

Issue #435 (Aug. 25, 2020)

#435 – AUGUST 25, 2020
View in Browser »

The PyCoder’s Weekly Logo


Data Version Control With Python and DVC

In this tutorial, you’ll learn to use DVC, a powerful tool that solves many problems encountered in machine learning and data science. You’ll find out how data version control helps you to track your data, share development machines with your team, and create easily reproducible experiments!
REAL PYTHON

A Deep Dive Into the Official Docker Image for Python

Take a stroll through the Dockerfile for the official Python Docker image. Along the way you’ll see how the image uses a custom Python build and always includes the latest version of pip.
ITAMAR TURNER-TRAURING

Python Developers Are in Demand on Vettery

alt

Vettery is an online hiring marketplace that’s changing the way people hire and get hired. Ready for a bold career move? Make a free profile, name your salary, and connect with hiring managers from top employers today →
VETTERY sponsor

Never Run Python in Your Downloads Folder

Learn about security issues that exploit how Python interacts with PATH and why you should always think twice about your current working directory.
GLYPH LEFKOWITZ

Ask for Forgiveness or Look Before You Leap?

Is it faster to “ask for forgiveness” or “look before you leap” in Python? And when is it better to use one over the other?
SEBASTIAN WITOWSKI

Common Python Packaging Mistakes

Learn about common mistakes made in creating and building a Python package and how to avoid them.
JOHN WODDER

Discussions

Why Does np.inf // 2 Result in NaN and Not Infinity?

Wait, can you even divide infinity by anything?
STACK OVERFLOW

I Programmed Someone Out of a Job and Now I Feel Bad

REDDIT

Python Jobs

Senior Backend Software Engineer (Los Angeles, CA, USA)

Tatari, Inc.

Advanced Python Engineer (Newport Beach, CA, USA)

Research Affiliates LLC

Software Engineer (Remote)

PRI Technology

Python Software Engineer (Remote)

Technology Navigators

More Python Jobs >>>

Articles & Tutorials

Python mmap: Improved File I/O With Memory Mapping

In this tutorial, you’ll learn how to use Python’s mmap module to improve your code’s performance when you’re working with files. You’ll get a quick overview of the different types of memory before diving into how and why memory mapping with mmap can make your file I/O operations faster.
REAL PYTHON

How to Fit All Human Knowledge in a Box

How to use the power of Cython to unlock the potential of both Python and C++, and help streamline the way that knowledge can be packaged and shared all over the world – even without the internet.
TESS MACCREA AND JUAN DIEGO CABALLERO • Shared by Tess McCrea

Structured Concurrency in Python With AnyIO

One of AsyncIO’s weaknesses is that it doesn’t support structured concurrency, unlike competitive projects like Trio. That’s where AnyIO comes in!
MATT WESTCOTT

Constant Time LFU

Caching is a popular way to improve application performance. The LRU strategy is popular, even though it is sometimes sub-optimal compared to the LFU strategy, because it has constant-time performance. However, you can improve LFU to get constant time performance, with some memory overhead as a trade-off.
ARPIT BHAYANI

Rust for a Pythonista, Part III: Python Bindings

In the third part of this series, you’ll learn how to add Python bindings to a Rust crate as well as test, package, and release a project to PyPI.
DMITRY DYGALO • Shared by Dmitry Dygalo

Market Prediction with ETFs & Convolutional Networks

How well can convolutional neural networks and sector ETF data predict the direction of the Dow Jones Industrial Average in the future?
GREGORY JANESH • Shared by Gregory Janesch

High Demand for Python-Driven ML Tools to Boost Robot Farming

Robots driven by Python, PyTorch, and computer vision are being used to identify, map, and target weeds in a field, helping farmers increase yield and avoid widespread use of an herbicide. While this article is non-technical, it’s an interesting study of Python’s impact in the real world.
FINTECHDEMAND.COM

Which Programming Language Is Best for Economic Research?

The most widely used programming languages for economic research are Julia, Matlab, Python and R. Despite Python’s strengths, most notably its extensive ecosystem of packages, the authors settle on Julia as their preferred language.
ALVARO AGUIRRE AND JON DANIELSSON

R and Python: The Data Science Dynamic Duo

The R language has seen a big comeback this summer, rising sharply in the TIOBE index. But the future of the relationship between R and Python is less about “R vs. Python” and more about “R and Python.”
ALEX WOODIE

Python, Javascript, R and Julia Contribute Billions to GDP

Open-source programming languages, which are incredibly valuable, are not well accounted for in economic statistics. How much economic value do you think Python has?
DAN KOPF

Projects & Code

newspaper: News, Full-Text, and Article Metadata Extraction in Python 3

GITHUB.COM/CODELUCAS

anyio: High Level Compatibility Layer for Multiple Asynchronous Event Loop Implementations

GITHUB.COM/AGRONHOLM

mccabe: McCabe Complexity Checker for Python

GITHUB.COM/PYCQA

databay: Python Interface for Scheduled Data Transfer

GITHUB.COM/VOYZ

strawberry: A GraphQL Library for Python

GITHUB.COM/STRAWBERRY-GRAPHQL

present: A Terminal-Based Presentation Tool With Colors and Effects

GITHUB.COM/VINAYAK-MEHTA • Shared by Vinayak Mehta

Events

PyCon Japan 2020

August 28 to August 30, 2020
PYCON.JP


Happy Pythoning!
This was PyCoder’s Weekly Issue #435.
View in Browser »

alt

[ Subscribe to 🐍 PyCoder’s Weekly 💌 – Get the best Python news, articles, and tutorials delivered to your inbox once a week >> Click here to learn more ]

August 25, 2020 07:30 PM UTC


PSF GSoC students blogs

Google Summer of Code Final Work Product

Google Summer of Code Final Work Product

Proposed Objectives

Modified Objectives

Objectives Completed

A combobox is a commonly used graphical user interface widget. Traditionally, it is a combination of a drop-down list or list box and a single-line textbox, allowing the user to select a value from the list. The term "combo box" is sometimes used to mean "drop-down list". Respective components, tests and tutorials were created. 

Pull Requests:

 

In interface design, a tabbed document interface or Tab is a graphical control element that allows multiple documents or panels to be contained within a single window, using tabs as a navigational widget for switching between sets of documents. Respective components, tests and tutorials were created.

Pull Requests:

 

Double click callbacks aren't implemented in VTK by default so they need to be implemented manually. With my mentor's help I was able to implement double click callbacks for all the three mouse buttons successfully.

Pull Requests:

 

The previous implementation of `TextBlock2D` was lacking a few features such as size arguments and text overflow. There was no specific way to create Texts occupying a said height or width area. Apart from that UI components like `ListBoxItem2D`, `FileMenu2D` etc had an issue where text would overflow from their specified width. In order to tackle these problems, a modification was done to `TextBlock2D` to accept size as an argument and a new method was added to clip overflowing text based on a specified width and to replace the overflowing characters with `...`.

Pull Requests:

 

Optional support for Physics engine integration of Pybullet was added to Fury. Pybullet's engine was used for the simulations and FURY was used for rendering the said simulations. Exhaustive examples were added to demonstrate various types of physics simulations possible using pybullet and fury. The said examples are as follows:

Apart from that, a document was created to explain the integration process between pybullet and fury in detail.

Pull Requests:

Objectives in Progress

The previous implementation of scrollbars were hard coded into `ListBox2D`. Therefore, it was not possible to use scrollbars with any other UI component. Apart from that, scrollbars in terms of design were limited. Creating a horizontal scrollbar was not possible. The objective of this PR is to make scrollbars separate so that other UI elements can also make use of it.

Currently, the skeletal and design aspects of the scrollbars are implemented but the combination of scrollbars with other UI components are still in progress.

Pull Requests:

 

Currently, we have access to `FileMenu2D` which allows us to navigate through the filesystem but it does not provide a user friendly Dialog to read and write files in Fury. Hence the idea is to create a file dialog which can easily open or save file at runtime. As of now, `Open` and `Save` operations are implemented. Corresponding tests and tutorials are in progress.

Pull Requests:

Other Objectives

The objects for Radio button and Checkbox tutorial were rendered using VTK's method by a fellow contributor so I decided to replace them with native FURY API. The methods were rewritten keeping the previous commits intact.

Pull Requests:

 

Weekly blogs were added for FURY's Website.

Pull Requests:

 

Timeline

Date Description Blog Link
30-05-2020 Welcome to my GSoC Blog!! Weekly Check-in #1
07-06-2020 First Week of Coding!! Weekly Check-in #2
14-06-2020 ComboBox2D Progress!! Weekly Check-in #3
21-06-2020 TextBlock2D Progress!! Weekly Check-in #4
28-06-2020 May the Force be with you!! Weekly Check-in #5
05-07-2020 Translation, Reposition, Rotation. Weekly Check-in #6
12-07-2020 Orientation, Sizing, Tab UI. Weekly Check-in #7
19-07-2020 ComboBox2D, TextBlock2D, Clipping Overflow. Weekly Check-in #8
26-07-2020 Tab UI, TabPanel2D, Tab UI Tutorial. Weekly Check-in #9
02-08-2020 Single Actor, Physics, Scrollbars. Weekly Check-in #10
09-08-2020 Chain Simulation, Scrollbar Refactor, Tutorial Update. Weekly Check-in #11
16-08-2020 Wrecking Ball Simulation, Scrollbars Update, Physics Tutorials. Weekly Check-in #12
23-08-2020 Part of the Journey is the end unless its Open Source! Weekly Check-in #13

Detailed weekly tasks and work done can be found here.

August 25, 2020 04:17 PM UTC

Weekly Check-in #12

What did I do this week?

I finished documenting sklearn operations. 

Did I get stuck somewhere?

No, everything worked out easily.

What's next?

Will make any changes as per mentor's review.

August 25, 2020 04:08 PM UTC

Week 12 Check-in

What did you do this week?

This week I started a new PR that adds multimethods for NumPy's random module. This continues the work started by one of my mentors by revising some multimethods and adding new ones as well as important classes like RandomState and Generator along with their methods. The multimethods added so far are manifold and so I won't extend the length of this blog post by enumerating them. You can however read the full list of multimethods in the PR link provided above. I also patched one ongoing PR that added multimethods for statistical functions by refactoring some default implementations. The defaults were redundantly using, in most but not all cases, a helper function for reducing the array argument's dimensions. This was brought to my attention by one of my mentors which resulted in a simple refactoring of the defaults.

What is coming up next?

The final week of GSoC is ahead of me and with that said now is the time to finish my project and write the final report. As for the first, this means mostly concluding the random module PR. If I have time I will also work on the documentation of unumpy's multimethods by linking it to NumPy's documentation through Sphinx. Although this might be an easy task, since this is my first time working with Sphinx I don't have an estimate for how long it will take me to do it. Ultimately, this might have to be done after GSoC has ended. More importantly, this upcoming week I want to focus on writing a good final report to showcase the work done during the program. This is of the utmost importance, as it is a necessary step in order to successfully pass the third and final evaluation.

Did you get stuck anywhere?

There were no blockers this week.

August 25, 2020 03:15 PM UTC


Real Python

Django Redirects

When you build web applications in Python using the Django framework, you’ll likely need to redirect the user from one URL to another. This course covers what you need to know about redirecting in Django. All the way from the low-level details of the HTTP protocol to the high-level way of dealing with them in Django.

By the end of this course you’ll learn:


[ Improve Your Python With 🐍 Python Tricks 💌 – Get a short & sweet Python Trick delivered to your inbox every couple of days. >> Click here to learn more and see examples ]

August 25, 2020 02:00 PM UTC


Artem Golubin

How to turn an ordinary gzip archive into a database

This article demonstrates how specially crafted but ordinary gzip archives can be used as a database like storage. It also introduces a Python package and explains how it works.

gzip is a popular file compression format to store large amounts of raw data. It has a good data compression ratio, but relatively slow compression/decompression speed.

Many companies use it in Big data applications when they need to store compressed CSV or JSON lines files. Such file formats are row-oriented and usually processed line by line. gzip can save a lot of space, especially when you have repetitive field names in JSON files.

Unfortunately, a compressed file can only be accessed in the streaming[....]

August 25, 2020 01:02 PM UTC


Zato Blog

Installing Zato under Mac

Zato and Mac logo

The next Zato release will offer a native Mac installer while for now an installation from source is needed - read on for details on how to set up Zato today using Homebrew.

The fundamental idea behind supporting non-Linux environments is that of making it easier for developers to work on their API services before the code is shipped to Linux test and production environments. That is, Linux is the final destination for code but that should not prevent one from using a non-Linux system during development.

One aspect to keep in mind is that the Mac version is still a technology preview - it is not a stable release yet and there may be some changes to core Zato before the final release is published. Be sure to keep your source updated.

Source installation

Creating a quickstart environment

A few screenshots

August 25, 2020 09:48 AM UTC


Codementor

What are the major differences between Python and R for data science?

Both Python and R have vast software ecosystems and communities, so either language is suitable for almost any data science task. That said, there are some areas in which one is stronger than the other.

August 25, 2020 09:04 AM UTC


PSF GSoC students blogs

Weekly Check-In #7 (16th Aug - 23rd Aug)

Hi, so we have almost reached the end of the program. It's time to wrap up all the work and polish it.

What did you do this week ?
I created the PR for date-parser incorporation and nearly all the test-cases seem to work , so that's good. On testing I did come across a bug for hindi language in number-parser , because of the different tokenization method for hindi, which I plan to fix this week.

Did you get stuck anywhere ?
Nothing major as such.

What is coming up next ?
Most of the coding part is pretty much wrapped up , just need to finalize the code , documentation etc for the final submission.

August 25, 2020 07:46 AM UTC


IslandT

Python for loop example – solving drone path

In this example, we will use the Python for loop with the range function to show the drone’s path by lighting up lamps on the path of the drone.

You will be given two strings: lamps and drone. lamps represents a row of lamps, currently off, each represented by x. When these lamps are on, they should be represented by o.

The drone string represents the position of the drone T and its flight path up until this point =. The drone always flies left to right, and always begins at the start of the row of lamps. Anywhere the drone has flown, including its current position, will result in the lamp at that position switching on.

We will create the below function to return the resulting lamps string.

def fly_by(lamps, drone):
    
    lamps = list(lamps)

    for i in range(0, len(drone)):
        if (i < len(lamps)):
            lamps[i] = 'o'
        else:
            break
    return "".join(lamps)

As an example, if you enter these strings into the above function you will get the following outcome.

fly_by('xxxxxx', '====T') # return 'ooooox'

As you can see we are using the string length of the drone in the for loop to change the state of the lamp.

Hello, please only leave the comment which is related to the above topic, I will approve all comments that are related to the above topic but the comment must comes with a python solution.

August 25, 2020 06:56 AM UTC


Podcast.__init__

Working In The Code Mines: Mining Software Repositories With PyDriller

A large portion of the software industry has standardized on Git as the version control sytem of choice. But have you thought about all of the information that you are generating with your branches, commits, and code changes? Davide Spadini created the PyDriller framework to simplify the work of mining software repositories to perform research on the technical and social aspects of software engineering. In this episode he shares some of the insights that you can gain by exploring the history of your code, the complexities of building a framework to interact with Git, and some of the interesting ways that PyDriller can be used to inform your own development practices.

Summary

A large portion of the software industry has standardized on Git as the version control sytem of choice. But have you thought about all of the information that you are generating with your branches, commits, and code changes? Davide Spadini created the PyDriller framework to simplify the work of mining software repositories to perform research on the technical and social aspects of software engineering. In this episode he shares some of the insights that you can gain by exploring the history of your code, the complexities of building a framework to interact with Git, and some of the interesting ways that PyDriller can be used to inform your own development practices.

Announcements

  • Hello and welcome to Podcast.__init__, the podcast about Python and the people who make it great.
  • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $60 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
  • You listen to this show to learn and stay up to date with the ways that Python is being used, including the latest in machine learning and data analysis. For more opportunities to stay up to date, gain new skills, and learn from your peers there are a growing number of virtual events that you can attend from the comfort and safety of your home. Go to pythonpodcast.com/conferences to check out the upcoming events being offered by our partners and get registered today!
  • Your host as usual is Tobias Macey and today I’m interviewing Davide Spadini about PyDriller, a framework for mining software repositories

Interview

  • Introductions
  • How did you get introduced to Python?
  • Can you start by describing what PyDriller is and how the project got started?
    • How is Pydriller different from other Git frameworks?
  • What kinds of information can you discover by mining a software repository?
    • Where and how might the collected information be used?
  • What are the limitations of the capabilities offered by Git for investigating the repository?
  • What are the additional metrics that you are able to extract using PyDriller?
  • Can you describe how PyDriller itself is implemented?
    • How has the project evolved since you first began working on it?
  • I noticed that for testing PyDriller you crafted a set of repositories to serve as test cases. What has been the most complex or challenging aspect of writing meaningful tests to ensure a reasonable coverage of this problem domain?
  • What would be required to add support for other version control systems?
  • How have you used PyDriller in your own research?
  • What are some of the most interesting, unexpected, or innovative ways that you have seen PyDriller used?
  • What are some of the most interesting, unexpected, or challenging lessons that you have learned while working on and with PyDriller?
  • What do you have planned for the future of PyDriller?

Keep In Touch

Picks

Closing Announcements

  • Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
  • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
  • If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@podcastinit.com) with your story.
  • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
  • Join the community in the new Zulip chat workspace at pythonpodcast.com/chat

Links

The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA

August 25, 2020 12:32 AM UTC