Sitelet https://github.com/NVIDIA/accelerated-computing-hub/tree/main/tutorials/accelerated-python
Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 
 
 
 
 

README.md

Accelerated Python Tutorial

This modular tutorial contains content on all things related to accelerated Python:

  • Notebooks containing lessons and exercises organized by topic, including CUDA Tile and the PyHPC material on CuPy, CUDA kernels, MPI, JAX, PyOMP, and Python/C++ interoperability. They are intended for self-paced or instructor-led learning and can be run on NVIDIA Brev or Google Colab.
  • Slides containing the lecture content for the lessons.
  • Syllabi that select a subset of the notebooks for a particular learning objective.
  • Docker Images and Docker Compose files for creating Brev Launchables or running locally.

Brev Launchables of this tutorial should use:

  • L40S, L4, or T4 instances (for non-distributed notebooks other than CUDA Tile).
  • A10G or newer Ampere, Ada, or Blackwell instances for CUDA Tile notebooks.
  • 4xL4 or 2xL4 instances (for distributed notebooks).
  • A host driver with CUDA 13 support.
  • Crusoe or any other provider with Flexible Ports.

Syllabi

The two-day PyHPC syllabus and the eight-hour CUDA Python course select lessons from the same topic directories as every other Accelerated Python course. The eight-hour course contains only the CuPy and Numba CUDA kernel-authoring foundations followed by CUDA Tile lessons 44 through 47. Self-contained lessons with a Colab badge can run on Google Colab. C++ interoperability and the Shallow Water Equations applications require the shared tutorial image and checked-in source files.

Applications 81 through 87 form one ordered case study. They solve the same 1D Shallow Water Equations problem with NumPy, JAX, PyOMP, nanobind, CppJIT/CUB, and mpi4py. Notebooks 81 through 86 write measurements to timings.json; run them before notebook 87, which compares the results. CppJIT is built from the course's pinned ISC 2026 branch and is available in the tutorial image rather than from the public alpha package.

The image has one system Python environment. The Python 3 (PyHPC) and Nsight profiler kernels only select course-specific startup and MPI settings: MPICH is used for local multi-rank exercises, while the default Python kernel continues to use OpenMPI. For CSCS Alps/Daint deployment, follow the CSCS launch guide.

Upgrading an existing deployment

The merged course initializes a new Docker repository volume named accelerated-python_accelerated-computing-hub-v2. Existing volumes named accelerated-python_accelerated-computing-hub or pyhpc_accelerated-computing-hub are deliberately left untouched. Before removing either old volume, copy any edited notebooks from the former Accelerated Python tree and move any former tutorials/pyhpc/notebooks work into the matching fundamentals, kernels, distributed, or applications directory under tutorials/accelerated-python/notebooks in the new deployment. Do not run docker compose down --volumes against the old deployment until that work has been backed up and verified.

Notebooks

Fundamentals

# Exercise Link Solution
01 NumPy Intro: ndarray Basics
02 NumPy Linear Algebra: SVD Reconstruction
03 NumPy to CuPy: ndarray Basics
04 NumPy to CuPy: SVD Reconstruction
05 Memory Spaces: Power Iteration
06 Asynchrony: Power Iteration
07 CUDA Core: Devices, Streams and Memory

Libraries

# Exercise Link Solution
20 cuDF: NYC Parking Violations
21 cudf.pandas: NYC Parking Violations
22 cuML
23 CUDA CCCL: Customizing Algorithms
24 nvmath-python: Interop
25 nvmath-python: Kernel Fusion
26 nvmath-python: Stateful APIs
27 nvmath-python: Scaling
28 PyNVML

Kernels

# Exercise Link Solution
40 Kernel Authoring: Copy
41 Kernel Authoring: Book Histogram
42 Kernel Authoring: Gaussian Blur
43 Kernel Authoring: Black and White
44 cuTile Python Intro: Vector Add
45 cuTile Python: Matrix Multiplication
46 cuTile Python: Transpose
47 cuTile Python: Activation Functions

Distributed

# Exercise Link Solution
60 mpi4py
61 Dask
62 mpi4py: Heat Equation

Applications

# Exercise Link Solution
80 C++ Interoperability
81 Shallow Water Equations: NumPy Baseline
82 Shallow Water Equations: JAX
83 Shallow Water Equations: PyOMP
84 Shallow Water Equations: nanobind
85 Shallow Water Equations: CppJIT and CUB
86 Shallow Water Equations: mpi4py
87 Shallow Water Equations: Synthesis