This modular tutorial contains content on all things related to accelerated Python:
- Notebooks containing lessons and exercises organized by topic, including CUDA Tile and the PyHPC material on CuPy, CUDA kernels, MPI, JAX, PyOMP, and Python/C++ interoperability. They are intended for self-paced or instructor-led learning and can be run on NVIDIA Brev or Google Colab.
- Slides containing the lecture content for the lessons.
- Syllabi that select a subset of the notebooks for a particular learning objective.
- Docker Images and Docker Compose files for creating Brev Launchables or running locally.
Brev Launchables of this tutorial should use:
- L40S, L4, or T4 instances (for non-distributed notebooks other than CUDA Tile).
- A10G or newer Ampere, Ada, or Blackwell instances for CUDA Tile notebooks.
- 4xL4 or 2xL4 instances (for distributed notebooks).
- A host driver with CUDA 13 support.
- Crusoe or any other provider with Flexible Ports.
- CUDA Python - CuPy, cuDF, CCCL, & Kernels - 8 Hours.
- CUDA Python Tutorial - CuPy, SIMT, & Tile - 8 Hours.
- CUDA Python - cuda.core & CCCL - 2 Hours
- CUDA Tile Tutorial
- PyHPC - NumPy, CuPy, & mpi4py - 4 Hours
- PyHPC - CuPy, Kernels, MPI, JAX, OMP, Interop - 2 Days
The two-day PyHPC syllabus and the eight-hour CUDA Python course select lessons from the same topic directories as every other Accelerated Python course. The eight-hour course contains only the CuPy and Numba CUDA kernel-authoring foundations followed by CUDA Tile lessons 44 through 47. Self-contained lessons with a Colab badge can run on Google Colab. C++ interoperability and the Shallow Water Equations applications require the shared tutorial image and checked-in source files.
Applications 81 through 87 form one ordered case study. They solve the same 1D
Shallow Water Equations problem with NumPy, JAX, PyOMP, nanobind, CppJIT/CUB,
and mpi4py. Notebooks 81 through 86 write measurements to timings.json; run
them before notebook 87, which compares the results. CppJIT is built from the
course's pinned ISC 2026 branch and is available in the tutorial image rather
than from the public alpha package.
The image has one system Python environment. The Python 3 (PyHPC) and Nsight
profiler kernels only select course-specific startup and MPI settings: MPICH is
used for local multi-rank exercises, while the default Python kernel continues
to use OpenMPI. For CSCS Alps/Daint deployment, follow the CSCS launch
guide.
The merged course initializes a new Docker repository volume named
accelerated-python_accelerated-computing-hub-v2. Existing volumes named
accelerated-python_accelerated-computing-hub or
pyhpc_accelerated-computing-hub are deliberately left untouched. Before
removing either old volume, copy any edited notebooks from the former
Accelerated Python tree and move any former tutorials/pyhpc/notebooks work
into the matching fundamentals, kernels, distributed, or applications
directory under tutorials/accelerated-python/notebooks in the new deployment.
Do not run docker compose down --volumes against the old deployment until
that work has been backed up and verified.
| # | Exercise | Link | Solution |
|---|---|---|---|
| 60 | mpi4py | ||
| 61 | Dask | ||
| 62 | mpi4py: Heat Equation |