Sitelet https://github.com/DanuserLab/u-segment3D/pull/16
Skip to content

Support CUDA 12 and NVIDIA Blackwell GPUs - #16

Merged
DanuserLab merged 1 commit into
DanuserLab:masterfrom
kevin-m-dean:kdean/cuda12-blackwell
Aug 25, 2026
Merged

DanuserLab merged 1 commit into
DanuserLab:masterfrom
kevin-m-dean:kdean/cuda12-blackwell

Conversation

@kevin-m-dean

Copy link
Copy Markdown
Contributor

Summary

Replace the non-macOS cupy-cuda11x dependency with cupy-cuda12x and update the installation guidance for CUDA 12 GPU runtimes.

This is needed for NVIDIA Blackwell devices. In our production integration, the CUDA 11 CuPy wheel failed immediately on an RTX PRO 6000 Blackwell with CUDA_ERROR_NO_BINARY_FOR_GPU. The same u-Segment3D workflow ran successfully after changing only the CuPy distribution to cupy-cuda12x.

The documentation also keeps an explicit CUDA 12.6 PyTorch path for Pascal GPUs such as the P40. Newer PyTorch/CUDA builds can be used for Ampere and Blackwell, but the P40 requires a build that still contains Pascal kernels.

Quantitative validation

We tested one real microscopy field with shape 9 x 2000 x 2000 x 4, using the same cyto3 weights, u-Segment3D configuration, image bytes, and segmentation parameters across runs.

GPU CuPy package PyTorch CUDA build Labels Foreground voxels Runtime
A100 cupy-cuda11x==13.6.0 cu130 309 3,334,997 514 s
A100 cupy-cuda12x==13.6.0 cu130 309 3,334,997 513 s
RTX PRO 6000 Blackwell cupy-cuda12x==13.6.0 cu130 304 3,338,014 356 s
A100 cupy-cuda12x==13.6.0 cu126 305 3,335,662 498 s
P40 cupy-cuda12x==13.6.0 cu126 307 3,337,938 474 s

The A100 CUDA-11 and CUDA-12 outputs were byte-identical, including the label-array SHA-256:

25a89a...

Cross-architecture comparisons were also highly concordant:

Comparison Dice IoU Boundary F1 ARI
A100 cu130 vs Blackwell cu130 0.999433 0.998867 0.997581 0.999308
A100 cu130 vs A100 cu126 0.999724 0.999448 0.999212 0.999677
A100 cu126 vs P40 cu126 0.999547 0.999094 0.997855 0.999449
Blackwell cu130 vs P40 cu126 0.999580 0.999161 0.999305 0.999555

Instance counts differ slightly across GPU architectures and PyTorch builds, so we do not claim byte-identical cross-device output. However, foreground and instance-label agreement are above 0.997 on every reported metric. The A100 serves as a bridge showing that the CuPy CUDA-11 to CUDA-12 dependency change itself is byte-identical under the same PyTorch build.

Changes

  • Use cupy-cuda12x on non-macOS platforms.
  • Explain that only one CuPy distribution should be installed in an environment.
  • Update Linux and Windows examples to CUDA 12-compatible PyTorch installation commands.
  • Document that Blackwell needs CUDA 12.8 or newer and that CUDA 12.6 remains useful for Pascal compatibility.

Verification

Built the wheel from this branch and inspected its generated METADATA:

Requires-Dist: cupy-cuda12x; sys_platform != "darwin"

No CUDA-11 CuPy dependency remains in the built package metadata.

Replace the non-macOS cupy-cuda11x dependency with cupy-cuda12x so installations can run natively on NVIDIA Blackwell devices, which require CUDA 12.8 or newer.

Update Linux and Windows installation guidance to select a CUDA-compatible PyTorch build, retain a CUDA 12.6 path for Pascal GPUs, and warn against installing multiple CuPy distributions in one environment.

Verified by building the wheel and inspecting its METADATA for the exact cupy-cuda12x environment-marked requirement.
@DanuserLab
DanuserLab merged commit 34ea57b into DanuserLab:master Aug 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants