Support CUDA 12 and NVIDIA Blackwell GPUs - #16
Merged
DanuserLab merged 1 commit intoAug 25, 2026
Merged
Conversation
Replace the non-macOS cupy-cuda11x dependency with cupy-cuda12x so installations can run natively on NVIDIA Blackwell devices, which require CUDA 12.8 or newer. Update Linux and Windows installation guidance to select a CUDA-compatible PyTorch build, retain a CUDA 12.6 path for Pascal GPUs, and warn against installing multiple CuPy distributions in one environment. Verified by building the wheel and inspecting its METADATA for the exact cupy-cuda12x environment-marked requirement.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Replace the non-macOS
cupy-cuda11xdependency withcupy-cuda12xand update the installation guidance for CUDA 12 GPU runtimes.This is needed for NVIDIA Blackwell devices. In our production integration, the CUDA 11 CuPy wheel failed immediately on an RTX PRO 6000 Blackwell with
CUDA_ERROR_NO_BINARY_FOR_GPU. The same u-Segment3D workflow ran successfully after changing only the CuPy distribution tocupy-cuda12x.The documentation also keeps an explicit CUDA 12.6 PyTorch path for Pascal GPUs such as the P40. Newer PyTorch/CUDA builds can be used for Ampere and Blackwell, but the P40 requires a build that still contains Pascal kernels.
Quantitative validation
We tested one real microscopy field with shape
9 x 2000 x 2000 x 4, using the same cyto3 weights, u-Segment3D configuration, image bytes, and segmentation parameters across runs.cupy-cuda11x==13.6.0cupy-cuda12x==13.6.0cupy-cuda12x==13.6.0cupy-cuda12x==13.6.0cupy-cuda12x==13.6.0The A100 CUDA-11 and CUDA-12 outputs were byte-identical, including the label-array SHA-256:
25a89a...Cross-architecture comparisons were also highly concordant:
Instance counts differ slightly across GPU architectures and PyTorch builds, so we do not claim byte-identical cross-device output. However, foreground and instance-label agreement are above 0.997 on every reported metric. The A100 serves as a bridge showing that the CuPy CUDA-11 to CUDA-12 dependency change itself is byte-identical under the same PyTorch build.
Changes
cupy-cuda12xon non-macOS platforms.Verification
Built the wheel from this branch and inspected its generated
METADATA:No CUDA-11 CuPy dependency remains in the built package metadata.