Update code to Optimise-away small GPU allocations for projector & mu… unitaryHACK26 - #783
Conversation
|
This is a wonderful diff - I'm kicking myself for not noticing Can you please share the mentioned driver/scripts for benchmarking? Can either whack it into a comment here, or include it into the diff (which we can delete later - changes will be squashed so it won't pollute your work). |
|
Note to self The template parameters of the below functions and functors are now redundant:
They can all be removed, along with the parameter dispatch in |
…review)
After the bitmask reformulation the projector no longer specialises on the target
count, so its numTargs template is dead code:
- drop the template from functor_projectStateVec/functor_projectDensMatr,
thrust_{statevec,densmatr}_multiQubitProjector_sub and
gpu_{statevec,densmatr}_multiQubitProjector_sub, and their
INSTANTIATE_FUNC_OPTIMISED_FOR_NUM_TARGS instantiations;
- apply the same bitmask reformulation to the CPU projector
(cpu_{statevec,densmatr}_multiQubitProjector_sub) so its template goes too;
- simplify accel_{statevec,densmatr}_multiQubitProjector_sub to a plain
isGpuAccelerated ? gpu_ : cpu_ branch (no GET_CPU_OR_GPU_FUNC dispatch).
Shared dispatch macros and the calcProb* template chain are untouched. Unit tests
pass on CPU/CPU+OMP/GPU/GPU+OMP (57146 assertions). Adds a throw-away benchmarks/
driver (not wired into CMake/CI); safe to squash/drop on merge.
Thanks @TysonRayJones! Glad it's useful. 🎉 I've added the driver to the PR under cmake -S . -B build_bench \
-D QUEST_ENABLE_CUDA=ON -D CMAKE_CUDA_ARCHITECTURES=120 \
-D CMAKE_BUILD_TYPE=Release \
-D USER_SOURCE_NAMES=benchmarks/benchmark_749.cpp \
-D USER_OUTPUT_EXE_NAME=bench_749
cmake --build build_bench --target bench_749 -j
./build_bench/bench_749 4 20 3 2000 # [minQ maxQ numTargs reps]It forces the single-GPU path ( On my machine (RTX PRO 6000 Blackwell, CUDA 13.0, sm_120): CUDA runtime API counts (N = 8…12, 200 reps, both ops, via
Per-call wall time (µs, numTargs = 3, 2000 reps):
(The residual ~1000 |
Done — I went ahead and removed them (pushed in a follow-up commit). Summary:
One thing to confirm: your note listed the GPU-side functions, but the A couple of notes for the record:
Verification (RTX PRO 6000 Blackwell, CUDA 13.0, sm_120, Release): rebuilt Also confirming: I'm fine with the |
|
Hi there Amon, It may take us a little while to integrate this - we have an idea how to further optimise a final remaining In any case, @JPRichings has profiled your solution and confirmed that you have satisfied the unitaryHACK challenge 🎉 🎉 Please comment on the issue (#749) so that we can assign it to you, awarding you the challenge. Nice work! |
…-Kit#717) Open the one-bit gap with a single shared low mask + a shift-by-one of the high bits, instead of the equivalent right/left-shift + nested concatenateBits. Algebraically identical (cannot change results); shaves a couple of ops off every unrolled insertBits iteration, benefiting all callers. Op-count/readability cleanup, not a measured speed win. Verified against the full unit-test suite (2,293,050 assertions) across CPU/CPU+OMP/GPU/GPU+OMP; the one unrelated failure (setQuESTNumGpuThreadsPerBlock) is pre-existing on devel.
|
Thanks so much @TysonRayJones, and thanks @JPRichings for profiling it — delighted it satisfies the challenge! 🎉 On the further optimisation: I dug into it and wanted to share what I found, since there are actually two separate things bundled in there.
I verified the One small unrelated heads-up from running the whole suite on this box (RTX PRO 6000 Blackwell, CUDA 13): the |
|
Thanks, for this additional work. we have also spotted the |
|
Thanks @JPRichings! Great to hear it's already fixed — I'll skip opening a separate PR for Thanks again to you both for the thorough review and for profiling the solution. |
|
Oh good call about the remaining Alas I'm not convinced your |
QuEST-Kit#717)" This reverts commit 9b87ec3. Reverted at the maintainer's request (review on PR QuEST-Kit#783): with inlining and compiler optimisation the rewritten insertBit produces identical bit operations to the original concatenateBits-based version, so the change only affected readability. Restoring the original keeps bitwise.hpp untouched; the thrust::reduce temp-allocation pooling idea will be pursued separately (re-using the persistent GPU cache in gpu_config.cpp).
|
Done — reverted in And nice — re-using the persistent, user-clearable |
* gpu_thrust.cuh: removed thrust::[unary|binary]_function to fix compatibility with CUDA 13. Fixes #693. * Simplify installation path configuration (#697) * Simplify installation path configuration Removed unnecessary path normalization and appending for installation. * Updates CMake config for conditional installs Modifies CMakeLists to conditionally build shared libraries and install binaries only at the top-level project. Introduces the INSTALL_BINARIES option to control the inclusion of example binaries in the installation process. Corrects a typo from 'RATH' to 'RPATH' for build configurations. * Doc binary install. (#701) * docs/cmake.md: fixed formatting of non-default options for mt and distribution * cmake: wrapped user source install in if(INSTALL_BINARIES) * docs/cmake.md: added INSTALL_BINARIES option * Thrust Overflow Fix (#699) * gpu_thrust.cuh: modified initial thrust counting iterator declarations to use long long to avoid overflow at >30 qubits. Fixes #698. * patched test of rightapplyCompMatr distributed validation The operation validation tests previously always uses a statevector to test the "targeted amps fit in node" validation, though the rightapply*() functions cannot accept statevectors, instead only density matrices. Because the "was given a density matrix" validation happens before "targeted amps fit in node" validation, the latter intended triggered error was beaten out by the earlier unintended one. Now, we are careful to pass a density matrix Qureg to the validation of "targeted amps fit in node" when triggered by a function which 'right-applies' (and is ergo only compatible with density matrices) * changed literals to defensive type --------- Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com> * patched validation test of rightapplyCompMatr (#703) addresses #700 * Random permutation of strings in Trotterisation (#702) * Implement PauliStrSum random permutations inspired by [arXiv:1805.08385](https://arxiv.org/abs/1805.08385) * Add randomisation to Trotter functions * Document random Pauli permutations for Trotterisation * Add test for Trotter randomisation * Added unit tests requested by Quantum Motion (#706) * Added unit tests requested by Quantum Motion * tests/unit/trotterisation.cpp: updated time evo calls to new API * tests/unit/trotterisation.cpp: updated authorlist * Fixed valgrind errors * tests/unit/trotterisation.cpp: tuned floating-point comparison epsilon to account for worst-case scenario which is single precision, single thread * added get-arbitrary-qureg test util since it will be used frequently by new input validation * removed Qureg creation in validation tests so that test failures do not cause memory leaks and e.g. add to valgrind noise. Tests now instead use getArbitraryCachedStatevec() or getArbitraryCachedDensmatr() to obtain an existing qureg with an arbitrary deployment. * restoring missing-validation comments since the validation for these functions wasn't added. Such functions have additional tests to their tested counterparts; for example, validating that matrix elements are non-zero when given a negative exponent * fixing test category * added missing operation validation tests * fixed indentation * making spacing consistent and adding a missing Hermiticity validation to applyTrotterizedUnitaryTimeEvolution test * added warning about untested deployments * removed defunct signature * patching C++ validation err msg Previously, an error message of the C++ API was not substituting in values for its placeholder variables. This affected the C++ variants of the below functions when passing vectors for the targets and outcome parameters of mismatching length: - calcProbOfMultiQubitOutcome - leftapplyMultiQubitProjector - rightapplyMultiQubitProjector - applyMultiQubitProjector - applyForcedMultiQubitMeasurement * added missing C++ API signatures * added C++-API validation tests * updated doc warnings * added Vasco to Trotter API authorlist * merged Tyson's patches --------- Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com> Co-authored-by: Maurice Jamieson <m.jamieson@epcc.ed.ac.uk> Co-authored-by: Oliver Thomson Brown <otbrown@users.noreply.github.com> Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com> * Add inverse to QFT implementation (#705) * Fix missing inverse argument to QFT validation tests * Implement Pauli sorting (#708) * Fix HIP install failures in CI (#718) CI updated to use latest AMD ROCm install instructions. As of this commit corresponding to ROCm 7.2. --------- Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com> * Fix applyMultiStateControlledSqrtSwap argument list (#738) Remove numControls argument from applyMultiStateControlledSqrtSwap overloaded definition taking std::vector<int> (cherry picked from commit 9c20792) Co-authored-by: D-Exposito <dexposito@cesga.es> * Trotterisation test update (#728) * tests/unit/trotterisation.cpp: updated to use REQUIRE_AGREE and cached statevecs and densmats, and both permutePaulis options * tests/utils/compare.hpp/cpp: added setters for test epsilon * tests/unit/trotterisation.cpp: adjusted test epsilon for quad precision imaginary time evolution tests * tests/unit/trotterisation.cpp: moved unitary time evo test to REQUIRE_AGREE * tests/utils/cache.hpp/cpp: added additional utilities for creating and destroying temp caches (which I guess makes them not caches?) with a set number of qubits * tests/unit/trotterisation.cpp: updated unitary time evo test to test across deployments * tests/unit/trotterisation.cpp: reduced number of qubits and increased number of steps to admit the possibility of testing density matrices too * tests/unit/trotterisation.cpp: added density matrix tests * reduce test precision to lazily pass CPU clang quad-precision * skip Trotter tests in paid CI * changing varname convention * renaming cache funcs --------- Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com> Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com> * added Daniel Patino to authorlist * CMake warn when non-release build (#742) --------- Co-authored-by: Oliver Thomson Brown <otbrown@users.noreply.github.com> * Stop Trotter funcs mutating PauliStrSum (#740) Formerly, the Trotter functions (such as applyTrotterizedPauliStrSumGadget()), when passed permutePaulis=true, would randomly permutate the order of the passed PauliStrSum, mutating it and affecting the outputs of subsequent functions like reportPauliStrSum(). The function also contained superfluous memory allocs/copies equal in size to the PauliStrSum. Now, the PauliStrSum is never mutated, and an internally allocated ordering list keeps track of the randomised permutation. We also updated the doc, renamed permutePaulis to permuteTerms, and improved validation. Note that 'permuteTerms' had not yet reached main/release, so these changes do not need to be documented in the v4.3 release notes. * Created custom backend complex types (#729) Created cpu_qcomp and gpu_qcomp (from a shared base_qcomp) to avoid std::complex arithmetic operators in hot loops which caused performance issues. Removed all prior compiler flags and related scaffolding attempting to mitigate the performance issue. Also gave MSVC build the params `/Zc:preprocessor -Xcompiler=/Zc:preprocessor /bigobj` as needed for compilation of the unit tests on my windows machines. * Replace vector<int> with SmallList (a stack array) (#743) This is to circumvent the std::vector performance overheads visible in few-qubit simulation (responsible for a performance regression from v3; see #720), and also so that qubit lists can be passed directly to CUDA kernels without conversion (as explored in #739). * Added few-qubit optimisations (#750) Optimisations include: - Adopted SmallView (const SmallList&) to avoid superfluous SmallList copies - Made internally created matrices static - Change accelerator dynamic function vectors to static arrays - Exit all validators early when validation is disabled Additional cleanup includes: - Tidied accelerator macros (replaced param-specific macros like "numCtrls" and "numTargs" with "param") - Fill ctrlStates vectors with default before localiser - Renamed getBitsFromInteger to setToBitsOfInteger - Adopted const in bitwise.hpp to better express intent Note that the naming of SmallList and SmallView will be subsequently changed to List64 and ConstList64 * Renamed debug API functions to contain "QuEST" (#752) * Renamed environment variables to begin with"QUEST" (#755) * Renamed CMake vars and preprocessors (#756) such that they all begin with QUEST, but some have additional changes * Renamed Small(List|View) to (Const)List64 (#757) * Defer Catch2 test discovery so that we can compile MPI tests on systems which cannot actually run with MPI, because they are missing an MPI or UCX library file, as is witnessed in the CI (when compiling with MPICH). It's generally irksome too to trigger an execution of the test binary (which itself initialises QuEST) during build when on a HPC platform with distinct submit and compute nodes * Enable user to take ownership of MPI (#722) * Added ENABLE_SUBCOMM build option * Moved from MPI_COMM_WORLD to mpiQuestComm * Decided passing *MPI_Comm was probably overly cautious, and updated function name to comm_getMpiComm * environment.cpp: added methods to reset rank and numNodes, and reporting for subcomm compiled * comm_config.hpp/cpp: added comm_setMpiComm * CMakeLists.txt: PUBLIC MPI::MPI_CXX turned out to be unhelpful, even for SubComm, because of course it enforces CXX * Added new custom QuESTEnv initialiser which allow user to positively declare that they take ownership of MPI * validation.cpp: updated comm_end call * comm_config.hpp: added config.h include so COMPILE_MPI is actually defined * subcommunicator.h/cpp: implemented QuESTEnv initialiser with custom MPI_Comm * CMake: added subcommunicator.cpp * comm_config.hpp: added missing config.h include... * comm_config.cpp: explicitly initialise mpiCommQuest to MPI_COMM_NULL, updated setComm for init only workflow * quest.h: added subcommunicator header * CMake: added MPI to application binaries when SUBCOMM is enabled * comm_routines.cpp: post Irecv before Isend which probably won't do anything but it makes MPI library implementers less nervous * tests: added new env test for initCustomMpiQuESTEnv * Added error throws to comm_config to cover new scenarios of badness with user owned MPI * subcommunicator.cpp: updated var names to match QuEST style * tests/unit/initialisations.cpp: slightly modified setQuregAmps test to avoid unexpected test failure due to range checking when compild in Debug configuration * Updated validation in comm_setMpiComm Co-authored-by: iarejula-bsc <inigo.arejula@bsc.es> * userOwnsMpi int->bool * comm_config.cpp: corrected call to MPI_Comm_free * subcommunicator.cpp: userOwnsMpi int->bool * subcommunicator.cpp: added comm_isInit guard around comm_setMpiComm * environment.cpp: USER_OWNS_MPI -> userOwnsMpi * comm_init: fixed case where useDistrib = 0 and userOwnsMpi = true * comm_init: moved (recently) misplaced MPI_Init * AUTHORS.txt: added iarejula-bsc * Added placeholder docstrings to new initialisers * docs/cmake.md: added ENABLE_SUBCOMM to list of QuEST CMake vars * Newly added COMPILE_MPI -> QUEST_COMPILE_MPI * ENABLE_SUBCOMM -> QUEST_ENABLE_SUBCOMM * CMake: corrected OpenMP and subcommunicator pre-processor definitions --------- Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com> Co-authored-by: iarejula-bsc <inigo.arejula@bsc.es> * Add flush and sync around prints (#763) to reduce the likelihood of users printing from non-root nodes interrupting QuEST root output. This is not bullet-proof; we sync the active communicator rather than MPI_COMM_WORLD so the user-controlled non-participating processes may still be printing. Furthermore, even if all processes participate, some may have outstanding non-root prints that are not aggregated to the user screen by the time MPI_Barrier finishes. But these syncs greatly reduce the change of corruption, and are effectively free! * Add GPU-aware MPICH detection This enables CRAY MPICH platforms to leverage GPU-awareness, greatly accelerating distributed GPU simulation Co-authored-by: JPRichings <james.richings@ed.ac.uk> * Cleanup custom MPI flow (#762) Important changes: - permit user initialisation of MPI when QuEST is not distributed - changed QuESTEnv fields bool from int (e.g. isMultithreaded) - add user-input validation for custom MPI calls - disambiguated comm_config.cpp concepts of "MPI is initialised" (comm_isMpiInit) from "QuEST communication is active" (comm_isActive) - refactored comm_config.cpp flow, especially related to pre-quest-init flow (during validation) - added Oliver's custom-MPI examples (from #712) - moved new API functions to experimental.h - tweaked reportQuESTEnv output grouping * Added user-control of GPU num threads per block (#736) Added: - QUEST_DEFAULT_NUM_GPU_THREADS_PER_BLOCK CMake option - QUEST_DEFAULT_NUM_GPU_THREADS_PER_BLOCK environment variable - setQuESTNumGpuThreadsPerBlock() API function - getQuESTNumGpuThreadsPerBlock() API function - set_num_gpu_threads examples in examples/extended --------- Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com> Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com> * Fix compiler warnings (#770) Beware this included removing the superfluous `numControls` argument from the C++only `std::vector` overload of `applyMultiStateControlledCompMatr2`, which is technically a teeny tiny API break ¯\_(ツ)_/¯ * tests/unit/debug.cpp: updated setQuESTSeeds validation tests to include new validation (#771) Updated number of seeds test to use a valid pointer and added a separate NULL pointer test. * Fix Windows CI test_free.yml: added Release config to ctest commands (#773) * Create v4.3 release candidate (#776) * Fix applyMultiStateControlledSqrtSwap argument list (#738) Remove numControls argument from applyMultiStateControlledSqrtSwap overloaded definition taking std::vector<int> (cherry picked from commit 9c20792) Co-authored-by: D-Exposito <dexposito@cesga.es> * Trotterisation test update (#728) * tests/unit/trotterisation.cpp: updated to use REQUIRE_AGREE and cached statevecs and densmats, and both permutePaulis options * tests/utils/compare.hpp/cpp: added setters for test epsilon * tests/unit/trotterisation.cpp: adjusted test epsilon for quad precision imaginary time evolution tests * tests/unit/trotterisation.cpp: moved unitary time evo test to REQUIRE_AGREE * tests/utils/cache.hpp/cpp: added additional utilities for creating and destroying temp caches (which I guess makes them not caches?) with a set number of qubits * tests/unit/trotterisation.cpp: updated unitary time evo test to test across deployments * tests/unit/trotterisation.cpp: reduced number of qubits and increased number of steps to admit the possibility of testing density matrices too * tests/unit/trotterisation.cpp: added density matrix tests * reduce test precision to lazily pass CPU clang quad-precision * skip Trotter tests in paid CI * changing varname convention * renaming cache funcs --------- Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com> Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com> * added Daniel Patino to authorlist * CMake warn when non-release build (#742) --------- Co-authored-by: Oliver Thomson Brown <otbrown@users.noreply.github.com> * Stop Trotter funcs mutating PauliStrSum (#740) Formerly, the Trotter functions (such as applyTrotterizedPauliStrSumGadget()), when passed permutePaulis=true, would randomly permutate the order of the passed PauliStrSum, mutating it and affecting the outputs of subsequent functions like reportPauliStrSum(). The function also contained superfluous memory allocs/copies equal in size to the PauliStrSum. Now, the PauliStrSum is never mutated, and an internally allocated ordering list keeps track of the randomised permutation. We also updated the doc, renamed permutePaulis to permuteTerms, and improved validation. Note that 'permuteTerms' had not yet reached main/release, so these changes do not need to be documented in the v4.3 release notes. * Created custom backend complex types (#729) Created cpu_qcomp and gpu_qcomp (from a shared base_qcomp) to avoid std::complex arithmetic operators in hot loops which caused performance issues. Removed all prior compiler flags and related scaffolding attempting to mitigate the performance issue. Also gave MSVC build the params `/Zc:preprocessor -Xcompiler=/Zc:preprocessor /bigobj` as needed for compilation of the unit tests on my windows machines. * Replace vector<int> with SmallList (a stack array) (#743) This is to circumvent the std::vector performance overheads visible in few-qubit simulation (responsible for a performance regression from v3; see #720), and also so that qubit lists can be passed directly to CUDA kernels without conversion (as explored in #739). * Added few-qubit optimisations (#750) Optimisations include: - Adopted SmallView (const SmallList&) to avoid superfluous SmallList copies - Made internally created matrices static - Change accelerator dynamic function vectors to static arrays - Exit all validators early when validation is disabled Additional cleanup includes: - Tidied accelerator macros (replaced param-specific macros like "numCtrls" and "numTargs" with "param") - Fill ctrlStates vectors with default before localiser - Renamed getBitsFromInteger to setToBitsOfInteger - Adopted const in bitwise.hpp to better express intent Note that the naming of SmallList and SmallView will be subsequently changed to List64 and ConstList64 * Renamed debug API functions to contain "QuEST" (#752) * Renamed environment variables to begin with"QUEST" (#755) * Renamed CMake vars and preprocessors (#756) such that they all begin with QUEST, but some have additional changes * Renamed Small(List|View) to (Const)List64 (#757) * Defer Catch2 test discovery so that we can compile MPI tests on systems which cannot actually run with MPI, because they are missing an MPI or UCX library file, as is witnessed in the CI (when compiling with MPICH). It's generally irksome too to trigger an execution of the test binary (which itself initialises QuEST) during build when on a HPC platform with distinct submit and compute nodes * Enable user to take ownership of MPI (#722) * Added ENABLE_SUBCOMM build option * Moved from MPI_COMM_WORLD to mpiQuestComm * Decided passing *MPI_Comm was probably overly cautious, and updated function name to comm_getMpiComm * environment.cpp: added methods to reset rank and numNodes, and reporting for subcomm compiled * comm_config.hpp/cpp: added comm_setMpiComm * CMakeLists.txt: PUBLIC MPI::MPI_CXX turned out to be unhelpful, even for SubComm, because of course it enforces CXX * Added new custom QuESTEnv initialiser which allow user to positively declare that they take ownership of MPI * validation.cpp: updated comm_end call * comm_config.hpp: added config.h include so COMPILE_MPI is actually defined * subcommunicator.h/cpp: implemented QuESTEnv initialiser with custom MPI_Comm * CMake: added subcommunicator.cpp * comm_config.hpp: added missing config.h include... * comm_config.cpp: explicitly initialise mpiCommQuest to MPI_COMM_NULL, updated setComm for init only workflow * quest.h: added subcommunicator header * CMake: added MPI to application binaries when SUBCOMM is enabled * comm_routines.cpp: post Irecv before Isend which probably won't do anything but it makes MPI library implementers less nervous * tests: added new env test for initCustomMpiQuESTEnv * Added error throws to comm_config to cover new scenarios of badness with user owned MPI * subcommunicator.cpp: updated var names to match QuEST style * tests/unit/initialisations.cpp: slightly modified setQuregAmps test to avoid unexpected test failure due to range checking when compild in Debug configuration * Updated validation in comm_setMpiComm Co-authored-by: iarejula-bsc <inigo.arejula@bsc.es> * userOwnsMpi int->bool * comm_config.cpp: corrected call to MPI_Comm_free * subcommunicator.cpp: userOwnsMpi int->bool * subcommunicator.cpp: added comm_isInit guard around comm_setMpiComm * environment.cpp: USER_OWNS_MPI -> userOwnsMpi * comm_init: fixed case where useDistrib = 0 and userOwnsMpi = true * comm_init: moved (recently) misplaced MPI_Init * AUTHORS.txt: added iarejula-bsc * Added placeholder docstrings to new initialisers * docs/cmake.md: added ENABLE_SUBCOMM to list of QuEST CMake vars * Newly added COMPILE_MPI -> QUEST_COMPILE_MPI * ENABLE_SUBCOMM -> QUEST_ENABLE_SUBCOMM * CMake: corrected OpenMP and subcommunicator pre-processor definitions --------- Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com> Co-authored-by: iarejula-bsc <inigo.arejula@bsc.es> * Add flush and sync around prints (#763) to reduce the likelihood of users printing from non-root nodes interrupting QuEST root output. This is not bullet-proof; we sync the active communicator rather than MPI_COMM_WORLD so the user-controlled non-participating processes may still be printing. Furthermore, even if all processes participate, some may have outstanding non-root prints that are not aggregated to the user screen by the time MPI_Barrier finishes. But these syncs greatly reduce the change of corruption, and are effectively free! * Add GPU-aware MPICH detection This enables CRAY MPICH platforms to leverage GPU-awareness, greatly accelerating distributed GPU simulation Co-authored-by: JPRichings <james.richings@ed.ac.uk> * Cleanup custom MPI flow (#762) Important changes: - permit user initialisation of MPI when QuEST is not distributed - changed QuESTEnv fields bool from int (e.g. isMultithreaded) - add user-input validation for custom MPI calls - disambiguated comm_config.cpp concepts of "MPI is initialised" (comm_isMpiInit) from "QuEST communication is active" (comm_isActive) - refactored comm_config.cpp flow, especially related to pre-quest-init flow (during validation) - added Oliver's custom-MPI examples (from #712) - moved new API functions to experimental.h - tweaked reportQuESTEnv output grouping * Added user-control of GPU num threads per block (#736) Added: - QUEST_DEFAULT_NUM_GPU_THREADS_PER_BLOCK CMake option - QUEST_DEFAULT_NUM_GPU_THREADS_PER_BLOCK environment variable - setQuESTNumGpuThreadsPerBlock() API function - getQuESTNumGpuThreadsPerBlock() API function - set_num_gpu_threads examples in examples/extended --------- Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com> Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com> * Fix compiler warnings (#770) Beware this included removing the superfluous `numControls` argument from the C++only `std::vector` overload of `applyMultiStateControlledCompMatr2`, which is technically a teeny tiny API break ¯\_(ツ)_/¯ * tests/unit/debug.cpp: updated setQuESTSeeds validation tests to include new validation (#771) Updated number of seeds test to use a valid pointer and added a separate NULL pointer test. * Fix Windows CI test_free.yml: added Release config to ctest commands (#773) --------- Co-authored-by: D-Exposito <dexposito@cesga.es> Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com> Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com> Co-authored-by: iarejula-bsc <inigo.arejula@bsc.es> Co-authored-by: JPRichings <james.richings@ed.ac.uk> * Version bump to 4.3 * LICENCE.txt: updated year * Optimised away Thrust qubit-list allocations As part of unitaryHACK26 (challenge #783), the Thrust backend has been optimised in the small-Qureg regime, by removing superfluous device-memory allocations due to passing qubit lists. These were seen exclusively in the multiQubitProjector functions. The CPU multiQubitProjector functions were also incidentally optimised by removing getValueOfBits() from the hot loops, which itself contained a run-time loop bounded by a now defunct template parameter. * Optimised away all remaining GPU qubit-list allocations (#739) This accelerates the QuEST GPU backend in the small-Qureg regime, where device memory allocation costs become non-negligible. Also commit also pins the Windows CI version to 2022, to circumvent a CI regression. * Added Qureg checkpointing via ADIOS2 (#780) Added new functions saveQuregToFile() and createQuregFromFile(), which are optionally enabled through ADIOS2 with new CMake option QUEST_ENABLE_ADIOS2. Ashmit added the core functionality, doc and prototyped tests. Tyson added input validation, bug fixes, and auto ADIOS2 downloading. --------- Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com> * Optimised bitwise ops with BMI2 PEXT/PDEP (#796) When the new QUEST_ENABLE_BMI2 option is ON, the (new) bitwise functions getValueOfPossiblySortedBits() and (a new overload of) insertBitsWithMaskedValues() will use the BMI2 PEXT and PDEP instructions. This accelerates the CPU amplitude indexing logic, likely visible when simulation small Quregs, where memory costs do not dominate. The GPU backend does not make use of these instructions. This PR was the winning submission to UnitaryHACK 2026, challenge #717. This PR also made changes to QUEST_OUTPUT_LIB_NAME when QUEST_APPEND_CONFIG_TO_LIB_NAME is ON: - "+mt" was renamed to "+omp" - "+bmi2", "+adios2", and "+subcomm" were added (missed in previous commits) PoJen added the intrinsics, performed all benchmarking, and prototyped the extensive cpu_subroutines.cpp changes. Tyson refactored, and updated the build, macros and CI. --------- Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com> * Updated authorlist * Refactored compile-time-variable optimisations to make use of a safer and more readable inlined templated function, rather than an unsightly macro. Addresses #802 * Extended tests to 8-qubit edgecases (#808) to (only partially) address #801 * Devel (#822) * add AMD bug guard (#817) * Doc all the things (#811) * Move API and Tests doc links to top-level * Complete decoherence doc except the C++ overloads * decoherence doc touchups * added seeding doc * added validation-handling doc * added set-reporter doc * added GPU cache doc * hide C++ overloads doc since they clutter the API doc pages, and C++ users will probably see them in autocomplete, else can easily intuit them (or just use the C API) * rename matr to matrix in operators.h signatures and fixed some other typos * added operations.h multi-state-controlled doc stubs by redirection to the non-controlled function, and the multi-state-controlled CompMatr1. Also added stubs for some base functions * added notes about target ordering efficiency * remove "more" from "for more information" my brain is poop * added singly-controlled doc stubs by redirecting to controlled compmatr1 * remove ctrl and targ abbreviations from operations.h So that the pasted/duplicated redirections are always correct, and also for reader clarity / consistency * added multi-controlled operator doc stubs by redirection to applyMultiControlledCompMatr1 * added SuperOp and KrausMap struct doc * added note about epsilon-dependent struct fields * linked SuperOp and KrausMap init examples * added matrix struct doc * fix matrices.h typos * added doc for apply(Multi)(State)ControlledCompMatr1 since they are the "main, core" doc'd operations to which all other operators are currently hyperlinked * doc channel non-inline funcs * doc channel inline funcs * fix debug doc * fix hyperlink doxygen's inconsistent discovery of the inline links is concerning * bah * replace local links * added env doc * doc'd initState funcs * patch init-API examples which reveals quite a terrible user pitfall - that the init functions do not modify the user-touchable Qureg.cpuAmps when the Qureg happens to be distributed! * doc setQuregAmps which is the "main" function of its family, and ergo has the most sophisticated doc * doc setDensityQureg(Flat)Amps * V4.3 final updates (#821) * README.md: added QASC to Related Projects * Updated version number in CMake and Doxygen --------- Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com> --------- Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com> Co-authored-by: Erich Essmann <eessmann@users.noreply.github.com> Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com> Co-authored-by: Vasco Ferreira <140641650+vaferreiQMT@users.noreply.github.com> Co-authored-by: Maurice Jamieson <maurice.jamieson@ed.ac.uk> Co-authored-by: Maurice Jamieson <m.jamieson@epcc.ed.ac.uk> Co-authored-by: D-Exposito <dexposito@cesga.es> Co-authored-by: iarejula-bsc <inigo.arejula@bsc.es> Co-authored-by: JPRichings <james.richings@ed.ac.uk> Co-authored-by: Amon K. <amon06251994@gmail.com> Co-authored-by: Ashmit JaiSarita Gupta <43639341+ashmitjsg@users.noreply.github.com> Co-authored-by: PoJen Wang <62788145+nez0b@users.noreply.github.com>
Profile and optimise-away small GPU allocations
Closes #749
Summary
The single-GPU backend copied the qubit-index list from host to device
(
cudaMalloc+cudaMemcpyAsync+cudaFree) on every call of severalmulti-qubit operations, via the
getDevInts()helper. For small Quregs thisfixed allocation latency dominates the actual kernel runtime — exactly the
overhead issue #749 asks us to profile and remove.
This PR eliminates that per-call copy for the two operations agreed in scope,
following the register-resident-bitmask pattern already used by
thrust_statevec_calcExpecAnyTargZ_sub:applyMultiQubitProjectorthrust_{statevec,densmatr}_multiQubitProjector_subcalcProbOfMultiQubitOutcomethrust_{statevec,densmatr}_calcProbOfMultiQubitOutcome_subAll changes are confined to a single file:
quest/src/gpu/gpu_thrust.cuh(52 insertions, 52 deletions). No public API, dispatch, or
thrust_*signatureschange — only the internals.
Design
Projectors — bitmask reformulation (removes the copy for all sizes)
A projector keeps an amplitude iff its target qubits match the requested
outcomes. The original device-array test
getValueOfBits(n, targetsPtr, numBits) == retainValueis mathematically identical to the register-only primitive
where
qubitMask = util_getBitMask(qubits)flags the target positions andvalueMask = util_getBitMask(qubits, outcomes)holds the desired outcome bits.Both masks are plain
qindexscalars passed as kernel arguments, so no devicearray is allocated at all. For the density-matrix functor the same masks are
applied to both the row and column substates:
Outcome probability — pass the list by value (removes the copy for all sizes)
calcProbOfMultiQubitOutcomeusesfunctor_insertBits, whoseinsertBitsWithMaskedValuesis a scatter that needs the actual sorted qubitpositions (a bitmask is insufficient). Instead of allocating a
device_vector, the functor now stores the positions in aList64by value— a trivially-copyable, fixed-size (
int[64]), CUDA-kernel-compatible structthat already exists in the codebase precisely for this purpose
(
quest/src/core/lists.hpp). The list rides along as a kernel argument, socudaMalloc/cudaMemcpydisappear for every qubit count (the previous code onlyavoided the copy never — it always called
getDevInts).getDevInts()itself is untouched and still used by ~20 other operations ingpu_subroutines.cppthat are out of scope for this issue.Results (RTX PRO 6000 Blackwell, CUDA 13.0, sm_120)
Nsight Systems — CUDA API call counts
Trace of N = 8…12, 200 reps each, both operations, captured with
Nsight Systems 2025.3.2 (
nsys profile+ per-call CUDA runtime API counts):cudaMalloccudaFreecudaMemcpyAsynccudaLaunchKernelThe ~2000 eliminated allocations are exactly the two per-call qubit-list copies
(one in the projector, one in the probability path). The residual ~1000
cudaMallocin the optimised build isthrust::reduce's own internal temporaryin
calcProb— inherent to Thrust and out of scope.cudaLaunchKernelisunchanged, confirming the kernels themselves are untouched.
Wall-clock per call (microseconds),
numTargs = 3, 2000 repsIn the small-Qureg regime the projector is consistently ~1.8× faster and the
probability calc ~1.4× faster, purely from removing the allocation. (The
large baseline jump at N≥17 is the allocation interacting with the now-larger
state kernels; removing it also removes that cliff.)
Correctness
Built with
-D QUEST_ENABLE_CUDA=ON -D QUEST_BUILD_TESTS=ON -D CMAKE_CUDA_ARCHITECTURES=120 -D CMAKE_BUILD_TYPE=Releaseagainst CUDA 13.0.The unit tests for the affected operations pass across all four deployments
(CPU, CPU+OpenMP, GPU, GPU+OpenMP):
(A correct sm_120 build is also implicitly validated, since a wrong architecture
silently corrupts GPU results.)
Re-verified end-to-end on a freshly re-cloned
devel(HEADb9830592) with thesingle-file change re-applied: configure, build, and the test command above all
pass unchanged.
How to reproduce the measurements
The numbers above come from a small standalone driver (kept out of this PR to
preserve the single-file diff) that, against both a clean
origin/develbuild andthis branch:
-D QUEST_ENABLE_CUDA=ON -D CMAKE_CUDA_ARCHITECTURES=120 -D CMAKE_BUILD_TYPE=Release;N, allocates a Qureg, then times a loop ofapplyMultiQubitProjectorand
calcProbOfMultiQubitOutcome(numTargs = 3) with a CUDA-synchronisedwall clock over many reps (per-call µs in the table);
nsys profileand tallies CUDA runtime API calls(
cudaMalloc/cudaFree/cudaMemcpyAsync/cudaLaunchKernel) for the table above.Happy to share the driver/scripts separately if useful for CI; they are not part of
this change.
Notes for reviewers
devel(the active unitaryHACK branch and the only one thatbuilds on CUDA 13 —
main/v4.2 still usesthrust::binary_function, removedin CUDA 13's libcu++).
cuQuantum is not enabled). cuStateVec is a separate backend, not what Profile and optimise-away small GPU allocations #749
concerns; it was left disabled for these measurements. (For the record, as of
2026 cuStateVec does support CUDA 13 and Blackwell, so this is a scope choice,
not a compatibility workaround.)
getDevIntsingpu_thrust.cuh.AI usage disclosure
Per unitaryHACK's AI guide ("human-in-the-loop";
honesty required for bounty eligibility): an AI coding assistant (Anthropic Claude,
via Claude Code) was used as a co-pilot for parts of this work — to help survey the
gpu_thrust.cuhcode paths, brainstorm the bitmask/List64-by-value reformulation,and draft this PR description and the profiling methodology. It
was not the author of record: every change was reviewed, compiled, and tested by
me on real hardware (RTX PRO 6000 Blackwell, CUDA 13.0, sm_120). The diff is a single
file (+52/−52), the algebraic equivalence of the bitmask reformulation was checked by
hand, and correctness was confirmed by the upstream unit tests passing across all four
deployments (CPU / CPU+OpenMP / GPU / GPU+OpenMP). No unverified or copy-pasted AI
output is included.
unitaryHACK 2026 checklist
Closes #749).