Sitelet https://github.com/QuEST-Kit/QuEST/pull/796
Skip to content

Optimise getValueOfBits and insertBits with BMI2 PEXT/PDEP (#717) - #796

Merged
TysonRayJones merged 16 commits into
QuEST-Kit:develfrom
nez0b:bmi2-bitwise-717
Jun 28, 2026
Merged

TysonRayJones merged 16 commits into
QuEST-Kit:develfrom
nez0b:bmi2-bitwise-717

Conversation

@nez0b

@nez0b nez0b commented Jun 15, 2026

Copy link
Copy Markdown
Contributor

Closes #717.

Summary

This adds guarded x86 BMI2 fast paths for QuEST's hot bit gather/scatter helpers, it hoists the loop-invariant work out of the 2^N statevector loops so each call collapses to a single instruction:

  • insertBitsWithMaskedValues(...) (the gate-application path): I build the position mask once per
    gate
    and use _pdep_u64, instead of rebuilding it on every amplitude.
  • getValueOfBits(...) (measurement / diagonal-matrix / projector): _pext_u64 when the qubits are
    strictly increasing, falling back to the original loop otherwise.
  • CMakeLists.txt adds a QUEST_ENABLE_BMI2 option (off by default): a default build stays
    portable scalar — no BMI2 in the binary, so it can't SIGILL on a pre-BMI2 CPU — and
    -DQUEST_ENABLE_BMI2=ON wires -mbmi2 (PEXT/PDEP) through the library (which then requires a
    BMI2-capable CPU at runtime).

The earlier PRs rebuilt the mask / re-checked sortedness inside
the per-amplitude function; here those are loop-invariant and computed once.

A note on impact. In practice this is a modest optimisation, and I want to be upfront about
that. The per-call speedup is large (6–12× in cache), but non-trivial state-vector simulation is
memory-bandwidth-bound, so most of that win is hidden: end-to-end I measure only ~1.0–1.3×
single-threaded, collapsing toward 1× (occasionally below) once the state is DRAM-bound or the run
saturates memory bandwidth across threads. It is probably not a large win for
big production runs.

Benchmarks

"Per call" means one invocation of the bit gather/scatter helper (getValueOfBits /
insertBitsWithMaskedValues). The plot below isolates that instruction. it's the bit-op run in a tight
loop with the result kept in a register, no statevector traffic.

Pure bit-op speedup: PDEP/PEXT are about 6–12× faster than the scalar loops in cache (k=6, single thread),
compressing to ~4–7× at the DRAM wall (one thread can't saturate memory):

pr_speedup

That isolated number is the ceiling, not what a gate sees — inside a real kernel the bit-op only
computes an index, and the strided complex-amplitude load/store it gates dominates the cost (see Thread
scaling below). The three views form one memory-roofline ladder — isolated instruction ~6–12× →
inside a gate ~2.5× → whole circuit ~1.0–1.3×
— each diluted by progressively more memory traffic.

End-to-end on whole circuits (AI assisted benchmark generation) (Xeon 6448H, single thread, all bit-identical to the scalar build):

circuit nq=12 nq=16
QFT (controlled-phase ladder) 1.26× 1.01×
random / supremacy 1.16× 1.27×
Grover (many-controlled) 1.17× 1.27×
VQE + measurement 1.15× 1.28×

So the end-to-end win is modest — roughly 1.0–1.3× single-threaded, biggest where index math
dominates (controlled-phase-heavy circuits) on cache-resident states, and it shrinks toward 1× (and
occasionally below — e.g. 0.84× at nq=20 / 32 threads) once the state is DRAM-bound or threads
saturate bandwidth.

Thread scaling — the bit-op inside a real gate kernel (Xeon 6448H, L3-resident q=18). Unlike the
isolated plot above, each call here also reads/writes the strided complex amplitudes it indexes (~64 B of
scattered memory per call vs the plot's single sequential 8-byte read), so memory, not the instruction,
sets the pace:

threads 1 8 32 64 128
PDEP (scatter) 2.53× 1.30× 1.15× 1.11× 1.01×
PEXT (sorted gather) 3.89× 3.95× 2.52× 2.02× 1.22×

Even at 1 thread this is only 2.5–3.9× — not the ~7–8× the isolated plot shows at q=18 — because the
strided amplitude traffic already dominates; it then erodes toward ~1× by 128 threads as bandwidth
saturates. (The 1–32-thread runs are pinned to one idle-gated socket; 64/128 spill across this shared
node's other sockets, so read those two columns as trend rather than precise figures.)

I also tried VPSHUFBITQMB (AVX-512)

Because getValueOfBits can be handed unsorted qubits, I benchmarked the AVX-512 arbitrary-order
gather (VPSHUFBITQMB, BITALG) as an alternative to PEXT. Per-call, in cache (gather):

gather, in-cache ns/call vs scalar needs
scalar loop 4.36 1.0× —
PEXT (sorted only) 0.49 8.9× BMI2
VPSHUFBITQMB (any order) 0.69 6.3× AVX-512 BITALG
VPSHUFBITQMB ×8 (any order, batched) 0.41 10.6× AVX-512 BITALG + loop rewrite

I didn't use it, for three reasons: (1) it needs AVX-512 BITALG (Ice Lake-SP+ / Zen 4+), so it would
need runtime CPU dispatch rather than a compile-time guard; (2) for the common sorted case PEXT is
already faster per call; (3) the only thing it uniquely buys — the unsorted gather — is a tiny slice
of the work. Instrumenting the four circuit families, scatter (PDEP) outweighs gather by ~58:1,
and the unsorted part is ~0.5% of bitwise work. BMI2-only keeps the diff small and portable for the
whole win.

Tests

  • tests/unit/bitwise.cpp — checks the new mask-accepting helpers are bit-identical to the original
    getValueOfBits / insertBitsWithMaskedValues over exhaustive-small and randomised inputs, plus
    deterministic boundary cases at bits 31/32/61/62/63 (incl. the int64 sign bit), and
    isStrictlyIncreasing. Built with -DQUEST_ENABLE_BMI2=ON it's compiled with -mbmi2, so it
    exercises the real PEXT/PDEP path; 41,849 assertions pass (and pass equally on the scalar fallback).
  • examples/automated/benchmark_bitwise_bmi2.cpp — a quick benchmark that prints the timings and
    which path was compiled in
    , so a non-x86 / non--mbmi2 CI runner just prints the scalar timings
    rather than hitting SIGILL. Runs in well under a second.

AI assistance

I used an AI assistant to help survey the candidate instructions (PEXT/PDEP, VPSHUFBITQMB, GFNI), draft
the hoisting and the compile guards, and assemble the benchmarks. I reviewed the diff myself, kept the
arbitrary-order semantics of getValueOfBits, and ran the validation above before opening this.

Closes #717

nez0b and others added 2 commits June 15, 2026 20:25
…#717)

Replace the per-amplitude looped bit gather/scatter in the CPU statevector/
density-matrix kernels with x86 BMI2 PEXT/PDEP, hoisting the loop-invariant
masks out of the 2^N loops so each per-amplitude call becomes one instruction:

- insertBitsWithMaskedValues sites: compute the position mask once per gate and
  use _pdep_u64 (order-invariant scatter; unconditionally correct).
- getValueOfBits sites: _pext_u64 when the qubits are strictly increasing,
  falling back to the original scalar loop otherwise (order is preserved).

Portability: BMI2 is opt-in via a new QUEST_ENABLE_BMI2 CMake option, OFF by
default, so a default build stays portable scalar (no BMI2 in the binary, no
SIGILL on pre-BMI2 CPUs). Enabling it wires -mbmi2 through the library; the
intrinsics are additionally guarded to x86 host TUs (never CUDA/HIP device
code). The scalar fallback is byte-identical, and QUEST_BITWISE_FORCE_SCALAR
forces it on a BMI2-capable host.

Tests/benchmark: tests/unit/bitwise.cpp asserts the new helpers are
bit-identical to the originals over exhaustive-small, randomised, and boundary
inputs (bits 31/32/61/62/63 incl. the int64 sign bit); examples/automated adds
a cross-platform benchmark that prints timings and which path was compiled in.

Bit-identical to the scalar path (verified by unit tests, the QuEST suite for
the touched kernels, and amplitude hashes of QFT/random/Grover/VQE circuits).

Closes QuEST-Kit#717
@TysonRayJones

Copy link
Copy Markdown
Member

Wew nice work! 🎉 Great to see thought put into the call sites!

I have tailored the CI to run with the new instrinsic pathway, and will manually test on another Intel machine.

Interestingly, this diff reveals two kinds of callers of (insert|get)Bits() - those which know the qubit list will be unconditionally sorted (and can ergo always make use of the instrinsic when available), and those which have to check (and so can only maybe make use of the intrinsic).

  • Unconditional example (here):
     qindex qubitsPosMask = getBitMask(sortedQubits.data(), numQubitBits);
    
     for (qindex n=0; n<numIts; n++) {
    
         qindex i0 = insertBitsWithMaskedValuesAndPosMask(n, qubitStateMask, qubitsPosMask, sortedQubits.data(), numQubitBits);
  • Conditional example (here)
     qindex qubitsPosMask = getBitMask(qubits.data(), numBits);  // loop-invariant: hoisted out of the per-amplitude loop
     bool qubitsSorted = isStrictlyIncreasing(qubits.data(), numBits);  // likewise loop-invariant (order checked once per gate)
    
     for (qindex n=0; n<numIts; n++) {
         qreal prob = norm(amps[n]);
         qindex i = concatenateBits(qureg.rank, n, qureg.logNumAmpsPerNode);
    
         qindex j = (qubitsSorted ? getValueOfBitsFromSortedPosMask(i, qubitsPosMask, qubits.data(), numBits) : getValueOfBits(i, qubits.data(), numBits));

This is a very important difference. In the un-conditional case, everything can be fully inlined, and getValueOfBitsFromSortedPosMask could become a | _pdep_u64(b, ~c) at compile-time. But in the conditional case, we have a tension. We can either...

  • (current design) Decide whether or not to (attemptedly) leverage the intrinsic within each iteration of the loop. This ternary introduces branching in a hot-loop, though since its value is fixed (which we can make more explicit using const), I believe that either...
    • The compiler will unroll the entire loop into two loops, with distinct bodies based on qubitsSorted.
    • Otherwise, branch-prediction will become so simple/quick, that after a few iterations, the ternary will be free.
  • unroll the loop ourselves (very ugly)
  • prior resolve a function pointer (disables inlining)
  • use some templating trick
  • unconditionally sort the qubits and change the algorithm logic

I believe the current design is actually simultaneously the cleanest and best performing (when the compiler isn't stoopid, and when the function doesn't have a natural always-sorted formulation). I'm not above mostly for my own reference/history!

I will make some cleanup changes to this PR, and possibly make an extension (e.g. an O(1) fallback alternative to the instrinsic). No more changes are necessary for unitaryHACK however! 🎉 Please comment on issue #717 so I can assign it to you!

Comment thread quest/src/cpu/cpu_subroutines.cpp Outdated
Comment thread quest/src/core/bitwise.hpp Outdated
Comment thread quest/src/core/bitwise.hpp Outdated
Comment thread quest/src/core/bitwise.hpp Outdated
Comment thread quest/src/cpu/cpu_subroutines.cpp Outdated
Comment thread quest/src/cpu/cpu_subroutines.cpp Outdated
Comment thread tests/unit/bitwise.cpp Outdated
Comment thread tests/unit/CMakeLists.txt Outdated
Comment thread CMakeLists.txt Outdated
Comment thread CMakeLists.txt Outdated
Comment thread CMakeLists.txt
Comment thread quest/src/core/bitwise.hpp Outdated
since it is pointless to attempt to compile-time-guard against BMI2 intrinsics; that will not avoid the error when precompiling on one system (upon which the compiler recognises BMI2), and running on another (where the instruction is not recognised by the CPU)

So, when QUEST_COMPILE_BMI2 is set,  compilation should assume the intrinsic exists; when it does not, let compilation failed. Such a circumstance should be impossible anyway, since the CMake build now performs a compile probe
which moves the is-sorted checking inside the bitwise function. This should not preclude the loop/branch unrolling, since forcefully inlined
Comment thread quest/src/cpu/cpu_subroutines.cpp Outdated
changed to be consistent with mask preparation in other functions; above SET_VAR_AT_COMPILE_TIME, and using util_getBitMask() rather than the lower-level bitwise.hpp function
…lues overload

and fixed messy mask preparation placement (as well as use of bitwise.hpp getBitMask with util_getBitMask)
@TysonRayJones TysonRayJones mentioned this pull request Jun 28, 2026
@TysonRayJones TysonRayJones mentioned this pull request Jun 28, 2026
since we reached 540 configs, exceeding the github max of 256
@TysonRayJones

Copy link
Copy Markdown
Member

I just accidentally nailed the Github Actions matrix size limit of 256! 🎉 Life's little treats

@TysonRayJones
TysonRayJones merged commit 3a85f47 into QuEST-Kit:devel Jun 28, 2026
311 checks passed
otbrown added a commit that referenced this pull request Sep 24, 2026
* gpu_thrust.cuh: removed thrust::[unary|binary]_function to fix compatibility with CUDA 13. Fixes #693.

* Simplify installation path configuration (#697)

* Simplify installation path configuration

Removed unnecessary path normalization and appending for installation.

* Updates CMake config for conditional installs

Modifies CMakeLists to conditionally build shared libraries and
install binaries only at the top-level project. Introduces the
INSTALL_BINARIES option to control the inclusion of example
binaries in the installation process. Corrects a typo from
'RATH' to 'RPATH' for build configurations.

* Doc binary install.  (#701)

* docs/cmake.md: fixed formatting of non-default options for mt and distribution

* cmake: wrapped user source install in if(INSTALL_BINARIES)

* docs/cmake.md: added INSTALL_BINARIES option

* Thrust Overflow Fix (#699)

* gpu_thrust.cuh: modified initial thrust counting iterator declarations to use long long to avoid overflow at >30 qubits. Fixes #698.

* patched test of rightapplyCompMatr distributed validation

The operation validation tests previously always uses a statevector to test the "targeted amps fit in node" validation, though the rightapply*() functions cannot accept statevectors, instead only density matrices. Because the "was given a density matrix" validation happens before "targeted amps fit in node" validation, the latter intended triggered error was beaten out by the earlier unintended one.

Now, we are careful to pass a density matrix Qureg to the validation of "targeted amps fit in node" when triggered by a function which 'right-applies' (and is ergo only compatible with density matrices)

* changed literals to defensive type

---------

Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com>

* patched validation test of rightapplyCompMatr (#703)

addresses #700

* Random permutation of strings in Trotterisation (#702)

* Implement PauliStrSum random permutations inspired by [arXiv:1805.08385](https://arxiv.org/abs/1805.08385)

* Add randomisation to Trotter functions

* Document random Pauli permutations for Trotterisation

* Add test for Trotter randomisation

* Added unit tests requested by Quantum Motion (#706)

* Added unit tests requested by Quantum Motion

* tests/unit/trotterisation.cpp: updated time evo calls to new API

* tests/unit/trotterisation.cpp: updated authorlist

* Fixed valgrind errors

* tests/unit/trotterisation.cpp: tuned floating-point comparison epsilon to account for worst-case scenario which is single precision, single thread

* added get-arbitrary-qureg test util

since it will be used frequently by new input validation

* removed Qureg creation in validation tests

so that test failures do not cause memory leaks and e.g. add to valgrind noise. Tests now instead use getArbitraryCachedStatevec() or getArbitraryCachedDensmatr() to obtain an existing qureg with an arbitrary deployment.

* restoring missing-validation comments

since the validation for these functions wasn't added. Such functions have additional tests to their tested counterparts; for example, validating that matrix elements are non-zero when given a negative exponent

* fixing test category

* added missing operation validation tests

* fixed indentation

* making spacing consistent

and adding a missing Hermiticity validation to applyTrotterizedUnitaryTimeEvolution test

* added warning about untested deployments

* removed defunct signature

* patching C++ validation err msg

Previously, an error message of the C++ API was not substituting in values for its placeholder variables. This affected the C++ variants of the below functions when passing vectors for the targets and outcome parameters of mismatching length:
- calcProbOfMultiQubitOutcome
- leftapplyMultiQubitProjector
- rightapplyMultiQubitProjector
- applyMultiQubitProjector
- applyForcedMultiQubitMeasurement

* added missing C++ API signatures

* added C++-API validation tests

* updated doc warnings

* added Vasco to Trotter API authorlist

* merged Tyson's patches

---------

Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com>
Co-authored-by: Maurice Jamieson <m.jamieson@epcc.ed.ac.uk>
Co-authored-by: Oliver Thomson Brown <otbrown@users.noreply.github.com>
Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com>

* Add inverse to QFT implementation (#705)

* Fix missing inverse argument to QFT validation tests

* Implement Pauli sorting (#708)

* Fix HIP install failures in CI (#718)

CI updated to use latest AMD ROCm install instructions. As of this commit corresponding to ROCm 7.2.

---------

Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com>

* Fix applyMultiStateControlledSqrtSwap argument list (#738)

Remove numControls argument from applyMultiStateControlledSqrtSwap overloaded definition taking std::vector<int>

(cherry picked from commit 9c20792)

Co-authored-by: D-Exposito <dexposito@cesga.es>

* Trotterisation test update (#728)

* tests/unit/trotterisation.cpp: updated to use REQUIRE_AGREE and cached statevecs and densmats, and both permutePaulis options

* tests/utils/compare.hpp/cpp: added setters for test epsilon

* tests/unit/trotterisation.cpp: adjusted test epsilon for quad precision imaginary time evolution tests

* tests/unit/trotterisation.cpp: moved unitary time evo test to REQUIRE_AGREE

* tests/utils/cache.hpp/cpp: added additional utilities for creating and destroying temp caches (which I guess makes them not caches?) with a set number of qubits

* tests/unit/trotterisation.cpp: updated unitary time evo test to test across deployments

* tests/unit/trotterisation.cpp: reduced number of qubits and increased number of steps to admit the possibility of testing density matrices too

* tests/unit/trotterisation.cpp: added density matrix tests

* reduce test precision

to lazily pass CPU clang quad-precision

* skip Trotter tests in paid CI

* changing varname convention

* renaming cache funcs

---------

Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com>
Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com>

* added Daniel Patino to authorlist

* CMake warn when non-release build (#742)


---------

Co-authored-by: Oliver Thomson Brown <otbrown@users.noreply.github.com>

* Stop Trotter funcs mutating PauliStrSum (#740)

Formerly, the Trotter functions (such as applyTrotterizedPauliStrSumGadget()), when passed permutePaulis=true, would randomly permutate the order of the passed PauliStrSum, mutating it and affecting the outputs of subsequent functions like reportPauliStrSum(). The function also contained superfluous memory allocs/copies equal in size to the PauliStrSum.

Now, the PauliStrSum is never mutated, and an internally allocated ordering list keeps track of the randomised permutation. We also updated the doc, renamed permutePaulis to permuteTerms, and improved validation. Note that 'permuteTerms' had not yet reached main/release, so these changes do not need to be documented in the v4.3 release notes.

* Created custom backend complex types (#729)

Created cpu_qcomp and gpu_qcomp (from a shared base_qcomp) to avoid std::complex arithmetic operators in hot loops which caused performance issues. Removed all prior compiler flags and related scaffolding attempting to mitigate the performance issue.

Also gave MSVC build the params `/Zc:preprocessor -Xcompiler=/Zc:preprocessor /bigobj` as needed for compilation of the unit tests on my windows machines.

* Replace vector<int> with SmallList (a stack array) (#743)

This is to circumvent the std::vector performance overheads visible in few-qubit simulation (responsible for a performance regression from v3; see #720), and also so that qubit lists can be passed directly to CUDA kernels without conversion (as explored in #739).

* Added few-qubit optimisations (#750)

Optimisations include:
- Adopted SmallView (const SmallList&) to avoid superfluous SmallList copies
- Made internally created matrices static
- Change accelerator dynamic function vectors to static arrays
- Exit all validators early when validation is disabled

Additional cleanup includes:
- Tidied accelerator macros (replaced param-specific macros like "numCtrls" and "numTargs" with "param")
- Fill ctrlStates vectors with default before localiser
- Renamed getBitsFromInteger to setToBitsOfInteger
- Adopted const in bitwise.hpp to better express intent

Note that the naming of SmallList and SmallView will be subsequently changed to List64 and ConstList64

* Renamed debug API functions to contain "QuEST" (#752)

* Renamed environment variables to begin with"QUEST" (#755)

* Renamed CMake vars and preprocessors (#756)

such that they all begin with QUEST, but some have additional changes

* Renamed Small(List|View) to (Const)List64 (#757)

* Defer Catch2 test discovery

so that we can compile MPI tests on systems which cannot actually run with MPI, because they are missing an MPI or UCX library file, as is witnessed in the CI (when compiling with MPICH). It's generally irksome too to trigger an execution of the test binary (which itself initialises QuEST) during build when on a HPC platform with distinct submit and compute nodes

* Enable user to take ownership of MPI (#722)

* Added ENABLE_SUBCOMM build option

* Moved from MPI_COMM_WORLD to mpiQuestComm

* Decided passing *MPI_Comm was probably overly cautious, and updated function name to comm_getMpiComm

* environment.cpp: added methods to reset rank and numNodes, and reporting for subcomm compiled

* comm_config.hpp/cpp: added comm_setMpiComm

* CMakeLists.txt: PUBLIC MPI::MPI_CXX turned out to be unhelpful, even for SubComm, because of course it enforces CXX

* Added new custom QuESTEnv initialiser which allow user to positively declare that they take ownership of MPI

* validation.cpp: updated comm_end call

* comm_config.hpp: added config.h include so COMPILE_MPI is actually defined

* subcommunicator.h/cpp: implemented QuESTEnv initialiser with custom MPI_Comm

* CMake: added subcommunicator.cpp

* comm_config.hpp: added missing config.h include...

* comm_config.cpp: explicitly initialise mpiCommQuest to MPI_COMM_NULL, updated setComm for init only workflow

* quest.h: added subcommunicator header

* CMake: added MPI to application binaries when SUBCOMM is enabled

* comm_routines.cpp: post Irecv before Isend which probably won't do anything but it makes MPI library implementers less nervous

* tests: added new env test for initCustomMpiQuESTEnv

* Added error throws to comm_config to cover new scenarios of badness with user owned MPI

* subcommunicator.cpp: updated var names to match QuEST style

* tests/unit/initialisations.cpp: slightly modified setQuregAmps test to avoid unexpected test failure due to range checking when compild in Debug configuration

* Updated validation in comm_setMpiComm

Co-authored-by: iarejula-bsc <inigo.arejula@bsc.es>

* userOwnsMpi int->bool

* comm_config.cpp: corrected call to MPI_Comm_free

* subcommunicator.cpp: userOwnsMpi int->bool

* subcommunicator.cpp: added comm_isInit guard around comm_setMpiComm

* environment.cpp: USER_OWNS_MPI -> userOwnsMpi

* comm_init: fixed case where useDistrib = 0 and userOwnsMpi = true

* comm_init: moved (recently) misplaced MPI_Init

* AUTHORS.txt: added iarejula-bsc

* Added placeholder docstrings to new initialisers

* docs/cmake.md: added ENABLE_SUBCOMM to list of QuEST CMake vars

* Newly added COMPILE_MPI -> QUEST_COMPILE_MPI

* ENABLE_SUBCOMM -> QUEST_ENABLE_SUBCOMM

* CMake: corrected OpenMP and subcommunicator pre-processor definitions

---------

Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com>
Co-authored-by: iarejula-bsc <inigo.arejula@bsc.es>

* Add flush and sync around prints (#763)

to reduce the likelihood of users printing from non-root nodes interrupting QuEST root output. This is not bullet-proof; we sync the active communicator rather than MPI_COMM_WORLD so the user-controlled non-participating processes may still be printing. Furthermore, even if all processes participate, some may have outstanding non-root prints that are not aggregated to the user screen by the time MPI_Barrier finishes. But these syncs greatly reduce the change of corruption, and are effectively free!

* Add GPU-aware MPICH detection

This enables CRAY MPICH platforms to leverage GPU-awareness, greatly accelerating distributed GPU simulation

Co-authored-by: JPRichings <james.richings@ed.ac.uk>

* Cleanup custom MPI flow (#762)

Important changes:
- permit user initialisation of MPI when QuEST is not distributed
- changed QuESTEnv fields bool from int (e.g. isMultithreaded)
- add user-input validation for custom MPI calls
- disambiguated comm_config.cpp concepts of "MPI is initialised" (comm_isMpiInit) from "QuEST communication is active" (comm_isActive)
- refactored comm_config.cpp flow, especially related to pre-quest-init flow (during validation)
- added Oliver's custom-MPI examples (from #712)
- moved new API functions to experimental.h
- tweaked reportQuESTEnv output grouping

* Added user-control of GPU num threads per block (#736)

Added:
- QUEST_DEFAULT_NUM_GPU_THREADS_PER_BLOCK CMake option
- QUEST_DEFAULT_NUM_GPU_THREADS_PER_BLOCK environment variable
- setQuESTNumGpuThreadsPerBlock() API function
- getQuESTNumGpuThreadsPerBlock() API function
- set_num_gpu_threads examples in examples/extended

---------

Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com>
Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com>

* Fix compiler warnings (#770)

Beware this included removing the superfluous `numControls` argument from the C++only `std::vector` overload of `applyMultiStateControlledCompMatr2`, which is technically a teeny tiny API break ¯\_(ツ)_/¯

* tests/unit/debug.cpp: updated setQuESTSeeds validation tests to include new validation (#771)

Updated number of seeds test to use a valid pointer and added a separate NULL pointer test.

* Fix Windows CI

test_free.yml: added Release config to ctest commands (#773)

* Create v4.3 release candidate (#776)

* Fix applyMultiStateControlledSqrtSwap argument list (#738)

Remove numControls argument from applyMultiStateControlledSqrtSwap overloaded definition taking std::vector<int>

(cherry picked from commit 9c20792)

Co-authored-by: D-Exposito <dexposito@cesga.es>

* Trotterisation test update (#728)

* tests/unit/trotterisation.cpp: updated to use REQUIRE_AGREE and cached statevecs and densmats, and both permutePaulis options

* tests/utils/compare.hpp/cpp: added setters for test epsilon

* tests/unit/trotterisation.cpp: adjusted test epsilon for quad precision imaginary time evolution tests

* tests/unit/trotterisation.cpp: moved unitary time evo test to REQUIRE_AGREE

* tests/utils/cache.hpp/cpp: added additional utilities for creating and destroying temp caches (which I guess makes them not caches?) with a set number of qubits

* tests/unit/trotterisation.cpp: updated unitary time evo test to test across deployments

* tests/unit/trotterisation.cpp: reduced number of qubits and increased number of steps to admit the possibility of testing density matrices too

* tests/unit/trotterisation.cpp: added density matrix tests

* reduce test precision

to lazily pass CPU clang quad-precision

* skip Trotter tests in paid CI

* changing varname convention

* renaming cache funcs

---------

Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com>
Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com>

* added Daniel Patino to authorlist

* CMake warn when non-release build (#742)


---------

Co-authored-by: Oliver Thomson Brown <otbrown@users.noreply.github.com>

* Stop Trotter funcs mutating PauliStrSum (#740)

Formerly, the Trotter functions (such as applyTrotterizedPauliStrSumGadget()), when passed permutePaulis=true, would randomly permutate the order of the passed PauliStrSum, mutating it and affecting the outputs of subsequent functions like reportPauliStrSum(). The function also contained superfluous memory allocs/copies equal in size to the PauliStrSum.

Now, the PauliStrSum is never mutated, and an internally allocated ordering list keeps track of the randomised permutation. We also updated the doc, renamed permutePaulis to permuteTerms, and improved validation. Note that 'permuteTerms' had not yet reached main/release, so these changes do not need to be documented in the v4.3 release notes.

* Created custom backend complex types (#729)

Created cpu_qcomp and gpu_qcomp (from a shared base_qcomp) to avoid std::complex arithmetic operators in hot loops which caused performance issues. Removed all prior compiler flags and related scaffolding attempting to mitigate the performance issue.

Also gave MSVC build the params `/Zc:preprocessor -Xcompiler=/Zc:preprocessor /bigobj` as needed for compilation of the unit tests on my windows machines.

* Replace vector<int> with SmallList (a stack array) (#743)

This is to circumvent the std::vector performance overheads visible in few-qubit simulation (responsible for a performance regression from v3; see #720), and also so that qubit lists can be passed directly to CUDA kernels without conversion (as explored in #739).

* Added few-qubit optimisations (#750)

Optimisations include:
- Adopted SmallView (const SmallList&) to avoid superfluous SmallList copies
- Made internally created matrices static
- Change accelerator dynamic function vectors to static arrays
- Exit all validators early when validation is disabled

Additional cleanup includes:
- Tidied accelerator macros (replaced param-specific macros like "numCtrls" and "numTargs" with "param")
- Fill ctrlStates vectors with default before localiser
- Renamed getBitsFromInteger to setToBitsOfInteger
- Adopted const in bitwise.hpp to better express intent

Note that the naming of SmallList and SmallView will be subsequently changed to List64 and ConstList64

* Renamed debug API functions to contain "QuEST" (#752)

* Renamed environment variables to begin with"QUEST" (#755)

* Renamed CMake vars and preprocessors (#756)

such that they all begin with QUEST, but some have additional changes

* Renamed Small(List|View) to (Const)List64 (#757)

* Defer Catch2 test discovery

so that we can compile MPI tests on systems which cannot actually run with MPI, because they are missing an MPI or UCX library file, as is witnessed in the CI (when compiling with MPICH). It's generally irksome too to trigger an execution of the test binary (which itself initialises QuEST) during build when on a HPC platform with distinct submit and compute nodes

* Enable user to take ownership of MPI (#722)

* Added ENABLE_SUBCOMM build option

* Moved from MPI_COMM_WORLD to mpiQuestComm

* Decided passing *MPI_Comm was probably overly cautious, and updated function name to comm_getMpiComm

* environment.cpp: added methods to reset rank and numNodes, and reporting for subcomm compiled

* comm_config.hpp/cpp: added comm_setMpiComm

* CMakeLists.txt: PUBLIC MPI::MPI_CXX turned out to be unhelpful, even for SubComm, because of course it enforces CXX

* Added new custom QuESTEnv initialiser which allow user to positively declare that they take ownership of MPI

* validation.cpp: updated comm_end call

* comm_config.hpp: added config.h include so COMPILE_MPI is actually defined

* subcommunicator.h/cpp: implemented QuESTEnv initialiser with custom MPI_Comm

* CMake: added subcommunicator.cpp

* comm_config.hpp: added missing config.h include...

* comm_config.cpp: explicitly initialise mpiCommQuest to MPI_COMM_NULL, updated setComm for init only workflow

* quest.h: added subcommunicator header

* CMake: added MPI to application binaries when SUBCOMM is enabled

* comm_routines.cpp: post Irecv before Isend which probably won't do anything but it makes MPI library implementers less nervous

* tests: added new env test for initCustomMpiQuESTEnv

* Added error throws to comm_config to cover new scenarios of badness with user owned MPI

* subcommunicator.cpp: updated var names to match QuEST style

* tests/unit/initialisations.cpp: slightly modified setQuregAmps test to avoid unexpected test failure due to range checking when compild in Debug configuration

* Updated validation in comm_setMpiComm

Co-authored-by: iarejula-bsc <inigo.arejula@bsc.es>

* userOwnsMpi int->bool

* comm_config.cpp: corrected call to MPI_Comm_free

* subcommunicator.cpp: userOwnsMpi int->bool

* subcommunicator.cpp: added comm_isInit guard around comm_setMpiComm

* environment.cpp: USER_OWNS_MPI -> userOwnsMpi

* comm_init: fixed case where useDistrib = 0 and userOwnsMpi = true

* comm_init: moved (recently) misplaced MPI_Init

* AUTHORS.txt: added iarejula-bsc

* Added placeholder docstrings to new initialisers

* docs/cmake.md: added ENABLE_SUBCOMM to list of QuEST CMake vars

* Newly added COMPILE_MPI -> QUEST_COMPILE_MPI

* ENABLE_SUBCOMM -> QUEST_ENABLE_SUBCOMM

* CMake: corrected OpenMP and subcommunicator pre-processor definitions

---------

Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com>
Co-authored-by: iarejula-bsc <inigo.arejula@bsc.es>

* Add flush and sync around prints (#763)

to reduce the likelihood of users printing from non-root nodes interrupting QuEST root output. This is not bullet-proof; we sync the active communicator rather than MPI_COMM_WORLD so the user-controlled non-participating processes may still be printing. Furthermore, even if all processes participate, some may have outstanding non-root prints that are not aggregated to the user screen by the time MPI_Barrier finishes. But these syncs greatly reduce the change of corruption, and are effectively free!

* Add GPU-aware MPICH detection

This enables CRAY MPICH platforms to leverage GPU-awareness, greatly accelerating distributed GPU simulation

Co-authored-by: JPRichings <james.richings@ed.ac.uk>

* Cleanup custom MPI flow (#762)

Important changes:
- permit user initialisation of MPI when QuEST is not distributed
- changed QuESTEnv fields bool from int (e.g. isMultithreaded)
- add user-input validation for custom MPI calls
- disambiguated comm_config.cpp concepts of "MPI is initialised" (comm_isMpiInit) from "QuEST communication is active" (comm_isActive)
- refactored comm_config.cpp flow, especially related to pre-quest-init flow (during validation)
- added Oliver's custom-MPI examples (from #712)
- moved new API functions to experimental.h
- tweaked reportQuESTEnv output grouping

* Added user-control of GPU num threads per block (#736)

Added:
- QUEST_DEFAULT_NUM_GPU_THREADS_PER_BLOCK CMake option
- QUEST_DEFAULT_NUM_GPU_THREADS_PER_BLOCK environment variable
- setQuESTNumGpuThreadsPerBlock() API function
- getQuESTNumGpuThreadsPerBlock() API function
- set_num_gpu_threads examples in examples/extended

---------

Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com>
Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com>

* Fix compiler warnings (#770)

Beware this included removing the superfluous `numControls` argument from the C++only `std::vector` overload of `applyMultiStateControlledCompMatr2`, which is technically a teeny tiny API break ¯\_(ツ)_/¯

* tests/unit/debug.cpp: updated setQuESTSeeds validation tests to include new validation (#771)

Updated number of seeds test to use a valid pointer and added a separate NULL pointer test.

* Fix Windows CI

test_free.yml: added Release config to ctest commands (#773)

---------

Co-authored-by: D-Exposito <dexposito@cesga.es>
Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com>
Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com>
Co-authored-by: iarejula-bsc <inigo.arejula@bsc.es>
Co-authored-by: JPRichings <james.richings@ed.ac.uk>

* Version bump to 4.3

* LICENCE.txt: updated year

* Optimised away Thrust qubit-list allocations

As part of unitaryHACK26 (challenge #783), the Thrust backend has been optimised in the small-Qureg regime, by removing superfluous device-memory allocations due to passing qubit lists. These were seen exclusively in the multiQubitProjector functions.

The CPU multiQubitProjector functions were also incidentally optimised by removing getValueOfBits() from the hot loops, which itself contained a run-time loop bounded by a now defunct template parameter.

* Optimised away all remaining GPU qubit-list allocations (#739)

This accelerates the QuEST GPU backend in the small-Qureg regime, where device memory allocation costs become non-negligible.

Also commit also pins the Windows CI version to 2022, to circumvent a CI regression.

* Added Qureg checkpointing via ADIOS2 (#780)

Added new functions saveQuregToFile() and createQuregFromFile(), which are optionally enabled through ADIOS2 with new CMake option QUEST_ENABLE_ADIOS2.

Ashmit added the core functionality, doc and prototyped tests. Tyson added input validation, bug fixes, and auto ADIOS2 downloading.

---------

Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com>

* Optimised bitwise ops with BMI2 PEXT/PDEP (#796)

When the new QUEST_ENABLE_BMI2 option is ON, the (new) bitwise functions getValueOfPossiblySortedBits() and (a new overload of) insertBitsWithMaskedValues() will use the BMI2 PEXT and PDEP instructions. This accelerates the CPU amplitude indexing logic, likely visible when simulation small Quregs, where memory costs do not dominate. The GPU backend does not make use of these instructions.

This PR was the winning submission to UnitaryHACK 2026, challenge #717.

This PR also made changes to QUEST_OUTPUT_LIB_NAME when QUEST_APPEND_CONFIG_TO_LIB_NAME is ON:
  - "+mt" was renamed to "+omp"
  - "+bmi2", "+adios2", and "+subcomm" were added (missed in previous commits)

PoJen added the intrinsics, performed all benchmarking, and prototyped the extensive cpu_subroutines.cpp changes.
Tyson refactored, and updated the build, macros and CI.

---------

Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com>

* Updated authorlist

* Refactored compile-time-variable optimisations

to make use of a safer and more readable inlined templated function, rather than an unsightly macro. Addresses #802

* Extended tests to 8-qubit edgecases (#808)

to (only partially) address #801

* Devel (#822)

* add AMD bug guard (#817)

* Doc all the things (#811)

* Move API and Tests doc links to top-level

* Complete decoherence doc

except the C++ overloads

* decoherence doc touchups

* added seeding doc

* added validation-handling doc

* added set-reporter doc

* added GPU cache doc

* hide C++ overloads doc

since they clutter the API doc pages, and C++ users will probably see them in autocomplete, else can easily intuit them (or just use the C API)

* rename matr to matrix in operators.h signatures

and fixed some other typos

* added operations.h multi-state-controlled doc stubs

by redirection to the non-controlled function, and the multi-state-controlled CompMatr1. Also added stubs for some base functions

* added notes about target ordering efficiency

* remove "more" from "for more information"

my brain is poop

* added singly-controlled doc stubs

by redirecting to controlled compmatr1

* remove ctrl and targ abbreviations from operations.h

So that the pasted/duplicated redirections are always correct, and also for reader clarity / consistency

* added multi-controlled operator doc stubs

by redirection to applyMultiControlledCompMatr1

* added SuperOp and KrausMap struct doc

* added note about epsilon-dependent struct fields

* linked SuperOp and KrausMap init examples

* added matrix struct doc

* fix matrices.h typos

* added doc for apply(Multi)(State)ControlledCompMatr1

since they are the "main, core" doc'd operations to which all other operators are currently hyperlinked

* doc channel non-inline funcs

* doc channel inline funcs

* fix debug doc

* fix hyperlink

doxygen's inconsistent discovery of the inline links is concerning

* bah

* replace local links

* added env doc

* doc'd initState funcs

* patch init-API examples

which reveals quite a terrible user pitfall - that the init functions do not modify the user-touchable Qureg.cpuAmps when the Qureg happens to be distributed!

* doc setQuregAmps

which is the "main" function of its family, and ergo has the most sophisticated doc

* doc setDensityQureg(Flat)Amps

* V4.3 final updates (#821)

* README.md: added QASC to Related Projects

* Updated version number in CMake and Doxygen

---------

Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com>

---------

Co-authored-by: Oliver Thomson Brown <8394906+otbrown@users.noreply.github.com>
Co-authored-by: Erich Essmann <eessmann@users.noreply.github.com>
Co-authored-by: Tyson Jones <tyson.jones.input@gmail.com>
Co-authored-by: Vasco Ferreira <140641650+vaferreiQMT@users.noreply.github.com>
Co-authored-by: Maurice Jamieson <maurice.jamieson@ed.ac.uk>
Co-authored-by: Maurice Jamieson <m.jamieson@epcc.ed.ac.uk>
Co-authored-by: D-Exposito <dexposito@cesga.es>
Co-authored-by: iarejula-bsc <inigo.arejula@bsc.es>
Co-authored-by: JPRichings <james.richings@ed.ac.uk>
Co-authored-by: Amon K. <amon06251994@gmail.com>
Co-authored-by: Ashmit JaiSarita Gupta <43639341+ashmitjsg@users.noreply.github.com>
Co-authored-by: PoJen Wang <62788145+nez0b@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants