QuEST-Kit/QuEST: v4.3.0
Abstract
Overview This release accelerates simulation of (especially) few-qubit Quregs, adds control and performance tuning utilities for MPI and GPU superusers, adds randomised Trotter simulation, and reduces the likelihood of QuEST symbols colliding with other software stacks. Optimisations QuEST's simulation of few-qubit Qureg has been accelerated, by reducing memory transfer overheads (between function stacks, and between CPU and GPU memory), and through optional BMI2 intrinsics. QuEST's CPU and multithreaded backend has been accelerated (at all scales) on specific platforms / compilers, by switching to custom complex arithmetic. QuEST's detection of CUDA-aware MPI has been expanded to MPICH CRAY systems, greatly accelerating multi-GPU simulation. New features Added experimental initCustomMpiCommQuESTEnv() and initCustomMpiCommQuESTEnv() functions which permit users to retain control of MPI during QuEST simulation, and dedicate only a sub-communicator (e.g. only some MPI processes) to QuEST. Users can now also disable QuEST distribution through initCustomQuESTEnv() while still themselves using MPI, even when QuEST was compiled with MPI. Added an experimental setQuESTNumGpuThreadsPerBlock() function to override QuEST's GPU parallelisation granularity at runtime, permitting simple performance tuning. This accompanies a getter (get...), an environment variable QUEST_DEFAULT_NUM_GPU_THREADS_PER_BLOCK to override the parallelisation at launch time, and a CMake option of the same name to override at build-time. Added experimental functions saveQuregToFile() and createQuregFromFile(), which use ADIOS2 for Qureg checkpointing. Added sortPauliStrSumLexicographic() and sortPauliStrSumMagnitude() to reorder the terms within a PauliStrSum, affecting the numerical accuracy of Trotterisation. The Trotter functions now accept an additional permuteTerms boolean, which when true, sees every Trotter repetition randomise the ordering of the Trotter terms, often improving numerical accuracy. This affects: The QFT functions now accept an inverse boolean, to apply the inverse QFT. Improved the CMake build: Added QUEST_INSTALL_BINARIES to enable including examples and user-source in the QuEST installation directory. Added a warning when CMake configures a (unoptimised) non-release build. Removed brittle, platform-specific compiler flags (to work around a since-resolved performance problem). Added QUEST_ prefix to (almost) all CMake options, to avoid collision with other projects (see API breaks below). During compilation of the tests, the compiled executable is no longer run (for Catch2 test discovery), improving ease of compilation on supercomputers (e.g. systems with distinct job-submission and job-run nodes). API breaks As above, the Trotter functions now accept an additional permuteTerms boolean. This affects: apply(Multi)(State)(Controlled)TrotterizedPauliStrSumGadget() applyTrotterizedNonUnitaryPauliStrSumGadget() applyTrotterized(Unitary|Imaginary|Noisy)TimeEvolution() As above, the QFT functions now accept an inverse boolean. This affects: apply(Full)QuantumFourierTransform() A superfluous numControls argument was removed from the std::vector (C++ only) overloads of applyMultiStateControlledSqrtSwap() and applyMultiStateControlledCompMatr2(). All debug functions now explicitly contain the word QuEST (for example, setQuESTSeeds()) as the second word, indicated by [QuEST] below. This affects functions: (set|get)[QuEST]Seeds(ToDefault)() get[QuEST]NumSeeds() set[QuEST]InputErrorHandler() set[QuEST]Validation(On|Off)() set[QuEST]ValidationEpsilon(ToDefault)() set[QuEST]MaxNumReportedItems() set[QuEST]MaxNumReportedSigFigs() set[QuEST]NumReportedNewlines() set[QuEST]ReportedPauli(Chars|StrStyle)() get[QuEST]GpuCacheSize() clear[QuEST]GpuCache() get[QuEST]EnvironmentString() All environment variables now begin with QUEST_, as indicated by [QUEST_] below. This affects environment variables: [QUEST_]PERMIT_NODES_TO_SHARE_GPU [QUEST_]DEFAULT_VALIDATION_EPSILON [QUEST_]TEST_NUM_QUBITS_IN_QUREG [QUEST_]TEST_MAX_NUM_QUBIT_PERMUTATIONS [QUEST_]TEST_MAX_NUM_SUPEROP_TARGETS [QUEST_]TEST_NUM_MIXED_DEPLOYMENT_REPETITIONS TEST_ALL_DEPLOYMENTS has become QUEST_TEST_TRY_ALL_DEPLOYMENTS All CMake options have been renamed, to now begin with QUEST_ or USER_ (to disambiguate whether they relate to the QuEST library or the user's optional source files), and several have been made more explicit. The changes are: USER_SOURCE -> USER_SOURCE_NAMES OUTPUT_EXE -> USER_OUTPUT_EXE_NAME LIB_NAME -> QUEST_OUTPUT_LIB_NAME VERBOSE_LIB_NAME -> QUEST_APPEND_CONFIG_TO_LIB_NAME FLOAT_PRECISION -> QUEST_FLOAT_PRECISION BUILD_EXAMPLES -> QUEST_BUILD_EXAMPLES ENABLE_MULTITHREADING -> QUEST_ENABLE_OMP ENABLE_DISTRIBUTION -> QUEST_ENABLE_MPI ENABLE_TESTING -> QUEST_BUILD_TESTS DOWNLOAD_CATCH2 -> QUEST_TESTS_DOWNLOAD_CATCH2 [QUEST_]ENABLE_CUDA [QUEST_]ENABLE_CUQUANTUM [QUEST_]ENABLE_HIP [QUEST_]ENABLE_DEPRECATED_API [QUEST_]DISABLE_DEPRECATION_WARNINGS Minor changes The output of reportQuESTEnv() has been rearranged and reordered. setSeeds (now called setQuESTSeeds()) now validates that its given list of integers is non-null. The int fields of the QuESTEnv struct (such as isMultithreaded) have been changed to type bool (exposed by in C). Added extended examples set_num_gpu_threads.(c|cpp) and user_owned_(sub)mpi.(c|cpp) Patches #729 Restored compatibility with ROCm 6 and beyond, and removed the need for -Ofast to be applied to CPU subroutines. #699 Patched an overflow bug in the GPU (Thrust) backend affecting simulation using more than 64 GiB of memory in a single GPU (i.e. a non-distributed 32-qubit statevector, or 16-qubit density matrix). Formerly, the below functions would induce a crash, or set all amplitudes to zero, or output zero. (set|create)FullStateDiagMatrFromPauliStrSum() setQuregToPauliStrSum() calcTotalProb() calcProbOf(Multi)QubitOutcome() calcFidelity() calcExpecPauliStr() calcExpecPauliStrSum() calcExpecFullStateDiagMatr() apply(Multi)QubitProjector() initRandomPureState() #693 Restored compatibility with CUDA 13 (solving compilation failure). #817 Added guards to avoid bug on AMD GPUs with >= 2^32 amplitudes per GPU. Expect to patch underlying issue in v4.3.1. Notable internal changes The CPU and GPU backends no longer use std::complex arithmetic overloads, and instead make use of custom cpu_qcomp and gpu_qcomp types. The internal functions no longer pass qubit lists as copies of heap-based std::vector . They now instead use List64 - a custom, stack-based, light-weight, fixed-capacity list - and pass constant references thereof (``ConstList64`) where possible. Internal MPI calls now use a dedicated MPI communicator rather than MPI_COMM_WORLD to avoid collisions with user messages. New contributors This release contained contributions from new contributors: Vasco Ferreira in #702, #705, #708 Maurice Jamieson in #706 Daniel Expósito Patiño in #738 Íñigo Aréjula Aísa in #722 Amon K. in #783 Ashmit JaiSarita Gupta in #780 PoJen Wang in #796 Mukul Kumar in #791 (unmerged, but contains useful benchmarking)