Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 14 additions & 0 deletions releasenotes/notes/amd-gpu-support-5c81e3a749f206bd.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
---
features:
- |
AMD GPUs are now supported through the OpenMP target-offload backend. The same
extension source builds under either vendor -- ``nvc++ -mp=gpu`` for NVIDIA,
``amdclang++ --offload-arch=gfx*`` for AMD -- so the ``'gpu-omp'`` device is
vendor-neutral. The aliases ``'gpu-amd-omp'``, ``'gpu-rocm-omp'`` and
``'rocm'`` resolve to it, alongside the existing ``'gpu-omp-offload'``,
``'gpu-nvhpc-omp'`` and ``'gpu-nvidia-omp'``, and
``sbd.get_backend('gpu-omp').__sbd_offload_target__`` reports the target a
given installation was actually built for.

The Thrust backend (``'gpu'``) remains NVIDIA-only, since upstream wires it to
``nvc++ -cuda`` and there is no rocThrust configuration to build.
18 changes: 18 additions & 0 deletions releasenotes/notes/backend-introspection-2c98d4e15b07f36a.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
---
features:
- |
Three functions have been added for inspecting how the compiled backends are
behaving in a given process: :func:`sbd.loaded_backends` lists the backends
actually imported so far, :func:`sbd.backend_load_errors` maps each unusable
backend to the reason it is unusable -- distinguishing "not built" from a
missing shared library, an architecture mismatch or a failed import -- and
:func:`sbd.has_backend_conflict` reports whether this process has loaded a
combination of backends that silently disables GPU offload.

:func:`sbd.available_backends` now judges each candidate extension statically,
by checking that it is present with no unresolved shared-library dependencies
and a matching architecture. It deliberately does not import it: importing a
backend runs ``MPI_Init``, which on an MPI without PMIx support hangs when the
process was not started under a launcher. A listed backend is therefore
present and structurally sound rather than guaranteed to load, since an
ABI mismatch only surfaces on a real import.
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
---
upgrade:
- |
The ``SBD_BUILD_BACKEND=both`` build option has been replaced by
``SBD_BUILD_BACKEND=all``; the old spelling is now rejected with an error
rather than silently misinterpreted. The accepted values are ``auto`` (the
default), ``all``, ``cpu``, ``gpu`` (alias ``gpu_thrust``) and
``gpu_omp_offload``.

``auto`` now builds every backend the detected toolchain supports, which on
NVHPC means all three rather than just CPU and Thrust, so one installation can
serve CPU and both GPU paths. Since Thrust is NVIDIA-only,
``SBD_BUILD_BACKEND=gpu`` now fails on an AMD host instead of quietly building
the CPU backend alone; use ``gpu_omp_offload`` or ``auto`` there.
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
---
upgrade:
- |
:func:`sbd.sbd_solver.solve_sci` and
:func:`sbd.sbd_solver.solve_sci_batch` now configure SBD with
``carryover_type = 0`` instead of ``1``. SBD's carryover is its own iterative
mechanism for choosing the next subspace, whereas an SQD loop selects the next
subspace itself from the returned amplitudes -- the solver discards SBD's
carryover lists entirely, so computing them was pure work, which for
``carryover_type = 2`` includes building singles-extended determinant lists.

Energies are bit-identical across carryover types, so this is a speedup rather
than a change in results. Pass ``sbd_config={"carryover_type": 1}`` to restore
the previous setting.
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
---
upgrade:
- |
``pybind11`` is no longer a runtime dependency. It is needed only to compile
the extension modules and remains declared under ``[build-system] requires``,
but nothing in the installed package imports it, so every installation had
been pulling it in for nothing. The runtime dependencies are now ``mpi4py``
and ``numpy``.
28 changes: 28 additions & 0 deletions releasenotes/notes/examples-reworked-6e03b8a51d97c42f.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
---
features:
- |
The example drivers shipped under ``python/examples`` have been substantially
reworked. Their command-line options are now grouped by the layer they act on,
with the SBD solver's options prefixed to distinguish them from the SQD loop's
-- several names previously read as one layer and acted on the other -- and
every former spelling is kept as an alias. Three SQD options that had no
command-line spelling at all, the energy and occupancy convergence tolerances
and the loop's carryover threshold, now reach the SQD loop instead of being
stuck at its defaults.

The Hartree-Fock occupancy seed has been dropped from the SQD driver. Passing
it replaced postselection with configuration recovery in every run, so at the
first iteration -- where the only occupancies available are mean-field, with
every virtual orbital at exactly zero -- the subspace was largely synthesized
by bit flips against that prior rather than built from sampled
configurations. The driver now postselects at the first iteration and uses
solver-derived occupancies thereafter, which is what published SQD does.
- |
A new example driver, ``run_sqd_enlarge_subspace_sbd.py``, grows its own
subspace between diagonalizations. It drives the SQD loop one iteration at a
time and, after each solve, expands the dominant determinants through
same-spin single excitations using ``qiskit-addon-sqd``'s own excitation
utilities, feeding the result forward as the next round's configurations. It
stops when the expanded set adds nothing new -- the subspace is closed under
single-excitation connectivity -- or when both the energy and the occupancies
stop moving.
17 changes: 17 additions & 0 deletions releasenotes/notes/fix-amplitude-labels-9c2f60b81a4e37d5.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
---
fixes:
- |
Wavefunction amplitudes are now labeled with the determinants that were
actually diagonalized rather than with the CI strings passed in. The two can
differ: the input is sorted into SBD's canonical order and deduplicated before
diagonalization, so amplitudes and labels are now taken from one list by
construction.

In practice this fixes duplicate CI strings, which the sorting step documents
that it removes. A caller who passed them got a ``RuntimeError`` complaining
that the amplitude count did not match the subspace size, rejecting a run that
had in fact succeeded; such a call now completes, returning the amplitudes for
the distinct determinants. Ordering was not observably wrong before, since the
canonical order happens to coincide with ascending integer order, but nothing
upstream guarantees that -- and if it ever stopped holding, the result would
have been correct amplitudes against the wrong labels, with no error raised.
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
---
fixes:
- |
The determinant word size used by :func:`sbd.sbd_solver.solve_sci` and
:func:`sbd.sbd_solver.solve_sci_batch` no longer defaults to 64, which was
undefined behavior. SBD packs determinant bitstrings into words of
``bit_length`` bits and builds its masks by shifting a 64-bit word left by
that amount; a shift of 64 is undefined, and in practice the shift count is
masked to zero, collapsing the mask. The affected routine is reached from the
multi-rank determinant redistribution and sorting paths. The default is now
20, matching upstream's documented default, and the maximum usable value is
63.
- |
A caller-supplied ``bit_length`` in ``sbd_config`` is now honored end to end.
The value previously reached the C++ engine but not the code that packs the
determinants, which kept using 64, so the two disagreed about how to interpret
the same words and the determinants were silently corrupted. The configuration
is now built before packing, so one value is used throughout.
10 changes: 10 additions & 0 deletions releasenotes/notes/fix-device-config-auto-5e8c31f70a9b26d4.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
---
fixes:
- |
``DeviceConfig.auto()`` now accounts for which backends were actually built,
not only for what hardware is present. It previously returned the Thrust GPU
device whenever any GPU was detected, which selected a backend that cannot
exist on an AMD host, and on a CPU-only build running on a machine with a GPU
it selected a backend that had not been compiled. The available devices are
now intersected with the backends present, so a CPU-only installation resolves
to the CPU and an AMD host resolves to the OpenMP target-offload backend.
20 changes: 20 additions & 0 deletions releasenotes/notes/fix-fabricated-amplitudes-0b7d4e8a26c19f53.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
---
fixes:
- |
The wavefunction amplitudes returned by :func:`sbd.sbd_solver.solve_sci` and
:func:`sbd.sbd_solver.solve_sci_batch` are no longer replaced by a uniform
array. SBD writes the amplitudes for the full input subspace -- the product of
the alpha and beta determinant lists -- regardless of its carryover settings,
but the solver sized its read against the carryover counts instead. Under the
former default those sizes never matched, and the mismatch was answered by
substituting a uniform array: correctly shaped, correctly normalized, and
carrying no information at all.

This was silent and load-bearing. ``qiskit-addon-sqd`` selects each
iteration's determinants by amplitude magnitude and weights them by the
squared magnitude, so every iteration after the first was seeded from uniform
weights over an already-truncated list of determinants. The amplitudes are now
read over the full subspace and paired with the determinants that span it, and
a missing dump or a size mismatch raises rather than being papered over, since
either means the run did not do what was asked. Energies and orbital
occupancies were never affected; they come from SBD directly.
15 changes: 15 additions & 0 deletions releasenotes/notes/fix-rocm-detection-2a4b96e0d75c831f.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
---
fixes:
- |
AMD GPU detection and reporting have been corrected. The ``rocm-smi`` probe ran
with a two-second timeout, which is shorter than the tool takes to answer on a
populated node -- so device detection flaked and could resolve to the CPU on a
machine with eight GPUs. The probe now allows enough time, and requires a
per-device line in the output rather than treating a successful exit as proof
that a GPU is present, since ``rocm-smi`` exits successfully with none.
- |
``get_device_info()`` no longer overcounts AMD GPUs. It counted every output
line mentioning a GPU, reporting 40 devices on a node with eight, and now
counts distinct device indices. Relatedly, ``print_device_info()`` no longer
reports the CPU as available while :func:`sbd.available_backends` reports that
nothing was compiled; it now leads with the backends that were actually built.
11 changes: 11 additions & 0 deletions releasenotes/notes/fix-temp-dir-deadlock-7d19b5e82c034af6.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
---
fixes:
- |
A ``temp_dir`` that does not yet exist is now created, along with any missing
parents, instead of hanging a multi-rank run. Handing such a path to
:func:`sbd.sbd_solver.solve_sci` raised ``FileNotFoundError`` on rank 0, and it
did so before the broadcast that hands the created directory to the other
ranks -- so the remaining ranks waited in that broadcast forever and the job
hung with no output rather than reporting the error. Directory setup is now
wrapped so that its outcome is broadcast: on failure every rank raises
together.
16 changes: 16 additions & 0 deletions releasenotes/notes/get-device-id-8f25c091e7b43ad6.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
---
features:
- |
A new :func:`sbd.get_device_id` function reports the GPU index the calling
rank will use, or ``-1`` when there is none -- a CPU-only build, or no visible
devices. It applies the same ``rank % num_gpus`` rule SBD's own
diagonalization uses, asking the backend rather than recomputing it, so
another library placed on the same card cannot drift out of step with SBD.
The device count comes from the GPU runtime rather than from parsing
``nvidia-smi`` output.

The query selects no device and creates no context, so it is safe to call
before any GPU work and before another framework initializes its own backend.
That is what it is for: a framework such as JAX, initialized on a multi-GPU
node without being told which device belongs to this rank, will otherwise
claim a card SBD is about to use.
23 changes: 23 additions & 0 deletions releasenotes/notes/lazy-backend-loading-9a17c4be250d3f86.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
---
upgrade:
- |
Compiled backends are now imported lazily, one per process: only the backend
for the device actually in use is loaded. This is what makes it safe for all
three backends to be installed side by side, and it is not merely an
optimization. When the CPU extension and the OpenMP target-offload extension
are both loaded into one process, the OpenMP runtime is left initialized
host-only, after which offload regions run on the host while device queries
still report a GPU -- so the run looks accelerated, returns the correct
energy, and exits successfully. Code that imports the ``_core_*`` extension
modules directly should import only one; :func:`sbd.has_backend_conflict`
reports the bad combination.
- |
``OMP_TARGET_OFFLOAD`` is now set to ``MANDATORY`` before the OpenMP
target-offload backend is imported, unless it is already set to something
non-empty. Under the OpenMP default of ``DEFAULT``, a rank with no visible
device falls back to the host and returns a plausible answer with no
indication that nothing was offloaded -- reachable from an empty
``CUDA_VISIBLE_DEVICES``, a mispinning launcher, or a GPU-less node.
``MANDATORY`` turns each of those into an immediate error. Set
``OMP_TARGET_OFFLOAD=DEFAULT`` explicitly to restore host fallback, for
example to smoke-test on a machine with no GPU.
10 changes: 10 additions & 0 deletions releasenotes/notes/macos-support-7b4f92c05e83a1d6.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
---
features:
- |
macOS is now a supported platform, declared in the trove classifiers and
covered by CI on both Linux and macOS. Building on macOS requires an OpenMP
runtime, which Apple's clang does not ship -- install one first, for example
with ``brew install libomp``. Homebrew's prefix is discovered by asking
``brew --prefix`` rather than assuming ``/opt/homebrew``, so Intel and Apple
silicon Macs are both handled, and the built extensions carry the rpath
entries needed to resolve keg-only libraries at import time.
17 changes: 17 additions & 0 deletions releasenotes/notes/populate-sci-result-rdms-1d6a05f8c43b297e.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
---
features:
- |
:func:`sbd.sbd_solver.solve_sci` and
:func:`sbd.sbd_solver.solve_sci_batch` now populate the ``rdm1`` and ``rdm2``
fields of the returned ``SCIResult`` when RDMs were requested with
``sbd_config={"do_rdm": 1}``. SBD computed these all along and the bindings
returned them, but the solver never read the keys, so the fields were left at
their ``None`` default even for a caller who had asked for them. The
spin-summed tensors use the index convention that ``SCIResult.rdm1`` and
``rdm2`` are contracted with elsewhere in ``qiskit-addon-sqd``, verified
against PySCF's own ``make_rdm1``/``make_rdm2`` on all three backends.

``do_rdm = 0`` remains the default, and in that case both fields stay ``None``
exactly as before. The conversion is also available on its own as
:func:`sbd.sbd_solver.assemble_rdms`, for callers working with the raw result
dictionary returned by :func:`sbd.tpb_diag` and :func:`sbd.gdb_diag`.
16 changes: 16 additions & 0 deletions releasenotes/notes/remove-h-comm-size-4f2a91c7d3b0e85a.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
---
upgrade:
- |
The ``h_comm_size`` attribute has been removed from the ``TPB_SBD`` and
``GDB_SBD`` configuration objects. Upstream SBD declares the field but never
reads it: ``diag()`` shadows it with a local
``h_comm_size = mpi_size / (task_comm_size * base_comm_size)`` and passes that
to the determinant-basis communicator, so the attribute could only ever read
back as 1 and silently ignored whatever was assigned to it.

Code that assigns ``sbd_data.h_comm_size`` now raises ``AttributeError``.
Note also that an ``"h_comm_size"`` key in the ``sbd_config`` dictionary
accepted by :func:`sbd.sbd_solver.solve_sci` is now skipped without warning,
since unknown keys are ignored. To change the helper dimension of the MPI
grid, change the number of ranks or one of ``adet_comm_size``,
``bdet_comm_size`` and ``task_comm_size``.
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
---
upgrade:
- |
Multi-iteration SQD results computed through :func:`sbd.sbd_solver.solve_sci`
or :func:`sbd.sbd_solver.solve_sci_batch` will differ from those of earlier
releases. The wavefunction amplitudes returned to the SQD loop were
previously replaced by a uniform array in the common case, and
``qiskit-addon-sqd`` selects each iteration's determinants from those
amplitudes; the amplitudes are now correct, so the subspaces explored after
the first iteration differ. See the corresponding entry under Bug Fixes.

Single-iteration energies and orbital occupancies are unaffected: those come
from SBD directly and were never derived from the amplitudes.
13 changes: 13 additions & 0 deletions releasenotes/notes/thrust-mpi-build-options-4a6d17b9c8e052f3.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
---
features:
- |
Two build options have been exposed for running the Thrust backend on an MPI
that cannot address device memory: ``SBD_NON_CUDA_AWARE_MPI=1`` stages every
transfer through host memory, and ``SBD_THRUST_SAFE_MPI_ALLREDUCE=1`` stages
only the allreduce, which is the cheaper of the two. Without one of these, a
GPU-aware MPI is a hard requirement of that backend.

Both are compile-time defines, so they must be set before installing, and
changing one means rebuilding. Setting them in the environment of an existing
installation does nothing. Expect either to be slower than a GPU-aware MPI,
since staging adds a copy per transfer.