GDB: sharded determinant basis, numpy marshalling, drivers and first tests - #46
Open
Sophia Wen (hfwen0502) wants to merge 22 commits into
Open
Sophia Wen (hfwen0502) wants to merge 22 commits into
Sophia Wen (hfwen0502) wants to merge 22 commits into
Conversation
Sophia Wen (hfwen0502)
force-pushed
the
gdb-sharding
branch
from
September 30, 2026 13:56
4e07e7b to
88f8bd7
Compare
Jim Garrison (garrison)
added this pull request to stack #47
September 30, 2026 16:55
Sophia Wen (hfwen0502)
force-pushed
the
gdb-sharding
branch
from
September 30, 2026 17:45
f13cf79 to
e440f7a
Compare
…d lists GDB spans a subspace with an explicit determinant list rather than the Cartesian product TPB uses, which is what makes it the right solver for a sparse subspace. But the binding required every rank to pass the WHOLE list, pinning b_comm_size to 1, and that cascaded into three ceilings: - b_comm is the only dimension that divides the basis. h_comm is a row stride within a block (gdb/mult.h:70) closed by an allreduce (:208) and MakeHelpers ignores it, so at b=1 every rank held the entire basis AND the entire excitation lookup however many ranks were used. Memory did not scale at all. - t_comm_size was forced to 1 too: GDB runs one task per basis-ring station and a single block has one station. - GPU GDB was capped at ONE rank. gdb/mult_thrust.h:310-314 throws unless h_comm_size == 1, and h = ranks/(t*b), so with b=t=1 the helper dimension took every rank. Multi-GPU GDB was unreachable. So b_comm_size now selects the contract: 1 keeps the previous meaning, >1 means each rank passes its own shard. What upstream does not check is checked here, collectively so the verdict is unanimous before any rank raises: shards identical across a b_comm position (the in-memory path never broadcasts the list, unlike the file overload at gdb/sbdiag.h:764), globally sorted and disjoint via the neighbour exchange that load_basis_from_files uses, t <= b, t*b dividing the rank count exactly (upstream's integer division otherwise yields communicators of unequal size), a non-empty shard for the schemes that index config[0] unconditionally, and h == 1 on Thrust. Completeness cannot be checked, so global_dim is returned for the caller to assert on. All six of the app's placement schemes are exposed as one determinant_distribution string. The three legacy booleans are deliberately not: they encode a four-way choice with a priority rule where two are unreachable unless do_redist_alpha_eq is explicitly zeroed. do_shuffle is dropped from GDB_SBD for the same reason h_comm_size was -- upstream parses it and never reads it, so the attribute could only mislead. Determinants now cross as an (n, words) uint64 array, built into det_vector with a single memcpy instead of three copies, and from_strings packs in bulk rather than once per determinant across the boundary. A nested list still works via forcecast. The shape is passed as an explicit std::vector<py::ssize_t>: a braced initializer list is ambiguous under g++ between array_t's ShapeContainer constructor and its copy/move, which nvc++ and clang accept, so the Linux CPU build would otherwise not compile. method 2 and 3 are rejected. They are valid for TPB, where they select Lanczos; GDB has none, and gdb::diag assigns energy only inside its method 0 and 1 branches (gdb/sbdiag.h:418, :501), so passing 2 returned an uninitialized double with no diagonalization having run. Closes half of #22: h_comm_size was already unexposed, and the b_comm_size == 1 restriction it flagged is lifted here using the redistribution APIs the 93ebabec bump brought in. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
GDB was untested in this wrapper and, at the pinned upstream commit, upstream's only GDB test exercises the distribution helpers rather than a diagonalization. The anchor needs no pinned number: TPB's subspace is one GDB can also express, so interleaving every (alpha, beta) pair into a full determinant points both solvers at exactly the same Hilbert space and their energies must agree. On h2o they agree to 1.4e-14, and the full 275^2 interleave reproduces the -76.23594663 published for that alpha list, which test_reference_energies.py asserts through TPB. Also covered: the sharded path against a TPB run of the same subspace; the task dimension becoming usable once b_comm_size allows it; all six placement schemes agreeing, since placement is a load-balancing decision and must not move the answer; and each guard with its own case -- non-disjoint shards, a rank count the grid cannot tile, unknown and misconfigured placement schemes, duplicate determinants, and Lanczos methods. from_strings is checked elementwise against a per-determinant from_string loop, and a nested list is checked to still work. The Fe4S4 case runs in a subprocess: det_vector's row width is fixed per process, and h2o packs into one 64-bit word where Fe4S4 needs two, so they cannot share one. That is the constraint gdb_diag reports, and forking is the workaround it prescribes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two drivers, both staying on the in-memory entry point -- determinant text is read in Python and handed to the binding as an array, never through SBD's file-based overload. run_gdb_diag.py runs a single diagonalization over an explicit determinant list. It takes full-determinant files (defaulting to upstream's four Fe4S4 files) or forms the product basis from an alpha list, which is the subspace TPB would build and so the cross-check against it. With --b_comm_size it shards: when the number of files is any multiple of b, each rank reads its own contiguous block of files, which has to be contiguous rather than strided because shard i must hold a strictly lower range than shard i+1. run_gdb_heatbath.py iterates: diagonalize, let SBD expand the subspace from the resulting wavefunction, diagonalize the larger subspace, repeat. carryover_type 2 and 3 return the parents together with the new candidates, so one round's result is the next round's subspace and the loop needs nothing extra. It is structured as a cutoff LADDER because the expansion reaches a self-consistent size for a fixed heatbath_cutoff and then stops growing -- rounds run at one cutoff until the energy settles or the subspace stops growing, then the next rung starts, with --max_dim capping the whole run. --seed picks the starting subspace (determinant files, the Hartree-Fock determinant alone, an interleaved alpha list, or an arbitrary bitstring file) so one driver produces every row of a seed comparison. The README documents what each path does and the constraints that are easy to get wrong: b_comm_size is the only dimension that divides memory; t <= b because there is one task per basis-ring station; t*b must divide the rank count; the helper dimension must be 1 on Thrust, which is why multi-GPU GDB needs b > 1; expansion and carryover run on the host even in a GPU build, so OMP_NUM_THREADS should stay generous; and heatbath_truncation prunes parents rather than admitting candidates, so leaving it at 0 is almost always right. It also records that upstream's Fe4S4 "GDB" data is the full 244x244 product of AlphaDets.txt rather than a sparse subspace, which is worth knowing before reading it as a GDB benchmark, and points at apps/gen_dets for producing pre-split shard files rather than shipping a second tool for the same job. Deliberately no timing or hardware figures: the point is which paths exist and what each costs in memory and constraints, and users should measure on their own systems rather than inherit ours. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
On an install with more than one backend built, every test that builds objects from the fixture and then calls a module-level entry point failed: TypeError: tpb_diag(): incompatible function arguments. Invoked with: ..., <sbd._core_gpu_thrust.TPB_SBD object>, <sbd._core_gpu_thrust.FCIDump object>, ... The fixture calls get_backend(None), which auto-resolves and PREFERS Thrust when it is built and a GPU is present. The module-level tpb_diag/gdb_diag then reach _ensure_initialized(), whose default is 'cpu'. So the FCIDump and config came from one extension module and the diagonalization was dispatched in another. Pre-existing and not GDB-specific: test_reference_energies.py fails 2/2 the same way. It stayed invisible because a CPU-only build has nothing to disagree with, and that is what CI and a laptop give you. It shows up on any machine where the default build produced cpu + gpu + gpu-omp, i.e. exactly the hardware the GPU paths are meant for. Found while confirming the g++ CPU build on a box with both cpu and Thrust modules present. Pinning the default to whatever the fixture hands out fixes both test files and keeps SBD_TEST_DEVICE working as before. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three gaps a reader would hit, now that GDB is more than a footnote. The Overview presented the bindings as TPB's, with GDB in a trailing sentence. It now names both methods by the property that decides which you want -- TPB's subspace is a product of two half-determinant lists, GDB's is the explicit list you pass and so can be sparse -- and says outright to prefer TPB for a product subspace, since it represents that case with the two lists rather than every product element and carries no extra constraints. The qiskit-addon-sqd section read as though either solver could be plugged in. It cannot, and not for want of plumbing: the addon's interface is a product subspace by construction (ci_strings is a (strings_a, strings_b) pair, SCIState.amplitudes an |a| x |b| matrix), so a sparse determinant list cannot be expressed without padding back to the full product and discarding the reason to use GDB. Stated plainly, with a pointer to calling gdb_diag directly. The API reference mentioned gdb_diag without its data contract. It now says det may be a whole basis or this rank's shard depending on b_comm_size, lists the constraints a sharded run carries, and notes they are checked rather than assumed -- then defers to examples/gdb/README.md for the decomposition, the placement schemes and which returned values are replicated. Deliberately no copy of that detail here: duplicating it guarantees the two drift. Also corrects the savename description, which this branch had made wrong: with the basis split it is one file per b_comm position, not a single ...000000.bin. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Stripping the benchmark figures replaced the passage from "Verified on 8x H100..." onward, but the paragraph before it already made the same point, so the section stated "multi-GPU GDB requires --b_comm_size" twice in consecutive paragraphs. Merged into one, keeping the mult_thrust.h reference and the worked example. The substance was right and is unchanged: GDB has Thrust kernels only, so GPU GDB is NVIDIA-only and AMD GDB is CPU-only; the Thrust path needs helper == 1, which is why more than one GPU requires splitting the basis. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The backend bullet listed cpu, gpu and gpu-omp without qualification, which is right for TPB and misleading for GDB: GDB has Thrust kernels only, so under 'gpu-omp' it pins a device and then runs on the host. A reader of the landing page alone would have concluded AMD GDB has GPU support. Adds the distinction and a pointer, without pulling the GPU detail up from examples/gdb/README.md. That was the last place in the docs where the claim was ambiguous; the GDB and shared example READMEs already stated it, and the driver warns at runtime when --device resolves to gpu-omp. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…licated global_dim() summed local determinant counts over MPI_COMM_WORLD. Ranks sharing a b_comm position hold identical shards by contract, so with t_comm_size or the helper dimension above 1 every shard was counted once per replica: Fe4S4's 59,536-determinant basis reported as 119,072 at --b_comm_size 2 --t_comm_size 2 on four ranks. Energies were unaffected -- the binding computes its own global_dim over b_comm, which is what run_gdb_diag reports -- but the driver's dimension column was wrong and --max_dim compared against the inflated figure, so a ladder would have stopped early. It hid because every run until now used b == ranks with t = 1, where the world sum and the b_comm sum coincide. Found by deliberately exercising the b != ranks reshard path. Fixed by contributing only from the ranks that are one-per-b-position: with the layout rank = h*(b*t) + t_index*b + b_index those are exactly ranks 0..b_comm_size-1, so the b_comm sum needs no sub-communicator. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This branch lifts that restriction, and the prose further down already described the shard contract, but the configuration table was left saying "must be 1 for gdb_diag" -- actively wrong for the branch that changes it. I had edited the surrounding prose and missed the table. Replaced with what the three dimensions now are and whose constraint each one is, since they are different in kind: b_comm_size is free and is the only dimension that divides memory; t_comm_size <= b_comm_size is upstream's algorithm (one task per basis-ring station, unchecked there, so gdb_diag rejects it up front); and their product must divide the rank count, with the derived helper dimension taking the quotient and required to be 1 on Thrust, which is why more than one GPU needs a split basis. Depth stays in examples/gdb/README.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…t a limitation The entry inherited from the reorg branch says multi-rank GPU GDB is impossible because b_comm_size is pinned to 1. This branch lifts that, so the same symptom now has a resolution: give every rank to the basis, e.g. -np 4 with b_comm_size 4, leaving a helper dimension of 1. Also notes that gdb_diag checks it up front rather than letting the Thrust kernel throw mid-launch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The intro and See Also deferred to examples/README.md for backend selection, which no longer exists. Replaced with what a GDB user needs in place: the --device values, with 'gpu' flagged as the only GPU backend that has GDB kernels; available_backends() and loaded_backends(); and the 'cpu'/'gpu-omp' interaction, where the CPU module initializes the shared OpenMP runtime host-only and the offload backend then runs silently on the host. Links to the TPB README for the bundled test data rather than restating it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… examples/gdb The Configuration section had grown three paragraphs on how the GDB communicators relate, all of which examples/gdb/README.md already covers -- its dimension table, the ring explanation for t <= b, "only b shards memory", and the Thrust helper restriction. Replaced with the same shape TPB uses: the rank arithmetic in one line, plus the one API fact a caller cannot do without (above b_comm_size 1, det is this rank's shard), and a pointer for the rest. The gdb_diag prose further down is trimmed the same way. Two stale claims fixed while there. The "Shares ... with TPB_SBD" line listed method, whose range differs -- GDB has only 0 and 1 -- so method now appears in the GDB table explicitly, as it does in TPB's. It also listed do_shuffle, which this branch removed from GDB_SBD because upstream parses it and never reads it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
gdb_diag does not return the wavefunction, which invites the question of how an iterative driver can work without it. It can because the selection happens inside SBD: the amplitudes drive it -- weight truncation keeps determinants by |c| and heatbath scoring is essentially |c_i . H_ij| -- but WeightTruncation and HeatbathExpansion consume them in C++ and return only the expanded determinant list, which is the next subspace. Nothing but determinants crosses the Python boundary. Also says where amplitudes would be needed: Python-side selection, as the TPB enlarge-subspace driver does, which for GDB would mean reading them back from the per-shard savename files. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…/gdb The top README carried the per-shard file layout, which is GDB detail and belongs with the rest of it. It also read as a requirement when it is not: nothing in the normal flow needs the wavefunction -- energy, density and RDMs come back directly, a heatbath ladder iterates on carryover_det, and GDB is not wired into qiskit-addon-sqd, whose SCIState would be the usual consumer. Top README now says amplitudes are not returned, that this rarely matters, and where to look if you do want them. examples/gdb/README.md gains a short "Getting the amplitudes, if you want them" section holding the file layout, and its outputs table stops repeating it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…drivers Three gaps, all found by reading the README as a new user would. The commands that pass no input flags looked like they ran on no data at all. They run upstream's Fe4S4 case, because --fcidump and --detfiles default to it; say so in a new "The default data" section, and show one command with the paths written out so the --detfiles syntax is copyable. Add --device gpu examples. GDB's only GPU backend is Thrust, which requires helper == 1, so multi-GPU GDB cannot omit --b_comm_size -- worth showing next to the CPU runs rather than only in the GPU section further down. Document run_gdb_heatbath.py's parameters as tables, grouped seed / ladder / solver, matching examples/tpb/README.md. Cross-checked against argparse: no invented flags, every option covered, aliases named. test/test_gdb_drivers.py is new: the drivers had no test, so argument parsing, determinant reading, the sharding arithmetic and the default paths could all break with every library test still green. Runs each driver as a subprocess, so __main__ and the defaults are exercised rather than bypassed. The h2o case asserts the same -76.0588897208 that test_gdb_equivalence.py anchors against TPB, tying driver to solver; the slow case asserts the flagless run really is Fe4S4 at 59,536 determinants, which is what keeps the README's claim honest. Both guardrails are covered, t > b especially -- unguarded it segfaults instead of erroring. 6 passed in 16 s by default, 7 in 37 s with --run-slow. Verified non-vacuous: perturbing the anchor by 1e-8 fails the test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nput flags differ The backend block listed all four --device values as if they were alternatives. For GDB only two are: gpu-omp has no GDB kernels and runs on the host, and auto can resolve to gpu-omp and do the same. List cpu and gpu, then one sentence on why the other two are accepted but not GPU paths. Also state plainly that --fcidump and --detfiles are identical across the two drivers, and that two input flags are not: the alpha list is --alpha-file in the heatbath driver but --from-alpha in the diagonalization one, and --seed selects the subspace source here while there it is an integer RNG seed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ha-file spelling The two drivers had grown a flag collision and a gratuitous difference. --seed took a string here (files|hf|from-alpha|strings) and an int in run_gdb_diag.py, where it is the RNG seed for a random initial vector. Same flag, incompatible types, so a command could not be moved between the drivers and `--seed hf` failed confusingly in one of them. Renamed to --subspace-from, which says what it does. --seed still works and prints a deprecation notice to stderr; argparse's own deprecated= needs 3.13 and this package supports 3.10, so the old spelling is detected in argv instead. The alpha determinant list was --from-alpha in one driver and --alpha-file in the other, for no reason. Both drivers now accept both spellings. Tests cover all three paths: either alpha spelling gives the same 576-determinant energy, and the deprecated --seed still runs while emitting the notice. 9 passed / 1 slow-skipped in the driver file, 26 passed in the serial suite, 13 under mpirun -n 2. README parameter tables re-checked against argparse: no invented flags, none undocumented. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…missed The branch is unmerged, so nothing depends on the old spelling: --seed is gone rather than deprecated, along with its argv-sniffing warning and the test for it. `--seed hf` now exits 2 with "unrecognized arguments". The previous commit's rename was pattern-by-pattern and missed four lines of the module docstring plus two messages, so `--help` still advertised --seed while argparse had moved on. Swept the whole file this time and asserted no mention survives; the only remaining --seed in examples/gdb is the README line contrasting it with run_gdb_diag.py's integer RNG seed, which is deliberate. 9 passed with --run-slow, 25 in the serial suite, 13 under mpirun -n 2. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…subspace-from --strings-file takes one determinant per line, not a counts JSON, and the conversion has a step that is easy to miss: qiskit-addon-sqd emits [beta | alpha] concatenated while GDB wants the bits interleaved. Document both steps with a snippet built on the driver's own interleave(), and note that a set/sorted() handles the duplicate rejection and ordering the shard contract expects. Skipping the interleave is worth spelling out because it does not fail cleanly: the concatenated strings diagonalize to a plausible number (-66.8043 where the correct conversion gives -76.0724) and only abort later inside the heatbath expansion with std::out_of_range. The electron-count check cannot catch it -- permuting bits preserves how many are set -- so the density still sums to 10. Snippet and command both run as written; the snippet reproduces the same file byte-for-byte as the conversion that was tested. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both drivers now compare every determinant's per-spin electron count against the FCIDUMP's NELEC/MS2 and refuse a mismatch, naming how many determinants are wrong, the first offender, and the two likely causes: a concatenated [beta | alpha] bit order where GDB wants them interleaved, or samples that were never postselected. This is worth a guard because the failure is not self-announcing. Feeding qiskit-addon-sqd's concatenated strings straight in diagonalizes to -66.8043 where the correct conversion gives -76.0724, then aborts inside the heatbath expansion with std::out_of_range -- a crash whose message says nothing about bit order. The occupation density cannot catch it: permuting bits preserves how many are set, so it still sums to the right electron count. Cost is not a concern. The popcount runs on the already-packed words via a 256-entry uint8 table, 37 ms per million determinants measured in isolation, and end to end it disappears into run variance: a 1M-determinant run took 76.5/76.9/76.6 s without the check and 77.5/76.6/76.6 s with it. numpy.bitwise_count would be ~9x faster but needs numpy 2.0 and this package's floor is 1.19, so the table is the portable choice. --skip-weight-check opts out for a deliberately mixed-sector subspace. 28 passed serial, 12 with --run-slow, 13 under mpirun -n 2. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`mpirun -np 4 -x OMP_NUM_THREADS=8 ...` is Open MPI syntax. On MPICH, Hydra
rejects it outright:
[mpiexec] match_arg (lib/utils/args.c:166): unrecognized argument x
so that example could not run at all on an MPICH system -- including the h100 box
we validate on. Set the variable in the shell instead, which both implementations
honour for a single-node run, and say why rather than leaving the next person to
rediscover it.
Verified on both: Open MPI 5.0.10 and MPICH 5.0.0 give -76.0588897208 with 144
determinants per rank at --b_comm_size 4.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…cannot shard Three things a real run hit that the drivers said nothing about. SBD's Davidson can return without iterating. If the start vector has no coupling to the rest of the subspace the residual is zero immediately, so it reports one determinant's diagonal element as the ground state energy: a valid upper bound, but not the subspace's eigenvalue, and the run exits 0 with a plausible number. The occupancies give it away -- a one-determinant wavefunction has every occupancy 0 or 1, where a correlated one is fractional -- so both drivers check that and say what happened. Guarded on dimension > 1, since a single-determinant subspace legitimately looks that way; the heatbath driver only checks its seed round, because after an expansion the subspace contains its own excitation neighbours. --subspace-from hf with --b_comm_size > 1 is a trap: one determinant cannot be divided, so rank 0 owns it, the other ranks get empty shards, and a single parent leaves OpenMP nothing to split either -- the first rounds crawl on one core. The driver now warns, and the README example runs serially. Also note in the README that the last cutoff dominates a ladder's cost, and that --max_dim is what makes an unreachable rung a clean stop rather than an out-of-memory failure mid-round. Verified: detector true on [1,1,0,0], false on fractional occupancies; a healthy h2o run emits nothing. 34 passed serial, 13 under mpirun -n 2. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sophia Wen (hfwen0502)
force-pushed
the
gdb-sharding
branch
from
September 30, 2026 18:03
e440f7a to
f36798e
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Makes GDB usable for a sparse, distributed subspace. Experimental — GDB moves fast upstream and this will follow it.
Why
gdb_diagrequired every rank to pass the whole determinant list, sob_comm_sizehad to be 1. That cascaded:bis the only dimension that shards memory (h_commis a row stride plus an allreduce;MakeHelpersignores it),t <= bis structural, and Thrust throws unlessh_comm_size == 1— so multi-GPU GDB was unreachable and memory did not scale at all. Upstream's ownrun.shuses--b_comm_size 2; the restriction was ours, not SBD's.What changed
gdb_diagtakes a per-rank shard.b_comm_size == 1keeps today's meaning (whole basis);> 1means this rank's shard, verified globally rather than trusted — shard identity acrossh_comm, local sort, cross-shard disjointness. All validation votes throughMPI_Allreducebefore any rank throws, so a bad input cannot deadlock.(N, words)in and out, viac_style | forcecastso existing list callers still work. Removes two of three copies; addsfrom_stringsfor bulk packing instead of onefrom_stringcall per determinant (issue Duplicate-determinant handling differs between TPB and GDB; cost of the #29 label conversion at scale #31 flagged that at 39k).determinant_distributionstring. The three legacy booleans stay unexposed: they encode a four-way choice with a priority rule, anddo_redist_alpha_eqdefaults true so the others are dead unless it is zeroed.t > bsegfaults upstream (a starved rank dereferences an emptyexidx) andt*bnot dividing the rank count silently builds ragged communicators. Both now refuse up front.run_gdb_diag.pyandrun_gdb_heatbath.py(a cutoff ladder;carryover_type2/3 returns parents with candidates, so one round's result is the next subspace). In-memory throughout, no file round-trip.[beta | alpha]strings where GDB wants them interleaved used to diagonalize to a plausible wrong energy and then abort inside the expansion withstd::out_of_range. Now refused with the cause named. The occupation density cannot catch it — permuting bits preserves how many are set.Verification
GDB had no test anywhere before this, in the wrapper or upstream. Now:
-76.23594663, h2o-1em4-76.24295848, n2-1em4, n2-1em5, up to 16.3M determinants.b1/2/4/8 andt1/2, across all six schemes — placement must not change the answer.28 passedserial,12with--run-slow,13undermpirun -n 2.Known gaps
RDMs (
do_rdm=1) andsavenameare untested atb > 1; no multi-node run;DetBasisCommunicatorleaks four comms per call upstream, which a long loop would feel.gpu-omphas no GDB kernels at all, so GDB on AMD is CPU-only — documented rather than worked around.🤖 Generated with Claude Code