Conversation
Signed-off-by: Zack Meeks <zmeeks@nvidia.com>
This PR replaces [cuVS-Lucene #195](NVIDIA/cuvs-lucene#195) as cuVS-Lucene has been merged into cuVS. This PR builds on the out-of-core host-streaming refactor from #2476 by @nvzm123. This PR subsumes #2476 to avoid stacking the PRs. On top of #2476, this PR adds the following. **Improvements** - Native flat buffering: stream vectors directly into a native host matrix during indexing instead of buffering them as a heap List<float[]>, cutting peak host memory from ~2x to ~1x and eliminating the per-vector matrix-assembly copy. Opt-in; requires all input vectors to be indexed in the original order, i.e. not supported for sorted, merged, or filtered index segments; not currently supported for segments with quantized fields. - Parallelized CAGRA-to-HNSW graph conversion: materializing the CAGRA adjacency into on-heap NeighborArrays was a serial per-node loop; parallelized under the existing `writerThreads` knob. - Parallelized level-0 HNSW graph serialization: level-0 (all N nodes) is now delta/VInt-encoded in parallel across memory-bounded waves of threads, then concatenated to the `IndexOutput` in node order according to on-disk format in the serial path. **New example: OptimizedCagraHnswBuildExample** A reference pattern for building a large accelerated HNSW index with all ingest- and build-side optimizations, showcasing: - Streaming, prefetched, bounded-memory ingestion: open the source file once, read it front-to-back in large sequential chunks; hold at most two chunks in memory; fill the next chunk on a background thread while the ingest thread drains the current one, hiding disk read behind indexing; unpack into a caller-reused float[] allocation (no per-vector allocation, safe because Lucene copies the value eagerly inside `addDocument`). - Native flat buffering sized exactly to each segment's vector count via `withNumInputVectors`, avoiding the heap-buffered assembly copy. - Overlapping multi-segment build: - Default: K single-segment passes appended to one directory; peak host memory is one slice (N/K). - Overlapped: a bounded pool (`PIPELINE_DEPTH`) builds segments concurrently into their own directories with the GPU commit serialized on a semaphore (ingest overlaps a prior segment's GPU commit), then combines the finished per-segment indexes by hardlinking their files into the final directory via `HardlinkCopyDirectoryWrapper` + `addIndexes` (no bulk copy of vector data). Peak host memory is up to `PIPELINE_DEPTH * N / K`. - An `enableRMMAsyncMemory()` call (with a note that it must not be used with CPU-only codecs) to opt-in to RMM-managed memory resources, plus exposing the primary tuning knobs (withMaxConn, withBeamWidth, withCuvsDistanceType, withWriterThreads). Authors: - James Xia (https://github.com/jamxia155) - https://github.com/nvzm123 - Igor Motov (https://github.com/imotov) Approvers: - James Lamb (https://github.com/jameslamb) - Igor Motov (https://github.com/imotov) URL: #2481
Acquire device streams before allocation, preserve ownership across builder cleanup, and keep CAGRA persistence on the original device matrix. Expand lifecycle, fallback, merge, and graph-integrity regression coverage.
Keep host-backed CAGRA-to-HNSW inputs, exact live-vector merge sizing, trivial merge handling, and the compatible upper-layer bridge. Restore the GPU-search codec to the target-branch device-input behavior and defer broader lifecycle, graph-integrity, and quantized-merge hardening to a follow-up.
|
/ok to test 8faae50 |
imotov
left a comment
There was a problem hiding this comment.
That looks much better. I think we are really close. Left a few minor comments.
| import org.apache.lucene.util.InfoStream; | ||
|
|
||
| /** Observes the real codec writer, not IndexWriter's cached buffering counters. Requires cuVS. */ | ||
| public class TestAcceleratedHNSWHostInputMemory extends LuceneTestCase { |
There was a problem hiding this comment.
Could you ensure that this an all other tests that require GPU to run as skipped on CPU-only machines?
There was a problem hiding this comment.
Done! The GPU-dependent tests added or updated here now skip when cuVS isn’t available. If cuVS reports support, they still require an accelerated writer, so silent CPU fallback can’t produce a false pass.
| /** Counts the live vectors that the merge iterator will actually yield. */ | ||
| private static int countMergedVectors(FieldInfo fieldInfo, MergeState mergeState) | ||
| throws IOException { | ||
| FloatVectorValues mergedVectors = |
There was a problem hiding this comment.
Nit: I think these iterations can be skipped and we can got with total of original sizes if there were no deletes or updates (all liveDocs are empty).
There was a problem hiding this comment.
Great suggestion, thanks — this now uses mergedVectors.size() when every liveDocs entry is null and only iterates when a source has deletions.
| writer.commit(); | ||
|
|
||
| try (DirectoryReader sourceReader = DirectoryReader.open(writer)) { | ||
| assertEquals("the test requires three source segments", 3, sourceReader.leaves().size()); |
There was a problem hiding this comment.
That fails for me with seed 5A68E1972548005F:36610B2E43B74A02
To reproduce run
LD_LIBRARY_PATH=$PWD/../../cpp/build/c:$PWD/../../cuvs/cpp/build${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH} \
mvn -B test \
-Dtest=TestAcceleratedHNSWDeletedDocuments \
-Dtests.method=testForceMergeCountsOnlyLiveSparseVectors \
-Dtests.seed=5A68E1972548005F:36610B2E43B74A02
For this check we should probably disable auto flush like we do in other tests so it doesn't interfere with segments.
There was a problem hiding this comment.
Fixed! The fixture now disables document-count auto-flush and uses a 256 MB RAM buffer with NoMergePolicy, so explicit commits control the source segments. The reported seed passed, along with 10 repeated class iterations..
|
/ok to test 3b0f875 |
@imotov, there was an error processing your request: See the following link for more information: https://docs.gha-runners.nvidia.com/cpr/e/2/ |
|
/ok to test 83a5f97 |
@imotov, there was an error processing your request: See the following link for more information: https://docs.gha-runners.nvidia.com/cpr/e/2/ |
|
/ok to test 83a5f97 |
@imotov, there was an error processing your request: See the following link for more information: https://docs.gha-runners.nvidia.com/cpr/e/2/ |
|
/ok to test 83a5f97 |
Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Summary
This ports NVIDIA/cuvs-lucene#173 into
java/cuvs-lucenefollowing the move into the cuVS monorepo.The scope is limited to the accelerated-HNSW memory-pressure change:
ramBytesUsed()and emit allocation diagnostics throughInfoStream. This observability does not update Lucene's cachedIndexWritercounters or impose a native-memory limit.List<float[]>.The resulting codec still uses GPU CAGRA construction followed by Lucene CPU HNSW search.
Merge and ownership safety
The memory change includes the safety required by its new allocation and replay path:
CuVSMatrix.Builderis closeable. The built-in builders release untransferred storage when closed, transfer ownership only after a successfulbuild(), reject use after close or build, and allow repeatedclose()calls.close()preserves compatibility for providers compiled against the earlier builder interface.AcceleratedHNSWUtils.createMultiLayerHnswGraph(...)list-based descriptor remains available; the native-matrix overload used internally is package-private.-1padding is not serialized as an HNSW neighbor.Scope boundaries
CuVS2510GPUSearchCodecremains device-backed and is otherwise unchanged by this PR.Validation
Validation completed during development of this branch:
cuvs-java: 5 unit tests and 111 integration tests; 0 failures, 0 errors, and 1 inherited skipcuvs-lucene: 348 tests; 0 failures, 0 errors, and 30 inherited skipsgit diff --checkDeep1B 1M, 96 dimensions, using fresh index directories:
checkIntegrity(), 1,000,000 documents and vectors, a complete unique source-ID domain, graph level membership and nesting, duplicate-free adjacency, finite search results, and ground-truth overlap.These Deep1B runs are correctness sanity checks. Page-cache eviction was unavailable in the validation container, so the recorded timings are not performance evidence.
Retest Results
Retested the live PR-2476 head
eb6c03e77on Deep1B 100M at 96 dimensions. Both layouts used the historical cold-source configuration,forceMerge=0, and 1,000 measured queries.1 -> 16c1755021 -> 1eb6c03e774 -> 46c1755024 -> 4eb6c03e77