Skip to content

perf(parquet codec): write one row group per 8192 rows - #49

Closed
vaclav-dvorak wants to merge 2 commits into
v0.54.0-exaforcefrom
sundar-review/parquet-row-group-chunking
Closed

vaclav-dvorak wants to merge 2 commits into
v0.54.0-exaforcefrom
sundar-review/parquet-row-group-chunking

Conversation

@vaclav-dvorak

@vaclav-dvorak vaclav-dvorak commented Sep 23, 2026 •

Copy link
Copy Markdown

Summary

  • encode() opens one row group for the whole batch, so each of the 78 CloudTrail columns re-walks every event in it. This writes a row group per 8192 rows instead.
  • +16% to +33% at 1-4 threads, neutral (-1% to -3%) at 8-12. Memory is unchanged -- still one column buffered at a time, which matters because most tenants run vector with a 10Gi limit.
  • It does not fix the multi-core ceiling. Encoding stops scaling past ~4-8 threads whatever the row group size. That is the more important result and it is unresolved.

Details

Two benches, both against the real 78-column cloudtrail.schema copied from chronoforge.

parquet_encode (criterion, single-threaded, per-iteration format! corpus):

rows per row group 10k batch 50k batch scaling
unbounded (current) 58.8 K/s 49.7 K/s degrades with batch size
8192 65.1 K/s 55.8 K/s degrades
1024 106.2 K/s 103.8 K/s flat

parquet_threads (thread scaling, events backed by ONE shared Bytes blob):

threads unbounded 8192 1024
1 68,630 86,507 131,132
2 112,667 130,693 200,051
4 148,126 196,582 224,582
8 163,076 161,539 120,082
12 146,021 142,205 119,522

1024 looks much better single-threaded and is a 26% regression at 8 threads, so it is not the value to ship. Small row groups mean far more page and dictionary flushes, and that allocation churn appears to contend across threads. 8192 keeps the low-thread win without regressing higher counts.

The shared-blob corpus matters: events parsed from one S3 object hold Bytes that all slice the same allocation, so clone/drop during encoding hits one refcount. The criterion corpus builds each value with format!, giving every value its own refcount, and cannot show this at all.

Motivation: gainsight's vector wedged on 2026-09-23, all 32 cores at zero throughput with a 145k SQS backlog. perf record -F 99 -a on the live pod:

34.18%  vector  __aarch64_ldadd8_rel
30.95%  vector  __aarch64_ldadd8_relax
 9.39%  vector  <[T]>::to_vec_in::ConvertVec>::to_vec
 2.27%  vector  bytes::bytes::shared_clone
 2.21%  vector  <vrl::value::value::Value as Clone>::clone
 2.18%  vector  drop_in_place<vrl::value::value::Value>

32 threads in state R with wchan=0, 32 voluntary context switches for the entire process lifetime, flat read_bytes: no I/O, no blocking, pure coherence traffic. Capping vector to 8 threads stopped the livelock but left the tenant under capacity.

Caveats:

  • Measured on a 12-core laptop (6 performance + 6 efficiency cores), not a 32-core node. The 8-thread inversion may sit elsewhere on server hardware.
  • Read side unmeasured: 8192-row groups give ~12 row groups per 100k-event file; footer metadata grows and DuckDB/Athena read performance was not tested.
  • Synthetic corpus. Column cardinality is set per field (30 regions, 200 event sources, unique IDs) after a first version used unique-per-row strings, which inflated dictionary growth with row count and contaminated the scaling measurement.

All 13 existing parquet codec tests pass, including round-trip and compression.

The parquet.rs diff is 17 semantic lines; the rest is cargo fmt re-indenting the block that gained a nesting level (git diff -w).

@vaclav-dvorak vaclav-dvorak changed the title perf(parquet codec): write one row group per 1024 rows perf(parquet codec): write one row group per 8192 rows Sep 23, 2026
@vaclav-dvorak

Copy link
Copy Markdown
Author

Closing: the thread-scaling bench shows this is a regression where it matters. At 1024 rows it is 26% slower at 8 threads; 8192 is only neutral at 8-12. All measurements were on a 12-core laptop, while the fleet runs vector with up to 64 threads on superset nodes, so there is no basis to claim it is safe at that width.

Throughput peaks at 4-8 threads and declines after, regardless of row group size. Capping vector's thread count is the fix; the row group change does not address the ceiling.

@vaclav-dvorak
vaclav-dvorak deleted the sundar-review/parquet-row-group-chunking branch September 23, 2026 18:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant