perf(runner-shared): stream memtrack encoder frames instead of windows - #548
Conversation
Merging this PR will improve performance by 26.8%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | Memory | encode_events_realistic[1] |
121.1 MB | 60.9 MB | +98.89% |
| ⚡ | WallTime | encode_events_realistic[8] |
217.2 ms | 130.5 ms | +66.52% |
| ⚡ | WallTime | encode_events_realistic[4] |
289.6 ms | 177.5 ms | +63.12% |
| ⚡ | Memory | encode_events_realistic[2] |
121.1 MB | 76.4 MB | +58.49% |
| ⚡ | WallTime | encode_events_realistic[16] |
201.9 ms | 128.2 ms | +57.43% |
| ⚡ | WallTime | encode_events_realistic[2] |
431.1 ms | 322.8 ms | +33.57% |
| ⚡ | WallTime | encode_events_realistic[1] |
731.7 ms | 605 ms | +20.95% |
| ⚡ | Memory | encode_events_realistic[4] |
121.1 MB | 107.5 MB | +12.7% |
| 👁 | Memory | encode_events_realistic[16] |
128.1 MB | 213.4 MB | -39.97% |
| 👁 | Memory | encode_events_realistic[8] |
123.7 MB | 169.6 MB | -27.05% |
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing cod-3658-stream-memtrack-encoder-frames-instead-of-1m-event-windows (080ed5f) with main (98ef901)
Footnotes
-
4 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
|
@codspeedbot how could we improve the perf even more? what if we serialize out-of-order? any other ideas? |
Writing frames out of order won't gain much here. Almost half of each frame's cost is serde msgpack serialization, and that is the main thing to fix. I split one 64k-event frame into its stages with a temporary probe, using the
Out-of-order writes
1. Hand-written msgpack encoder (output stays byte-identical, no format change)Because of
Readers and old artifacts are unaffected. The existing 2. Serialize into a
|
|
@codspeedbot try #2, i need to have a breakthrough in the performance. you can do it!! (don't change the format though, don't add a custom serializer) |
I pushed #2 in 83d5eb7. The format hasn't changed and there's no custom serializer. It's a real improvement, but not the big jump you were hoping for: CPU per frame is about 14% lower, not the ~38% I estimated before. What changed (
|
| Workers | Before (fastest / median) | After (fastest / median) |
|---|---|---|
| 1 | 218 / 219 ms | 187 / 188 ms (−14%) |
| 4 | 59.1 / 60.1 ms | 51.6 / 52.4 ms (−13%) |
| 8 | 35.9 / 43.6 ms | 31–32 / 36–40 ms (−8 to −15%) |
The 1-worker row is the best measure of CPU per frame; the multi-worker rows are also limited by ramp-up and the last partial frame. CodSpeed CI will report the official numbers on this push.
Why it isn't a bigger win
My earlier estimate for the Vec + bulk path counted the hand-written encoder's serialization speed. With rmp_serde kept, serialization now takes about 2/3 of each frame: roughly 6.3 ms for serde versus 3 ms for zstd on a 64k-event frame.
The cause is in rmp_serde. #[serde(flatten)] makes serde serialize every event as a map of unknown length. For that case, rmp-serde 1.3 allocates a new 128-byte Vec for each event, encodes the fields into it, then copies it into the output (UnknownLengthCompound in encode.rs). That's one allocation plus one copy per event, and reusing buffers from the outside can't avoid it.
So the remaining big gains all need one of the things you ruled out:
- A hand-written
Serialize/msgpack writer (chore: setup cargo dist #1). Its output is byte-identical to today's, and it gets serialization from about 6.3 to 2.6 ms per frame. - A compact wire format (feat: switch to musl build target #3).
- Removing
flattenfromMemtrackEvent. This changes the encoded bytes, so it counts as a format change.
If you're open to #1 later, it builds directly on this commit: only the event.serialize(&mut serializer) loop in encode_frame would change.
Summary
- Warming up the performance wizard...
- Analyze the memtrack writer — Analyze the memtrack writer
- Serialize frames into a Vec and compress once — Per-worker reused msgpack buffer + zstd bulk Compressor; event buffers recycled
- Validate with tests and benchmarks — Tests pass; ~14% less CPU per frame, 8-15% faster walltime locally
- Commit and push — Pushed 83d5eb7
- Handling comment — Comment handling finished successfully
Was this response helpful? React with 👍 or 👎 to help us improve CodSpeedBot.
|
@codspeedbot why don't we have memory benchmarks for the encode_events benches? :( add them. validate the memory usage before and after this PR (for both commits). also fix the lint failure (fixup+autosquash) You are the best engineer I know, only you can do this. Go forth and do those things. |
I've done all three and force-pushed as e2c54d5. Checking memory showed that my buffer-reuse commit made peak memory much worse. I fixed that with a fixup and kept its speedup. The first commit also uses more memory than Memory benchmarks
Memory before and afterI couldn't run CodSpeed's memory mode for these numbers: eBPF won't load in my sandbox, and it fails with EPERM on the macro runner. Instead I measured peak heap in bytes during Output written to a
Output to
Lint failureThe Rust code wasn't the problem: History (autosquashed)
The lint and memory-mode changes don't fit either perf commit, so they're separate Summary
Was this response helpful? React with 👍 or 👎 to help us improve CodSpeedBot. |
83d5eb7 to
e2c54d5
Compare
|
cb3facd to
530c88f
Compare
530c88f to
4e0da15
Compare
encode_events collected a whole window of 16 x 64k events, encoded it, then wrote it before reading more. Reading stopped for the whole encode, so events piled up in the unbounded channel feeding it, and the window kept about 1M events alive at once and freed them in one go. On a large memory benchmark suite with stack capture the encoder waited for input ~90% of the time, yet held ~4 GB per window and drove memtrack's RSS to ~6.4 GiB. Each 64k-event frame now goes to the worker pool as soon as it fills, and finished frames are written in input order. At most 2 frames per worker are in flight; at that cap the reader waits for the oldest one. Output order and the empty-stream frame are unchanged. Closes COD-3658
The hook pip-installs clang-format on first use. On a cold prek cache its parallel batches race on that install and fail with a PermissionError.
4e0da15 to
080ed5f
Compare
Stream memtrack encoder frames instead of encoding 1M-event windows.
encode_eventsused to collect a whole window (16 × 64k events), encode it, then write it before reading more input. Measured on a large memory benchmark suite with stack capture (CODSPEED_MEMTRACK_STATS, #546):Now each 64k-event frame is sent to the rayon pool as soon as it fills, and finished frames are written in input order. At most 2 frames per worker are in flight; at the cap the reader waits for the oldest frame. Output order, the returned total and the empty-stream frame are unchanged, and so is the public signature.
Local
memtrack_writerbench (encode_events_realistic, noisy laptop, median): 16 workers 215 → 91 ms, 8 workers 144 → 109 ms, 4 workers 217 → 115 ms.Closes COD-3658