Conversation
Merging this PR will regress 1 benchmark
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | random_i16[0.95] |
81.5 µs | 99.7 µs | -18.26% |
| ⚡ | WallTime | arrow_checked_add_u32_neon[16384] |
20.5 µs | 12.3 µs | +66.78% |
| ⚡ | Simulation | decode_primitives[f32, (1000, 512)] |
64.2 µs | 42 µs | +52.92% |
| ⚡ | Simulation | random_i8[0.5] |
96 µs | 72.4 µs | +32.69% |
| ⚡ | Simulation | decompress[u64, (4000, 1024)] |
86.9 µs | 71 µs | +22.34% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing mk/bitpacked-stack-04-offset-child (cfa9850) with mk/bitpacked-stack-03-width-child (bcd5b9c)2
Footnotes
-
218 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
No successful run was found on
mk/bitpacked-stack-03-width-child(d740f8f) during the generation of this report, so 553f2c9 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩
Signed-off-by: "Matt Katz" <mhkatz97@gmail.com>
355e292 to
cfa9850
Compare
Store u64 byte boundaries as a second layout child, including the trailing boundary. Scalar access reads only the needed child scalars; bulk kernels materialize and validate the layout once. Slices retain their original offset origin and validate boundaries before unpacking. Include malformed-offset and child-shape coverage.
Part 4/9 of the bitpacked-v2 stack (437 added + removed lines). Merge in order; each draft targets the preceding branch, starting from
develop.1 → 2 → 3 → 4 → 5 → 6 → 7 → 8 → 9
Validation: All 11 focused plugin, serde, and chunk-layout tests passed on this commit, including the regression that rejects direct vtable serde. Changed Rust files were formatted with nightly;
git diff --checkpassed.GPU execution was not tested. CUDA validation and workspace-wide all-feature Clippy were previously blocked by the nvCOMP 5.1 SDK download in this environment.