Conversation
Merging this PR will regress 2 benchmarks
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | bitpacked_compress_u32 |
36.5 µs | 47.1 µs | -22.52% |
| ❌ | Simulation | random_i16[0.8] |
79.4 µs | 97.8 µs | -18.85% |
| ⚡ | WallTime | arrow_checked_add_u32_neon[16384] |
20.4 µs | 12.3 µs | +66.78% |
| ⚡ | Simulation | decode_primitives[f32, (1000, 512)] |
64.2 µs | 42 µs | +52.92% |
| ⚡ | Simulation | random_i8[0.5] |
96 µs | 72.4 µs | +32.6% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing mk/bitpacked-stack-03-width-child (d740f8f) with mk/bitpacked-stack-02-cpu-layout (9737c00)2
Footnotes
-
218 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
No successful run was found on
mk/bitpacked-stack-02-cpu-layout(7390ffe) during the generation of this report, so 64a967c was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩
Signed-off-by: "Matt Katz" <mhkatz97@gmail.com>
bcd5b9c to
d740f8f
Compare
Replace the in-memory scalar width with a non-nullable u8 child. Uniform encoders use a constant child; bulk kernels prepare the layout once per operation. V1 serialization omits the child. Update callers and CUDA's Rust adapters; CUDA continues to accept only uniform widths.
Part 3/9 of the bitpacked-v2 stack (700 added + removed lines). Merge in order; each draft targets the preceding branch, starting from
develop.1 → 2 → 3 → 4 → 5 → 6 → 7 → 8 → 9
Validation: All 6 focused plugin, serde, and chunk-layout tests passed on this commit, including the regression that rejects direct vtable serde. Changed Rust files were formatted with nightly;
git diff --checkpassed.GPU execution was not tested. CUDA validation and workspace-wide all-feature Clippy were previously blocked by the nvCOMP 5.1 SDK download in this environment.