perf/inline blake3 round - #557
Draft
gabriel-barrett wants to merge 8 commits into
Draft
Conversation
gabriel-barrett
force-pushed
the
perf/inline-blake3-round
branch
3 times, most recently
from
August 13, 2026 16:25
2b2a4fd to
1617d60
Compare
gabriel-barrett
force-pushed
the
perf/inline-blake3-round
branch
from
August 13, 2026 17:21
1617d60 to
cba177f
Compare
gabriel-barrett
force-pushed
the
perf/inline-blake3-round
branch
2 times, most recently
from
August 13, 2026 17:24
1f5a367 to
2b83e35
Compare
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
Plus the trailing integration: the aiur test toplevels migrate off the removed builtins (u8_chain_rotr7/4 wrappers and the u32_rotr7 fixture become u8_xor_split7/4 wrappers and a two-operand u32_xor_rotr7, with freshly computed vectors), the new xor_split gadget helpers get clippy-clean byte casts, codegen regenerated, and all 71 kernel FFT pins + the shard aggregate re-measured (median -0.4%, shard -1.3%). aiur (cargo + prove + cross agreement), ixvm (pins + parity), multi-stark and recursive-verifier suites pass; fmt, clippy, deny clean.
gabriel-barrett
force-pushed
the
perf/inline-blake3-round
branch
from
August 13, 2026 21:20
2b83e35 to
3d03094
Compare
Member
Author
|
!benchmark aiur-recursive fresh |
|
| constant | recursive-prove-time (main) | recursive-prove-time (PR) | Δ% | recursive-peak-ram (main) | recursive-peak-ram (PR) | Δ% | recursive-proof-size (main) | recursive-proof-size (PR) | Δ% | recursive-verify-time (main) | recursive-verify-time (PR) | Δ% | recursive-execute-time (main) | recursive-execute-time (PR) | Δ% | recursive-fft-cost (main) | recursive-fft-cost (PR) | Δ% | prove-time (main) | prove-time (PR) | Δ% | proof-size (main) | proof-size (PR) | Δ% | verify-time (main) | verify-time (PR) | Δ% | peak-ram (main) | peak-ram (PR) | Δ% |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
String.split |
37.804 s | 32.754 s | -13.4% (1.15× faster) 🟢 | 107.69 GiB | 81.72 GiB | -24.1% (1.32× smaller) 🟢 | 5.30 MiB | 6.27 MiB | +18.3% (1.18× larger) |
33.5 ms | 39.9 ms | +19.4% (1.19× slower) |
6.031 s | 5.052 s | -16.2% (1.19× faster) 🟢 | 180.38B | 143.00B | -20.7% (1.26× fewer) 🟢 | 16.609 s | 15.583 s | -6.2% (1.07× faster) 🟢 | 11.21 MiB | 12.14 MiB | +8.3% (1.08× larger) |
69.3 ms | 72.5 ms | +4.6% |
32.79 GiB | 29.81 GiB | -9.1% (1.10× smaller) 🟢 |
Vector.append |
34.530 s | 29.134 s | -15.6% (1.19× faster) 🟢 | 95.91 GiB | 72.66 GiB | -24.2% (1.32× smaller) 🟢 | 5.30 MiB | 6.23 MiB | +17.5% (1.18× larger) |
32.5 ms | 37.2 ms | +14.6% (1.15× slower) |
5.182 s | 4.353 s | -16.0% (1.19× faster) 🟢 | 157.02B | 126.74B | -19.3% (1.24× fewer) 🟢 | 4.029 s | 3.930 s | -2.5% | 10.42 MiB | 11.34 MiB | +8.9% (1.09× larger) |
66.3 ms | 70.2 ms | +5.9% (1.06× slower) |
6.95 GiB | 7.01 GiB | +0.8% |
Nat.add_comm |
22.163 s | 18.671 s | -15.8% (1.19× faster) 🟢 | 58.49 GiB | 45.34 GiB | -22.5% (1.29× smaller) 🟢 | 5.30 MiB | 6.23 MiB | +17.5% (1.18× larger) |
35.0 ms | 36.5 ms | +4.2% |
3.845 s | 3.160 s | -17.8% (1.22× faster) 🟢 | 113.85B | 92.98B | -18.3% (1.22× fewer) 🟢 | 1.048 s | 1.080 s | +3.1% |
8.93 MiB | 9.86 MiB | +10.4% (1.10× larger) |
53.1 ms | 61.2 ms | +15.3% (1.15× slower) |
3.90 GiB | 3.55 GiB | -9.1% (1.10× smaller) 🟢 |
Member
Author
|
!benchmark aiur fresh |
|
| constant | prove-time (main) | prove-time (PR) | Δ% | throughput (const/s) (main) | throughput (const/s) (PR) | Δ% | peak-ram (main) | peak-ram (PR) | Δ% | execute-time (main) | execute-time (PR) | Δ% | verify-time (main) | verify-time (PR) | Δ% | proof-size (main) | proof-size (PR) | Δ% | fft-cost (main) | fft-cost (PR) | Δ% |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Array.extract_append |
36.925 s | 36.527 s | -1.1% | 41.760 | 42.210 | +1.1% | 73.67 GiB | 70.62 GiB | -4.1% 🟢 | 9.436 s | 9.428 s | -0.1% | 135.7 ms | 159.6 ms | +17.7% (1.18× slower) |
21.59 MiB | 23.38 MiB | +8.3% (1.08× larger) |
136.02B | 130.81B | -3.8% 🟢 |
ByteArray.utf8DecodeChar?_utf8EncodeChar_append |
36.311 s | 34.193 s | -5.8% (1.06× faster) 🟢 | 74.110 | 78.700 | +6.2% (1.06× faster) 🟢 | 73.14 GiB | 67.22 GiB | -8.1% (1.09× smaller) 🟢 | 9.603 s | 9.278 s | -3.4% 🟢 | 135.8 ms | 162.4 ms | +19.6% (1.20× slower) |
21.69 MiB | 23.47 MiB | +8.2% (1.08× larger) |
138.80B | 126.89B | -8.6% (1.09× fewer) 🟢 |
Char.ofOrdinal_le_of_le |
28.193 s | 28.101 s | -0.3% | 93.850 | 94.160 | +0.3% | 58.52 GiB | 58.52 GiB | +0.0% | 6.665 s | 6.326 s | -5.1% (1.05× faster) 🟢 | 129.5 ms | 144.4 ms | +11.5% (1.12× slower) |
21.64 MiB | 23.42 MiB | +8.3% (1.08× larger) |
99.42B | 90.63B | -8.8% (1.10× fewer) 🟢 |
Vector.extract_append._proof_2 |
20.704 s | 20.059 s | -3.1% 🟢 | 62.890 | 64.910 | +3.2% 🟢 | 39.72 GiB | 38.18 GiB | -3.9% 🟢 | 5.131 s | 5.009 s | -2.4% | 134.0 ms | 147.8 ms | +10.3% (1.10× slower) |
21.30 MiB | 23.08 MiB | +8.4% (1.08× larger) |
77.08B | 73.31B | -4.9% (1.05× fewer) 🟢 |
_private.Init.Data.Range.Polymorphic.SInt.0.Int64.instRxcHasSize_eq |
16.764 s | 15.781 s | -5.9% (1.06× faster) 🟢 | 108.090 | 114.820 | +6.2% (1.06× faster) 🟢 | 34.80 GiB | 31.83 GiB | -8.5% (1.09× smaller) 🟢 | 3.499 s | 3.391 s | -3.1% 🟢 | 138.4 ms | 137.8 ms | -0.4% | 21.47 MiB | 23.26 MiB | +8.3% (1.08× larger) |
54.85B | 49.57B | -9.6% (1.11× fewer) 🟢 |
String.split |
15.766 s | 14.647 s | -7.1% (1.08× faster) 🟢 | 111.950 | 120.500 | +7.6% (1.08× faster) 🟢 | 32.71 GiB | 29.74 GiB | -9.1% (1.10× smaller) 🟢 | 3.325 s | 3.109 s | -6.5% (1.07× faster) 🟢 | 131.1 ms | 139.3 ms | +6.3% (1.06× slower) |
21.68 MiB | 23.46 MiB | +8.2% (1.08× larger) |
50.25B | 45.59B | -9.3% (1.10× fewer) 🟢 |
List.mergeSort |
11.706 s | 11.153 s | -4.7% 🟢 | 123.700 | 129.830 | +5.0% 🟢 | 23.76 GiB | 22.27 GiB | -6.3% (1.07× smaller) 🟢 | 2.335 s | 2.305 s | -1.3% | 127.8 ms | 138.9 ms | +8.7% (1.09× slower) |
21.53 MiB | 23.32 MiB | +8.3% (1.08× larger) |
36.10B | 32.44B | -10.1% (1.11× fewer) 🟢 |
Vector.append |
3.865 s | 3.759 s | -2.8% | 125.480 | 129.040 | +2.8% | 6.58 GiB | 7.02 GiB | +6.6% (1.07× larger) |
668.4 ms | 639.4 ms | -4.3% 🟢 | 131.2 ms | 133.4 ms | +1.7% | 20.13 MiB | 21.92 MiB | +8.9% (1.09× larger) |
8.02B | 7.24B | -9.7% (1.11× fewer) 🟢 |
Nat.gcd_comm |
3.135 s | 2.975 s | -5.1% (1.05× faster) 🟢 | 124.410 | 131.070 | +5.4% (1.05× faster) 🟢 | 5.96 GiB | 6.11 GiB | +2.6% | 512.5 ms | 497.8 ms | -2.9% | 123.3 ms | 127.2 ms | +3.1% |
19.90 MiB | 21.68 MiB | +9.0% (1.09× larger) |
5.25B | 4.60B | -12.4% (1.14× fewer) 🟢 |
String.append |
2.339 s | 2.250 s | -3.8% 🟢 | 129.970 | 135.110 | +4.0% 🟢 | 4.66 GiB | 4.58 GiB | -1.8% | 392.4 ms | 379.2 ms | -3.4% 🟢 | 122.6 ms | 130.0 ms | +6.1% (1.06× slower) |
19.17 MiB | 20.96 MiB | +9.3% (1.09× larger) |
2.90B | 2.55B | -12.1% (1.14× fewer) 🟢 |
Int.gcd |
1.954 s | 1.956 s | +0.1% | 106.460 | 106.340 | -0.1% | 4.55 GiB | 4.92 GiB | +8.2% (1.08× larger) |
343.9 ms | 334.3 ms | -2.8% | 119.6 ms | 119.1 ms | -0.4% | 18.73 MiB | 20.51 MiB | +9.5% (1.10× larger) |
1.87B | 1.64B | -11.9% (1.14× fewer) 🟢 |
Nat.sub_le_of_le_add |
1.763 s | 1.754 s | -0.5% | 96.410 | 96.940 | +0.5% | 4.96 GiB | 4.54 GiB | -8.5% (1.09× smaller) 🟢 | 341.8 ms | 328.7 ms | -3.8% 🟢 | 113.1 ms | 126.9 ms | +12.3% (1.12× slower) |
19.08 MiB | 20.87 MiB | +9.4% (1.09× larger) |
1.60B | 1.42B | -11.5% (1.13× fewer) 🟢 |
Nat.add_comm |
1.051 s | 1.024 s | -2.5% | 39.970 | 41 | +2.6% | 3.92 GiB | 4.17 GiB | +6.2% (1.06× larger) |
245.7 ms | 241.0 ms | -1.9% | 103.9 ms | 109.7 ms | +5.6% (1.06× slower) |
17.25 MiB | 19.04 MiB | +10.4% (1.10× larger) |
271.37M | 249.57M | -8.0% (1.09× fewer) 🟢 |
gabriel-barrett
force-pushed
the
perf/inline-blake3-round
branch
from
August 14, 2026 10:56
3d03094 to
af38029
Compare
This reverts commit ac1cc6d.
Rebased onto the fused xor-rotation branch: the seven inline calls now compose with the two-operand u32_xor_rotr7/12 forms. Codegen regenerated and all 71 kernel FFT pins + the shard aggregate re-measured on top of the new base (median -5.4%, range -11.4%..-0.2%; shard 6.08B -> 5.21B, -14.4%). aiur (cargo + prove + cross), ixvm (pins + parity), multi-stark and recursive-verifier suites pass; fmt, clippy, deny clean.
gabriel-barrett
force-pushed
the
perf/inline-blake3-round
branch
from
August 14, 2026 11:35
af38029 to
d6f0857
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Inline
blake3_compress_inner_j. This has the effect of increasing the total width, though not at the same proportion as what we gain on efficiency, for the lookup itself is costly, especially when the input arguments are big.Eventually we will have to decrease the proof size of the recursive verifier, but there are ways around it see #551
Since
blake3_compress