Skip to content

lowram: Keep w1 in packed form during signature generation - #1345

Open
mkannwischer wants to merge 1 commit into
mainfrom
lowram-sign-pack-w1
Open

lowram: Keep w1 in packed form during signature generation#1345
mkannwischer wants to merge 1 commit into
mainfrom
lowram-sign-pack-w1

Conversation

@mkannwischer

@mkannwischer mkannwischer commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

The signing attempt held w1 as a full mld_polyveck, but w1 is only needed as the w1Encode input to H and, coefficient-wise, as the a1 argument of MakeHint. Both work from the packed encoding, so Decompose and w1Encode now run one polynomial at a time into a MLDSA_K * MLDSA_POLYW1_PACKEDBYTES buffer, and mld_pack_sig_h recovers each row with the new mld_polyw1_unpack.

With w1 gone, the matrix-vector scratch is the sole remaining user of the buffer the two shared. In REDUCE_RAM mode that scratch is a single polynomial rather than a polyvecl, and it now shares storage with z, so the buffer disappears.

Signing allocation in bytes:

default REDUCE_RAM
ML-DSA-44 44704 -> 44448 13120 -> 9792
ML-DSA-65 69312 -> 68032 17248 -> 11872
ML-DSA-87 108224 -> 107200 21344 -> 14176

mld_polyveck_decompose and mld_polyveck_pack_w1 have no other callers and are replaced by the fused mld_polyveck_decompose_pack_w1.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mac Mini (M1, 2020) benchmarks (opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 46502 cycles 46534 cycles 1.00
ML-DSA-44 sign 131401 cycles 131164 cycles 1.00
ML-DSA-44 verify 47327 cycles 47342 cycles 1.00
ML-DSA-65 keypair 81701 cycles 81716 cycles 1.00
ML-DSA-65 sign 215390 cycles 215441 cycles 1.00
ML-DSA-65 verify 79329 cycles 79334 cycles 1.00
ML-DSA-87 keypair 132436 cycles 132496 cycles 1.00
ML-DSA-87 sign 277552 cycles 277435 cycles 1.00
ML-DSA-87 verify 133580 cycles 133546 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mac Mini (M1, 2020) benchmarks (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 112924 cycles 112777 cycles 1.00
ML-DSA-44 sign 403184 cycles 401308 cycles 1.00
ML-DSA-44 verify 119532 cycles 119390 cycles 1.00
ML-DSA-65 keypair 193857 cycles 192926 cycles 1.00
ML-DSA-65 sign 651728 cycles 649924 cycles 1.00
ML-DSA-65 verify 193025 cycles 193010 cycles 1.00
ML-DSA-87 keypair 318845 cycles 318891 cycles 1.00
ML-DSA-87 sign 831120 cycles 828752 cycles 1.00
ML-DSA-87 verify 321851 cycles 321752 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 3rd gen (c6a)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 52243 cycles 52005 cycles 1.00
ML-DSA-44 sign 156689 cycles 155314 cycles 1.01
ML-DSA-44 verify 54300 cycles 54379 cycles 1.00
ML-DSA-65 keypair 89232 cycles 89891 cycles 0.99
ML-DSA-65 sign 252045 cycles 255341 cycles 0.99
ML-DSA-65 verify 89122 cycles 89591 cycles 0.99
ML-DSA-87 keypair 143951 cycles 142178 cycles 1.01
ML-DSA-87 sign 310297 cycles 310461 cycles 1.00
ML-DSA-87 verify 139225 cycles 139090 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton5

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 51564 cycles 51457 cycles 1.00
ML-DSA-44 sign 162252 cycles 161886 cycles 1.00
ML-DSA-44 verify 54672 cycles 54811 cycles 1.00
ML-DSA-65 keypair 90155 cycles 90177 cycles 1.00
ML-DSA-65 sign 266116 cycles 268015 cycles 0.99
ML-DSA-65 verify 89823 cycles 89977 cycles 1.00
ML-DSA-87 keypair 145851 cycles 145957 cycles 1.00
ML-DSA-87 sign 335809 cycles 335685 cycles 1.00
ML-DSA-87 verify 145427 cycles 144930 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 4th gen (c7a)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 47026 cycles 46740 cycles 1.01
ML-DSA-44 sign 139819 cycles 145382 cycles 0.96
ML-DSA-44 verify 49294 cycles 51476 cycles 0.96
ML-DSA-65 keypair 82252 cycles 82956 cycles 0.99
ML-DSA-65 sign 227024 cycles 229646 cycles 0.99
ML-DSA-65 verify 82074 cycles 82585 cycles 0.99
ML-DSA-87 keypair 130391 cycles 129744 cycles 1.00
ML-DSA-87 sign 280889 cycles 280564 cycles 1.00
ML-DSA-87 verify 129532 cycles 128041 cycles 1.01

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A76 (Raspberry Pi 5) benchmarks (opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 112309 cycles 112045 cycles 1.00
ML-DSA-44 sign 354052 cycles 353393 cycles 1.00
ML-DSA-44 verify 117053 cycles 117282 cycles 1.00
ML-DSA-65 keypair 194750 cycles 194574 cycles 1.00
ML-DSA-65 sign 583585 cycles 583697 cycles 1.00
ML-DSA-65 verify 192910 cycles 193189 cycles 1.00
ML-DSA-87 keypair 320601 cycles 320560 cycles 1.00
ML-DSA-87 sign 747320 cycles 747028 cycles 1.00
ML-DSA-87 verify 318394 cycles 318140 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton5 (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 100288 cycles 100193 cycles 1.00
ML-DSA-44 sign 369472 cycles 369630 cycles 1.00
ML-DSA-44 verify 110068 cycles 110334 cycles 1.00
ML-DSA-65 keypair 176009 cycles 175189 cycles 1.00
ML-DSA-65 sign 596526 cycles 597631 cycles 1.00
ML-DSA-65 verify 177530 cycles 177767 cycles 1.00
ML-DSA-87 keypair 286614 cycles 286470 cycles 1.00
ML-DSA-87 sign 758280 cycles 759318 cycles 1.00
ML-DSA-87 verify 295698 cycles 296310 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 3rd gen (c6a) (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 133444 cycles 133190 cycles 1.00
ML-DSA-44 sign 520204 cycles 518479 cycles 1.00
ML-DSA-44 verify 146679 cycles 146665 cycles 1.00
ML-DSA-65 keypair 223925 cycles 224540 cycles 1.00
ML-DSA-65 sign 847147 cycles 844646 cycles 1.00
ML-DSA-65 verify 234409 cycles 234394 cycles 1.00
ML-DSA-87 keypair 367457 cycles 367714 cycles 1.00
ML-DSA-87 sign 1062489 cycles 1061302 cycles 1.00
ML-DSA-87 verify 380562 cycles 381497 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 4th gen (c7i)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 43230 cycles 43292 cycles 1.00
ML-DSA-44 sign 130048 cycles 130069 cycles 1.00
ML-DSA-44 verify 45248 cycles 45143 cycles 1.00
ML-DSA-65 keypair 75909 cycles 75764 cycles 1.00
ML-DSA-65 sign 213758 cycles 213624 cycles 1.00
ML-DSA-65 verify 74418 cycles 74335 cycles 1.00
ML-DSA-87 keypair 123062 cycles 122992 cycles 1.00
ML-DSA-87 sign 271631 cycles 271188 cycles 1.00
ML-DSA-87 verify 120766 cycles 120668 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 3rd gen (c6i)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 61412 cycles 61655 cycles 1.00
ML-DSA-44 sign 189046 cycles 188845 cycles 1.00
ML-DSA-44 verify 66231 cycles 66290 cycles 1.00
ML-DSA-65 keypair 109628 cycles 112698 cycles 0.97
ML-DSA-65 sign 312172 cycles 318478 cycles 0.98
ML-DSA-65 verify 109663 cycles 111904 cycles 0.98
ML-DSA-87 keypair 170175 cycles 171213 cycles 0.99
ML-DSA-87 sign 379730 cycles 379008 cycles 1.00
ML-DSA-87 verify 170594 cycles 170855 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 4th gen (c7a) (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 118667 cycles 118359 cycles 1.00
ML-DSA-44 sign 457628 cycles 458607 cycles 1.00
ML-DSA-44 verify 131501 cycles 129909 cycles 1.01
ML-DSA-65 keypair 201384 cycles 200886 cycles 1.00
ML-DSA-65 sign 740800 cycles 744497 cycles 1.00
ML-DSA-65 verify 208939 cycles 208873 cycles 1.00
ML-DSA-87 keypair 331948 cycles 331103 cycles 1.00
ML-DSA-87 sign 936713 cycles 938146 cycles 1.00
ML-DSA-87 verify 343584 cycles 345239 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton4

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 67496 cycles 67139 cycles 1.01
ML-DSA-44 sign 198509 cycles 198230 cycles 1.00
ML-DSA-44 verify 70148 cycles 70284 cycles 1.00
ML-DSA-65 keypair 119095 cycles 119469 cycles 1.00
ML-DSA-65 sign 325504 cycles 326490 cycles 1.00
ML-DSA-65 verify 116584 cycles 116939 cycles 1.00
ML-DSA-87 keypair 196377 cycles 196410 cycles 1.00
ML-DSA-87 sign 421551 cycles 421431 cycles 1.00
ML-DSA-87 verify 193276 cycles 192954 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 4th gen (c7i) (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 92216 cycles 91733 cycles 1.01
ML-DSA-44 sign 349576 cycles 351109 cycles 1.00
ML-DSA-44 verify 99758 cycles 99446 cycles 1.00
ML-DSA-65 keypair 154324 cycles 153920 cycles 1.00
ML-DSA-65 sign 567577 cycles 571127 cycles 0.99
ML-DSA-65 verify 159730 cycles 160330 cycles 1.00
ML-DSA-87 keypair 255364 cycles 255096 cycles 1.00
ML-DSA-87 sign 725239 cycles 721477 cycles 1.01
ML-DSA-87 verify 263819 cycles 264366 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 3rd gen (c6i) (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 154136 cycles 154277 cycles 1.00
ML-DSA-44 sign 588309 cycles 588660 cycles 1.00
ML-DSA-44 verify 168888 cycles 169423 cycles 1.00
ML-DSA-65 keypair 262317 cycles 267849 cycles 0.98
ML-DSA-65 sign 958311 cycles 979945 cycles 0.98
ML-DSA-65 verify 271809 cycles 277179 cycles 0.98
ML-DSA-87 keypair 431619 cycles 432934 cycles 1.00
ML-DSA-87 sign 1211290 cycles 1215004 cycles 1.00
ML-DSA-87 verify 446915 cycles 447425 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A76 (Raspberry Pi 5) benchmarks (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 211677 cycles 211596 cycles 1.00
ML-DSA-44 sign 761225 cycles 759367 cycles 1.00
ML-DSA-44 verify 228934 cycles 229045 cycles 1.00
ML-DSA-65 keypair 375171 cycles 375197 cycles 1.00
ML-DSA-65 sign 1245358 cycles 1247878 cycles 1.00
ML-DSA-65 verify 371103 cycles 371441 cycles 1.00
ML-DSA-87 keypair 600704 cycles 600145 cycles 1.00
ML-DSA-87 sign 1585805 cycles 1584430 cycles 1.00
ML-DSA-87 verify 616453 cycles 615831 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton4 (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 128011 cycles 127916 cycles 1.00
ML-DSA-44 sign 441684 cycles 441335 cycles 1.00
ML-DSA-44 verify 136368 cycles 136365 cycles 1.00
ML-DSA-65 keypair 223088 cycles 221710 cycles 1.01
ML-DSA-65 sign 714012 cycles 714085 cycles 1.00
ML-DSA-65 verify 220581 cycles 220606 cycles 1.00
ML-DSA-87 keypair 364536 cycles 365306 cycles 1.00
ML-DSA-87 sign 914776 cycles 916248 cycles 1.00
ML-DSA-87 verify 371022 cycles 370929 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton3

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 71530 cycles 71457 cycles 1.00
ML-DSA-44 sign 209907 cycles 208907 cycles 1.00
ML-DSA-44 verify 74817 cycles 75005 cycles 1.00
ML-DSA-65 keypair 125984 cycles 125992 cycles 1.00
ML-DSA-65 sign 344118 cycles 345143 cycles 1.00
ML-DSA-65 verify 124046 cycles 124143 cycles 1.00
ML-DSA-87 keypair 206428 cycles 206682 cycles 1.00
ML-DSA-87 sign 439635 cycles 444029 cycles 0.99
ML-DSA-87 verify 204505 cycles 204086 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton3 (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 138537 cycles 138455 cycles 1.00
ML-DSA-44 sign 486610 cycles 486114 cycles 1.00
ML-DSA-44 verify 149235 cycles 149248 cycles 1.00
ML-DSA-65 keypair 241159 cycles 242184 cycles 1.00
ML-DSA-65 sign 790421 cycles 791620 cycles 1.00
ML-DSA-65 verify 241516 cycles 241529 cycles 1.00
ML-DSA-87 keypair 395244 cycles 396160 cycles 1.00
ML-DSA-87 sign 1011999 cycles 1013640 cycles 1.00
ML-DSA-87 verify 404107 cycles 403824 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton2

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 112444 cycles 112659 cycles 1.00
ML-DSA-44 sign 354470 cycles 354364 cycles 1.00
ML-DSA-44 verify 117205 cycles 117531 cycles 1.00
ML-DSA-65 keypair 195775 cycles 195098 cycles 1.00
ML-DSA-65 sign 587160 cycles 584870 cycles 1.00
ML-DSA-65 verify 194231 cycles 193581 cycles 1.00
ML-DSA-87 keypair 320708 cycles 321321 cycles 1.00
ML-DSA-87 sign 747632 cycles 747338 cycles 1.00
ML-DSA-87 verify 318503 cycles 318952 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton2 (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 212383 cycles 212400 cycles 1.00
ML-DSA-44 sign 761080 cycles 760457 cycles 1.00
ML-DSA-44 verify 229961 cycles 229630 cycles 1.00
ML-DSA-65 keypair 376776 cycles 375851 cycles 1.00
ML-DSA-65 sign 1247361 cycles 1248638 cycles 1.00
ML-DSA-65 verify 372562 cycles 371900 cycles 1.00
ML-DSA-87 keypair 602074 cycles 602038 cycles 1.00
ML-DSA-87 sign 1586617 cycles 1591124 cycles 1.00
ML-DSA-87 verify 618303 cycles 618524 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A72 (Raspberry Pi 4) benchmarks (opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 218089 cycles 213397 cycles 1.02
ML-DSA-44 sign 599197 cycles 587414 cycles 1.02
ML-DSA-44 verify 217607 cycles 213073 cycles 1.02
ML-DSA-65 keypair 379744 cycles 376687 cycles 1.01
ML-DSA-65 sign 986692 cycles 979172 cycles 1.01
ML-DSA-65 verify 364473 cycles 363486 cycles 1.00
ML-DSA-87 keypair 639246 cycles 647130 cycles 0.99
ML-DSA-87 sign 1311852 cycles 1349859 cycles 0.97
ML-DSA-87 verify 619801 cycles 632424 cycles 0.98

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A72 (Raspberry Pi 4) benchmarks (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 300220 cycles 297307 cycles 1.01
ML-DSA-44 sign 1133706 cycles 1124831 cycles 1.01
ML-DSA-44 verify 331808 cycles 327013 cycles 1.01
ML-DSA-65 keypair 549132 cycles 547240 cycles 1.00
ML-DSA-65 sign 1869739 cycles 1861858 cycles 1.00
ML-DSA-65 verify 530042 cycles 527234 cycles 1.01
ML-DSA-87 keypair 849544 cycles 841613 cycles 1.01
ML-DSA-87 sign 2351348 cycles 2350913 cycles 1.00
ML-DSA-87 verify 880873 cycles 873805 cycles 1.01

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A55 (Snapdragon 888) benchmarks (opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 273189 cycles 271602 cycles 1.01
ML-DSA-44 sign 813853 cycles 811152 cycles 1.00
ML-DSA-44 verify 273807 cycles 273547 cycles 1.00
ML-DSA-65 keypair 465037 cycles 466624 cycles 1.00
ML-DSA-65 sign 1332507 cycles 1341019 cycles 0.99
ML-DSA-65 verify 453524 cycles 452548 cycles 1.00
ML-DSA-87 keypair 799137 cycles 798726 cycles 1.00
ML-DSA-87 sign 1842572 cycles 1835055 cycles 1.00
ML-DSA-87 verify 787003 cycles 773446 cycles 1.02

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A55 (Snapdragon 888) benchmarks (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 466035 cycles 465745 cycles 1.00
ML-DSA-44 sign 2138837 cycles 2140830 cycles 1.00
ML-DSA-44 verify 556780 cycles 558445 cycles 1.00
ML-DSA-65 keypair 783328 cycles 784167 cycles 1.00
ML-DSA-65 sign 3498055 cycles 3497460 cycles 1.00
ML-DSA-65 verify 868982 cycles 867417 cycles 1.00
ML-DSA-87 keypair 1266853 cycles 1272563 cycles 1.00
ML-DSA-87 sign 4330173 cycles 4331369 cycles 1.00
ML-DSA-87 verify 1392739 cycles 1392417 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SpacemiT K1 8 (Banana Pi F3) benchmarks (no-opt)

Details
Benchmark suite Current: 7c8238d Previous: 41c00da Ratio
ML-DSA-44 keypair 757718 cycles 760195 cycles 1.00
ML-DSA-44 sign 3137744 cycles 3141418 cycles 1.00
ML-DSA-44 verify 857097 cycles 859027 cycles 1.00
ML-DSA-65 keypair 1288623 cycles 1288559 cycles 1.00
ML-DSA-65 sign 5092434 cycles 5082523 cycles 1.00
ML-DSA-65 verify 1368044 cycles 1368131 cycles 1.00
ML-DSA-87 keypair 2112748 cycles 2109622 cycles 1.00
ML-DSA-87 sign 6377925 cycles 6359264 cycles 1.00
ML-DSA-87 verify 2226700 cycles 2225228 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot

oqs-bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-87, REDUCE-RAM)

⚠️ Attention Required

Proof Status Current Previous Change
compute_pack_t0_t1 ⚠️ 27s 12s +125%
mld_attempt_signature_generation ⚠️ 151s 34s +344%
sign_keypair_internal ⚠️ 47s 6s +683%
sign_pk_from_sk ⚠️ 61s 6s +917%
sign_verify_internal ⚠️ 195s 47s +315%
Full Results (212 proofs)
Proof Status Current Previous Change
**TOTAL** 1336s 1436s -7.0%
sign_verify_internal ⚠️ 195s 47s +315%
mld_attempt_signature_generation ⚠️ 151s 34s +344%
polyvec_matrix_pointwise_montgomery_yvec 115s 196s -41%
poly_pointwise_montgomery_c 63s 125s -50%
sign_pk_from_sk ⚠️ 61s 6s +917%
mld_invntt_layer 54s 111s -51%
sign_keypair_internal ⚠️ 47s 6s +683%
compute_pack_t0_t1 ⚠️ 27s 12s +125%
fqmul 22s 39s -44%
mld_ntt_layer 20s 44s -55%
sign_signature_internal 14s 3s +367%
keccakf1600x4_permute_native 12s 25s -52%
mld_ntt_butterfly_block 12s 23s -48%
sig_unpack_hints 12s 4s +200%
poly_ntt_c 10s 19s -47%
polyt0_unpack 9s 14s -36%
rej_uniform_c 9s 18s -50%
poly_chknorm_c 8s 13s -38%
poly_uniform_eta_4x 8s 12s -33%
polyeta_unpack 8s 13s -38%
rej_uniform_native_x86_64 8s - new
sign_keypair 7s 4s +75%
mld_sign_resume 6s - new
poly_add 6s 8s -25%
poly_invntt_tomont_c 6s 9s -33%
polyvecl_ntt 6s 8s -25%
rej_uniform 6s 7s -14%
keccak_absorb_once_x4 5s 8s -38%
mld_check_pct 5s 15s -67%
mld_ct_cmask_nonzero_u8 5s 2s +150%
mld_keccakf1600_permute_c 5s 8s -38%
mld_sample_s1_s2_serial 5s 6s -17%
poly_caddq 5s 5s +0%
poly_caddq_native_aarch64 5s 3s +67%
polyvec_matrix_pointwise_montgomery_row 5s 13s -62%
polyveck_chknorm 5s 9s -44%
polyvecl_chknorm 5s 38s -87%
sign_signature_pre_hash_shake256 5s 3s +67%
caddq 4s 3s +33%
keccak_absorb 4s 4s +0%
keccak_f1600_x1_native_aarch64_v84a 4s 2s +100%
mld_polymat_expand_entry 4s 4s +0%
pointwise_acc_native_aarch64 4s 5s -20%
pointwise_acc_native_x86_64 4s 8s -50%
poly_decompose_c 4s 6s -33%
poly_power2round 4s 7s -43%
polyt1_unpack 4s 3s +33%
polyvec_matrix_expand 4s 3s +33%
polyveck_decompose_pack_w1 4s - new
polyveck_reduce 4s 6s -33%
polyw1_unpack_32 4s - new
polyz_unpack_c 4s 7s -43%
rej_eta_c 4s 4s +0%
sign_signature_pre_hash_internal 4s 2s +100%
unpack_sk 4s 4s +0%
decompose 3s 2s +50%
fqscale 3s 1s +200%
intt_native_x86_64 3s 4s -25%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 3s 2s +50%
keccakf1600x4_extract_bytes 3s 1s +200%
keccakf1600x4_xor_bytes 3s 1s +200%
make_hint 3s 2s +50%
mld_ct_get_optblocker_i64 3s 2s +50%
mld_ct_sel_int32 3s 2s +50%
mld_h 3s 2s +50%
mld_sample_s1_s2 3s 8s -62%
mld_sign_finish 3s - new
ntt_native_x86_64 3s 2s +50%
nttunpack_native_x86_64 3s 3s +0%
pack_sig_c 3s 2s +50%
pointwise_native_aarch64 3s 4s -25%
poly_challenge 3s 4s -25%
poly_chknorm_native 3s 1s +200%
poly_invntt_tomont_native 3s 2s +50%
poly_ntt 3s 2s +50%
poly_ntt_native 3s 4s -25%
poly_pointwise_montgomery_native 3s 3s +0%
poly_reduce 3s 4s -25%
poly_shiftl 3s 4s -25%
poly_sub 3s 5s -40%
poly_uniform_eta 3s 5s -40%
poly_use_hint_native_aarch64 3s 3s +0%
polyeta_pack 3s 3s +0%
polyveck_caddq 3s 8s -62%
polyvecl_pointwise_acc_montgomery_native 3s 2s +50%
polyvecl_uniform_gamma1 3s 3s +0%
polyw1_pack_32 3s 3s +0%
polyz_pack 3s 4s -25%
polyz_unpack_19_native_aarch64 3s 5s -40%
rej_uniform_eta_native_aarch64 3s 3s +0%
rej_uniform_native 3s 5s -40%
shake128_finalize 3s 2s +50%
shake256_finalize 3s 3s +0%
shake256_release 3s 5s -40%
sign_signature 3s 4s -25%
sign_signature_extmu 3s 4s -25%
sign_verify 3s 2s +50%
sign_verify_pre_hash_shake256 3s 7s -57%
sk_t0hat_get_poly 3s 3s +0%
use_hint 3s 3s +0%
keccak_init 2s 1s +100%
keccak_squeeze 2s 5s -60%
keccakf1600_permute 2s 2s +0%
keccakf1600_permute_native 2s 3s -33%
keccakf1600_xor_bytes 2s 1s +100%
keccakf1600_xor_bytes (big endian) 2s 2s +0%
keccakf1600x4_permute 2s 2s +0%
keccakf1600x4_xor_bytes_native 2s 2s +0%
mld_ct_get_optblocker_u32 2s 2s +0%
mld_ct_memcmp 2s 1s +100%
mld_keccakf1600_extract_bytes 2s 1s +100%
mld_keccakf1600x4_extract_bytes_c 2s 3s -33%
mld_keccakf1600x4_xor_bytes_c 2s 1s +100%
mld_prepare_domain_separation_prefix 2s 3s -33%
mld_value_barrier_u32 2s 3s -33%
mld_value_barrier_u8 2s 2s +0%
pack_sig_h 2s 3s -33%
pack_sk_rho_key_tr_s2 2s 2s +0%
pack_sk_s1 2s 2s +0%
poly_caddq_c 2s 3s -33%
poly_caddq_native 2s 3s -33%
poly_caddq_native_x86_64 2s 3s -33%
poly_chknorm 2s 4s -50%
poly_chknorm_native_aarch64 2s 5s -60%
poly_chknorm_native_x86_64 2s 2s +0%
poly_decompose 2s 2s +0%
poly_decompose_native 2s 2s +0%
poly_decompose_native_x86_64 2s 3s -33%
poly_uniform 2s 4s -50%
poly_uniform_4x 2s 3s -33%
poly_uniform_gamma1 2s 3s -33%
poly_uniform_gamma1_4x 2s 3s -33%
poly_use_hint 2s 4s -50%
poly_use_hint_native_x86_64 2s - new
polyt0_pack 2s 3s -33%
polyt1_pack 2s 5s -60%
polyveck_ntt 2s 3s -33%
polyveck_pack_eta 2s 5s -60%
polyvecl_pointwise_acc_montgomery_c 2s 2s +0%
polyvecl_uniform_gamma1_serial 2s 2s +0%
polyvecl_unpack_eta 2s 3s -33%
polyvecl_unpack_z 2s 3s -33%
polyw1_pack 2s 2s +0%
polyw1_pack_88 2s 2s +0%
polyw1_unpack 2s - new
polyz_unpack_17_native_aarch64 2s 4s -50%
reduce32 2s 3s -33%
rej_eta_native 2s 4s -50%
rej_uniform_eta_native_x86_64 2s - new
rej_uniform_native_aarch64 2s 3s -33%
shake128_init 2s 2s +0%
shake128_release 2s 3s -33%
shake128_squeeze 2s 1s +100%
shake128x4_squeezeblocks 2s 1s +100%
shake256 2s 3s -33%
shake256_squeeze 2s 2s +0%
shake256x4_absorb_once 2s 5s -60%
sign_verify_extmu 2s 4s -50%
sk_s1hat_get_poly 2s 3s -33%
sk_s2hat_get_poly 2s 2s +0%
sys_check_capability 2s 2s +0%
unpack_sk_s2hat 2s 3s -33%
yvec_get_poly 2s 2s +0%
yvec_init 2s 4s -50%
intt_native_aarch64 1s 4s -75%
keccak_f1600_x1_native_aarch64 1s 2s -50%
keccak_f1600_x4_native_aarch64_v84a 1s 1s +0%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 1s 2s -50%
keccak_f1600_x4_native_avx2 1s 3s -67%
keccak_finalize 1s 1s +0%
keccak_squeezeblocks_x4 1s 4s -75%
keccakf1600_extract_bytes (big endian) 1s 3s -67%
keccakf1600x4_extract_bytes_native 1s 4s -75%
mld_compute_pack_z 1s 6s -83%
mld_ct_abs_i32 1s 1s +0%
mld_ct_cmask_neg_i32 1s 2s -50%
mld_ct_cmask_nonzero_u32 1s 5s -80%
mld_ct_get_optblocker_u8 1s 3s -67%
mld_sign_attempt 1s - new
mld_value_barrier_i64 1s 2s -50%
montgomery_reduce 1s 2s -50%
ntt_native_aarch64 1s 4s -75%
pack_sig_z 1s 3s -67%
pointwise_native_x86_64 1s 5s -80%
poly_decompose_32_native_aarch64 1s 3s -67%
poly_decompose_88_native_aarch64 1s 2s -50%
poly_invntt_tomont 1s 4s -75%
poly_permute_bitrev_to_custom_optional 1s 2s -50%
poly_permute_bitrev_to_custom_optional_native 1s 3s -67%
poly_pointwise_montgomery 1s 3s -67%
poly_use_hint_c 1s 4s -75%
poly_use_hint_native 1s 1s +0%
polyvec_matrix_expand_serial 1s 3s -67%
polyveck_invntt_tomont 1s 4s -75%
polyveck_unpack_eta 1s 3s -67%
polyvecl_pack_eta 1s 2s -50%
polyvecl_pointwise_acc_montgomery 1s 2s -50%
polyw1_unpack_88 1s - new
polyz_unpack 1s 3s -67%
polyz_unpack_native 1s 1s +0%
polyz_unpack_native_x86_64 1s 3s -67%
power2round 1s 3s -67%
rej_eta 1s 2s -50%
shake128_absorb 1s 2s -50%
shake128x4_absorb_once 1s 4s -75%
shake256_absorb 1s 3s -67%
shake256_init 1s 3s -67%
shake256x4_squeezeblocks 1s 4s -75%
sign_verify_pre_hash_internal 1s 3s -67%
unpack_pk_t1 1s 2s -50%
unpack_sk_s1hat 1s 3s -67%
unpack_sk_t0hat 1s 4s -75%

@oqs-bot

oqs-bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-65, REDUCE-RAM)

⚠️ Attention Required

Proof Status Current Previous Change
compute_pack_t0_t1 ⚠️ 32s 7s +357%
mld_attempt_signature_generation ⚠️ 113s 31s +265%
sign_keypair_internal ⚠️ 29s 3s +867%
sign_pk_from_sk ⚠️ 49s 5s +880%
sign_verify_internal ⚠️ 145s 71s +104%
Full Results (212 proofs)
Proof Status Current Previous Change
**TOTAL** 1161s 1482s -21.7%
sign_verify_internal ⚠️ 145s 71s +104%
mld_attempt_signature_generation ⚠️ 113s 31s +265%
polyvec_matrix_pointwise_montgomery_yvec 73s 201s -64%
poly_pointwise_montgomery_c 66s 116s -43%
mld_invntt_layer 54s 108s -50%
sign_pk_from_sk ⚠️ 49s 5s +880%
compute_pack_t0_t1 ⚠️ 32s 7s +357%
sign_keypair_internal ⚠️ 29s 3s +867%
fqmul 23s 40s -43%
mld_ntt_layer 21s 43s -51%
keccakf1600x4_permute_native 11s 22s -50%
mld_ntt_butterfly_block 11s 25s -56%
sign_signature_internal 11s 6s +83%
sig_unpack_hints 10s 1s +900%
poly_ntt_c 8s 22s -64%
poly_uniform_eta_4x 8s 13s -38%
polyt0_unpack 8s 13s -38%
polyveck_chknorm 8s 36s -78%
poly_chknorm_c 7s 13s -46%
rej_uniform_native_x86_64 7s - new
poly_decompose_c 6s 8s -25%
polyveck_decompose_pack_w1 6s - new
rej_uniform_c 6s 16s -62%
sign_keypair 6s 4s +50%
sign_signature_pre_hash_shake256 6s 7s -14%
keccak_absorb_once_x4 5s 10s -50%
keccakf1600_permute_native 5s 2s +150%
pointwise_acc_native_aarch64 5s 6s -17%
poly_caddq_native_aarch64 5s 2s +150%
poly_invntt_tomont_c 5s 10s -50%
poly_power2round 5s 7s -29%
polyveck_caddq 5s 7s -29%
sign_verify 5s 6s -17%
keccakf1600_xor_bytes 4s 4s +0%
keccakf1600x4_permute 4s 1s +300%
mld_check_pct 4s 12s -67%
mld_compute_pack_z 4s 5s -20%
mld_keccakf1600_permute_c 4s 7s -43%
mld_sign_resume 4s - new
montgomery_reduce 4s 1s +300%
ntt_native_x86_64 4s 5s -20%
pointwise_native_aarch64 4s 3s +33%
poly_caddq_c 4s 5s -20%
poly_uniform_4x 4s 4s +0%
poly_uniform_eta 4s 5s -20%
polyeta_unpack 4s 4s +0%
polyvec_matrix_expand 4s 6s -33%
polyvec_matrix_pointwise_montgomery_row 4s 8s -50%
polyveck_reduce 4s 6s -33%
polyvecl_chknorm 4s 43s -91%
polyvecl_ntt 4s 7s -43%
polyw1_pack_88 4s 3s +33%
polyz_unpack_c 4s 9s -56%
rej_eta 4s 5s -20%
sign_signature_extmu 4s 4s +0%
unpack_pk_t1 4s 2s +100%
unpack_sk 4s 3s +33%
keccak_f1600_x1_native_aarch64_v84a 3s 2s +50%
keccak_squeezeblocks_x4 3s 4s -25%
mld_prepare_domain_separation_prefix 3s 4s -25%
pack_sig_h 3s 2s +50%
pack_sig_z 3s 1s +200%
pack_sk_s1 3s 3s +0%
pointwise_acc_native_x86_64 3s 5s -40%
poly_caddq 3s 3s +0%
poly_caddq_native 3s 4s -25%
poly_caddq_native_x86_64 3s 1s +200%
poly_chknorm 3s 3s +0%
poly_chknorm_native_x86_64 3s 4s -25%
poly_decompose_native_x86_64 3s 2s +50%
poly_invntt_tomont_native 3s 5s -40%
poly_pointwise_montgomery_native 3s 4s -25%
poly_shiftl 3s 3s +0%
polyeta_pack 3s 1s +200%
polyt0_pack 3s 5s -40%
polyveck_pack_eta 3s 3s +0%
polyveck_unpack_eta 3s 3s +0%
polyvecl_uniform_gamma1_serial 3s 1s +200%
polyw1_pack_32 3s 2s +50%
polyw1_unpack 3s - new
polyz_pack 3s 3s +0%
polyz_unpack_19_native_aarch64 3s 5s -40%
polyz_unpack_native_x86_64 3s 3s +0%
rej_uniform_native 3s 4s -25%
shake256_init 3s 4s -25%
sign_signature_pre_hash_internal 3s 3s +0%
sign_verify_pre_hash_internal 3s 5s -40%
sign_verify_pre_hash_shake256 3s 6s -50%
sk_s1hat_get_poly 3s 3s +0%
unpack_sk_s1hat 3s 1s +200%
yvec_get_poly 3s 3s +0%
yvec_init 3s 5s -40%
decompose 2s 4s -50%
fqscale 2s 3s -33%
intt_native_aarch64 2s 2s +0%
intt_native_x86_64 2s 4s -50%
keccak_absorb 2s 3s -33%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 2s 3s -33%
keccak_finalize 2s 2s +0%
keccakf1600x4_extract_bytes_native 2s 1s +100%
keccakf1600x4_xor_bytes 2s 2s +0%
keccakf1600x4_xor_bytes_native 2s 2s +0%
make_hint 2s 3s -33%
mld_ct_abs_i32 2s 6s -67%
mld_ct_cmask_neg_i32 2s 3s -33%
mld_ct_cmask_nonzero_u32 2s 3s -33%
mld_ct_get_optblocker_u32 2s 3s -33%
mld_ct_get_optblocker_u8 2s 4s -50%
mld_ct_memcmp 2s 4s -50%
mld_ct_sel_int32 2s 2s +0%
mld_h 2s 3s -33%
mld_keccakf1600_extract_bytes 2s 2s +0%
mld_keccakf1600x4_extract_bytes_c 2s 2s +0%
mld_sign_attempt 2s - new
mld_value_barrier_u32 2s 4s -50%
mld_value_barrier_u8 2s 1s +100%
ntt_native_aarch64 2s 3s -33%
nttunpack_native_x86_64 2s 3s -33%
pack_sig_c 2s 3s -33%
pointwise_native_x86_64 2s 2s +0%
poly_add 2s 7s -71%
poly_chknorm_native_aarch64 2s 2s +0%
poly_decompose 2s 2s +0%
poly_decompose_32_native_aarch64 2s 4s -50%
poly_decompose_88_native_aarch64 2s 2s +0%
poly_decompose_native 2s 5s -60%
poly_invntt_tomont 2s 5s -60%
poly_ntt 2s 2s +0%
poly_ntt_native 2s 2s +0%
poly_permute_bitrev_to_custom_optional_native 2s 3s -33%
poly_sub 2s 4s -50%
poly_uniform 2s 2s +0%
poly_uniform_gamma1 2s 3s -33%
poly_use_hint_native 2s 7s -71%
poly_use_hint_native_aarch64 2s 4s -50%
polyt1_pack 2s 3s -33%
polyt1_unpack 2s 3s -33%
polyvec_matrix_expand_serial 2s 3s -33%
polyveck_invntt_tomont 2s 7s -71%
polyveck_ntt 2s 4s -50%
polyvecl_pack_eta 2s 2s +0%
polyvecl_pointwise_acc_montgomery 2s 3s -33%
polyvecl_pointwise_acc_montgomery_c 2s 3s -33%
polyvecl_pointwise_acc_montgomery_native 2s 3s -33%
polyvecl_uniform_gamma1 2s 4s -50%
polyvecl_unpack_eta 2s 2s +0%
polyvecl_unpack_z 2s 1s +100%
polyw1_pack 2s 3s -33%
polyw1_unpack_32 2s - new
polyz_unpack 2s 3s -33%
polyz_unpack_native 2s 3s -33%
reduce32 2s 2s +0%
rej_eta_c 2s 4s -50%
rej_uniform 2s 7s -71%
rej_uniform_eta_native_x86_64 2s - new
rej_uniform_native_aarch64 2s 5s -60%
shake128_absorb 2s 2s +0%
shake128_finalize 2s 2s +0%
shake128_init 2s 6s -67%
shake128_release 2s 2s +0%
shake128_squeeze 2s 3s -33%
shake128x4_absorb_once 2s 2s +0%
shake256_finalize 2s 1s +100%
shake256_release 2s 3s -33%
shake256x4_squeezeblocks 2s 2s +0%
sign_signature 2s 5s -60%
sk_t0hat_get_poly 2s 1s +100%
unpack_sk_t0hat 2s 4s -50%
caddq 1s 2s -50%
keccak_f1600_x1_native_aarch64 1s 2s -50%
keccak_f1600_x4_native_aarch64_v84a 1s 2s -50%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 1s 3s -67%
keccak_f1600_x4_native_avx2 1s 2s -50%
keccak_init 1s 4s -75%
keccak_squeeze 1s 2s -50%
keccakf1600_extract_bytes (big endian) 1s 3s -67%
keccakf1600_permute 1s 4s -75%
keccakf1600_xor_bytes (big endian) 1s 2s -50%
keccakf1600x4_extract_bytes 1s 4s -75%
mld_ct_cmask_nonzero_u8 1s 2s -50%
mld_ct_get_optblocker_i64 1s 1s +0%
mld_keccakf1600x4_xor_bytes_c 1s 2s -50%
mld_polymat_expand_entry 1s 3s -67%
mld_sample_s1_s2 1s 6s -83%
mld_sample_s1_s2_serial 1s 3s -67%
mld_sign_finish 1s - new
mld_value_barrier_i64 1s 2s -50%
pack_sk_rho_key_tr_s2 1s 2s -50%
poly_challenge 1s 5s -80%
poly_chknorm_native 1s 4s -75%
poly_permute_bitrev_to_custom_optional 1s 3s -67%
poly_pointwise_montgomery 1s 2s -50%
poly_reduce 1s 4s -75%
poly_uniform_gamma1_4x 1s 3s -67%
poly_use_hint 1s 3s -67%
poly_use_hint_c 1s 2s -50%
poly_use_hint_native_x86_64 1s - new
polyw1_unpack_88 1s - new
polyz_unpack_17_native_aarch64 1s 2s -50%
power2round 1s 2s -50%
rej_eta_native 1s 3s -67%
rej_uniform_eta_native_aarch64 1s 5s -80%
shake128x4_squeezeblocks 1s 2s -50%
shake256 1s 2s -50%
shake256_absorb 1s 4s -75%
shake256_squeeze 1s 4s -75%
shake256x4_absorb_once 1s 2s -50%
sign_verify_extmu 1s 2s -50%
sk_s2hat_get_poly 1s 1s +0%
sys_check_capability 1s 1s +0%
unpack_sk_s2hat 1s 3s -67%
use_hint 1s 2s -50%

@oqs-bot

oqs-bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-87)

⚠️ Attention Required

Proof Status Current Previous Change
compute_pack_t0_t1 ⚠️ 190s 19s +900%
mld_attempt_signature_generation ⚠️ 243s 55s +342%
sig_unpack_hints ⚠️ 29s 3s +867%
sign_keypair_internal ⚠️ 44s 6s +633%
sign_pk_from_sk ⚠️ 61s 6s +917%
sign_signature_internal ⚠️ 310s 42s +638%
sign_verify_internal ⚠️ 509s 97s +425%
Full Results (212 proofs)
Proof Status Current Previous Change
**TOTAL** 2458s 2151s +14.3%
sign_verify_internal ⚠️ 509s 97s +425%
sign_signature_internal ⚠️ 310s 42s +638%
mld_attempt_signature_generation ⚠️ 243s 55s +342%
polyvecl_pointwise_acc_montgomery_c 218s 353s -38%
compute_pack_t0_t1 ⚠️ 190s 19s +900%
polyvec_matrix_expand 83s 324s -74%
sign_pk_from_sk ⚠️ 61s 6s +917%
mld_invntt_layer 58s 115s -50%
poly_pointwise_montgomery_c 45s 144s -69%
sign_keypair_internal ⚠️ 44s 6s +633%
sig_unpack_hints ⚠️ 29s 3s +867%
fqmul 24s 43s -44%
polyvec_matrix_expand_serial 24s 38s -37%
mld_ntt_layer 22s 45s -51%
mld_ntt_butterfly_block 13s 24s -46%
polyvec_matrix_pointwise_montgomery_yvec 13s 18s -28%
keccakf1600x4_permute_native 12s 23s -48%
rej_uniform 11s 15s -27%
poly_ntt_c 10s 21s -52%
rej_uniform_c 10s 19s -47%
poly_chknorm_c 8s 15s -47%
poly_uniform_eta_4x 8s 11s -27%
polyt0_unpack 8s 14s -43%
poly_invntt_tomont_c 7s 12s -42%
poly_uniform_4x 7s 13s -46%
polyveck_decompose_pack_w1 7s - new
rej_uniform_native_x86_64 7s - new
polyt1_unpack 6s 4s +50%
sign_keypair 6s 4s +50%
poly_chknorm_native 5s 4s +25%
polyeta_unpack 5s 17s -71%
polyveck_caddq 5s 8s -38%
polyvecl_ntt 5s 7s -29%
polyvecl_unpack_eta 5s 5s +0%
shake256_absorb 5s 2s +150%
sign_signature_pre_hash_shake256 5s 6s -17%
keccak_absorb_once_x4 4s 9s -56%
mld_check_pct 4s 16s -75%
mld_compute_pack_z 4s 8s -50%
mld_sample_s1_s2 4s 6s -33%
mld_sample_s1_s2_serial 4s 8s -50%
mld_sign_finish 4s - new
mld_value_barrier_i64 4s 1s +300%
ntt_native_x86_64 4s 2s +100%
pointwise_acc_native_aarch64 4s 6s -33%
poly_decompose_c 4s 5s -20%
poly_power2round 4s 3s +33%
poly_use_hint_native_x86_64 4s - new
polyt1_pack 4s 3s +33%
polyveck_invntt_tomont 4s 9s -56%
polyveck_ntt 4s 11s -64%
polyw1_pack_32 4s 4s +0%
polyw1_unpack_88 4s - new
shake128x4_absorb_once 4s 5s -20%
shake256x4_absorb_once 4s 2s +100%
sign_signature 4s 5s -20%
sign_signature_pre_hash_internal 4s 5s -20%
sys_check_capability 4s 3s +33%
unpack_pk_t1 4s 1s +300%
unpack_sk_s1hat 4s 2s +100%
caddq 3s 4s -25%
fqscale 3s 3s +0%
keccak_f1600_x1_native_aarch64_v84a 3s 1s +200%
keccak_f1600_x4_native_aarch64_v84a 3s 4s -25%
keccak_squeezeblocks_x4 3s 4s -25%
keccakf1600_permute_native 3s 4s -25%
keccakf1600x4_extract_bytes_native 3s 2s +50%
keccakf1600x4_xor_bytes_native 3s 3s +0%
mld_ct_cmask_neg_i32 3s 4s -25%
mld_ct_cmask_nonzero_u8 3s 3s +0%
mld_keccakf1600_permute_c 3s 7s -57%
mld_keccakf1600x4_xor_bytes_c 3s 4s -25%
mld_polymat_expand_entry 3s 3s +0%
pack_sig_c 3s 4s -25%
pack_sig_h 3s 4s -25%
pack_sk_rho_key_tr_s2 3s 3s +0%
pointwise_acc_native_x86_64 3s 6s -50%
pointwise_native_aarch64 3s 5s -40%
poly_add 3s 6s -50%
poly_caddq 3s 2s +50%
poly_caddq_c 3s 3s +0%
poly_caddq_native_aarch64 3s 3s +0%
poly_decompose_32_native_aarch64 3s 5s -40%
poly_decompose_native_x86_64 3s 2s +50%
poly_invntt_tomont_native 3s 5s -40%
poly_uniform 3s 4s -25%
poly_uniform_gamma1_4x 3s 3s +0%
poly_use_hint 3s 1s +200%
poly_use_hint_native 3s 3s +0%
polyeta_pack 3s 4s -25%
polyvec_matrix_pointwise_montgomery_row 3s 3s +0%
polyveck_pack_eta 3s 6s -50%
polyvecl_chknorm 3s 7s -57%
polyvecl_unpack_z 3s 4s -25%
polyw1_unpack 3s - new
polyz_pack 3s 5s -40%
polyz_unpack 3s 2s +50%
polyz_unpack_c 3s 3s +0%
power2round 3s 4s -25%
rej_eta_c 3s 5s -40%
rej_eta_native 3s 6s -50%
rej_uniform_native 3s 8s -62%
shake128_absorb 3s 4s -25%
shake128_finalize 3s 2s +50%
shake128_init 3s 1s +200%
shake256_release 3s 3s +0%
sign_signature_extmu 3s 4s -25%
sign_verify 3s 5s -40%
sign_verify_pre_hash_internal 3s 3s +0%
sign_verify_pre_hash_shake256 3s 3s +0%
sk_s2hat_get_poly 3s 3s +0%
unpack_sk_t0hat 3s 7s -57%
use_hint 3s 3s +0%
decompose 2s 3s -33%
intt_native_x86_64 2s 3s -33%
keccak_absorb 2s 3s -33%
keccak_f1600_x1_native_aarch64 2s 1s +100%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 2s 4s -50%
keccak_f1600_x4_native_avx2 2s 2s +0%
keccak_init 2s 3s -33%
keccakf1600_extract_bytes (big endian) 2s 3s -33%
keccakf1600_permute 2s 3s -33%
keccakf1600_xor_bytes 2s 3s -33%
keccakf1600_xor_bytes (big endian) 2s 2s +0%
keccakf1600x4_permute 2s 3s -33%
keccakf1600x4_xor_bytes 2s 3s -33%
mld_ct_abs_i32 2s 3s -33%
mld_ct_cmask_nonzero_u32 2s 2s +0%
mld_ct_get_optblocker_i64 2s 3s -33%
mld_ct_get_optblocker_u8 2s 3s -33%
mld_ct_sel_int32 2s 3s -33%
mld_h 2s 3s -33%
mld_keccakf1600_extract_bytes 2s 4s -50%
mld_prepare_domain_separation_prefix 2s 4s -50%
mld_sign_attempt 2s - new
mld_sign_resume 2s - new
mld_value_barrier_u32 2s 2s +0%
mld_value_barrier_u8 2s 3s -33%
ntt_native_aarch64 2s 6s -67%
pack_sig_z 2s 5s -60%
pack_sk_s1 2s 1s +100%
pointwise_native_x86_64 2s 2s +0%
poly_caddq_native 2s 4s -50%
poly_caddq_native_x86_64 2s 2s +0%
poly_chknorm_native_aarch64 2s 4s -50%
poly_chknorm_native_x86_64 2s 2s +0%
poly_decompose 2s 3s -33%
poly_decompose_88_native_aarch64 2s 2s +0%
poly_decompose_native 2s 2s +0%
poly_invntt_tomont 2s 1s +100%
poly_ntt 2s 3s -33%
poly_permute_bitrev_to_custom_optional 2s 3s -33%
poly_permute_bitrev_to_custom_optional_native 2s 5s -60%
poly_pointwise_montgomery_native 2s 5s -60%
poly_reduce 2s 2s +0%
poly_shiftl 2s 4s -50%
poly_sub 2s 3s -33%
poly_uniform_eta 2s 4s -50%
poly_uniform_gamma1 2s 4s -50%
poly_use_hint_c 2s 2s +0%
polyt0_pack 2s 4s -50%
polyveck_chknorm 2s 3s -33%
polyveck_reduce 2s 3s -33%
polyveck_unpack_eta 2s 4s -50%
polyvecl_pointwise_acc_montgomery 2s 5s -60%
polyvecl_pointwise_acc_montgomery_native 2s 3s -33%
polyvecl_uniform_gamma1 2s 2s +0%
polyvecl_uniform_gamma1_serial 2s 4s -50%
polyw1_pack_88 2s 3s -33%
polyw1_unpack_32 2s - new
polyz_unpack_17_native_aarch64 2s 3s -33%
polyz_unpack_19_native_aarch64 2s 4s -50%
reduce32 2s 3s -33%
rej_eta 2s 4s -50%
rej_uniform_eta_native_aarch64 2s 5s -60%
rej_uniform_eta_native_x86_64 2s - new
rej_uniform_native_aarch64 2s 2s +0%
shake128_squeeze 2s 2s +0%
shake128x4_squeezeblocks 2s 3s -33%
shake256 2s 2s +0%
shake256_finalize 2s 2s +0%
shake256_init 2s 2s +0%
shake256x4_squeezeblocks 2s 1s +100%
sign_verify_extmu 2s 3s -33%
sk_s1hat_get_poly 2s 3s -33%
unpack_sk 2s 5s -60%
unpack_sk_s2hat 2s 3s -33%
yvec_get_poly 2s 4s -50%
yvec_init 2s 4s -50%
intt_native_aarch64 1s 3s -67%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 1s 3s -67%
keccak_finalize 1s 1s +0%
keccak_squeeze 1s 2s -50%
keccakf1600x4_extract_bytes 1s 2s -50%
make_hint 1s 3s -67%
mld_ct_get_optblocker_u32 1s 2s -50%
mld_ct_memcmp 1s 3s -67%
mld_keccakf1600x4_extract_bytes_c 1s 2s -50%
montgomery_reduce 1s 3s -67%
nttunpack_native_x86_64 1s 3s -67%
poly_challenge 1s 5s -80%
poly_chknorm 1s 3s -67%
poly_ntt_native 1s 3s -67%
poly_pointwise_montgomery 1s 4s -75%
poly_use_hint_native_aarch64 1s 2s -50%
polyvecl_pack_eta 1s 3s -67%
polyw1_pack 1s 4s -75%
polyz_unpack_native 1s 4s -75%
polyz_unpack_native_x86_64 1s 3s -67%
shake128_release 1s 4s -75%
shake256_squeeze 1s 2s -50%
sk_t0hat_get_poly 1s 4s -75%

@oqs-bot

oqs-bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-DSA-44)

⚠️ Attention Required

Proof Status Current Previous Change
compute_pack_t0_t1 ⚠️ 94s 50s +88%
mld_attempt_signature_generation ⚠️ 155s 58s +167%
polyveck_chknorm ⚠️ 51s 5s +920%
sign_pk_from_sk ⚠️ 29s 5s +480%
sign_signature_internal ⚠️ 90s 26s +246%
sign_verify_internal ⚠️ 251s 114s +120%
Full Results (212 proofs)
Proof Status Current Previous Change
**TOTAL** 1561s 1514s +3.1%
sign_verify_internal ⚠️ 251s 114s +120%
mld_attempt_signature_generation ⚠️ 155s 58s +167%
polyvecl_pointwise_acc_montgomery_c 137s 131s +5%
compute_pack_t0_t1 ⚠️ 94s 50s +88%
sign_signature_internal ⚠️ 90s 26s +246%
mld_invntt_layer 53s 107s -50%
polyveck_chknorm ⚠️ 51s 5s +920%
poly_pointwise_montgomery_c 43s 114s -62%
sign_pk_from_sk ⚠️ 29s 5s +480%
fqmul 26s 39s -33%
mld_ntt_layer 21s 42s -50%
sig_unpack_hints 18s 2s +800%
polyvec_matrix_expand 17s 28s -39%
sign_keypair_internal 17s 4s +325%
keccakf1600x4_permute_native 13s 22s -41%
mld_ntt_butterfly_block 13s 23s -43%
polyvec_matrix_pointwise_montgomery_yvec 11s 15s -27%
poly_invntt_tomont_c 9s 11s -18%
poly_ntt_c 9s 19s -53%
rej_uniform 9s 18s -50%
rej_uniform_c 8s 15s -47%
poly_uniform_eta_4x 7s 11s -36%
polyeta_unpack 7s 14s -50%
polyt0_unpack 7s 15s -53%
rej_uniform_native_x86_64 7s - new
poly_chknorm_c 6s 17s -65%
poly_uniform_4x 6s 14s -57%
sign_signature_pre_hash_shake256 6s 3s +100%
keccak_absorb_once_x4 5s 9s -44%
mld_check_pct 5s 14s -64%
mld_sample_s1_s2_serial 5s 3s +67%
poly_pointwise_montgomery_native 5s 2s +150%
polyvec_matrix_expand_serial 5s 8s -38%
polyveck_pack_eta 5s 2s +150%
sign_keypair 5s 4s +25%
unpack_sk_s1hat 5s 1s +400%
unpack_sk_t0hat 5s 3s +67%
mld_compute_pack_z 4s 7s -43%
mld_ct_cmask_nonzero_u32 4s 2s +100%
mld_ct_get_optblocker_u8 4s 1s +300%
mld_keccakf1600_permute_c 4s 8s -50%
mld_sample_s1_s2 4s 4s +0%
ntt_native_aarch64 4s 3s +33%
poly_add 4s 7s -43%
poly_chknorm 4s 4s +0%
poly_invntt_tomont 4s 4s +0%
poly_use_hint_native_x86_64 4s - new
polyvecl_unpack_z 4s 2s +100%
polyz_pack 4s 2s +100%
polyz_unpack_c 4s 11s -64%
power2round 4s 3s +33%
rej_eta_native 4s 6s -33%
rej_uniform_native 4s 4s +0%
shake128_absorb 4s 2s +100%
shake128_squeeze 4s 1s +300%
shake256_init 4s 3s +33%
sign_signature_extmu 4s 4s +0%
fqscale 3s 3s +0%
keccak_absorb 3s 5s -40%
keccak_f1600_x1_native_aarch64_v84a 3s 3s +0%
keccak_f1600_x4_native_aarch64_v84a 3s 2s +50%
keccak_f1600_x4_native_avx2 3s 3s +0%
keccakf1600x4_permute 3s 4s -25%
mld_ct_cmask_nonzero_u8 3s 2s +50%
mld_h 3s 4s -25%
mld_keccakf1600_extract_bytes 3s 2s +50%
mld_sign_resume 3s - new
pack_sig_z 3s 4s -25%
pointwise_acc_native_aarch64 3s 4s -25%
pointwise_acc_native_x86_64 3s 7s -57%
pointwise_native_aarch64 3s 5s -40%
poly_caddq 3s 2s +50%
poly_decompose_32_native_aarch64 3s 2s +50%
poly_decompose_native 3s 3s +0%
poly_reduce 3s 3s +0%
poly_shiftl 3s 3s +0%
poly_sub 3s 3s +0%
poly_uniform_eta 3s 5s -40%
poly_uniform_gamma1 3s 4s -25%
poly_use_hint_c 3s 3s +0%
poly_use_hint_native 3s 4s -25%
polyt1_pack 3s 4s -25%
polyveck_reduce 3s 5s -40%
polyveck_unpack_eta 3s 2s +50%
polyvecl_chknorm 3s 10s -70%
polyvecl_pack_eta 3s 3s +0%
polyvecl_uniform_gamma1 3s 3s +0%
polyw1_pack 3s 3s +0%
polyw1_pack_32 3s 2s +50%
polyw1_unpack_88 3s - new
polyz_unpack_native 3s 2s +50%
rej_eta 3s 3s +0%
rej_uniform_eta_native_x86_64 3s - new
shake128x4_squeezeblocks 3s 1s +200%
shake256_finalize 3s 2s +50%
shake256x4_absorb_once 3s 3s +0%
sign_signature 3s 3s +0%
sign_verify 3s 4s -25%
sign_verify_extmu 3s 3s +0%
sign_verify_pre_hash_internal 3s 3s +0%
sign_verify_pre_hash_shake256 3s 5s -40%
sk_t0hat_get_poly 3s 2s +50%
unpack_sk 3s 3s +0%
use_hint 3s 2s +50%
caddq 2s 2s +0%
decompose 2s 2s +0%
intt_native_aarch64 2s 9s -78%
intt_native_x86_64 2s 4s -50%
keccak_finalize 2s 3s -33%
keccak_init 2s 3s -33%
keccak_squeeze 2s 1s +100%
keccak_squeezeblocks_x4 2s 4s -50%
keccakf1600_extract_bytes (big endian) 2s 2s +0%
keccakf1600_permute_native 2s 3s -33%
keccakf1600x4_extract_bytes_native 2s 3s -33%
keccakf1600x4_xor_bytes_native 2s 3s -33%
make_hint 2s 3s -33%
mld_ct_abs_i32 2s 2s +0%
mld_ct_cmask_neg_i32 2s 3s -33%
mld_ct_get_optblocker_i64 2s 4s -50%
mld_ct_get_optblocker_u32 2s 1s +100%
mld_ct_memcmp 2s 3s -33%
mld_ct_sel_int32 2s 2s +0%
mld_keccakf1600x4_extract_bytes_c 2s 2s +0%
mld_sign_attempt 2s - new
mld_sign_finish 2s - new
mld_value_barrier_u32 2s 2s +0%
mld_value_barrier_u8 2s 3s -33%
ntt_native_x86_64 2s 3s -33%
pack_sig_c 2s 2s +0%
pack_sig_h 2s 3s -33%
pack_sk_rho_key_tr_s2 2s 3s -33%
pack_sk_s1 2s 2s +0%
pointwise_native_x86_64 2s 4s -50%
poly_caddq_c 2s 2s +0%
poly_caddq_native 2s 5s -60%
poly_caddq_native_aarch64 2s 4s -50%
poly_caddq_native_x86_64 2s 4s -50%
poly_chknorm_native_x86_64 2s 2s +0%
poly_decompose 2s 3s -33%
poly_decompose_c 2s 4s -50%
poly_decompose_native_x86_64 2s 3s -33%
poly_ntt 2s 3s -33%
poly_ntt_native 2s 5s -60%
poly_permute_bitrev_to_custom_optional 2s 2s +0%
poly_power2round 2s 4s -50%
poly_uniform 2s 6s -67%
poly_uniform_gamma1_4x 2s 5s -60%
poly_use_hint 2s 4s -50%
polyeta_pack 2s 3s -33%
polyvec_matrix_pointwise_montgomery_row 2s 2s +0%
polyveck_decompose_pack_w1 2s - new
polyveck_invntt_tomont 2s 5s -60%
polyvecl_pointwise_acc_montgomery 2s 4s -50%
polyvecl_unpack_eta 2s 2s +0%
polyw1_pack_88 2s 1s +100%
polyw1_unpack_32 2s - new
reduce32 2s 2s +0%
rej_eta_c 2s 4s -50%
rej_uniform_eta_native_aarch64 2s 2s +0%
shake128_finalize 2s 2s +0%
shake128_init 2s 3s -33%
shake256_release 2s 2s +0%
shake256_squeeze 2s 2s +0%
shake256x4_squeezeblocks 2s 4s -50%
sign_signature_pre_hash_internal 2s 6s -67%
sk_s1hat_get_poly 2s 3s -33%
sk_s2hat_get_poly 2s 2s +0%
sys_check_capability 2s 4s -50%
unpack_pk_t1 2s 5s -60%
unpack_sk_s2hat 2s 4s -50%
keccak_f1600_x1_native_aarch64 1s 2s -50%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 1s 3s -67%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 1s 1s +0%
keccakf1600_permute 1s 2s -50%
keccakf1600_xor_bytes 1s 1s +0%
keccakf1600_xor_bytes (big endian) 1s 4s -75%
keccakf1600x4_extract_bytes 1s 2s -50%
keccakf1600x4_xor_bytes 1s 2s -50%
mld_keccakf1600x4_xor_bytes_c 1s 2s -50%
mld_polymat_expand_entry 1s 3s -67%
mld_prepare_domain_separation_prefix 1s 4s -75%
mld_value_barrier_i64 1s 2s -50%
montgomery_reduce 1s 3s -67%
nttunpack_native_x86_64 1s 1s +0%
poly_challenge 1s 3s -67%
poly_chknorm_native 1s 4s -75%
poly_chknorm_native_aarch64 1s 4s -75%
poly_decompose_88_native_aarch64 1s 4s -75%
poly_invntt_tomont_native 1s 4s -75%
poly_permute_bitrev_to_custom_optional_native 1s 4s -75%
poly_pointwise_montgomery 1s 3s -67%
poly_use_hint_native_aarch64 1s 2s -50%
polyt0_pack 1s 2s -50%
polyt1_unpack 1s 5s -80%
polyveck_caddq 1s 3s -67%
polyveck_ntt 1s 5s -80%
polyvecl_ntt 1s 2s -50%
polyvecl_pointwise_acc_montgomery_native 1s 2s -50%
polyvecl_uniform_gamma1_serial 1s 2s -50%
polyw1_unpack 1s - new
polyz_unpack 1s 3s -67%
polyz_unpack_17_native_aarch64 1s 5s -80%
polyz_unpack_19_native_aarch64 1s 5s -80%
polyz_unpack_native_x86_64 1s 3s -67%
rej_uniform_native_aarch64 1s 5s -80%
shake128_release 1s 1s +0%
shake128x4_absorb_once 1s 3s -67%
shake256 1s 1s +0%
shake256_absorb 1s 3s -67%
yvec_get_poly 1s 3s -67%
yvec_init 1s 3s -67%

@mkannwischer

Copy link
Copy Markdown
Contributor Author

Interesting! What motivated this idea? Was it just from poking around in the area?

The motivation is to reduce memory consumption during signing. For many consumers that's the biggest issue for ML-DSA so every KiB we can save is going to make someone more happy.

The idea isn't new. We've done that on embedded platforms for a while, see https://eprint.iacr.org/2022/323.
I was curious what the performance impact would be if we implemented it unconditionally and I'm happy to see that it seems to be negligible.

I still have to go through the implementation to see if something can be simplified. I'd be very happy to see this merged for v2.1 though. If this makes it into the release, we should also revisit keygen as that then needs more memory than signing which is only because I stopped optimizing after it was below signing memory.

@mkannwischer
mkannwischer force-pushed the lowram-sign-pack-w1 branch 2 times, most recently from 578d61f to 7c8238d Compare August 11, 2026 06:03
@mkannwischer mkannwischer changed the title [TEST] sign: Keep w1 in packed form during signature generation lowram: Keep w1 in packed form during signature generation Aug 11, 2026
@mkannwischer
mkannwischer marked this pull request as ready for review August 11, 2026 06:39
@mkannwischer
mkannwischer requested a review from a team as a code owner August 11, 2026 06:39
Comment thread mldsa/src/packing.h Outdated
Comment thread mldsa/src/poly_kl.h Outdated
Comment thread mldsa/src/poly.h
Comment thread mldsa/src/poly.h
Comment thread mldsa/src/poly_kl.h Outdated
Comment thread mldsa/src/polyvec.h Outdated
Comment thread mldsa/src/polyvec.h Outdated
Comment thread mldsa/src/polyvec.h Outdated
Comment thread mldsa/src/polyvec_lazy.h Outdated
Comment thread mldsa/src/polyvec_lazy.h Outdated
Comment thread mldsa/src/sign.c
Comment thread mldsa/src/sign.c Outdated
Comment thread mldsa/src/sign.c Outdated

@hanno-becker hanno-becker left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is an interesting and impactful optimization, thank you for proposing it @mkannwischer.

Overall, I think this is worth implementing.

However, we are diverging more and more from the spec and reference implementation here. Especially with the lazy/eager split for yvec, we have to take utmost care to keep the documentation of both source and its relation to FIPS-204 accurate and crisp in terms of providing exactly the information that is needed to reason about the code at any point -- and nothing more to, to avoid unnecessary cognitive/reasoning load. I left a few comments in this direction.

@hanno-becker

Copy link
Copy Markdown
Contributor

We've done that on embedded platforms for a while, see https://eprint.iacr.org/2022/323.

We should cite the relevant papers.

The signing attempt held w1 as a full mld_polyveck, but w1 is only
needed as the w1Encode input to H and, coefficient-wise, as the a1
argument of MakeHint. Both work from the packed encoding, so Decompose
and w1Encode now run one polynomial at a time into a
MLDSA_K * MLDSA_POLYW1_PACKEDBYTES buffer, and mld_pack_sig_h recovers
each row with the new mld_polyw1_unpack.

With w1 gone, the matrix-vector scratch is the sole remaining user of
the buffer the two shared. In REDUCE_RAM mode that scratch is a single
polynomial rather than a polyvecl, and it now shares storage with z, so
the buffer disappears.

Signing allocation in bytes:

              default           REDUCE_RAM
  ML-DSA-44   44704 -> 44448    13120 ->  9792
  ML-DSA-65   69312 -> 68032    17248 -> 11872
  ML-DSA-87  108224 -> 107200   21344 -> 14176

mld_polyveck_decompose and mld_polyveck_pack_w1 have no other callers
and are replaced by the fused mld_polyveck_decompose_pack_w1.

Signed-off-by: Matthias J. Kannwischer <matthias@zerorisc.com>
Comment thread mldsa/src/polyvec_lazy.h
Comment on lines +514 to +516
void mld_polyvec_matrix_pointwise_montgomery_yvec_eager(
mld_polyveck *w, mld_polymat_eager *mat, const mld_yvec_eager *y,
mld_yvec_scratch_eager *scratch)

@hanno-becker hanno-becker Aug 12, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Have you considered whether the scratch can be embedded into mld_yvec, which anyway is an abstract type? Needs some checking of lifetimes in sign.c.

This would simplify the interface.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants