[NFC] Simplify lexer and move to header by tlively · Pull Request #8597 · WebAssembly/binaryen

tlively · 2026-04-13T23:58:00Z

The lexer previously used its own internal LexerCtx abstraction that allowed it to consume the characters that made up a token without changing the lexer state, then update the state at once when committing to consuming the characters. However, manually resetting the lexer to the original position when giving up on parsing a token is simple enough that this abstraction was not holding its weight. Simplify the lexer by removing internal contexts, and move the simplified method bodies to lexer.h. Generally we try to avoid putting lots of code in headers, but in this case making the code available to the inliner, along with removing the extra layer of abstraction, makes the parser about 20% faster.

The lexer previously used its own internal `LexerCtx` abstraction that allowed it to consume the characters that made up a token without changing the lexer state, then update the state at once when committing to consuming the characters. However, manually resetting the lexer to the original position when giving up on parsing a token is simple enough that this abstraction was not holding its weight. Simplify the lexer by removing internal contexts, and move the simplified method bodies to lexer.h. Generally we try to avoid putting lots of code in headers, but in this case making the code available to the inliner, along with removing the extra layer of abstraction, makes the parser about 20% faster.

The first parser pass is responsible for two things: finding the locations of definitions of top-level module items like globals and functions and finding the locations of implicit function type definitions. It previously accomplished the latter by fully parsing every instruction in each function. But the IR is not constructed in this phase of parsing, so fully parsing every instruction was largely wasted work. Optimize the parser by parsing only the instructions that might have implicit type definitions and otherwise just blindly match parentheses to skip the function body. Combined with #8597, this speeds up parsing by 30-40%.

kripken

Do I want to know how this affects our compile times? 😄

tlively · 2026-04-14T15:29:06Z

No difference! (at least on a very unscientific experiment with N=1)

tlively requested a review from kripken April 13, 2026 23:58

tlively requested a review from a team as a code owner April 13, 2026 23:58

fixes

b2528ad

tlively mentioned this pull request Apr 14, 2026

[NFC] Skip parsing instructions in first parser pass #8601

Open

kripken approved these changes Apr 14, 2026

View reviewed changes

tlively merged commit 54f9f7a into main Apr 14, 2026
16 checks passed

tlively deleted the parser-slowdown branch April 14, 2026 15:29

tlively mentioned this pull request Apr 14, 2026

~+700% regression in wasm-opt --nm performance from Emscripten 3.1.38 -> 4.0.19 #8406

Open

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[NFC] Simplify lexer and move to header#8597

[NFC] Simplify lexer and move to header#8597
tlively merged 2 commits intomainfrom
parser-slowdown

tlively commented Apr 13, 2026

Uh oh!

kripken left a comment

Uh oh!

tlively commented Apr 14, 2026

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

Conversation

tlively commented Apr 13, 2026

Uh oh!

kripken left a comment

Choose a reason for hiding this comment

Uh oh!

tlively commented Apr 14, 2026

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants