Skip to content

feat(engine): single-source logic engine generating TypeScript, Python, Go and Rust - #587

Closed
hyanmandian wants to merge 42 commits into
mainfrom
claude/generic-utilities-engine-h2qm1o
Closed

hyanmandian wants to merge 42 commits into
mainfrom
claude/generic-utilities-engine-h2qm1o

Conversation

@hyanmandian

@hyanmandian hyanmandian commented Sep 21, 2026 •

Copy link
Copy Markdown
Member

What does this PR do?

Adds a single-source logic engine: a utility's logic is written once, in a restricted and semantically typed subset of TypeScript, and a native, idiomatic implementation is generated for TypeScript, Python, Go and Rust.

No shared runtime package, no WASM, no FFI, no bridge, no interpreter at run time, and no external dependencies in the generated code — the generated Rust crate's Cargo.lock lists exactly one package, itself.

Two new directories, and nothing in src/ changes:

  • engine/ — the compiler. Domain neutral and self-contained, so it can be extracted into its own repository later: one runtime dependency (oxc-parser), no build step (Node runs the TypeScript directly), its own test suite, and a generic example project (examples/generic, Luhn and a slugifier) that a test uses to keep any Brazilian Utils assumption from leaking into the compiler.
  • core/ — the first project that uses it: the pilot utilities, their conformance harness, the measured contracts and the cross-language benchmark.

How it works

source (restricted TypeScript)
   │  frontend: parse, resolve, reject what is outside the subset
   ▼
Semantic HIR ── the frontend contract
   │  checker: types, refinements, ranges, effects → Core
   ▼
Core IR ── what the program means, in no particular language
   │  link, comptime, optimize, capability threading
   ▼
per target: lowering selection → Target AST → printer → that language's formatter
   ▼
TypeScript · Python · Go · Rust

Types carry proofs — an integer's range, a string's character class and length — and a native lowering is selected only when its preconditions hold. The clearest example is one line of source:

str.compare(left, right)
  • Python: < compares code points, which is the Core's order → native, no precondition.
  • Go and Rust: strings.Compare and Ord for str compare UTF-8 bytes, same order → native.
  • TypeScript: < compares UTF-16 code units, so an astral scalar sorts below U+E000 → native only when both sides are proven ASCII, otherwise the portable implementation from the engine's own source-language standard library.

out/<target>/LOWERING.md records every selection and the rule that decided it.

The engine generates the core, never the public API: the handwritten DX in each language keeps its coercion, defaults and naming. core/docs/contracts.md records that split for every pilot, measured from src/ rather than assumed (including that formatCurrency emits an ordinary space after R$ because the package replaces Intl's NBSP, and that Intl rounds the shortest decimal representation, so 1.005 is "1,01" while toFixed(2) is "1.00").

The generated output is committed, under core/out/, so it can be read in review without running anything. CI fails if it drifts from what the engine emits for the committed sources.

Pilots

utility what it proves
isValidCpf strings, refinements, ranges, check-digit arithmetic
isValidCnpj numeric and alphanumeric, regex, ASCII, character values
formatCnpj shared helper, options record, enum option, pattern formatting
getHolidays stable sort, records, civil dates, Easter as source library code
isBusinessDay CivilDate, holiday lookup
getAddressInfoByCep Http capability, race, retry written in source, async colouring, error hierarchy, a JSON field reader written in the subset
formatCurrency exact Decimal, pt-BR formatting as source library code
generateCpf Random capability, rejection sampling written in source, bounded loops
generateCnpj the same draw reused, with the check-digit rule it shares with isValidCnpj

The two generators are the case where "written once" has to mean bit for bit: everything past the single random.nextU32 intrinsic is written in the subset, because a % bound bias would otherwise have to match digit for digit across four standard libraries to stay invisible. The conformance run also feeds whatever the generators drew back through isValidCpf / isValidCnpj, so four targets agreeing on an invalid document would still fail.

Rust was the falsification, and it held

engine/docs/targets/rust-sketch.md walked every Core construct through Rust on paper and predicted two frictions, both from std being smaller than the other three standard libraries. Both held, and neither needed a crate: no HTTP becomes a generated Capabilities trait the caller implements (Sync-bounded, because task.race is std::thread::scope), and no JSON becomes a hand-written codec used only by the differential driver.

The sketch's central plan did not survive, which is the more interesting result: it assumed &str and &[T] parameters, and borrowing cannot be decided where this pipeline decides things — a lowering runs before any function scope exists, and a printer sees one module without its callees' signatures. Every heap value is owned instead. ADR 0009 records that, and the sketch now says where reality diverged from it rather than being quietly corrected.

The lowering mix is the number worth reading: 250 native / 54 library / 2 portable against Go's 277/27/2 and TypeScript's 304/0/2. That column measures each standard library, not the engine.

Verification — cd core && npm run verify

Also runs on this PR, in the new Engine workflow.

check result
differential conformance 4256/4256 through the reference interpreter and all four targets, in both idiom modes (--no-idioms too)
against the published package 4247/4247 on every case reproducible offline
determinism byte-identical regeneration
generated code tsc --strict --noEmit, python3 -m compileall, go vet, cargo build + cargo clippy -D warnings clean; formatted by Prettier, ruff format, gofmt and rustfmt, each pinned
engine suite 64 tests: subset rejections, refinement acceptance, loop soundness, intrinsic vectors, regex vs RegExp, translation validation, import boundaries, stage snapshots

Checklist

  • My commit/PR title follows Conventional Commits.
  • I added or updated tests covering this change — the engine's own suite plus the differential harness (cd core && npm run verify, wired into CI). The package's npm test is unaffected: vite.config.ts excludes the new directories from the package's own format, lint and test passes, since they carry their own toolchain.
  • I updated the documentation — engine/docs/ (specification, generated intrinsic reference, per-target notes, 9 ADRs, how to add a utility or a target, measured metrics) and core/docs/ (contracts, survey).
  • docs/utilities.md / docs/pt-br/utilities.md — not applicable: no published utility changes.
  • npm run check — unaffected; the new directories are excluded and verified by their own pipeline.
  • npm run build:llms — not applicable, docs/utilities.md untouched.
  • Not a breaking change: nothing in src/ is touched and nothing is published from here yet.
  • Nothing added to the package's dependencies, runtime or development. The engine has its own package.json: oxc-parser to parse, and TypeScript and Prettier for development. Generated code has none, in any of the four languages.

Additional context

Benchmarks, language by language (core/bench/) — each generated core against the implementation that language's community actually ships, not against something written for the occasion. Every pair has to agree on every input before either is timed; none disagreed.

language handwritten baseline generated, against it
TypeScript this package's src/ 0.49x to 0.76x
Python brazilian-utils/python 1.02x to 1.05x
Go brazilian-utils/go 0.13x to 0.30x
Rust brazilian-utils/rust 3.59x and 10.89x

Four defects this found, and the reason the cross-language benchmark was worth building:

  • The checker treated break and continue as function exits, so the state they carry was discarded instead of reaching the loop head. A local assigned just before a continue re-entered the next iteration with its pre-loop type, the fixpoint converged on a range the loop exceeds, and every proof resting on that range — an unchecked seq.get among them — was unsound. Reproduced against the reference interpreter before fixing; engine/tests/loops.spec.ts covers it.
  • A break that ended a switch case was lowered as a loop break, so the generated Python left the enclosing loop where the source left the switch.
  • The Go target compiled each regex inside the function that used it, and regexp has no compilation cache, so a Unicode-class pattern was rebuilt from its source string on every call — 79.6x the cost of the match itself. JavaScript caches a compiled literal and Python caches inside re, so only Go ever paid it and the TypeScript benchmark had been green throughout. That is the argument for measuring against another language's real implementation instead of against yourself.
  • The Rust target interpreted a pattern tree at run time, allocating per node per position. It now emits a straight-line scanner decided at generation time, which took isValidCpf from 50x to 10.9x and isValidCnpj from 24x to 3.6x.

Rust's remaining gap is measured rather than assumed, and it is not the regex: cpf_check_digit calls a string-taking helper inside its loop and every value is owned, so it clones an eleven-character string once per weight. Closing it means inferring which parameters are only read, which supersedes ADR 0009 rather than extending it, so it is recorded instead of half-done.

How much of the package could be authored once: core/docs/survey.md classifies all 138 utilities by the feature they need. 120 use only features the engine supports today; the other 18 need Map/Set, discriminated unions or Unicode normalization, all of which are rejected today with a diagnostic rather than mistranslated.

Deliberately not done: a tree-shaking bundle gate, a baked Dataset type, Map/Set, discriminated unions, recursion, and any publishing. Nothing here is wired into src/ yet either — the generated core is read, verified and benchmarked, but no published utility is served by it. This is a validation of the idea, and it is a draft for that reason.

Where the engine deviates from the plan it was built to, the deviation is an ADR with the evidence: engine/docs/decisions/.

🤖 Generated with Claude Code

https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf

…n and Go

Write a utility's logic once, in a restricted and semantically typed subset of
TypeScript, and generate a native, idiomatic implementation for each target.
No shared runtime package, no WASM, no FFI, no bridge, no interpreter at run
time, and no external dependencies in the generated code.

The engine (`engine/`) is domain neutral and self contained, so it can be
extracted into its own repository later: one dependency (the parser), no build
step, its own tests, and a generic example project that keeps any Brazilian
Utils assumption from leaking into the compiler.

`core/` is the first project that uses it: seven pilots — isValidCpf,
isValidCnpj, formatCnpj, getHolidays, isBusinessDay, getAddressInfoByCep and
formatCurrency — covering strings, refinements, ranges, regex, stable sorting,
civil dates, exact decimals, HTTP with retry and a race, and a domain error
hierarchy.

How it works: source -> Semantic HIR -> Core IR -> per target lowering
selection -> Target AST -> printer -> that language's own formatter. Types
carry proofs (an integer's range, a string's character class and length), and a
native lowering is selected only when its preconditions hold — which is why
`str.compare` is native in Python and Go, where comparison is code point order,
and portable in TypeScript, where `<` compares UTF-16 code units.

Verified by `core: npm run verify`: 4252 differential cases through the
reference interpreter and all three targets, in both idiom modes, plus 4245
cases compared against the published package; byte-identical regeneration;
`tsc --strict`, `compileall` and `go vet` clean over the output. The three
benchmarked pilots are faster than the handwritten implementations they
replace (0.51x to 0.82x).

Nothing in `src/` changes. `vite.config.ts` excludes the new directories from
the package's own format, lint and test passes, since they carry their own
toolchain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
@coderabbitai

coderabbitai Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: f06e27d2-2f6f-419b-bb9b-e49d0170dec2

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@vercel

vercel Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
brazilian-utils Error Error Sep 22, 2026 2:28am UTC

The engine carries its own toolchain, so it runs in its own workflow: Node
runs its TypeScript directly, and the generated Python and Go are compiled by
the runner's preinstalled toolchains. The job runs `verify`, which regenerates
every target in both idiom modes, diffs a second generation for determinism,
runs each target's linters over the output, and runs the differential
conformance harness.

The root `tsconfig.json` now excludes those directories, the way the format,
lint and test passes already do: the package's type checking stops at its own
sources.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
@github-actions

Copy link
Copy Markdown
Contributor

Tree-shaking report

✅ No bundle size impact. All 155 exports are the same size as on the base branch (full import 648.9 KB, gzip 166.2 KB).

All exports (155)
Export Base Head Δ gzip
⚪ GetAddressInfoByCepError 966 B 966 B 0 B 600 B
⚪ GetAddressInfoByCepNotFoundError 1.0 KB 1.0 KB 0 B 618 B
⚪ GetAddressInfoByCepServiceError 1.0 KB 1.0 KB 0 B 617 B
⚪ GetAddressInfoByCepValidationError 1.0 KB 1.0 KB 0 B 620 B
⚪ GetCepInfoByAddressError 966 B 966 B 0 B 600 B
⚪ GetCepInfoByAddressNotFoundError 1.0 KB 1.0 KB 0 B 618 B
⚪ GetCepInfoByAddressValidationError 1.0 KB 1.0 KB 0 B 620 B
⚪ addBusinessDays 6.8 KB 6.8 KB 0 B 2.8 KB
⚪ capitalize 2.5 KB 2.5 KB 0 B 1.3 KB
⚪ convertCurrencyToWords 2.8 KB 2.8 KB 0 B 1.5 KB
⚪ convertDateToWords 3.2 KB 3.2 KB 0 B 1.7 KB
⚪ convertLicensePlateToMercosul 1.3 KB 1.3 KB 0 B 807 B
⚪ convertNumberToWords 2.4 KB 2.4 KB 0 B 1.3 KB
⚪ differenceInBusinessDays 6.9 KB 6.9 KB 0 B 2.9 KB
⚪ formatBoleto 1.4 KB 1.4 KB 0 B 837 B
⚪ formatCEP 1.2 KB 1.2 KB 0 B 778 B
⚪ formatCNPJ 1.4 KB 1.4 KB 0 B 854 B
⚪ formatCPF 1.3 KB 1.3 KB 0 B 806 B
⚪ formatCaepf 1.3 KB 1.3 KB 0 B 787 B
⚪ formatCei 1.3 KB 1.3 KB 0 B 785 B
⚪ formatCep 1.2 KB 1.2 KB 0 B 778 B
⚪ formatCertidao 1.3 KB 1.3 KB 0 B 789 B
⚪ formatCnae 1.2 KB 1.2 KB 0 B 782 B
⚪ formatCnh 1.3 KB 1.3 KB 0 B 780 B
⚪ formatCno 1.3 KB 1.3 KB 0 B 786 B
⚪ formatCnpj 1.4 KB 1.4 KB 0 B 854 B
⚪ formatCns 1.3 KB 1.3 KB 0 B 780 B
⚪ formatCpf 1.3 KB 1.3 KB 0 B 806 B
⚪ formatCurrency 1.8 KB 1.8 KB 0 B 1.0 KB
⚪ formatIban 1.1 KB 1.1 KB 0 B 696 B
⚪ formatLegalNature 1.2 KB 1.2 KB 0 B 777 B
⚪ formatLicensePlate 1.2 KB 1.2 KB 0 B 738 B
⚪ formatNcm 1.2 KB 1.2 KB 0 B 780 B
⚪ formatNfeKey 1.3 KB 1.3 KB 0 B 784 B
⚪ formatPassport 1.0 KB 1.0 KB 0 B 643 B
⚪ formatPhone 2.8 KB 2.8 KB 0 B 1.5 KB
⚪ formatPis 1.3 KB 1.3 KB 0 B 781 B
⚪ formatProcessoJuridico 1.3 KB 1.3 KB 0 B 785 B
⚪ formatVoterId 1.3 KB 1.3 KB 0 B 821 B
⚪ generateBoleto 2.0 KB 2.0 KB 0 B 1.1 KB
⚪ generateCNPJ 1.6 KB 1.6 KB 0 B 965 B
⚪ generateCPF 1.4 KB 1.4 KB 0 B 878 B
⚪ generateCep 984 B 984 B 0 B 610 B
⚪ generateCnh 1.4 KB 1.4 KB 0 B 828 B
⚪ generateCnpj 1.6 KB 1.6 KB 0 B 965 B
⚪ generateCpf 1.4 KB 1.4 KB 0 B 878 B
⚪ generateLegalNature 5.9 KB 5.9 KB 0 B 2.1 KB
⚪ generateLicensePlate 1.1 KB 1.1 KB 0 B 692 B
⚪ generatePassport 1.1 KB 1.1 KB 0 B 656 B
⚪ generatePhone 1.5 KB 1.5 KB 0 B 900 B
⚪ generatePis 1.2 KB 1.2 KB 0 B 744 B
⚪ generatePixPayload 6.3 KB 6.3 KB 0 B 2.8 KB
⚪ generateProcessoJuridico 1.4 KB 1.4 KB 0 B 870 B
⚪ generateRenavam 1.2 KB 1.2 KB 0 B 760 B
⚪ generateVoterId 1.7 KB 1.7 KB 0 B 1021 B
⚪ getAddressInfoByCep 4.1 KB 4.1 KB 0 B 1.9 KB
⚪ getAreaCodeInfo 3.9 KB 3.9 KB 0 B 1.4 KB
⚪ getAreaCodesByState 1.6 KB 1.6 KB 0 B 917 B
⚪ getBankByCode 38.6 KB 38.6 KB 0 B 9.8 KB
⚪ getBankByIspb 38.6 KB 38.6 KB 0 B 9.8 KB
⚪ getBanks 38.4 KB 38.4 KB 0 B 9.6 KB
⚪ getBoletoInfo 3.1 KB 3.1 KB 0 B 1.6 KB
⚪ getCbo 119.1 KB 119.1 KB 0 B 30.7 KB
⚪ getCepInfoByAddress 2.7 KB 2.7 KB 0 B 1.4 KB
⚪ getCertidaoInfo 1.8 KB 1.8 KB 0 B 1.0 KB
⚪ getCfop 68.9 KB 68.9 KB 0 B 6.9 KB
⚪ getCities 154.3 KB 154.3 KB 0 B 49.9 KB
⚪ getCnae 93.9 KB 93.9 KB 0 B 21.2 KB
⚪ getFormatLicensePlate 1.1 KB 1.1 KB 0 B 692 B
⚪ getHolidays 6.1 KB 6.1 KB 0 B 2.6 KB
⚪ getIbanInfo 1.6 KB 1.6 KB 0 B 955 B
⚪ getLegalNature 6.3 KB 6.3 KB 0 B 2.3 KB
⚪ getLegalNatures 5.9 KB 5.9 KB 0 B 2.1 KB
⚪ getLegalNaturesByCategory 6.5 KB 6.5 KB 0 B 2.4 KB
⚪ getMunicipalities 156.4 KB 156.4 KB 0 B 50.3 KB
⚪ getMunicipality 154.9 KB 154.9 KB 0 B 50.3 KB
⚪ getMunicipalityByCode 156.5 KB 156.5 KB 0 B 50.4 KB
⚪ getNfeKeyInfo 2.7 KB 2.7 KB 0 B 1.5 KB
⚪ getPixKeyInfo 4.5 KB 4.5 KB 0 B 2.0 KB
⚪ getPixPayloadInfo 2.9 KB 2.9 KB 0 B 1.4 KB
⚪ getStateByIbgeCode 3.2 KB 3.2 KB 0 B 1.1 KB
⚪ getStateCodeByName 3.2 KB 3.2 KB 0 B 1.1 KB
⚪ getStateNameByCode 3.1 KB 3.1 KB 0 B 1.0 KB
⚪ getStates 3.0 KB 3.0 KB 0 B 1017 B
⚪ getTimezoneByState 1.6 KB 1.6 KB 0 B 809 B
⚪ isBusinessDay 6.5 KB 6.5 KB 0 B 2.7 KB
⚪ isHoliday 6.4 KB 6.4 KB 0 B 2.7 KB
⚪ isValidBankAccount 7.4 KB 7.4 KB 0 B 2.8 KB
⚪ isValidBoleto 2.4 KB 2.4 KB 0 B 1.3 KB
⚪ isValidCEP 984 B 984 B 0 B 610 B
⚪ isValidCNPJ 1.6 KB 1.6 KB 0 B 914 B
⚪ isValidCPF 1.3 KB 1.3 KB 0 B 805 B
⚪ isValidCaepf 1.5 KB 1.5 KB 0 B 913 B
⚪ isValidCbo 119.2 KB 119.2 KB 0 B 30.7 KB
⚪ isValidCei 1.5 KB 1.5 KB 0 B 899 B
⚪ isValidCep 984 B 984 B 0 B 610 B
⚪ isValidCertidao 1.6 KB 1.6 KB 0 B 938 B
⚪ isValidCfop 68.9 KB 68.9 KB 0 B 6.9 KB
⚪ isValidCnae 94.0 KB 94.0 KB 0 B 21.2 KB
⚪ isValidCnh 1.4 KB 1.4 KB 0 B 856 B
⚪ isValidCno 1.5 KB 1.5 KB 0 B 901 B
⚪ isValidCnpj 1.6 KB 1.6 KB 0 B 914 B
⚪ isValidCns 1.5 KB 1.5 KB 0 B 925 B
⚪ isValidCpf 1.3 KB 1.3 KB 0 B 805 B
⚪ isValidCreditCard 1.4 KB 1.4 KB 0 B 868 B
⚪ isValidCsosn 1.2 KB 1.2 KB 0 B 737 B
⚪ isValidCst 1.8 KB 1.8 KB 0 B 1.0 KB
⚪ isValidEmail 1.0 KB 1.0 KB 0 B 622 B
⚪ isValidIE 5.7 KB 5.7 KB 0 B 2.1 KB
⚪ isValidIban 1.3 KB 1.3 KB 0 B 836 B
⚪ isValidIe 5.7 KB 5.7 KB 0 B 2.1 KB
⚪ isValidLandlinePhone 1.5 KB 1.5 KB 0 B 933 B
⚪ isValidLegalNature 5.8 KB 5.8 KB 0 B 2.1 KB
⚪ isValidLicensePlate 1.1 KB 1.1 KB 0 B 702 B
⚪ isValidMobilePhone 1.6 KB 1.6 KB 0 B 971 B
⚪ isValidNcm 114.2 KB 114.2 KB 0 B 24.6 KB
⚪ isValidNfeKey 2.7 KB 2.7 KB 0 B 1.5 KB
⚪ isValidPIS 1.2 KB 1.2 KB 0 B 785 B
⚪ isValidPassport 1.0 KB 1.0 KB 0 B 654 B
⚪ isValidPhone 2.6 KB 2.6 KB 0 B 1.3 KB
⚪ isValidPis 1.2 KB 1.2 KB 0 B 785 B
⚪ isValidPixKey 4.6 KB 4.6 KB 0 B 2.1 KB
⚪ isValidPixPayload 2.9 KB 2.9 KB 0 B 1.5 KB
⚪ isValidProcessoJuridico 1.3 KB 1.3 KB 0 B 787 B
⚪ isValidRegistroProfissional 1.6 KB 1.6 KB 0 B 964 B
⚪ isValidRenavam 1.3 KB 1.3 KB 0 B 815 B
⚪ isValidServicePhone 1.5 KB 1.5 KB 0 B 846 B
⚪ isValidVin 1.6 KB 1.6 KB 0 B 995 B
⚪ isValidVoterId 1.6 KB 1.6 KB 0 B 900 B
⚪ parseBoleto 1020 B 1020 B 0 B 634 B
⚪ parseCaepf 1003 B 1003 B 0 B 621 B
⚪ parseCbo 1002 B 1002 B 0 B 620 B
⚪ parseCei 1003 B 1003 B 0 B 619 B
⚪ parseCep 1002 B 1002 B 0 B 620 B
⚪ parseCertidao 1003 B 1003 B 0 B 621 B
⚪ parseCfop 1002 B 1002 B 0 B 620 B
⚪ parseCnae 1002 B 1002 B 0 B 620 B
⚪ parseCnh 1003 B 1003 B 0 B 621 B
⚪ parseCno 1003 B 1003 B 0 B 619 B
⚪ parseCnpj 1.1 KB 1.1 KB 0 B 669 B
⚪ parseCns 1003 B 1003 B 0 B 621 B
⚪ parseCpf 1003 B 1003 B 0 B 621 B
⚪ parseCurrency 1.4 KB 1.4 KB 0 B 881 B
⚪ parseIban 1.0 KB 1.0 KB 0 B 638 B
⚪ parseLegalNature 1002 B 1002 B 0 B 620 B
⚪ parseLicensePlate 1.0 KB 1.0 KB 0 B 638 B
⚪ parseNcm 1002 B 1002 B 0 B 620 B
⚪ parseNfeKey 1.0 KB 1.0 KB 0 B 659 B
⚪ parsePassport 1.0 KB 1.0 KB 0 B 637 B
⚪ parsePhone 1.1 KB 1.1 KB 0 B 707 B
⚪ parsePis 1003 B 1003 B 0 B 621 B
⚪ parseProcessoJuridico 1003 B 1003 B 0 B 621 B
⚪ parseVoterId 1.0 KB 1.0 KB 0 B 650 B
⚪ removeAccents 953 B 953 B 0 B 593 B
⚪ subBusinessDays 6.9 KB 6.9 KB 0 B 2.9 KB
How this is measured

Every export is imported alone into an esbuild consumer bundle (minified, tree-shaken) built from the head and from the base of this pull request; the sizes are the resulting bundles, gzip is their gzipped size. 🔴 marks a regression: a pre-existing export that grew more than 20% and more than 256 B, or the bundle importing every pre-existing export growing more than 5%. 🟡 is growth under the threshold, 🟢 a decrease, ⚪ no change, 🆕 an export that does not exist on the base (never a regression), 🗑️ an export that was removed. An intentional increase is accepted with the tree-shaking: accepted label.

Comment thread engine/src/targets/typescript/index.ts Fixed
Comment thread engine/src/targets/typescript/index.ts Fixed
Comment thread engine/src/targets/typescript/index.ts Fixed
@codecov

codecov Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 100.00%. Comparing base (69e9b1f) to head (ab4e22e).

Additional details and impacted files
@@            Coverage Diff            @@
##              main      #587   +/-   ##
=========================================
  Coverage   100.00%   100.00%           
=========================================
  Files          183       183           
  Lines         2069      2069           
  Branches       612       612           
=========================================
  Hits          2069      2069           
Flag Coverage Δ
node 100.00% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

`some(unwrap(x))` is `x`: the checker inserts the unwrap where it proved the
value present, and re-wrapping it was a round trip every target printed. The
Go backend now emits the absence check as a comparison node as well, so
negating it reads `response != nil" rather than `!(response == nil)`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
An enum's rendered type contains `|`, which ended the column and broke every
row that mentioned one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
The checker treated `break` and `continue` the way it treats `return`: as
statements that leave the enclosing function, so the state they carry was
discarded instead of reaching the loop head and the code after the loop.

A local assigned just before a `continue` therefore re-entered the next
iteration with its pre-loop type, and one assigned before a `break` was read
after the loop with the same stale type. The fixpoint then converged on a
range the loop actually exceeds, and every proof resting on that range, an
unchecked `seq.get` among them, was unsound. Reproduced against the reference
interpreter: the checker proved `Int[0..0]` for an accumulator the interpreter
answers `15` for, and accepted an index the loop drives past the end.

The loop body now reports the three ways it can be left, and the fixpoint
merges the fall-through, `continue` and `break` states into the head on every
pass; the state after the loop is the head joined with fall-through and the
`break` states. The same repro is now rejected with the real range, `Int[0..15]`.

A `break` that ends a `switch` case was also lowered as a loop break, so the
generated Python left the enclosing loop where the source left the switch. The
trailing case break is now dropped when the case is translated, and any other
break inside a case is rejected with `E_SWITCH_BREAK` rather than silently
mistranslated.

`tests/loops.spec.ts` covers all of it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
The TypeScript target wrote `./capabilities.ts` into every module that needs
a capability type, but the file sits at the output root, so a module under
`lib/` pointed at a path that does not exist. Nothing caught it until a
library module started taking a capability, because until then only root
modules imported it.

The support import now travels as the same `SUPPORT` sentinel the Python
target already uses, and the printer resolves it against the module's own
path, so `lib/cnpj.ts` reaches for `../capabilities.ts`.

Also moves `groupThousands` and `easterSunday` out of the utilities that
happened to define them and into `lib/`, restoring the one-export-per-file
rule the source root is supposed to keep: the date construction both Easter
and the fixed holidays need is now `lib/civil.ts`, so `get-holidays.ts` and
`format-currency.ts` export a single utility again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
Reading what the engine produces is the point of the project, so `out/` now
lives in the repository: the TypeScript, Python and Go each utility lowers
to, the `LOWERING.md` that explains every choice, and the `API.json` and
`SOURCEMAP.json` that go with them.

The `--no-idioms` trees stay out, being a comparison the verification builds
rather than something anyone reads, and CI now fails when the committed
output drifts from what the engine emits for the committed sources.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
Both were allowed and used, but the specification listed neither, so the rule
the checker now enforces had nowhere to be read: the state a `break` or a
`continue` carries belongs to the loop's fixpoint, and a `break` inside a
`switch` case may only be the one that closes the case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
`generateCpf` and `generateCnpj` are the first utilities that draw, so they
are the first test of whether a random-using utility can be written once and
still agree across three languages. Everything past the single `random.nextU32`
intrinsic is written in the subset — `lib/random.ts` derives a uniform value
below a bound by rejection sampling rather than `% bound`, because the modulo
bias of a fixed-width draw would otherwise have to match digit for digit
across three unrelated standard libraries to stay invisible.

The subset has no unbounded `while`, so both the rejection search and the
repeated-base redraw are counted loops whose fallback probability is worked
out in the comment that states it: under 2^-32 for the draw, under 10^-64 and
10^-88 for the bases.

A base is built as straight-line concatenation rather than in a loop: a loop's
fixpoint proves a range over every iteration, `Digits[0..12]`, not the exact
length the check-digit functions require, while `str.concat` sums the bounds
exactly.

The three generated differential drivers used to answer 0 to every draw, and
built one capability object for a whole batch. They now carry the reference
PCG32, with the interpreter's own constants and default seed, constructed per
request line the way the reference builds a fresh interpreter per case — so
the comparison is call for call. The conformance run also derives an
`isValidCpf`/`isValidCnpj` case from whatever the interpreter actually drew,
which is what would catch three targets agreeing bit for bit on an invalid
document.

Conformance: 4256/4256 in every target, in both idiom modes; 4247/4247 against
the published package.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
Three things the committed output made obvious, all in the Python target and
none in what the program means.

`opt.orElse` printed its operand twice. Over a checked accessor that produced
a conditional nested inside its own condition, six lines where one belongs;
over a call it ran the call twice, so `get_address_info_by_cep` parsed each
JSON field two times. The checked lowerings now build the conditional as a
node instead of as text, so a default drops straight into the `else` branch,
and anything else is bound once with a walrus.

`opt.isNone` was text too, so negating it gave `not response is None`. It is
a comparison node now, and the printer negates it to `response is not None`.

Docstrings kept only the first line of the source doc, which cut sentences in
half wherever the prose wraps. A doc that spans lines now keeps all of them,
in the shape PEP 257 describes.

Conformance is unchanged at 4256/4256 in every target and both idiom modes;
the output is regenerated in the same commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
Every number in `progress.md` comes from `scripts/metrics.ts`, the conformance
run or the benchmark, so they are re-derived rather than edited: nine
utilities, the lowering mix, the per-part line counts, the per-utility
expansion, 4256/4256 in every target and 4247/4247 against the published
package, and a fresh benchmark at 0.61x, 0.78x and 0.55x.

The marginal-cost section now records what the two generators cost: one more
intrinsic, `random.nextU32`, and no compiler change beyond it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
The drift check failed on its first run, and the reason is worth keeping: a
missing formatter was a silent skip, so a checkout without prettier or ruff
generated the same program with different bytes. CI had neither, produced
unformatted output, and reported it as a diff against the committed files.

Prettier is now a pinned dev dependency of the engine, resolved from the
engine's own `node_modules` and run with `--no-config` so an ancestor
configuration cannot change the result; ruff is installed in the workflow at
the version `toolchain.lock.json` records; gofmt comes with Go.

`verify` now starts by checking all three are installed and fails if one is
not, and the CLI warns on stderr when it leaves a target unformatted. The
generated output is unchanged by all of this, which is the point.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
`semantics.md` listed only the engine's namespaces, which are now the older
of two accepted spellings. Section 7.1 carries the mapping and the five forms
that are refused rather than mapped, each with why JavaScript's meaning has no
counterpart in the other three languages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
Before this, `core/source` compiled and did not run. `tsc` accepted it because
the engine ships ambient declarations, but importing a utility and calling it
failed with `ReferenceError: re is not defined` — `str`, `re`, `seq` and `int`
were names with nothing behind them. "The author writes TypeScript" was not
true, and saying it anyway would have been a lie to whoever used this.

The source now uses the ordinary spelling wherever the frontend reads it:
`.charCodeAt(i)`, `.replace(/[^0-9]/g, "")`, `PATTERN.test(s)`, `.trim()`,
`.padStart(n, c)`, `String(n)`, `Math.min`/`max`, `xs[i]`. `isValidCpf` and
`isValidCnpj` now execute in plain Node with no engine involved and answer
correctly.

`conformance (source, no engine)` is a new verify step that proves it rather
than claiming it: it imports the utilities and runs the same vectors against
them, with no compilation and no generation. `isValidCpf` matches 1519 of
1519. `isValidCnpj` matches 1216 of 2032, the rest blocked on one named gap,
and the step fails if either number moves in either direction — a gap that
closes silently and a utility that stops running are both things worth
knowing.

Seven utilities do not run yet, each for a reason the step names: `dec.*`
because JavaScript has no exact decimal, `random.nextU32` and `http.request`
because they are effects, `date.*` because `new Date` has no
proleptic-Gregorian equivalent, and `formatCnpj` because `formatWithPattern`
reaches a checked accessor on every call. `core/docs/idiomatic-migration.md`
assesses, honestly and without implementing any of it, whether importing a
real module could give each of them an ordinary spelling — `Temporal` is
plausible for dates, a decimal library is harder because the scale is a type
parameter here and a value there, and randomness, HTTP and `race` are not
spellings of the same shape at all.

Three call sites deliberately keep the namespace form, each with a comment:
the index is provable there, so the bracket spelling would select the
unchecked accessor — a different Core, not the same meaning respelled.

The generated output does not move: the entire diff is four `SOURCEMAP.json`
files, whose byte offsets into the source necessarily shift when the source
text does. No generated program logic changed, in any target, in either idiom
mode, and conformance holds at 4256/4256 throughout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
A separate branch, pull request #580, explored the same problem with a
different answer: a compiler to seven targets whose output leans on a
per-language runtime shipped beside it. That branch is closing, and one part
of it is worth more than the code — the measurements that answer a question
this engine never addressed, which is why generate source at all instead of
shipping one binary core and binding to it.

Two findings decide it, and only one is about speed. A tree-shakeable package
cannot take a binary core: one utility as generated source was 509 bytes that
disappear when unused, against 9,291 indivisible and asynchronous ones. Go's
cost for cgo is cross compilation, static binaries and `CGO_ENABLED=0`, not
its 58 ns. Neither is a preference between acceptable options.

The finding that does not decide it is kept too, because it is the one that
makes the position honest: for Python, Ruby, C# and Java a binding really is
cheaper than generated source, and that document's first version got it wrong
by measuring the bindings a script reaches for rather than the ones a package
ships. So the answer is scoped — JavaScript and Go must have generated source
on grounds unrelated to speed, and the other four are a choice this engine
makes for readability and the absence of a runtime rather than for throughput.

ADR 0012 records that, including what it costs: four emitters to maintain and
each language's semantics encoded four times. The document carries a header
saying the harness behind its numbers was retired with the branch, so they
should be reproduced before anything rests on them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
… case

The frontend chose between the unchecked and the checked accessor by
provability alone: `xs[i]` became `seq.get` when the index was proven in range
and `seq.at` when it was not. That rule is why three call sites in
`core/source` had to keep the engine's spelling — there the index *is*
provable, but the author wants the `Option` anyway, and the bracket spelling
would have quietly selected the unchecked accessor and changed the Core.

Provability is the wrong signal on its own. In TypeScript under
`noUncheckedIndexedAccess`, `xs[i]` is `T | undefined` whatever the compiler
can prove, and someone writing `xs[i] ?? fallback` has said in the language's
own terms that the absent case is wanted. So the `??` decides: it selects the
checked accessor unconditionally, and a bare `xs[i]` keeps the old rule.

That unblocks what had no ordinary spelling. `formatCnpj` goes from running
none of its cases to all 160 — `formatWithPattern` reached a checked accessor
on every call, so it was not partially runnable, it was not runnable at all.
`isValidCnpj` goes from 1216 of 2032 to 1627. `isValidCpf` stays at 1519 of
1519. The three call sites that carried a comment explaining why they could
not move are now written the ordinary way, with the comment explaining why the
`??` is load-bearing rather than decorative.

Generated output does not move: the entire diff is the four `SOURCEMAP.json`
files, whose byte offsets follow the source text. Conformance holds at
4256/4256 per target in both idiom modes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
The rule that `xs[i] ?? fallback` always selects the checked accessor needs
its reasoning recorded, because the alternative looks defensible until it is
written out: under `noUncheckedIndexedAccess` real TypeScript types a bracket
index `T | undefined` whatever the engine can prove about the index, since
`tsc` has no access to that proof. So `?? fallback` is the author's statement
in the language's own terms, not a claim about provability.

Deciding it by provability instead would mean the same source text lowering
to a different Core depending on a fact about the value that the author cannot
see from where they are standing — which is the kind of surprise this whole
frontier exists to remove.

Also records `s[i]?.charCodeAt(0)` as the checked numeric accessor, since
`s.charCodeAt(i)` answers `NaN` past the end rather than `undefined` and so
has no `??` form of its own, and brings `core/docs/idiomatic-migration.md` up
to date with what now runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
`number` was refused outright, and it was the single biggest barrier between
"write TypeScript" and the truth: of twelve ordinary-TypeScript probes, eight
hit `E_BARE_NUMBER`, and for four of them it was the only diagnostic that got
out at all — a frontend check that cut the compile short before the real gap
was ever reached.

It is accepted now, with the range inferred rather than demanded, from the two
places the author already wrote it.

A helper takes its contract from **its call sites**: the checker already
specializes one per call site, which is how `digitAt` is checked against an
eleven-digit CPF and a fourteen-digit CNPJ today, so a `number` parameter
there simply starts at the platform-safe default and the caller's proven type
is substituted in. Nothing new was needed.

An exported utility has no call site, so it takes its contract from **a guard
in its own body**. Every read of the parameter outside a guard's own condition
is tracked, and the union of what was proven at those reads becomes the
published type. The distinction matters: a guard's condition proves nothing
about a use that has not happened yet, only its result narrows anything.

What is left is an exported function that guards nothing, and that is refused
rather than assumed or silently degraded:

    E_BARE_NUMBER: `n` is a bare `number`, and `isPositive` never narrows it
    before using it
      help: add a guard before n is used, for example
      `if (n < 0 || n > 99) return …;` — the range a guard like that proves
      becomes n's published contract; write `n: Int` instead if the full range
      really is what is meant

Five annotations in `core/source` were only ever a way around the old ban and
are gone. The thirty-eight that remain are contract: `year: IntRange<1900,
2099>` is real-world knowledge no guard in that body states, and `Digits` and
`Ascii` refine a string rather than a number.

Re-running the twelve probes found two claims in `subset-gaps.md` that were
wrong. Method call syntax no longer needs a recognizer — that landed earlier.
And an optional parameter is not refused at all: `params()` never reads the
`?`, so `suffix?: string` is silently taken as a required `string`, which is a
mistranslation rather than a gap, and worse than one. Both corrected there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
… does

The benchmark's bar moved from 1.5x to 1.0x — generated code equal to or
faster than the implementation each community ships by hand. Three shapes
accounted for the rows above it, and each got its own answer.

**A candidate nobody wrote.** TypeScript's `getHolidays` spent 78% of its call
in `civilDate`, computing a date and then verifying it by a round trip through
three floor-division helpers, while the host has `new Date` right there. The
first attempt used `Date.UTC` and measured *slower* than the round trip it
replaced — object construction costs more than the arithmetic — so what
shipped is a days-in-month check and a single forward computation. 1.77x to
0.93x, and `isBusinessDay` follows it down to 0.66x.

**Per-call overhead, with a budget declared per target.** A small pure
function in a loop is free in V8 and in LLVM and expensive in CPython, so
inlining is a cost decision per target like every other choice here: Python
aggressive at twelve statements, TypeScript conservative at six — budget one
was tried first and barely moved, which is itself the evidence that V8 was not
the bottleneck at that layer — and Rust none at all, because the inliner has
no notion of a borrow and spliced an owned local where ADR 0010 had removed
the clone.

**Allocation per step.** `str.concatAll` assembles a chain into one buffer
instead of a fresh string per link, and `re.retain` and the code-point
conversions work on ASCII bytes rather than decoding UTF-8.

Ten of twenty-seven rows are still above 1.0x and each is recorded with why,
because a row that cannot beat a hand-written implementation for a structural
reason is a finding rather than an omission. Python's `formatCurrency` is an
author's explicit character loop against one C-level formatter; its generators
are a rejection-sampling loop run nine to twelve times against a single
`randint`. Rust's remaining cost is spread thin across the scanner, the mask
trim and one allocation, with no dominant term left — `#[inline]` hints were
tried, measured to change nothing over six runs, and reverted rather than kept
on faith.

TypeScript's inlining costs bytes and the number is here rather than buried:
the generated tree grows 20%, from 156,322 to 187,760, partly offset by
`lib/random.ts` disappearing and `std/date.ts` shrinking. The package is
tree-shakeable and that is a trade to state.

A checkpoint during this work caught a real bug it had introduced: a
double-evaluated capacity computation inside `str.concatAll` corrupted the
draw sequence, and `rust-plain` produced a perfectly valid CPF that was not
the one the reference produced. It compiled, it linted, the document
validated, and only the differential comparison against the interpreter saw
it. Fixed, and the whole suite re-run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
The TypeScript target's output is shipped to a browser by a tree-shakeable
package, so bytes over the wire are a result it is judged on. The inlining
budget had no notion of that: it sized the callee and never priced the copy.
Measured the way a consumer's bundler would (one single-import entry point per
utility, esbuild, minified, gzipped) it cost +75% across every export and +148%
on generateCnpj alone, for two benchmark rows it moved by about a tenth each --
inside the noise of a generator whose own retry loop is random.

- `InlineBudget.maxDuplicatedNodes` caps what one inline may add. A splice that
  takes a callee's last call site is refunded its whole definition, which falls
  out of the dependency closure, so a sole-call-site helper is absorbed whatever
  its size and `0` means "take the inlines that pay for themselves".
- A callee that is one `return <expr>` is substituted as an expression, and a
  straight-line callee is spliced without the early-return sentinel; a parameter
  passed a literal or a local is substituted rather than bound.
- A read the checker proved to be one integer is printed as that integer, which
  turns a specialized helper's arithmetic into constants; the bindings and
  parameters that leaves unread are removed, since Go and Rust reject both.
- `backend/fold.ts` folds the lowered target AST, where a lowering's own
  expansion of a now-constant argument lives, and TypeScript's `date.fromYmd`
  folds a literal month directly.
- `writeFiles` removes generated files a later run stopped producing. Two stale
  files were in the repository, typechecked and committed as though they were
  output.
- A generated module now exports whatever another generated module imports: a
  surviving specialization is not the declaration it came from and carried that
  flag as false.

Every export is now smaller than with inlining disabled (3,175 bytes gzipped
across all nine, against 3,224 off and 5,644 under the old budget), Go's
generators moved 0.57x/0.44x to 0.39x/0.32x, and the generated TypeScript reads
like its source again. `engine/scripts/size.ts` measures it and `verify`'s new
`typescript size` step holds it there. ADR 0013 has the reasoning, and
`core/bench/README.md` the numbers.

verify 17/17; conformance 4256/4256 on all four targets in both idiom modes,
interpreter vs npm 4247/4247; fuzz full --seed 555 --count 200 clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
…percent

Every target's optimizations are paid for in the currency that target is judged
in, and for this one that currency is bytes a browser downloads: an
optimization here is only an optimization if the bundle does not grow for it.
The 5% tolerance the size gate started with was a number with nothing behind
it. Gzipped output is deterministic, so zero is not noisy either.

A change that needs more room is a change whose trade has to be argued,
measured and written down, and then accepted with `--write`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
Rust was the worst target by some distance. The cause was not one thing, so
each of these was timed on a scratch crate built against the generated one
before being written as a candidate:

- `re.retain` builds its result with a `for` loop into one `String` instead of
  `.filter().collect::<Vec<u8>>()` handed to a UTF-8 validator that re-walks
  it. Every byte pushed is already a range-tested ASCII byte. 11.4-12.3ms to
  5.4-7.1ms per 200k calls.
- `re_take_fixed`/`re_take_class` test the leading byte directly and decode a
  full `char` only above 0x7F. These two are the whole of every generated
  chain-pattern scanner, so this speeds up every regex-shaped validator the
  engine will ever emit, not just the rows that motivated it. 6.7ms to 4.4ms.
- `str.padStart`, gated on both the value and the pad being proven ASCII,
  where a scalar count is a byte count and neither `Vec<char>` is needed.
  22.8-24.2ms to 10.4-10.5ms. This is most of `formatCurrency`'s move.
- `str.codePoints`/`str.fromCodePoints`, ASCII-gated, mirroring what Go
  already had and Rust was missing.
- `str.fromInt` for a value the checker proved is one decimal digit: one ASCII
  byte, not the general integer formatter.

Measured over three runs, generated over handwritten:

  isValidCpf     2.73x -> 1.44x
  isValidCnpj    1.43x -> 0.96x, now faster than the crate it compares against
  formatCurrency 1.81x -> 1.05x
  generateCpf    1.66x -> 1.51x
  generateCnpj   1.22x -> 1.15x

Two results that are not speedups, kept because they are worth knowing. An
`unsafe from_utf8_unchecked` variant of `keep_digits` measured 5.4-5.6ms
against the safe loop's 5.4-5.5ms, so the generated code stays `unsafe`-free
for a measured reason rather than a stylistic one. And the first version of
the `str.fromInt` candidate returned a `raw` text fragment, which stringifies
its argument immediately and so hid a list literal from `hoistConstantTables`:
a weight table that had been a module-level `const` silently became a `vec![]`
allocated per call. Conformance did not catch it, since the answers were
identical; `cargo clippy`'s `useless_vec` did. The candidate now builds a
structured `call` node like the others.

The remaining generator gap is the per-digit `String` allocation, which is
roughly 40% of each draw. Removing it needs a return type that is not owned,
which ADR 0009/0010 deliberately declined, so it is recorded rather than
worked around.

verify 17/17; the other three targets' output is byte-identical.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
…er parses

Gzip alone is the wrong single number. Its window makes locally repeated text
almost free, so an encoding can be meaningfully shorter raw and larger
gzipped — measured, not hypothetical: replacing nine unrolled concatenations
with a counted loop removed 66 minified bytes from `generateCpf` and added 19
gzipped ones, because what it removed was near-perfect redundancy that LZ77
was already compressing for nothing.

So the harness reports three numbers and gates two. `minified` is what the
browser parses and the engine holds; it is not a transfer size, and trading it
against one is an argument to have rather than a threshold to trip, so it is
reported only. `gzip` and `brotli` are both transfer sizes and they disagree
about which source is smaller, and a consumer gets whichever their CDN
negotiates, so both are gated at zero growth.

Brotli is measured at quality 11 with a size hint, since the streaming
defaults would understate what a CDN serving a static asset produces.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
…irmed against the parser

The row said the frontend `params()` "never reads a parameter's `?`" as an
inference from the diagnostic that came out. Checked against the parser
directly: `b?: string` parses as an `Identifier` carrying `optional: true`
and `params()` reads only `name` and `typeAnnotation`, so the flag is
dropped. `c: boolean = false` parses as an `AssignmentPattern` and is
rejected loudly.

So one half of this row is a refusal and the other is a silent miscompile,
which is the worst shape a gap can have. Recorded, with what closing it should
look like for each half.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf
…a botched splice

Two things.

The finding: `randomCpfBase` builds its result from nine textually identical
`randomDigit()` calls, and a candidate lowering can re-roll N structurally
identical pieces of a concat chain into one counted loop. It is sound (the
call still runs N times, in order, so the draw sequence the conformance
protocol fixes is untouched), it is faster (both generator rows go from noisy
around 1.0x to 0.68-0.91x), and it removes 66 and 112 minified bytes. It was
still declined, because both transfer encodings grow: +19/+23 gzip and
+19/+19 brotli, +17 on the combined bundle under either.

The reason is worth the space it takes, because the next idea for shrinking
generated TypeScript by removing repeated call text will hit the same wall:
text repeated eight or eleven times is nearly free to an LZ77-family coder,
which spends a short back-reference per repetition instead of the literal
bytes, while the loop that replaces it is shorter but unique and has nothing
to point at. Brotli was measured specifically to test whether this was a gzip
artifact. It is not; both codecs agree on direction.

The botch: the two sections added in the previous two commits were spliced
into the middle of a sentence in the intro, because the marker they were
inserted before also occurs inside backticks there and the edit replaced every
occurrence rather than the heading. Seven thousand characters of duplicated
section have been removed and the sentence restored. Each section now appears
exactly once.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013i7T7KEWrhJFNZ6qaVqJWf

Copy link
Copy Markdown
Member Author

Closing this: the work moved to two repositories of its own, so nothing here is meant to merge.

  • brazilian-utils/engine — the compiler, split out of engine/ with its history. It is generic: its CI verifies it against its own examples/generic project and it knows nothing about this library. A project is any directory with an engine.config.json.
  • brazilian-utils/core — the logic, split out of core/ with its history, plus this repository's src/ copied under javascript/ with its own 555 commits and 33 authors intact. That copy is both the material still to migrate and the reference every conformance case is checked against.

Nothing in this repository was changed. src/ is untouched; the copy was taken with git subtree split, so authorship is preserved rather than flattened into one commit.

Where it got to. Seven targets generate from one source — TypeScript, Python, Go, Rust, Ruby, F# and Erlang — each at 4256/4256 differential conformance in both idiom modes, against the reference interpreter and against this package. Erlang was the interesting one: the Core IR is imperative and Erlang has no loops and no mutable variables, so every reassignment becomes a fresh binding and every loop a self-recursive fun.

Two findings worth carrying over even if the rest is ignored:

  1. [\--/], the character class the mask patterns use, is read as the range 0x2D–0x2F by JavaScript, Python and RE2, and as the two-element set {-, /} by .NET — so . does not match there. Verified against all three engines. A hand-written .NET port of isValidCpf would reject 123.456.789-09 and nowhere else would.
  2. The generated output directory accumulated files the engine had stopped producing; two stale ones were committed here and were being typechecked and benchmarked as though they were output.

Generated by Claude Code

This branch had an error being deployed

1 failed deployment
Preview — ab4e22e7 Deployed Sep 22, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants