Partial protovalidate support for pure Lean 4, focused on CEL annotation support, consuming and extending rules_lean and grpc-lean.
Rather than a procedural CEL evaluator, each CEL-annotated field gets a Lean subtype: a refinement of the plain field type by the proposition the CEL expression denotes. The proto rule
string name = 1 [(buf.validate.field).cel = {
id: "name.len", expression: "this.size() <= 100 && this.startsWith('foo')"
}];corresponds to the field type
{ x : String // x.size ≤ 100 ∧ x.startsWith "foo" }The base types are exactly the types grpc-lean's protobuf codegen produces
(Int32/Int64/UInt32/UInt64, Float/Float32, Bool, String,
ByteArray, Array α for repeated, Std.HashMap κ ν for maps, Option α
for explicit presence, verbatim snake_case field names), and the proposition
is compositional at CEL granularity — the constraint a Lean author would have
written by hand, never an opaque "satisfies this CEL string" predicate.
load("@rules_lean_grpc//:defs.bzl", "lean_proto_library")
load("@protovalidate_lean//protovalidate:defs.bzl", "lean_protovalidate_library")
proto_library(
name = "shipping_proto",
srcs = ["shipping.proto"],
deps = ["@protovalidate//proto/protovalidate/buf/validate:validate_proto"],
)
lean_proto_library(name = "shipping_lean", proto = ":shipping_proto")
lean_protovalidate_library(
name = "shipping_valid",
proto = ":shipping_proto",
lean_proto = ":shipping_lean", # base types the refinements wrap
)lean_protovalidate_library runs protoc-gen-lean-protovalidate (a Go protoc
plugin using celtolean as a library) and compiles its output against the
base types. For every message with protovalidate rules it emits, in a parallel
<package>.Valid namespace:
- a structure whose annotated fields are subtypes of the base field types and whose message-level rules are dependent Prop fields over the value fields;
ValidPred : Base → Prop, the declarative validity predicate: a Prop structure with one field per rule conjunct, phrased over the base value;validate : Base → Except Protovalidate.Violation Valid, deciding each rule (an evidence-carryingDecision (ValidPred b)under the hood, so the proofs flow into the structure) or reporting the first violation with its rule id, plus aDecidable (ValidPred b)instance;- kernel-checked soundness/completeness theorems tying the two together (see below);
toBase : Valid → Base, forgetting refinements;decodeValid : ByteArray → Except String Valid, composing the base wire decoder with validation.
Every generated message additionally carries M.ValidPred : Base → Prop — a
declarative Prop structure with one named field per rule conjunct (presence
conjuncts for required Option shapes, ∀ x, b.f = some x → … for optional
rules, ∀ x ∈ b.f, Elem.ValidPred x for repeated-message elements, nested
ValidPreds for embedded messages, and the message-level CEL Props) — and
kernel-checked theorems relating it to the procedural validator:
theorem M.validate_sound : M.validate b = .ok v → M.ValidPred b
theorem M.validate_complete : M.ValidPred b → (M.validate b).isOk
theorem M.decodeValid_sound : M.decodeValid bytes = .ok v →
∃ b, Base.decode bytes = .ok b ∧ M.validate b = .ok vThe scheme: M.checkPred b : Protovalidate.Decision (M.ValidPred b) decides
the whole predicate at once — each step either extends the proof or refutes
it while reporting the first violation in protovalidate rule order — and
validate is its Except view, so the theorems are instances of the generic
Decision lemmas with no per-message proof search. Soundness hands back the
conjuncts over the input base value ((M.validate_sound h).items_elems),
which downstream code can project and rewrite along without re-running any
checks. Test/PredSchemeTest.lean keeps a hand-written mirror of the emission
scheme compiling.
Five exact-policy lean_assurance_test targets audit axiom closures at
compile time (no sorryAx, standard axioms only), covering both the instances
and the scheme they instantiate. Each target also pins the definitional type
of every principal theorem and the complete in-scope inventory of unsafe,
partial, opaque, @[implemented_by], @[extern], and native dependencies.
Additions and removals fail. Every generated example in the repo is audited:
| target | principal theorems |
|---|---|
//lean:runtime_assurance |
the generic Decision lemmas (pred_of_toExcept_ok, toExcept_isOk, decodeThenValidate_sound), the presence plumbing, size_mapArray — plus a scan of the whole Protovalidate.Cel/Runtime surface the generated propositions are stated over |
//lean:format_laws_assurance |
the Cel.Format recognizer laws (see format assurance), headed by ipv4Value?_lt |
//examples/shipping:shipping_valid_assurance |
shipping.v1.Valid.{Order,Item} (soundness, completeness, decode) and the validated Order.payer_Type sum |
//examples/shipping:inventory_valid_assurance |
catalog.v1.Valid.{Product,Payment} and the Payment.method_Type sum (standard rules, enums, element rules, validated oneof) |
//examples/common:money_valid_assurance |
common.v1.Valid.Money — the cross-target embed, audited at its source |
A repeated-message field's own rules are stated over the validated array
(Protovalidate.mapArray b.f Elem.ofPred h.f_elems), because those rules may
mention the element type's refinements. Protovalidate.size_mapArray moves
the structural ones back to b.f ((mapArray …).size = xs.size); it is a
propositional rewrite, not a definitional one, which is why the conjunct is
not phrased over b.f directly. toBase is likewise not a left inverse of
ofPred — it drops the base value's unknown wire fields — so message rules
traversing a validated field cannot be restated over the base value at all
(Test/PredSchemeTest.lean witnesses both facts).
The standard rule vocabulary lowers onto the same refinement machinery (mostly by synthesizing the equivalent CEL; enum and element rules assemble Lean directly):
- numerics (all 12 int/uint/float kinds):
const,lt/lte/gt/gte(including protovalidate's inverted-range disjunction),in/not_in,finite; - string:
const, len, min_len, max_len, len_bytes, min_bytes, max_bytes, pattern, prefix, suffix, contains, not_contains, in, not_inand the well-known formatsemail, hostname, ip, ipv4, ipv6, uri, uri_ref, address, uuid, tuuid, ip_with_prefixlen (+v4/v6), ip_prefix (+v4/v6), host_and_port; - bytes:
const, len, min_len, max_len, prefix, suffix, contains, in, not_in; - bool:
const; enum:const, defined_only, in, not_in(as constructor tests on the generated open inductive, e.g.¬(x matches .«Unknown.Value» _)); - repeated:
min_items, max_items, unique, and scalaritemsrules as∀ v ∈ x, …; map:min_pairs, max_pairs, and scalarkeys/valuesrules quantified overx.keys/x.values; required: explicit-presence fields are unwrapped — the validated field drops itsOption(validate reportsrequiredonnone,toBasere-wraps withsome); implicit-presence fields gain a non-zero rule;ignore:IGNORE_ALWAYSdrops the field's rules;IGNORE_IF_ZERO_VALUEbecomes a zero-escape disjunct ({ x : String // x.isEmpty ∨ (x.isEmail) }) with a fast path in validate.
string.pattern regexes are gated at codegen time against the runtime
engine's own grammar (a Go port, plus RE2 validity), then #guard-checked at
Lean compile time — both for acceptance and against a per-pattern battery of
RE2-labelled probe strings (see
regex assurance) — and
decided at runtime by the pure-Lean engine in
Protovalidate.Cel.Regex (a Pike-VM NFA: classes, Perl classes, anchors,
alternation, bounded quantifiers; linear time, total). The
email/hostname/ip/uri/... formats are real implementations in
Protovalidate.Cel.Format following protovalidate's grammars, differentially
tested at Lean compile time against protovalidate's own conformance suite (see
format assurance).
timestamp('...')/duration('...') literals fold to
Cel.Timestamp.mk/Cel.Duration.mk constants at codegen time, with RFC
3339/duration parsers and timestamp/duration arithmetic (now - this <= duration('24h')) available at runtime.
From the example's generated code (examples/shipping/):
structure Order where
/-- CEL: `this.size() == 8` (id.len); `this.startsWith('ord-') || ...` (id.prefix) -/
id : { x : String // x.size = 8 ∧ (x.startsWith "ord-" ∨ x.startsWith "ORD-") }
status : String
tracking : Option String
/-- CEL: `this.size() > 0` (items.nonempty) -/
items : { x : Array shipping.v1.Valid.Item // x.size > 0 }
/-- CEL: `this >= 0 && this % 100 == 0` (discount.range) -/
discount_cents : Option { x : Int64 // x ≥ 0 ∧ x % 100 = 0 }
labels : Std.HashMap String String
payer : Option (shipping.v1.Valid.Order.payer_Type) -- validated sum: customer_id min_len
/-- CEL: `this.status == 'SHIPPED' ? has(this.tracking) : true` (order.tracking) -/
order_tracking : status = "SHIPPED" → tracking.isSomeNote the composition rules visible above: message fields recursively use the
validated variant when they carry no field rules; repeated-message fields
with list rules refine the array of validated elements (structural rules
like size() compose with per-element validation); optional constrained
fields refine inside the Option (protovalidate evaluates rules only when
set); constrained oneofs use validated sums with refined constructor
payloads (unconstrained ones ride along as the base sum); message-level rules
become Props over the (refined) sibling fields — the CEL ternary idiom
c ? p : true lands as the implication c → p.
The CEL → Lean expression converter used by the code generator is also usable standalone:
$ bazel run //cmd/cel2lean -- 'this.size() <= 100 && !this.contains("bad")'
x.size ≤ 100 ∧ ¬Cel.contains x "bad"
$ bazel run //cmd/cel2lean -- -var s -json 'this == "" || this.isEmail()'
{"lean":"s = \"\" ∨ s.isEmail","var":"s","kind":"prop"}
It is written in Go around the official cel-go parser (with macro
expansion disabled, so all/exists arrive as calls and can be rendered as
binders). That choice anchors the hardest part — CEL grammar, escapes,
precedence, macro shapes — on the reference implementation; the translator
itself is a compact syntax-directed mapping (celtolean/). Go is a
codegen-time dependency only: nothing generated depends on Go at runtime. The
protoc plugin reads the buf.validate descriptor extensions through Go's
first-class protobuf library support.
Translation consumes only the CEL string (plus a binder name). This works because the output leans on Lean's own elaboration against the field's base type:
- numeric literals stay polymorphic (
this >= 1→x ≥ 1for any ordered numeric field); - CEL functions map to dot-notation resolved by the receiver's type
(
size()→.size,startsWith→.startsWith); - ill-typed CEL (wrong field names, mismatched types) becomes a Lean compile error in the generated code — deliberately deferred to the compiler.
Consequently the tool never generates the full subtype def; the codegen
layer assembles { x : T // <emitted proposition> } itself.
| CEL | Lean | note |
|---|---|---|
&& || ! |
∧ ∨ ¬ |
Props; &&/|| chains flattened |
== != |
= ≠ |
↔ when an operand is itself propositional |
< <= > >= |
< ≤ > ≥ |
|
e in c |
e ∈ c |
list/array/map-key membership |
+ - * / % |
+ - * / % |
+ on strings/bytes/lists via Add shims; int/uint ops carry Cel.addOk-style overflow guards (see below) |
c ? p : true |
c → p |
also c ? p : false → c ∧ p, c ? true : p → c ∨ p |
c ? a : b |
if c then a else b |
decidable ite |
'lit', b'lit', 123, 1.5, null |
"lit", "lit".toUTF8/ByteArray.mk #[..], 123, 1.5, none |
|
[1, 2] / {'k': v} |
#[1, 2] / Std.HashMap.ofList [("k", v)] |
repeated fields are Array |
this / this.f / l[i] |
binder / x.f / guarded (l[i]?).get h |
binder auto-renamed on collision; indexing guarded by isSome (CEL errors out of range) |
size(e), e.size() |
e.size |
String.size/List.size shims |
contains |
Cel.contains e a |
typeclass: substring vs element vs subslice |
startsWith endsWith |
.startsWith .endsWith |
+ ByteArray shims |
matches |
Cel.regexMatch e r |
pure-Lean RE2-subset engine (Cel.Regex); literal patterns gated at generation time + #guarded at Lean compile time |
has(e.f) |
e.f.isSome |
explicit-presence fields |
l.all(v, p) / l.exists(v, p) |
∀ v ∈ l, p / ∃ v ∈ l, p |
|
l.exists_one(v, p) |
l.countP (fun v => p) = 1 |
counts duplicates like CEL (not ∃!) |
l.map(v, e) / l.filter(v, p) |
l.map (fun v => e) / l.filter (fun v => p) |
Props decide-wrapped in Bool positions |
unique() |
l.Nodup |
Array.Nodup shim, decidable |
isNan isInf |
.isNaN .isInf/.isInfSign |
|
isEmail isHostname isUri isUriRef isIp isIpPrefix isHostAndPort |
.isEmail ... |
real grammars in Cel.Format (HTML-standard email, RFC 1034, RFC 4291 + RFC 4007 zones, RFC 3986 + RFC 6874, CIDR) |
int() uint() double() bool() |
Cel.toInt ... |
typeclass conversions |
string(e) / bytes(e) / dyn(e) |
toString e / e.toUTF8 / e |
|
timestamp() duration() now |
folded Cel.Timestamp.mk s n constants / Cel.now |
literals fold at codegen; runtime RFC 3339 + duration parsers; ts - ts, ts ± dur, duration accessors |
Emitted propositions are Decidable end to end (asserted by the test
corpus), so runtime validation can construct subtype values by deciding the
refinement after decoding.
An erroring CEL rule is not satisfied; the translation encodes exactly that (error ⇒ the proposition is false), so the kernel can never prove a proposition whose CEL evaluation would have errored:
- Integer division/modulo agree exactly with CEL. grpc-lean maps proto
ints to fixed-width
Int32/Int64etc., whose Lean//%truncate toward zero exactly like CEL's Go semantics. (UnboundedIntwould not have matched — it uses Euclidean conventions.) - Overflow: CEL integer arithmetic is 64-bit and errors on overflow.
Every translated
+/-/*/unary-on error-capable types carries aCel.addOk/subOk/mulOk/negOkside-condition computing the exact result in unboundedInt/Natand demanding the CEL domain, discharged asif h : Cel.mulOk x 2 then … else False— overflow falsifies the proposition. Where the plugin resolves the operand's proto type (see descriptor-aware numerics below) the arithmetic is lifted to CEL's own width first — anint32field'sthis * 100emitsx.toInt64 * 100underCel.mulOk x.toInt64 100— so the guard is exactly CEL's condition and an intermediate outside the 32-bit range is not an overflow, matching CEL. Division and modulo widen too, soint32Min / -1follows CEL semantics instead of wrapping asInt32. Comprehension binders and indexed elements are typed too (see below), so this reaches insideall/existsbodies; only values with no proto type at all (free identifiers) fall back to the operand's own Lean type, whereInt32/UInt32operands demand the 32-bit range — conservatively stricter than CEL, sound but not exact.Natsize arithmetic guards truncating subtraction. Constant arithmetic folds and range-checks at generation time (an always-erroring rule is rejected); concatenation and timestamp/duration arithmetic are total and unguarded. - Descriptor-aware numerics: Lean silently wraps an out-of-range literal
(
(3000000000 : Int32)elaborates to-1294967296), so the plugin range-checks every numeric literal against the domain of the proto value it meets and rejects the rule at generation time with a source-located error naming the field's Lean type and bounds. This coversthisin a scalar field rule,this-rooted select paths in message rules (including map-selection leaves and enum ints, which CEL compares asint32), and custom CEL onrepeated.items/map.keys/map.valueselements; it also rejects negative literals against unsigned fields. Literals meeting a widened arithmetic result are checked at 64-bit width instead, sothis * 10000 <= 5000000000on anint32field is accepted — the product isInt64. The same typing makes mixed-width comparisons elaborate (this.small < this.bigacrossint32/int64widens the narrower side). - Container elements are typed: a path may descend into a container, so
the value a comprehension binds (
this.all(v, …)— a repeated field's elements, a map's keys) and the value indexing yields (this.items[0]— an element, a map's value) carry the same descriptor typing a named field does. Paths compose through binders, sothis.all(o, o.lots.all(n, n > 3e9))is range-checked at the innermost element's type. Only free identifiers lack descriptor typing. - Indexing: CEL errors on an out-of-range index or missing map key, so
l[i]/m['k']translate to(l[i]?).get hunder a pending(l[i]?).isSomeguard rather than a panickingl[i]!. Guards discharge at exactly the boundaries CEL's error absorption allows (&&/||operands, ternary branches,all/existsbodies — in positive polarity); in error-strict or negated contexts they float outward and close at the root, so!x[0] > 5on an empty array is false, matching CEL's error. The few shapes with no sound discharge point (index proofs insideexists_one/map/filterbodies or negated quantifier bodies) are source-located generation errors. - Enums work as ints: the plugin resolves every
thispath against the descriptors, so enum-typed values (field rules, message-rule leaves, repeated/map element rules) render through the generated enum's.toInt32view, matching CEL's enum-as-int semantics — including enum values a comprehension binds or indexing reaches (this.tier_history[0] >= 1), under the index guard. Standard enum rules keep their constructor-test form. - Map selection sugar works:
this.labels.prioritybecomes guarded key indexing (missing key ⇒ false, CEL's error), andhas(this.labels.priority)becomes key presence — at any path depth. - Warnings (stderr / JSON) flag translations with caveats:
has()presence semantics beyondOptionfields, evaluation-time-dependentnow, placeholder timestamp/duration types, and free identifiers kept verbatim. floatfields are compared at double precision, as CEL does. A protofloatis LeanFloat32, but CEL converts it to a double before any numeric operation, so a typedfloatvalue is emitted asx.toFloatwherever it meets a literal, another numeric field, or arithmetic (this <= 1.1on afloatfield becomesx.toFloat ≤ 1.1). Widening Float32 → Float is exact, so this only removes a rounding: without it the literal would elaborate atFloat32andthis == 1.1would be satisfiable in Lean (by theFloat32nearest 1.1) while being unsatisfiable in CEL — the kernel proving a proposition whose CEL evaluation is false. Standardfloat.*rules are left narrow deliberately: their bounds come from the descriptor and are already exactlyFloat32values, so the comparison cannot round either way. Afloat/doublefield comparison elaborates by widening thefloatside.string(e)usestoString, whose formatting may diverge from CEL's for floats.
- Field rules apply to
this= the field value: whole array/map for repeated/map fields, the unwrapped value for explicit-presence fields (checked only when present, matching protovalidate). - Message rules may reference sibling fields (
this.f) with CEL's value semantics, to any depth: heads go through generated per-field view functions ((Order.featuredD featured)— base-typed, default instance when unset), intermediate hops select the base codegen's<field>Dgetters (every non-final hop of a CEL path is message-typed), and final selects stay plain.this.featured.dims.weight_grams <= 50000lands as(Order.featuredD featured).dimsD.weight_grams ≤ 50000;this.trackingover an optional scalar reads as(tracking.getD ""). Refined plain fields becomef.val, so the Prop states exactly what the structure carries.has(this.f)follows CEL's proto presence semantics per field:isSomefor explicit presence, non-default for implicit scalars (!f.isEmpty,f != 0, non-zero enum case), non-empty for repeated/map, a case test (o.any (· matches .member _), oro matches .member _for required oneofs) for oneof members, andtrueforrequiredfields — so a message rule likehas(this.refund_iban) ? has(this.iban) : truebecomes an implication between case tests across oneofs. Value access to oneof members follows CEL's value-or-default semantics through generated<member>_or_defaultaccessors (returning the base payload, the member type's zero value when the case is inactive):this.iban.startsWith('DE')on a required oneof lands asmethod.iban_or_default.startsWith "DE". Paths through nested oneof members work too: base messages carry generated member getters (m.iban/m.ibanD, value-or-default), sothis.payment.store_credit_product.sku != 'banned'lands as(Order.paymentD payment).store_credit_productD.sku ≠ "banned". Nestedhas()delegates to generatedhas_<field>/has_<member>predicates encoding CEL's per-category semantics (case tests,isSome, non-empty, non-default), sohas(this.payment.cash_note) ? … : truelands as(Order.paymentD payment).has_cash_note → …. - Oneofs with constrained members generate a validated sum type whose
constructors carry the refined payloads:
| iban : { x : String // x.size ≥ 15 ∧ x.size ≤ 34 } → Payment.method_Type, with case-wisevalidate/toBase; unannotated message members embed the member type's validated variant.(buf.validate.oneof).requiredunwraps theOption(validate reportsrequiredon an unset oneof).requiredon an individual member is ignored with a note (the active case already implies presence). Unconstrained oneofs use the base sum. Every real oneof additionally gets per-member<member>_or_defaultaccessors (on the validated sum, or extending the base sum's namespace for unconstrained oneofs), shaped to the field: bare-sum dot form for required oneofs,Option-taking otherwise. - A message whose validated variant would be recursive (through validated message fields) is rejected at generation time — Lean structures cannot be recursive.
- Cross-target validated references work through
valid_deps(paired withproto_depson the baselean_proto_library): message fields typed by another target's protos use that target's validated variants, and violations surface through the embed. Map values always keep base types. - Singular message fields with their own CEL rules refine the base message type (their CEL inspects base fields); element-wise validation composes automatically only for repeated-message fields.
google.protobuf.Timestamp/Durationfields are supported end to end: grpc-lean maps their imports to the sharedProtobuf.WellKnownmodule, andCel.Timestamp/Cel.Durationare those generated types (comparisons, arithmetic, and==ignore unknown fields; propositional=on them is not decidable — use==/ordering in rules). Standard rulestimestamp.const/lt/lte/gt/gteandduration.const/lt/lte/gt/gte/in/not_inlower like the numeric families;lt_now/gt_now/withinare rejected becauseCel.nowis opaque; evaluation-time constraints are outside the refinement.- The regex engine rejects unsupported RE2 constructs (
(?i)flags,\b, named classes) at parse time — such patterns would match nothing at runtime, which would be unsound under negation. See regex assurance for the layers that prevent it and their exact scope.validatereports the first violation (fieldPath, ruleId, message);toBasedrops unknown fields. - A field's rules are decided as one dependent-
ifchain drawing hypothesis names from a fixed 64-name pool; a field carrying more than 64 rules is a generation-time error rather than Lean with a captured binder.
string.pattern and CEL matches are decided at runtime by
Protovalidate.Cel.Regex, a pure-Lean Pike VM. It is trusted code: no
theorem relates it to a mathematical semantics of regular expressions. What
protects it is a generation-time gate plus two compile-time batteries, and it
is worth being exact about their reach.
Enforced. For every literal pattern:
celtolean.RegexAccepted(a hand-written Go port of the Lean engine's grammar) must accept it, and Go's RE2 (regexp.Compile) must accept it — a source-located generation error otherwise. The port is deliberately stricter where RE2 and the Lean parser would silently diverge (POSIX named classes, octal escapes,[]-a]).- The generated Lean file carries
#guard Cel.Regex.accepts "<pattern>". Lean evaluates it while elaborating the module, so if the Go gate accepted a pattern the Lean parser rejects, the build fails. This is what rules out the match-nothing fallback (and with it the unsoundness of!this.matches(p)degenerating totrue). - A differential battery per pattern: probe strings walked out of the
pattern's own RE2 syntax tree (the first and the last alternative, the
minimum and one extra repetition, class representatives) plus
perturbations of those, each labelled
by Go's RE2 engine — the reference
matchessemantics — and emitted as#guard Cel.regexMatch "<probe>" "<pattern>"/#guard !(Cel.regexMatch …). Lean re-decides every one at elaboration time, so a language-level disagreement between RE2 and the Lean engine on an emitted probe fails the build. Up to 10 probes per pattern; the same batteries are emitted for theTest/cel_corpus.tsvrows.
Not enforced — do not read more into the guards than this.
- The batteries are differential testing, not equivalence: agreement on
finitely many probe strings does not prove
Cel.Regexand RE2 accept the same language for that pattern, let alone in general. There is no equivalence theorem and none is claimed; the Go side is not formalized. #guardis kernel evaluation of a closed Bool, not a proof carried in any theorem's axiom closure. A generatedvalidate_sounddoes not depend on it.- Dynamic patterns (
this.matches(this.other)) are rejected by generation: the generator cannot establish that their runtime value belongs to the supported subset, and accepting them would be unsound under negation. - Hand-written callers with an untrusted pattern should use
Cel.regexMatchChecked, which reports parser failure separately;Cel.regexMatchretains its generated-predicateBoolinterface and maps an invalid pattern tofalseonly after the generator has established the literal-pattern precondition. - The Go gate being stricter than the Lean parser is unchecked (and harmless: it only rejects patterns the engine could have handled). Only the other direction — Go accepts, Lean rejects — is unsound, and that is exactly what guard 2 catches.
//Conformance:supported_slice_test runs the official protovalidate v1.2.2
Go harness, fetched hermetically through Bazel, against a pure-Lean executor.
The executor decodes the harness envelope, dispatches a closed registry of
Any.type_url values to generated Lean decoders/validators, and emits the
official violation structure. The executable supported slice consists of all
six cases in standard_rules/bool, compared with both
--strict_error and --strict_message; the expected-failure file is empty.
The upstream bool vectors exercise exact result classification, violation
count, field path, rule path, and rule ID. They do not supply expected
violation-message text or compilation/runtime-error cases, so the two strict
flags do not test those dimensions. This is a supported-slice gate, not a
claim of whole-suite conformance. Because the
upstream runner exits 0 when a suite filter matches nothing, the gate also
asserts the harness summary reports exactly six executed, passing cases; a
corpus bump that renames or empties the suite fails instead of passing
vacuously.
//Conformance:registry_freshness_test separately classifies every one of the
32 .proto files in the pinned official case corpus. It fails if upstream's
inventory changes, if the compiling source target is silently narrowed, or if
a case lacks a status/reason. The compiling set consists of bool.proto and
the generation-only filename-with-dash.proto.
Conformance/corpus_registry.tsv records the blockers for every other file:
single-constructor/oneof pattern generation, missing scalar wrapper WKTs,
protobuf Int32 versus CEL Int coercions, and fixed-width/float numeric
literal and decidability gaps.
Protovalidate.Cel.Format implements protovalidate's well-known string
formats (email, hostname, address, ip, ipv4, ipv6, ip_prefix, ip_with_prefixlen, host_and_port, uri, uri_ref) as hand-written grammars.
Like the regex engine it is trusted code: no theorem relates it to the RFCs.
Enforced. //Test:format_corpus_test compiles Test/format_corpus.tsv
into a module of #guards and Lean re-decides every one at elaboration time.
The corpus is upstream protovalidate's own conformance suite — the
cases_is_*.go / cases_strings.go expectations in the protovalidate bazel
module this build already depends on, which are normative for every
protovalidate runtime — extracted by cmd/formatcorpus -extract (a dev-time
step; the build only reads the TSV). 731 labelled inputs, covering the RFC
edge cases the suite exercises: trailing dots, all-digit labels, 63/64-char
labels, IPv6 zone identifiers, bracketed and zone-bearing hosts, IPvFuture,
percent-encoded UTF-8 in hosts, port ranges, leading zeros, uppercase hex,
IDN labels, empty parts. Each row is checked through the expression codegen
emits (x.isIpPrefix 4 true, the isHostname || isIp disjunction
string.address lowers to, the Cel.regexMatch calls string.uuid /
string.tuuid lower to), so a wrapper-level slip is caught too.
Not enforced — do not read more into the battery than this.
- It is differential testing against a finite corpus, not a proof that
Cel.Formatimplements the RFCs, and not a proof that it agrees with protovalidate on inputs the corpus does not contain. There is no equivalence theorem and none is claimed. #guardis kernel evaluation of a closedBool, not a proof carried in any theorem's axiom closure — a generatedvalidate_sounddoes not depend on it.- The corpus is a vendored snapshot of one
-extractrun; the build does not detect additions to the upstream corpus. - Formats protovalidate has that this repo does not lower at all (
ulid,well_known_regex) are rejected at generation time, so they have no corpus rows.
Proved, not sampled. Cel.Format's recognizer laws
(lean/Protovalidate/Cel/FormatLaws.lean, audited by
//lean:format_laws_assurance) carry the complementary kind of evidence —
properties holding for every input rather than for corpus points. The
load-bearing one is ipv4Value?_lt: the value an accepted IPv4 address denotes
really is below 2 ^ 32, which is exactly what isIpPrefix's host-bit masking
(v % 2 ^ (32 - len)) assumes. The others pin documented shape claims
(ipv4Value?_length, isHostname_length, isPort_le, isIp_version — only
versions 0/4/6 are ever accepted). The URI, email and IPv6 grammars have no
corresponding recognizer laws; their evidence is the corpus alone.
celtolean/ Go library: CEL AST → Lean expression (golden tests)
cmd/cel2lean/ CLI: single expression, -json, -batch corpus→Lean
cmd/protoc-gen-lean-protovalidate/ protoc plugin entry point
leanvalidate/ plugin codegen: descriptors + buf.validate rules → Lean
lean/Protovalidate/ Lean runtime (modules Protovalidate.Cel, .Runtime) plus
Cel/FormatLaws.lean (recognizer theorems, not runtime)
protovalidate/ Starlark API: lean_protovalidate_library (defs.bzl)
Test/ corpus TSV + genrule compiling every translation with
Decidable assertions (build_test) + stand-in msg types.
A row's flags may carry `dom=<proto kind>` to supply the
descriptor typing the plugin would derive.
format_corpus.tsv holds protovalidate's conformance
expectations for the string formats (see format
assurance above).
Conformance/ pinned official v1.2.2 supported-slice executor, corpus
classification, expected-failure and freshness gates
cmd/formatcorpus/ extracts that corpus from an upstream protovalidate
checkout, and renders it as the Lean #guard battery
examples/shipping/ end-to-end pipeline example + runtime lean_test
(Lean sources sit under lean/ — with the prefix stripped from module names —
because Protovalidate/ would collide with the protovalidate/ Starlark
package on case-insensitive filesystems.)
bazel test //... # Go golden tests, corpus compile test, e2e runtime test
- The executable official-conformance claim covers
standard_rules/bool. The complete corpus is freshness-classified, and each case outside the supported slice has a generator or runtime blocker recorded inConformance/corpus_registry.tsv. - Predefined user-extension rules, the
(?i)and\bregex constructs, and evaluation-time-dependent rules such astimestamp.lt_nowandwithinare unsupported.Cel.nowhas no validation-time interpretation in a refinement. - Some translations are sound but conservative or rejected at generation
time:
- free identifiers carry no descriptor type, so their literals are unchecked and their 32-bit guards use the operand's Lean domain;
- error absorption under negation (
!(err && false)) translates to failure; - index proofs inside
exists_one/map/filterbodies and negated quantifier bodies have no sound discharge point and are rejected.
- No formal relation connects
Protovalidate.Cel.Regexto RE2 beyond the per-pattern differential batteries (see regex assurance). No formal relation connectsCel.Format's grammars to their RFCs beyond the conformance battery and the IPv4/hostname/port recognizer laws (see format assurance). Both corpora are vendored snapshots of upstream references; the URI, email, and IPv6 grammars have no corresponding recognizer laws.