Skip to content

PERCENTAGE_SPLIT: implement engine's 1-in-9999 recursion edge case in SQL #4

Description

@khvn26

The pure-SQL PERCENTAGE_SPLIT hash diverges from the engine on the ~1/9999 inputs where the bare hash mod 9999 == 9998 (the engine recurses with doubled input; this implementation skips).

Real-world impact at typical thresholds: ~0.005% false-negative rate on the count of identities matching a percentage-split segment. At 870M with threshold 50, that's ~22k false-negatives across the env — a rounding error on a count-badge UI but a measurable bias if anyone uses the count for billing or contract decisions.

What to ship

Implement the recursion as a CASE WHEN bare_hash_mod = 9998 THEN <recursive hash> ELSE <main hash> END wrapper in the inline SQL. The recursive hash uses doubled input: seg_key || ',' || value || ',' || seg_key || ',' || value. Cap at 2-3 iterations (engine recurses arbitrarily but in practice the second iteration almost always lands at non-9998).

Why deferred

Sub-0.005% bias on a UI count is below the threshold that customers care about. Defer until a customer reports a discrepancy or until the engine's bucketing semantics change.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions