From d98c5c8ac6b86ac752b30fdb59b6c9ded2d190f0 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Wed, 19 Aug 2026 14:56:08 -0400 Subject: [PATCH 1/7] Funder expansion page: policybench.org/expand Packages for regional, population, and program slices with published price anchors, the compute-at-cost disclosure, sample sizing by confidence width, and the coverage-never-scores integrity rules. Mirrors the Robin Hood one-pager's structure; hero stats recomputed on the 30-model board (owed-SNAP: 42% answered $0, 1/600 exact). Footer link + sitemap entry. Co-Authored-By: Claude Fable 5 --- app/src/App.tsx | 6 ++ app/src/app/expand/page.tsx | 152 ++++++++++++++++++++++++++++++++++++ app/src/app/sitemap.ts | 5 ++ 3 files changed, 163 insertions(+) create mode 100644 app/src/app/expand/page.tsx diff --git a/app/src/App.tsx b/app/src/App.tsx index cce42b0..b912c87 100644 --- a/app/src/App.tsx +++ b/app/src/App.tsx @@ -386,6 +386,12 @@ export default function App() { PolicyBench.org {" "} ·{" "} + + Expand + +
+ {title} +
+
+ {price} +
+
+ {children} +
+ + ); +} + +export default function ExpandPage() { + return ( +
+ +
+
PolicyBench · for funders
+

+ Expand PolicyBench to your region, population, or program +

+ +

+ Families already ask AI about the questions that decide their month: + Do I qualify for SNAP? How much is my credit? Will this job cost me + Medicaid? The{" "} + + public board + {" "} + tests 30 frontier models on 100 real households. + The best model computes 88.7% of amounts within $1. On SNAP cases + where the family is owed benefits, models answer exactly $0 in 42% + of cells, and no model gets more than 1 case in 20 right. A family + told “$0” does not apply. Those are national numbers — + nobody measures this for your region. +

+ +

+ A PolicyBench slice measures it. We draw households from certified + survey microdata weighted to your population, cover your programs, + benchmark the models your people actually use, and compute every + reference from the law with PolicyEngine. Every miss gets a + diagnosed failure mode. You get a public slice leaderboard, a + written analysis of where models fail your population, and a + briefing. +

+ +
+ + One program family — SNAP, Medicaid, child care, tax credits — + across all 30 board models. Per-model accuracy, diagnosed failure + modes, written analysis, and a briefing. Fast: the board already + holds the raw material. + + + New households weighted to your area and program mix. A published + slice leaderboard beside the national board, the full audit, the + analysis, and a briefing for your team or grantees. + + + Your slice stays current. New models fold in as they ship, + quarterly refresh and re-analysis, and technical assistance to + grantees building AI tools for your population. + +
+ +

+ Model inference is the small part. A 100-household sweep of all 30 + models costs about $200 in tokens; adding one new model this week + cost $1.44. The price buys scenario curation, certified references, + a diagnosed audit of every miss, publication, and the analysis. We + size samples to the confidence width your question needs, from your + population — not a fixed household count. +

+ +
+
+ What funding buys — and what it never buys +
+

+ Funders buy coverage: households, programs, regions, refresh + cadence. Funding never buys scores, rankings, or placement. We + take no money from model vendors. Every slice stays public — + prompts, references, predictions, and diagnoses. +

+
+ +

+ Write to Max Ghenis at{" "} + + max@policyengine.org + + . We also brief funder networks — one slice presented to a convened + room goes further than a dozen pitches, and we are glad to present + at yours. +

+
+
+ ); +} diff --git a/app/src/app/sitemap.ts b/app/src/app/sitemap.ts index e934726..816f7cc 100644 --- a/app/src/app/sitemap.ts +++ b/app/src/app/sitemap.ts @@ -21,6 +21,11 @@ export default function sitemap(): MetadataRoute.Sitemap { changeFrequency: "monthly", priority: 0.7, }, + { + url: "https://policybench.org/expand", + changeFrequency: "monthly", + priority: 0.5, + }, ...modelEntries, ]; } From d2f5a95a97a7bbe01ec00ff9dd7ccc8e0e5cf899 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Wed, 19 Aug 2026 14:56:51 -0400 Subject: [PATCH 2/7] Eyebrow names the work, not the audience Co-Authored-By: Claude Fable 5 --- app/src/app/expand/page.tsx | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/app/src/app/expand/page.tsx b/app/src/app/expand/page.tsx index 86a82f8..108433a 100644 --- a/app/src/app/expand/page.tsx +++ b/app/src/app/expand/page.tsx @@ -64,7 +64,7 @@ export default function ExpandPage() { }} />
-
PolicyBench · for funders
+
PolicyBench · expansion

Expand PolicyBench to your region, population, or program

From 14de91a80478c6beb101638c74e47c69d4b9bb79 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Wed, 19 Aug 2026 15:23:34 -0400 Subject: [PATCH 3/7] Expand page edits: cut the inference-cost paragraph, contact CTA, vendor-funding clarity - The compute-at-cost paragraph invited arithmetic nobody asked for; the sizing sentence stays. - Contact button mails contact@policybench.org (Cloudflare route created, forwards like max@policybench.org). - Integrity block: no model vendor pays for evaluation; vendor funding for separate projects never touches the benchmark. Co-Authored-By: Claude Fable 5 --- app/src/app/expand/page.tsx | 29 +++++++++++++++-------------- 1 file changed, 15 insertions(+), 14 deletions(-) diff --git a/app/src/app/expand/page.tsx b/app/src/app/expand/page.tsx index 108433a..3d4df8c 100644 --- a/app/src/app/expand/page.tsx +++ b/app/src/app/expand/page.tsx @@ -114,12 +114,8 @@ export default function ExpandPage() {

- Model inference is the small part. A 100-household sweep of all 30 - models costs about $200 in tokens; adding one new model this week - cost $1.44. The price buys scenario curation, certified references, - a diagnosed audit of every miss, publication, and the analysis. We - size samples to the confidence width your question needs, from your - population — not a fixed household count. + We size samples to the confidence width your question needs, from + your population — not a fixed household count.

@@ -128,21 +124,26 @@ export default function ExpandPage() {

Funders buy coverage: households, programs, regions, refresh - cadence. Funding never buys scores, rankings, or placement. We - take no money from model vendors. Every slice stays public — + cadence. Funding never buys scores, rankings, or placement. No + model vendor pays for evaluation — vendor funding for separate + projects never touches the benchmark. Every slice stays public — prompts, references, predictions, and diagnoses.

-

- Write to Max Ghenis at{" "} +

- max@policyengine.org + Contact us - . We also brief funder networks — one slice presented to a convened + + contact@policybench.org + +
+

+ We also brief funder networks — one slice presented to a convened room goes further than a dozen pitches, and we are glad to present at yours.

From 6a6b59b2d2ea7923db2034e256c1d26b01fdbd9a Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Wed, 19 Aug 2026 15:25:03 -0400 Subject: [PATCH 4/7] Expand page: answer the funder-simulation findings Medicaid stat joins the SNAP evidence (median model misclassifies 1 in 15), reference-validation sentence with the paper link, consumer-tool claim softened, scaled-engagement and grantee-tool lines, 501(c)(3) grantability line, jargon trimmed. Co-Authored-By: Claude Fable 5 --- app/src/app/expand/page.tsx | 28 +++++++++++++++++++++------- 1 file changed, 21 insertions(+), 7 deletions(-) diff --git a/app/src/app/expand/page.tsx b/app/src/app/expand/page.tsx index 3d4df8c..65859ba 100644 --- a/app/src/app/expand/page.tsx +++ b/app/src/app/expand/page.tsx @@ -78,20 +78,29 @@ export default function ExpandPage() { {" "} tests 30 frontier models on 100 real households. The best model computes 88.7% of amounts within $1. On SNAP cases - where the family is owed benefits, models answer exactly $0 in 42% - of cells, and no model gets more than 1 case in 20 right. A family + where the family is owed benefits, models answer exactly $0 in 42% of + answers, and no model gets more than 1 case in 20 right. + On Medicaid eligibility, the median model misclassifies 1 person in + 15; the weakest, nearly 1 in 3. A family told “$0” does not apply. Those are national numbers — nobody measures this for your region.

- A PolicyBench slice measures it. We draw households from certified + A PolicyBench slice measures it. We draw households from survey microdata weighted to your population, cover your programs, - benchmark the models your people actually use, and compute every + benchmark the models behind the tools your people use, and compute every reference from the law with PolicyEngine. Every miss gets a diagnosed failure mode. You get a public slice leaderboard, a written analysis of where models fail your population, and a - briefing. + briefing. The answer key checks itself in public: references come + from open-source code, cross-checked against other calculators + where they exist, and challenged values get adjudicated against + the statute — the{" "} + + methodology and adjudication record + {" "} + are published.

@@ -115,7 +124,10 @@ export default function ExpandPage() {

We size samples to the confidence width your question needs, from - your population — not a fixed household count. + your population — not a fixed household count. Multi-state + observatories and portfolio-wide coverage are scoped directly. + Grantee tools that expose an API can run the same households as the + board — ask us.

@@ -143,7 +155,9 @@ export default function ExpandPage() {

- We also brief funder networks — one slice presented to a convened + PolicyBench is a project of PolicyEngine, a 501(c)(3) nonprofit — + engagements work as grants or contracts. We also brief funder + networks — one slice presented to a convened room goes further than a dozen pitches, and we are glad to present at yours.

From faaab79fa85aa29d3443d22bcf92aee6114bb437 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Wed, 19 Aug 2026 15:34:04 -0400 Subject: [PATCH 5/7] Fourth package card: national or portfolio scope, unpriced The scaled tier moves from a buried sentence to a card, SaaS-style minus the word enterprise. Grantee-tool evaluation lives there too. Co-Authored-By: Claude Fable 5 --- app/src/app/expand/page.tsx | 12 +++++++----- 1 file changed, 7 insertions(+), 5 deletions(-) diff --git a/app/src/app/expand/page.tsx b/app/src/app/expand/page.tsx index 65859ba..29bf9cb 100644 --- a/app/src/app/expand/page.tsx +++ b/app/src/app/expand/page.tsx @@ -103,7 +103,7 @@ export default function ExpandPage() { are published.

-
+
One program family — SNAP, Medicaid, child care, tax credits — across all 30 board models. Per-model accuracy, diagnosed failure @@ -120,14 +120,16 @@ export default function ExpandPage() { quarterly refresh and re-analysis, and technical assistance to grantees building AI tools for your population. + + A 50-state observatory, or coverage across a whole grantee + portfolio — including your grantees’ own tools, run through + the same households as the board. We scope these directly. +

We size samples to the confidence width your question needs, from - your population — not a fixed household count. Multi-state - observatories and portfolio-wide coverage are scoped directly. - Grantee tools that expose an API can run the same households as the - board — ask us. + your population — not a fixed household count.

From fe03dc7ef8f5044cd50760c60a2fc28ba16eb2b8 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Wed, 19 Aug 2026 15:34:51 -0400 Subject: [PATCH 6/7] Three-card ladder: the unpriced tier replaces standing coverage Standing coverage becomes an add-on any slice can attach, priced with the slice. Co-Authored-By: Claude Fable 5 --- app/src/app/expand/page.tsx | 12 +++++------- 1 file changed, 5 insertions(+), 7 deletions(-) diff --git a/app/src/app/expand/page.tsx b/app/src/app/expand/page.tsx index 29bf9cb..cdbdbfa 100644 --- a/app/src/app/expand/page.tsx +++ b/app/src/app/expand/page.tsx @@ -103,7 +103,7 @@ export default function ExpandPage() { are published.

-
+
One program family — SNAP, Medicaid, child care, tax credits — across all 30 board models. Per-model accuracy, diagnosed failure @@ -115,11 +115,6 @@ export default function ExpandPage() { slice leaderboard beside the national board, the full audit, the analysis, and a briefing for your team or grantees. - - Your slice stays current. New models fold in as they ship, - quarterly refresh and re-analysis, and technical assistance to - grantees building AI tools for your population. - A 50-state observatory, or coverage across a whole grantee portfolio — including your grantees’ own tools, run through @@ -129,7 +124,10 @@ export default function ExpandPage() {

We size samples to the confidence width your question needs, from - your population — not a fixed household count. + your population — not a fixed household count. Any slice can add + standing coverage: quarterly refresh and re-analysis, new models as + they ship, and technical assistance to grantees building AI tools — + priced with the slice.

From 0feff9674147b9b98f5037d942755550ccc68e3c Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Wed, 19 Aug 2026 15:47:03 -0400 Subject: [PATCH 7/7] Contact button: white text via the background token The site's unlayered 'a { color: inherit }' beats every layered Tailwind utility, so the inline token style carries the color. Co-Authored-By: Claude Fable 5 --- app/src/app/expand/page.tsx | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/app/src/app/expand/page.tsx b/app/src/app/expand/page.tsx index cdbdbfa..b94d5bf 100644 --- a/app/src/app/expand/page.tsx +++ b/app/src/app/expand/page.tsx @@ -146,7 +146,8 @@ export default function ExpandPage() {