Skip to content

Backfill: re-run PCA publish for both SOTUs with structured provenance#40

Merged
aRealGem merged 1 commit into
mainfrom
claude/pca-provenance-backfill
Jul 19, 2026
Merged

Backfill: re-run PCA publish for both SOTUs with structured provenance#40
aRealGem merged 1 commit into
mainfrom
claude/pca-provenance-backfill

Conversation

@aRealGem

Copy link
Copy Markdown
Owner

Regenerates site-pca via a fresh live truthbot publish --engine pca over Trump 2026 + Biden 2022, so the published reports carry the VerdictProvenance + reconciled-judge render shipped in PR-E (6850da1). The 7/16 site predated persistence and was provenance-less; this backfills it and generates real vote data.

What changed

  • site-pca/ regenerated — both reports now render the provenance strip (Layer A → PCA panel: Misleading ×2 → CRM-114: MISLEADING→FALSE), and genuinely-unconverged split claims show No single verdict instead of rendering blank.
  • metrics/pca_runs/<run_id>.json ×2 — replay artifacts holding raw per-seat panel votes for all 276 claims. This is the vote data that was previously unrecoverable, and it unblocks the strength_from_votes recalibration.

Results (dev roster, open-book Brave+FactCheck, CRM-114)

Trump 2026 Biden 2022
Check-worthy 166 / 722 110 / 480
True / Misleading / False / Unverifiable 66 / 53 / 18 / 23 90 / 7 / 0 / 11
Non-converged (of escalated) 6 (49) 2 (14)

The 49 Trump panel-splits are now surfaced (were collapsed to ~5–7 before) — the core provenance win.

Notes / follow-ups (not in scope here)

  • Cost telemetry undercount: pipeline self-reported $0.163 + $0.062 but real proxy spend was ~$1.83 (~7–8×). Worth a fix to the cost_usd fold so operators can budget accurately.
  • Biden severity-softening: still skews 82% True — known open-book thread, unchanged.

Data-only change; suite 944 green.

🤖 Generated with Claude Code

Regenerates site-pca via a fresh live `truthbot publish --engine pca` over
Trump 2026 + Biden 2022, so the published reports carry the VerdictProvenance
+ reconciled-judge render shipped in PR-E (6850da1) — the 7/16 site predated
persistence and was provenance-less.

Adds the two replay artifacts (metrics/pca_runs/<run_id>.json) holding the raw
per-seat panel votes for all 276 claims — the vote data that was previously
unrecoverable, which unblocks the strength_from_votes recalibration.

Results (dev roster, open-book Brave+FactCheck, CRM-114):
- Trump 2026: 166/722 check-worthy — True66/Misleading53/False18/Unverifiable23,
  6 non-converged; 49 panel-split claims now surfaced (were collapsed before).
- Biden 2022: 110/480 check-worthy — True90/Misleading7/Unverifiable11,
  2 non-converged. (Biden still skews True — known severity-softening, separate.)

Note: pipeline self-reported cost ($0.163 + $0.062) is a ~7-8x undercount of
real proxy spend (~$1.83 total this run). Cost-telemetry fix is a follow-up.

Data-only change; suite 944 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@cursor

cursor Bot commented Jul 17, 2026

Copy link
Copy Markdown

Bugbot is not enabled for your account, so this pull request was not reviewed.

Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs.

aRealGem pushed a commit that referenced this pull request Jul 19, 2026
…ong to PR #40

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@aRealGem
aRealGem merged commit 43f3c2b into main Jul 19, 2026
@aRealGem
aRealGem deleted the claude/pca-provenance-backfill branch July 19, 2026 19:12
aRealGem added a commit that referenced this pull request Jul 19, 2026
… HOLD for jackie (#45)

* data(gold): +14 FALSE/MISLEADING boundary rows, 21 -> 35 (P67 Phase 2)

Targeted expansion to make the severity-softening fix measurable: FALSE
4 -> 11, MISLEADING 6 -> 11, TRUE 10 -> 11, UNVERIFIABLE 1 -> 2. Rows
selected on the FALSE/MISLEADING boundary (absolute quantifiers,
superlatives, exaggerated numbers) per the Phase-1 diagnosis; every
decidable row is anchored to published fact-checks / primary statistics
researched live (AP, FactCheck.org, PolitiFact, Poynter, WGBH, BEA, DIA
via FactCheck.org, Lead Stories, The Assembly NC). Claim text verbatim
from claim_set.train.jsonl; TRAIN-only (I6-safe); suite 944 green.

All new rows needs_review=true (annotator claude-expand-p67) — verdict
judgments held for jackie's adjudication per the no-auto-merge rule.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: drop metrics/pca_runs artifacts accidentally staged — they belong to PR #40

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Cass (Claude Code agent) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants