-
Notifications
You must be signed in to change notification settings - Fork 186
pstack: strengthen blinded eval bakeoffs #158
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
1114f44
b0bb156
8d53b5c
98957c3
e330ed5
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -2,26 +2,31 @@ | |
|
|
||
| **You own the experiment design. Plan, blind, run, synthesize.** | ||
|
|
||
| Evals test how a change affects agent behavior before promoting it: a new skill variant, a structural change, a prompt tweak. The failure mode is the observer effect. An agent that knows it's being evaluated behaves differently, so candidates must run blind. | ||
| Evals are blinded, one-shot bakeoffs for deciding whether to promote or reject a change: a new skill variant, a structural change, a prompt tweak. Each trial gets one clean attempt with no feedback or repair. This is not a standing regression suite or a CI merge gate. | ||
|
|
||
| **Non-negotiables for blinding:** | ||
| **Non-negotiables for blinding and isolation:** | ||
|
|
||
| - No `eval`, `test`, `judge`, `experiment`, `rubric`, `score`, `compare`, `benchmark`, `candidate`, or `arena` in any directory, file, or prompt the candidate sees. | ||
| - The candidate prompt looks like an organic user request. State the goal, not the meta. "build me a small todo cli" not "show me how you follow the principles chain". | ||
| - Candidate prompts must not name a skill, give a skill path, or say to apply/use/follow a skill. Plant skills in the workspace; the candidate discovers them. | ||
| - No chain-eliciting cues. Don't ask the candidate to list which skills, principles, or files they applied; that meta-prompt inflates citation behavior. Ask for design notes generally and grade chain-following from code shape, not self-report. | ||
| - Sanitize directory and slug names. Use project-shaped names a user might pick, not labels like `candidate-1` or `agent-a`. | ||
| - Sanitize directory, slug, and arm names. Use project-shaped names a user might pick, not labels like `candidate-1`, `agent-a`, `control`, or `skill-off`. | ||
| - Don't tell the candidate other candidates exist. | ||
| - The judge can know it's judging but sees outputs by sanitized label only, never by model name. | ||
| - Comparing two variants: one judge scores both sets in a single pass on one scale, blind to which set each came from. Two judge runs with different prompts don't compare, the calibration drifts. | ||
| - Start each trial in a fresh workspace and preferably a new session. Clear prior chat and give it no sibling memory. Never plant prior transcripts, judge notes, or sibling outputs. | ||
|
|
||
| **Steps:** | ||
|
|
||
| 1. **Frame.** State what variant is under test and what behavior counts as success. Write the rubric (3-6 concrete criteria) for the judge only. Hold it back from candidates. | ||
| 2. **Set up sanitized environments.** Per-candidate working dir with the variant in place. Plant any context an organic task would have: a project skeleton, the skills the candidate would naturally read. | ||
| 3. **Author one organic prompt.** What a user would type. No leakage of what's being measured. | ||
| 4. **Spawn N parallel candidates** on different models per the **arena** skill's Phase B. Each works in its own sanitized dir; same prompt to each. | ||
| 5. **Spawn one blinded judge** on a different model family per the **arena** skill's Phase C. Judge sees outputs by sanitized label and the rubric, never a model name. | ||
| 6. **Verify the chain from transcripts, not self-report.** Read each candidate's local transcript under the active workspace's `agent-transcripts/` directory (the system prompt names this path). Do not glob across `~/.cursor/projects/*/`; that crosses workspace boundaries and reads private chats from unrelated projects. Look at which files each candidate actually opened. Citing a principle is not reading its leaf skill, and reading it is not applying it. Grade chain-following from the files it really read plus the shape of the code, never from the candidate's own claims. | ||
| 7. **Read every candidate output yourself** end to end. Compare to the judge's verdict. Disagreement means a model is biased or the rubric is ambiguous. Synthesize. | ||
| 1. **Frame.** State the variant and the promote-or-reject claim. Write a judge-only rubric with 3-6 concrete criteria. Grade task success and the intended behavioral shape. Never make a turn-1 skill load, a particular file read, a citation, or "did the skill trigger?" a pass condition. | ||
| 2. **Author an organic prompt set.** Include at least one task where the behavior should apply. If the variant changes a description, routing, sticky behavior, or when-to-apply rule, include at least one task where it should not engage and add false-positive cost to the rubric. Write what a user would type. Never name the behavioral tell the rubric grades, and never name or path a skill (if you measure dated headings, do not say "dated note" in the prompt). No other leakage of what is measured. If the task prompt itself is the target, write matched current and proposed versions here; otherwise every arm gets the same prompt. | ||
| 3. **Build comparison arms.** Before editing, snapshot any prior skill contents the control will need. Variant gets the proposed skill, structure, or prompt. Control gets the current version. For skill presence or content changes, run both a prior-version control and a skill-absent arm unless absence is impossible. Never plant the ablation-target skill into a skill-absent control. Hold the project skeleton, model mix, and every non-target input constant. Controls are other sanitized labels. Promote only when the variant beats the prior control on the rubric without looking worse than absent on false positives. | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. FP gate lacks negative promptsMedium Severity · Logic Bug Step 3’s promote rule requires the variant not look worse than the skill-absent arm on false positives for skill presence or content changes, but step 2 only authors negative organic prompts and false-positive rubric cost when the variant changes description, routing, sticky, or when-to-apply behavior. Pure content edits therefore hit an FP-vs-absent gate with no prompt set or rubric criteria that measure it. Additional Locations (1)Reviewed by Cursor Bugbot for commit e330ed5. Configure here. |
||
| 4. **Set up isolated trials.** Fresh per-trial workspace with only that arm's variant and organic-task context. Identical project skeleton across arms. A fresh workspace does not clear skills from workspace `.cursor/skills/`, user `~/.cursor/skills/`, or plugin installs: for a skill-absent arm, use a workspace-local isolation (or disable that is restored before any other arm runs). Never apply a shared user/plugin disable that also strips the skill from the variant arm. Preflight resolved sources and fail setup if the skill remains visible on a skill-absent arm or missing on a variant arm. For sticky, mode, description, or other always-on triggers, preflight that the variant reaches the candidate the way production does (reminder in context, description always loaded, and so on). If the harness cannot inject it that way, stop: the bakeoff is invalid for that variant class. Record each trial's workspace path and transcript ID as orchestrator-only metadata. Cheap deterministic preflights aid synthesis only; they never replace the blinded rubric. | ||
|
cursor[bot] marked this conversation as resolved.
|
||
| 5. **Run 2-3 one-shot trials per prompt and arm.** Launch each runner directly in its recorded workspace with that arm's isolated context. Fan out in parallel with no shared grounding and no candidate-visible files across workspaces. Match model and trial pairings across arms. If the skill ships across models, use at least two model families; matched pairings on one family are not enough. Ask only for the organic task output, not a graft rationale. Missing output fails the trial. No retries, coaching, or repair. When budget binds, prefer 2 trials on fewer models over 1 on many. | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Budget clashes with model ruleMedium Severity · Logic Bug Step 5 newly requires at least two model families when a skill ships across models, while the same step still says that under budget pressure agents should prefer two trials on fewer models. There is no tie-break, so a verbatim reader can collapse to one family with two matched trials, then promote on evidence that does not generalize. Reviewed by Cursor Bugbot for commit 98957c3. Configure here. |
||
| 6. **Spawn one blinded judge** on a different model family after every trial finishes. In one pass, score every output by randomized sanitized label against the same rubric. Mark each criterion and output pass or fail. Programmatic checks may filter obvious fails before the judge; they do not decide promote or reject. Do not run the arena pick/graft workflow. This bakeoff ends at arm-level scoring. | ||
| 7. **Inspect transcripts after scoring to explain how, not to decide pass or fail.** Read only the recorded transcript for each trial from that workspace's transcript directory (normally `~/.cursor/projects/<trial-workspace-slug>/agent-transcripts/`), using the session or transcript ID from setup. Derive the slug from the recorded workspace path. Do not glob across `~/.cursor/projects/*/` or open unregistered workspaces. Transcripts verify isolation and explain the output. They are not a pass gate. | ||
| 8. **Read the outputs yourself.** At small N, read every output end to end. At large N, read every fail plus a stated random sample of passes; silent skim is not enough. Report pass rates by arm and prompt, then compare with the judge. Promote only when the variant beats the control overall without adding false positives. Otherwise reject. Explain disagreements as judge bias, contamination, or rubric ambiguity. | ||
|
cursor[bot] marked this conversation as resolved.
|
||
|
|
||
| **Reply:** variant under test, rubric, per-candidate notes, judge's verdict, your synthesis, and a recommendation for whether to promote the variant. | ||
| **Related:** Shipped skills may keep a separate standing regression pack of 5-20 cases. It is distinct from this bakeoff. | ||
|
|
||
| **Reply:** variant and control, prompt set, rubric, trial pass rates, per-candidate notes, judge's verdict, your synthesis, and the promote-or-reject decision. | ||


There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Final promote bar contradicts ablation
High Severity · Logic Bug
Step 3 now requires evaluating against both a prior-version control and a skill-absent arm, with promotion gated on beating the prior control without looking worse than absent on false positives. However, Step 8 and the 'Reply' section still refer to a singular 'control,' which could lead to the full evaluation criteria from Step 3 being overlooked in the final synthesis.
Reviewed by Cursor Bugbot for commit e330ed5. Configure here.