Harden review instructions against self-supplied evidence - #18
Conversation
Fable 5 adversarial review — round 1MUST-FIX 1: MUST-FIX 2: SHOULD-FIX 3: The blanket weakening-language regexes scan the whole generated task body, risking keyword-theater false positives outside the provenance guidance. SHOULD-FIX 4: Conditional requirements 2.2–2.4 and 3.2 read like generation-time boundary inference, while implementation correctly asks the reviewer to conditionally apply always-present guidance. SHOULD-FIX 5: The REQ 6.1 test asserts the basename of a temp directory that it created itself; remove this self-supplied assertion. NOTES: Magic lower-bound counts are loose; the 7.1 test intentionally couples to committed verdict state; hardcoded calibration file:line pointers can rot; text requirements remain instruction change detectors whose behavioral effect ultimately depends on reviewer compliance. The reviewer also found the issue #17 scope boundary, literal-lint rejection, and calibration 014 internally consistent. |
Sol 5.6 adversarial review — round 1MUST-FIX: No other actionable findings identified. |
Fable 5 adversarial review — round 2MUST-FIX: none. The accepted fixes hold. SHOULD-FIX 1: SHOULD-FIX 2: SHOULD-FIX 3: Calibration 012/013 production pointers use hardcoded line numbers that can silently drift; add stable test-title anchors or mechanical validation. NOTES: The session-context length cap is a tripwire; the lint test intentionally proves only the deferred syntax-lint boundary; non-vacuity floors and committed-verdict coupling fail loudly; narrowing calibration 013 away from filename loudness is intentional. Overall conclusion: the spec/test story is honest, direct judgments and audits are covered, and issue #17 cleanly owns mechanical set enumeration. |
Sol 5.6 adversarial review — round 2No actionable findings. The prior gaps are genuinely closed:
The reviewer independently verified |
Final dual-review triageAccepted and fixed
All accepted fixes were followed by the full suite (17 files, 168 tests), Rejected or deferred
Suggested placeholder issuesThese are recommendations for the operator to endorse, reject, or edit at the PR approval gate. Hermetic execution of version-pinned installed hooks. Preserve the end-to-end property that tests execute the exact Mechanically validate calibration provenance anchors. Calibration 012 and 013 currently cite accurate file:line ranges, but ordinary edits can move those lines without changing the prose. A follow-up should bind each provenance citation to a stable test-title or machine-readable anchor and fail loudly when the cited source disappears or no longer contains the claimed controls; empty or unmatched enumeration must itself fail. |
Summary
Specification
Published task artifact: panopticon://tasks/e28cde1041a24c99b3b26a5b72183247/artifacts/self-supplied-evidence-spec.md
Focused verification