feat(analytics): issue task IDs from the CLI - #316
Merged
Merged
Conversation
The CLI refused an agent's command until it passed --analytics-task-id and --analytics-intent, and answered on the stderr of the failed command with an instruction block: generate a UUID, run this `node -e` line, never use a placeholder. An agent meeting that cold cannot tell it from a prompt injection, and agents supplied placeholders anyway. The options are now --task-id and --user-intent. The old names still work. `prismic task-id` issues an ID, so the CLI decides the format and a value it did not issue is refused: a placeholder cannot be passed. The error names that command instead of explaining how to invent an ID. The word "analytics" left the option names because an agent reading it in a flag it types on every command treats the call as data collection and says so to the user. The MCP server's userIntent carries the same data under a name that does not, and no agent has ever raised it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Not published. The eval reads it with EVAL_SKILL_FILE so a CLI change and the skill change it needs can be tested together. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Two bugs in the task-grouping eval made a working candidate look broken.
The second judge interpolated a literal `${DATE_FIELD}` rather than the
request, so it graded intents against a placeholder. And requiring one
distinct intent string per request failed agents that reworded a correct
paraphrase between commands, while passing an agent that sent "test" on
every call.
Analytics groups by task ID, so the intent is a label for the group, not a
key. Judge each distinct value for being a real paraphrase of the request
and drop the equality check; the one-ID-per-request assertions stand. The
failure message now prints intents alongside IDs, so an intent failure no
longer reports IDs.
Also bring the gate's unit tests to the current design: IDs are minted by
`prismic task-id`, a UUID is no longer accepted, and the old option names
still work but are hidden from help.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
2 of 3 tasks
The alphabet has 32 characters, so `byte % 32` was already unbiased, but the expression reads as the biased pattern and breaks silently if the alphabet ever changes length. Mask the low five bits instead, which is exact by construction. Also drop two imports the new gate message left unused. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
The branch carried a copy of the Prismic skill so the eval could run against wording that is not published yet. A copy drifts. The eval already pins a skill commit and fetches it, so point that pin at the branch instead and drop both the copy and the environment override that read it. Fold the single-request test into the two-request one, whose first request already covers it, and drop the UUID the task ID pattern still allowed: the CLI issues every ID now, so a UUID can never be one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
The eval's own comment said a refused call loses no analytics and is not asserted, but every call reached the assertions: the wrapper logs argv before the CLI runs. An agent that passed a malformed ID, was refused, then used one good ID for the rest of the request failed a trial whose analytics were correct. Keep the diagnostic listing every call, refusals included, and assert on the ones that carried what the gate wants. Judge intents over the same set, which had the same fault. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
The test held six pairs of first/second variables, hid its two requests behind a standalone function, and explained itself in comments that were harder to read than the code. Collect the requests into one list and assert over it. Register the two skill states with describe.for, which lets the test body be inline and lets vitest infer its context. Name the helpers for what they select, so needsTaskId and accepted say what a comment used to. Rename mintTaskId to genTaskId. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
They were named constants when the judge referred to them by name. The judge now reads the prompt from the request it is judging, leaving one use. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
The condition mixed whether a command needs a task ID with whether it has one. Name each half, so the gate reads as the question it asks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Collecting the requests into a list needed a helper to build them and another to print them, for a test that has exactly two. Ask for each one by name and let the assertions repeat. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Two helpers existed to keep the judge criterion and the failure summary from being written twice, for a test with two requests. Write them twice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Splitting the requests into a test each lost what the file is named for: a second request needs a new ID, and that only happens inside one session. Ask for both in one test, with the prompts written where the agent is asked, so the two requests and the ID comparison between them read in order. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
…turing The fallback read the options out and then chose between them two lines later. Default each new option to its old name where the options are read, as the repository name already does. Name the old options in the test that covers them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
angeloashmore
marked this pull request as ready for review
September 22, 2026 21:32
levimykel
approved these changes
Sep 23, 2026
The gate printed "missing --task-id" for three different failures: no ID, an ID the CLI did not issue, and a good ID without --user-intent. An agent that passed an ID could read that as the CLI contradicting itself, and one missing only the intent could go and get a second ID for the same request. Name the option that failed, and ask for one ID per user request. Keep the ID format in one place, src/lib/task-id.ts, for the command that generates IDs and the gate that checks them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Each agent() call started a new conversation, so the agent handling the field never saw the slice's ID and could not reuse it. The check that the second request gets its own ID passed without testing anything. agent() now returns one turn, with continue() to send the next message in the same conversation, and calls holds only that turn's CLI calls. The eval asks for the field with continue(), so the first ID is in the agent's context when it decides whether to start a new one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Resolves:
Description
Before this PR, an agent that ran a command without
--analytics-task-idwas refused and told to generate a UUID itself with an inlinenode -ecommand. Agents read that as an instruction injected through a failed command. In one recorded run, Claude Code refused to continue and told the developer the CLI could not be trusted.After this PR, the CLI issues the IDs.
prismic task-idprints one ID for one user request, and the refusal names that command instead of dictating a recipe. The options are--task-idand--user-intent, so nothing an agent types on every command reads as telemetry.Everything here is agent-only, keyed on
detectAgent(): thetask-idcommand, both options, theAGENTSsection of the help, and the refusal itself. A human sees no new command, no new options and no refusal. The old--analytics-*names keep working for both, hidden from help.The order matters: release this to npm before prismicio/skills#12 merges, or agents are told to run a command their installed CLI does not have.
Checklist
Preview
How to QA 1
Run a command with an agent environment variable set, for example
CLAUDECODE=1 prismic whoami, and follow the refusal. Run the same command without one to see nothing change.🤖 Generated with Claude Code
https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Note
Medium Risk
Changes agent-only CLI entry validation and analytics option names; legacy flags remain, but agents must adopt the new id format and command flow.
Overview
Agents get CLI-issued task IDs instead of being told to mint UUIDs themselves. New hidden command
prismic task-idprints apt_…id fromgenTaskId(); agent runs must pass--task-idand--user-intent(replacing--analytics-*, which still parse but are hidden from help).Validation now requires the
pt_format, so arbitrary UUIDs or placeholders fail. Refusal text points atprismic task-idrather than an inlinenode -erecipe;task-iditself is exempt from the check.Evals now assert one task id per user request (not one id shared across unrelated requests), with and without the Prismic skill, and bump the skill ref for
prismic task-iddocs.Reviewed by Cursor Bugbot for commit afa7171. Bugbot is set up for automated code reviews on this repo. Configure here.
Generated by Claude Code
Footnotes
Please use these labels when submitting a review:
⚠️ #issue: Strongly suggest a change.
❓ #ask: Ask a question.
💡 #idea: Suggest an idea.
🎉 #nice: Share a compliment. ↩