Skip to content

feat(analytics): issue task IDs from the CLI - #316

Merged
angeloashmore merged 16 commits into
mainfrom
claude/analytics-minted-task-id
Sep 23, 2026
Merged

angeloashmore merged 16 commits into
mainfrom
claude/analytics-minted-task-id

Conversation

@angeloashmore

@angeloashmore angeloashmore commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

Resolves:

Description

Before this PR, an agent that ran a command without --analytics-task-id was refused and told to generate a UUID itself with an inline node -e command. Agents read that as an instruction injected through a failed command. In one recorded run, Claude Code refused to continue and told the developer the CLI could not be trusted.

After this PR, the CLI issues the IDs. prismic task-id prints one ID for one user request, and the refusal names that command instead of dictating a recipe. The options are --task-id and --user-intent, so nothing an agent types on every command reads as telemetry.

Everything here is agent-only, keyed on detectAgent(): the task-id command, both options, the AGENTS section of the help, and the refusal itself. A human sees no new command, no new options and no refusal. The old --analytics-* names keep working for both, hidden from help.

The order matters: release this to npm before prismicio/skills#12 merges, or agents are told to run a command their installed CLI does not have.

Checklist

  • If my changes require tests, I added them.
  • If my changes affect backward compatibility, it has been discussed.
  • If my changes require an update to the CONTRIBUTING.md guide, I updated it.

Preview

$ npx prismic@pr-316 whoami                   # as a human
Not logged in. Run `prismic login` first.

$ CLAUDECODE=1 npx prismic@pr-316 whoami      # as an agent
error: missing --task-id
run `prismic task-id` once per user request, then pass --task-id <id> --user-intent "<what the user asked for>" on every command until the user asks for something else

$ npx prismic@pr-316 task-id
pt_sewgrz7vydac0ksb

How to QA 1

Run a command with an agent environment variable set, for example CLAUDECODE=1 prismic whoami, and follow the refusal. Run the same command without one to see nothing change.

🤖 Generated with Claude Code

https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6


Note

Medium Risk
Changes agent-only CLI entry validation and analytics option names; legacy flags remain, but agents must adopt the new id format and command flow.

Overview
Agents get CLI-issued task IDs instead of being told to mint UUIDs themselves. New hidden command prismic task-id prints a pt_… id from genTaskId(); agent runs must pass --task-id and --user-intent (replacing --analytics-*, which still parse but are hidden from help).

Validation now requires the pt_ format, so arbitrary UUIDs or placeholders fail. Refusal text points at prismic task-id rather than an inline node -e recipe; task-id itself is exempt from the check.

Evals now assert one task id per user request (not one id shared across unrelated requests), with and without the Prismic skill, and bump the skill ref for prismic task-id docs.

Reviewed by Cursor Bugbot for commit afa7171. Bugbot is set up for automated code reviews on this repo. Configure here.


Generated by Claude Code

Footnotes

  1. Please use these labels when submitting a review:
    ❓ #ask: Ask a question.
    💡 #idea: Suggest an idea.
    ⚠️ #issue: Strongly suggest a change.
    🎉 #nice: Share a compliment. ↩

Angelo Ashmore and others added 3 commits September 18, 2026 19:50
The CLI refused an agent's command until it passed --analytics-task-id and
--analytics-intent, and answered on the stderr of the failed command with an
instruction block: generate a UUID, run this `node -e` line, never use a
placeholder. An agent meeting that cold cannot tell it from a prompt
injection, and agents supplied placeholders anyway.

The options are now --task-id and --user-intent. The old names still work.
`prismic task-id` issues an ID, so the CLI decides the format and a value it
did not issue is refused: a placeholder cannot be passed. The error names that
command instead of explaining how to invent an ID.

The word "analytics" left the option names because an agent reading it in a
flag it types on every command treats the call as data collection and says so
to the user. The MCP server's userIntent carries the same data under a name
that does not, and no agent has ever raised it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Not published. The eval reads it with EVAL_SKILL_FILE so a CLI change and the
skill change it needs can be tested together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Two bugs in the task-grouping eval made a working candidate look broken.

The second judge interpolated a literal `${DATE_FIELD}` rather than the
request, so it graded intents against a placeholder. And requiring one
distinct intent string per request failed agents that reworded a correct
paraphrase between commands, while passing an agent that sent "test" on
every call.

Analytics groups by task ID, so the intent is a label for the group, not a
key. Judge each distinct value for being a real paraphrase of the request
and drop the equality check; the one-ID-per-request assertions stand. The
failure message now prints intents alongside IDs, so an intent failure no
longer reports IDs.

Also bring the gate's unit tests to the current design: IDs are minted by
`prismic task-id`, a UUID is no longer accepted, and the old option names
still work but are hidden from help.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Comment thread src/tracking.ts Fixed
Angelo Ashmore and others added 11 commits September 22, 2026 19:18
The alphabet has 32 characters, so `byte % 32` was already unbiased, but
the expression reads as the biased pattern and breaks silently if the
alphabet ever changes length. Mask the low five bits instead, which is
exact by construction.

Also drop two imports the new gate message left unused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
The branch carried a copy of the Prismic skill so the eval could run against
wording that is not published yet. A copy drifts. The eval already pins a
skill commit and fetches it, so point that pin at the branch instead and drop
both the copy and the environment override that read it.

Fold the single-request test into the two-request one, whose first request
already covers it, and drop the UUID the task ID pattern still allowed: the
CLI issues every ID now, so a UUID can never be one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
The eval's own comment said a refused call loses no analytics and is not
asserted, but every call reached the assertions: the wrapper logs argv
before the CLI runs. An agent that passed a malformed ID, was refused, then
used one good ID for the rest of the request failed a trial whose analytics
were correct.

Keep the diagnostic listing every call, refusals included, and assert on the
ones that carried what the gate wants. Judge intents over the same set, which
had the same fault.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
The test held six pairs of first/second variables, hid its two requests
behind a standalone function, and explained itself in comments that were
harder to read than the code.

Collect the requests into one list and assert over it. Register the two
skill states with describe.for, which lets the test body be inline and
lets vitest infer its context. Name the helpers for what they select, so
needsTaskId and accepted say what a comment used to.

Rename mintTaskId to genTaskId.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
They were named constants when the judge referred to them by name. The
judge now reads the prompt from the request it is judging, leaving one use.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
The condition mixed whether a command needs a task ID with whether it has
one. Name each half, so the gate reads as the question it asks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Collecting the requests into a list needed a helper to build them and
another to print them, for a test that has exactly two. Ask for each one
by name and let the assertions repeat.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Two helpers existed to keep the judge criterion and the failure summary
from being written twice, for a test with two requests. Write them twice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Splitting the requests into a test each lost what the file is named for: a
second request needs a new ID, and that only happens inside one session.

Ask for both in one test, with the prompts written where the agent is asked,
so the two requests and the ID comparison between them read in order.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
…turing

The fallback read the options out and then chose between them two lines
later. Default each new option to its old name where the options are read,
as the repository name already does.

Name the old options in the test that covers them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
@angeloashmore
angeloashmore marked this pull request as ready for review September 22, 2026 21:32

@levimykel levimykel left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pre-approved with a few comments.

Comment thread evals/group-commands-by-task.eval.ts Outdated
Comment thread src/index.ts Outdated
Comment thread src/tracking.ts Outdated
Angelo Ashmore and others added 2 commits September 23, 2026 20:53
The gate printed "missing --task-id" for three different failures: no ID,
an ID the CLI did not issue, and a good ID without --user-intent. An agent
that passed an ID could read that as the CLI contradicting itself, and one
missing only the intent could go and get a second ID for the same request.

Name the option that failed, and ask for one ID per user request.

Keep the ID format in one place, src/lib/task-id.ts, for the command that
generates IDs and the gate that checks them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
Each agent() call started a new conversation, so the agent handling the
field never saw the slice's ID and could not reuse it. The check that the
second request gets its own ID passed without testing anything.

agent() now returns one turn, with continue() to send the next message in
the same conversation, and calls holds only that turn's CLI calls. The eval
asks for the field with continue(), so the first ID is in the agent's
context when it decides whether to start a new one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANFuYqfG55YAJUNNREVgE6
@angeloashmore
angeloashmore merged commit 1213a3e into main Sep 23, 2026
16 of 18 checks passed
@angeloashmore
angeloashmore deleted the claude/analytics-minted-task-id branch September 23, 2026 21:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants