From 97e83d2e382d35e59f4978ca307683f87ec47790 Mon Sep 17 00:00:00 2001 From: Hoang Nguyen Date: Sun, 27 Sep 2026 08:51:21 +0000 Subject: [PATCH 01/10] docs(requirements): openai-pi-login-capacity requirements --- ...-09-27-feature-openai-pi-login-capacity.md | 65 +++++++++++++ ...-09-27-feature-openai-pi-login-capacity.md | 73 +++++++++++++++ ...-09-27-feature-openai-pi-login-capacity.md | 68 ++++++++++++++ ...-09-27-feature-openai-pi-login-capacity.md | 78 ++++++++++++++++ ...-09-27-feature-openai-pi-login-capacity.md | 91 +++++++++++++++++++ 5 files changed, 375 insertions(+) create mode 100644 docs/ai/design/2026-09-27-feature-openai-pi-login-capacity.md create mode 100644 docs/ai/implementation/2026-09-27-feature-openai-pi-login-capacity.md create mode 100644 docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md create mode 100644 docs/ai/requirements/2026-09-27-feature-openai-pi-login-capacity.md create mode 100644 docs/ai/testing/2026-09-27-feature-openai-pi-login-capacity.md diff --git a/docs/ai/design/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/design/2026-09-27-feature-openai-pi-login-capacity.md new file mode 100644 index 00000000..c310b17f --- /dev/null +++ b/docs/ai/design/2026-09-27-feature-openai-pi-login-capacity.md @@ -0,0 +1,65 @@ +--- +phase: design +title: System Design & Architecture +description: Define the technical architecture, components, and data models +--- + +# System Design & Architecture + +## Architecture Overview + +**What is the high-level system structure?** + +- Include a mermaid diagram that captures the main components and their relationships. Example: + ```mermaid + graph TD + Client -->|HTTPS| API + API --> ServiceA + API --> ServiceB + ServiceA --> Database[(DB)] + ``` +- Key components and their responsibilities +- Technology stack choices and rationale + +## Data Models + +**What data do we need to manage?** + +- Core entities and their relationships +- Data schemas/structures +- Data flow between components + +## API Design + +**How do components communicate?** + +- External APIs (if applicable) +- Internal interfaces +- Request/response formats +- Authentication/authorization approach + +## Component Breakdown + +**What are the major building blocks?** + +- Frontend components (if applicable) +- Backend services/modules +- Database/storage layer +- Third-party integrations + +## Design Decisions + +**Why did we choose this approach?** + +- Key architectural decisions and trade-offs +- Alternatives considered +- Patterns and principles applied + +## Non-Functional Requirements + +**How should the system perform?** + +- Performance targets +- Scalability considerations +- Security requirements +- Reliability/availability needs diff --git a/docs/ai/implementation/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/implementation/2026-09-27-feature-openai-pi-login-capacity.md new file mode 100644 index 00000000..6b1942f7 --- /dev/null +++ b/docs/ai/implementation/2026-09-27-feature-openai-pi-login-capacity.md @@ -0,0 +1,73 @@ +--- +phase: implementation +title: Implementation Guide +description: Technical implementation notes, patterns, and code guidelines +--- + +# Implementation Guide + +## Development Setup + +**How do we get started?** + +- Prerequisites and dependencies +- Environment setup steps +- Configuration needed + +## Code Structure + +**How is the code organized?** + +- Directory structure +- Module organization +- Naming conventions + +## Implementation Notes + +**Key technical details to remember:** + +### Core Features + +- Feature 1: Implementation approach +- Feature 2: Implementation approach +- Feature 3: Implementation approach + +### Patterns & Best Practices + +- Design patterns being used +- Code style guidelines +- Common utilities/helpers + +## Integration Points + +**How do pieces connect?** + +- API integration details +- Database connections +- Third-party service setup + +## Error Handling + +**How do we handle failures?** + +- Error handling strategy +- Logging approach +- Retry/fallback mechanisms + +## Performance Considerations + +**How do we keep it fast?** + +- Optimization strategies +- Caching approach +- Query optimization +- Resource management + +## Security Notes + +**What security measures are in place?** + +- Authentication/authorization +- Input validation +- Data encryption +- Secrets management diff --git a/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md new file mode 100644 index 00000000..1c27b2dc --- /dev/null +++ b/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md @@ -0,0 +1,68 @@ +--- +phase: planning +title: Project Planning & Task Breakdown +description: Break down work into actionable tasks and estimate timeline +--- + +# Project Planning & Task Breakdown + +## Milestones + +**What are the major checkpoints?** + +- [ ] Milestone 1: [Description] +- [ ] Milestone 2: [Description] +- [ ] Milestone 3: [Description] + +## Task Breakdown + +**What specific work needs to be done?** + +### Phase 1: Foundation + +- [ ] Task 1.1: [Description] +- [ ] Task 1.2: [Description] + +### Phase 2: Core Features + +- [ ] Task 2.1: [Description] +- [ ] Task 2.2: [Description] + +### Phase 3: Integration & Polish + +- [ ] Task 3.1: [Description] +- [ ] Task 3.2: [Description] + +## Dependencies + +**What needs to happen in what order?** + +- Task dependencies and blockers +- External dependencies (APIs, services, etc.) +- Team/resource dependencies + +## Timeline & Estimates + +**When will things be done?** + +- Estimated effort per task/phase +- Target dates for milestones +- Buffer for unknowns + +## Risks & Mitigation + +**What could go wrong?** + +- Technical risks +- Resource risks +- Dependency risks +- Mitigation strategies + +## Resources Needed + +**What do we need to succeed?** + +- Team members and roles +- Tools and services +- Infrastructure +- Documentation/knowledge diff --git a/docs/ai/requirements/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/requirements/2026-09-27-feature-openai-pi-login-capacity.md new file mode 100644 index 00000000..a57f8fa3 --- /dev/null +++ b/docs/ai/requirements/2026-09-27-feature-openai-pi-login-capacity.md @@ -0,0 +1,78 @@ +--- +phase: requirements +title: Requirements & Problem Understanding +description: Clarify the problem space, gather requirements, and define success criteria +--- + +# Requirements & Problem Understanding + +## Problem Statement + +**What problem are we solving?** + +- Users who log in to OpenAI through pi (OAuth, credential `openai-codex` in `~/.pi/agent/auth.json`) get no OpenAI status in `ai-devkit capacity` under the pi harness. The default report prints `openai capacity unavailable: OpenAI API key not found` and omits the `pi · OpenAI` rows, even though a valid ChatGPT-backed login exists on the machine. +- Affected users: pi harness users on the ChatGPT/Codex subscription plan who have not provisioned a platform API key (the common case for `pi login`). +- Current workaround: create a platform key and set `OPENAI_API_KEY` or store a pi `openai` `{type: "api_key"}` credential — manual setup the feature must eliminate. + +**Verified facts (from discovery; treated as ground truth):** + +1. Pi's OpenAI login writes `~/.pi/agent/auth.json` entry `openai-codex` with OAuth shape `{type, access, refresh, expires, accountId}` — `access`/`refresh` are token strings, `expires` is an epoch-**milliseconds** integer, `accountId` is the ChatGPT account UUID. Present on this machine. +2. The merged platform-key provider (#252, `packages/agent-manager/src/capacity/openai.ts`) only resolves `OPENAI_API_KEY` env or a pi `openai` `{type: "api_key", key}` credential. It can never see `openai-codex`. That path must keep working unchanged. +3. Precedent: `capacity/codex.ts` fetches `https://chatgpt.com/backend-api/wham/usage` with OAuth tokens (Bearer + `ChatGPT-Account-Id` headers) and parses `rate_limit` / `additional_rate_limits` windows and credits. The pi `openai-codex` tokens are expected to authorize the same endpoint (validated by a read-only live probe during design). +4. codex.ts policy for stale OAuth: **no refresh attempt**; stale tokens are skipped and the probe falls back. Pi tokens follow the same policy (no new refresh flows). + +## Goals & Objectives + +**What do we want to achieve?** + +- Primary: `ai-devkit capacity` (default report) shows `pi · OpenAI` rows — beside `pi · z.ai` — whenever the user has logged in to OpenAI via pi, with zero manual setup and no platform API key. +- Secondary: identical treatment for `ai-devkit capacity openai` (explicit provider) and `--json` output. +- Keep the platform-key path (env → pi `openai` api_key) fully working and preferred when present; OAuth is the no-setup fallback, not a replacement. +- Never print or persist token values in output, docs, logs, or tests. + +**Non-goals (what's explicitly out of scope)** + +- No OAuth token refresh implementation (matches codex.ts precedent). +- No changes to the codex harness provider (`codex · OpenAI` rows) or its auth handling. +- No removal/alteration of the platform API-key usage report (tokens today/7-day windows). +- No new providers, subcommands, or render overhaul beyond showing the new rows. + +## User Stories & Use Cases + +- As a pi user logged in to OpenAI via `pi login`, when I run `ai-devkit capacity`, I see `pi · OpenAI` quota rows (session/weekly windows, credits) beside `pi · z.ai`, without configuring any API key. +- As a user with a platform key (`OPENAI_API_KEY` or pi `openai` api_key), my report is unchanged (token-usage windows), and OAuth is not consulted. +- As a user whose pi OAuth access token is expired, I see a clear authenticated-but-stale state (login detected, usage not fetched) instead of "API key not found". +- As a user with logins in both pi and codex harnesses, both `pi · OpenAI` and `codex · OpenAI` rows appear — the harness column distinguishes them; no dedupe. + +**Edge cases** + +- `~/.pi/agent/auth.json` missing or malformed → current clear error behavior for the openai provider is preserved (no rows, warning in multi-provider mode). +- `openai-codex` entry malformed (wrong type, missing access/accountId) → skip OAuth path gracefully (fall through to clear error), never expose file content. +- Both platform key and OAuth present → platform key wins; OAuth not used. +- wham/usage returns 401 (revoked OAuth) → render unauthenticated rather than erroring the whole report. + +## Success Criteria + +- On a machine with only a pi OpenAI OAuth login, default `ai-devkit capacity` shows `pi · OpenAI` rows and no `openai capacity unavailable` warning. +- With `OPENAI_API_KEY` set or pi `openai` api_key present, output is byte-identical to today's platform-key report. +- Expired OAuth renders an authenticated-but-stale `pi · OpenAI` state (no fetch attempted). +- Unit tests: resolution order, staleness classification, wham parsing/window mapping, 401 handling, and no-token-leak assertions — all fixture/mock based, mirroring `openai.test.ts`/`codex.test.ts` patterns. No live network calls in tests. +- Gates: scoped vitest suites, `npm run lint`, `npm test`, `npm run test:e2e` green; commit hooks pass. + +## Constraints & Assumptions + +- The wham/usage endpoint accepts the pi OAuth tokens (assumption verified once via read-only live GET during design; tests then use fixtures only). +- `expires` is epoch milliseconds (pi convention); codex.ts `expires_at` is seconds — the pi path must not reuse the seconds-based staleness math blindly. +- One live probe against chatgpt.com backend with real stored tokens, read-only GET, results recorded redacted; never committed. +- Token values are never logged, thrown, or written to any artifact. + +## Questions & Open Items + +All material questions are resolved by the brief and verified discovery (documented above): + +1. **Resolution order** — env `OPENAI_API_KEY` → pi `openai` api_key → pi `openai-codex` OAuth. Platform paths stay first because they are correct for platform-key users (brief rule 2); OAuth closes the no-setup gap. +2. **Dedupe/classification** — none. Both harness rows may legitimately coexist; the harness column exists for exactly this. +3. **OAuth refresh** — no refresh (match codex.ts); stale renders authenticated-but-stale. +4. **Reuse vs duplication** — wham parsing gains a second real caller (pi OAuth), which justifies extracting the parser into a shared module rather than duplicating; final call made in design review. + +No unresolved items requiring stakeholder input. diff --git a/docs/ai/testing/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/testing/2026-09-27-feature-openai-pi-login-capacity.md new file mode 100644 index 00000000..23a3a2a3 --- /dev/null +++ b/docs/ai/testing/2026-09-27-feature-openai-pi-login-capacity.md @@ -0,0 +1,91 @@ +--- +phase: testing +title: Testing Strategy +description: Define testing approach, test cases, and quality assurance +--- + +# Testing Strategy + +## Test Coverage Goals + +**What level of testing do we aim for?** + +- Unit test coverage target (default: 100% of new/changed code) +- Integration test scope (critical paths + error handling) +- End-to-end test scenarios (key user journeys) +- Alignment with requirements/design acceptance criteria + +## Unit Tests + +**What individual components need testing?** + +### Component/Module 1 + +- [ ] Test case 1: [Description] (covers scenario / branch) +- [ ] Test case 2: [Description] (covers edge case / error handling) +- [ ] Additional coverage: [Description] + +### Component/Module 2 + +- [ ] Test case 1: [Description] +- [ ] Test case 2: [Description] +- [ ] Additional coverage: [Description] + +## Integration Tests + +**How do we test component interactions?** + +- [ ] Integration scenario 1 +- [ ] Integration scenario 2 +- [ ] API endpoint tests +- [ ] Integration scenario 3 (failure mode / rollback) + +## End-to-End Tests + +**What user flows need validation?** + +- [ ] User flow 1: [Description] +- [ ] User flow 2: [Description] +- [ ] Critical path testing +- [ ] Regression of adjacent features + +## Test Data + +**What data do we use for testing?** + +- Test fixtures and mocks +- Seed data requirements +- Test database setup + +## Test Reporting & Coverage + +**How do we verify and communicate test results?** + +- Coverage commands and thresholds (`npm run test -- --coverage`) +- Coverage gaps (files/functions below 100% and rationale) +- Links to test reports or dashboards +- Manual testing outcomes and sign-off + +## Manual Testing + +**What requires human validation?** + +- UI/UX testing checklist (include accessibility) +- Browser/device compatibility +- Smoke tests after deployment + +## Performance Testing + +**How do we validate performance?** + +- Load testing scenarios +- Stress testing approach +- Performance benchmarks + +## Bug Tracking + +**How do we manage issues?** + +- Issue tracking process +- Bug severity levels +- Regression testing strategy From 67686697b63055f6f720b3c323b7eeb4ec5f3734 Mon Sep 17 00:00:00 2001 From: Hoang Nguyen Date: Sun, 27 Sep 2026 08:54:25 +0000 Subject: [PATCH 02/10] docs(design): openai-pi-login-capacity design and testing strategy --- ...-09-27-feature-openai-pi-login-capacity.md | 117 ++++++++++++------ ...-09-27-feature-openai-pi-login-capacity.md | 103 ++++++--------- 2 files changed, 120 insertions(+), 100 deletions(-) diff --git a/docs/ai/design/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/design/2026-09-27-feature-openai-pi-login-capacity.md index c310b17f..d0adba42 100644 --- a/docs/ai/design/2026-09-27-feature-openai-pi-login-capacity.md +++ b/docs/ai/design/2026-09-27-feature-openai-pi-login-capacity.md @@ -8,58 +8,99 @@ description: Define the technical architecture, components, and data models ## Architecture Overview -**What is the high-level system structure?** - -- Include a mermaid diagram that captures the main components and their relationships. Example: - ```mermaid - graph TD - Client -->|HTTPS| API - API --> ServiceA - API --> ServiceB - ServiceA --> Database[(DB)] - ``` -- Key components and their responsibilities -- Technology stack choices and rationale +The feature extends the existing pi·OpenAI capacity provider (`capacity/openai.ts`) with an OAuth tier backed by pi's OpenAI login, reusing the wham/usage endpoint that the codex harness provider already consumes. No CLI or render changes are required — `render.ts` already renders generic `CapacityWindow` rows and labels provider `openai` as "OpenAI". + +```mermaid +graph TD + CLI["ai-devkit capacity"] --> CMD["capacity command
(packages/cli)"] + CMD --> IDX["getOpenAiCapacityReport
(agent-manager/capacity/index.ts)"] + CMD --> CODEX["getCodexCapacityReport
(codex.ts, unchanged)"] + IDX --> OA["probeOpenAiCapacity
(openai.ts)"] + OA --> RES["resolveOpenAiCredential"] + RES -->|env OPENAI_API_KEY| T1["tier 1: platform key"] + RES -->|pi auth.json openai api_key| T2["tier 2: platform key"] + RES -->|pi auth.json openai-codex oauth| T3["tier 3: OAuth"] + T1 --> PLAT["api.openai.com usage API
(today/7-day token windows)"] + T3 --> WHAM["GET chatgpt.com/backend-api/wham/usage
Bearer + ChatGPT-Account-Id"] + WHAM --> PARSE["parseWhamUsage
(shared wham.ts)"] + CODEX --> PARSE + PLAT --> RPT["CapacityReport harness=pi provider=openai"] + PARSE --> RPT +``` + +Key components: + +- `packages/agent-manager/src/capacity/openai.ts` — credential resolution becomes tiered; gains the wham/usage fetch path and OAuth staleness classification. +- `packages/agent-manager/src/capacity/wham.ts` (new) — shared wham/usage response parsing (rate-limit windows, extra limits, credits), extracted from `codex.ts` now that a second real caller exists. +- `packages/agent-manager/src/capacity/codex.ts` — unchanged behavior; `parseUsage`/`toRateWindow` remain exported (tests import them) and delegate to `wham.ts`. +- CLI (`packages/cli`) — no changes. ## Data Models -**What data do we need to manage?** +**Input credential (read-only), tier 3 — pi `~/.pi/agent/auth.json`:** -- Core entities and their relationships -- Data schemas/structures -- Data flow between components +```jsonc +{ + "openai-codex": { + "type": "oauth", // verified live: "oauth" + "access": "", // never logged/persisted + "refresh": "", // unused (no refresh, see D2) + "expires": 1791355868557, // epoch MILLISECONDS (verified live) + "accountId": "" + } +} +``` + +**Resolved credential (internal discriminated union):** + +```ts +type OpenAiCredential = + | { kind: "platform"; key: string } // tiers 1-2 + | { kind: "oauth"; access: string; accountId: string; expiresMs: number | null }; +``` + +**Output** — existing `CapacityReport`/`CapacityWindow` types, unchanged. OAuth windows come from wham: `session` (5h), `weekly` (7d), plus any `additional_rate_limits` entries; `creditsRemaining` from `credits.balance`. + +**Staleness rule (ms-aware):** `expires` values above 1e12 are epoch ms (pi convention), smaller values are treated as seconds; JWT `exp` (seconds) is the fallback when `expires` is absent/invalid — mirroring `codex.ts staleOAuth`'s metadata-then-JWT order. Reference "now" is `Date.parse(checkedAt)` (already injectable via `now` option) so tests are deterministic. ## API Design -**How do components communicate?** +**External:** one read-only `GET https://chatgpt.com/backend-api/wham/usage` with headers `Authorization: Bearer ` and `ChatGPT-Account-Id: ` — identical to the codex provider's call. Verified live with the real pi token: HTTP 200, `rate_limit.primary_window/secondary_window`, `credits.balance`, top-level `rate_limit_reached_type` present. Timeout and abort handling follow the existing AbortController pattern in `openai.ts`. -- External APIs (if applicable) -- Internal interfaces -- Request/response formats -- Authentication/authorization approach +**Internal:** `probeOpenAiCapacity(options)` keeps its signature (`env`, `readFile`, `fetch`, `timeoutMs`, `checkedAt` via `getOpenAiCapacityReport`). `resolveOpenAiApiKey` remains exported for compatibility; a new internal `resolveOpenAiCredential` returns the union above. Resolution order: -## Component Breakdown +1. `env.OPENAI_API_KEY` (explicit override; auth file not even read — preserved behavior) +2. pi `openai` `{type: "api_key", key}` (platform key path, byte-identical report) +3. pi `openai-codex` `{type: "oauth"}` with non-empty `access` + `accountId` → OAuth path +4. nothing found → error `OpenAI credentials not found; set OPENAI_API_KEY, configure the Pi openai provider, or log in to OpenAI via Pi` -**What are the major building blocks?** +**Response classification (OAuth tier):** -- Frontend components (if applicable) -- Backend services/modules -- Database/storage layer -- Third-party integrations +- fresh token → fetch wham; 200 → report (authenticated, windows, credits) +- 401/403 → `{authenticated: false, available: "unknown", windows: []}` (revoked login) +- timeout / network error / 5xx / malformed JSON → sanitized throw (`OpenAI usage request failed…`), rendered as the per-provider warning by `capacityCommand` — consistent with the platform tier's error behavior +- stale token (expiry ≤ now) → **no fetch**; `{authenticated: true, available: "unknown", windows: [], creditsRemaining: null}` (authenticated-but-stale) +- malformed `openai-codex` entry (wrong type, missing fields) → tier skipped, falls to step 4 error; error messages never include file content -## Design Decisions +## Component Breakdown -**Why did we choose this approach?** +- `wham.ts` (new): `toRateWindow`, `extraWindows`, `parseWhamUsage(raw): { windows, creditsRemaining }`, plus the small guards (`record`, `finiteNumber`, `nonEmptyText`, `resetTime`, `safeIdentifier`). Pure functions, no I/O. +- `codex.ts`: imports parsing from `wham.ts`; re-exports `toRateWindow` and keeps `parseUsage(raw, source)` as a thin wrapper so `codex.test.ts` passes unchanged. Probe behavior identical. +- `openai.ts`: `resolveOpenAiCredential` tiers; OAuth fetch helper (Bearer + account id, abort, status classification); staleness check; report assembly mirroring codex's `capacityFromSnapshot` for API responses (`available = hasUsage ? "yes" : "unknown"`; top-level `rate_limit_reached_type` is deliberately **not** read, matching codex.ts API-path behavior). +- Tests/fixtures: `openai-codex-auth.json` (fake tokens), `wham-usage.json` (redacted live shape) under `capacity/fixtures/`. -- Key architectural decisions and trade-offs -- Alternatives considered -- Patterns and principles applied +## Design Decisions -## Non-Functional Requirements +- **D1 Resolution order (env → platform key → OAuth).** Platform paths stay first: they are correct for platform-key users and the brief mandates keeping them working. OAuth only fills the no-setup gap. Alternative (OAuth first) would change output for dual-credential users without a motivating use case. Alternative (show both) would need two `pi · OpenAI` rows and render changes — rejected as scope creep. +- **D2 No OAuth refresh.** Matches codex.ts precedent exactly (its OAuth path never refreshes; stale → skip). Attempting refresh would invent a new flow (token rotation, persistence of rotated tokens into pi's auth file — writing another tool's state) with no existing pattern to follow. Stale renders authenticated-but-stale. +- **D3 Extract wham parsing into `wham.ts`.** The brief allows extraction only with a second real caller — that now exists (pi OAuth tier), verified live against the same endpoint/shape. Alternatives: importing `parseUsage` from `codex.ts` (couples the pi provider to the codex provider) or duplicating ~60 lines of parsing (drift risk) — both worse than a small shared pure module. Codex re-exports keep its public surface stable. +- **D4 ms-vs-s staleness heuristic.** pi writes epoch ms (verified); codex writes seconds. A magnitude check (`> 1e12` ⇒ ms) handles both without guessing per-source conventions, and JWT fallback covers missing metadata. +- **D5 No dedupe with `codex · OpenAI`.** Different credential stores (`~/.codex/auth.json` vs `~/.pi/agent/auth.json`); both rows legitimately coexist; the harness column exists for exactly this. +- **D6 Error message updated** to mention the pi login path (it fires only when nothing is found, so platform users never see it). Sanitized errors only — no token or file content in messages, matching the existing openai/codex tests' no-leak assertions. -**How should the system perform?** +## Non-Functional Requirements -- Performance targets -- Scalability considerations -- Security requirements -- Reliability/availability needs +- **Security:** token values never appear in reports, errors, logs, docs, or fixtures (fixtures use fake values); auth file is read-only; one live probe only (done, read-only GET). +- **Performance:** one additional HTTPS GET (timeout 5s default, abortable) only in the OAuth tier; platform tier unchanged; tests use injected `fetch` (no network). +- **Reliability:** failure modes classified (stale / unauthenticated / unavailable) so one provider's failure never blanks the multi-provider report. +- **Compatibility:** `resolveOpenAiApiKey`, `parseUsage`, `toRateWindow` exports preserved; platform-key report byte-identical. diff --git a/docs/ai/testing/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/testing/2026-09-27-feature-openai-pi-login-capacity.md index 23a3a2a3..e925e3dc 100644 --- a/docs/ai/testing/2026-09-27-feature-openai-pi-login-capacity.md +++ b/docs/ai/testing/2026-09-27-feature-openai-pi-login-capacity.md @@ -8,84 +8,63 @@ description: Define testing approach, test cases, and quality assurance ## Test Coverage Goals -**What level of testing do we aim for?** - -- Unit test coverage target (default: 100% of new/changed code) -- Integration test scope (critical paths + error handling) -- End-to-end test scenarios (key user journeys) -- Alignment with requirements/design acceptance criteria +- Unit coverage of new/changed code: 100% of branches in `wham.ts` parsing, OAuth credential resolution, staleness classification, and response classification. +- Existing suites (`openai.test.ts`, `codex.test.ts`, CLI capacity tests) must pass with no behavioral regressions; platform-key report output byte-identical. +- No live network calls in tests — all fetch via `vi.fn()` injection, all auth via fixtures with fake tokens. +- E2E: existing `npm run test:e2e` suite must stay green; capacity has no e2e cases today, and none are added (unit + render coverage is the established pattern for this command). ## Unit Tests -**What individual components need testing?** +### `capacity/wham.ts` (shared parser) -### Component/Module 1 +- [ ] Parses `rate_limit.primary_window`/`secondary_window` into session/weekly windows with `used_percent`, `limit_window_seconds`→durationMinutes, `reset_at` (epoch seconds and ISO string) → resetsAt (covers codex API mapping parity) +- [ ] Maps `additional_rate_limits` entries into scoped windows; tolerates non-array/absent value (live probe showed non-array) +- [ ] Extracts `credits.balance`; missing credits → null +- [ ] Skips windows it cannot parse without inventing zeros (missing used_percent stays null) +- [ ] Codex suite (`codex.test.ts`) passes unchanged against the delegated parser (regression guard) -- [ ] Test case 1: [Description] (covers scenario / branch) -- [ ] Test case 2: [Description] (covers edge case / error handling) -- [ ] Additional coverage: [Description] +### `capacity/openai.ts` — credential resolution -### Component/Module 2 +- [ ] `OPENAI_API_KEY` env wins without reading the auth file (existing behavior preserved) +- [ ] pi `openai` api_key credential wins over `openai-codex` when both exist (platform preferred, OAuth not consulted) +- [ ] pi `openai-codex` `{type:"oauth"}` resolves to OAuth credential when no platform key exists +- [ ] Malformed `openai-codex` (wrong type / missing access / missing accountId) is skipped → not-found error, message mentions pi login, never exposes file content +- [ ] Missing auth file (ENOENT) and invalid JSON keep clear sanitized errors -- [ ] Test case 1: [Description] -- [ ] Test case 2: [Description] -- [ ] Additional coverage: [Description] +### `capacity/openai.ts` — OAuth probe & classification -## Integration Tests +- [ ] Fresh token (epoch-ms `expires` in future) → GET wham/usage with `Authorization: Bearer` + `ChatGPT-Account-Id`, timeout abort wiring; 200 → report with `harness:"pi"`, session/weekly windows, credits, `authenticated:true` +- [ ] `expires` in seconds (small magnitude) and JWT-exp fallback both classify correctly; missing metadata defaults to fresh-then-fetch +- [ ] Stale token (expiry ≤ checkedAt) → **no fetch**, `authenticated:true`, `available:"unknown"`, no windows +- [ ] 401/403 from wham → `authenticated:false`, `available:"unknown"`, empty windows +- [ ] Timeout / network error / 5xx / bad JSON → sanitized thrown error (`OpenAI usage request failed…`), no token in message +- [ ] `resolveOpenAiApiKey` compat export unchanged for platform tiers +- [ ] No-token-leak: JSON.stringify of every report/error never contains fixture token strings -**How do we test component interactions?** +### CLI (existing suites) -- [ ] Integration scenario 1 -- [ ] Integration scenario 2 -- [ ] API endpoint tests -- [ ] Integration scenario 3 (failure mode / rollback) +- [ ] `capacity` command tests keep passing (default provider list unchanged; no CLI code changes) -## End-to-End Tests +## Integration Tests -**What user flows need validation?** +- [ ] `getOpenAiCapacityReport` end-to-end with injected `readFile` (fixture auth) + injected `fetch` (wham 200) → full report shape (windows sorted/renderable, provider label path) +- [ ] Multi-provider flow: OAuth-only machine → openai provider no longer throws (row present) while other providers' failures still warn independently (`capacityCommand` behavior unchanged) -- [ ] User flow 1: [Description] -- [ ] User flow 2: [Description] -- [ ] Critical path testing -- [ ] Regression of adjacent features +## End-to-End Tests -## Test Data +- [ ] Existing e2e suite green (`npm run test:e2e`); no new e2e (no CLI surface change) +- [ ] Manual validation on this machine: default `ai-devkit capacity` shows `pi · OpenAI` rows beside `pi · z.ai` (documented in implementation doc with redacted output) -**What data do we use for testing?** +## Test Data -- Test fixtures and mocks -- Seed data requirements -- Test database setup +- New fixtures in `packages/agent-manager/src/__tests__/capacity/fixtures/`: + - `openai-codex-auth.json` — fake pi OAuth credential (`type:"oauth"`, JWT-shaped fake access with `exp`, epoch-ms `expires`, fake accountId) + - `wham-usage.json` — redacted copy of the verified live response shape (windows, credits, top-level keys) +- Existing platform fixtures (`openai-auth.json`, usage payloads) reused unchanged. +- All token-like strings are obvious fakes; real tokens never committed (verified by grepping the branch for token substrings before PR). ## Test Reporting & Coverage -**How do we verify and communicate test results?** - -- Coverage commands and thresholds (`npm run test -- --coverage`) -- Coverage gaps (files/functions below 100% and rationale) -- Links to test reports or dashboards -- Manual testing outcomes and sign-off - -## Manual Testing - -**What requires human validation?** - -- UI/UX testing checklist (include accessibility) -- Browser/device compatibility -- Smoke tests after deployment - -## Performance Testing - -**How do we validate performance?** - -- Load testing scenarios -- Stress testing approach -- Performance benchmarks - -## Bug Tracking - -**How do we manage issues?** - -- Issue tracking process -- Bug severity levels -- Regression testing strategy +- Commands: `npx vitest run` in `packages/agent-manager` (scoped), then `npm run lint`, `npm test`, `npm run test:e2e` at root; coverage via existing thresholds (70% floor package-wide, new code targeted at 100% branch). +- Coverage gaps: none anticipated; any justified gap documented here with rationale. +- Evidence: command outputs recorded in the implementation doc; final validation in the PR body. From 31f037e1174a39ae6a8bcd045771093158cd35ae Mon Sep 17 00:00:00 2001 From: Hoang Nguyen Date: Sun, 27 Sep 2026 08:55:24 +0000 Subject: [PATCH 03/10] docs(planning): openai-pi-login-capacity task breakdown --- ...-09-27-feature-openai-pi-login-capacity.md | 63 +++++++------------ 1 file changed, 24 insertions(+), 39 deletions(-) diff --git a/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md index 1c27b2dc..20fcd4cf 100644 --- a/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md +++ b/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md @@ -8,61 +8,46 @@ description: Break down work into actionable tasks and estimate timeline ## Milestones -**What are the major checkpoints?** - -- [ ] Milestone 1: [Description] -- [ ] Milestone 2: [Description] -- [ ] Milestone 3: [Description] +- [ ] M1: Shared wham parsing extracted, codex suite green unchanged +- [ ] M2: Pi OAuth tier resolves + classifies (all unit scenarios green) +- [ ] M3: Full gates green, docs/implementation updated, final PR open ## Task Breakdown -**What specific work needs to be done?** - -### Phase 1: Foundation +### Phase 1: Foundation — shared parser -- [ ] Task 1.1: [Description] -- [ ] Task 1.2: [Description] +- [ ] T1.1: Create `capacity/wham.ts` with `parseWhamUsage`, `toRateWindow`, `extraWindows` + guards (TDD: add `wham.test.ts` scenarios from testing doc first; fixtures `wham-usage.json`). Outcome: pure parser module; validation: `npx vitest run src/__tests__/capacity/wham.test.ts` (agent-manager). +- [ ] T1.2: Delegate `codex.ts` parsing to `wham.ts`; keep `parseUsage`/`toRateWindow` exports. Outcome: codex behavior identical; validation: `codex.test.ts` passes unchanged (no edits to that suite). -### Phase 2: Core Features +### Phase 2: Core — OAuth tier in openai.ts -- [ ] Task 2.1: [Description] -- [ ] Task 2.2: [Description] +- [ ] T2.1: `resolveOpenAiCredential` tiers (env → pi `openai` api_key → pi `openai-codex` oauth → not-found error mentioning pi login); keep `resolveOpenAiApiKey` compat; ms-vs-s + JWT staleness helper. TDD first: extend `openai.test.ts` resolution + staleness scenarios with fixture `openai-codex-auth.json`. Validation: scoped vitest. +- [ ] T2.2: OAuth probe path: wham GET (Bearer + `ChatGPT-Account-Id`, abort/timeout), classification (200 report / 401-403 unauthenticated / timeout-5xx-badJSON sanitized throw / stale no-fetch), report assembly mirroring codex API-path semantics; no-leak assertions. Validation: scoped vitest; JSON.stringify scan for fake tokens. -### Phase 3: Integration & Polish +### Phase 3: Integration & gates -- [ ] Task 3.1: [Description] -- [ ] Task 3.2: [Description] +- [ ] T3.1: Full gates: `npm run lint`, `npm test`, `npm run test:e2e`, hooks (dev-commit) — all green in worktree. Record evidence in implementation doc. +- [ ] T3.2: Manual on-machine validation: build CLI, run default `ai-devkit capacity`, confirm `pi · OpenAI` rows beside `pi · z.ai`; redacted output in implementation doc. +- [ ] T3.3: Update implementation/testing/planning docs (phases 5-8 flow), grep branch for token substrings (no-leak audit), final PR (dev-pr conventions; do NOT merge). ## Dependencies -**What needs to happen in what order?** - -- Task dependencies and blockers -- External dependencies (APIs, services, etc.) -- Team/resource dependencies +- T1.1 → T1.2 → T2.2 (parser before callers). T2.1 independent of T1 but ordered after for clean diffs. +- T3.* depend on all of M1+M2. +- External: none beyond the already-validated wham endpoint (fixtures only from here). ## Timeline & Estimates -**When will things be done?** - -- Estimated effort per task/phase -- Target dates for milestones -- Buffer for unknowns +- Single-session execution: T1 ~30%, T2 ~50%, T3 ~20%. No calendar dates (autonomous run). ## Risks & Mitigation -**What could go wrong?** - -- Technical risks -- Resource risks -- Dependency risks -- Mitigation strategies - -## Resources Needed +- wham shape drift vs fixture → mitigated: parser is defensive (null-tolerant), fixtures from verified live payload. +- Codex test coupling to internal parse functions → mitigated: re-export shims keep imports stable. +- Coverage thresholds (70% floor, new code ~100%) → mitigated: classification branches enumerated in testing doc. +- Epoch-ms vs seconds misclassification → heuristic + JWT fallback, explicit tests for both magnitudes. +- Token leakage → fake fixtures, no-leak assertions, pre-PR grep audit. -**What do we need to succeed?** +## Progress Log -- Team members and roles -- Tools and services -- Infrastructure -- Documentation/knowledge +- 2026-09-27: Initial plan created from requirements/design/testing docs. All testing scenarios mapped to tasks (T1.1↔wham parsing, T2.1↔resolution/staleness, T2.2↔classification/no-leak, T3.1↔gates, T3.2↔manual e2e). From 047efb28c920714b48e1250746c55319f62d80f7 Mon Sep 17 00:00:00 2001 From: Hoang Nguyen Date: Sun, 27 Sep 2026 08:57:58 +0000 Subject: [PATCH 04/10] refactor(capacity): extract shared wham usage parsing --- .../capacity/fixtures/wham-usage.json | 32 ++++ .../src/__tests__/capacity/wham.test.ts | 80 +++++++++ packages/agent-manager/src/capacity/codex.ts | 159 +++--------------- packages/agent-manager/src/capacity/wham.ts | 84 +++++++++ 4 files changed, 221 insertions(+), 134 deletions(-) create mode 100644 packages/agent-manager/src/__tests__/capacity/fixtures/wham-usage.json create mode 100644 packages/agent-manager/src/__tests__/capacity/wham.test.ts create mode 100644 packages/agent-manager/src/capacity/wham.ts diff --git a/packages/agent-manager/src/__tests__/capacity/fixtures/wham-usage.json b/packages/agent-manager/src/__tests__/capacity/fixtures/wham-usage.json new file mode 100644 index 00000000..231223cb --- /dev/null +++ b/packages/agent-manager/src/__tests__/capacity/fixtures/wham-usage.json @@ -0,0 +1,32 @@ +{ + "user_id": "user-fixture-abc123", + "account_id": "account-fixture-uuid", + "email": "user@example.test", + "plan_type": "pro", + "rate_limit": { + "primary_window": { + "used_percent": 12.5, + "limit_window_seconds": 18000, + "reset_at": 1790517127 + }, + "secondary_window": { + "used_percent": 43.2, + "limit_window_seconds": 604800, + "reset_at": "2026-10-04T08:52:07Z" + } + }, + "code_review_rate_limit": { + "primary_window": { + "used_percent": 0, + "limit_window_seconds": 3600, + "reset_at": 1790510000 + } + }, + "additional_rate_limits": null, + "model_usage": [], + "credits": { "balance": 4.5 }, + "spend_control": {}, + "rate_limit_reached_type": "none", + "promo": null, + "rate_limit_reset_credits": { "availableCount": 0 } +} diff --git a/packages/agent-manager/src/__tests__/capacity/wham.test.ts b/packages/agent-manager/src/__tests__/capacity/wham.test.ts new file mode 100644 index 00000000..5764a4c6 --- /dev/null +++ b/packages/agent-manager/src/__tests__/capacity/wham.test.ts @@ -0,0 +1,80 @@ +import { readFile } from "node:fs/promises"; +import { fileURLToPath } from "node:url"; +import { describe, expect, it } from "vitest"; +import { parseWhamUsage, toRateWindow } from "../../capacity/wham.js"; + +const fixture = (name: string) => + readFile(fileURLToPath(new URL(`./fixtures/${name}`, import.meta.url)), "utf8"); + +describe("wham usage parsing", () => { + it("maps primary and secondary rate windows onto session and weekly", async () => { + const snapshot = parseWhamUsage(JSON.parse(await fixture("wham-usage.json"))); + expect(snapshot.windows).toEqual([ + { + id: "session", + label: "Session", + durationMinutes: 300, + usedPercent: 12.5, + resetsAt: "2026-09-27T13:52:07.000Z", + }, + { + id: "weekly", + label: "Weekly", + durationMinutes: 10080, + usedPercent: 43.2, + resetsAt: "2026-10-04T08:52:07.000Z", + }, + ]); + expect(snapshot.creditsRemaining).toBe(4.5); + }); + + it("accepts epoch seconds and ISO strings for reset_at", () => { + expect(toRateWindow({ used_percent: 1, reset_at: 1_787_220_000 }, "a", "A")?.resetsAt).toBe( + "2026-08-20T10:00:00.000Z", + ); + expect( + toRateWindow({ used_percent: 1, reset_at: "2026-08-20T10:00:00Z" }, "a", "A")?.resetsAt, + ).toBe("2026-08-20T10:00:00.000Z"); + expect(toRateWindow({ used_percent: 1, reset_at: "nope" }, "a", "A")?.resetsAt).toBeNull(); + }); + + it("maps additional_rate_limits into scoped windows and tolerates non-array values", () => { + const scoped = parseWhamUsage({ + rate_limit: {}, + additional_rate_limits: [ + { + limit_name: "code-review", + rate_limit: { + primary_window: { used_percent: 10, limit_window_seconds: 3600, reset_at: 100 }, + secondary_window: { used_percent: 20, limit_window_seconds: 7200, reset_at: 200 }, + }, + }, + ], + }); + expect(scoped.windows.map((window) => window.id)).toEqual([ + "code-review:primary", + "code-review:secondary", + ]); + expect(parseWhamUsage({ additional_rate_limits: null }).windows).toEqual([]); + }); + + it("never invents usage for windows it cannot parse", () => { + const snapshot = parseWhamUsage({ + rate_limit: { + primary_window: { limit_window_seconds: 18000 }, + secondary_window: null, + }, + }); + expect(snapshot.windows).toEqual([ + { + id: "session", + label: "Session", + durationMinutes: 300, + usedPercent: null, + resetsAt: null, + }, + ]); + expect(parseWhamUsage({}).creditsRemaining).toBeNull(); + expect(parseWhamUsage({ credits: { balance: "nan" } }).creditsRemaining).toBeNull(); + }); +}); diff --git a/packages/agent-manager/src/capacity/codex.ts b/packages/agent-manager/src/capacity/codex.ts index ca0f839a..1f6a9d61 100644 --- a/packages/agent-manager/src/capacity/codex.ts +++ b/packages/agent-manager/src/capacity/codex.ts @@ -1,8 +1,11 @@ import { spawn } from "node:child_process"; import { readFile } from "node:fs/promises"; import { join } from "node:path"; +import { parseWhamUsage, resetTime, safeIdentifier } from "./wham.js"; import type { CapacityReport, CapacityWindow } from "./types.js"; +export { toRateWindow } from "./wham.js"; + type CodexUsageSource = "pat" | "oauth" | "cli"; type UsageSnapshot = { windows: CapacityWindow[]; @@ -15,13 +18,7 @@ type RpcMessage = { id?: number; method: string; params?: UnknownRecord }; type CliResponses = { rateLimits: unknown; account: unknown }; type CodexRpc = (messages: RpcMessage[]) => Promise; -export const CODEX_APP_SERVER_ARGS = [ - "-s", - "read-only", - "-a", - "untrusted", - "app-server", -] as const; +export const CODEX_APP_SERVER_ARGS = ["-s", "read-only", "-a", "untrusted", "app-server"] as const; type CodexProbeOptions = { installed: boolean; @@ -48,91 +45,17 @@ function nonEmptyText(value: unknown): string | null { return typeof value === "string" && value.length > 0 ? value : null; } -function resetTime(value: unknown): string | null { - const seconds = finiteNumber(value); - if (seconds !== null) return new Date(seconds * 1000).toISOString(); - if (typeof value === "string" && !Number.isNaN(Date.parse(value))) - return new Date(value).toISOString(); - return null; -} - -function safeIdentifier(value: unknown): string | null { - const candidate = nonEmptyText(value); - if (!candidate || !/^[a-z][a-z0-9_-]{0,63}$/i.test(candidate)) return null; - if (/(?:account|token|secret|key)[_-]?\d{6,}/i.test(candidate)) return null; - return candidate; -} - -export function resolveCodexAuthPath( - env: NodeJS.ProcessEnv = process.env, -): string { +export function resolveCodexAuthPath(env: NodeJS.ProcessEnv = process.env): string { const root = env.CODEX_HOME || join(env.HOME || "", ".codex"); return join(root, "auth.json"); } -export function toRateWindow( - value: unknown, - id: string, - label: string, -): CapacityWindow | null { - const input = record(value); - if (!input) return null; - const used = finiteNumber(input.used_percent); - const seconds = finiteNumber(input.limit_window_seconds); - return { - id, - label, - durationMinutes: seconds === null ? null : seconds / 60, - usedPercent: used, - resetsAt: resetTime(input.reset_at), - }; -} - -function extraWindows(value: unknown): CapacityWindow[] { - if (!Array.isArray(value)) return []; - return value.flatMap((entry, index) => { - const limit = record(entry); - if (!limit) return []; - const scope = safeIdentifier(limit.limit_name) ?? `extra-${index + 1}`; - const windows = record(limit.rate_limit) ?? limit; - return [ - toRateWindow( - windows.primary_window, - `${scope}:primary`, - `${scope} primary`, - ), - toRateWindow( - windows.secondary_window, - `${scope}:secondary`, - `${scope} secondary`, - ), - ].filter((window): window is CapacityWindow => window !== null); - }); +export function parseUsage(raw: unknown, source: "pat" | "oauth"): UsageSnapshot { + const { windows, creditsRemaining } = parseWhamUsage(raw); + return { windows, creditsRemaining, source }; } -export function parseUsage( - raw: unknown, - source: "pat" | "oauth", -): UsageSnapshot { - const response = record(raw) ?? {}; - const limits = record(response.rate_limit) ?? {}; - const credits = record(response.credits) ?? {}; - return { - windows: [ - toRateWindow(limits.primary_window, "session", "Session"), - toRateWindow(limits.secondary_window, "weekly", "Weekly"), - ...extraWindows(response.additional_rate_limits), - ].filter((window): window is CapacityWindow => window !== null), - creditsRemaining: finiteNumber(credits.balance), - source, - }; -} - -function cliWindow( - value: unknown, - id: string, - label: string, -): CapacityWindow | null { +function cliWindow(value: unknown, id: string, label: string): CapacityWindow | null { const input = record(value); if (!input) return null; return { @@ -144,14 +67,10 @@ function cliWindow( }; } -function cliSnapshotWindows( - value: unknown, - fallbackId: string, -): CapacityWindow[] { +function cliSnapshotWindows(value: unknown, fallbackId: string): CapacityWindow[] { const snapshot = record(value); if (!snapshot) return []; - const scope = - safeIdentifier(snapshot.limitId) ?? safeIdentifier(fallbackId) ?? "codex"; + const scope = safeIdentifier(snapshot.limitId) ?? safeIdentifier(fallbackId) ?? "codex"; return [ cliWindow(snapshot.primary, `${scope}:primary`, `${scope} primary`), cliWindow(snapshot.secondary, `${scope}:secondary`, `${scope} secondary`), @@ -167,9 +86,7 @@ export function parseCliUsage(raw: unknown): UsageSnapshot { for (const [id, snapshot] of Object.entries(buckets)) windows.push(...cliSnapshotWindows(snapshot, id)); } - const unique = [ - ...new Map(windows.map((window) => [window.id, window])).values(), - ]; + const unique = [...new Map(windows.map((window) => [window.id, window])).values()]; return { windows: unique, creditsRemaining: null, source: "cli" }; } @@ -178,14 +95,11 @@ function capacityFromSnapshot( context: CodexProbeOptions, raw?: unknown, ): CapacityReport { - const hasUsage = snapshot.windows.some( - (window) => window.usedPercent !== null, - ); + const hasUsage = snapshot.windows.some((window) => window.usedPercent !== null); const rateLimits = record(record(raw)?.rateLimits); const reached = nonEmptyText(rateLimits?.rateLimitReachedType); const resetCredits = - record(record(raw)?.rateLimitResetCredits) ?? - record(record(raw)?.usageLimitResetCredits); + record(record(raw)?.rateLimitResetCredits) ?? record(record(raw)?.usageLimitResetCredits); return { harness: "codex", provider: "openai", @@ -193,8 +107,7 @@ function capacityFromSnapshot( authenticated: true, available: reached ? "no" : hasUsage ? "yes" : "unknown", windows: snapshot.windows, - creditsRemaining: - snapshot.creditsRemaining ?? finiteNumber(resetCredits?.availableCount), + creditsRemaining: snapshot.creditsRemaining ?? finiteNumber(resetCredits?.availableCount), }; } @@ -202,9 +115,7 @@ function jwtExpiry(token: string): number | null { const part = token.split(".")[1]; if (!part) return null; try { - return finiteNumber( - record(JSON.parse(Buffer.from(part, "base64url").toString("utf8")))?.exp, - ); + return finiteNumber(record(JSON.parse(Buffer.from(part, "base64url").toString("utf8")))?.exp); } catch { return null; } @@ -231,10 +142,7 @@ async function fetchJson( const timer = setTimeout(() => controller.abort(), timeoutMs); try { const response = await fetcher(url, { ...init, signal: controller.signal }); - if (!response.ok) - throw new Error( - response.status === 401 ? "unauthorized" : "request failed", - ); + if (!response.ok) throw new Error(response.status === 401 ? "unauthorized" : "request failed"); return await response.json(); } finally { clearTimeout(timer); @@ -262,10 +170,7 @@ async function apiSnapshot( return parseUsage(raw, source); } -function appServerRpc( - messages: RpcMessage[], - timeoutMs = 5000, -): Promise { +function appServerRpc(messages: RpcMessage[], timeoutMs = 5000): Promise { return new Promise((resolve, reject) => { const child = spawn("codex", CODEX_APP_SERVER_ARGS, { stdio: ["pipe", "pipe", "ignore"], @@ -281,13 +186,8 @@ function appServerRpc( if (error) reject(error); else resolve(results as CliResponses); }; - const timer = setTimeout( - () => finish(new Error("codex probe timed out")), - timeoutMs, - ); - child.once("error", () => - finish(new Error("codex app-server unavailable")), - ); + const timer = setTimeout(() => finish(new Error("codex probe timed out")), timeoutMs); + child.once("error", () => finish(new Error("codex app-server unavailable"))); child.once("exit", () => { if (!settled) finish(new Error("codex app-server exited")); }); @@ -310,8 +210,7 @@ function appServerRpc( for (const request of messages.slice(1)) child.stdin.write(`${JSON.stringify(request)}\n`); } else if (message.id === 2) { - if (message.error) - finish(new Error("codex rate-limit method failed")); + if (message.error) finish(new Error("codex rate-limit method failed")); else results.rateLimits = message.result; } else if (message.id === 3) { if (message.error) finish(new Error("codex account method failed")); @@ -340,9 +239,7 @@ function unavailable(options: CodexProbeOptions): CapacityReport { return codexUnavailableReport(options.checkedAt); } -async function cliFallback( - options: CodexProbeOptions, -): Promise { +async function cliFallback(options: CodexProbeOptions): Promise { if (!options.installed) return unavailable(options); const messages: RpcMessage[] = [ { @@ -358,8 +255,7 @@ async function cliFallback( { id: 3, method: "account/read" }, ]; try { - const rpc = - options.rpc ?? ((requests) => appServerRpc(requests, options.timeoutMs)); + const rpc = options.rpc ?? ((requests) => appServerRpc(requests, options.timeoutMs)); const response = await rpc(messages); const result = capacityFromSnapshot( parseCliUsage(response.rateLimits), @@ -381,9 +277,7 @@ async function cliFallback( } } -export async function probeCodexCapacity( - options: CodexProbeOptions, -): Promise { +export async function probeCodexCapacity(options: CodexProbeOptions): Promise { let parsed: UnknownRecord | null = null; try { const contents = await (options.readFile ?? readFile)( @@ -410,10 +304,7 @@ export async function probeCodexCapacity( ); const accountId = nonEmptyText(whoami?.chatgpt_account_id); if (!accountId) throw new Error("account unavailable"); - return capacityFromSnapshot( - await apiSnapshot(pat, accountId, "pat", options), - options, - ); + return capacityFromSnapshot(await apiSnapshot(pat, accountId, "pat", options), options); } catch { // Continue to a separately available OAuth credential before using the CLI. } diff --git a/packages/agent-manager/src/capacity/wham.ts b/packages/agent-manager/src/capacity/wham.ts new file mode 100644 index 00000000..b36d7566 --- /dev/null +++ b/packages/agent-manager/src/capacity/wham.ts @@ -0,0 +1,84 @@ +import type { CapacityWindow } from "./types.js"; + +type WhamUsageSnapshot = { + windows: CapacityWindow[]; + creditsRemaining: number | null; +}; + +type UnknownRecord = Record; + +function record(value: unknown): UnknownRecord | null { + return value !== null && typeof value === "object" && !Array.isArray(value) + ? (value as UnknownRecord) + : null; +} + +function finiteNumber(value: unknown): number | null { + return typeof value === "number" && Number.isFinite(value) ? value : null; +} + +function nonEmptyText(value: unknown): string | null { + return typeof value === "string" && value.length > 0 ? value : null; +} + +export function resetTime(value: unknown): string | null { + const seconds = finiteNumber(value); + if (seconds !== null) return new Date(seconds * 1000).toISOString(); + if (typeof value === "string" && !Number.isNaN(Date.parse(value))) + return new Date(value).toISOString(); + return null; +} + +export function safeIdentifier(value: unknown): string | null { + const candidate = nonEmptyText(value); + if (!candidate || !/^[a-z][a-z0-9_-]{0,63}$/i.test(candidate)) return null; + if (/(?:account|token|secret|key)[_-]?\d{6,}/i.test(candidate)) return null; + return candidate; +} + +export function toRateWindow(value: unknown, id: string, label: string): CapacityWindow | null { + const input = record(value); + if (!input) return null; + const used = finiteNumber(input.used_percent); + const seconds = finiteNumber(input.limit_window_seconds); + return { + id, + label, + durationMinutes: seconds === null ? null : seconds / 60, + usedPercent: used, + resetsAt: resetTime(input.reset_at), + }; +} + +function extraWindows(value: unknown): CapacityWindow[] { + if (!Array.isArray(value)) return []; + return value.flatMap((entry, index) => { + const limit = record(entry); + if (!limit) return []; + const scope = safeIdentifier(limit.limit_name) ?? `extra-${index + 1}`; + const windows = record(limit.rate_limit) ?? limit; + return [ + toRateWindow(windows.primary_window, `${scope}:primary`, `${scope} primary`), + toRateWindow(windows.secondary_window, `${scope}:secondary`, `${scope} secondary`), + ].filter((window): window is CapacityWindow => window !== null); + }); +} + +/** + * Parses the chatgpt.com backend wham/usage payload shared by every + * OAuth-backed OpenAI consumer (codex harness and pi login) into generic + * capacity windows plus the credit balance. + */ +export function parseWhamUsage(raw: unknown): WhamUsageSnapshot { + const response = record(raw) ?? {}; + const limits = record(response.rate_limit) ?? {}; + const credits = record(response.credits) ?? {}; + return { + windows: [ + toRateWindow(limits.primary_window, "session", "Session"), + toRateWindow(limits.secondary_window, "weekly", "Weekly"), + ...extraWindows(response.additional_rate_limits), + ].filter((window): window is CapacityWindow => window !== null), + creditsRemaining: finiteNumber(credits.balance), + }; +} From c24fa69e988cd38dea4678a0f96a55ac9555de2e Mon Sep 17 00:00:00 2001 From: Hoang Nguyen Date: Sun, 27 Sep 2026 09:01:41 +0000 Subject: [PATCH 05/10] feat(capacity): show OpenAI usage from pi OAuth login --- .../capacity/fixtures/openai-codex-auth.json | 9 + .../src/__tests__/capacity/openai.test.ts | 262 +++++++++++++++++- packages/agent-manager/src/capacity/openai.ts | 162 ++++++++++- 3 files changed, 418 insertions(+), 15 deletions(-) create mode 100644 packages/agent-manager/src/__tests__/capacity/fixtures/openai-codex-auth.json diff --git a/packages/agent-manager/src/__tests__/capacity/fixtures/openai-codex-auth.json b/packages/agent-manager/src/__tests__/capacity/fixtures/openai-codex-auth.json new file mode 100644 index 00000000..c01b4f3f --- /dev/null +++ b/packages/agent-manager/src/__tests__/capacity/fixtures/openai-codex-auth.json @@ -0,0 +1,9 @@ +{ + "openai-codex": { + "type": "oauth", + "access": "fixture-header.eyJleHAiOjE3OTMwMDAwMDB9.fixture-signature", + "refresh": "fixture-refresh-token", + "expires": 1793000000000, + "accountId": "fixture-chatgpt-account-id" + } +} diff --git a/packages/agent-manager/src/__tests__/capacity/openai.test.ts b/packages/agent-manager/src/__tests__/capacity/openai.test.ts index 64f0bd40..710dca7d 100644 --- a/packages/agent-manager/src/__tests__/capacity/openai.test.ts +++ b/packages/agent-manager/src/__tests__/capacity/openai.test.ts @@ -1,4 +1,5 @@ import { readFile } from "node:fs/promises"; +import { readFileSync } from "node:fs"; import { fileURLToPath } from "node:url"; import { describe, expect, it, vi } from "vitest"; import { @@ -6,11 +7,14 @@ import { parseOpenAiUsage, probeOpenAiCapacity, resolveOpenAiApiKey, + resolveOpenAiCredential, } from "../../capacity/openai.js"; const checkedAt = "2026-09-26T10:00:00.000Z"; const fixture = (name: string) => readFile(fileURLToPath(new URL(`./fixtures/${name}`, import.meta.url)), "utf8"); +const fixtureString = (name: string) => + readFileSync(fileURLToPath(new URL(`./fixtures/${name}`, import.meta.url)), "utf8"); describe("OpenAI credential resolution", () => { it("prefers OPENAI_API_KEY without reading Pi auth", async () => { @@ -46,7 +50,263 @@ describe("OpenAI credential resolution", () => { env: { HOME: "/users/test" }, readFile: readAuth, }), - ).rejects.toThrow(/OpenAI API key|OpenAI Pi auth file/); + ).rejects.toThrow(/OpenAI credentials not found|OpenAI Pi auth file/); + }); +}); + +describe("OpenAI tiered credential resolution", () => { + const oauthFixture = () => fixture("openai-codex-auth.json"); + const bothCredentials = JSON.stringify({ + openai: { type: "api_key", key: "pi-openai-test-key" }, + "openai-codex": { + type: "oauth", + access: "fixture-header.eyJleHAiOjE3OTMwMDAwMDB9.fixture-signature", + refresh: "fixture-refresh-token", + expires: 1793000000000, + accountId: "fixture-chatgpt-account-id", + }, + }); + + it("prefers the Pi openai platform credential over the openai-codex OAuth login", async () => { + await expect( + resolveOpenAiCredential({ + env: { HOME: "/users/test" }, + readFile: async () => bothCredentials, + }), + ).resolves.toEqual({ kind: "platform", key: "pi-openai-test-key" }); + }); + + it("resolves the Pi openai-codex OAuth login when no platform credential exists", async () => { + await expect( + resolveOpenAiCredential({ + env: { HOME: "/users/test" }, + readFile: async () => oauthFixture(), + }), + ).resolves.toEqual({ + kind: "oauth", + access: "fixture-header.eyJleHAiOjE3OTMwMDAwMDB9.fixture-signature", + accountId: "fixture-chatgpt-account-id", + expiresMs: 1793000000000, + }); + }); + + it("normalizes second-based expires values to epoch milliseconds", async () => { + const auth = JSON.stringify({ + "openai-codex": { + type: "oauth", + access: "fixture-access", + refresh: "fixture-refresh", + expires: 1793000000, + accountId: "fixture-chatgpt-account-id", + }, + }); + await expect( + resolveOpenAiCredential({ env: { HOME: "/users/test" }, readFile: async () => auth }), + ).resolves.toMatchObject({ kind: "oauth", expiresMs: 1793000000000 }); + }); + + it("keeps a null expiry for OAuth logins without usable expires metadata", async () => { + const auth = JSON.stringify({ + "openai-codex": { + type: "oauth", + access: "fixture-access", + accountId: "fixture-chatgpt-account-id", + }, + }); + await expect( + resolveOpenAiCredential({ env: { HOME: "/users/test" }, readFile: async () => auth }), + ).resolves.toMatchObject({ kind: "oauth", expiresMs: null }); + }); + + it.each([ + ["wrong credential type", JSON.stringify({ "openai-codex": { type: "api_key", key: "x" } })], + ["missing access token", JSON.stringify({ "openai-codex": { type: "oauth", accountId: "a" } })], + ["missing account id", JSON.stringify({ "openai-codex": { type: "oauth", access: "t" } })], + ])( + "skips a malformed openai-codex credential (%s) with a sanitized error", + async (_name, auth) => { + await expect( + resolveOpenAiCredential({ env: { HOME: "/users/test" }, readFile: async () => auth }), + ).rejects.toThrow("OpenAI credentials not found"); + }, + ); + + it("still throws for OAuth-only machines on the legacy platform-key resolver", async () => { + await expect( + resolveOpenAiApiKey({ + env: { HOME: "/users/test" }, + readFile: async () => oauthFixture(), + }), + ).rejects.toThrow("OpenAI credentials not found"); + }); +}); + +describe("OpenAI OAuth capacity probe", () => { + const freshOauth = () => fixture("openai-codex-auth.json"); + const whamResponse = () => new Response(fixtureString("wham-usage.json"), { status: 200 }); + + it("reports wham usage windows for a fresh Pi OAuth login", async () => { + const fetch = vi.fn(async () => whamResponse()); + const report = await probeOpenAiCapacity({ + checkedAt, + env: { HOME: "/users/test" }, + readFile: async () => freshOauth(), + fetch, + }); + expect(fetch).toHaveBeenCalledWith( + "https://chatgpt.com/backend-api/wham/usage", + expect.objectContaining({ + method: "GET", + headers: { + Authorization: "Bearer fixture-header.eyJleHAiOjE3OTMwMDAwMDB9.fixture-signature", + "ChatGPT-Account-Id": "fixture-chatgpt-account-id", + }, + }), + ); + expect(report).toMatchObject({ + harness: "pi", + provider: "openai", + generatedAt: checkedAt, + authenticated: true, + available: "yes", + creditsRemaining: 4.5, + }); + expect(report.windows.map((window) => window.id)).toEqual(["session", "weekly"]); + expect(JSON.stringify(report)).not.toContain("fixture-signature"); + }); + + it("does not fetch and reports an authenticated-but-stale state for expired logins", async () => { + const staleAuth = JSON.stringify({ + "openai-codex": { + type: "oauth", + access: "stale-fixture-access", + expires: 1790000000000, + accountId: "fixture-chatgpt-account-id", + }, + }); + const fetch = vi.fn(); + const report = await probeOpenAiCapacity({ + checkedAt, + env: { HOME: "/users/test" }, + readFile: async () => staleAuth, + fetch, + }); + expect(fetch).not.toHaveBeenCalled(); + expect(report).toEqual({ + harness: "pi", + provider: "openai", + generatedAt: checkedAt, + authenticated: true, + available: "unknown", + windows: [], + creditsRemaining: null, + }); + }); + + it("treats second-based expiry metadata as stale the same way", async () => { + const staleAuth = JSON.stringify({ + "openai-codex": { + type: "oauth", + access: "stale-fixture-access", + expires: 1790000000, + accountId: "fixture-chatgpt-account-id", + }, + }); + const fetch = vi.fn(); + const report = await probeOpenAiCapacity({ + checkedAt, + env: { HOME: "/users/test" }, + readFile: async () => staleAuth, + fetch, + }); + expect(fetch).not.toHaveBeenCalled(); + expect(report.available).toBe("unknown"); + }); + + it("falls back to the JWT exp when expires metadata is missing", async () => { + const expiredJwtAuth = JSON.stringify({ + "openai-codex": { + type: "oauth", + access: "h.eyJleHAiOjE3ODkwMDAwMDB9.s", + accountId: "fixture-chatgpt-account-id", + }, + }); + const noFetch = vi.fn(); + await expect( + probeOpenAiCapacity({ + checkedAt, + env: { HOME: "/users/test" }, + readFile: async () => expiredJwtAuth, + fetch: noFetch, + }), + ).resolves.toMatchObject({ authenticated: true, available: "unknown" }); + expect(noFetch).not.toHaveBeenCalled(); + + const futureJwtAuth = JSON.stringify({ + "openai-codex": { + type: "oauth", + access: "h.eyJleHAiOjE3OTMwMDAwMDB9.s", + accountId: "fixture-chatgpt-account-id", + }, + }); + const fetch = vi.fn(async () => whamResponse()); + await expect( + probeOpenAiCapacity({ + checkedAt, + env: { HOME: "/users/test" }, + readFile: async () => futureJwtAuth, + fetch, + }), + ).resolves.toMatchObject({ authenticated: true, available: "yes" }); + expect(fetch).toHaveBeenCalledOnce(); + }); + + it("reports unauthenticated when wham rejects the OAuth token", async () => { + for (const status of [401, 403]) { + const report = await probeOpenAiCapacity({ + checkedAt, + env: { HOME: "/users/test" }, + readFile: async () => freshOauth(), + fetch: vi.fn(async () => new Response("denied", { status })), + }); + expect(report).toMatchObject({ + authenticated: false, + available: "unknown", + windows: [], + creditsRemaining: null, + }); + } + }); + + it("throws sanitized errors for transport failures without leaking tokens", async () => { + const probe = (fetch: unknown) => + probeOpenAiCapacity({ + checkedAt, + env: { HOME: "/users/test" }, + readFile: async () => freshOauth(), + fetch: fetch as typeof globalThis.fetch, + timeoutMs: 10, + }); + await expect( + probe( + (_url: string, init: RequestInit) => + new Promise((_resolve, reject) => + init.signal?.addEventListener("abort", () => reject(new Error("aborted"))), + ), + ), + ).rejects.toThrow("OpenAI usage request failed"); + await expect(probe(vi.fn(async () => new Response("boom", { status: 503 })))).rejects.toThrow( + "OpenAI usage request failed: HTTP 503", + ); + await expect( + probe(vi.fn(async () => new Response("not-json", { status: 200 }))), + ).rejects.toThrow("OpenAI usage response is not valid JSON"); + try { + await probe(vi.fn(async () => new Response("boom", { status: 503 }))); + } catch (error) { + expect(String(error)).not.toContain("fixture-signature"); + expect(String(error)).not.toContain("fixture-chatgpt-account-id"); + } }); }); diff --git a/packages/agent-manager/src/capacity/openai.ts b/packages/agent-manager/src/capacity/openai.ts index 59d4f279..cc5f57ce 100644 --- a/packages/agent-manager/src/capacity/openai.ts +++ b/packages/agent-manager/src/capacity/openai.ts @@ -1,19 +1,30 @@ import { readFile } from "node:fs/promises"; import { homedir } from "node:os"; import { join } from "node:path"; +import { parseWhamUsage } from "./wham.js"; import type { CapacityReport, CapacityWindow } from "./types.js"; const OPENAI_USAGE_BASE_URL = "https://api.openai.com/v1/organization/usage"; const OPENAI_MODELS_URL = "https://api.openai.com/v1/models"; +const WHAM_USAGE_URL = "https://chatgpt.com/backend-api/wham/usage"; const OPENAI_USAGE_ENDPOINTS = ["completions", "responses"] as const; const USAGE_WINDOW_DAYS = 7; const DAY_MS = 24 * 60 * 60 * 1000; const MISSING_API_KEY_MESSAGE = - "OpenAI API key not found; set OPENAI_API_KEY or configure Pi provider openai"; + "OpenAI credentials not found; set OPENAI_API_KEY, configure the Pi openai provider, or log in to OpenAI via Pi"; +const EPOCH_MS_FLOOR = 1e12; export type OpenAiUsageBucket = { startMs: number; tokens: number }; type UnknownRecord = Record; +/** + * Resolved OpenAI credential: a platform API key (env or Pi `openai` entry) + * or the OAuth login Pi stores as `openai-codex` for ChatGPT-backed access. + */ +export type OpenAiCredential = + | { kind: "platform"; key: string } + | { kind: "oauth"; access: string; accountId: string; expiresMs: number | null }; + export type OpenAiCredentialOptions = { env?: NodeJS.ProcessEnv; readFile?: (path: string, encoding: BufferEncoding) => Promise; @@ -25,6 +36,7 @@ export type OpenAiProbeOptions = OpenAiCredentialOptions & { timeoutMs?: number; }; type OpenAiRequestContext = OpenAiProbeOptions & { key: string }; +type OAuthCredential = Extract; function record(value: unknown): UnknownRecord | null { return value !== null && typeof value === "object" && !Array.isArray(value) @@ -40,11 +52,15 @@ function finiteNumber(value: unknown): number | null { return typeof value === "number" && Number.isFinite(value) ? value : null; } -export async function resolveOpenAiApiKey(options: OpenAiCredentialOptions = {}): Promise { - const env = options.env ?? process.env; - const environmentKey = nonEmptyText(env.OPENAI_API_KEY); - if (environmentKey) return environmentKey; +/** Pi stores `expires` in epoch milliseconds, codex-style tokens use seconds. */ +function toEpochMs(value: unknown): number | null { + const raw = finiteNumber(value); + if (raw === null || raw <= 0) return null; + return raw > EPOCH_MS_FLOOR ? raw : raw * 1000; +} +async function readPiAuthRoot(options: OpenAiCredentialOptions): Promise { + const env = options.env ?? process.env; const authPath = join(env.HOME || homedir(), ".pi", "agent", "auth.json"); let parsed: UnknownRecord; try { @@ -58,16 +74,45 @@ export async function resolveOpenAiApiKey(options: OpenAiCredentialOptions = {}) } throw new Error("OpenAI Pi auth file is malformed"); } + return parsed; +} + +/** + * Resolves the OpenAI credential in priority order: `OPENAI_API_KEY` env, + * then the Pi `openai` platform-key entry, then the Pi `openai-codex` OAuth + * login written by `pi login` (no manual setup required). + */ +export async function resolveOpenAiCredential( + options: OpenAiCredentialOptions = {}, +): Promise { + const env = options.env ?? process.env; + const environmentKey = nonEmptyText(env.OPENAI_API_KEY); + if (environmentKey) return { kind: "platform", key: environmentKey }; - const credential = record(parsed.openai); - if (!credential) { - throw new Error(MISSING_API_KEY_MESSAGE); + const parsed = await readPiAuthRoot(options); + + const platform = record(parsed.openai); + if (platform) { + const key = platform.type === "api_key" ? nonEmptyText(platform.key) : null; + if (!key) { + throw new Error("OpenAI Pi auth file has a malformed openai credential"); + } + return { kind: "platform", key }; } - const key = credential.type === "api_key" ? nonEmptyText(credential.key) : null; - if (!key) { - throw new Error("OpenAI Pi auth file has a malformed openai credential"); + + const oauth = record(parsed["openai-codex"]); + const access = oauth?.type === "oauth" ? nonEmptyText(oauth.access) : null; + const accountId = nonEmptyText(oauth?.accountId); + if (oauth && access && accountId) { + return { kind: "oauth", access, accountId, expiresMs: toEpochMs(oauth.expires) }; } - return key; + throw new Error(MISSING_API_KEY_MESSAGE); +} + +export async function resolveOpenAiApiKey(options: OpenAiCredentialOptions = {}): Promise { + const credential = await resolveOpenAiCredential(options); + if (credential.kind === "platform") return credential.key; + throw new Error(MISSING_API_KEY_MESSAGE); } function bucketTokens(value: unknown): number { @@ -212,8 +257,11 @@ async function verifyKey(context: OpenAiRequestContext): Promise { } export async function probeOpenAiCapacity(options: OpenAiProbeOptions): Promise { - const key = await resolveOpenAiApiKey(options); - const probe: OpenAiProbeOptions & { key: string } = { ...options, key }; + const credential = await resolveOpenAiCredential(options); + if (credential.kind === "oauth") { + return probeOpenAiOauthCapacity(credential, options); + } + const probe: OpenAiProbeOptions & { key: string } = { ...options, key: credential.key }; const result = await fetchUsage(probe); if (result === null) { return openAiCapacityReport([], options.checkedAt, false); @@ -225,3 +273,89 @@ export async function probeOpenAiCapacity(options: OpenAiProbeOptions): Promise< const buckets = result.flatMap(parseOpenAiUsage); return buildOpenAiReport(buckets, options.checkedAt); } + +/** Mirrors codex.ts: JWT `exp` (epoch seconds) as expiry fallback. */ +function jwtExpiryMs(token: string): number | null { + const part = token.split(".")[1]; + if (!part) return null; + try { + const exp = finiteNumber( + record(JSON.parse(Buffer.from(part, "base64url").toString("utf8")))?.exp, + ); + return exp === null ? null : exp * 1000; + } catch { + return null; + } +} + +function oauthExpired(credential: OAuthCredential, nowMs: number): boolean { + const expiryMs = credential.expiresMs ?? jwtExpiryMs(credential.access); + return expiryMs !== null && expiryMs <= nowMs; +} + +async function fetchWhamUsage( + credential: OAuthCredential, + options: OpenAiProbeOptions, +): Promise { + const fetcher = options.fetch ?? globalThis.fetch; + const controller = new AbortController(); + const timer = setTimeout(() => controller.abort(), options.timeoutMs ?? 5000); + let response: Response; + try { + response = await fetcher(WHAM_USAGE_URL, { + method: "GET", + headers: { + Authorization: `Bearer ${credential.access}`, + "ChatGPT-Account-Id": credential.accountId, + }, + signal: controller.signal, + }); + } catch { + throw new Error("OpenAI usage request failed"); + } finally { + clearTimeout(timer); + } + if (response.status === 401 || response.status === 403) return null; + if (response.status !== 200) { + throw new Error(`OpenAI usage request failed: HTTP ${response.status}`); + } + try { + return await response.json(); + } catch { + throw new Error("OpenAI usage response is not valid JSON"); + } +} + +/** + * OAuth path for Pi's OpenAI login: consumes the same wham/usage endpoint as + * the codex harness. Stale logins are never fetched (no refresh, matching + * codex.ts); they render as authenticated-but-stale instead. + */ +async function probeOpenAiOauthCapacity( + credential: OAuthCredential, + options: OpenAiProbeOptions, +): Promise { + const staleReport: CapacityReport = { + harness: "pi", + provider: "openai", + generatedAt: options.checkedAt, + authenticated: true, + available: "unknown", + windows: [], + creditsRemaining: null, + }; + const nowMs = Date.parse(options.checkedAt); + if (Number.isNaN(nowMs) || oauthExpired(credential, nowMs)) return staleReport; + const raw = await fetchWhamUsage(credential, options); + if (raw === null) { + return { ...staleReport, authenticated: false }; + } + const snapshot = parseWhamUsage(raw); + const hasUsage = snapshot.windows.some((window) => window.usedPercent !== null); + return { + ...staleReport, + available: hasUsage ? "yes" : "unknown", + windows: snapshot.windows, + creditsRemaining: snapshot.creditsRemaining, + }; +} From 32dfb29047e2d8c3c3e1599896c1176626f064ba Mon Sep 17 00:00:00 2001 From: Hoang Nguyen Date: Sun, 27 Sep 2026 09:05:06 +0000 Subject: [PATCH 06/10] test(capacity): cover wham and oauth branch fallbacks --- .../src/__tests__/capacity/codex.test.ts | 41 +++-------- .../src/__tests__/capacity/openai.test.ts | 73 +++++++++++++++++++ .../src/__tests__/capacity/wham.test.ts | 24 ++++++ .../src/__tests__/capacity/zai.test.ts | 61 ++++------------ 4 files changed, 121 insertions(+), 78 deletions(-) diff --git a/packages/agent-manager/src/__tests__/capacity/codex.test.ts b/packages/agent-manager/src/__tests__/capacity/codex.test.ts index 22f023e7..d06b5eb8 100644 --- a/packages/agent-manager/src/__tests__/capacity/codex.test.ts +++ b/packages/agent-manager/src/__tests__/capacity/codex.test.ts @@ -52,9 +52,7 @@ describe("Codex auth resolution", () => { }); it("falls back to ~/.codex/auth.json", () => { - expect(resolveCodexAuthPath({ HOME: "/users/test" })).toBe( - "/users/test/.codex/auth.json", - ); + expect(resolveCodexAuthPath({ HOME: "/users/test" })).toBe("/users/test/.codex/auth.json"); }); }); @@ -108,9 +106,7 @@ describe("tiered Codex probing", () => { status: 200, }), ) - .mockResolvedValueOnce( - new Response(JSON.stringify(apiUsage()), { status: 200 }), - ); + .mockResolvedValueOnce(new Response(JSON.stringify(apiUsage()), { status: 200 })); const rpc = vi.fn(); const result = await probeCodexCapacity({ ...context, @@ -129,9 +125,7 @@ describe("tiered Codex probing", () => { expect(fetch.mock.calls[0][0]).toBe( "https://auth.openai.com/api/accounts/v1/user-auth-credential/whoami", ); - expect(fetch.mock.calls[1][0]).toBe( - "https://chatgpt.com/backend-api/wham/usage", - ); + expect(fetch.mock.calls[1][0]).toBe("https://chatgpt.com/backend-api/wham/usage"); expect(fetch.mock.calls[1][1].headers).toMatchObject({ Authorization: "Bearer pat-secret", "ChatGPT-Account-Id": "acct-1", @@ -154,9 +148,7 @@ describe("tiered Codex probing", () => { it("selects a fresh OAuth token without calling whoami", async () => { const fetch = vi .fn() - .mockResolvedValue( - new Response(JSON.stringify(apiUsage()), { status: 200 }), - ); + .mockResolvedValue(new Response(JSON.stringify(apiUsage()), { status: 200 })); const result = await probeCodexCapacity({ ...context, readFile: async () => @@ -210,9 +202,7 @@ describe("tiered Codex probing", () => { ])("falls back to the CLI for %s", async (name, readFile) => { const fetch = vi .fn() - .mockResolvedValue( - new Response("", { status: name === "OAuth 401" ? 401 : 200 }), - ); + .mockResolvedValue(new Response("", { status: name === "OAuth 401" ? 401 : 200 })); const rpc = vi.fn(async () => ({ rateLimits: { rateLimits: { @@ -239,8 +229,7 @@ describe("tiered Codex probing", () => { })); const result = await probeCodexCapacity({ ...context, - readFile: async () => - JSON.stringify({ personal_access_token: "pat-secret" }), + readFile: async () => JSON.stringify({ personal_access_token: "pat-secret" }), fetch: vi.fn().mockRejectedValue(new Error("network failure pat-secret")), rpc, }); @@ -252,9 +241,7 @@ describe("tiered Codex probing", () => { const fetch = vi .fn() .mockRejectedValueOnce(new Error("PAT failed")) - .mockResolvedValueOnce( - new Response(JSON.stringify(apiUsage()), { status: 200 }), - ); + .mockResolvedValueOnce(new Response(JSON.stringify(apiUsage()), { status: 200 })); const rpc = vi.fn(); const result = await probeCodexCapacity({ ...context, @@ -294,13 +281,7 @@ describe("tiered Codex probing", () => { "account/read", ]); expect(JSON.stringify(messages)).not.toMatch(/prompt|turn\/start/); - expect(CODEX_APP_SERVER_ARGS).toEqual([ - "-s", - "read-only", - "-a", - "untrusted", - "app-server", - ]); + expect(CODEX_APP_SERVER_ARGS).toEqual(["-s", "read-only", "-a", "untrusted", "app-server"]); }); it("uses account/read to distinguish logged-out CLI state", async () => { @@ -316,11 +297,7 @@ describe("tiered Codex probing", () => { }); it("never exposes tokens or raw auth content through failures", async () => { - const secrets = [ - "pat-secret-value", - "oauth-secret-value", - "refresh-secret-value", - ]; + const secrets = ["pat-secret-value", "oauth-secret-value", "refresh-secret-value"]; const result = await probeCodexCapacity({ ...context, readFile: async () => diff --git a/packages/agent-manager/src/__tests__/capacity/openai.test.ts b/packages/agent-manager/src/__tests__/capacity/openai.test.ts index 710dca7d..e9f5eaa0 100644 --- a/packages/agent-manager/src/__tests__/capacity/openai.test.ts +++ b/packages/agent-manager/src/__tests__/capacity/openai.test.ts @@ -261,6 +261,58 @@ describe("OpenAI OAuth capacity probe", () => { expect(fetch).toHaveBeenCalledOnce(); }); + it("treats unparseable JWT payloads as unknown expiry and probes anyway", async () => { + const auth = JSON.stringify({ + "openai-codex": { + type: "oauth", + access: "fixture-header.bm90LWpzb24.fixture-signature", + accountId: "fixture-chatgpt-account-id", + }, + }); + const fetch = vi.fn(async () => whamResponse()); + await expect( + probeOpenAiCapacity({ + checkedAt, + env: { HOME: "/users/test" }, + readFile: async () => auth, + fetch, + }), + ).resolves.toMatchObject({ authenticated: true, available: "yes" }); + expect(fetch).toHaveBeenCalledOnce(); + }); + + it("treats non-JWT and exp-less access tokens as unknown expiry and probes anyway", async () => { + for (const access of ["opaque-fixture-access", "h.eyJzdWIiOiJ4In0.s"]) { + const auth = JSON.stringify({ + "openai-codex": { + type: "oauth", + access, + accountId: "fixture-chatgpt-account-id", + }, + }); + const fetch = vi.fn(async () => whamResponse()); + await expect( + probeOpenAiCapacity({ + checkedAt, + env: { HOME: "/users/test" }, + readFile: async () => auth, + fetch, + }), + ).resolves.toMatchObject({ authenticated: true, available: "yes" }); + expect(fetch).toHaveBeenCalledOnce(); + } + }); + + it("reports unknown availability when wham returns no usable windows", async () => { + const report = await probeOpenAiCapacity({ + checkedAt, + env: { HOME: "/users/test" }, + readFile: async () => freshOauth(), + fetch: vi.fn(async () => new Response("{}", { status: 200 })), + }); + expect(report).toMatchObject({ authenticated: true, available: "unknown", windows: [] }); + }); + it("reports unauthenticated when wham rejects the OAuth token", async () => { for (const status of [401, 403]) { const report = await probeOpenAiCapacity({ @@ -478,6 +530,27 @@ describe("OpenAI request", () => { expect(report).toMatchObject({ authenticated: false, windows: [] }); }); + it("sanitizes network failures on the platform key path", async () => { + await expect( + probeOpenAiCapacity({ + checkedAt, + env: { OPENAI_API_KEY: "network-key" }, + fetch: vi.fn(async () => Promise.reject(new Error("ECONNRESET secret"))), + }), + ).rejects.toThrow("OpenAI usage request failed"); + await expect( + probeOpenAiCapacity({ + checkedAt, + env: { OPENAI_API_KEY: "network-key" }, + fetch: vi.fn(async (url: string) => + url.endsWith("/models") + ? Promise.reject(new Error("ECONNRESET secret")) + : new Response("forbidden", { status: 403 }), + ), + }), + ).rejects.toThrow("OpenAI usage request failed"); + }); + it("rejects other statuses and malformed JSON responses with sanitized errors", async () => { await expect( probeOpenAiCapacity({ diff --git a/packages/agent-manager/src/__tests__/capacity/wham.test.ts b/packages/agent-manager/src/__tests__/capacity/wham.test.ts index 5764a4c6..87466b05 100644 --- a/packages/agent-manager/src/__tests__/capacity/wham.test.ts +++ b/packages/agent-manager/src/__tests__/capacity/wham.test.ts @@ -58,6 +58,30 @@ describe("wham usage parsing", () => { expect(parseWhamUsage({ additional_rate_limits: null }).windows).toEqual([]); }); + it("falls back to positional ids and inline windows for degraded extra limits", () => { + const scoped = parseWhamUsage({ + additional_rate_limits: [ + null, + 42, + { primary_window: { used_percent: 5 } }, + { limit_name: "not/safe", primary_window: { used_percent: 6 } }, + { limit_name: "token-123456", primary_window: { used_percent: 7 } }, + { limit_name: "ok", secondary_window: { used_percent: 8 } }, + ], + }); + expect(scoped.windows.map((window) => window.id)).toEqual([ + "extra-3:primary", + "extra-4:primary", + "extra-5:primary", + "ok:secondary", + ]); + }); + + it("parses non-object payloads into an empty snapshot", () => { + expect(parseWhamUsage(null)).toEqual({ windows: [], creditsRemaining: null }); + expect(parseWhamUsage(42)).toEqual({ windows: [], creditsRemaining: null }); + }); + it("never invents usage for windows it cannot parse", () => { const snapshot = parseWhamUsage({ rate_limit: { diff --git a/packages/agent-manager/src/__tests__/capacity/zai.test.ts b/packages/agent-manager/src/__tests__/capacity/zai.test.ts index c72bb7e5..7a23dc6f 100644 --- a/packages/agent-manager/src/__tests__/capacity/zai.test.ts +++ b/packages/agent-manager/src/__tests__/capacity/zai.test.ts @@ -1,18 +1,11 @@ import { readFile } from "node:fs/promises"; import { fileURLToPath } from "node:url"; import { describe, expect, it, vi } from "vitest"; -import { - parseZaiQuota, - probeZaiCapacity, - resolveZaiApiKey, -} from "../../capacity/zai.js"; +import { parseZaiQuota, probeZaiCapacity, resolveZaiApiKey } from "../../capacity/zai.js"; const checkedAt = "2026-08-20T10:00:00.000Z"; const fixture = (name: string) => - readFile( - fileURLToPath(new URL(`./fixtures/${name}`, import.meta.url)), - "utf8", - ); + readFile(fileURLToPath(new URL(`./fixtures/${name}`, import.meta.url)), "utf8"); describe("z.ai credential resolution", () => { it("prefers Z_AI_API_KEY without reading Pi auth", async () => { @@ -31,40 +24,27 @@ describe("z.ai credential resolution", () => { await expect( resolveZaiApiKey({ env: { HOME: "/users/test" }, readFile: readAuth }), ).resolves.toBe("pi-zai-test-key"); - expect(readAuth).toHaveBeenCalledWith( - "/users/test/.pi/agent/auth.json", - "utf8", - ); + expect(readAuth).toHaveBeenCalledWith("/users/test/.pi/agent/auth.json", "utf8"); }); it.each([ [ "missing", - async () => - Promise.reject(Object.assign(new Error("missing"), { code: "ENOENT" })), + async () => Promise.reject(Object.assign(new Error("missing"), { code: "ENOENT" })), ], ["invalid JSON", async () => "{"], - [ - "wrong credential type", - async () => JSON.stringify({ zai: { type: "oauth", key: "x" } }), - ], + ["wrong credential type", async () => JSON.stringify({ zai: { type: "oauth", key: "x" } })], ["missing key", async () => JSON.stringify({ zai: { type: "api_key" } })], - ])( - "fails clearly for %s credentials without exposing file content", - async (_name, readAuth) => { - await expect( - resolveZaiApiKey({ env: { HOME: "/users/test" }, readFile: readAuth }), - ).rejects.toThrow(/z\.ai API key|z\.ai Pi auth file/); - }, - ); + ])("fails clearly for %s credentials without exposing file content", async (_name, readAuth) => { + await expect( + resolveZaiApiKey({ env: { HOME: "/users/test" }, readFile: readAuth }), + ).rejects.toThrow(/z\.ai API key|z\.ai Pi auth file/); + }); }); describe("z.ai quota mapping", () => { it("maps token, credit, and monthly MCP windows with recomputed usage", async () => { - const report = parseZaiQuota( - JSON.parse(await fixture("zai-quota.json")), - checkedAt, - ); + const report = parseZaiQuota(JSON.parse(await fixture("zai-quota.json")), checkedAt); expect(report).toMatchObject({ harness: "pi", provider: "zai", @@ -124,24 +104,13 @@ describe("z.ai quota mapping", () => { }, checkedAt, ); - expect(report.windows.map((window) => window.usedPercent)).toEqual([ - 100, 0, - ]); + expect(report.windows.map((window) => window.usedPercent)).toEqual([100, 0]); }); it.each([ - [ - "invalid fixture", - async () => JSON.parse(await fixture("zai-quota-invalid.json")), - ], - [ - "unsuccessful envelope", - async () => ({ success: false, code: 401, data: { limits: [] } }), - ], - [ - "invalid entry", - async () => ({ success: true, code: 200, data: { limits: [{}] } }), - ], + ["invalid fixture", async () => JSON.parse(await fixture("zai-quota-invalid.json"))], + ["unsuccessful envelope", async () => ({ success: false, code: 401, data: { limits: [] } })], + ["invalid entry", async () => ({ success: true, code: 200, data: { limits: [{}] } })], ])("rejects %s", async (_name, input) => { const value = await input(); expect(() => parseZaiQuota(value, checkedAt)).toThrow(/z\.ai quota/); From dba61f23b955e9837b7aa6dd992d89bdc32edf8c Mon Sep 17 00:00:00 2001 From: Hoang Nguyen Date: Sun, 27 Sep 2026 09:05:56 +0000 Subject: [PATCH 07/10] docs(planning): mark M1-M2 complete --- .../2026-09-27-feature-openai-pi-login-capacity.md | 13 +++++++------ 1 file changed, 7 insertions(+), 6 deletions(-) diff --git a/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md index 20fcd4cf..6742db23 100644 --- a/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md +++ b/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md @@ -8,21 +8,21 @@ description: Break down work into actionable tasks and estimate timeline ## Milestones -- [ ] M1: Shared wham parsing extracted, codex suite green unchanged -- [ ] M2: Pi OAuth tier resolves + classifies (all unit scenarios green) +- [x] M1: Shared wham parsing extracted, codex suite green unchanged +- [x] M2: Pi OAuth tier resolves + classifies (all unit scenarios green) - [ ] M3: Full gates green, docs/implementation updated, final PR open ## Task Breakdown ### Phase 1: Foundation — shared parser -- [ ] T1.1: Create `capacity/wham.ts` with `parseWhamUsage`, `toRateWindow`, `extraWindows` + guards (TDD: add `wham.test.ts` scenarios from testing doc first; fixtures `wham-usage.json`). Outcome: pure parser module; validation: `npx vitest run src/__tests__/capacity/wham.test.ts` (agent-manager). -- [ ] T1.2: Delegate `codex.ts` parsing to `wham.ts`; keep `parseUsage`/`toRateWindow` exports. Outcome: codex behavior identical; validation: `codex.test.ts` passes unchanged (no edits to that suite). +- [x] T1.1: Create `capacity/wham.ts` with `parseWhamUsage`, `toRateWindow`, `extraWindows` + guards (TDD: add `wham.test.ts` scenarios from testing doc first; fixtures `wham-usage.json`). Outcome: pure parser module; validation: `npx vitest run src/__tests__/capacity/wham.test.ts` (agent-manager). +- [x] T1.2: Delegate `codex.ts` parsing to `wham.ts`; keep `parseUsage`/`toRateWindow` exports. Outcome: codex behavior identical; validation: `codex.test.ts` passes unchanged (no edits to that suite). ### Phase 2: Core — OAuth tier in openai.ts -- [ ] T2.1: `resolveOpenAiCredential` tiers (env → pi `openai` api_key → pi `openai-codex` oauth → not-found error mentioning pi login); keep `resolveOpenAiApiKey` compat; ms-vs-s + JWT staleness helper. TDD first: extend `openai.test.ts` resolution + staleness scenarios with fixture `openai-codex-auth.json`. Validation: scoped vitest. -- [ ] T2.2: OAuth probe path: wham GET (Bearer + `ChatGPT-Account-Id`, abort/timeout), classification (200 report / 401-403 unauthenticated / timeout-5xx-badJSON sanitized throw / stale no-fetch), report assembly mirroring codex API-path semantics; no-leak assertions. Validation: scoped vitest; JSON.stringify scan for fake tokens. +- [x] T2.1: `resolveOpenAiCredential` tiers (env → pi `openai` api_key → pi `openai-codex` oauth → not-found error mentioning pi login); keep `resolveOpenAiApiKey` compat; ms-vs-s + JWT staleness helper. TDD first: extend `openai.test.ts` resolution + staleness scenarios with fixture `openai-codex-auth.json`. Validation: scoped vitest. +- [x] T2.2: OAuth probe path: wham GET (Bearer + `ChatGPT-Account-Id`, abort/timeout), classification (200 report / 401-403 unauthenticated / timeout-5xx-badJSON sanitized throw / stale no-fetch), report assembly mirroring codex API-path semantics; no-leak assertions. Validation: scoped vitest; JSON.stringify scan for fake tokens. ### Phase 3: Integration & gates @@ -51,3 +51,4 @@ description: Break down work into actionable tasks and estimate timeline ## Progress Log - 2026-09-27: Initial plan created from requirements/design/testing docs. All testing scenarios mapped to tasks (T1.1↔wham parsing, T2.1↔resolution/staleness, T2.2↔classification/no-leak, T3.1↔gates, T3.2↔manual e2e). +- 2026-09-27: M2 done via TDD. T2.1: `resolveOpenAiCredential` union + tiers, not-found message now mentions pi login (CLI test uses injected errors — unaffected; openai.test.ts regex updated accordingly). T2.2: OAuth probe with wham fetch/classification; stale no-fetch; JWT fallback. Coverage: openai.ts & wham.ts 100% lines; remaining uncovered branches are the never-network `?? globalThis.fetch`/default-timeout injections, consistent with codex.ts/zai.ts norms. 70 capacity tests + full agent-manager suite green; hooks passed. Next: T3.1 root gates. From 55cc1ba7e0e2c9e36f1bf594f2908b91260cd8a7 Mon Sep 17 00:00:00 2001 From: Hoang Nguyen Date: Sun, 27 Sep 2026 09:09:12 +0000 Subject: [PATCH 08/10] chore(capacity): revert unrelated test formatting churn --- .../src/__tests__/capacity/codex.test.ts | 41 ++++++++++--- .../src/__tests__/capacity/zai.test.ts | 61 ++++++++++++++----- 2 files changed, 78 insertions(+), 24 deletions(-) diff --git a/packages/agent-manager/src/__tests__/capacity/codex.test.ts b/packages/agent-manager/src/__tests__/capacity/codex.test.ts index d06b5eb8..22f023e7 100644 --- a/packages/agent-manager/src/__tests__/capacity/codex.test.ts +++ b/packages/agent-manager/src/__tests__/capacity/codex.test.ts @@ -52,7 +52,9 @@ describe("Codex auth resolution", () => { }); it("falls back to ~/.codex/auth.json", () => { - expect(resolveCodexAuthPath({ HOME: "/users/test" })).toBe("/users/test/.codex/auth.json"); + expect(resolveCodexAuthPath({ HOME: "/users/test" })).toBe( + "/users/test/.codex/auth.json", + ); }); }); @@ -106,7 +108,9 @@ describe("tiered Codex probing", () => { status: 200, }), ) - .mockResolvedValueOnce(new Response(JSON.stringify(apiUsage()), { status: 200 })); + .mockResolvedValueOnce( + new Response(JSON.stringify(apiUsage()), { status: 200 }), + ); const rpc = vi.fn(); const result = await probeCodexCapacity({ ...context, @@ -125,7 +129,9 @@ describe("tiered Codex probing", () => { expect(fetch.mock.calls[0][0]).toBe( "https://auth.openai.com/api/accounts/v1/user-auth-credential/whoami", ); - expect(fetch.mock.calls[1][0]).toBe("https://chatgpt.com/backend-api/wham/usage"); + expect(fetch.mock.calls[1][0]).toBe( + "https://chatgpt.com/backend-api/wham/usage", + ); expect(fetch.mock.calls[1][1].headers).toMatchObject({ Authorization: "Bearer pat-secret", "ChatGPT-Account-Id": "acct-1", @@ -148,7 +154,9 @@ describe("tiered Codex probing", () => { it("selects a fresh OAuth token without calling whoami", async () => { const fetch = vi .fn() - .mockResolvedValue(new Response(JSON.stringify(apiUsage()), { status: 200 })); + .mockResolvedValue( + new Response(JSON.stringify(apiUsage()), { status: 200 }), + ); const result = await probeCodexCapacity({ ...context, readFile: async () => @@ -202,7 +210,9 @@ describe("tiered Codex probing", () => { ])("falls back to the CLI for %s", async (name, readFile) => { const fetch = vi .fn() - .mockResolvedValue(new Response("", { status: name === "OAuth 401" ? 401 : 200 })); + .mockResolvedValue( + new Response("", { status: name === "OAuth 401" ? 401 : 200 }), + ); const rpc = vi.fn(async () => ({ rateLimits: { rateLimits: { @@ -229,7 +239,8 @@ describe("tiered Codex probing", () => { })); const result = await probeCodexCapacity({ ...context, - readFile: async () => JSON.stringify({ personal_access_token: "pat-secret" }), + readFile: async () => + JSON.stringify({ personal_access_token: "pat-secret" }), fetch: vi.fn().mockRejectedValue(new Error("network failure pat-secret")), rpc, }); @@ -241,7 +252,9 @@ describe("tiered Codex probing", () => { const fetch = vi .fn() .mockRejectedValueOnce(new Error("PAT failed")) - .mockResolvedValueOnce(new Response(JSON.stringify(apiUsage()), { status: 200 })); + .mockResolvedValueOnce( + new Response(JSON.stringify(apiUsage()), { status: 200 }), + ); const rpc = vi.fn(); const result = await probeCodexCapacity({ ...context, @@ -281,7 +294,13 @@ describe("tiered Codex probing", () => { "account/read", ]); expect(JSON.stringify(messages)).not.toMatch(/prompt|turn\/start/); - expect(CODEX_APP_SERVER_ARGS).toEqual(["-s", "read-only", "-a", "untrusted", "app-server"]); + expect(CODEX_APP_SERVER_ARGS).toEqual([ + "-s", + "read-only", + "-a", + "untrusted", + "app-server", + ]); }); it("uses account/read to distinguish logged-out CLI state", async () => { @@ -297,7 +316,11 @@ describe("tiered Codex probing", () => { }); it("never exposes tokens or raw auth content through failures", async () => { - const secrets = ["pat-secret-value", "oauth-secret-value", "refresh-secret-value"]; + const secrets = [ + "pat-secret-value", + "oauth-secret-value", + "refresh-secret-value", + ]; const result = await probeCodexCapacity({ ...context, readFile: async () => diff --git a/packages/agent-manager/src/__tests__/capacity/zai.test.ts b/packages/agent-manager/src/__tests__/capacity/zai.test.ts index 7a23dc6f..c72bb7e5 100644 --- a/packages/agent-manager/src/__tests__/capacity/zai.test.ts +++ b/packages/agent-manager/src/__tests__/capacity/zai.test.ts @@ -1,11 +1,18 @@ import { readFile } from "node:fs/promises"; import { fileURLToPath } from "node:url"; import { describe, expect, it, vi } from "vitest"; -import { parseZaiQuota, probeZaiCapacity, resolveZaiApiKey } from "../../capacity/zai.js"; +import { + parseZaiQuota, + probeZaiCapacity, + resolveZaiApiKey, +} from "../../capacity/zai.js"; const checkedAt = "2026-08-20T10:00:00.000Z"; const fixture = (name: string) => - readFile(fileURLToPath(new URL(`./fixtures/${name}`, import.meta.url)), "utf8"); + readFile( + fileURLToPath(new URL(`./fixtures/${name}`, import.meta.url)), + "utf8", + ); describe("z.ai credential resolution", () => { it("prefers Z_AI_API_KEY without reading Pi auth", async () => { @@ -24,27 +31,40 @@ describe("z.ai credential resolution", () => { await expect( resolveZaiApiKey({ env: { HOME: "/users/test" }, readFile: readAuth }), ).resolves.toBe("pi-zai-test-key"); - expect(readAuth).toHaveBeenCalledWith("/users/test/.pi/agent/auth.json", "utf8"); + expect(readAuth).toHaveBeenCalledWith( + "/users/test/.pi/agent/auth.json", + "utf8", + ); }); it.each([ [ "missing", - async () => Promise.reject(Object.assign(new Error("missing"), { code: "ENOENT" })), + async () => + Promise.reject(Object.assign(new Error("missing"), { code: "ENOENT" })), ], ["invalid JSON", async () => "{"], - ["wrong credential type", async () => JSON.stringify({ zai: { type: "oauth", key: "x" } })], + [ + "wrong credential type", + async () => JSON.stringify({ zai: { type: "oauth", key: "x" } }), + ], ["missing key", async () => JSON.stringify({ zai: { type: "api_key" } })], - ])("fails clearly for %s credentials without exposing file content", async (_name, readAuth) => { - await expect( - resolveZaiApiKey({ env: { HOME: "/users/test" }, readFile: readAuth }), - ).rejects.toThrow(/z\.ai API key|z\.ai Pi auth file/); - }); + ])( + "fails clearly for %s credentials without exposing file content", + async (_name, readAuth) => { + await expect( + resolveZaiApiKey({ env: { HOME: "/users/test" }, readFile: readAuth }), + ).rejects.toThrow(/z\.ai API key|z\.ai Pi auth file/); + }, + ); }); describe("z.ai quota mapping", () => { it("maps token, credit, and monthly MCP windows with recomputed usage", async () => { - const report = parseZaiQuota(JSON.parse(await fixture("zai-quota.json")), checkedAt); + const report = parseZaiQuota( + JSON.parse(await fixture("zai-quota.json")), + checkedAt, + ); expect(report).toMatchObject({ harness: "pi", provider: "zai", @@ -104,13 +124,24 @@ describe("z.ai quota mapping", () => { }, checkedAt, ); - expect(report.windows.map((window) => window.usedPercent)).toEqual([100, 0]); + expect(report.windows.map((window) => window.usedPercent)).toEqual([ + 100, 0, + ]); }); it.each([ - ["invalid fixture", async () => JSON.parse(await fixture("zai-quota-invalid.json"))], - ["unsuccessful envelope", async () => ({ success: false, code: 401, data: { limits: [] } })], - ["invalid entry", async () => ({ success: true, code: 200, data: { limits: [{}] } })], + [ + "invalid fixture", + async () => JSON.parse(await fixture("zai-quota-invalid.json")), + ], + [ + "unsuccessful envelope", + async () => ({ success: false, code: 401, data: { limits: [] } }), + ], + [ + "invalid entry", + async () => ({ success: true, code: 200, data: { limits: [{}] } }), + ], ])("rejects %s", async (_name, input) => { const value = await input(); expect(() => parseZaiQuota(value, checkedAt)).toThrow(/z\.ai quota/); From 8dc2a37053a415d021ea502c840c393531acfd03 Mon Sep 17 00:00:00 2001 From: Hoang Nguyen Date: Sun, 27 Sep 2026 09:11:42 +0000 Subject: [PATCH 09/10] docs(openai-pi-login-capacity): implementation and testing evidence --- ...-09-27-feature-openai-pi-login-capacity.md | 69 ++++++------------- ...-09-27-feature-openai-pi-login-capacity.md | 10 +-- ...-09-27-feature-openai-pi-login-capacity.md | 53 ++++++++------ 3 files changed, 58 insertions(+), 74 deletions(-) diff --git a/docs/ai/implementation/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/implementation/2026-09-27-feature-openai-pi-login-capacity.md index 6b1942f7..29d290fb 100644 --- a/docs/ai/implementation/2026-09-27-feature-openai-pi-login-capacity.md +++ b/docs/ai/implementation/2026-09-27-feature-openai-pi-login-capacity.md @@ -8,66 +8,39 @@ description: Technical implementation notes, patterns, and code guidelines ## Development Setup -**How do we get started?** - -- Prerequisites and dependencies -- Environment setup steps -- Configuration needed +- Worktree `.worktrees/feature-openai-pi-login-capacity`, branch `feature/openai-pi-login-capacity` (brief-mandated name), base `b10c0f6` (main). +- `npm ci` at root; unit tests run TS directly via vitest (`packages/agent-manager`); build CLI+agent-manager (`npx nx run-many -t build -p cli agent-manager`) before manual capacity runs. ## Code Structure -**How is the code organized?** - -- Directory structure -- Module organization -- Naming conventions +- `packages/agent-manager/src/capacity/wham.ts` (new) — shared pure parser for `chatgpt.com/backend-api/wham/usage`: `parseWhamUsage`, `toRateWindow`, `resetTime`, `safeIdentifier` (exported for codex's CLI-fallback path). +- `packages/agent-manager/src/capacity/openai.ts` — tiered `resolveOpenAiCredential` (env → pi `openai` api_key → pi `openai-codex` OAuth → sanitized not-found error), `toEpochMs` (>1e12 ⇒ ms), `jwtExpiryMs` fallback, `fetchWhamUsage` (Bearer + `ChatGPT-Account-Id`, AbortController timeout, 401/403 → null), `probeOpenAiOauthCapacity` classification. `resolveOpenAiApiKey` kept as platform-only compat wrapper. +- `packages/agent-manager/src/capacity/codex.ts` — parsing delegated to `wham.ts`; `parseUsage`/`toRateWindow` re-exported unchanged so `codex.test.ts` passes untouched. +- Tests/fixtures: `wham.test.ts`, extended `openai.test.ts`, `fixtures/openai-codex-auth.json` (fake tokens), `fixtures/wham-usage.json` (redacted live shape). ## Implementation Notes -**Key technical details to remember:** - ### Core Features -- Feature 1: Implementation approach -- Feature 2: Implementation approach -- Feature 3: Implementation approach +- **OAuth tier:** fresh token → wham GET → windows (`session`/`weekly`/extras) + credits; stale (`expiresMs ?? JWT exp` ≤ checkedAt) → no fetch, `authenticated:true, available:"unknown"`; 401/403 → `authenticated:false`; transport/5xx/bad-JSON → sanitized throw rendered as per-provider warning. +- **Platform path untouched:** same fetches, same report, byte-identical; platform key wins over OAuth when both exist (tier order). +- **No refresh, no dedupe:** matches codex.ts policy; `pi · OpenAI` and `codex · OpenAI` rows coexist (harness column). ### Patterns & Best Practices -- Design patterns being used -- Code style guidelines -- Common utilities/helpers - -## Integration Points - -**How do pieces connect?** - -- API integration details -- Database connections -- Third-party service setup - -## Error Handling - -**How do we handle failures?** - -- Error handling strategy -- Logging approach -- Retry/fallback mechanisms - -## Performance Considerations - -**How do we keep it fast?** +- TDD throughout (red → green per behavior; codex suite was the regression net for the parser extraction). +- Token hygiene: fixtures use `fixture-*` fakes; no-leak assertions on reports and errors; errors carry status codes only. -- Optimization strategies -- Caching approach -- Query optimization -- Resource management +## Validation Evidence -## Security Notes +- Scoped: 70 capacity tests green; agent-manager full suite green (682+ tests). +- Coverage: `openai.ts` / `wham.ts` 100% lines; uncovered branches are only the never-network default-injection arms (`?? globalThis.fetch`, `?? 5000`), consistent with codex.ts/zai.ts. +- Root gates: `npm run lint` ✓ (6 projects), `npm test` ✓ (6 projects), `npm run test:e2e` ✓ (42 tests). Husky pre-commit (`npm run lint` + `npm test`) passed on every commit. +- Live validation (this machine, real login, redacted): default `ai-devkit capacity` now shows `pi · OpenAI` Session/Weekly rows beside `pi · z.ai`, no unavailable warning; `codex · OpenAI` rows unchanged. `capacity openai --json` clean of `access`/`refresh`/`accountId` values (scripted check). Live `creditsRemaining` renders null when the account's `credits.balance` is non-numeric — parser tolerates. -**What security measures are in place?** +## Deviations & Cleanups -- Authentication/authorization -- Input validation -- Data encryption -- Secrets management +- Not-found error text now mentions the pi login path (only shown when nothing resolves; platform users unaffected). One existing openai.test.ts regex updated accordingly. +- Reverted accidental `oxfmt` reformatting of `codex.test.ts`/`zai.test.ts` (unrelated churn) — suites pass with pristine files. +- `lint --feature` branch-name check expects `feature-openai-pi-login-capacity`; the brief mandates `feature/openai-pi-login-capacity` — documented deviation, docs checks pass. +- Follow-up (out of scope): wham `rate_limit_reached_type` is not surfaced as `available:"no"` for API responses (parity with codex.ts API path); could be revisited for both harnesses together. diff --git a/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md index 6742db23..dcbfb077 100644 --- a/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md +++ b/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md @@ -10,7 +10,7 @@ description: Break down work into actionable tasks and estimate timeline - [x] M1: Shared wham parsing extracted, codex suite green unchanged - [x] M2: Pi OAuth tier resolves + classifies (all unit scenarios green) -- [ ] M3: Full gates green, docs/implementation updated, final PR open +- [x] M3: Full gates green, docs/implementation updated, final PR open ## Task Breakdown @@ -26,8 +26,8 @@ description: Break down work into actionable tasks and estimate timeline ### Phase 3: Integration & gates -- [ ] T3.1: Full gates: `npm run lint`, `npm test`, `npm run test:e2e`, hooks (dev-commit) — all green in worktree. Record evidence in implementation doc. -- [ ] T3.2: Manual on-machine validation: build CLI, run default `ai-devkit capacity`, confirm `pi · OpenAI` rows beside `pi · z.ai`; redacted output in implementation doc. +- [x] T3.1: Full gates: `npm run lint`, `npm test`, `npm run test:e2e`, hooks (dev-commit) — all green in worktree. Record evidence in implementation doc. +- [x] T3.2: Manual on-machine validation: build CLI, run default `ai-devkit capacity`, confirm `pi · OpenAI` rows beside `pi · z.ai`; redacted output in implementation doc. - [ ] T3.3: Update implementation/testing/planning docs (phases 5-8 flow), grep branch for token substrings (no-leak audit), final PR (dev-pr conventions; do NOT merge). ## Dependencies @@ -51,4 +51,6 @@ description: Break down work into actionable tasks and estimate timeline ## Progress Log - 2026-09-27: Initial plan created from requirements/design/testing docs. All testing scenarios mapped to tasks (T1.1↔wham parsing, T2.1↔resolution/staleness, T2.2↔classification/no-leak, T3.1↔gates, T3.2↔manual e2e). -- 2026-09-27: M2 done via TDD. T2.1: `resolveOpenAiCredential` union + tiers, not-found message now mentions pi login (CLI test uses injected errors — unaffected; openai.test.ts regex updated accordingly). T2.2: OAuth probe with wham fetch/classification; stale no-fetch; JWT fallback. Coverage: openai.ts & wham.ts 100% lines; remaining uncovered branches are the never-network `?? globalThis.fetch`/default-timeout injections, consistent with codex.ts/zai.ts norms. 70 capacity tests + full agent-manager suite green; hooks passed. Next: T3.1 root gates. +- 2026-09-27: M1 done. T1.1 wham.ts + wham.test.ts (TDD red→green; note: scoped-window ids reject spaces by safeIdentifier design — test uses `code-review`). T1.2 codex.ts delegates parsing; exported `resetTime`/`safeIdentifier` from wham.ts because codex's CLI-fallback path uses them (caught by 3 codex CLI-fallback tests going red). All 50 capacity tests + 682 agent-manager tests green; hooks passed. +- 2026-09-27: M2 done via TDD. T2.1: `resolveOpenAiCredential` union + tiers, not-found message now mentions pi login (CLI test uses injected errors — unaffected; openai.test.ts regex updated accordingly). T2.2: OAuth probe with wham fetch/classification; stale no-fetch; JWT fallback. Coverage: openai.ts & wham.ts 100% lines; remaining uncovered branches are the never-network `?? globalThis.fetch`/default-timeout injections, consistent with codex.ts/zai.ts norms. 70 capacity tests + full agent-manager suite green; hooks passed. +- 2026-09-27: M3 in progress. T3.1 gates green (lint 6 projects, npm test 6 projects, e2e 42). T3.2 live validation green (`pi · OpenAI` rows render; json no-leak scripted check; unrelated oxfmt churn in codex/zai tests reverted). Implementation + testing docs finalized. Remaining: T3.3 no-leak grep + PR. diff --git a/docs/ai/testing/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/testing/2026-09-27-feature-openai-pi-login-capacity.md index e925e3dc..6b7d0edd 100644 --- a/docs/ai/testing/2026-09-27-feature-openai-pi-login-capacity.md +++ b/docs/ai/testing/2026-09-27-feature-openai-pi-login-capacity.md @@ -17,43 +17,43 @@ description: Define testing approach, test cases, and quality assurance ### `capacity/wham.ts` (shared parser) -- [ ] Parses `rate_limit.primary_window`/`secondary_window` into session/weekly windows with `used_percent`, `limit_window_seconds`→durationMinutes, `reset_at` (epoch seconds and ISO string) → resetsAt (covers codex API mapping parity) -- [ ] Maps `additional_rate_limits` entries into scoped windows; tolerates non-array/absent value (live probe showed non-array) -- [ ] Extracts `credits.balance`; missing credits → null -- [ ] Skips windows it cannot parse without inventing zeros (missing used_percent stays null) -- [ ] Codex suite (`codex.test.ts`) passes unchanged against the delegated parser (regression guard) +- [x] Parses `rate_limit.primary_window`/`secondary_window` into session/weekly windows with `used_percent`, `limit_window_seconds`→durationMinutes, `reset_at` (epoch seconds and ISO string) → resetsAt (covers codex API mapping parity) +- [x] Maps `additional_rate_limits` entries into scoped windows; tolerates non-array/absent value (live probe showed non-array) +- [x] Extracts `credits.balance`; missing credits → null +- [x] Skips windows it cannot parse without inventing zeros (missing used_percent stays null) +- [x] Codex suite (`codex.test.ts`) passes unchanged against the delegated parser (regression guard) ### `capacity/openai.ts` — credential resolution -- [ ] `OPENAI_API_KEY` env wins without reading the auth file (existing behavior preserved) -- [ ] pi `openai` api_key credential wins over `openai-codex` when both exist (platform preferred, OAuth not consulted) -- [ ] pi `openai-codex` `{type:"oauth"}` resolves to OAuth credential when no platform key exists -- [ ] Malformed `openai-codex` (wrong type / missing access / missing accountId) is skipped → not-found error, message mentions pi login, never exposes file content -- [ ] Missing auth file (ENOENT) and invalid JSON keep clear sanitized errors +- [x] `OPENAI_API_KEY` env wins without reading the auth file (existing behavior preserved) +- [x] pi `openai` api_key credential wins over `openai-codex` when both exist (platform preferred, OAuth not consulted) +- [x] pi `openai-codex` `{type:"oauth"}` resolves to OAuth credential when no platform key exists +- [x] Malformed `openai-codex` (wrong type / missing access / missing accountId) is skipped → not-found error, message mentions pi login, never exposes file content +- [x] Missing auth file (ENOENT) and invalid JSON keep clear sanitized errors ### `capacity/openai.ts` — OAuth probe & classification -- [ ] Fresh token (epoch-ms `expires` in future) → GET wham/usage with `Authorization: Bearer` + `ChatGPT-Account-Id`, timeout abort wiring; 200 → report with `harness:"pi"`, session/weekly windows, credits, `authenticated:true` -- [ ] `expires` in seconds (small magnitude) and JWT-exp fallback both classify correctly; missing metadata defaults to fresh-then-fetch -- [ ] Stale token (expiry ≤ checkedAt) → **no fetch**, `authenticated:true`, `available:"unknown"`, no windows -- [ ] 401/403 from wham → `authenticated:false`, `available:"unknown"`, empty windows -- [ ] Timeout / network error / 5xx / bad JSON → sanitized thrown error (`OpenAI usage request failed…`), no token in message -- [ ] `resolveOpenAiApiKey` compat export unchanged for platform tiers -- [ ] No-token-leak: JSON.stringify of every report/error never contains fixture token strings +- [x] Fresh token (epoch-ms `expires` in future) → GET wham/usage with `Authorization: Bearer` + `ChatGPT-Account-Id`, timeout abort wiring; 200 → report with `harness:"pi"`, session/weekly windows, credits, `authenticated:true` +- [x] `expires` in seconds (small magnitude) and JWT-exp fallback both classify correctly; missing metadata defaults to fresh-then-fetch +- [x] Stale token (expiry ≤ checkedAt) → **no fetch**, `authenticated:true`, `available:"unknown"`, no windows +- [x] 401/403 from wham → `authenticated:false`, `available:"unknown"`, empty windows +- [x] Timeout / network error / 5xx / bad JSON → sanitized thrown error (`OpenAI usage request failed…`), no token in message +- [x] `resolveOpenAiApiKey` compat export unchanged for platform tiers +- [x] No-token-leak: JSON.stringify of every report/error never contains fixture token strings ### CLI (existing suites) -- [ ] `capacity` command tests keep passing (default provider list unchanged; no CLI code changes) +- [x] `capacity` command tests keep passing (default provider list unchanged; no CLI code changes) ## Integration Tests -- [ ] `getOpenAiCapacityReport` end-to-end with injected `readFile` (fixture auth) + injected `fetch` (wham 200) → full report shape (windows sorted/renderable, provider label path) -- [ ] Multi-provider flow: OAuth-only machine → openai provider no longer throws (row present) while other providers' failures still warn independently (`capacityCommand` behavior unchanged) +- [x] `getOpenAiCapacityReport` end-to-end with injected `readFile` (fixture auth) + injected `fetch` (wham 200) → full report shape (windows sorted/renderable, provider label path) +- [x] Multi-provider flow: OAuth-only machine → openai provider no longer throws (row present) while other providers' failures still warn independently (`capacityCommand` behavior unchanged) ## End-to-End Tests -- [ ] Existing e2e suite green (`npm run test:e2e`); no new e2e (no CLI surface change) -- [ ] Manual validation on this machine: default `ai-devkit capacity` shows `pi · OpenAI` rows beside `pi · z.ai` (documented in implementation doc with redacted output) +- [x] Existing e2e suite green (`npm run test:e2e`); no new e2e (no CLI surface change) +- [x] Manual validation on this machine: default `ai-devkit capacity` shows `pi · OpenAI` rows beside `pi · z.ai` (documented in implementation doc with redacted output) ## Test Data @@ -65,6 +65,15 @@ description: Define testing approach, test cases, and quality assurance ## Test Reporting & Coverage +- **Results (2026-09-27, worktree `feature/openai-pi-login-capacity`):** all 22 scenarios covered and green. + - `packages/agent-manager/src/__tests__/capacity/wham.test.ts` — 6 tests (parsing parity, extras, degraded extras, non-object payload). + - `packages/agent-manager/src/__tests__/capacity/openai.test.ts` — 36 tests (16 pre-existing platform tests + resolution tiers, staleness ms/s/JWT, classification, no-leak, sanitized errors). + - `packages/agent-manager/src/__tests__/capacity/codex.test.ts` — unchanged from main, passes against delegated parser (regression guard). + - CLI capacity command suites unchanged and green. +- **Commands/evidence:** scoped `npx vitest run src/__tests__/capacity/` → 70 passed; full agent-manager suite green; root gates `npm run lint` ✓, `npm test` ✓, `npm run test:e2e` ✓ (42); husky pre-commit green on all commits. +- **Coverage:** `openai.ts` and `wham.ts` at 100% lines; branch gaps limited to never-network default-injection arms (`?? globalThis.fetch`, `?? 5000`) — same arms uncovered in codex.ts/zai.ts; documented as accepted (package convention, network-free tests). +- **Manual/live:** default report shows `pi · OpenAI` rows on this machine; `--json` verified free of token/accountId values via scripted check. + - Commands: `npx vitest run` in `packages/agent-manager` (scoped), then `npm run lint`, `npm test`, `npm run test:e2e` at root; coverage via existing thresholds (70% floor package-wide, new code targeted at 100% branch). - Coverage gaps: none anticipated; any justified gap documented here with rationale. - Evidence: command outputs recorded in the implementation doc; final validation in the PR body. From 3f51f5054e2a7437aaf5567bf6d7a734fe8376f4 Mon Sep 17 00:00:00 2001 From: Hoang Nguyen Date: Sun, 27 Sep 2026 09:14:43 +0000 Subject: [PATCH 10/10] docs(planning): close out openai-pi-login-capacity lifecycle --- .../planning/2026-09-27-feature-openai-pi-login-capacity.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md b/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md index dcbfb077..0957e4aa 100644 --- a/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md +++ b/docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md @@ -28,7 +28,7 @@ description: Break down work into actionable tasks and estimate timeline - [x] T3.1: Full gates: `npm run lint`, `npm test`, `npm run test:e2e`, hooks (dev-commit) — all green in worktree. Record evidence in implementation doc. - [x] T3.2: Manual on-machine validation: build CLI, run default `ai-devkit capacity`, confirm `pi · OpenAI` rows beside `pi · z.ai`; redacted output in implementation doc. -- [ ] T3.3: Update implementation/testing/planning docs (phases 5-8 flow), grep branch for token substrings (no-leak audit), final PR (dev-pr conventions; do NOT merge). +- [x] T3.3: Update implementation/testing/planning docs (phases 5-8 flow), grep branch for token substrings (no-leak audit), final PR (dev-pr conventions; do NOT merge). ## Dependencies @@ -53,4 +53,4 @@ description: Break down work into actionable tasks and estimate timeline - 2026-09-27: Initial plan created from requirements/design/testing docs. All testing scenarios mapped to tasks (T1.1↔wham parsing, T2.1↔resolution/staleness, T2.2↔classification/no-leak, T3.1↔gates, T3.2↔manual e2e). - 2026-09-27: M1 done. T1.1 wham.ts + wham.test.ts (TDD red→green; note: scoped-window ids reject spaces by safeIdentifier design — test uses `code-review`). T1.2 codex.ts delegates parsing; exported `resetTime`/`safeIdentifier` from wham.ts because codex's CLI-fallback path uses them (caught by 3 codex CLI-fallback tests going red). All 50 capacity tests + 682 agent-manager tests green; hooks passed. - 2026-09-27: M2 done via TDD. T2.1: `resolveOpenAiCredential` union + tiers, not-found message now mentions pi login (CLI test uses injected errors — unaffected; openai.test.ts regex updated accordingly). T2.2: OAuth probe with wham fetch/classification; stale no-fetch; JWT fallback. Coverage: openai.ts & wham.ts 100% lines; remaining uncovered branches are the never-network `?? globalThis.fetch`/default-timeout injections, consistent with codex.ts/zai.ts norms. 70 capacity tests + full agent-manager suite green; hooks passed. -- 2026-09-27: M3 in progress. T3.1 gates green (lint 6 projects, npm test 6 projects, e2e 42). T3.2 live validation green (`pi · OpenAI` rows render; json no-leak scripted check; unrelated oxfmt churn in codex/zai tests reverted). Implementation + testing docs finalized. Remaining: T3.3 no-leak grep + PR. +- 2026-09-27: M3 done. No-leak audit clean (branch diff + tracked files scanned against real token values). Final gates fresh on HEAD: lint 0 / npm test 0 / e2e 0 (42). Rebased (up to date) onto origin/main, pushed, PR opened: https://github.com/codeaholicguy/ai-devkit/pull/254 (left unmerged for review). Lifecycle complete.