Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
106 changes: 106 additions & 0 deletions docs/ai/design/2026-09-27-feature-openai-pi-login-capacity.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,106 @@
---
phase: design
title: System Design & Architecture
description: Define the technical architecture, components, and data models
---

# System Design & Architecture

## Architecture Overview

The feature extends the existing pi·OpenAI capacity provider (`capacity/openai.ts`) with an OAuth tier backed by pi's OpenAI login, reusing the wham/usage endpoint that the codex harness provider already consumes. No CLI or render changes are required — `render.ts` already renders generic `CapacityWindow` rows and labels provider `openai` as "OpenAI".

```mermaid
graph TD
CLI["ai-devkit capacity"] --> CMD["capacity command<br/>(packages/cli)"]
CMD --> IDX["getOpenAiCapacityReport<br/>(agent-manager/capacity/index.ts)"]
CMD --> CODEX["getCodexCapacityReport<br/>(codex.ts, unchanged)"]
IDX --> OA["probeOpenAiCapacity<br/>(openai.ts)"]
OA --> RES["resolveOpenAiCredential"]
RES -->|env OPENAI_API_KEY| T1["tier 1: platform key"]
RES -->|pi auth.json openai api_key| T2["tier 2: platform key"]
RES -->|pi auth.json openai-codex oauth| T3["tier 3: OAuth"]
T1 --> PLAT["api.openai.com usage API<br/>(today/7-day token windows)"]
T3 --> WHAM["GET chatgpt.com/backend-api/wham/usage<br/>Bearer + ChatGPT-Account-Id"]
WHAM --> PARSE["parseWhamUsage<br/>(shared wham.ts)"]
CODEX --> PARSE
PLAT --> RPT["CapacityReport harness=pi provider=openai"]
PARSE --> RPT
```

Key components:

- `packages/agent-manager/src/capacity/openai.ts` — credential resolution becomes tiered; gains the wham/usage fetch path and OAuth staleness classification.
- `packages/agent-manager/src/capacity/wham.ts` (new) — shared wham/usage response parsing (rate-limit windows, extra limits, credits), extracted from `codex.ts` now that a second real caller exists.
- `packages/agent-manager/src/capacity/codex.ts` — unchanged behavior; `parseUsage`/`toRateWindow` remain exported (tests import them) and delegate to `wham.ts`.
- CLI (`packages/cli`) — no changes.

## Data Models

**Input credential (read-only), tier 3 — pi `~/.pi/agent/auth.json`:**

```jsonc
{
"openai-codex": {
"type": "oauth", // verified live: "oauth"
"access": "<jwt string>", // never logged/persisted
"refresh": "<opaque>", // unused (no refresh, see D2)
"expires": 1791355868557, // epoch MILLISECONDS (verified live)
"accountId": "<chatgpt account uuid>"
}
}
```

**Resolved credential (internal discriminated union):**

```ts
type OpenAiCredential =
| { kind: "platform"; key: string } // tiers 1-2
| { kind: "oauth"; access: string; accountId: string; expiresMs: number | null };
```

**Output** — existing `CapacityReport`/`CapacityWindow` types, unchanged. OAuth windows come from wham: `session` (5h), `weekly` (7d), plus any `additional_rate_limits` entries; `creditsRemaining` from `credits.balance`.

**Staleness rule (ms-aware):** `expires` values above 1e12 are epoch ms (pi convention), smaller values are treated as seconds; JWT `exp` (seconds) is the fallback when `expires` is absent/invalid — mirroring `codex.ts staleOAuth`'s metadata-then-JWT order. Reference "now" is `Date.parse(checkedAt)` (already injectable via `now` option) so tests are deterministic.

## API Design

**External:** one read-only `GET https://chatgpt.com/backend-api/wham/usage` with headers `Authorization: Bearer <access>` and `ChatGPT-Account-Id: <accountId>` — identical to the codex provider's call. Verified live with the real pi token: HTTP 200, `rate_limit.primary_window/secondary_window`, `credits.balance`, top-level `rate_limit_reached_type` present. Timeout and abort handling follow the existing AbortController pattern in `openai.ts`.

**Internal:** `probeOpenAiCapacity(options)` keeps its signature (`env`, `readFile`, `fetch`, `timeoutMs`, `checkedAt` via `getOpenAiCapacityReport`). `resolveOpenAiApiKey` remains exported for compatibility; a new internal `resolveOpenAiCredential` returns the union above. Resolution order:

1. `env.OPENAI_API_KEY` (explicit override; auth file not even read — preserved behavior)
2. pi `openai` `{type: "api_key", key}` (platform key path, byte-identical report)
3. pi `openai-codex` `{type: "oauth"}` with non-empty `access` + `accountId` → OAuth path
4. nothing found → error `OpenAI credentials not found; set OPENAI_API_KEY, configure the Pi openai provider, or log in to OpenAI via Pi`

**Response classification (OAuth tier):**

- fresh token → fetch wham; 200 → report (authenticated, windows, credits)
- 401/403 → `{authenticated: false, available: "unknown", windows: []}` (revoked login)
- timeout / network error / 5xx / malformed JSON → sanitized throw (`OpenAI usage request failed…`), rendered as the per-provider warning by `capacityCommand` — consistent with the platform tier's error behavior
- stale token (expiry ≤ now) → **no fetch**; `{authenticated: true, available: "unknown", windows: [], creditsRemaining: null}` (authenticated-but-stale)
- malformed `openai-codex` entry (wrong type, missing fields) → tier skipped, falls to step 4 error; error messages never include file content

## Component Breakdown

- `wham.ts` (new): `toRateWindow`, `extraWindows`, `parseWhamUsage(raw): { windows, creditsRemaining }`, plus the small guards (`record`, `finiteNumber`, `nonEmptyText`, `resetTime`, `safeIdentifier`). Pure functions, no I/O.
- `codex.ts`: imports parsing from `wham.ts`; re-exports `toRateWindow` and keeps `parseUsage(raw, source)` as a thin wrapper so `codex.test.ts` passes unchanged. Probe behavior identical.
- `openai.ts`: `resolveOpenAiCredential` tiers; OAuth fetch helper (Bearer + account id, abort, status classification); staleness check; report assembly mirroring codex's `capacityFromSnapshot` for API responses (`available = hasUsage ? "yes" : "unknown"`; top-level `rate_limit_reached_type` is deliberately **not** read, matching codex.ts API-path behavior).
- Tests/fixtures: `openai-codex-auth.json` (fake tokens), `wham-usage.json` (redacted live shape) under `capacity/fixtures/`.

## Design Decisions

- **D1 Resolution order (env → platform key → OAuth).** Platform paths stay first: they are correct for platform-key users and the brief mandates keeping them working. OAuth only fills the no-setup gap. Alternative (OAuth first) would change output for dual-credential users without a motivating use case. Alternative (show both) would need two `pi · OpenAI` rows and render changes — rejected as scope creep.
- **D2 No OAuth refresh.** Matches codex.ts precedent exactly (its OAuth path never refreshes; stale → skip). Attempting refresh would invent a new flow (token rotation, persistence of rotated tokens into pi's auth file — writing another tool's state) with no existing pattern to follow. Stale renders authenticated-but-stale.
- **D3 Extract wham parsing into `wham.ts`.** The brief allows extraction only with a second real caller — that now exists (pi OAuth tier), verified live against the same endpoint/shape. Alternatives: importing `parseUsage` from `codex.ts` (couples the pi provider to the codex provider) or duplicating ~60 lines of parsing (drift risk) — both worse than a small shared pure module. Codex re-exports keep its public surface stable.
- **D4 ms-vs-s staleness heuristic.** pi writes epoch ms (verified); codex writes seconds. A magnitude check (`> 1e12` ⇒ ms) handles both without guessing per-source conventions, and JWT fallback covers missing metadata.
- **D5 No dedupe with `codex · OpenAI`.** Different credential stores (`~/.codex/auth.json` vs `~/.pi/agent/auth.json`); both rows legitimately coexist; the harness column exists for exactly this.
- **D6 Error message updated** to mention the pi login path (it fires only when nothing is found, so platform users never see it). Sanitized errors only — no token or file content in messages, matching the existing openai/codex tests' no-leak assertions.

## Non-Functional Requirements

- **Security:** token values never appear in reports, errors, logs, docs, or fixtures (fixtures use fake values); auth file is read-only; one live probe only (done, read-only GET).
- **Performance:** one additional HTTPS GET (timeout 5s default, abortable) only in the OAuth tier; platform tier unchanged; tests use injected `fetch` (no network).
- **Reliability:** failure modes classified (stale / unauthenticated / unavailable) so one provider's failure never blanks the multi-provider report.
- **Compatibility:** `resolveOpenAiApiKey`, `parseUsage`, `toRateWindow` exports preserved; platform-key report byte-identical.
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
---
phase: implementation
title: Implementation Guide
description: Technical implementation notes, patterns, and code guidelines
---

# Implementation Guide

## Development Setup

- Worktree `.worktrees/feature-openai-pi-login-capacity`, branch `feature/openai-pi-login-capacity` (brief-mandated name), base `b10c0f6` (main).
- `npm ci` at root; unit tests run TS directly via vitest (`packages/agent-manager`); build CLI+agent-manager (`npx nx run-many -t build -p cli agent-manager`) before manual capacity runs.

## Code Structure

- `packages/agent-manager/src/capacity/wham.ts` (new) — shared pure parser for `chatgpt.com/backend-api/wham/usage`: `parseWhamUsage`, `toRateWindow`, `resetTime`, `safeIdentifier` (exported for codex's CLI-fallback path).
- `packages/agent-manager/src/capacity/openai.ts` — tiered `resolveOpenAiCredential` (env → pi `openai` api_key → pi `openai-codex` OAuth → sanitized not-found error), `toEpochMs` (>1e12 ⇒ ms), `jwtExpiryMs` fallback, `fetchWhamUsage` (Bearer + `ChatGPT-Account-Id`, AbortController timeout, 401/403 → null), `probeOpenAiOauthCapacity` classification. `resolveOpenAiApiKey` kept as platform-only compat wrapper.
- `packages/agent-manager/src/capacity/codex.ts` — parsing delegated to `wham.ts`; `parseUsage`/`toRateWindow` re-exported unchanged so `codex.test.ts` passes untouched.
- Tests/fixtures: `wham.test.ts`, extended `openai.test.ts`, `fixtures/openai-codex-auth.json` (fake tokens), `fixtures/wham-usage.json` (redacted live shape).

## Implementation Notes

### Core Features

- **OAuth tier:** fresh token → wham GET → windows (`session`/`weekly`/extras) + credits; stale (`expiresMs ?? JWT exp` ≤ checkedAt) → no fetch, `authenticated:true, available:"unknown"`; 401/403 → `authenticated:false`; transport/5xx/bad-JSON → sanitized throw rendered as per-provider warning.
- **Platform path untouched:** same fetches, same report, byte-identical; platform key wins over OAuth when both exist (tier order).
- **No refresh, no dedupe:** matches codex.ts policy; `pi · OpenAI` and `codex · OpenAI` rows coexist (harness column).

### Patterns & Best Practices

- TDD throughout (red → green per behavior; codex suite was the regression net for the parser extraction).
- Token hygiene: fixtures use `fixture-*` fakes; no-leak assertions on reports and errors; errors carry status codes only.

## Validation Evidence

- Scoped: 70 capacity tests green; agent-manager full suite green (682+ tests).
- Coverage: `openai.ts` / `wham.ts` 100% lines; uncovered branches are only the never-network default-injection arms (`?? globalThis.fetch`, `?? 5000`), consistent with codex.ts/zai.ts.
- Root gates: `npm run lint` ✓ (6 projects), `npm test` ✓ (6 projects), `npm run test:e2e` ✓ (42 tests). Husky pre-commit (`npm run lint` + `npm test`) passed on every commit.
- Live validation (this machine, real login, redacted): default `ai-devkit capacity` now shows `pi · OpenAI` Session/Weekly rows beside `pi · z.ai`, no unavailable warning; `codex · OpenAI` rows unchanged. `capacity openai --json` clean of `access`/`refresh`/`accountId` values (scripted check). Live `creditsRemaining` renders null when the account's `credits.balance` is non-numeric — parser tolerates.

## Deviations & Cleanups

- Not-found error text now mentions the pi login path (only shown when nothing resolves; platform users unaffected). One existing openai.test.ts regex updated accordingly.
- Reverted accidental `oxfmt` reformatting of `codex.test.ts`/`zai.test.ts` (unrelated churn) — suites pass with pristine files.
- `lint --feature` branch-name check expects `feature-openai-pi-login-capacity`; the brief mandates `feature/openai-pi-login-capacity` — documented deviation, docs checks pass.
- Follow-up (out of scope): wham `rate_limit_reached_type` is not surfaced as `available:"no"` for API responses (parity with codex.ts API path); could be revisited for both harnesses together.
56 changes: 56 additions & 0 deletions docs/ai/planning/2026-09-27-feature-openai-pi-login-capacity.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
---
phase: planning
title: Project Planning & Task Breakdown
description: Break down work into actionable tasks and estimate timeline
---

# Project Planning & Task Breakdown

## Milestones

- [x] M1: Shared wham parsing extracted, codex suite green unchanged
- [x] M2: Pi OAuth tier resolves + classifies (all unit scenarios green)
- [x] M3: Full gates green, docs/implementation updated, final PR open

## Task Breakdown

### Phase 1: Foundation — shared parser

- [x] T1.1: Create `capacity/wham.ts` with `parseWhamUsage`, `toRateWindow`, `extraWindows` + guards (TDD: add `wham.test.ts` scenarios from testing doc first; fixtures `wham-usage.json`). Outcome: pure parser module; validation: `npx vitest run src/__tests__/capacity/wham.test.ts` (agent-manager).
- [x] T1.2: Delegate `codex.ts` parsing to `wham.ts`; keep `parseUsage`/`toRateWindow` exports. Outcome: codex behavior identical; validation: `codex.test.ts` passes unchanged (no edits to that suite).

### Phase 2: Core — OAuth tier in openai.ts

- [x] T2.1: `resolveOpenAiCredential` tiers (env → pi `openai` api_key → pi `openai-codex` oauth → not-found error mentioning pi login); keep `resolveOpenAiApiKey` compat; ms-vs-s + JWT staleness helper. TDD first: extend `openai.test.ts` resolution + staleness scenarios with fixture `openai-codex-auth.json`. Validation: scoped vitest.
- [x] T2.2: OAuth probe path: wham GET (Bearer + `ChatGPT-Account-Id`, abort/timeout), classification (200 report / 401-403 unauthenticated / timeout-5xx-badJSON sanitized throw / stale no-fetch), report assembly mirroring codex API-path semantics; no-leak assertions. Validation: scoped vitest; JSON.stringify scan for fake tokens.

### Phase 3: Integration & gates

- [x] T3.1: Full gates: `npm run lint`, `npm test`, `npm run test:e2e`, hooks (dev-commit) — all green in worktree. Record evidence in implementation doc.
- [x] T3.2: Manual on-machine validation: build CLI, run default `ai-devkit capacity`, confirm `pi · OpenAI` rows beside `pi · z.ai`; redacted output in implementation doc.
- [x] T3.3: Update implementation/testing/planning docs (phases 5-8 flow), grep branch for token substrings (no-leak audit), final PR (dev-pr conventions; do NOT merge).

## Dependencies

- T1.1 → T1.2 → T2.2 (parser before callers). T2.1 independent of T1 but ordered after for clean diffs.
- T3.* depend on all of M1+M2.
- External: none beyond the already-validated wham endpoint (fixtures only from here).

## Timeline & Estimates

- Single-session execution: T1 ~30%, T2 ~50%, T3 ~20%. No calendar dates (autonomous run).

## Risks & Mitigation

- wham shape drift vs fixture → mitigated: parser is defensive (null-tolerant), fixtures from verified live payload.
- Codex test coupling to internal parse functions → mitigated: re-export shims keep imports stable.
- Coverage thresholds (70% floor, new code ~100%) → mitigated: classification branches enumerated in testing doc.
- Epoch-ms vs seconds misclassification → heuristic + JWT fallback, explicit tests for both magnitudes.
- Token leakage → fake fixtures, no-leak assertions, pre-PR grep audit.

## Progress Log

- 2026-09-27: Initial plan created from requirements/design/testing docs. All testing scenarios mapped to tasks (T1.1↔wham parsing, T2.1↔resolution/staleness, T2.2↔classification/no-leak, T3.1↔gates, T3.2↔manual e2e).
- 2026-09-27: M1 done. T1.1 wham.ts + wham.test.ts (TDD red→green; note: scoped-window ids reject spaces by safeIdentifier design — test uses `code-review`). T1.2 codex.ts delegates parsing; exported `resetTime`/`safeIdentifier` from wham.ts because codex's CLI-fallback path uses them (caught by 3 codex CLI-fallback tests going red). All 50 capacity tests + 682 agent-manager tests green; hooks passed.
- 2026-09-27: M2 done via TDD. T2.1: `resolveOpenAiCredential` union + tiers, not-found message now mentions pi login (CLI test uses injected errors — unaffected; openai.test.ts regex updated accordingly). T2.2: OAuth probe with wham fetch/classification; stale no-fetch; JWT fallback. Coverage: openai.ts & wham.ts 100% lines; remaining uncovered branches are the never-network `?? globalThis.fetch`/default-timeout injections, consistent with codex.ts/zai.ts norms. 70 capacity tests + full agent-manager suite green; hooks passed.
- 2026-09-27: M3 done. No-leak audit clean (branch diff + tracked files scanned against real token values). Final gates fresh on HEAD: lint 0 / npm test 0 / e2e 0 (42). Rebased (up to date) onto origin/main, pushed, PR opened: https://github.com/codeaholicguy/ai-devkit/pull/254 (left unmerged for review). Lifecycle complete.
Loading
Loading