crew: count a pinned checker in the panel's coverage warning - #1501
Merged
Merged
Conversation
… open their pages (#1231)
…operties, with the verdict basis recorded (#1220) The worker declares its checks, and a holds verdict rests on a recorded, auditable, zero-exit run of every declared command. The worker records each command's own exit status, an absent status reads as unknown and earns nothing, and a record opens with a line that marks it as one that records exits, so a declared check that was never observed is refused, including when the record is missing or empty. A record written before exits were recorded still holds by reading and names the checks it did not observe. A does-not-hold verdict is ungated, a task with no declaration gets a reading verdict, and the verdict basis is persisted. One quote-aware reader of "is this one command" lives in internal/approval and is shared by the proposal door, the checker's runner gate and the store.
… callback ports (#1237)
…cadence, and the harness that measured it (#1235)
…ver the conversation (#1244)
…ntroducing itself again (#1236)
…her's keys (#1248) A profile write now takes a cross-process lock on a stable config.json.lock for the whole read, copy and rename, inside the in-process write mutex. The lock rides internal/filelock, so it holds on every platform the repo builds for. A write waits at most two seconds for another process and then fails with a plain timeout error. The lock file is never unlinked: a crashed holder is released by the operating system closing its file, so there is no stale lock to reclaim. Tests run two real processes forced to contend and show both keys survive, a bounded timeout, and recovery after a holder exits without unlocking. Those forced tests are unix-only; their clocks are sanity bounds, not load bounds.
…ump (#1249) Both tests dumped the whole environment through the capped output collector and read variables out of it. On a machine with a large ambient environment the dump is cut before the TMUX_TMPDIR line, so the test read it absent while the job shell had it, and the two checks for stripped variables passed without proving anything. The shell now prints the three variables on one short line with an explicit word for unset, through the same login shell road the product uses, and the test fails if that line is missing from the captured output. Test only; the product was already correct on every platform.
…llows on a beat (#1251) The run's summary is a read nobody pressed for, and its refresh can wait on a model for seconds. It was asked through the one ordered line that carries a person's gestures to the engine, so a press on a run's row, or a stop, waited behind it: 2.4 to 10.1 seconds on a real screen, 0.09 to 0.18 after. The summary now goes beside the line. What may leave the line is decided by property: the ask was not a gesture and nothing a person does next depends on the engine having seen it first. A law lists every door that stands in the line and fails when one is added without a reason. An open page on a task that can still move was re-read on every paint tick once the last read answered, 509 wire reads in ninety seconds for one page. It now follows on the rail's own beat, counted from the last read for any reason, and a page on an ended or held task is never read again. The follow stays in the line because its fold replaces the page and must not overtake a note.
A run under the task belt was started on a context nothing could cut, and no cancel was kept. Its rows wear task numbers, but a stop by number went only to the task graph, which has never held a run's rows. Measured on the real binary in hosted mode: the stop card answered that there was no such task, the run's own page answered with the store's sentence about who owns the root, ctrl+c and /quit closed the window, and the run carried on to its own landing every time. The id a surface already sends is now resolved to whoever owns the row, so an older window stops a run through a newer engine. A stop of a run ends the run's task and everything open under it in one store write, then cuts the run's context so workers and their calls end at once, then commits what was made on the run's own branch and gives the copy back. Nothing goes into the person's folder, no model turn is bought, and the person is told where the work is and how to bring it in or drop it. A hand-off that joined a run is stopped alone. A stopped run buys no further summary. On the surface: x on the run's own page raises the same stop card, x typed while that page is still loading reaches it too, the page no longer offers a hold the store would refuse, and ctrl+c is read above both the loading page and the open page. A law lists every place that publishes a running row and fails until each has a test that a stop on that row ends its owner. What ctrl+c, /quit and a closed window do to a live run is unchanged here.
FINDINGS.md at the repository root is a working note from the cell behind the unread config key notice. It is not product documentation and nothing reads it. A cell's findings belong in its run folder, outside the repository.
The test printed the whole PATH through the capped output collector and read its first entry. On a machine with a long PATH the capture can be cut, and a cut capture could read as a pass or as a false failure. The shell now prints the first PATH entry on one short marker line, and the test fails when that line is missing from the capture. Test only.
) make check built and tested only the host platform, so a symbol defined for one platform family passed the gate and would have broken another platform's build; only a nightly would have caught it. A new build-cross step reads the shipped platform list from the release workflow, builds ./... for each with the shared build cache, names the target that failed, and prints its own duration: about 11 seconds on a warm box, at most about 111 on a cold one.
…es (#1256) The standing ticker was a package-level goroutine started once and never stopped, so it could write under the home's v3/standing folder after Close returned; one of the temp directory cleanup failures in the full check named that folder. The ticker now belongs to the process: Close stops it and waits for its loop, and a second process in the same binary starts its own. Quitting never waits out background work: the ticker holds one context that the stop cancels, and a pass derives its 120 second ceiling from it, so a pass in flight ends at once when the process closes. A cancelled pass leaves no half-written file, because standing writes are temp then rename or single short appends and a pass checks its context between steps. Tests: no write under v3/standing after Close on the real close path, a second process starts its own ticker, and Close returns promptly with a pass held in flight, ordered by channels.
The side list read a run's rows from the engine inside the frame: on the first reading, when one of the window's own rows moved, and when the beat came due. Over a connection that read is a call to another process, so a slow link froze the frame for as long as the read took. Measured on Spark with 250 ms injected on the read, real remote client and engine: 254 to 258 ms per frame on trunk, 0.1 to 0.7 ms after, with no read inside the frame. The rows are now asked for from the update loop, held on the surface per conversation, and the frame draws what is held; a verb that lands while a read is out is answered by one more read. The law that lists every door still on the loop could not see a door reached through the plan reader; it now can, so every plan door is under it. The hosted drive script opens the side list with its own key when the copied profile starts with it hidden.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comments in internal/crewroute (route.go, classify.go, prior.go, route_test.go) now state the rule the router follows, not the figures behind it: narrow fixes use the cheapest qualified crew, and open-ended work gets a strong checker. The knee comment says what the knee does (a step up is bought only above [Knee] points per dollar) and no longer lists per-crew deltas and costs. Four test failure messages in route_test.go now say what was expected instead of citing evidence. prior.json keeps all of its data; its source and costs_source notes now read "measured crews per class of work; see the Pareto Crewing paper". The spend guard's comment and its test lead-in describe an overrun without the incident's figures. The manual's crew estimate paragraph and the #1436 change entry refer to the router's table and the design paper instead of a trial. No behaviour or user-facing string changes; gofmt and go vet are clean. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The paper's LaTeX source, bibliography, generated tables (tab-*.tex), figure PDFs (fig-*.pdf) and the Makefile that built them are removed. docs/design/model-pool keeps pareto-crewing.pdf, RUNBOOK.md and data/seed-cells.csv, which internal/pool/index/cmd/seedgen reads. Nothing else in the tree referenced the removed sources; only the removed files referenced each other. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…5 limit Replace the shipped per-model table and hand-set scoring in internal/crewroute with a learned prior read from catalog metadata. crewroute: - prior.json now holds fitted weights only: a joint Gaussian over ability and catalog features (log prices, log context, release date, open weights, the three published indexes, arena Elo), family offsets, a per-(class, seat) linear link from ability to quality, per-seat token shapes and cost scales. No per-model rows ship. - quality() conditions on the fields a row has and returns a mean and a standard deviation; missing fields widen the variance. credible() needs a finite-variance score (evidence ratio) and an upper bound over the floor. - Cost comes from catalog prices x seat shape x class scale x the install factor. costFloor is the cheapest credible priced model; unpriced models are weighed at a priced model of equal ability. - Remove Snapshot/Measured/IsMeasured, seatWeight, unseenMargin, unseenShrink, the measured tie-break and the price ceiling. Every catalog model is a candidate. - Request.Learned applies per-install offsets; Request.TaskCap keeps the pick under the per-task limit (estimate x install factor), --best included. Note reads "held under the $X task limit". - Model gains ArenaElo and Released; ability memoised per model. router/config: - CrewRecord.Learned and CrewLog.Quality: accepted, kept, redo and failed outcomes move a per-(class, seat, model) offset by a bounded step (rate 0.1, bound 1 quality point), recorded on the decision row. - catalog reads `created`; crewModelOf sets Released from it or the slug date. Rows with indexes added are re-scored on the next refresh. - Per-task spend limit models.crew.task_cap (default $5): pre-call guard on every task call, stop line "this task reached its $5 limit · raise it in /crew", /crew cap row "per task $5 · daily $X", `/crew cap task <$>`, `codeaf do` holds it and -yes-spend does not lift it. tui3: Spending tab's per task row reads the /crew limit. Crew panel fixtures carry release dates and Elo; snapshots follow the new picks. Tests: replace the table-reproduction and effort-evidence tests with behaviour tests (stronger-indexed similar-priced model wins a seat, no metadata is not picked unless pinned, install evidence moves but does not freeze a score, bounded outcome learning, --best under the task limit, decision under 2 ms, weights under 2 MB). Manual and changelog updated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ach on one-paragraph asks crewroute ability: - A row that publishes any index (AA intelligence/coding/agentic, arena Elo) is conditioned on those indexes only; price, context, date, licence and family no longer move it. A row with none is conditioned on context, date and licence plus the family offset, and its mean is capped at the population mean with the excess added to the variance. Price is not an ability feature. prior.json v2 carries index_features/base_features. - Seats score the solve rate at theta - 1 sd (riskKappa); quality is no longer clamped to [0, 10], so ordering is preserved. - The checker needs mean u >= u_floor whenever any candidate reaches it (seatCredible); Gaps uses u + sd against the open-ended u_ref. - Link refitted with seat pooling within a class (support slopes shrunk toward the worker's), so no seat has a zero slope; --best buys the top model in every seat under the task cap. classify: - Reach signals read the title when a task has no body. - New reach signal: security fixes (traversal, injection, bypass, ...); "N repros" counts; keepPassing allows words between "existing" and "tests". - Title failure words add wrongly/rejected/instead of; a rename or version bump reads as a simple fix. Tests: dominance property test (no dearer and >= on every shared index never loses, 400 random pairs plus the v4.1-flash/glm-5.3-flash case), fix checker not the cheapest thin model, checker floor, --best strongest per seat, index-less row capped at the population mean, one-paragraph reach (T03/T11 texts), defect words and renames. Knee test now asserts open-ended buys a stronger checker at prices a fix does not. Config and tui3 expectations follow the new picks. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…and cheap tolerance - checkerFirst: on open-ended and other work at λ > 0, when the planner and checker slopes are within one SD of each other, the stronger of the two support picks sits the checker seat and the weaker the planner. Pins are never moved; a model that cannot sit the other seat stays. --best (λ = 0) is unchanged. - seatCredible: the worker, like the checker, needs mean u >= u_floor whenever any candidate reaches it (abilityFloor). Fallback ladders are exempt, so a seat that cannot start still moves to a similar cost. - --cheap: strongerWithin trades the worker pick for the strongest eligible model costing at most 1.5x (cheapTolerance). The knee comparison for the effort note runs without it. Tests: the checker is never weaker than the planner at the knee on open-ended and other work, and a pinned planner stays; a cheap worker is credible (glm-5.3-flash over a coding-only model) and the tolerance buys a stronger worker at 1.2x cost but not at 2.4x. Manual and changelog updated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
dev now carries santos/dev2 (#1410) and #1388, #1426, #1440 and others. The PR's net change is taken from its fork point bdb08cf; the earlier dev2 merge and its revert are left out, so dev's skills (#1396) and worker step-boundary work stay as dev has them. Conflicts: config/crew.go, config/seats.go and tui3/crew.go take the routed-crew versions; /crew keeps dev's notice event. do.go keeps both the kept-record line and the crew outcome log. task_run_belt.go keeps both the machine-hold rail and the crew completer. home.md keeps dev's rows without the preset forms and adds the crew-panel row. truth_test.go keeps dev's new facts and counts the crew as three seats. back_test.go stays deleted as on dev. CHANGELOG.md keeps dev's released history. #1440 on the routed crew: reflex and small work keep their free defaults on a balance read as low; the three crew seats see the OpenRouter account as out of credit before any call (not re-probed per task) and route to free pools with the free-routes notice. Tests follow. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With HOME and CODEAF_HOME empty, all three seats are routed, nothing is pinned, the allowed rule is all, a task is held to $5 and there is no daily crew cap. With nothing routable, no crew seat falls back to a built-in model and resolving a crew is an error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Names the four kinds of work (bugfix, complex fix, open-ended, other), the candidates (every model a connected provider serves) and learned weights; draws the panel's cap row as "per task $5 · daily none" and says the seats' "usually" reads the last eight tasks; adds a route health section, the /crew cap task row, the per-task limit on codeaf do (-yes-spend does not lift it), and the crew's limits in LIMITS.md. openrouter-credits.md says how a low balance reaches the routed crew. The changelog entry names the classes, the #1440 change and the redo decay as the code has it. Onboarding comments no longer mention a crew step. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The two-millisecond budget was checked against one batch's mean, which on a loaded shared machine also counts scheduler preemption. It now takes the fastest of five batches of forty. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
crewroute.Gaps asked only whether any allowed, seatable model met the strong-checker test. A pinned checker always runs, so the answer ignored the seat that actually checks: a strong pin over a weak set (or a pin whose route is quarantined and so missing from the candidates) kept the "no strong checker among the models you allow" warning, and a weak pin over a strong set showed none. Gaps now takes the pinned seats. With the checker pinned, the pin alone is tested (read from the candidates by lineage when they carry it, else from the catalog figures config passes): strong gives no gap, weak gives "checker pinned to <model> · open-ended work will be checked weakly". An unpinned checker is judged over the allowed models as before. The checker is the only gap kind today, so there is nothing else to adjust. config.CrewGapsAt reads the pins and resolves each through the catalog. The manual names both warning lines. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
santoshkumarradha
marked this pull request as ready for review
September 25, 2026 14:59
Member
Author
|
@AbirAbbas small follow-up to #1436, ready for review. |
TestCrewSnapshotWeakPinnedChecker was pinned against the panel as it stood before #1518, whose cap row now says `crew daily cap none` and names the daily limit on a second line. Integrated onto dev, the warning under the rows was right and only the cap rows above it had moved, so the frame takes the rows TestCrewSnapshotPanel already pins. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Collaborator
|
Picking this up per your Slack note. I'm pushing three plain commits on top of
Checked in the real binary with a throwaway home and zero spend. All pass:
Should-fixes, not blocking:
|
AbirAbbas
approved these changes
Sep 27, 2026
AbirAbbas
left a comment
Collaborator
There was a problem hiding this comment.
Verified in the real binary and merged into dev; details in the comment above.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What was wrong.
crewroute.Gapsasked only whether any allowed, seatable model passed the strong-checker test. A pinned checker always runs, but the check never looked at it. Pinning a strong checker over a weak set, or pinning one whose route is quarantined and so absent from the candidates, leftno strong checker among the models you allow · open-ended work will be checked weaklyon the/crewpanel. Pinning a weak checker over a strong set showed no warning at all.Fix.
Gaps(candidates, pins map[Seat]Model): when the checker is pinned, only the pin is tested. It is read from the candidates by lineage when they carry it, the waypinnedreads it, and otherwise from the figures passed in. A strong pin gives no gap. A weak pin giveschecker pinned to <model> · open-ended work will be checked weakly. An unpinned checker is judged the same way as before. The checker is the only gap kind today, so no other kind needed the same change.config.CrewGapsAtreads the pins and resolves each one through the catalog, so the result does not depend on the pin's route health.models-and-cost.md) now names both warning lines. The changelog entry is1501-crew-pinned-checker-gap.md.Tests
crewroute/TestGapsCountAPinnedCheckercovers four cases: a strong pin over a weak set gives no gap, whether the candidates carry the pin or not; a weak pin over a strong set gives a gap that names the pin; with only other seats pinned, the unpinned checker behaves as before.config/TestTheGapCountsAPinnedCheckerworks end to end on a profile: a strong pin whose route is quarantined gives no gap, a weak pin gives the named line, and unpinning brings back today's answer.tui3/TestCrewSnapshotWeakPinnedCheckeris a panel snapshot with the named warning under the rows. The existingTestCrewSnapshotPanelalready covers a strong pin with no warning.All of these were run on
b1f125a, withOPENROUTER_API_KEYunset (#1489), and all passed:make vet fmt-checkgo testforcrewroute,config,tui3,manualandcmd/codeafmake manual-gates test-laws changelog-check test-packed-manualThe diff touches only those packages, so the full
make checkwas not needed.Stacked on #1436; retarget to dev once #1436 merges.
🤖 Generated with Claude Code