Repository navigation
service-job: type: 'once' schedules on DbJobAdapter get no leader election either — same routing limb as #13686, deliberately left out of its scope #13918
Description
Activity
- addedbugSomething isn't workingSomething isn't workingpriority:p2Medium: important, M3Medium: important, M3
on Aug 31, 2026 裁决:修 ——
once委托 leader-elected 路径;崩溃语义裁 at-most-once(维护者 2026-09-01,总监批 #24)项目总监席 · session
session_01KGtaLpkW1mycWgkbSb3H6t· 维护者对本批逐字:「同意」。- 修法照卡:
DbJobAdapter.schedule()把type:'once'委托给this.cron(CronJobAdapter.schedule()已有once分支,走同一条runScheduled()+lock.acquire发射路径 —— service-job: DbJobAdapter 的 interval 型调度在多副本下无 leader-election(cron 型有)—— #2219 声明的 interval 半边缺失,竞态实锤重复执行 #13686 对 interval 的同款,一枝之隔);inner侧走register()存不armed,保住trigger()/replay()/getExecutions()/listJobs(); - 崩溃语义裁定:at-most-once per cluster 即可 —— 今天单节点
setTimeout本就不落盘、崩溃同样丢,收敛到 leader-elected 不使任何场景变差(今天的现实是「每副本各跑一遍」的重复灾);「锁成功才释放」的重投机制 ⛔ 不建(无实测消费者,不为设想场景造持久化);语义写进 docblock 一句; - 消费者普查随实施顺手做(grep
type: 'once'注册点),结果记 PR 正文 —— 零消费者也照修(小、关洞、保险); - 复现钉照卡的 repro sketch(两个
DbJobAdapter栈共享一把锁,今天两行sys_job_run,修后一行); - Clause-②:预期 no(services 实现面),实施者按实际 diff 复declare。
状态转移(同笔)
needs-user-decision→pm:queue;bugpriority:p2domain:services不动。
Generated by Claude Code
- 修法照卡:
Claim:
domain:servicesPM seat, sessionsession_01AUF1NoViznQK32gqpK8wS8(GitHubos-sales), 2026-09-02 ~12:36Z.- Session:
session_01AUF1NoViznQK32gqpK8wS8 - Branch:
claude/issue-13918-once-schedule-leader-election - Worktree:
/home/user/objectstack-issue-13918(dev-created, worktree-first) - Domain:
domain:services(packages/services/service-job) - File surface:
packages/services/service-job/src/db-job-adapter.ts(schedule()— thetype: 'once'limb; theintervallimb landed by service-job: DbJobAdapter 的 interval 型调度在多副本下无 leader-election(cron 型有)—— #2219 声明的 interval 半边缺失,竞态实锤重复执行 #13686 is the template one branch over) and, only if theoncebranch there needs its docblock sentence,packages/services/service-job/src/cron-job-adapter.ts(comment only); tests underpackages/services/service-job/src/*.test.ts(natural home: a sibling ofdb-job-adapter.interval-leader.test.ts); one changeset (@objectstack/service-job). ⛔ Not touched:packages/spec/**,content/docs/releases/**, the lock implementation, any re-arm / re-try mechanism (ruled out: at-most-once). - Serial constraints cleared: every open PR branch diffed against its merge-base at ~12:33Z — no open PR holds any file under
packages/services/service-job/. - Container & model: in-process
os-devsubagent, model opus.node scripts/pm/dispatch-gates.mjs --tierfloor for services paths: no path-derived mandate; Clause-② judged from content — expectednoper the ruling's point 5 (services implementation face), the dev re-declares from the actual diff. - Ruling of record: maintainer 2026-09-01 via the director seat, 13918#issuecomment-5494528879 — 「同意」 on the five points (delegate
onceto the leader-electedcronpath withinner.register()keepingtrigger()/replay()/getExecutions()/listJobs()intact; crash semantics at-most-once per cluster, one docblock sentence, ⛔ no re-arm/persistence; consumer census oftype: 'once'registrations recorded in the PR body, fix regardless of a zero; the card's repro sketch as the pin — two adapters, one lock, twosys_job_runrows today, one after; Clause-② expectedno).
Race re-read follows this comment before the dev is launched.
Generated by Claude Code
- Session:
Dev interruption record: the dev subagent working this card was terminated by the platform's session usage limit (HTTP 429, "session limit resets 14:20 UTC") at ~13:08Z, after reproducing P3 (the handler ran twice across two replicas) and before the fix. Not a maintainer stop and not a dead claim: the branch
claude/issue-13918-once-schedule-leader-electionis pushed at270c6355eand the worktree/home/user/objectstack-issue-13918holds further local work. The same agent was resumed at ~14:22Z with an idempotent order (re-read state, mergeorigin/main2a2653619, continue from the fix; thescripts/engine-double-contract.pinned.jsonchange in its local commit is to be kept only if the gate requires it and declared as an adjacent mechanical fix, merge-tree checked against open PR #14528 which edits the same ledger). Claim, branch and file surface unchanged.
Generated by Claude Code
- added a commit that references this issue
on Sep 2, 2026 os-dev-report
{ "issue": 13918, "status": "done", "branch": "claude/issue-13918-once-schedule-leader-election", "pr": "https://github.com/objectstack-ai/objectstack/pull/14618", "head_sha": "51fb2de6f", "premise_still_valid": true, "summary": "DbJobAdapter.schedule() routed `type: 'once'` registrations to the inner IntervalJobAdapter — a bare setTimeout with no cluster lock — so a one-shot job ran once per replica instead of once per cluster, the last limb left after #13686 did the same for `interval`. One condition now routes `once` to `this.cron` (CronJobAdapter, whose own `once` branch already fires through the leader-electing runScheduled()) whenever a cron adapter is assembled, with inner.register() keeping the registration so trigger()/replay()/getExecutions()/listJobs() are unaffected; with no cron adapter assembled the inner.schedule() fallback is unchanged. Crash semantics are at-most-once per cluster per the 2026-09-01 ruling, stated in the schedule() docblock; no re-arm, retry or persistence was built. Twelve pins added in a new db-job-adapter.once-leader.test.ts, plus a patch changeset. Merged origin/main (2a2653619) into the branch; merge-tree against the current origin/main (2514d49f3) is clean. Note for the PM: this run was killed by the HTTP 429 session limit at ~13:08Z and resumed at ~14:22Z on the same claim, branch and worktree — no work was lost and nothing was redone.", "hypotheses": { "H1": "holds — the whole fix is the `once` limb mirroring the `interval` limb: `(schedule.type === 'interval' || schedule.type === 'once') && this.cron`, then cron.schedule + inner.register. The no-cron assembly keeps inner.schedule, pinned unchanged.", "H2": "holds — trigger() (while a peer denies the lock), replay() + getExecutions(), and listJobs() all answer for a delegated `once` job; three pins.", "H3": "holds — cancel() before `at` gives zero executions, zero lock acquires, empty listJobs() and sys_job.active === false.", "H4": "holds — at-most-once by construction: the lease is released by the existing runScheduled finally, no re-arm exists and none was added, and one docblock paragraph states the semantics. The diff adds no timer, no retry and no store." }, "premise_checks": { "P1": "holds — `once` fell through to the else limb and reached this.inner.schedule(...) on origin/main.", "P2": "holds — CronJobAdapter.schedule() still has its `type === 'once' && schedule.at` branch arming setTimeout to runScheduled, and runScheduled still takes lock.acquire('job:' + name, { waitMs: 0 }).", "P3": "holds — reproduced BEFORE any source edit: the card's two-replica/one-lock pin executed the handler 2 times instead of 1.", "P4": "holds — packages/spec/** and content/docs/releases/** untouched; the diff is four files." }, "tests": "All runs under scripts/pm/os-verify-lock.sh, exit codes captured before any pipe, on head 51fb2de6f (post-merge of origin/main). RED half, before the source edit: `vitest run src/db-job-adapter.once-leader.test.ts` -> 'Test Files 1 failed (1) / Tests 4 failed | 8 passed (12)', 'AssertionError: one deadline must execute the job once across the cluster, not once per replica: expected \"vi.fn()\" to be called 1 times, but got 2 times'. GREEN, whole package: `pnpm --filter @objectstack/service-job test` -> 'Test Files 10 passed (10) / Tests 106 passed (106)'. Typecheck: `pnpm --filter @objectstack/service-job exec tsc --noEmit --listFiles` exit 0, and coverage MEASURED rather than assumed — the --listFiles output (407 files) names both edited files, so the clean typecheck really does cover the new test file.", "ablation": "On the committed tree, mutation = restore the pre-fix routing. Mutation proved on disk before anything was measured, by occurrence counts anchored on the exact text plus the blob hash: PRE fixed-form 1 / broken-form 0 / hash 4d58eb1fa64a5b342643ab5eea7c2ab46354705b (= the HEAD blob); POST fixed-form 0 / broken-form 1 / hash 6e2c49ae1803e7fc275e5d346725b3680e9d00b4. No rebuild is owed on either leg and none was done: the pins import the subject relatively (from './db-job-adapter.js'), so vitest resolves it to TypeScript source, never through the package exports to dist/ — demonstrated in this run, since the suite went red then green across a source edit with no service-job build in between. Mutated leg exit 1: 'Tests 4 failed | 8 passed (12)', reporting all three of the card's facts in one run (handler called 2 times not 1; fence.acquire called 0 times — unrouted, the lock is never consulted at all; and 'expected [ 2 rows ] to have a length of 1 but got 2', which is the ruling's point-4 'two sys_job_run rows today, one after'). Mutated-leg CONTROL, the #13686 interval pins: exit 0, 'Tests 10 passed (10)' — the mutation touches only the once limb. Restore leg proved by state, not by exit code: hash back to 4d58eb1f... and `git diff HEAD` empty for the file; restored leg exit 0, 'Tests 12 passed (12)'. The script carried `trap ... EXIT INT TERM` with absolute paths so a container cap kill could not leave a mutated tree.", "gates": "Derived with `node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack --commands` on the FINAL head (4 paths vs merge base 2a2653619), re-derived after the ledger row entered the diff — that added seven gates, which were then run too. All 44 run on that head: 41 exit 0, 0 red, 3 exit 3 = NOT MEASURED by their own printed verdicts (check-test-completeness 'PREREQUISITE NOT MET — grades a saved turbo run test log'; check:dual-build-cjs-loads 'PREREQUISITE NOT MET — reads built output, and some package has no dist/'; check:type-check-debt 'PREREQUISITE NOT MET — --re-measure refuses without the built workspace closure'). check:type-check-coverage itself is green. check:engine-double-contract is green after the ledger row, verdict line: 'update doubles: 347 in 313 test file(s) — 247 pinned to ObjectQL.update's dispatch predicate, 100 in the shrink-only baseline'. Repo-wide `pnpm lint` was NOT run locally and no narrowing is claimed for it — CI owns it; this is the 'not run' case, not a proven narrowing.", "clause_2": "no — declared from the actual diff, not from the expectation. `git diff -U0 origin/main...HEAD` filtered to added/removed lines containing 'export' returns nothing; the only occurrences are hunk-header context for the unchanged `export class DbJobAdapter`. No export added, removed or renamed; no accept-set move; packages/spec/** untouched. Matches the ruling's point 5.", "adjacent_change_declared": "scripts/engine-double-contract.pinned.json — outside the claimed file surface, kept only because the gate demanded it and declared in the PR body with the gate's own verdict line ('x RETAINED [update]: ... pins 1 engine double(s) that the pinned ledger does not record ... Run `node scripts/check-engine-double-contract.mjs --write` and commit'). Regenerated with that exact command: '694 (file, verb) row(s), 1 added or grown, 0 lost' — coverage growth only, the grow-only direction. PR #14528 also edits this ledger; `git merge-tree --write-tree --name-only origin/main HEAD` against the current origin/main (2514d49f3) lists no file, so there is no overlap today.", "consumer_census": "NOT zero — four live production registration paths reach IJobService.schedule with { type: 'once', at }: service-automation wait-node.ts:259 (a wait node arming its timer resume, one per suspended flow run), wait-node.ts:440 (rearmSuspendedWaitTimers on cold boot — one per suspended run on EVERY replica's boot, the sharpest of the four), trigger-schedule schedule-trigger.ts:140 (a schedule-triggered flow declaring `at`), and runtime job-schedule.ts:62 via app-plugin.ts:1017 (an app-declared once job). Remaining hits are tests, docs and changelog prose. Positive control: the same method run for `type: 'interval'` returns the registrations #13686 was about (plugin-approvals approvals-plugin.ts:329, plugin-reports reports-plugin.ts:155), so a zero would have been a real zero. Full table in the PR body.", "mcp_calls": "3 — create_pull_request, issue_write (the finding), add_issue_comment (this report). Everything else went through git and unauthenticated repo-scoped REST reads, which this seat probed working (HTTP 200) before choosing the channel.", "open_questions": [], "out_of_scope_findings": [ "filed as #14619: service-job scheduler leader election excludes for the DURATION OF THE FIRE, not for the deadline — runScheduled releases the lease in `finally`, so the mutual-exclusion window equals the handler's runtime and replica clock skew larger than it defeats election. Pre-existing, shared by all three schedule types, neither introduced nor worsened here; labelled `finding` + `domain:services`, unassigned. Dedup: 480 open issues pulled from the REST list endpoint and grepped locally over titles AND bodies, with a positive control; nearest neighbour #14501 records across-replica duplication as handled by #13686 and does not discuss the lease window." ] }Generated by Claude Code
Generated by Claude Code
PM ACCEPT — PR #14618 at
51fb2de6f(Clause-②no, verified on the tree)Collected the
os-dev-reportand verified every load-bearing claim againstorigin/main2514d49f3and the branch, not against the report.Ruling conformance (2026-09-01, 13918#issuecomment-5494528879). Point 1 is the whole code change: the non-comment delta in
packages/services/service-job/src/db-job-adapter.tsis one line —schedule.type === 'interval' && this.cronbecomes(schedule.type === 'interval' || schedule.type === 'once') && this.cron, sooncetakes the samecron.schedule+inner.registerpath theintervallimb has taken since #13686. Point 2 (at-most-once per cluster, no re-arm, no persistence, semantics in the docblock) is honoured: the diff adds no timer, no retry and no store. Point 3's census is in the PR body and is not a zero — four live registration paths, the sharpest being the wait-node's cold-boot re-arm, which every replica ran. Point 4's repro is the pin, measured red before the source edit (handler called twice, twosys_job_runrows) and green after. Point 5 re-declared from the diff.Seat verification (measured, not read).
- File surface: exactly 4 files — the adapter, the new
db-job-adapter.once-leader.test.ts, thepatchchangeset for@objectstack/service-job, and the declared adjacentscripts/engine-double-contract.pinned.jsonrow. - Clause-② no:
git diff -U0 origin/main...HEAD | grep -E '^[+-].*\bexport\b'returns nothing on the branch. No export added, removed or renamed;packages/spec/**andcontent/docs/releases/**untouched. Noneeds:contract-reviewis owed on either carrier. - Changeset:
patch, and its text states the multi-replica behaviour change plus the at-most-once semantics in plain words, with the affected consumer packages named. - Serial constraints:
git merge-tree --write-tree --name-only origin/main HEADlists no file. The engine-double ledger is the one shared path in the diff; PR fix(plugin-sharing): let system writes materialize sharing rules — drop the isSystem skips in bindRuleHooks (#13533) #14528, the other open PR in this lane, does not touch it (measured by diffing its branch againstorigin/main), and CI's ownNo other open PR may claim the same single-writer pathis green. - Log levels: no new log site of any level. The deliberate asymmetry —
intervalwarns on the cron-less fallback andoncedoes not — is argued in the docblock from frequency (onceregistrations are per-occurrence, so the same line would be a per-run flood on a hot path). Accepted as reasoned, not as an omission.
Out-of-scope finding filed by the dev: #14619 (the lease is released in
runScheduled'sfinally, so the mutual-exclusion window equals the handler's runtime; pre-existing and shared by all three schedule types). Left with triage for first-touch grading.Landing: CI on
51fb2de6fwas still running when this was written (Build Core, Dogfood Verify CLI, Type Check source gates, Check Changeset, the queue guards all green; Test Core, the Dogfood gates,Lint & Repo Gatesand the remaining Type Check jobs in progress). On all green the seat flips the PR to ready, arms auto-merge and posts landing provenance; on MERGED the card closes andpm:dispatchedcomes off.
Generated by Claude Code
- File surface: exactly 4 files — the adapter, the new
Found while implementing #13686 (interval leader election). Filed rather than fixed there: #13686's scope, its field evidence and #2219's declared capability are all about
cron/interval, and widening the routing to a third schedule type is a behaviour change the reporter's cluster cannot confirm.The gap
DbJobAdapter.schedule()decides which adapter owns a scheduled fire, and only the adapter it picks decides whether that fire is leader-elected:type: 'cron'→this.cron(CronJobAdapter) — fires throughrunScheduled(), which takeslock.acquire('job:' + name, { waitMs: 0 })and skips when a peer holds it ✓type: 'interval'→this.cronas of service-job: DbJobAdapter 的 interval 型调度在多副本下无 leader-election(cron 型有)—— #2219 声明的 interval 半边缺失,竞态实锤重复执行 #13686 ✓type: 'once'→this.inner(IntervalJobAdapter) — a baresetTimeout, no lock anywhere in that file ✗So on a multi-replica deployment a one-shot job registered with
{ type: 'once', at }runs once per replica, not once per cluster. Same failure shape as #13686 and, for a one-shot, arguably a worse one: there is no later tick during which a de-duplication marker could win, so every replica's copy lands in the same short window.The fix is already sitting there
CronJobAdapter.schedule()handlestype: 'once'itself (else if (schedule.type === 'once' && schedule.at)) and arms it withsetTimeout(() => { void this.runScheduled(name); }, delay)— the same leader-elected fire path its cron and interval limbs use. So the repair is the one #13686 applied to the interval limb, one branch over: delegateoncetothis.cronwhen one is assembled, and register it oninnerviaIntervalJobAdapter.register()(stores without arming a timer) sotrigger()/replay()/getExecutions()/listJobs()are unaffected.What needs deciding before someone writes that
Not mechanical, which is why this is an issue and not a follow-up commit:
type: 'once'against a multi-replica assembly was not measured — worth measuring before spending the change, per the startup-scope-discipline axis.oncethere is no next tick: the lease expires with the work undone and nothing re-arms it, so "exactly once per cluster" becomes "at most once per cluster". That may be fine, or it may want the lock to be released only on success — a decision, not a detail.Repro sketch
Same instrument as
packages/services/service-job/src/db-job-adapter.interval-leader.test.ts: twoDbJobAdapterstacks over one fake engine and one shared lock, a{ type: 'once', at: <now + tick> }registration on each, oneadvanceTimersByTimeAsync. Today that executes the handler twice and writes twosys_job_runrows.Landing site:
packages/services/service-job/src/db-job-adapter.ts(schedule()).