Skip to content

service-job: type: 'once' schedules on DbJobAdapter get no leader election either — same routing limb as #13686, deliberately left out of its scope #13918

Description

@os-steve

Found while implementing #13686 (interval leader election). Filed rather than fixed there: #13686's scope, its field evidence and #2219's declared capability are all about cron/interval, and widening the routing to a third schedule type is a behaviour change the reporter's cluster cannot confirm.

The gap

DbJobAdapter.schedule() decides which adapter owns a scheduled fire, and only the adapter it picks decides whether that fire is leader-elected:

So on a multi-replica deployment a one-shot job registered with { type: 'once', at } runs once per replica, not once per cluster. Same failure shape as #13686 and, for a one-shot, arguably a worse one: there is no later tick during which a de-duplication marker could win, so every replica's copy lands in the same short window.

The fix is already sitting there

CronJobAdapter.schedule() handles type: 'once' itself (else if (schedule.type === 'once' && schedule.at)) and arms it with setTimeout(() => { void this.runScheduled(name); }, delay) — the same leader-elected fire path its cron and interval limbs use. So the repair is the one #13686 applied to the interval limb, one branch over: delegate once to this.cron when one is assembled, and register it on inner via IntervalJobAdapter.register() (stores without arming a timer) so trigger() / replay() / getExecutions() / listJobs() are unaffected.

What needs deciding before someone writes that

Not mechanical, which is why this is an issue and not a follow-up commit:

  1. Is there a real consumer? feat(service-job): leader-elect scheduled cron/interval jobs across the cluster #2219 declared cron/interval. Whether any shipped or app-level code registers type: 'once' against a multi-replica assembly was not measured — worth measuring before spending the change, per the startup-scope-discipline axis.
  2. Crash semantics differ for a one-shot. For interval, a leader that dies mid-fire costs one tick and the next tick re-elects. For once there is no next tick: the lease expires with the work undone and nothing re-arms it, so "exactly once per cluster" becomes "at most once per cluster". That may be fine, or it may want the lock to be released only on success — a decision, not a detail.

Repro sketch

Same instrument as packages/services/service-job/src/db-job-adapter.interval-leader.test.ts: two DbJobAdapter stacks over one fake engine and one shared lock, a { type: 'once', at: <now + tick> } registration on each, one advanceTimersByTimeAsync. Today that executes the handler twice and writes two sys_job_run rows.

Landing site: packages/services/service-job/src/db-job-adapter.ts (schedule()).

Activity

  1. huangyiirene commented on Sep 1, 2026

    @huangyiirene
    Collaborator

    裁决:修 —— once 委托 leader-elected 路径;崩溃语义裁 at-most-once(维护者 2026-09-01,总监批 #24)

    项目总监席 · session session_01KGtaLpkW1mycWgkbSb3H6t · 维护者对本批逐字:「同意」。

    1. 修法照卡:DbJobAdapter.schedule() 把 type:'once' 委托给 this.cron(CronJobAdapter.schedule() 已有 once 分支,走同一条 runScheduled() + lock.acquire 发射路径 —— service-job: DbJobAdapter 的 interval 型调度在多副本下无 leader-election(cron 型有)—— #2219 声明的 interval 半边缺失,竞态实锤重复执行 #13686 对 interval 的同款,一枝之隔);inner 侧走 register() 存不armed,保住 trigger()/replay()/getExecutions()/listJobs();
    2. 崩溃语义裁定:at-most-once per cluster 即可 —— 今天单节点 setTimeout 本就不落盘、崩溃同样丢,收敛到 leader-elected 不使任何场景变差(今天的现实是「每副本各跑一遍」的重复灾);「锁成功才释放」的重投机制 ⛔ 不建(无实测消费者,不为设想场景造持久化);语义写进 docblock 一句;
    3. 消费者普查随实施顺手做(grep type: 'once' 注册点),结果记 PR 正文 —— 零消费者也照修(小、关洞、保险);
    4. 复现钉照卡的 repro sketch(两个 DbJobAdapter 栈共享一把锁,今天两行 sys_job_run,修后一行);
    5. Clause-②:预期 no(services 实现面),实施者按实际 diff 复declare。

    状态转移(同笔)

    needs-user-decision → pm:queue;bug priority:p2 domain:services 不动。


    Generated by Claude Code

  2. self-assigned this
    on Sep 2, 2026
  3. os-sales commented on Sep 2, 2026

    @os-sales
    Collaborator

    Claim: domain:services PM seat, session session_01AUF1NoViznQK32gqpK8wS8 (GitHub os-sales), 2026-09-02 ~12:36Z.

    • Session: session_01AUF1NoViznQK32gqpK8wS8
    • Branch: claude/issue-13918-once-schedule-leader-election
    • Worktree: /home/user/objectstack-issue-13918 (dev-created, worktree-first)
    • Domain: domain:services (packages/services/service-job)
    • File surface: packages/services/service-job/src/db-job-adapter.ts (schedule() — the type: 'once' limb; the interval limb landed by service-job: DbJobAdapter 的 interval 型调度在多副本下无 leader-election(cron 型有)—— #2219 声明的 interval 半边缺失,竞态实锤重复执行 #13686 is the template one branch over) and, only if the once branch there needs its docblock sentence, packages/services/service-job/src/cron-job-adapter.ts (comment only); tests under packages/services/service-job/src/*.test.ts (natural home: a sibling of db-job-adapter.interval-leader.test.ts); one changeset (@objectstack/service-job). ⛔ Not touched: packages/spec/**, content/docs/releases/**, the lock implementation, any re-arm / re-try mechanism (ruled out: at-most-once).
    • Serial constraints cleared: every open PR branch diffed against its merge-base at ~12:33Z — no open PR holds any file under packages/services/service-job/.
    • Container & model: in-process os-dev subagent, model opus. node scripts/pm/dispatch-gates.mjs --tier floor for services paths: no path-derived mandate; Clause-② judged from content — expected no per the ruling's point 5 (services implementation face), the dev re-declares from the actual diff.
    • Ruling of record: maintainer 2026-09-01 via the director seat, 13918#issuecomment-5494528879 — 「同意」 on the five points (delegate once to the leader-elected cron path with inner.register() keeping trigger()/replay()/getExecutions()/listJobs() intact; crash semantics at-most-once per cluster, one docblock sentence, ⛔ no re-arm/persistence; consumer census of type: 'once' registrations recorded in the PR body, fix regardless of a zero; the card's repro sketch as the pin — two adapters, one lock, two sys_job_run rows today, one after; Clause-② expected no).

    Race re-read follows this comment before the dev is launched.


    Generated by Claude Code

  4. os-sales commented on Sep 2, 2026

    @os-sales
    Collaborator

    Dev interruption record: the dev subagent working this card was terminated by the platform's session usage limit (HTTP 429, "session limit resets 14:20 UTC") at ~13:08Z, after reproducing P3 (the handler ran twice across two replicas) and before the fix. Not a maintainer stop and not a dead claim: the branch claude/issue-13918-once-schedule-leader-election is pushed at 270c6355e and the worktree /home/user/objectstack-issue-13918 holds further local work. The same agent was resumed at ~14:22Z with an idempotent order (re-read state, merge origin/main 2a2653619, continue from the fix; the scripts/engine-double-contract.pinned.json change in its local commit is to be kept only if the gate requires it and declared as an adjacent mechanical fix, merge-tree checked against open PR #14528 which edits the same ledger). Claim, branch and file surface unchanged.


    Generated by Claude Code

  5. added a commit that references this issue on Sep 2, 2026
  6. os-sales commented on Sep 2, 2026

    @os-sales
    Collaborator

    os-dev-report

    {
      "issue": 13918,
      "status": "done",
      "branch": "claude/issue-13918-once-schedule-leader-election",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/14618",
      "head_sha": "51fb2de6f",
      "premise_still_valid": true,
      "summary": "DbJobAdapter.schedule() routed `type: 'once'` registrations to the inner IntervalJobAdapter — a bare setTimeout with no cluster lock — so a one-shot job ran once per replica instead of once per cluster, the last limb left after #13686 did the same for `interval`. One condition now routes `once` to `this.cron` (CronJobAdapter, whose own `once` branch already fires through the leader-electing runScheduled()) whenever a cron adapter is assembled, with inner.register() keeping the registration so trigger()/replay()/getExecutions()/listJobs() are unaffected; with no cron adapter assembled the inner.schedule() fallback is unchanged. Crash semantics are at-most-once per cluster per the 2026-09-01 ruling, stated in the schedule() docblock; no re-arm, retry or persistence was built. Twelve pins added in a new db-job-adapter.once-leader.test.ts, plus a patch changeset. Merged origin/main (2a2653619) into the branch; merge-tree against the current origin/main (2514d49f3) is clean. Note for the PM: this run was killed by the HTTP 429 session limit at ~13:08Z and resumed at ~14:22Z on the same claim, branch and worktree — no work was lost and nothing was redone.",
      "hypotheses": {
        "H1": "holds — the whole fix is the `once` limb mirroring the `interval` limb: `(schedule.type === 'interval' || schedule.type === 'once') && this.cron`, then cron.schedule + inner.register. The no-cron assembly keeps inner.schedule, pinned unchanged.",
        "H2": "holds — trigger() (while a peer denies the lock), replay() + getExecutions(), and listJobs() all answer for a delegated `once` job; three pins.",
        "H3": "holds — cancel() before `at` gives zero executions, zero lock acquires, empty listJobs() and sys_job.active === false.",
        "H4": "holds — at-most-once by construction: the lease is released by the existing runScheduled finally, no re-arm exists and none was added, and one docblock paragraph states the semantics. The diff adds no timer, no retry and no store."
      },
      "premise_checks": {
        "P1": "holds — `once` fell through to the else limb and reached this.inner.schedule(...) on origin/main.",
        "P2": "holds — CronJobAdapter.schedule() still has its `type === 'once' && schedule.at` branch arming setTimeout to runScheduled, and runScheduled still takes lock.acquire('job:' + name, { waitMs: 0 }).",
        "P3": "holds — reproduced BEFORE any source edit: the card's two-replica/one-lock pin executed the handler 2 times instead of 1.",
        "P4": "holds — packages/spec/** and content/docs/releases/** untouched; the diff is four files."
      },
      "tests": "All runs under scripts/pm/os-verify-lock.sh, exit codes captured before any pipe, on head 51fb2de6f (post-merge of origin/main). RED half, before the source edit: `vitest run src/db-job-adapter.once-leader.test.ts` -> 'Test Files 1 failed (1) / Tests 4 failed | 8 passed (12)', 'AssertionError: one deadline must execute the job once across the cluster, not once per replica: expected \"vi.fn()\" to be called 1 times, but got 2 times'. GREEN, whole package: `pnpm --filter @objectstack/service-job test` -> 'Test Files 10 passed (10) / Tests 106 passed (106)'. Typecheck: `pnpm --filter @objectstack/service-job exec tsc --noEmit --listFiles` exit 0, and coverage MEASURED rather than assumed — the --listFiles output (407 files) names both edited files, so the clean typecheck really does cover the new test file.",
      "ablation": "On the committed tree, mutation = restore the pre-fix routing. Mutation proved on disk before anything was measured, by occurrence counts anchored on the exact text plus the blob hash: PRE fixed-form 1 / broken-form 0 / hash 4d58eb1fa64a5b342643ab5eea7c2ab46354705b (= the HEAD blob); POST fixed-form 0 / broken-form 1 / hash 6e2c49ae1803e7fc275e5d346725b3680e9d00b4. No rebuild is owed on either leg and none was done: the pins import the subject relatively (from './db-job-adapter.js'), so vitest resolves it to TypeScript source, never through the package exports to dist/ — demonstrated in this run, since the suite went red then green across a source edit with no service-job build in between. Mutated leg exit 1: 'Tests 4 failed | 8 passed (12)', reporting all three of the card's facts in one run (handler called 2 times not 1; fence.acquire called 0 times — unrouted, the lock is never consulted at all; and 'expected [ 2 rows ] to have a length of 1 but got 2', which is the ruling's point-4 'two sys_job_run rows today, one after'). Mutated-leg CONTROL, the #13686 interval pins: exit 0, 'Tests 10 passed (10)' — the mutation touches only the once limb. Restore leg proved by state, not by exit code: hash back to 4d58eb1f... and `git diff HEAD` empty for the file; restored leg exit 0, 'Tests 12 passed (12)'. The script carried `trap ... EXIT INT TERM` with absolute paths so a container cap kill could not leave a mutated tree.",
      "gates": "Derived with `node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack --commands` on the FINAL head (4 paths vs merge base 2a2653619), re-derived after the ledger row entered the diff — that added seven gates, which were then run too. All 44 run on that head: 41 exit 0, 0 red, 3 exit 3 = NOT MEASURED by their own printed verdicts (check-test-completeness 'PREREQUISITE NOT MET — grades a saved turbo run test log'; check:dual-build-cjs-loads 'PREREQUISITE NOT MET — reads built output, and some package has no dist/'; check:type-check-debt 'PREREQUISITE NOT MET — --re-measure refuses without the built workspace closure'). check:type-check-coverage itself is green. check:engine-double-contract is green after the ledger row, verdict line: 'update doubles: 347 in 313 test file(s) — 247 pinned to ObjectQL.update's dispatch predicate, 100 in the shrink-only baseline'. Repo-wide `pnpm lint` was NOT run locally and no narrowing is claimed for it — CI owns it; this is the 'not run' case, not a proven narrowing.",
      "clause_2": "no — declared from the actual diff, not from the expectation. `git diff -U0 origin/main...HEAD` filtered to added/removed lines containing 'export' returns nothing; the only occurrences are hunk-header context for the unchanged `export class DbJobAdapter`. No export added, removed or renamed; no accept-set move; packages/spec/** untouched. Matches the ruling's point 5.",
      "adjacent_change_declared": "scripts/engine-double-contract.pinned.json — outside the claimed file surface, kept only because the gate demanded it and declared in the PR body with the gate's own verdict line ('x RETAINED [update]: ... pins 1 engine double(s) that the pinned ledger does not record ... Run `node scripts/check-engine-double-contract.mjs --write` and commit'). Regenerated with that exact command: '694 (file, verb) row(s), 1 added or grown, 0 lost' — coverage growth only, the grow-only direction. PR #14528 also edits this ledger; `git merge-tree --write-tree --name-only origin/main HEAD` against the current origin/main (2514d49f3) lists no file, so there is no overlap today.",
      "consumer_census": "NOT zero — four live production registration paths reach IJobService.schedule with { type: 'once', at }: service-automation wait-node.ts:259 (a wait node arming its timer resume, one per suspended flow run), wait-node.ts:440 (rearmSuspendedWaitTimers on cold boot — one per suspended run on EVERY replica's boot, the sharpest of the four), trigger-schedule schedule-trigger.ts:140 (a schedule-triggered flow declaring `at`), and runtime job-schedule.ts:62 via app-plugin.ts:1017 (an app-declared once job). Remaining hits are tests, docs and changelog prose. Positive control: the same method run for `type: 'interval'` returns the registrations #13686 was about (plugin-approvals approvals-plugin.ts:329, plugin-reports reports-plugin.ts:155), so a zero would have been a real zero. Full table in the PR body.",
      "mcp_calls": "3 — create_pull_request, issue_write (the finding), add_issue_comment (this report). Everything else went through git and unauthenticated repo-scoped REST reads, which this seat probed working (HTTP 200) before choosing the channel.",
      "open_questions": [],
      "out_of_scope_findings": [
        "filed as #14619: service-job scheduler leader election excludes for the DURATION OF THE FIRE, not for the deadline — runScheduled releases the lease in `finally`, so the mutual-exclusion window equals the handler's runtime and replica clock skew larger than it defeats election. Pre-existing, shared by all three schedule types, neither introduced nor worsened here; labelled `finding` + `domain:services`, unassigned. Dedup: 480 open issues pulled from the REST list endpoint and grepped locally over titles AND bodies, with a positive control; nearest neighbour #14501 records across-replica duplication as handled by #13686 and does not discuss the lease window."
      ]
    }

    Generated by Claude Code


    Generated by Claude Code

  7. os-sales commented on Sep 2, 2026

    @os-sales
    Collaborator

    PM ACCEPT — PR #14618 at 51fb2de6f (Clause-② no, verified on the tree)

    Collected the os-dev-report and verified every load-bearing claim against origin/main 2514d49f3 and the branch, not against the report.

    Ruling conformance (2026-09-01, 13918#issuecomment-5494528879). Point 1 is the whole code change: the non-comment delta in packages/services/service-job/src/db-job-adapter.ts is one line — schedule.type === 'interval' && this.cron becomes (schedule.type === 'interval' || schedule.type === 'once') && this.cron, so once takes the same cron.schedule + inner.register path the interval limb has taken since #13686. Point 2 (at-most-once per cluster, no re-arm, no persistence, semantics in the docblock) is honoured: the diff adds no timer, no retry and no store. Point 3's census is in the PR body and is not a zero — four live registration paths, the sharpest being the wait-node's cold-boot re-arm, which every replica ran. Point 4's repro is the pin, measured red before the source edit (handler called twice, two sys_job_run rows) and green after. Point 5 re-declared from the diff.

    Seat verification (measured, not read).

    • File surface: exactly 4 files — the adapter, the new db-job-adapter.once-leader.test.ts, the patch changeset for @objectstack/service-job, and the declared adjacent scripts/engine-double-contract.pinned.json row.
    • Clause-② no: git diff -U0 origin/main...HEAD | grep -E '^[+-].*\bexport\b' returns nothing on the branch. No export added, removed or renamed; packages/spec/** and content/docs/releases/** untouched. No needs:contract-review is owed on either carrier.
    • Changeset: patch, and its text states the multi-replica behaviour change plus the at-most-once semantics in plain words, with the affected consumer packages named.
    • Serial constraints: git merge-tree --write-tree --name-only origin/main HEAD lists no file. The engine-double ledger is the one shared path in the diff; PR fix(plugin-sharing): let system writes materialize sharing rules — drop the isSystem skips in bindRuleHooks (#13533) #14528, the other open PR in this lane, does not touch it (measured by diffing its branch against origin/main), and CI's own No other open PR may claim the same single-writer path is green.
    • Log levels: no new log site of any level. The deliberate asymmetry — interval warns on the cron-less fallback and once does not — is argued in the docblock from frequency (once registrations are per-occurrence, so the same line would be a per-run flood on a hot path). Accepted as reasoned, not as an omission.

    Out-of-scope finding filed by the dev: #14619 (the lease is released in runScheduled's finally, so the mutual-exclusion window equals the handler's runtime; pre-existing and shared by all three schedule types). Left with triage for first-touch grading.

    Landing: CI on 51fb2de6f was still running when this was written (Build Core, Dogfood Verify CLI, Type Check source gates, Check Changeset, the queue guards all green; Test Core, the Dogfood gates, Lint & Repo Gates and the remaining Type Check jobs in progress). On all green the seat flips the PR to ready, arms auto-merge and posts landing provenance; on MERGED the card closes and pm:dispatched comes off.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions