Skip to content

fix(publish-content): retry transient HANA pool-timeout 500s (#2286) - #2290

Merged
jung-thomas merged 1 commit into
DEVfrom
fix/2286-publish-begin-retry
Sep 14, 2026
Merged

jung-thomas merged 1 commit into
DEVfrom
fix/2286-publish-begin-retry

Conversation

@jung-thomas

Copy link
Copy Markdown
Contributor

Fixes #2286.

Root cause

beginSession (scripts/publish-content.ts) was the only publish HTTP call in the happy path with neither withRetry nor a try/catch. appendBatch, commit, and the three render-* phases are all wrapped; begin was not. So a transient HANA pool-acquire 500 (Pool resource could not be acquired within 1s) thrown by beginSession bubbled straight to the top-level main().catch → Fatal: HTTP 500, aborting the whole prod rebuild. This exactly matches the reported trace: the #672 short-circuit soft-fails first (disengaged), then the fatal 500.

The delta-detection fetchRemoteHashes (/content/hashes) had the same class of gap: a transient 500 there exits 1 before a session is even opened.

Fix

  1. Wrap beginSession in withRetry (5 attempts, [2000,5000,10000,20000], ±20% jitter) — same posture as the existing commit retry. 500/502/503/504 are already classified transient in publish-retry.ts, so no body-matching is needed.
  2. Wrap the delta fetchRemoteHashes in withRetry (4 attempts, [1000,3000,6000], jitter) inside its existing try/catch.
  3. Add cds.requires.db.pool config (acquireTimeoutMillis: 5000, max: 100, min: 0, fifo: true) for hana [hybrid]/[production]. There was no pool config → CAP/@cap-js/hana defaults incl. acquireTimeoutMillis: 1000 (the "within 1s"). Raising it lets brief contention wait instead of erroring. (Config shape verified against CAP docs.)

Tests

scripts/__tests__/publish-client.test.ts:

  • beginSession under withRetry recovers from a transient pool-timeout 500 (fails once → succeeds; 2 fetch calls).
  • beginSession under withRetry does not retry a permanent 409 (1 fetch call).

Full scripts/__tests__ suite: 937 passing.

beginSession was the one publish HTTP call in the happy path with neither
withRetry nor a try/catch, so a transient HANA pool-acquire 500 ('Pool
resource could not be acquired within 1s') bubbled to the top-level Fatal
catch and aborted the whole prod rebuild. Wrap begin (and the delta-detection
/content/hashes fetch) in withRetry with jittered backoff, mirroring the
existing commit posture; 500/502/503/504 are already classified transient.

Also add cds.requires.db.pool config (acquireTimeoutMillis 5000, max 100) for
hana [hybrid]/[production] so brief pool contention waits instead of erroring
at the 1s default.

Tests: beginSession under withRetry recovers from a transient pool-timeout
500 and does not retry a permanent 409.
@jung-thomas
jung-thomas merged commit ad82dac into DEV Sep 14, 2026
6 checks passed
@jung-thomas
jung-thomas deleted the fix/2286-publish-begin-retry branch September 14, 2026 16:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant