fix(publish-content): retry transient HANA pool-timeout 500s (#2286) - #2290
Merged
Merged
Conversation
beginSession was the one publish HTTP call in the happy path with neither
withRetry nor a try/catch, so a transient HANA pool-acquire 500 ('Pool
resource could not be acquired within 1s') bubbled to the top-level Fatal
catch and aborted the whole prod rebuild. Wrap begin (and the delta-detection
/content/hashes fetch) in withRetry with jittered backoff, mirroring the
existing commit posture; 500/502/503/504 are already classified transient.
Also add cds.requires.db.pool config (acquireTimeoutMillis 5000, max 100) for
hana [hybrid]/[production] so brief pool contention waits instead of erroring
at the 1s default.
Tests: beginSession under withRetry recovers from a transient pool-timeout
500 and does not retry a permanent 409.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #2286.
Root cause
beginSession(scripts/publish-content.ts) was the only publish HTTP call in the happy path with neitherwithRetrynor a try/catch.appendBatch,commit, and the threerender-*phases are all wrapped;beginwas not. So a transient HANA pool-acquire 500 (Pool resource could not be acquired within 1s) thrown bybeginSessionbubbled straight to the top-levelmain().catch→Fatal: HTTP 500, aborting the whole prod rebuild. This exactly matches the reported trace: the#672short-circuit soft-fails first (disengaged), then the fatal 500.The delta-detection
fetchRemoteHashes(/content/hashes) had the same class of gap: a transient 500 there exits 1 before a session is even opened.Fix
beginSessioninwithRetry(5 attempts,[2000,5000,10000,20000], ±20% jitter) — same posture as the existingcommitretry.500/502/503/504are already classified transient inpublish-retry.ts, so no body-matching is needed.fetchRemoteHashesinwithRetry(4 attempts,[1000,3000,6000], jitter) inside its existing try/catch.cds.requires.db.poolconfig (acquireTimeoutMillis: 5000,max: 100,min: 0,fifo: true) for hana[hybrid]/[production]. There was no pool config → CAP/@cap-js/hana defaults incl.acquireTimeoutMillis: 1000(the "within 1s"). Raising it lets brief contention wait instead of erroring. (Config shape verified against CAP docs.)Tests
scripts/__tests__/publish-client.test.ts:beginSessionunderwithRetryrecovers from a transient pool-timeout 500 (fails once → succeeds; 2 fetch calls).beginSessionunderwithRetrydoes not retry a permanent 409 (1 fetch call).Full
scripts/__tests__suite: 937 passing.