What happened
Prod content rebuild (single slug abap-environment-create-rap-business-events) failed fatally: run
[publish-content] #672 short-circuit disengaged: cannot reach /content/source-hashes: Error: HTTP 500 [status=500]
Fatal: HTTP 500: {"error":"Pool resource could not be acquired within 1s"}
The srv recovered within minutes — a later GET /content/source-hashes returned 200. So this was transient HANA connection-pool exhaustion, not a content or DB fault: a momentary load spike (rebuild fetches source-hashes + nav metadata for ~1400 tutorials + the publish PUT, colliding with live traffic / scheduled jobs) starved the pool, and the acquire timed out.
Root cause
- No DB pool config in the project → CAP/@cap-js/hana defaults, incl.
acquireTimeoutMillis: 1000 (the "within 1s"). Aggressive under load.
scripts/publish-content.ts has no retry/backoff on a 500, so a single transient pool-timeout aborts the whole rebuild. (The source-hashes fetch already fails soft — "short-circuit disengaged" → full delta — but the final publish PUT is fatal.)
Proposed fix
- Retry-with-backoff (e.g. 3 attempts, jittered) in
publish-content.ts on 500 whose body matches Pool resource could not be acquired (and similar transient acquire/timeout errors).
- Consider raising
cds.requires.db.pool.acquireTimeoutMillis (and/or max) for the srv so brief contention waits instead of erroring.
Impact
Prod rebuilds intermittently abort under load; requires manual re-run. Recurs whenever the pool briefly saturates.
What happened
Prod content rebuild (single slug
abap-environment-create-rap-business-events) failed fatally: runThe srv recovered within minutes — a later
GET /content/source-hashesreturned 200. So this was transient HANA connection-pool exhaustion, not a content or DB fault: a momentary load spike (rebuild fetches source-hashes + nav metadata for ~1400 tutorials + the publish PUT, colliding with live traffic / scheduled jobs) starved the pool, and the acquire timed out.Root cause
acquireTimeoutMillis: 1000(the "within 1s"). Aggressive under load.scripts/publish-content.tshas no retry/backoff on a 500, so a single transient pool-timeout aborts the whole rebuild. (The source-hashes fetch already fails soft — "short-circuit disengaged" → full delta — but the final publish PUT is fatal.)Proposed fix
publish-content.tson 500 whose body matchesPool resource could not be acquired(and similar transient acquire/timeout errors).cds.requires.db.pool.acquireTimeoutMillis(and/ormax) for the srv so brief contention waits instead of erroring.Impact
Prod rebuilds intermittently abort under load; requires manual re-run. Recurs whenever the pool briefly saturates.