Skip to content

The -W gate is blind in ci.yml, and a warning-only cache failure discards a completed execution #271

Description

@mmcky

Summary

Two independent weaknesses in cache.yml turned a one-line content defect into four days without a usable execution cache, and hid it from the check that should have caught it first.

What happened

Four cache.yml runs failed or were cancelled between 2026-08-20 and 2026-08-24, while the last success was 2026-08-17. The publish build for publish-2026aug24 consequently fell back on the eight-day-old artifact, which shipped a site with inconsistent chapter numbering (filed separately).

The causes were not a single recurring one:

Run Date Cause
32331172292 2026-08-20T04:14Z FileNotFoundError: '_fonts/SourceHanSerifSC-SemiBold.otf' in learning_approximation.md, plus the hoist_failure warning below — fixed by #262
— 2026-08-20T22:56Z hoist_failure warning alone
32433407779 2026-08-21T00:38Z cancelled by hand two minutes in, not a failure
32686259809 2026-08-24T03:23Z hoist_failure warning alone — fixed by #265

The hoist_failure cause was one Sphinx warning promoted to an error by -W: lectures/hoist_failure.md:34: WARNING: Document headings start at H2, not H1 [myst.header]. The stale duplicated frontmatter block left in the body by the #260 resync parsed as body text, and its trailing --- turned the block into a setext H2 ahead of the real # 故障树不确定性.

Both underlying defects were lecture-content defects introduced by translate seed/forward operations, and both are now fixed. What follows is why they cost four days.

Defect 1 — the -W gate never re-reads, so it fires late and in the wrong workflow

cache.yml runs the sphinx-tojupyter build and then the HTML build with -W. Because the doctrees survive between them, the -W step re-reads nothing, so the warning surfaces only on a cold run. The same shape exists in publish.yml, where the rm -r _build/.doctrees line is commented out.

The consequence is that a content defect that should fail a PR's CI instead fails the weekly cold cache build, days later and far from the change that caused it.

Suggested fix: clear _build/.doctrees between the two builds so the -W step re-reads every source. Smallest change with the strongest guarantee. Alternatively drop -W from the tojupyter step and rely on a re-reading HTML step.

Defect 2 — a warning-only failure discards 106 minutes of successful execution

On the 2026-08-24 run, all 142 notebooks executed successfully (142 Executed notebook in lines, zero CellExecutionError) over 1h46m, and the artifact was then thrown away because the upload step does not run on failure.

Suggested fix: add if: success() || failure() to the Upload "_build" folder (cache) step. jupyter-cache does not store failed executions, so the artifact stays correct even when a lecture genuinely errors, and the run's red status remains the signal. This change alone would have made the 2026-08-24 publish a cache hit despite the hoist_failure defect.

Also noted

cache.yml still carries bare apt-get install steps with no timeout-minutes and no retry — the weakness PR #263 would have hardened before it was closed. Not implicated in any of these four failures (apt was green throughout: graphviz 31s, texlive 3m28s), but it remains the open item behind the documented hang class.

Refs: QuantEcon/project-translation#48 (delivery integrity); #262, #265 (the two content fixes); #263 (closed apt hardening).

Activity

  1. mmcky commented on Aug 24, 2026

    @mmcky
    ContributorAuthor

    Correction — Defect 1 named the wrong workflow, and the real evidence is better than what I filed.

    As written, Defect 1 said cache.yml "runs the sphinx-tojupyter build and then the HTML build with -W" and that surviving doctrees make the -W step re-read nothing. That is wrong about cache.yml. cache.yml has exactly one build step (Build HTML, line 39) and downloads no artifact — it is a cold build from scratch. That is precisely why it did catch the hoist_failure warning, correctly, and failed. cache.yml was doing its job.

    The two-build shape is in ci.yml — the PR gate — and that is where the blindness lives. The evidence is the CI run for #260, the PR that introduced the duplicated frontmatter:

    Time Event
    02:14:18 rm -rf _build/.doctrees — ci.yml's clear step runs
    02:14:20 tojupyter build starts: jb build lectures -n -W --keep-going --builder=custom --custom-builder=jupyter, under set -eo pipefail
    02:14:25 updating environment: [new config] 143 added, 0 changed, 0 removed — it reads everything
    03:12:24 lectures/hoist_failure.md:34: WARNING: Document headings start at H2, not H1 [myst.header]
    03:12:54 next step starts — the tojupyter step exited 0 despite -W and the warning
    03:12:59 HTML build: updating environment: 0 added, 0 changed, 0 removed — reads nothing
    03:14:49 build succeeded. — no ##[error] anywhere; preview reported pass

    So two things compound, and both need naming:

    1. -W is not enforced on the custom/jupyter builder. The warning was emitted under -n -W --keep-going with set -eo pipefail in force, and jb build ... --builder=custom --custom-builder=jupyter still exited 0. Whatever the mechanism inside jupyter-book, the practical effect is that -W on that step is decorative.
    2. The one builder that does enforce -W never sees a warning. The HTML build reuses the doctrees the tojupyter build just wrote, so it re-reads nothing and there is nothing left to warn about. Clearing doctrees once after the download does not help — by the time the HTML build runs, the tojupyter build has repopulated them.

    The fix follows from (2) and is one line: clear _build/.doctrees between the two builds in ci.yml, not only after the download. Then the -W HTML build re-reads every source and the gate actually gates. Worth also confirming (1) independently, because if the custom builder silently drops -W then the same hole exists anywhere that builder runs.

    Note this applies to publish.yml too — after #272 its HTML build will likewise re-read nothing, so its -W is equally decorative. That does not affect what #272 fixes (the environment is fresh because the restored one is deleted, so the toctree and numbering are correct), but publish is not a warning gate and should not be mistaken for one.

    Defects 2 and 3 as filed are unaffected and stand: the upload step is not if: always(), discarding a completed 106-minute execution; and the four failures had two distinct causes, not one.

    The body above has been amended. I am leaving the error visible here rather than editing it away, because "the weekly cache build is the thing that catches content defects" is exactly the wrong lesson to draw — the PR gate should have caught this four days earlier, and the reason it did not is now evidenced.

  2. changed the title [-]cache.yml: the -W gate never re-reads, and a warning-only failure discards a completed execution[/-] [+]The -W gate is blind in ci.yml, and a warning-only cache failure discards a completed execution[/+] on Aug 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions