Skip to content

e2e: three TestTUIE2E subtests are red on dev, none new — a teardown race and two waits on a card foot the model does not always draw #1048

Description

@AbirAbbas

What happened

2026-09-15, dev@e2ae913b7 (and reproduced one commit earlier at 62bc4b7c5), make test-e2e-tui (go test -tags e2e -count=1 -timeout 40m -v -run '^TestTUIE2E$' ./internal/e2e/) on linux/amd64 with the profile's OpenRouter key, model deepseek/deepseek-v4-flash. 14 of 17 subtests pass; these three fail:

--- FAIL: TestTUIE2E/ask_here_end_to_end (36.19s)
    testing.go:1464: TempDir RemoveAll cleanup: unlinkat /tmp/TestTUIE2Eask_here_end_to_end4199179322/001/home/v3: directory not empty
--- FAIL: TestTUIE2E/a_refused_task_proposal_draws_no_schema_sentence (110.41s)
    refusedargs_e2e_test.go:58: waited 1m30s for ["1 yes, set it up · 0 no · c change"] and never saw it.
--- FAIL: TestTUIE2E/space_in_the_task_room_pages_the_card (114.52s)
    tui_e2e_test.go:1928: waited 30s for any of ["enter open its room" " · waiting in this conversation · alt+a"] and saw none.
    tui_e2e_test.go:1960: waited 30s for ["new conversation"] and never saw it.

In all three the law the subtest exists for passed and the failure is beside it:

  • ask_here_end_to_end: every screen assertion passed; something under the throwaway CODEAF_HOME is still writing home/v3 when t.TempDir cleans up. Intermittent: 3/3 green on re-run at e2ae913, 1 of 3 red at the parent with the same line.
  • a_refused_task_proposal_draws_no_schema_sentence: the log records the screen after the turn carries no repair sentence and Invalid arguments: is never drawn. The model declined to re-propose and ended the turn asking the person (not started · the call was refused, · stuck: propose_task was sent the same wrong argument 2 times, then prose ending "Which way?"), so there was no card and no foot to wait for.
  • space_in_the_task_room_pages_the_card: space paged the record exactly as pgdown does passed. The first wait met a third landing shape the test's comment does not accept — roster finished today · 1 folded away / ▸ completed hello.txt file write task, foot enter go to that conversation. The second wait's /history page reads tasks · 1 chat · 1 subtask · $0.0027 but lists only ✓ hello file under finished today: the untitled chat row is counted in the header and not drawn in the 14-row frame. That last part is a product question, not a test one.

Replication

Deterministic (no model). None yet for the first; for the other two the screens above are the evidence. A stub-model rig that ends the turn with a question instead of a card reproduces the second's shape.

Field (real models).

make build
go test -tags e2e -count=1 -timeout 40m -v \
  -run 'TestTUIE2E/(ask_here_end_to_end|a_refused_task_proposal_draws_no_schema_sentence|space_in_the_task_room_pages_the_card)' \
  ./internal/e2e/ > e2e.log 2>&1

Needs a key by any road config.APIKeyAt reads and tmux; about five minutes and a few cents. The two card-foot waits fire on most runs; the teardown race on roughly one in three.

Where

  • internal/e2e/tui_e2e_test.go, testAskHere (the TempDir cleanup) and the two waits in testTaskRoomKeepsSpace (search enter open its room and new conversation).
  • internal/e2e/refusedargs_e2e_test.go (search 1 yes, set it up).
  • The tasks page header count versus its drawn rows: internal/tui3 (search finished today).

The fix

  • The teardown: whatever still writes under home/v3 after the surface exits is stopped or awaited before the test returns; a TempDir that cannot be removed is a leak, not noise.
  • The two card-foot waits: wait on the fact the subtest is about (the turn ended without the repair sentence; the room was reached by whichever foot the landing drew), never on one spelling of a card foot a model may not draw. Widening the timeout is not a fix.
  • The /history header: a row the header counts is a row the page draws, or the header does not count it.

Acceptance

  • e2e: the three subtests above green three runs in a row on a clean dev, with -count=1.
  • e2e: make test-e2e-tui reports ok for the whole TestTUIE2E.
  • Unit: the tasks page's header count equals the rows it can draw for a fixture with one untitled chat and one finished task.
  • The manual page that covers the tasks page quotes any new wording, and the change entry's invalidates names what people believed before.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:testsThe suite itself — flakes, harnesses, laws, CI redsbugSomething the code does that it should notsev:papercutA wording, a hint, a small wrongness that costs a moment

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions