Skip to content
9 changes: 9 additions & 0 deletions docs/changes/unreleased/1612-standing-once-handoff.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
---
kind: fixed
title: one-time standing approvals preserve the action and report approval truthfully
pr: 1612
surface: [chat, engine]
invalidates:
- "Choosing Only now previously settled the card as done before work ran and returned a generic instruction. It now records approval without completion and hands the full approved action back with an explicit pending execution state."
- "A one-time scheduled task previously used reminder labels. Its card now says Run it then and describes running work; say-only reminders keep their existing wording."
---
2 changes: 1 addition & 1 deletion internal/manual/chat/asking-from-home.md
Original file line number Diff line number Diff line change
Expand Up @@ -307,7 +307,7 @@ on home says `? waiting on you` for as long as it does.

**The card stays after you answer it.** It does not disappear — it settles in place, greys
out, and its bottom edge carries what was decided in the same words a card in a conversation
uses: `Set it up · Mondays at 9am · set up`, `Only now, don't repeat · done now, nothing kept`,
uses: `Set it up · Mondays at 9am · set up`, `Only now, don't repeat · approved once, not scheduled`,
`Change… · you asked for something different`,
`not set up`, `ended · nothing was set up`. The answers go, so `1`, `3`, `0` and `o` are
ordinary characters again and can be typed into a follow-up. The only card that ever
Expand Down
20 changes: 16 additions & 4 deletions internal/manual/chat/keeping-an-eye.md
Original file line number Diff line number Diff line change
Expand Up @@ -939,10 +939,13 @@ piece of the ask still to do.
`the card was left unanswered — nothing was set up`. This is the opposite of a
task proposal, where silence starts the work: a task is bounded work somebody
is watching, and a standing item spends money at times nobody chose.
- **A "do it once" answer sets nothing up.** It answers
`Do it now as an ordinary step and report what happened. The person chose not to repeat it. Do not set it up again unless they ask. Do not investigate codeaf.`
and codeaf does the thing in front of you instead. A **one-off reminder's card
does not offer that answer** — see "Why is there no once on my reminder card".
- **A "do it once" answer approves immediate work, not a schedule.** The settled
card says `approved once, not scheduled`. That is an approval receipt, not
proof that the work has finished. The approved action, workspace, watch probe,
acceptance and spending limits are handed back to the current conversation;
it uses its ordinary tools and permissions and reports the actual result or
a blocker. No standing item is saved. A **one-off reminder's card does not
offer that answer** — see "Why is there no once on my reminder card".
- **It will not set a reminder for a moment that has already passed.** The stamp
is refused with the current time in it, and codeaf is asked to work it out
again from that.
Expand All @@ -959,3 +962,12 @@ piece of the ask still to do.
one: nothing that runs on its own may arm something else that runs on its own.
- **Nothing is armed by a matcher.** Nothing runs because a phrase looked like a
rule; every single one of these was a card you said yes to.

## Why does a one-time job say run it then instead of remind me?

A one-time card that will perform work says `wants to schedule work once`.
Its approval is `Run it then · <time>` with `Runs the work then. Nothing repeats.`
The decline is `Don't schedule it`: nothing is scheduled or run. It has no
`Only now, don't repeat` option, because this card approves the stated future
time. A say-only reminder still says `wants to remind you` and `Remind me`.
Both kinds show the action they will take before you approve.
22 changes: 18 additions & 4 deletions internal/session/answers.go
Original file line number Diff line number Diff line change
Expand Up @@ -417,16 +417,20 @@ func StandingOnceIsAnAnswer(item standing.Item) bool {
// The heads a standing card opens with. The person's own sentence is the next
// line, not this one: this line says what KIND of thing is being asked.
const (
StandingHeadReminder = "wants to remind you"
StandingHeadCheck = "wants to set up a repeating check"
StandingHeadWatch = "wants to watch for something"
StandingHeadRule = "wants to keep a rule"
StandingHeadReminder = "wants to remind you"
StandingHeadScheduledWork = "wants to schedule work once"
StandingHeadCheck = "wants to set up a repeating check"
StandingHeadWatch = "wants to watch for something"
StandingHeadRule = "wants to keep a rule"
)

// StandingHead is the card's first line for this item.
func StandingHead(item standing.Item) string {
switch item.CardKindOf() {
case standing.CardReminder:
if item.Does.Kind == standing.ActionTask {
return StandingHeadScheduledWork
}
return StandingHeadReminder
case standing.CardCheck:
return StandingHeadCheck
Expand Down Expand Up @@ -492,6 +496,16 @@ func StandingOptions(item standing.Item) []AnswerOption {
cadence := item.When.ShortWords()
switch item.CardKindOf() {
case standing.CardReminder:
if item.Does.Kind == standing.ActionTask {
yes := "Run it then"
if cadence != "" {
yes += standingCadenceMark + cadence
}
return []AnswerOption{
{Key: "1", Label: yes, Consequence: "Runs the work then. Nothing repeats."},
{Key: StandingNoKey, Label: "Don't schedule it", Consequence: "Nothing is scheduled or run.", Safe: true},
}
}
yes := "Remind me"
if cadence != "" {
yes = "Remind me " + cadence
Expand Down
18 changes: 18 additions & 0 deletions internal/session/answers_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -696,3 +696,21 @@ func TestAnAnswerLeftWhileTheWindowWasShutIsAppliedAsItOpens(t *testing.T) {
t.Fatalf("the doorstep was left behind: %v", err)
}
}

func TestScheduledTaskApprovalNamesWorkNotReminder(t *testing.T) {
item := standing.Item{When: standing.When{Kind: standing.WhenAt, Words: "tomorrow at 8"}, Does: standing.Action{Kind: standing.ActionTask, Brief: "write report"}}
if got := StandingHead(item); got != "wants to schedule work once" {
t.Fatal(got)
}
options := StandingOptions(item)
if len(options) != 2 || !strings.HasPrefix(options[0].Label, "Run it then") || options[0].Consequence != "Runs the work then. Nothing repeats." || options[1].Label != "Don't schedule it" {
t.Fatalf("options: %+v", options)
}
if StandingOnceIsAnAnswer(item) {
t.Fatal("future scheduled work unexpectedly offers now")
}
item.Does = standing.Action{Kind: standing.ActionSay, Say: "time for report"}
if StandingHead(item) != StandingHeadReminder || !strings.HasPrefix(StandingOptions(item)[0].Label, "Remind me") {
t.Fatal("say-only reminder changed")
}
}
63 changes: 44 additions & 19 deletions internal/session/standing_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -571,38 +571,63 @@ func TestStandingSurfacesAStoreThatWouldNotWrite(t *testing.T) {
}
}

// "DO IT ONCE" CREATES NOTHING and tells the model to do the thing here.
// The once decision carries the actual approved action, not a completion claim.
func TestStandingOnceCreatesNothing(t *testing.T) {
store := newFakeStanding(t)
completer := &scriptedCompleter{steps: []step{
standCall("s1", aReminder()),
finalText("doing it now"),
standCall("s1", `{"op":"propose","words":"weekly pantry report","when":{"kind":"every","every":"168h"},"does":{"kind":"task","brief":"write shopping.md from pantry.csv","acceptance":"report lists quantities"},"rails":{"per_run_usd":0.3}}`),
finalText("continuing the approved work"),
}}
agent := standingAgent(t, completer, store, nil)

events, err := agent.Submit(context.Background(), "remind me at 6")
events, err := agent.Submit(context.Background(), "prepare my weekly pantry report")
if err != nil {
t.Fatalf("Submit: %v", err)
t.Fatal(err)
}
var shown standing.Item
collected := drainAnsweringStanding(t, events, func(event Event) {
shown = event.Standing.Item
agent.ResolveStanding(event.Standing.ID, StandingAnswer{Once: true})
})
if len(store.created) != 0 {
t.Fatalf("a once answer created %d items", len(store.created))
t.Fatalf("once saved %d items", len(store.created))
}
output := toolOutput(t, collected, "stand")
for _, want := range []string{
"Do it now as an ordinary step and report what happened.",
"The person chose not to repeat it.",
"Do not set it up again unless they ask.",
"Do not investigate codeaf.",
} {
if !strings.Contains(output, want) {
t.Errorf("tool result missing %q\n%s", want, output)
}
output := strings.Split(toolOutput(t, collected, "stand"), "\nnow:")[0]
var result struct {
Decision string `json:"decision"`
Execution string `json:"execution"`
Saved bool `json:"standing_saved"`
Approved standing.Item `json:"approved_action"`
}
if err := json.Unmarshal([]byte(output), &result); err != nil {
t.Fatalf("handoff: %v: %s", err, output)
}
if result.Decision != "run_once_now" || result.Execution != "pending" || result.Saved {
t.Fatalf("decision: %+v", result)
}
want, _ := json.Marshal(shown)
got, _ := json.Marshal(result.Approved)
if string(got) != string(want) {
t.Fatalf("approved action changed: got %s want %s", got, want)
}
if result.Approved.Does.Brief != "write shopping.md from pantry.csv" || result.Approved.Rails.PerRunUSD != 0.3 {
t.Fatalf("missing action or limits: %+v", result.Approved)
}
}

func TestStandingRejectsOnceForReminder(t *testing.T) {
store := newFakeStanding(t)
completer := &scriptedCompleter{steps: []step{standCall("s1", aReminder()), finalText("not scheduled")}}
agent := standingAgent(t, completer, store, nil)
events, err := agent.Submit(context.Background(), "remind me later")
if err != nil {
t.Fatal(err)
}
collected := drainAnsweringStanding(t, events, func(event Event) { agent.ResolveStanding(event.Standing.ID, StandingAnswer{Once: true}) })
if len(store.created) != 0 {
t.Fatal("forged once created an item")
}
if strings.Contains(output, "\u2014") || strings.Contains(output, "\u2013") {
t.Errorf("tool result still has a dash: %q", output)
if out := toolOutput(t, collected, "stand"); !strings.Contains(out, "does not offer doing it once") {
t.Fatal(out)
}
}

Expand Down
27 changes: 17 additions & 10 deletions internal/session/steer_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -867,13 +867,15 @@ func TestASteerAdoptsAnOldBashAndItsExitArrivesLater(t *testing.T) {

func TestASteerWaitsForAYoungBash(t *testing.T) {
t.Parallel()
release := filepath.Join(t.TempDir(), "release-young-bash")
completer := &scriptedCompleter{steps: []step{
func(context.Context, []ai.Message) (*ai.Response, error) {
return toolResponse("young-bash", "bash", `{"command":"sleep 0.5; echo young-finished"}`), nil
},
bashCall("young-bash", "while [ ! -f "+shellQuoted(release)+" ]; do sleep 0.01; done; echo young-finished"),
func(context.Context, []ai.Message) (*ai.Response, error) { return textResponse("done"), nil },
}}
agent, _ := newTestAgent(t, completer, nil)
// The command stays observable until released; its age is a separate
// input, not a race against a half-second sleep on a busy test machine.
advanceSteerAge(agent, 0)
turn := mustSubmit(t, agent, "run the quick check")
waitFor(t, "young foreground bash to start", func() bool { return len(agent.inFlightBash.snapshot()) == 1 })
agent.mu.Lock()
Expand All @@ -882,18 +884,23 @@ func TestASteerWaitsForAYoungBash(t *testing.T) {
if generation != nil {
t.Fatal("a young bash still had a model generation to cut")
}
began := time.Now()
steered := mustSteer(t, agent, "then read the result")
collect(t, turn)
events := collect(t, steered)
if elapsed := time.Since(began); elapsed < 200*time.Millisecond {
t.Fatalf("young bash landed in %s, want the batch to finish first", elapsed)
select {
case event := <-steered:
if got := steerLanding([]Event{event}); got != "waiting for the running step" {
t.Fatalf("steer landing while bash is held = %q", got)
}
case <-time.After(10 * time.Second):
t.Fatal("the steer was not accepted while the bash was held")
}
if list := agent.jobs.list(); list != "No background jobs." {
t.Fatalf("young bash became a job: %q", list)
}
if got := steerLanding(events); got != "waiting for the running step" {
t.Fatalf("steer landing = %q", got)
writeFile(t, release, "finish now")
collect(t, turn)
collect(t, steered)
if got := roleText(completer.request(1), "tool"); !strings.Contains(got, "young-finished") {
t.Fatalf("the resumed turn did not receive the foreground result: %q", got)
}
}

Expand Down
3 changes: 1 addition & 2 deletions internal/session/task_phase_test.go
Original file line number Diff line number Diff line change
@@ -1,7 +1,6 @@
package session

import (
"context"
"strings"
"testing"
"time"
Expand Down Expand Up @@ -330,7 +329,7 @@ func TestTheHandoverSaysItIsBriefingAWorkerWhileItWritesTheBrief(t *testing.T) {
ran := make(ranNodes, 2)
stubbedGraph(agent, func(node *TaskNode) { ran <- node })

events, err := agent.Submit(context.Background(), "work through the four things I listed and report back")
events, err := agent.Submit(watchedContext(agent), "work through the four things I listed and report back")
if err != nil {
t.Fatalf("Submit: %v", err)
}
Expand Down
19 changes: 15 additions & 4 deletions internal/session/task_run_belt_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -164,11 +164,22 @@ func registerBeltRunEngine(t *testing.T, engine RunEngine) {
// directory's removal, and lost it under load as "directory not empty".
func endBeltRun(t *testing.T, agent *Agent, double *beltRunDouble) {
t.Helper()
agent.beltMu.Lock()
run := agent.beltRun
agent.beltMu.Unlock()
close(double.release)
beltRunWaitFor(t, "the run to end", func() bool {
agent.beltMu.Lock()
defer agent.beltMu.Unlock()
return agent.beltRun == nil
if run == nil {
return
}
// releaseBeltRun clears the agent's pointer before it closes the store.
// The run's completion channel covers that final write/close as well.
beltRunWaitFor(t, "the run and its store to close", func() bool {
select {
case <-run.over:
return true
default:
return false
}
})
}

Expand Down
5 changes: 3 additions & 2 deletions internal/session/task_run_orphan_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -210,8 +210,9 @@ func TestClosingUnderAProgramsRunEndsItInItsStoreFirst(t *testing.T) {
t.Fatalf("the second run is on root %q with brief %q", spec.Store.RootID(), spec.Store.Task(spec.Store.RootID()).Description)
}
endBeltRun(t, again, second)
close(first.release)
<-first.finished
// Start returning is not the end of the owner's bookkeeping. Join the
// closed agent's run before the fixture removes its store and workspace.
endBeltRun(t, agent, first)
}

// THE RESTORE ROAD: a conversation read back from disk whose run row comes
Expand Down
31 changes: 26 additions & 5 deletions internal/session/tools_standing.go
Original file line number Diff line number Diff line change
Expand Up @@ -494,11 +494,10 @@ func (a *Agent) standPropose(ctx context.Context, parsed standArguments) (string
}
switch {
case answer.Once:
// Nothing is created and nothing is scheduled. The person wanted the
// action, not the arrangement. The result says the next step in so
// many words, because "nothing was set up" sent a model off to read
// this program's source looking for a reminder that was never missing.
return "Do it now as an ordinary step and report what happened. The person chose not to repeat it. Do not set it up again unless they ask. Do not investigate codeaf.", false, nil
if !StandingOnceIsAnAnswer(item) {
return "this card does not offer doing it once now; nothing was set up or run", true, nil
}
return standingOnceHandoff(item), false, nil
case !answer.Approved:
if correction := strings.TrimSpace(answer.Change); correction != "" {
// AND THE CORRECTION MAY BE ABOUT ANY OF IT. The card's one change
Expand Down Expand Up @@ -533,6 +532,28 @@ func (a *Agent) standPropose(ctx context.Context, parsed standArguments) (string
return line, false, nil
}

// standingOnceHandoff records an approval, not execution. Keep the entire
// approved action on the tool boundary: the continuation must not reconstruct
// its brief, workspace, acceptance or watch probe from the scheduling request.
// Work stays in the ordinary turn under its existing tool permissions.
func standingOnceHandoff(item standing.Item) string {
payload := struct {
Decision string `json:"decision"`
Execution string `json:"execution"`
StandingSaved bool `json:"standing_saved"`
Instruction string `json:"next_step"`
Approved standing.Item `json:"approved_action"`
}{
Decision: "run_once_now",
Execution: "pending",
Instruction: "The person approved this action once now. This is not a decline. Execute approved_action in its workspace using the ordinary tools, respecting its grant, rails and the current permissions. For a watch, check its probe or condition once before the action. Ignore the future schedule: do not save or re-propose it. Report actual results or a concrete blocker; approval alone does not mean the work ran.",
Approved: item,
}
// standing.Item contains only validated JSON data from this tool's input.
raw, _ := json.Marshal(payload)
return string(raw)
}

// standingRatifiedLine is what a model is told the instant something stands,
// and it is the WHOLE of what it may say next.
//
Expand Down
4 changes: 3 additions & 1 deletion internal/session/tools_tasks_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -572,7 +572,9 @@ func TestHandingAYourCallToTheModelAndItsResolveAreOneRoad(t *testing.T) {
// what leaves a your-call landing in the PERSON'S hands: with nobody watching
// the model holds every one of them by policy ([Agent.settlePolicy]) and the
// card this test is about is never drawn.
agent, _ := newTestAgent(t, &scriptedCompleter{}, func(config *Config) {
// Keep the woken turn open until the test resolves the decision; a finished
// turn correctly hands unresolved work back to the person.
agent, _ := newTestAgent(t, &holdingCompleter{}, func(config *Config) {
config.Place, config.AskConsent = Place{Dir: mine}, true
})
graph := stubbedGraph(agent, func(node *TaskNode) {
Expand Down
2 changes: 1 addition & 1 deletion internal/tui3/homeband_answer.go
Original file line number Diff line number Diff line change
Expand Up @@ -438,7 +438,7 @@ func (a *app) answerHere(question session.PresenceQuestion, key string) (tea.Cmd
}
switch {
case action.Standing.Once:
return a.answerStanding(action.Standing, standOnceDone, standOnceWord), true
return a.answerStanding(action.Standing, standOnceApproved, standOnceWord), true
case action.Standing.Approved:
return a.answerStanding(action.Standing, standSetWord, standYesWord), true
default:
Expand Down
12 changes: 6 additions & 6 deletions internal/tui3/standing.go
Original file line number Diff line number Diff line change
Expand Up @@ -166,11 +166,11 @@ const (
// The verdicts a settled card keeps. They are sentences and not states,
// because the row is read once, later, by somebody reconstructing what
// happened.
standSetWord = "set up"
standChangedWord = "you asked for something different"
standOnceDone = "done now, nothing kept"
standNoWord = "not set up"
standExpiredWord = "ended · nothing was set up"
standSetWord = "set up"
standChangedWord = "you asked for something different"
standOnceApproved = "approved once, not scheduled"
standNoWord = "not set up"
standExpiredWord = "ended · nothing was set up"

// WHAT EACH ANSWER COSTS, one clause apiece, drawn beside its own word on
// the block's card ([session.AnswerOption.Consequence]).
Expand Down Expand Up @@ -633,7 +633,7 @@ func standVerdictOf(card *standingCard, answer session.Answer) (string, string)
word := standAnswerWord(card.item, key)
switch key {
case session.StandingOnceKey:
return standOnceDone, word
return standOnceApproved, word
case session.StandingNoKey:
// THE DECLINE KEEPS NO WORD BESIDE ITS VERDICT. `not set up · no` is the
// same fact twice, and the verdict is the half that says what happened.
Expand Down
Loading
Loading