Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
b617cc9
rescore_candidate: resolve the harbor_requirement build placeholder b…
varunursekar Sep 14, 2026
bb3a479
GAIA budget-blind build variant; rescore_candidate --build-file for r…
varunursekar Sep 15, 2026
5655c3a
rescore_candidate: --agent-env passthrough (harbor --ae); GAIA shell …
varunursekar Sep 15, 2026
6487e32
terminal-bench: non-functional shell seed (target-shell, build.shell.…
varunursekar Sep 15, 2026
5edfea4
terminal-bench shell seed: README variant table; shell-variant config…
varunursekar Sep 15, 2026
dc9536e
opencode as a full-harness seed for Terminal-Bench and GAIA
varunursekar Sep 15, 2026
3c693db
terminal-bench: register the opencode submodule (pinned e03db9bc); th…
varunursekar Sep 15, 2026
dbaf1b4
compiler: allow in-tree relative symlinks when snapshotting agent_rep…
varunursekar Sep 15, 2026
33631c2
gaia: e2e smoke variant of the opencode seed build
varunursekar Sep 15, 2026
0136d46
opencode seed: fetch the pinned Bun release with the standard library…
varunursekar Sep 15, 2026
4bc5b22
opencode seed: run against the public upstream from inside the task c…
varunursekar Sep 15, 2026
5fa2917
opencode builds: 4 h sandbox idle timeout — the verifier's test evalu…
varunursekar Sep 16, 2026
2b43d28
opencode builds: 90 min image build/start timeout
varunursekar Sep 16, 2026
aaa93a9
terminal-bench opencode seed: send the vendor-prefixed wire id (xai/g…
varunursekar Sep 16, 2026
08271f5
rescore_candidate --seed: skip node_modules/dist/.turbo when copying …
varunursekar Sep 16, 2026
ccf7fa7
rescore_candidate: --agent-setup-timeout-multiplier passthrough (larg…
varunursekar Sep 16, 2026
39ea0ec
CI: check out the vendored opencode submodules; the vendoring test as…
varunursekar Sep 17, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions .github/workflows/tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,16 @@ jobs:
steps:
- uses: actions/checkout@v4

# The opencode seeds vendor anomalyco/opencode as submodules and a config
# test checks the source is really there (an unpopulated submodule compiles
# to a hollow seed). Shallow, and only these two: the BrowseComp-Plus
# upstream corpus is not read by any test.
- name: Check out the vendored opencode source
run: |
git submodule update --init --depth 1 \
harness-opt-bench/terminal-bench/baseline/target-opencode/opencode \
harness-opt-bench/gaia/baseline/target-opencode/opencode

- name: Install uv
run: |
curl -LsSf https://astral.sh/uv/install.sh | sh
Expand Down
6 changes: 6 additions & 0 deletions .gitmodules
Original file line number Diff line number Diff line change
@@ -1,3 +1,9 @@
[submodule "browsecomp-plus-upstream"]
path = harness-opt-bench/browsecomp-plus/upstream
url = https://github.com/texttron/BrowseComp-Plus.git
[submodule "opencode-terminal-bench"]
path = harness-opt-bench/terminal-bench/baseline/target-opencode/opencode
url = https://github.com/anomalyco/opencode
[submodule "opencode-gaia"]
path = harness-opt-bench/gaia/baseline/target-opencode/opencode
url = https://github.com/anomalyco/opencode
4 changes: 3 additions & 1 deletion harness-opt-bench/gaia/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ files and images, run shell commands, and submit an exact answer.
| Development / validation / test | 33 / 66 / 66 |
| Split strategy | GAIA level and attachment presence |
| Target model | `gpt-5.4-mini` |
| Pinned seed baselines | 0.6205 working target; 0.0 shell target |
| Pinned seed baselines | 0.6205 working target; 0.0 shell target; opencode target unpinned (measure before use) |
| Scoring | Canonical task verifier |

Development cases expose full results and task resources. Validation exposes
Expand All @@ -25,6 +25,7 @@ aggregate scores, and test remains hidden until final scoring.
| `baseline/build.yaml` | Working tool-using agent | Measure harness improvement |
| `baseline/build.shell.yaml` | Minimal non-solving skeleton | Measure harness construction from scratch |
| `baseline/build.shell.e2e.yaml` | Skeleton with eight cases | Exercise the complete pipeline quickly |
| `baseline/build.opencode.yaml` | opencode, vendored as source and compiled per candidate | Measure improvement of a full coding harness |

The full and shell builds use the same tasks, model, budgets, and evaluation
policy. See [the baseline guide](baseline/README.md) for their editable surfaces
Expand Down Expand Up @@ -52,6 +53,7 @@ regenerate the partitions, and review the manifest before using new results.
| `baseline/build*.yaml` | Benchmark variants and evaluation settings |
| `baseline/target/` | Working editable agent |
| `baseline/target-shell/` | Minimal editable skeleton |
| `baseline/target-opencode/` | opencode source (git submodule) plus a thin wrapper |
| `partitions/` | Reported development, validation, and test split |
| `partitions-e2e/` | Small smoke-test split |
| `scripts/partition_gaia.py` | Split verification and regeneration |
149 changes: 149 additions & 0 deletions harness-opt-bench/gaia/baseline/build.opencode.e2e.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,149 @@
# Field meanings and the shared rules are in ../../CONFIGURATION.md; comments here
# cover only what is specific to this benchmark.
name: vero/optimize-gaia-opencode-e2e
description: >-
Improve opencode -- a full open-source coding agent, vendored as source under
opencode/ and compiled for each candidate -- as a GAIA agent, while preserving
the Harbor agent interface of the wrapper in src/. The target must answer
canonical GAIA tasks by writing the exact answer to /app/answer.txt; the wrapper
tells opencode so. Anything in the opencode source is editable: prompts, tools,
the agent loop, provider handling, defaults.
agent_repo: target-opencode
task_source: gaia/gaia@sha256:bbc356f476e0b70ba77da11a9be7d6345918d1e4a2daade0d6dfb82ee6f7b761
task_manifest: ../partitions/manifest.json
agent_import_path: gaia_agent.agent:GaiaAgent
harbor_requirement: ${harbor_requirement:-harbor[modal]==0.20.0}
vero_requirement: scaleapi-vero==0.6.0

# E2E SMOKE VARIANT of build.opencode.yaml. Same seed, same target, same gateway
# scopes; eight cases in total, tiny budgets, and a 25-minute agent clock so the
# run reaches finalization quickly. It exists to prove the opencode seed end to end
# (submodule snapshot, Bun build on the evaluation host, binary upload, run, session
# archive), not to measure anything. Never report its score.
partition_files:
development: ../partitions-e2e/development.json
validation: ../partitions-e2e/validation.json
test: ../partitions-e2e/test.json

agent_access:
- partition: development
disclosure: full
expose_case_resources: true
total_runs: 6
total_cases: 12
- partition: validation
disclosure: aggregate
expose_case_resources: false
min_aggregate_cases: 1
total_runs: 6
total_cases: 18

selection_partition: validation
targets:
- partition: test
reward_key: reward
# TODO(seed-complexity): measure with rescore_candidate.py --seed --build-file
# build.opencode.yaml (K=3) and pin here before quoting deltas.
baseline_reward: 0.0
failure_value: 0.0
max_attempts: 1
n_attempts: 3
aggregate_attempts: mean

evaluation_set_name: gaia
objective:
selector:
metric: score
direction: maximize
reward_mode: submit # agent picks; falls back to auto_best, then current version
baseline_floor: false # gates on validation while reward is on test; opt-in only
score_baseline: false
rescore_top_k: 1
rescore_attempts: 1

# GAIA includes image inputs, so the target must be multimodal. Keep this name
# aligned with the exact value sent by the agent.
model: gpt-5.4-mini
environment_name: ${inner_env:?set inner_env to a Harbor environment}
extra_harbor_args: ["--ek", "app_name=harness-opt-bench", "--ek", "sandbox_idle_timeout_secs=3600"]
harbor_python_version: "3.12"
n_attempts: 1
max_retries: 4
retry_max_wait_seconds: 120
infrastructure_max_attempts: 3
infrastructure_retry_delay_seconds: 5
aggregate_attempts: best
feedback_transcripts: true
feedback_max_bytes: 16000
expose_attempt_detail: false
# This smoke variant intentionally uses shorter overall limits.
timeout_seconds: 1800
case_timeout_seconds: 600
task_agent_timeout_seconds: 600
max_concurrency: 24
error_rate_threshold: 0.1
verifier_timeout_seconds: 3600
optimizer_sandbox_timeout_seconds: 14400
optimizer_sandbox_idle_timeout_seconds: 1800
# 25 minutes: long enough for the optimizer to build opencode once and run an
# evaluation or two, short enough that Harbor hands over to the verifier within
# the hour, which is the path this variant exists to exercise.
optimizer_agent_timeout_seconds: 1500
optimizer_allow_internet: true
optimizer_harness_versions:
claude-code: 2.1.220
codex: 0.146.0
opencode: 1.18.11
kimi-cli: 1.49.0
goose: 1.45.0
mini-swe-agent: 2.4.6

secrets:
- MODAL_TOKEN_ID
- MODAL_TOKEN_SECRET
- MODAL_ENVIRONMENT
- WANDB_API_KEY
- WANDB_BASE_URL

wandb:
project: harness-engineering-bench
group: gaia
name: ${wandb_run:-gaia-opencode-e2e}
tags: [gaia, opencode-seed, e2e-smoke]
log_traces: true

# opencode runs inside the task container, which cannot reach the compose-internal
# evaluation gateway. This hands the container the public upstream (OPENAI_*), as
# task-owned services already get; the metered gateway stays on VERO_AGENT_INFERENCE_*
# for in-process agents. Consequence: the target model is fixed by the wrapper, not by
# the gateway allow-list, and its use is audited from the upstream per-key request log.
task_services_use_upstream: true
# Required by the flag above (as in browsecomp-plus): the isolated harness user
# would otherwise see the raw upstream credential in its environment.
harness_user: null
agent_env:
BASH_MAX_TIMEOUT_MS: "3600000"
BASH_DEFAULT_TIMEOUT_MS: "3600000"
ENABLE_BACKGROUND_TASKS: "0"
FORCE_AUTO_BACKGROUND_TASKS: "0"
UV_TOOL_BIN_DIR: "/home/agent/.local/bin"

inference_gateway:
upstream_api_key_env: OPENAI_API_KEY
upstream_base_url_env: OPENAI_BASE_URL
request_log_attribution: true
producer:
allowed_models: ["${optimizer_model:-gpt-5.4}"]
max_concurrency: 8
evaluation:
allowed_models: [gpt-5.4-mini]
max_requests: 200000
max_tokens: 200000000 # 30 agent case-runs at most
max_concurrency: 64
finalization:
allowed_models: [gpt-5.4-mini]
max_requests: 200000
max_tokens: 2000000000 # 66 test cases x3 attempts + rescore headroom
max_concurrency: 64
instruct_multifidelity: true
instruct_exhaust_budget: true
144 changes: 144 additions & 0 deletions harness-opt-bench/gaia/baseline/build.opencode.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,144 @@
# Field meanings and the shared rules are in ../../CONFIGURATION.md; comments here
# cover only what is specific to this benchmark.
name: vero/optimize-gaia-opencode
description: >-
Improve opencode -- a full open-source coding agent, vendored as source under
opencode/ and compiled for each candidate -- as a GAIA agent, while preserving
the Harbor agent interface of the wrapper in src/. The target must answer
canonical GAIA tasks by writing the exact answer to /app/answer.txt; the wrapper
tells opencode so. Anything in the opencode source is editable: prompts, tools,
the agent loop, provider handling, defaults.
agent_repo: target-opencode
task_source: gaia/gaia@sha256:bbc356f476e0b70ba77da11a9be7d6345918d1e4a2daade0d6dfb82ee6f7b761
task_manifest: ../partitions/manifest.json
agent_import_path: gaia_agent.agent:GaiaAgent
harbor_requirement: ${harbor_requirement:-harbor[modal]==0.20.0}
vero_requirement: scaleapi-vero==0.6.0

partition_files:
development: ../partitions/development.json
validation: ../partitions/validation.json
test: ../partitions/test.json

agent_access:
- partition: development
disclosure: full
expose_case_resources: true
total_runs: 100
total_cases: 132
- partition: validation
disclosure: aggregate
expose_case_resources: false
min_aggregate_cases: 5
total_runs: 100
total_cases: 264

selection_partition: validation
targets:
- partition: test
reward_key: reward
# TODO(seed-complexity): measure with rescore_candidate.py --seed --build-file
# build.opencode.yaml (K=3) and pin here before quoting deltas.
baseline_reward: 0.0
failure_value: 0.0
max_attempts: 1
n_attempts: 3
aggregate_attempts: mean

evaluation_set_name: gaia
objective:
selector:
metric: score
direction: maximize
reward_mode: submit # agent picks; falls back to auto_best, then current version
baseline_floor: false # gates on validation while reward is on test; opt-in only
score_baseline: false
rescore_top_k: 3
rescore_attempts: 1

# GAIA includes image inputs, so the target must be multimodal. Keep this name
# aligned with the exact value sent by the agent.
model: gpt-5.4-mini
environment_name: ${inner_env:?set inner_env to a Harbor environment}
extra_harbor_args: ["--ek", "app_name=harness-opt-bench", "--ek", "sandbox_idle_timeout_secs=14400"]
harbor_python_version: "3.12"
n_attempts: 1
max_retries: 4
retry_max_wait_seconds: 120
infrastructure_max_attempts: 3
infrastructure_retry_delay_seconds: 5
aggregate_attempts: best
feedback_transcripts: true
feedback_max_bytes: 16000
expose_attempt_detail: false
# Above the 5,400-second worst case at the configured concurrency.
timeout_seconds: 7200
case_timeout_seconds: 600
task_agent_timeout_seconds: 600
max_concurrency: 24
error_rate_threshold: 0.1
verifier_timeout_seconds: 14400
optimizer_sandbox_timeout_seconds: 86400
optimizer_sandbox_idle_timeout_seconds: 14400
# The vendored opencode image takes 10-23 min to start on Modal; the 1800 s default
# timed out ~10 launches (EnvironmentStartTimeoutError).
build_timeout_seconds: 5400
optimizer_agent_timeout_seconds: 72000
optimizer_allow_internet: true
optimizer_harness_versions:
claude-code: 2.1.220
codex: 0.146.0
opencode: 1.18.11
kimi-cli: 1.49.0
goose: 1.45.0
mini-swe-agent: 2.4.6

secrets:
- MODAL_TOKEN_ID
- MODAL_TOKEN_SECRET
- MODAL_ENVIRONMENT
- WANDB_API_KEY
- WANDB_BASE_URL

wandb:
project: harness-engineering-bench
group: gaia
name: ${wandb_run:-gaia-opencode}
tags: [gaia, opencode-seed]
log_traces: true

# opencode runs inside the task container, which cannot reach the compose-internal
# evaluation gateway. This hands the container the public upstream (OPENAI_*), as
# task-owned services already get; the metered gateway stays on VERO_AGENT_INFERENCE_*
# for in-process agents. Consequence: the target model is fixed by the wrapper, not by
# the gateway allow-list, and its use is audited from the upstream per-key request log.
task_services_use_upstream: true
# Required by the flag above (as in browsecomp-plus): the isolated harness user
# would otherwise see the raw upstream credential in its environment.
harness_user: null
agent_env:
BASH_MAX_TIMEOUT_MS: "3600000"
BASH_DEFAULT_TIMEOUT_MS: "3600000"
ENABLE_BACKGROUND_TASKS: "0"
FORCE_AUTO_BACKGROUND_TASKS: "0"
UV_TOOL_BIN_DIR: "/home/agent/.local/bin"

inference_gateway:
upstream_api_key_env: OPENAI_API_KEY
upstream_base_url_env: OPENAI_BASE_URL
request_log_attribution: true
producer:
allowed_models: ["${optimizer_model:-gpt-5.4}"]
max_concurrency: 8
evaluation:
allowed_models: [gpt-5.4-mini]
max_requests: 200000
max_tokens: 2000000000 # 396 agent case-runs (132 dev + 264 validation)
max_concurrency: 64
finalization:
allowed_models: [gpt-5.4-mini]
max_requests: 200000
max_tokens: 2000000000 # 66 test cases x3 attempts + rescore headroom
max_concurrency: 64
instruct_multifidelity: true
instruct_exhaust_budget: true
Loading
Loading