Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
50 changes: 50 additions & 0 deletions harness-opt-bench/CONFIGURATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -97,6 +97,10 @@ not affect candidate selection.
| [BrowseComp-Plus](browsecomp-plus/baseline/build.yaml) | `deepseek-v4-flash` | 33 / 66 / 66 | 0.462 ± 0.028 |
| [Terminal-Bench](terminal-bench/baseline/build.yaml) | `grok-build-0.1` | 17 / 36 / 36 | 0.241 ± 0.013 |

Target models are named without a provider. Each build maps the name to a
deployment under the evaluation scope's `model_aliases`; pass
`--param target_model_route=<provider/model>` to use a different one.

Each baseline is the mean of three independent test rounds; the reported
uncertainty is the standard deviation of those round means. GAIA uses a
multimodal target because some held-out cases contain images. The reported GAIA
Expand Down Expand Up @@ -194,6 +198,51 @@ Some tasks depend on external services of their own and therefore use a reduced
isolation profile. That limitation must be declared in the benchmark configuration
and considered before running with an adversarial optimizer.

### Deployment settings

A few fields describe the services a run depends on rather than the benchmark
itself. They are the same in every build file, and the build files do not
repeat their meaning.

- `harbor_requirement` names the Harbor package the evaluation service installs.
It carries the extra for the evaluation environment and can be overridden with
`--param harbor_requirement=`.
- `secrets` lists the environment variable names the run needs for its execution
and telemetry services. Every name must be present in the env file; the
compiler refuses to compile otherwise. Edit the list and the env file together.
- `wandb` configures telemetry and is optional. Remove the block to run without it.
- `extra_harbor_args` passes options to the evaluation sandboxes. The defaults
group them under one application and reclaim idle ones; remove them for an
environment that does not accept those options.
- `inference_gateway.request_log_attribution` stamps each gateway request with the
trial it served, which is what makes per-trial usage attribution reliable.
- `agent_env` raises the optimizer's tool-call time limit so one full evaluation
can complete inside a single foreground call, and disables background tasks.

The remaining shared values, such as attempts, budgets, time limits and
isolation, follow the rules above; a build file comments only on what is
specific to its benchmark.

## Compiled tasks

`vero harbor build` turns a `build.yaml` into a self-contained task directory,
and two compiles of the same sources produce identical output. Nothing specific
to one run is written into the compiled task: the per-run access tokens and the
optimizer model are supplied by the launcher through the environment and read
once when the services start, so a running evaluation cannot be altered from
outside.

Each reported benchmark commits `baseline/compiled.manifest.json`, a checksum of
every compiled file. `vero harbor build --check <manifest>` recompiles and fails
if a checkout no longer reproduces the recorded task. The manifests are written
with `inner_env=modal`; pass the same parameter when checking. Regenerate the manifest
whenever the seed, the partitions, or the build configuration changes.

By default the compiled task carries a copy of the VeRO source tree so its
images can be built without a package index. Setting `vero_requirement` to an
exact published version (for example `scaleapi-vero==0.6.0`) installs that
release instead; the version must match the VeRO doing the compiling.

## Adding or changing a benchmark

1. Update the benchmark's `baseline/build.yaml`.
Expand All @@ -215,6 +264,7 @@ uv run pytest tests/test_v05_benchmark_configs.py

uv run vero harbor build \
--config ../harness-opt-bench/<benchmark>/baseline/build.yaml \
--param inner_env=<evaluation-environment> \
--output <output-directory>
```

Expand Down
18 changes: 10 additions & 8 deletions harness-opt-bench/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,10 +32,10 @@ Each benchmark pins its target, dataset, split, budgets, and scoring protocol in

| Benchmark | Editable target | Dataset | Split |
| --- | --- | --- | --- |
| [GAIA](gaia/baseline/) | Tool-using multimodal agent | GAIA | 20% / 40% / 40% |
| [OfficeQA](officeqa/baseline/) | Grounded document-QA agent | Treasury Bulletin corpus | 20% / 40% / 40% |
| [BrowseComp-Plus](browsecomp-plus/baseline/) | Fixed-corpus research agent | BrowseComp-Plus | 20% / 40% / 40% |
| [Terminal-Bench](terminal-bench/baseline/) | Shell-based terminal agent | Terminal-Bench 2.1 | 20% / 40% / 40% |
| [GAIA](gaia/) | Tool-using multimodal agent | GAIA | 20% / 40% / 40% |
| [OfficeQA](officeqa/) | Grounded document-QA agent | Treasury Bulletin corpus | 20% / 40% / 40% |
| [BrowseComp-Plus](browsecomp-plus/) | Fixed-corpus research agent | BrowseComp-Plus | 20% / 40% / 40% |
| [Terminal-Bench](terminal-bench/) | Shell-based terminal agent | Terminal-Bench 2.1 | 20% / 40% / 40% |

Additional benchmarks that were implemented but are not part of the reported
suite live under [`archive/`](archive/).
Expand Down Expand Up @@ -66,14 +66,15 @@ uv run vero harbor run \
--env-file <your>.env \
--agent <optimizer-harness> \
--model <optimizer-model> \
--param inner_env=<evaluation-environment> \
--param optimizer_model=<optimizer-model-as-the-harness-sends-it> \
-o ../runs/<run-name>/jobs
```

The env file holds your model-endpoint key, execution-environment tokens and
telemetry credentials; keep it outside the benchmark definition. The two model
arguments differ because some harnesses rewrite the model name before sending
it, and the gateway only accepts the name it was told to expect. The runbook,
The env file holds credentials required by your inference and execution
services; keep it outside the benchmark definition. Some harnesses rewrite the
model name before sending it, so `optimizer_model` must match that final value.
The runbook,
[`skills/run-benchmark/SKILL.md`](skills/run-benchmark/SKILL.md), covers that
rule, the preflight and the health checks. Before launching a full experiment,
use [`vero/examples/harness-conformance/`](../vero/examples/harness-conformance/)
Expand Down Expand Up @@ -101,6 +102,7 @@ inspect the generated task:
cd vero
VERO_SKIP_SECRET_CHECK=1 uv run vero harbor build \
--config ../harness-opt-bench/<benchmark>/baseline/build.yaml \
--param inner_env=<evaluation-environment> \
--output <output-directory>
```

Expand Down
28 changes: 15 additions & 13 deletions harness-opt-bench/archive/README.md
Original file line number Diff line number Diff line change
@@ -1,15 +1,17 @@
# Archived benchmarks

Three optimization tasks that were wired, run, and left out of the paper. They
still compile with `vero harbor build`, but they are not maintained to the
conventions in `../CONFIGURATION.md` and nothing in CI exercises them.

| benchmark | why it is not in the paper |
|---|---|
| `swe-atlas-qna/` | The binary held-out reward sits at the floor (seed 0.067, sd 0.025): 20 optimizer cells spanned 0.054 to 0.127 and no adjacent pair separated. The continuous `agg_score` the rubric also emits (seed 0.632) is more informative but is not surfaced as a selectable reward. |
| `tau3/` | Its user simulator and grader are both LLMs, so a single unchanged harness scored 0.800 and 0.547 on the same development set. Deltas under about 0.1 are unresolved, and only 4 of 16 grid cells recorded a usable baseline. |
| `swe-bench-pro/` | Never normalized to the shared conventions: 1x case budgets, single-shot held-out scoring, no pinned baseline, an agent clock at 0.6x the declared case timeout, and a seed that edits nothing. Its `build.yaml` is the full 731-instance split; `build.sample.yaml` is a nested 33/66/66 subsample. |

Runs that were made against these configs are under `runs/swe-atlas-qna/`,
`runs/tau3/` and `runs/conformance/`, and the write-ups that judged them are
`runs/AUDIT-swe-atlas-qna.md` and the tau3 rows of `runs/RESULTS-INDEX.md`.
These benchmarks are retained for reproducibility and future development but are
not part of the reported HarnessOpt-Bench suite.

| Benchmark | Why it was archived |
| --- | --- |
| [SWE-Atlas-QnA](swe-atlas-qna/) | Binary rewards did not separate optimizer results reliably |
| [tau3](tau3/) | Simulator and grader variability obscured small harness improvements |
| [SWE-bench-Pro](swe-bench-pro/) | The configuration was not normalized to the shared evaluation protocol |

The configurations remain useful as working examples, but results should not be
compared directly with the reported suite without first reviewing their pinned
baselines, budgets, attempts, and timeouts.

Each benchmark directory explains its task, split, editable target, and current
limitations.
44 changes: 27 additions & 17 deletions harness-opt-bench/archive/swe-atlas-qna/README.md
Original file line number Diff line number Diff line change
@@ -1,26 +1,36 @@
# SWE-Atlas-QnA

This benchmark optimizes an agent that investigates a checked-out software
repository and writes an evidence-backed answer to a deep codebase question.
Scoring is the canonical rubric judge shipped in each Harbor task.
SWE-Atlas-QnA evaluates an agent that investigates a checked-out repository and
writes an evidence-backed answer to a codebase question.

The pinned 124-task dataset is split 25/49/50 for development, validation, and
test. The split is deterministic and stratified by source repository so that
the ten represented codebases occur across the three partitions.
## At a glance

Regenerate or verify the committed split from an exported dataset:
| Item | Value |
| --- | --- |
| Cases | 124 |
| Development / validation / test | 25 / 49 / 50 |
| Editable harness | `baseline/target/` |
| Split strategy | Source repository |
| Scoring | Canonical rubric judge |
| Status | Archived; not part of the reported suite |

```bash
python harness-opt-bench/scripts/partition_dataset.py swe-atlas-qna \
--tasks-dir /path/to/exported/dataset \
--output-dir harness-opt-bench/candidates/swe-atlas-qna/partitions \
--fetch-registry
Development exposes complete task results and repositories. Validation is
aggregate-only, and test is held out for final scoring.

## Data and split

The split is deterministic and keeps all ten source repositories represented
across partitions. Verify it against an exported dataset:

~~~bash
python harness-opt-bench/scripts/partition_dataset.py swe-atlas-qna \
--tasks-dir /path/to/exported/dataset \
--output-dir harness-opt-bench/candidates/swe-atlas-qna/partitions \
--tasks-dir <exported-tasks> \
--output-dir harness-opt-bench/archive/swe-atlas-qna/partitions \
--check
```
~~~

Use `--fetch-registry` when intentionally refreshing the pinned package. Review
the task manifest and split before accepting new results.

`--fetch-registry` requires the pinned Harbor package and verifies every task
name and content digest against Harbor Hub.
See [the baseline guide](baseline/README.md) for the editable target and the two
target-model builds.
57 changes: 27 additions & 30 deletions harness-opt-bench/archive/swe-atlas-qna/baseline/README.md
Original file line number Diff line number Diff line change
@@ -1,39 +1,36 @@
# SWE-Atlas-QnA codebase agent
# SWE-Atlas-QnA editable target

This leaf benchmark optimizes a small Harbor-native agent that explores the
repository mounted at `/app` and writes its final answer to
`/logs/agent/answer.txt`. The editable program controls its prompt, search
strategy, shell tools, context management, and answer synthesis.
The seed explores the repository mounted at `/app` and writes its answer to
`/logs/agent/answer.txt`. The optimizer may change its prompt, search strategy,
shell tools, context management, answer synthesis, and dependencies.

Two pinned target builds share this seed, split, budgets, and access policy;
they differ only in target model, and each pins the seed floor measured on
**its own** model — a delta against the other build's floor is a model
comparison, not an optimization result:
Two builds share the same target, split, budgets, and access policy:

| build | target model | pinned `baseline_reward` |
| --- | --- | --- |
| `build.yaml` | `fireworks_ai/gpt-oss-120b` | 0.0667 (K=3, n=150) |
| `build.gpt54mini.yaml` | `gpt-5.4-mini` | 0.1216 (K=3, n=148) |
| Build | Target model | Pinned seed baseline |
| --- | --- | ---: |
| `build.yaml` | `gpt-oss-120b` | 0.0667 |
| `build.gpt54mini.yaml` | `gpt-5.4-mini` | 0.1216 |

Both floors were measured with `scripts/rescore_candidate.py --seed`, the same
path that produced every other pinned baseline in the suite. Trials the seed
itself killed score 0 (they are harness headroom, and finalization taxes a
candidate's dead attempts the same way); trials the platform killed are
excluded. For gpt-5.4-mini that prices in the seed's ~8% empty-completion
fail-fast — the largest single piece of fixable headroom in this seed.
Each baseline belongs to its configured model. Comparing a candidate against the
other build's baseline would mix harness improvement with a model change.
Candidate-caused failures score zero; evaluation-system failures are excluded.

The Harbor tasks retain their canonical rubric-based verifier. That verifier
needs `OPENAI_API_BASE`; the target agent uses `OPENAI_BASE_URL`. They may
point to the same OpenAI-compatible endpoint.
## Compile

Compile from the repository root:
From the repository root:

```bash
~~~bash
cd vero
VERO_SKIP_SECRET_CHECK=1 uv run vero harbor build \
--config ../harness-opt-bench/swe-atlas-qna/baseline/build.yaml \
--output ../harness-opt-bench/swe-atlas-qna/baseline/compiled
```

For a real run, provide `OPENAI_API_KEY`, `OPENAI_BASE_URL`,
`OPENAI_API_BASE`, and the Modal credentials declared in `build.yaml`.
--config ../harness-opt-bench/archive/swe-atlas-qna/baseline/build.yaml \
--param inner_env=<evaluation-environment> \
--output <output-directory>
~~~

Use `build.gpt54mini.yaml` to compile the alternate target-model build. The
`VERO_SKIP_SECRET_CHECK` setting is appropriate only for compile-time
validation.

Dataset and split details are in the
[SWE-Atlas-QnA overview](../README.md). For a real optimization run, follow the
shared [`run-benchmark` guide](../../../skills/run-benchmark/SKILL.md).
Loading
Loading