Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions docs/benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,3 +7,9 @@ Classified by what is being measured.
a long session, and behaviour through interrupts and crashes. Seven agent
CLIs, one harness, every table regenerable from the scripts beside the
results.
- [deepswe/](deepswe/): task-solving outcomes for senior-dev, the coding
harness vendored into this product, on the 113 scored tasks of the DeepSWE
v1.1 corpus. Three campaigns — a ten-harness comparison and two single-arm
model sweeps — as one per-task table each, with reward, F2P, P2P, wall time
and cost. The protocol for this repository's own `codeaf` runs on the same
corpus is separate, in [`bench/deepswe/`](../../bench/deepswe/).
Original file line number Diff line number Diff line change
@@ -0,0 +1,92 @@
# Harness comparison — DeepSWE 113, DeepSeek V4 Flash

Ten coding harnesses, same model, same 113 tasks, one attempt each.
**senior-dev finished first on reward and on F2P.**

| Harness | Solved | Reward | Mean F2P | Mean P2P | Spend | Per task | Mean time |
|---|---|---|---|---|---|---|---|
| [**senior-dev**](tasks/senior-dev.csv) | **62/113** | 54.87% | 88.64% | 99.42% | $24.73 | $0.22 | 54 min |
| [mini-swe-agent](tasks/mini-swe-agent.csv) | 56/113 | 49.56% | 88.17% | 99.66% | $43.09 | $0.38 | 44 min |
| [codex](tasks/codex.csv) | 51/113 | 45.13% | 86.81% | 98.76% | $42.23 | $0.37 | 45 min |
| [pi](tasks/pi.csv) | 42/113 | 37.17% | 71.63% | 89.09% | $39.61 | $0.35 | 51 min |
| [omp](tasks/omp.csv) | 31/113 | 27.43% | 71.91% | 85.26% | $55.98 | $0.50 | 48 min |
| [opencode](tasks/opencode.csv) | 30/113 | 26.55% | 71.13% | 86.79% | $56.96 | $0.50 | 47 min |
| [kilo](tasks/kilo.csv) | 30/113 | 26.55% | 69.03% | 90.51% | $54.65 | $0.48 | 53 min |
| [deepseek-harness](tasks/deepseek-harness.csv) | 16/113 | 14.16% | 53.39% | 91.45% | $169.47 | $1.50 | 93 min |
| [claude-code](tasks/claude-code.csv) | 16/113 | 14.16% | 39.72% | 91.60% | $21.72 | $0.19 | 31 min |
| [muse-code](tasks/muse-code.csv) | 3/113 | 2.65% | 5.14% | 99.01% | $13.54 | $0.12 | 15 min |

Reward is the binary verifier result. F2P and P2P are macro-averages over all 113
scheduled tasks — the five invalid outcomes contribute zero rather than being dropped.
Per task is spend ÷ 113, so invalid outcomes stay in the denominator too.

## Setup

| | |
|---|---|
| Benchmark | full DeepSWE set, 113 tasks, one seed per harness |
| Model | `deepseek/deepseek-v4-flash-0731` via OpenRouter |
| Verifiers | official DeepSWE at `0b9fabb` |
| Budget | 3 h per task |
| Isolation | four GCP shards per harness, dedicated OpenRouter key per harness |
| Ran | nine arms 2026-09-11; senior-dev arm 2026-09-12 03:10Z–08:20Z |
| Total spend | ≈ $521.98 |

**The senior-dev arm departs from the frozen sampling contract.** The other nine sent
temperature 1.0 and top-p 0.95; senior-dev sent neither, so OpenRouter and provider
defaults applied. Everything else matched. senior-dev was deliberately excluded from
the nine-harness campaign and run separately a day later.

## Invalid outcomes

Five tasks across four arms produced no usable verifier outcome. They have an empty
`reward` in `tasks.csv`, stay in the 113-task denominator and contribute zero.

| Harness | Task | Why |
|---|---|---|
| [codex](tasks/codex.csv) | `valibot-recursive-schema-composition` | patch not accepted |
| [pi](tasks/pi.csv) | `langchain-request-coalescing` | agent and verifier both timed out |
| [omp](tasks/omp.csv) | `actionlint-action-pinning-lint` | agent killed, exit 137, no verifier result |
| [omp](tasks/omp.csv) | `kombu-virtual-queue-dead-lettering` | verifier timeout, exit 124 |
| [opencode](tasks/opencode.csv) | `pwntools-tube-multiplexing` | verifier timeout, exit 124 |

A non-zero harness exit is not an invalid result — a task can time out and still
produce a patch the verifier grades. Only tasks with no verifier outcome count here.

## The files

`arms.csv` — one row per harness: solved, reward rate, valid grades, invalid count,
mean F2P/P2P, OpenRouter spend, cost per task, mean agent seconds, model.

`tasks/<harness>.csv` — one file per harness, 113 rows each, ordered by difficulty
rank. All ten share a column layout, so `cat` gives the whole campaign back:

```
awk 'FNR>1 || NR==1' tasks/*.csv > all-arms.csv
```

Columns:

| Column | Meaning |
|---|---|
| `harness` `task` `band` `difficulty_rank` `language` `dev_set` | which arm, which task, and how the task is classified |
| `reward` | 1 pass, 0 fail, empty = no verifier outcome |
| `f2p` `p2p` | fraction of issue tests now passing / pre-existing tests still passing |
| `f2p_passed` `f2p_total` `p2p_passed` `p2p_total` | the underlying test counts |
| `seconds` `started` `finished` | agent wall time and timestamps |
| `cost_usd` `model_calls` | per-attempt model spend and call count |
| `exit` `timed_out` | harness exit status |
| `model` `harness_version` `attempt_id` | what produced it |

**`cost_usd` and `model_calls` are empty for six arms** — omp, opencode, kilo,
deepseek-harness, claude-code and muse-code did not record per-attempt usage. Their
spend is known only at arm level, in `arms.csv`.

## Source

Published as [Harness Comparison — DeepSWE](https://app.notion.com/p/Harness-Comparison-DeepSWE-3d603ed71f1080a5bcf3fddf8d7b2925).

The per-attempt artifacts behind these rows — `events.ndjson`, `run.log`, traces,
verifier stdout, patches — are about 19 GB and stay on the machines that produced
them. These campaigns ran while the binary was still called `codeaf`; the values
here are rewritten to `senior-dev`.
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
harness,solved,tasks,reward_rate,valid_grades,invalid,mean_f2p,mean_p2p,openrouter_spend_usd,cost_per_task_usd,mean_agent_seconds,model
senior-dev,62,113,0.5487,113,0,0.8864,0.9942,24.73,0.2188,3258,openrouter/deepseek/deepseek-v4-flash-0731
mini-swe-agent,56,113,0.4956,113,0,0.8817,0.9966,43.09,0.3813,2651,openrouter/deepseek/deepseek-v4-flash-0731
codex,51,113,0.4513,112,1,0.8681,0.9876,42.23,0.3737,2746,deepseek/deepseek-v4-flash-0731
pi,42,113,0.3717,112,1,0.7163,0.8909,39.61,0.3505,3101,openrouter/deepseek/deepseek-v4-flash-0731
omp,31,113,0.2743,111,2,0.7191,0.8526,55.98,0.4954,2920,openrouter/deepseek/deepseek-v4-flash-0731
opencode,30,113,0.2655,112,1,0.7113,0.8679,56.96,0.5041,2871,openrouter/deepseek/deepseek-v4-flash-0731
kilo,30,113,0.2655,113,0,0.6903,0.9051,54.65,0.4836,3215,openrouter/deepseek/deepseek-v4-flash-0731
deepseek-harness,16,113,0.1416,113,0,0.5339,0.9145,169.47,1.4997,5613,openrouter/deepseek/deepseek-v4-flash-0731
claude-code,16,113,0.1416,113,0,0.3972,0.916,21.72,0.1922,1909,deepseek/deepseek-v4-flash-0731
muse-code,3,113,0.0265,113,0,0.0514,0.9901,13.54,0.1198,952,deepseek/deepseek-v4-flash-0731
Loading
Loading