Skip to content
← Blog

Mouse on FrontierHarness

Mouse scored 24/30 with the highest pass rate.

FrontierHarness is a benchmark that runs the same model, Kimi K3, through different coding agents on the same 30 tasks. Mouse passed 24 of 30, the highest pass rate and the fastest median time of any harness on the board.

Pass rate
80%
24 of 30 tasks
Rank
#1 of 13
pass rate and median time
Cost per pass
$3.13
$0.31 median
Median time
4:10
per pass
Tasks passed, 30 FrontierHarness tasks, Kimi K3
Mouse 80%
Codex 67%
DSH Creator 63%
Claude Code 63%
Pi 60%
DSH Standard 60%
DSH PTC 60%
Kimi Code 57%
DSH Minimal 57%
Oh My Pi 57%
Exo 53%
Hermes 50%
OpenCode 50%

The benchmark

FrontierHarness runs 30 tasks. 21 come from Terminal-Bench 2.1 and are command line jobs, such as recovering a corrupted SQLite database or building a Cython extension. The other 9 are DeepSWE tasks, which are bug fixes and features in open source repos, graded by hidden tests.

Every harness gets the same model (Kimi K3), the same tasks, and the same runtime. The harness is the only thing that changes, so the results show what the prompt, config, tools, and loop are actually worth.

How Mouse works

The Mouse agent harness is built on top of the OpenCode engine. We designed our system prompt, configuration, rules, and set of skills to steer a custom loop that outperforms on coding tasks.

The loop is the biggest differentiator within Mouse. Most agents finish when the model decides a task is done. With Mouse, we use a deterministic approach that mandates the workspace inspects, verifies, and ensures the evidence of the completed task holds up. This loop iterates until the desired outcome.

This approach shows up when long-horizon jobs are required, which is why OpenCode passed 0 of the DeepSWE tasks and Mouse passed 6.

This is the same loop that powers Night Shift, how we run agents on long time horizon jobs and overnight.

Leaderboard

Harness Pass Terminal DeepSWE $ / pass Median $ Cache Time
Mouse 24/30 18/21 6/9 $3.13 $0.31 84% 4:10
Codex 20/30 15/21 5/9 $3.47 $0.12 88% 6:43
DSH Creator 19/30 15/21 4/9 $3.28 $0.12 84% 6:43
Claude Code 19/30 16/21 3/9 $18.34 $0.29 68% 9:37
Pi 18/30 16/21 2/9 $2.43 $0.07 79% 7:33
DSH Standard 18/30 16/21 2/9 $3.46 $0.12 86% 6:16
DSH PTC 18/30 16/21 2/9 $4.58 $0.14 87% 7:43
Kimi Code 17/30 14/21 3/9 $3.65 $0.18 88% 7:55
DSH Minimal 17/30 15/21 2/9 $4.72 $0.12 85% 5:41
Oh My Pi 17/30 16/21 1/9 $4.75 $0.14 82% 6:46
Exo 16/30 16/21 0/9 $1.05 $0.07 70% 6:17
Hermes 15/30 13/21 2/9 $2.90 $0.17 86% 6:57
OpenCode 15/30 15/21 0/9 $3.24 $0.06 78% 6:27

Cost per pass is total spend on all 30 tasks divided by passes. Median cost, cache rate, and time cover passing tasks only. Mouse was the only harness to solve kv-store-grpc and scc-bounded-memory-spilling.

Cost

Cost per pass, total spend divided by passes, lower is better
Exo $1.05
Pi $2.43
Hermes $2.90
Mouse $3.13
OpenCode $3.24
DSH Creator $3.28
DSH Standard $3.46
Codex $3.47
Kimi Code $3.65
DSH PTC $4.58
DSH Minimal $4.72
Oh My Pi $4.75
Claude Code $18.34

Mouse is fourth on cost per pass and last on median cost. Each verification round is another model call, which increases the harness's costs. Our next move is to improve cache and compaction to bring these costs down.

Terminal-Bench

Task Mouse Steps Cost Time Field
regex-log Pass 21 $0.51 7:50 11/12
openssl-selfsigned-cert Pass 13 $0.20 2:44 12/12
polyglot-c-py Pass 14 $0.15 2:35 12/12
sqlite-db-truncate Pass 21 $0.32 4:17 12/12
git-leak-recovery Pass 15 $0.14 2:06 12/12
log-summary-date-ranges Pass 21 $0.25 3:32 12/12
constraints-scheduling Pass 15 $0.31 4:49 12/12
gcode-to-text Fail 27 $0.71 9:54 5/12
dna-insert Pass 18 $0.71 11:05 2/12
largest-eigenval Fail 20 $1.03 0/12
merge-diff-arc-agi-task Pass 19 $0.30 4:03 12/12
vulnerable-secret Pass 17 $0.18 2:15 12/12
extract-elf Fail 13 $0.49 7:49 2/12
build-cython-ext Pass 47 $0.93 13:08 12/12
kv-store-grpc Pass 17 $0.15 2:32 0/12
chess-best-move Pass 18 $0.59 6:13 5/12
db-wal-recovery Pass 13 $0.15 2:23 12/12
code-from-image Pass 8 $0.09 1:09 4/12
modernize-scientific-stack Pass 10 $0.21 3:27 12/12
multi-source-data-merger Pass 10 $0.19 3:07 12/12
sanitize-git-repo Pass 26 $0.47 3:58 10/12

Field is how many of the 12 published harnesses passed the task.

DeepSWE

Task Mouse Steps Cost Time Field
anko-typed-variable-bindings Fail 52 $2.64 21:01 4/12
arktype-json-schema-refs-dependencies Pass 190 $10.32 89:36 2/12
fastapi-deprecation-response-headers Pass 150 $7.08 85:32 4/12
httpx-multipart-response-parsing Pass 141 $6.06 70:48 5/12
expr-try-catch-errors Pass 134 $8.59 44:43 2/12
python-statemachine-state-data-scoping Pass 201 $13.49 73:21 7/12
katex-multicolumn-array-spans Fail 117 $5.65 44:52 1/12
scc-bounded-memory-spilling Pass 84 $4.65 52:44 0/12
meriyah-explicit-resource-declarations Fail 139 $8.47 65:49 1/12

Two of the three misses passed nearly all of their hidden tests, 92 of 94 and 46 of 49.

Setup

  • Harbor 0.22, the benchmark's runner, on a Google Cloud VM with 4 tasks in parallel.
  • Kimi K3 through OpenRouter, pinned to the Fireworks servers the benchmark uses.
  • One trial per task. Cost at list price. Time is wall clock per task.

Mouse is built on top of OpenCode. Run date: September 3, 2026.