CodexDSH CreatorClaude CodePiDSH PTCDSH StandardOh My PiKimi CodeDSH MinimalExo HarnessOpenCodeHermes
Pass rate
50.0%
55.6%
61.1%
66.7%
$1
$2
$5
$10
$20
Median cost per task

Pass rate

Median cost per successful task

Median cost per task

Median cache hit rate per successful task

Median time per successful task

Beyond the numbers

  1. 01
    OpenCode: failures excluded.

    It only covers 15 passes. Count failed attempts and the number becomes $3.24 per task.

  2. 02
    Cache hit rate is not cost.

    A cached 300-turn failure can still burn more than a short cache miss.

  3. 03
    Quality and cost can diverge.

    Claude Code passes 19 tasks, but reaches $18.34 in cost per task.

Run your harness on Runta.

If you want to test your own harness on Runta, we’ll give you $100 in credits to get started.

Tested harnesses

Codex
v0.148.0
DeepSeek Harness
v0.1.0-rc.8
Claude Code
v2.1.237
Pi
v0.84.2
Oh My Pi
v17.4.0
Kimi Code
v0.37.2
Exo Harness
v0.1.0
OpenCode
v1.18.19
Hermes
v0.20.4
  • FrontierHarness v1.0 focuses on software engineering contexts and terminal-based tasks. It may not generalize to other areas of knowledge work.
  • Evaluated on Runta agent runtimes. All harnesses and the task environment are prepared once as a golden checkpoint. Every run is a fresh restore with identical vCPU, memory, disk size, disk contents, and memory state.