Thursday Sep 3

Same Model, Same Task, 17x The Bill

2SEP
SAME MODEL$1.05$18.34

Runta ran nine agent harnesses on identical tasks with the same model. Pass rates landed within seventeen points of each other. Cost per completed task ranged from $1.05 to $18.34.

12 configurations, 9 harnesses, 360 runs, 2 billion tokens. One model, one runtime, one task set. The only variable was the scaffolding around it.

Success clustered between 50.0% and 66.7%. Median spend per solved job spanned 17x. The cheap end and the accurate end are not the same product.

Codex is the safe default: best success rate at $3.47, near the field median. Exo Harness is cheapest per finished job at $1.05.

full brief & sources

⚡ Why this matters

  • Everyone benchmarks models. Almost nobody benchmarks the wrapper around the model, which is where the money leaks.
  • A 17x spread on identical work means your unit economics are a scaffolding choice, not a model choice.
  • If you are negotiating an AI budget, this is the lever nobody in the room is pricing.

🔍 What happened

  • FrontierHarness Eval v1.0, published by Runta on September 2 and posted to Hacker News.
  • 12 configurations of 9 harnesses. Same model, same runtime, same software-engineering and terminal task set.
  • 360 runs, roughly 2 billion tokens.
  • Success rates: 50.0% to 66.7%. Cost per completed task: $1.05 to $18.34.
  • Named entrants include Codex, Claude Code, OpenCode, Kimi Code, Exo Harness, Hermes and several DSH variants.

💬 Smart takes

  • Runta's own read: Codex is the safe default because it tops the success rate while sitting near median cost at $3.47.
  • The cheapest option, Exo Harness at $1.05, does not top the accuracy table. The trade is explicit.
  • A parallel arXiv line of work argues harness adaptation can cut agent cost by 90% on smaller models. Same conclusion from the other direction.

🧭 Where this goes

  1. Likelyharness benchmarks become a standard procurement artifact next to model evals.
  2. Possiblevendors start publishing cost-per-pass instead of pass rate. That is a friendlier number to optimize.
  3. Wild Carda model provider bundles a tuned harness and the distinction stops being a buyer's choice.

🥄 The Spoon Take

This is the eval everyone should have run a year ago. The model is the commodity. The harness is where cost and reliability actually get decided, and almost nobody measures it. Go price your own agent stack per completed task, not per token. The number will surprise you.

🤔 Pushback

One task family, 360 runs, one vendor's benchmark. Software engineering is not every agent workload, and harness authors will contest the configurations.