Same Model, Same Task, 17x The Bill
2SEP
Runta ran nine agent harnesses on identical tasks with the same model. Pass rates landed within seventeen points of each other. Cost per completed task ranged from $1.05 to $18.34.
12 configurations, 9 harnesses, 360 runs, 2 billion tokens. One model, one runtime, one task set. The only variable was the scaffolding around it.
Success clustered between 50.0% and 66.7%. Median spend per solved job spanned 17x. The cheap end and the accurate end are not the same product.
Codex is the safe default: best success rate at $3.47, near the field median. Exo Harness is cheapest per finished job at $1.05.