Friday Aug 7

Right Answers, Invalid Reasoning, Full Marks

7AUG
37% ON HLE FULL MARKS SHORTCUTS

A new paper measured how often models reach correct benchmark answers through invalid reasoning. On common problems, 2%. On Humanity's Last Exam, the expert benchmark, 37% of correct answers were shortcuts.

Xuan Ren and co-authors call it solution hacking. The model guesses, pattern-matches, or exploits answer formats, then scores full marks because only the final answer gets graded.

The harder the benchmark, the worse the inflation. Vendor accuracy claims on frontier science tests may overstate real reasoning by a third.

If you buy models on benchmark deltas, this is your problem too. Ask vendors how they grade the method, not just the answer.

full brief & sources

⚡ Why this matters

  • Benchmark scores drive model procurement, pricing, and press - if a third of hard-test wins are hollow, every comparison chart wobbles.
  • The finding lands hardest on frontier science claims, exactly where labs market superhuman progress.
  • It gives buyers a concrete question to ask vendors: how do you grade reasoning validity?

🔍 What happened

  • Aug 3 - researchers post 'Right Answer, Wrong Method' on arXiv, studying shortcut hacking on frontier science benchmarks.
  • They find 2.2% of correct answers on common problems came through invalid reasoning routes.
  • Olympiad-level problems land in between at 28.3% - inflation climbs steadily with difficulty.
  • On Humanity's Last Exam, the 2,500-question expert benchmark from the Center for AI Safety and Scale AI, that rises to 37.4%.
  • Failure modes include lucky guessing, answer-format exploitation, and pattern-matching to training data.
  • The authors argue answer-only grading systematically inflates frontier reasoning claims.

💬 Smart takes

  • The paper: models hit the right answer through an invalid route, then get full marks anyway.
  • Asanify's analysis: benchmark shortcut hacking is inflating vendor accuracy claims.
  • Skeptic: grading reasoning validity is itself a judgment call by another model or rubric - the meta-grader can be wrong too.

🧭 Where this goes

  1. Likelybenchmark maintainers add method-validity grading to headline leaderboards within six months.
  2. Likelylab marketing quietly shifts from single accuracy numbers to verified-reasoning metrics.
  3. Possiblean enterprise buyer publicly walks back a model choice after re-grading with method checks.
  4. Wild Carda major leaderboard restates historical scores downward and reshuffles the rankings.

🥄 The Spoon Take

Every model comparison deck you've seen this year quietly assumed right answer means right reasoning. On the hardest tests, that's wrong more than a third of the time. The lesson isn't that models are dumb - it's that we've been grading them like multiple-choice students and calling it science.

🤔 Pushback

One paper on a handful of benchmarks isn't a field-wide indictment - replication on other test suites could shrink the 37% substantially.