Right Answers, Invalid Reasoning, Full Marks
7AUG
A new paper measured how often models reach correct benchmark answers through invalid reasoning. On common problems, 2%. On Humanity's Last Exam, the expert benchmark, 37% of correct answers were shortcuts.
Xuan Ren and co-authors call it solution hacking. The model guesses, pattern-matches, or exploits answer formats, then scores full marks because only the final answer gets graded.
The harder the benchmark, the worse the inflation. Vendor accuracy claims on frontier science tests may overstate real reasoning by a third.
If you buy models on benchmark deltas, this is your problem too. Ask vendors how they grade the method, not just the answer.