Thursday Aug 27

A 27B Model Beat Claude At Replication

27AUG
REPLICA TEST27B MODELFRONTIER

A twelve-person London startup left stealth with $50M. Its small open-weight system outperformed two frontier labs on a paper-reproduction test it wrote itself.

Inherent is four DeepMind and Reka alumni in King's Cross. Twelve people, growing to 25 by year-end. Seed led by Index and Radical Ventures.

Their benchmark, Replica, has 310 tasks drawn from 100 machine learning and AI-for-science papers. The model must reproduce the results without seeing the answers.

Faraday is built on a 27B Qwen base and uses GPT-5.5 Codex for the coding work. It scored above both frontier models on Replica.

full brief & sources

⚡ Why this matters

  • Reproducing a paper is a narrow, checkable task. Narrow tasks are where small models win.
  • If a 27B base beats frontier models here, task-specific tuning beats scale for this job.
  • Replication is the bottleneck in AI-for-science. Automating it compounds.

🔍 What happened

  • Out of stealth Aug 22 with a $50M seed.
  • Founders: Tantum Collins, Edward Hughes, Louis Kirsch, Kaloyan Aleksiev.
  • Replica: 310 tasks, 100 papers, no prior answers available to the model.
  • Faraday scored above Claude Opus 4.8 and GPT-5.5 on that benchmark.

💬 Smart takes

  • The benchmark is theirs. That is the whole caveat and it is a big one.
  • Using GPT-5.5 Codex inside the system muddies the claim of beating GPT-5.5.
  • Even so, the shape is right. Verification is more tractable than generation.

🧭 Where this goes

  1. LikelyReplica gets published and someone independent runs it.
  2. Possiblea frontier lab ships a replication mode and the moat closes.
  3. Wild Carda journal starts requiring an automated replication pass before review.

🥄 The Spoon Take

Every vendor-built benchmark should be read as a product claim, not a result. The interesting part is not the score. It is the bet that a small tuned model plus a frontier coder beats a frontier model alone. That architecture is cheap to copy, which means the moat has to be the benchmark itself.

🤔 Pushback

No independent verification. The benchmark is Inherent's own, and the system calls GPT-5.5 while claiming to beat it.