Sunday Jun 21

Startup Bets Weaker AI Stops Hallucinations

21JUN
99.99% TARGETBIG MODELCHECKER

Everyone chases smarter models. This startup went the other way. Probably raised $9M from a16z and Accel to catch AI errors with a smaller, dumber model. The bet: reliability is the missing piece.

Most labs fight hallucinations by making models bigger and smarter. Probably flips it - a small, narrow model double-checks the big model's output and flags likely errors before they ship.

The goal is 99.99% accuracy, the kind deterministic software hits but AI rarely does. The $9M seed was co-led by Andreessen Horowitz and Accel. It's early - one product, small round, no public benchmark yet.

If it works, reliability becomes a feature you buy, not a model you pray to. Watch whether enterprises pay for a 'correctness layer' on top of their existing AI.

full brief & sources

Why this matters

  • Reframes the hallucination problem - maybe the fix isn't a smarter model.
  • Reliability is the top blocker for enterprise AI deployment.
  • A 'correctness layer' could become its own product category.

🔍 What happened

  • Jun 16: Probably announces a $9M seed round.
  • Co-led by Andreessen Horowitz and Accel; Tokyo Black and Vermilion Cliffs Ventures joined.
  • Approach: a smaller, narrow model verifies a larger model's output.
  • Target: 99.99% accuracy, near deterministic-software levels.
  • Aims to catch factual errors before they reach the user.

💬 Smart takes

  • Probably (via TechCrunch): the fix for AI errors is a weaker model, not a smarter one.
  • Skeptic: a verifier model can hallucinate too - who checks the checker?

🧭 Where this goes

  1. Likelymore 'AI reliability' startups raise this year.
  2. Likelybig labs add their own verification layers natively.
  3. Possible'correctness-as-a-service' becomes a FinOps line item for AI.
  4. Wild Carda big lab acquires a reliability startup within 12 months.

🥄 The Spoon Take

The whole industry is in a horsepower race. Probably is selling brakes. The thing blocking enterprise AI isn't intelligence, it's trust. If a cheap second model catches the expensive one's lies, reliability becomes a product you buy - not a model you pray to.

🤔 Pushback

A verifier can hallucinate too - if it's wrong about being wrong, you've added cost and a false sense of safety.