Friday Jul 10

Nobody Gets An A On AI Safety

10JUL
9 LABSTOP: C+

An outside group just put numbers on AI safety talk. The industry's best score barely cleared passing. One lab led the pack, but no one is close to done.

Nine AI companies just sat for a safety exam. The best grade anyone got was a C-plus.

OpenAI and Google DeepMind both scored a plain C. Meta landed a D-plus; xAI, DeepSeek, and Mistral all failed with an F. The weakest area industry-wide was protecting against long-term existential risk.

Anthropic topped the list, then walked back a safety pledge in February. A report card is not the same thing as a rule.

full brief & sources

Why this matters

  • Buyers now have an independent scorecard for AI-vendor safety, not just marketing claims.
  • Even the top-ranked lab is walking back its own past safety commitments.

🔍 What happened

  • Future of Life Institute published its Summer 2026 AI Safety Index on July 7.
  • Nine major AI companies were graded across six categories: risk assessment, current harms, safety frameworks, existential safety, governance, and disclosure.
  • Anthropic ranked first, with a C+. OpenAI and Google DeepMind both scored C.
  • Meta scored D+. Z.ai and Alibaba Cloud scored D-. xAI, DeepSeek, and Mistral all scored F.
  • Existential safety was the weakest category industry-wide; no company scored above a C-.
  • The report flags companies quietly weakening past safety pledges, including Anthropic's February decision to drop a training-halt commitment.

💬 Smart takes

  • Future of Life Institute: found companies have "recently weakened or abandoned previous commitments to halt development" at certain risk levels.
  • TIME: framed the results bluntly as "nobody gets an A."
  • Skeptic: a nonprofit-run index with no regulatory teeth is a scorecard, not a constraint, on any of these labs.

🧭 Where this goes

  1. Likelyenterprise buyers start citing this index in vendor security reviews and RFPs.
  2. LikelyAnthropic uses its top ranking in sales and PR despite the pledge walk-back.
  3. Possiblethe index becomes an annual ritual that shifts grades slowly instead of forcing change.
  4. Wild Carda regulator cites the index directly in an enforcement action.

🥄 The Spoon Take

A C+ is now the best grade in the entire frontier AI industry. That should be the headline, not which lab won. Self-graded safety pledges are proving soft the moment they cost something, and the one lab bragging about the top score already broke one.

🤔 Pushback

The index is built by an advocacy nonprofit with a strong prior view on AI risk, so its grading rubric isn't a neutral instrument either.