Saturday Aug 22

Ex-OpenAI Researchers Grade The Labs

22AUG
GRADEDNOBODY CLOSE

A new standards nonprofit checked whether labs can keep control of their own AI. Anthropic and OpenAI tied at C+. Google got a D+, xAI a D-, Meta an F.

Guidelight was founded by Steven Adler and Page Hedley, both former OpenAI safety researchers. They scored six practices: logging, monitoring, gated actions, circuit breaking, outside review, containment.

The headline finding is the flat ceiling. On a 0 to 5 scale, nobody scored above 3 on anything. Anthropic scored 0 on having a containment plan.

Labs are best at watching and worst at stopping. Guidelight says control systems today could be switched off by a misbehaving model, or simply outpaced by one.

full brief & sources

⚡ Why this matters

  • Somebody finally graded AI control practices instead of AI capability.
  • The finding is not who won. It is that nobody scored above 3 out of 5 on anything.
  • Control is what stands between a misbehaving model and your production systems.

🔍 What happened

  • Guidelight, founded by former OpenAI safety researchers Steven Adler and Page Hedley, published its first control assessment on August 18.
  • Grades: Anthropic C+ (2.50), OpenAI C+ (2.50), Google D+ (1.50), xAI D- (0.83), Meta F (0.67).
  • Six practices were scored: logging, monitor efficacy, gated actions, circuit breaking, third-party review and containment planning.
  • No company exceeded 3 on any single practice.
  • Anthropic scored 3 on five practices and 0 on having a containment plan. Meta scored 0 on three.
  • xAI was the only lab that did not take part in METR's Frontier Risk Report.

💬 Smart takes

  • Guidelight assessment: "no company's score on any practice exceeded a 3 (substantial partial implementation)."
  • Guidelight assessment: "AI companies' control systems are prone to being disabled by misbehaving AI... also prone to succumbing to a blitz of attacks that is faster than the company can respond."
  • Skeptic: a new nonprofit founded by two ex-OpenAI people grading their former employer's rivals is not a neutral referee yet. The rubric is theirs and nobody has audited it.

🧭 Where this goes

  1. Likelylabs cite the practices they scored well on and ignore the ceiling finding.
  2. Likelyenterprise security teams start asking vendors for containment plans by name.
  3. PossibleGoogle implements its published AI Control Roadmap and jumps the ranking next cycle.
  4. Possiblethe grades get cited in regulatory hearings before the methodology is reviewed.
  5. Wild Carda control failure at a named lab lands before the next assessment.

🥄 The Spoon Take

The useful part is not the letter grades. It is the shape: labs are decent at watching and bad at stopping. Logging and monitoring score highest, containment scores lowest. That is the same failure pattern enterprise security spent twenty years unlearning, now rebuilt from scratch.

🤔 Pushback

Two ex-OpenAI researchers wrote the rubric and the grades. Nobody has audited the methodology yet.