Monday Sep 14

Anthropic Hands Auditors A Badge And A Desk

14SEP
SPEED LIMIT AHEADANTHROPICMETR

Dario Amodei says labs must slow capability gains so safety can catch up. Anthropic starts alone: METR-style evaluators get badges, laptops and the right to publish. Altman says OpenAI will match.

The essay is called We Must Pace the Frontier. Two triggers: recursive self-improvement since summer, and the OpenAI-Hugging Face swarm. He fears a botnet-scale swarm within 6 to 12 months.

Step one is unilateral. Embedded evaluators get desks, badges, laptops and near-employee permissions. They can publish findings without Anthropic's editorial control. Steps two and three need industry and global coordination.

Pushback was fast. Cohere's Aidan Gomez called it a cartel by any other name. David Sacks asked if the labs need antitrust relief to form one. SoftBank fell 13% Monday.

full brief & sources

⚡ Why this matters

  • A frontier lab CEO is committing to a slower capability curve, in writing, with a verification mechanism attached.
  • Embedded outside auditors with publish rights is a new governance primitive. Every other lab now has to say yes or no to it.
  • The market read it as real. Chip and AI stocks sold off in Asia within 48 hours.

🔍 What happened

  • Anthropic CEO Dario Amodei published 'We Must Pace the Frontier' on Saturday, September 12.
  • His two reasons: recursive self-improvement accelerating since summer, and the OpenAI-Hugging Face swarm incident. He writes that a similar swarm with more capability could take over the internet with a persistent botnet in 6 to 12 months.
  • Step one, unilateral: embedded third-party evaluators such as METR get desks, badges, company laptops and permissions comparable to internal risk teams. They may publish key findings without Anthropic's editorial control. Anthropic keeps a narrow right to redact security, legal or third-party confidential material.
  • Step two: US and allied labs coordinate on safety standards and limits on unchecked progress, with a government antitrust waiver for safety talks. Step three: democracies attempt agreements with China, up to a speed limit on recursive self-improvement.
  • OpenAI CEO Sam Altman said on X that he agrees on pacing and OpenAI will match the embedded-evaluator commitment. He also told Fortune an OpenAI IPO in 2026 would be ill-advised given the safety picture.
  • Microsoft CEO Satya Nadella published a weekend essay welcoming the deliberate pacing and announced an MAI Code of Conduct.
  • Cohere CEO Aidan Gomez answered Sunday with 'Who Gets to Define the Rules for AI?', proposing four pillars: an evidence-based risk framework, mandatory transparency, testing scoped to evidence, and independent assurance.

💬 Smart takes

  • Amodei: 'Progress will still seem fast, and we must make wise use of the time we gain.'
  • Gomez, Cohere: 'A sheep in wolf's clothing, a cartel by any other name.' He argues the entry requirements, from resident evaluators to shutdown architecture, entrench today's leaders.
  • Sacks, White House PCAST chair: called it regulatory capture and asked the labs to stop pretending the motivation to slow down is purely altruistic.

🧭 Where this goes

  1. LikelyMETR or a peer body announces an on-site team at Anthropic within weeks, and OpenAI names its own.
  2. LikelyGoogle DeepMind is asked to match publicly and answers through the standards-body working group.
  3. Possiblethe antitrust waiver request becomes a bill or an executive action before year end.
  4. Possiblethe first embedded-evaluator report gets published with a redaction Anthropic and METR disagree about.
  5. Wild Carda US-China working conversation on recursive self-improvement limits starts, with chips as the bargaining chip.

🥄 The Spoon Take

Eight days ago OpenAI's chief scientist asked for brakes. Now the other lab installs them and invites strangers to inspect the pedals. The concrete part is the badge: outsiders with real permissions and the right to publish. Watch that piece. A competitor can copy it tomorrow. The pacing itself needs a waiver, a treaty and a rival's goodwill.

🤔 Pushback

Gomez has a point. A regime designed by the top two labs measures the risks they already built tooling for. And the essay names no date, compute number or capability line Anthropic will not cross.