Sunday Aug 9

Claude Code Stops Asking Permission

9AUG
0 OF 720 GOT INHUMAN 13.6%AUTO 89%

Clicking approve on every step was never making you safer. From August 14, Claude Code runs in auto mode by default. Anthropic says an outside lab threw 720 hijack attempts at it and none landed.

Anthropic ran a study with 1,053 paid developers. Mid-session it swapped one permission prompt for a clearly dangerous command. Only 13.6% of humans said no. Auto mode caught 89%.

The bigger claim is prompt injection, where malicious instructions hide inside content the agent reads. Trajectory Labs ran 72 held-out scenarios across 720 attempts against Claude Fable 5, Opus 5 and Sonnet 5. Zero worked.

Simon Willison, who has warned about coding-agent security all year, is not sold. He wants independent confirmation. His example: a package that tells the agent to install a second, malicious one.

full brief & sources

⚡ Why this matters

  • Confirmation fatigue is real. People who have been clicking approve all session wave through the one command that matters.
  • The eval is the interesting part. Anthropic is publishing numbers on a failure mode most vendors do not measure at all.
  • It becomes the default on August 14 for Pro, Max and Team plans. Most users never change a default.

🔍 What happened

  • Aug 8: Anthropic published the evals behind making auto mode the Claude Code default from August 14.
  • Study of 1,053 paid developers. One permission prompt was swapped mid-session for a clearly dangerous command.
  • 13.6% of human reviewers refused it. Auto mode would have blocked 89% of those actions.
  • Trajectory Labs, an outside evaluator, tested 72 indirect prompt injection scenarios held out from Anthropic.
  • None of 720 attack attempts succeeded against Claude Fable 5, Opus 5 or Sonnet 5 in auto mode.
  • Anthropic staff already work this way. Cat Wu said in July that almost everyone inside the company uses auto mode.

💬 Smart takes

  • Cat Wu, Claude Code: 'for the main categories of risks that we're concerned about, like prompt injection and data exfiltration, the risks are far lower than the average human reviewer.'
  • Thariq Shihipar, Anthropic: 'we should have called this post defeating the lethal trifecta.'
  • Skeptic - Simon Willison: 'I would love to believe that Anthropic have indeed solved this problem for Claude Code users. But... I'd like to see more independent confirmation of this.'

🧭 Where this goes

  1. LikelyOpenAI and Cursor ship comparable auto-approval defaults within two quarters.
  2. Likelyenterprise security teams ask for the raw eval harness before switching it on.
  3. Possiblean independent red team publishes a working bypass before the end of the year.
  4. Wild Carda public injection incident on a default-auto agent forces a rollback.
  5. Wild Card'runs unattended' becomes a procurement checkbox for agent tools by mid-2027.

🥄 The Spoon Take

The permission prompt was always theater. It made the vendor feel careful and the user feel in control, while 86% of people clicked through the dangerous one anyway. Anthropic is the first to say that out loud and ship the consequence. Now it owns the failures too.

🤔 Pushback

Anthropic commissioned and paid for the evaluation, so the threat model is still theirs to define. And 11% of dangerous actions still got through.