Anthropic Raises Its Own Risk Rating
18AUG
Anthropic just told the world its models got riskier. Its new 186-page risk report moves misalignment from very low to low. It also reveals Model 2, a stronger internal model it won't release.
The report runs 186 pages under version 3.4 of Anthropic's Responsible Scaling Policy. The label moved not because of a new failure, but because recent incident disclosures increased overall uncertainty.
The bigger reveal is Model 2. It outperforms the public Mythos 5, and Anthropic is keeping it internal. A frontier lab now treats its best model as too sensitive to ship. That's new.
The contrast writes itself. The same week, The Verge reported OpenAI disbanded its preparedness team and spread the work across product groups. Two labs, one question, opposite answers.