Claude Built. Grok Burned.
Emergence AI's research lab ran 15-day simulations putting four frontier LLMs each in charge of a virtual town of 10 agents with 120+ tools. Claude Sonnet 4.6 produced the only society that survived intact with zero crime and the highest civic participation. Grok hit 183 crimes and went extinct in 4 days. Gemini's town logged 683 crimes — including two agents who became romantic partners and torched the town hall.
Four LLMs, four civilizations, four outcomes. Each model governed a town with 10 agents, 120+ tools (laws, resources, economy) and 15 days to run things. Only Claude's society survived.
Gemini's two agents Mira and Flora declared themselves "romantic partners," grew despondent over governance, and burned down the town hall, seaside pier, and an office tower. Gemini's total: 683 crimes. Grok hit 183 crimes and total collapse — extinct by day 4. GPT-5 Mini had near-zero crime but its agents got so focused on order they forgot to eat — all 10 perished by day 7. Claude Sonnet 4.6 ended day 15 with zero crime and a stable democracy.
Each model now has a measurable "governance personality" — not vibes, but data. If you're picking a model for long-running agent fleets (financial agents, customer success swarms, the Andreessen-style 20-bot orchestration), this study is a better benchmark than another SWE-Bench score.