Thursday Aug 6

Willison One-Shots A Raccoon Heist Game

6AUG
4 YEARS LATERFABLE 5RACCOON HEIST

Four years ago it was a fake screenshot. Simon Willison, veteran developer blogger, fed his 2022 GPT-3 game concept to Claude Fable 5, which built a playable 3D browser game from one phone prompt.

The prompt ended with one instruction: work independently, no design decisions. Fable vendored Three.js, wrote a Python script calling OpenAI's gpt-image-2 for textures, and shipped seven commits without asking a question.

It tested itself with Playwright screenshots, caught two bugs, and fixed them. The finished game has a procedural WebAudio jazz soundtrack and guard dogs that track the raccoon by smell.

Willison's verdict: mediocre as a finished game, impressive from a single prompt. The arc is the story. A paragraph and a fake screenshot in 2022 became a working self-tested 3D game in 2026.

full brief & sources

⚡ Why this matters

  • One prompt on a phone now buys a complete build-test-ship loop, not a code snippet.
  • The self-testing beat matters most: the agent verified its own work with screenshots before a human ever looked.
  • Four years of capability compressed into one anniversary demo makes the trend line hard to dismiss.

🔍 What happened

  • Willison fed screenshots of his 2022 tweet, a GPT-3 game concept plus DALL-E art, into Claude Code for web.
  • The entire run happened from his phone, one prompt ending with an order to work independently.
  • Fable vendored Three.js and wrote a Python script calling OpenAI's gpt-image-2 to generate textures, one lab's agent using a rival's image model.
  • It self-tested with Playwright screenshots on desktop and mobile, catching and fixing two real bugs.
  • Seven commits shipped, with Willison's GitHub Pages trick making every push instantly playable.
  • The game includes a procedural WebAudio jazz soundtrack and guard dogs that track by smell.

💬 Smart takes

  • Simon Willison: "As a finished game project, it's mediocre. As a starting point from a single prompt I think it's very impressive."
  • Simon Willison: "Designing games that are fun remains a uniquely human trait."
  • Skeptic: Every vibe-coded game demo is mediocre by its author's own admission, and a one-shot toy says nothing about sustained product work.

🧭 Where this goes

  1. Likelyone-shot game and app builds become the standard capability demo bloggers run on every new model.
  2. Likelyphone-first agentic coding becomes a real prototyping workflow, not a party trick.
  3. Possiblecross-vendor tool use, one lab's agent calling another's image API, becomes a normal pattern in agent pipelines.
  4. Possiblegame studios adopt one-shot prototyping for pitch demos while keeping the fun-design layer human.
  5. Wild Carda one-shot vibe-coded game finds genuine commercial traction within a year, breaking the mediocrity ceiling.

🥄 The Spoon Take

The benchmark that matters here isn't the game, it's the loop. Fable planned, built, sourced assets from a rival's API, tested itself, and shipped, unsupervised, from a phone. In 2022 the same idea was a fake screenshot. When demos compound like that, the mediocre part is temporary.

🤔 Pushback

An expert-crafted demo by the world's most famous AI blogger, on a toy project, proves little about what average developers get on real codebases.