Sunday Aug 30

Five Second Video In Three Seconds

30AUG
RENDEREDWATCHING

AI video generation just got faster than watching the video. fal, an AI inference company, shipped H3 Max this week. Generating is no longer the slow part.

H3 Max makes a 5-second clip in under 3 seconds. That is 35 times the throughput of the model it was post-trained from. Ethan Mollick, the Wharton professor, said a line was crossed.

Speed usually costs quality. H3 Max ranks first on human preference against twelve rival video models, including Veo 3.1 and Kling 3. Independent benchmarks from Artificial Analysis and Design Arena agree.

Three seconds turns video from a batch job into a live one. Pieter Levels, the indie builder, called it a historical moment. Watch what creative tools do when the render bar disappears.

full brief & sources

⚡ Why this matters

  • Generation is now faster than playback. The human eye becomes the bottleneck.
  • Interactive video tools become possible. Batch tools and live tools get used differently.
  • The speed-versus-quality tradeoff in generative video just weakened.

🔍 What happened

  • fal Research released H3 Max on August 27, a post-trained version of the open-weights MiniMax H3.
  • A 5-second 768p clip with synced audio renders in about 3 seconds.
  • Roughly 35x the throughput of the official MiniMax H3 endpoint, and 15x faster than comparable-quality models.
  • Ranked #1 on overall quality, prompt understanding, and aesthetics in head-to-head human preference tests against 12 models.
  • Trained and served entirely on NVIDIA GB200 NVL72 systems.
  • Available now on fal, at 50% off for the first week.

💬 Smart takes

  • Ethan Mollick, Wharton professor: a line in AI video was crossed - you can now make reasonably high quality video in less time than it takes to watch it.
  • Pieter Levels, indie builder: “Today is a very historical moment for AI video generation.”
  • Todd Jackson, First Round Capital: “When you give creative people tools like this that are so fast and good, it unlocks incredible new ways of storytelling.”
  • Skeptic: the headline benchmarks come from fal's own preference studies, and 5 seconds at 768p is still a clip, not a scene.

🧭 Where this goes

  1. Likelyrival inference providers ship sub-real-time endpoints for Veo, Kling and Wan within two quarters.
  2. Likelyvideo editors add live preview panels that regenerate as you type the prompt.
  3. Possiblead and social tools ship real-time video generation as a default feature, not a queue.
  4. Wild Carda game or streaming app ships generated video inside the render loop, frame by frame.

🥄 The Spoon Take

The interesting number is not 35x. It is 3 seconds. Under about five seconds, a tool stops feeling like a request and starts feeling like a cursor. Video generation just entered that range. The products built on top will look nothing like the batch tools we have now.

🤔 Pushback

Speed only matters if the clip is usable, and five seconds at 768p still needs a human to stitch anything watchable together.