Wednesday Sep 2
SOLARIS NO CODE

Runway announced a model that draws working app screens instead of code. Solaris renders each frame as you click. No HTML, no CSS, no JavaScript underneath.

It sits on Gen-4.5, the company's video system. Your clicks and drags become the prompt for the next image. The picture is the software.

In a 250-person study, testers preferred it over Claude Opus 5 for following instructions, 61% to 24%. On behaving naturally it won 71% to 21%.

Nobody can use it yet. Early access is a form. Banking and healthcare are out entirely, because there is no DOM for auditors or screen readers to read.

full brief & sources

⚡ Why this matters

  • A generated interface has no code artifact. The screen itself is the product.
  • If it holds, the gap between a convincing prototype and a shipped app narrows to a prompt.
  • It also breaks everything that expects a DOM: tests, accessibility tools, analytics.

🔍 What happened

  • Runway announced Solaris on August 31, 2026.
  • It builds on Gen-4.5, the company's video generation model.
  • User clicks and drags become the prompt for the next rendered frame.
  • There is no HTML, CSS, or JavaScript behind the interface.
  • A 250-person study preferred it over Claude Opus 5 on instruction-following, 61% to 24%.
  • On behaving naturally it won 71% to 21%.
  • Access is an early-access form; banking and healthcare uses are excluded.

💬 Smart takes

  • Runway: describes Solaris as rendering interfaces frame by frame in response to input, rather than generating code a browser runs.
  • Skeptic: with no DOM there is nothing for a screen reader, an automated test, or an auditor to read.

🧭 Where this goes

  1. Likelyearly access stays narrow through the end of 2026.
  2. Likelyrivals ship generated-UI demos within months, still code-backed underneath.
  3. Possiblefirst real uses are games, toys, and marketing pages rather than apps.
  4. Wild Cardsomeone ships it in a regulated flow and an accessibility complaint stops it.

🥄 The Spoon Take

Generated pixels can look perfect and still fail the boring parts. A stable layout on the fiftieth click. A form that submits the same way twice. A page a screen reader can read. Runway excluding banking and healthcare on day one is the honest tell about where this actually is.

🤔 Pushback

The preference numbers come from Runway's own study, not an independent benchmark. Nobody outside early access has used it, and short tests say nothing about long sessions.

Monday Aug 31
APACHE 2.0MICRODUCK

Open-source robotics just hit a hobbyist price. Hugging Face is selling Microduck, a 9.8-inch bipedal robot, for $399. CEO Clem Delangue calls it a robot you teach new tricks.

It walks, sits, kicks a ball, grabs things with its beak, and stands back up after falling. Fifteen motors, a wide-angle camera, LiDAR, mics and a speaker.

The whole stack is Apache 2.0. SDK, MuJoCo simulation, the reinforcement learning training code, and the seven shipped behaviors. You can inspect every policy and retrain it.

Pollen Robotics built it, the French company Hugging Face bought in 2025. First units ship before Christmas. Nvidia agreed to buy Hugging Face two days before this launch.

full brief & sources

⚡ Why this matters

  • Reinforcement learning on real hardware has been a lab sport because the hardware cost five figures. $399 changes who gets to try.
  • Shipping the training stack, not just the SDK, means the community can produce new behaviors the vendor never wrote.
  • It lands two days after Nvidia agreed to buy Hugging Face, so the open-hardware bet now sits inside a chip company.

🔍 What happened

  • Announced August 27. Preorder at $399 before tax and shipping, first deliveries before Christmas 2026.
  • 9.8 inches tall, bipedal, 15 motors.
  • Sensors: wide-angle camera, LiDAR, microphones, speaker, two IMUs, NFC, Wi-Fi, Bluetooth.
  • Seven pre-trained behaviors: walking, sitting and standing, kicking, grabbing, roller skating, recovering from a fall.
  • SDK, MuJoCo simulation environment and RL training stack published under Apache 2.0.
  • Built by Pollen Robotics, the French robotics firm Hugging Face acquired in 2025.

💬 Smart takes

  • Clem Delangue, Hugging Face CEO: an "open-source robot you can teach new tricks with reinforcement learning."
  • The Register: framed it as a teaching tool for people who want to poke at RL rather than a product with a job to do.
  • Skeptic: a duck that roller skates is a great demo and a weak business. Preorders shipping four months out have a habit of slipping.

🧭 Where this goes

  1. Likelyuniversity and bootcamp robotics courses adopt it as the default hardware inside two semesters.
  2. Likelycommunity-trained policies outnumber the seven shipped ones within six months.
  3. Possiblea Chinese manufacturer clones the form factor at half the price before the first units ship.
  4. PossibleNvidia folds it into a Jetson-branded learning kit after the acquisition closes.
  5. Wild Cardthe community policy library becomes the actual product, and the duck turns into a loss-leader for a robot behavior hub.

🥄 The Spoon Take

The price is the strategy. Hugging Face did this once already with models, giving away the weights until the hub became the place everyone went. A $399 body that anyone can retrain is the same move in hardware. Worth watching whether the policies, not the duck, end up being the asset.

🤔 Pushback

Cheap hardware with a real bill of materials is a thin-margin trap, and preorders shipping in December are not shipped.

Sunday Aug 30
RENDEREDWATCHING

AI video generation just got faster than watching the video. fal, an AI inference company, shipped H3 Max this week. Generating is no longer the slow part.

H3 Max makes a 5-second clip in under 3 seconds. That is 35 times the throughput of the model it was post-trained from. Ethan Mollick, the Wharton professor, said a line was crossed.

Speed usually costs quality. H3 Max ranks first on human preference against twelve rival video models, including Veo 3.1 and Kling 3. Independent benchmarks from Artificial Analysis and Design Arena agree.

Three seconds turns video from a batch job into a live one. Pieter Levels, the indie builder, called it a historical moment. Watch what creative tools do when the render bar disappears.

full brief & sources

⚡ Why this matters

  • Generation is now faster than playback. The human eye becomes the bottleneck.
  • Interactive video tools become possible. Batch tools and live tools get used differently.
  • The speed-versus-quality tradeoff in generative video just weakened.

🔍 What happened

  • fal Research released H3 Max on August 27, a post-trained version of the open-weights MiniMax H3.
  • A 5-second 768p clip with synced audio renders in about 3 seconds.
  • Roughly 35x the throughput of the official MiniMax H3 endpoint, and 15x faster than comparable-quality models.
  • Ranked #1 on overall quality, prompt understanding, and aesthetics in head-to-head human preference tests against 12 models.
  • Trained and served entirely on NVIDIA GB200 NVL72 systems.
  • Available now on fal, at 50% off for the first week.

💬 Smart takes

  • Ethan Mollick, Wharton professor: a line in AI video was crossed - you can now make reasonably high quality video in less time than it takes to watch it.
  • Pieter Levels, indie builder: “Today is a very historical moment for AI video generation.”
  • Todd Jackson, First Round Capital: “When you give creative people tools like this that are so fast and good, it unlocks incredible new ways of storytelling.”
  • Skeptic: the headline benchmarks come from fal's own preference studies, and 5 seconds at 768p is still a clip, not a scene.

🧭 Where this goes

  1. Likelyrival inference providers ship sub-real-time endpoints for Veo, Kling and Wan within two quarters.
  2. Likelyvideo editors add live preview panels that regenerate as you type the prompt.
  3. Possiblead and social tools ship real-time video generation as a default feature, not a queue.
  4. Wild Carda game or streaming app ships generated video inside the render loop, frame by frame.

🥄 The Spoon Take

The interesting number is not 35x. It is 3 seconds. Under about five seconds, a tool stops feeling like a request and starts feeling like a cursor. Video generation just entered that range. The products built on top will look nothing like the batch tools we have now.

🤔 Pushback

Speed only matters if the clip is usable, and five seconds at 768p still needs a human to stitch anything watchable together.

Thursday Aug 27
PLAUD ONENO PHONE

Plaud put a cellular connection in the charging case. The AI agent stays online with no phone nearby, records the room, and writes the follow-up email. $249, ships late this year.

The case is the interesting part. It has an eSIM with 4G, a speaker, and its own microphones. Leave the earbuds in, set the case on the table, and it captures the meeting.

The agent connects to Gmail, Calendar, Notion and Slack. After a conversation it can draft the follow-up or build the doc. Plaud says the data plan costs the buyer nothing extra.

Plaud is not a demo company. It sells recorders, has millions of users, and this is its third hardware product. The agent features arrive after launch, which is the part to watch.

full brief & sources

⚡ Why this matters

  • Every AI wearable so far has been a phone accessory. Putting the modem in the case makes the phone optional, which is a different product category.
  • The agent is not just transcribing. It has write access to Gmail, Calendar, Notion and Slack, so the output is an action, not a summary.
  • Plaud has real distribution already. This is not a Kickstarter promise from a company with no shipped hardware.

🔍 What happened

  • Aug 27 - Plaud unveils Plaud One at IFA. $249 pre-order, shipping in the fourth quarter of 2026.
  • The charging case carries a built-in eSIM with 4G LTE, provided at no extra cost to the buyer.
  • Earbuds handle calls, meetings, music and agent activation. The case adds in-person capture, cellular, speaker output and charging.
  • Recordings upload for transcription and summarization even when the phone is offline or out of range.
  • The agent connects to Gmail, Google Calendar, Notion and Slack, and can draft follow-up emails, documents and presentations.
  • Plaud says the most ambitious agent features land after launch, not at ship.

💬 Smart takes

  • TechCrunch: the eSIM lives in the case so earbuds and case stay connected when the phone is offline or out of range.
  • TechRadar: Plaud's framing is that the user stays in control, and the deeper agent features are coming later.
  • Skeptic: a recorder you set on the table is a consent problem in most meeting rooms, and a case that is always online makes it easier to forget one is there.

🧭 Where this goes

  1. Likelythe agent features slip past Q4 and land in 2027, given Plaud already says they come after launch.
  2. Likelyrivals add cellular to their own cases within two product cycles.
  3. Possibleenterprise buyers block it outright on recording-consent grounds before it ships.
  4. Possiblea phone maker responds by opening up on-device agent APIs to keep accessories dependent.
  5. Wild Cardthe case, not the earbuds, becomes the product, and the next version drops the buds entirely.

🥄 The Spoon Take

The quiet move here is the modem. Once the accessory has its own connection, the phone stops being the hub and becomes one more screen. That is a bigger shift than better transcription, and it is being shipped by a note-taking company rather than by anyone building a phone.

🤔 Pushback

The agent features that make this interesting are not in the box on day one, and Plaud says so.

Wednesday Aug 26
FREE MODELYOUR CODE

A free coding model showed up on model marketplace OpenRouter with no company name. It beat the big labs on a 10-task sample. The full benchmark says otherwise.

Developer Ben Davis ran 10 tasks from DeepSWE, a software engineering test. Ox Alpha hit 80 percent. Claude Fable 5 got 65. On the full 113-task run it scored 58.4.

Then developers went hunting. Tokenizer fingerprints and video encoder behavior point at Zhipu AI in Beijing. No lab has confirmed anything.

The free window closes around August 27. Prompts sent through OpenRouter are retained by a provider that will not name itself. Read the terms before you paste your codebase in.

full brief & sources

⚡ Why this matters

  • A frontier-class coding model is being handed out free by an operator nobody can name.
  • The gap between the 10-task headline and the full-benchmark number is a lesson in how AI benchmarks get sold.
  • Anyone piping a work codebase into a free endpoint is making a data decision without a counterparty.

🔍 What happened

  • Ox Alpha appeared on OpenRouter and OpenCode on August 20, 2026, free for roughly a week.
  • It offers a 1,048,576-token context window and accepts text, images and video.
  • Ben Davis ran a hand-picked 10-task DeepSWE sample: Ox Alpha 80%, Claude Fable 5 65%, GPT-5.6 Sol 52%.
  • On the full DeepSWE run of 113 tasks across 91 repositories, Ox Alpha scored 58.4%.
  • The official DeepSWE leaderboard lists Claude Opus 5 at 74% and GPT-5.6 Sol at 73%.
  • Tokenizer and video-encoder fingerprints point at Zhipu AI's unreleased GLM flagship. Zhipu has not confirmed.
  • The OpenRouter route retains prompts and completions; the OpenCode route states zero retention. Two routes, two policies.

💬 Smart takes

  • Ben Davis: ran the 10-task sample that produced the 80% number now quoted everywhere.
  • The benchmark read: a hand-picked 10 of 113 tasks can favor any model through selection alone.
  • Skeptic: retention by an unnamed provider means you cannot read the policy, cannot request deletion, and cannot name who holds your code.

🧭 Where this goes

  1. Likelya named lab claims Ox Alpha within weeks, as happened with earlier stealth models.
  2. Likelythe free window closes and the model reappears on paid pricing under a real name.
  3. PossibleZhipu confirms and this becomes a GLM launch story rather than a mystery.
  4. Wild Carda company discovers proprietary code went through the OpenRouter route and it turns into a compliance incident.

🥄 The Spoon Take

Two separate things got fused into one headline. A stealth model scoring well on someone's 10-task sample is not the same as beating Opus 5. The full run puts Ox Alpha 15 points behind. What is actually interesting is the setup: free frontier-grade inference, anonymous operator, prompts retained. That trade is the news, not the leaderboard.

🤔 Pushback

Stealth previews are a normal pre-launch practice and OpenRouter's terms are public. A 58.4% full run is respectable for an unannounced model. The hype is the coverage, not the model.

Monday Aug 24
VIA CURSORSTAYS ON

Each agent gets its own cloud computer that stays on. SpaceXAI opened Grok Bot to Cursor Pro+ and SuperGrok Plus subscribers. The distribution runs through Cursor, not xAI.

Grok Bot launched in beta on August 11. Ten days later access widened to Cursor Pro+ at $60 a month and Cursor Teams Standard at $40 a seat.

The agent keeps a persistent cloud machine and works across your existing tools. It is not a chat window. It is a teammate with its own desk.

SpaceX bought Cursor's parent Anysphere for $60 billion in June. This is the first product where that deal shows up on a pricing page.

full brief & sources

⚡ Why this matters

  • Persistent per-agent compute is a different product shape than a chat session.
  • The pricing page shows the acquisition working as distribution, not just as a headline.
  • Ten days from closed beta to paid tiers is fast even by 2026 standards.

🔍 What happened

  • Grok Bot shipped in beta on August 11, 2026, on Mac and iOS first, with Windows and Linux available and Android to follow.
  • August 21: access expanded to SuperGrok Plus, Cursor Pro+, and Cursor Teams Standard, plus a limited free trial for everyone else.
  • Each agent gets its own persistent cloud computer to run multi-step work across a user's existing tools.
  • Tiers now include Cursor Pro+ at $60 a month, Cursor Teams Standard at $40 a seat, SuperGrok Plus at $100, and SuperGrok Heavy near $300.
  • SpaceX acquired Anysphere, Cursor's parent, for $60 billion in an all-stock deal announced June 16.

💬 Smart takes

  • SpaceXAI product framing: an always-on AI teammate rather than an assistant you open.
  • 9to5Mac: the app arrives as a joint SpaceXAI and Cursor release, not an xAI one.
  • Skeptic: gating a flagship agent behind an IDE subscription narrows it to developers, which is the opposite of an always-on teammate for everyone.

🧭 Where this goes

  1. Likelyrivals ship persistent per-agent compute as a tier within two quarters.
  2. LikelyCursor becomes the main distribution channel for SpaceXAI consumer agents.
  3. Possibleper-agent machine hours appear as a separate billing line rather than a bundled perk.
  4. Wild Cardthe always-on model runs into cost reality and the free trial quietly disappears.

🥄 The Spoon Take

Bundling is the tell. SpaceXAI could have sold Grok Bot on its own and did not. It put the agent inside the subscription developers already pay for, which is the cheapest distribution in software. The $60 billion Cursor deal is starting to look like a channel purchase.

🤔 Pushback

Persistent cloud machines per agent are expensive to run, and the tier bundling may be a launch subsidy rather than the real price.

WASTED CALLSAGENT INDEX

Coding agents waste most tool calls hunting for docs. Firecrawl launched an index of 70 million READMEs, issues, and pull requests for agents. It beats general web search by 18 points on recall.

It holds artifacts, not web pages. Repo landing files, bug threads, merge requests, API specs. Refreshed every day. No source code is stored.

An open benchmark, DevDex, scores 1,179 real developer questions. The new layer hit 0.63. Parallel got 0.57. Context7 managed 0.17.

Claude Code, Cursor, and Codex can query it now. Two credits per ten results. No key needed to start.

full brief & sources

⚡ Why this matters

  • Retrieval quality is now the ceiling on coding-agent quality, not model quality.
  • The 0.45 to 0.63 recall gap is the size of the problem general web search leaves on the table.
  • A public benchmark means agent retrieval can finally be shopped and compared like a model.

🔍 What happened

  • August 20, 2026: Firecrawl shipped Developer Index, a retrieval layer built for coding agents.
  • 70 million-plus artifacts: READMEs, external docs, issues, pull requests, OpenAPI specs, skill files. Most refreshed daily.
  • Reached through /search/developer or standard /search with a developer category. Stable IDs like issue:owner/repo#123.
  • DevDex ships alongside: 1,179 developer queries across repo discovery, docs lookup, and issue resolution. Half the dataset and the harness are open source.
  • Scores: Firecrawl Developer Index 0.63, Firecrawl Search 0.58, Parallel 0.57, Mintlify and Exa 0.54, native web search 0.45, Context7 0.17.
  • No API key needed to start. Two credits per ten results.

💬 Smart takes

  • Neha Patil, research engineer at Firecrawl: a coding agent does not want a web page, it wants an artifact.
  • Firecrawl on the control group: the gap between no tools and every other row is the size of the retrieval problem.
  • Skeptic: the benchmark was built by the vendor that wins it, and Parallel still beats it on repository discovery at 0.82 to 0.76.

🧭 Where this goes

  1. Likelyrival providers publish DevDex numbers within a quarter, because the harness is open.
  2. Likelyagent harnesses start shipping a default retrieval provider the way they ship a default model.
  3. PossibleGitHub responds with its own agent-facing artifact API rather than letting a third party index it.
  4. Wild Cardretrieval recall becomes a line item in enterprise agent contracts alongside token price and latency.

🥄 The Spoon Take

The interesting move is not the index. It is the benchmark. Firecrawl published the scoreboard it currently leads, which invites everyone to beat it and makes agent retrieval a measured category instead of a vibe. That is how a commodity layer turns into a market.

🤔 Pushback

Vendor-built benchmarks flatter their authors, and a 0.18 recall gap may not survive contact with agents that just retry three times.

Monday Aug 17
XHIGH MODEPELICAN

Alibaba's new open model runs on a laptop and codes, sees, and calls tools. Simon Willison calls it the best local model he's tested. One catch: the default setting overthinks everything.

Qwen 3.8 27B is a 17GB file. It handles vision, tool use, and a 262,000-token context. Willison ran it on a MacBook and an NVIDIA Spark.

The default reasoning mode burns tokens on everything. One pelican drawing took 21 minutes and 22,000 thinking tokens. Turn reasoning down and the same job takes two minutes.

It also drove a real coding agent and nailed image bounding boxes. Willison's verdict: a miracle file, held back only by speed.

full brief & sources

⚡ Why this matters

  • A 17GB open file now covers vision, coding, tool use, and long context on consumer hardware.
  • Local models this capable change the math on privacy-sensitive and offline AI work.
  • Reasoning-effort defaults matter: the same model is brilliant or comically slow depending on one setting.

🔍 What happened

  • Alibaba's Qwen lab released Qwen 3.8 27B on Friday - Apache 2 licensed, vision-capable, 262K context.
  • Simon Willison ran the 17GB build on an M5 MacBook Pro and an NVIDIA DGX Spark.
  • At the default xhigh reasoning setting, one pelican SVG took 21 minutes and 22,276 reasoning tokens.
  • With reasoning off, the same prompt finished in just over two minutes.
  • The model drove the Pi coding agent through real repo questions and built a working bounding-box tool.
  • A llama.cpp multi-token-prediction build ran about 72% faster than the LM Studio default.

💬 Smart takes

  • Simon Willison: "The fact that a 17GB file can do all of this stuff on my home machines is a miracle."
  • Willison on the default: "This is a hilarious default. It's absolutely not a good way to run the model."
  • Skeptic: at 15-30 tokens per second, hosted APIs still answer several times faster - speed, not smarts, keeps local models off daily-driver duty.

🧭 Where this goes

  1. Likelycommunity MTP and MLX optimizations cut local inference times sharply within weeks.
  2. Likelyindependent benchmarks confirm the 27B beats Qwen's own closed 3.7-Plus on several tasks.
  3. Possiblereasoning-effort dials become a standard control across all open model releases.
  4. Wild Carda laptop-class open model becomes the default coding agent driver for small teams within a year.

🥄 The Spoon Take

The story isn't one model - it's the floor rising. A free 17GB file now does what expensive hosted setups did a year ago. When capability is solved at laptop scale, speed becomes the last moat for the API labs.

🤔 Pushback

Willison's 15-30 tokens per second is painful in practice - most users will drift back to fast hosted models within a week.

Monday Aug 10
ONE FORMAT SIX RIVALS

Every AI coding tool used to speak its own dialect. OpenAI and five rivals just agreed on one shared skill format. Build a skill once, and it now works everywhere.

Every AI vendor built its own plugin format. Skills built for one tool never worked in another.

Agent Plugins packages a skill and its MCP config into one folder. Vercel proposed it; OpenAI, AWS, Microsoft, GitHub, and Cursor signed on. Launch clients include ChatGPT, Codex, Copilot, and VS Code.

No single company controls the format going forward. Expect every major agent platform to adopt it within a year.

full brief & sources

⚡ Why this matters

  • Ends months of developers rebuilding the same skill for every AI tool.
  • First time OpenAI, Microsoft, and rivals ship shared infrastructure instead of competing on it.
  • Sets the packaging layer other agent tools will build on next.

🔍 What happened

  • Vercel drafted the proposal; AWS, Cursor, GitHub, Microsoft, and OpenAI refined it into Agent Plugins 1.0.
  • A plugin bundles Agent Skills (reusable instructions) and MCP server configs (tool connections) in one package.
  • Announced August 6, 2026, timed to GPT-5's first anniversary.
  • Launch clients: ChatGPT, Codex, Cursor, GitHub Copilot, Kiro, and VS Code.
  • Google joined as a core maintainer the same day.
  • The spec is openly licensed with a public technical steering committee, not one company's roadmap.

💬 Smart takes

  • Vercel: frames it as build a plugin once, use it across every compatible agent client.
  • AWS Open Source Blog: positions it as ending duplicate work for portable agent extensions.
  • Skeptic (AgentPatterns.ai): the spec deliberately skips distribution, permissions, and trust, so teams on a single client gain nothing.

🧭 Where this goes

  1. Likelyevery major coding assistant (Copilot, Cursor, Codex) supports Agent Plugins within 6 months.
  2. Likelya plugin marketplace emerges, run by one of the six backers, not a neutral body.
  3. PossibleGoogle's same-day maintainer seat means Gemini clients adopt it by year-end.
  4. Wild Cardthe standard fragments again within 18 months as vendors bolt on proprietary extensions.

🥄 The Spoon Take

Six rivals just agreed a plugin should look the same everywhere. That's rare in AI, where everyone competes on lock-in. The format is a truce, not a product. Whoever builds the marketplace on top of it wins, and the spec conveniently skips that part.

🤔 Pushback

The spec skips distribution, permissions, and trust, so single-client teams gain nothing and multi-client teams still need a marketplace layer.

Sunday Aug 9
0 OF 720 GOT INHUMAN 13.6%AUTO 89%

Clicking approve on every step was never making you safer. From August 14, Claude Code runs in auto mode by default. Anthropic says an outside lab threw 720 hijack attempts at it and none landed.

Anthropic ran a study with 1,053 paid developers. Mid-session it swapped one permission prompt for a clearly dangerous command. Only 13.6% of humans said no. Auto mode caught 89%.

The bigger claim is prompt injection, where malicious instructions hide inside content the agent reads. Trajectory Labs ran 72 held-out scenarios across 720 attempts against Claude Fable 5, Opus 5 and Sonnet 5. Zero worked.

Simon Willison, who has warned about coding-agent security all year, is not sold. He wants independent confirmation. His example: a package that tells the agent to install a second, malicious one.

full brief & sources

⚡ Why this matters

  • Confirmation fatigue is real. People who have been clicking approve all session wave through the one command that matters.
  • The eval is the interesting part. Anthropic is publishing numbers on a failure mode most vendors do not measure at all.
  • It becomes the default on August 14 for Pro, Max and Team plans. Most users never change a default.

🔍 What happened

  • Aug 8: Anthropic published the evals behind making auto mode the Claude Code default from August 14.
  • Study of 1,053 paid developers. One permission prompt was swapped mid-session for a clearly dangerous command.
  • 13.6% of human reviewers refused it. Auto mode would have blocked 89% of those actions.
  • Trajectory Labs, an outside evaluator, tested 72 indirect prompt injection scenarios held out from Anthropic.
  • None of 720 attack attempts succeeded against Claude Fable 5, Opus 5 or Sonnet 5 in auto mode.
  • Anthropic staff already work this way. Cat Wu said in July that almost everyone inside the company uses auto mode.

💬 Smart takes

  • Cat Wu, Claude Code: 'for the main categories of risks that we're concerned about, like prompt injection and data exfiltration, the risks are far lower than the average human reviewer.'
  • Thariq Shihipar, Anthropic: 'we should have called this post defeating the lethal trifecta.'
  • Skeptic - Simon Willison: 'I would love to believe that Anthropic have indeed solved this problem for Claude Code users. But... I'd like to see more independent confirmation of this.'

🧭 Where this goes

  1. LikelyOpenAI and Cursor ship comparable auto-approval defaults within two quarters.
  2. Likelyenterprise security teams ask for the raw eval harness before switching it on.
  3. Possiblean independent red team publishes a working bypass before the end of the year.
  4. Wild Carda public injection incident on a default-auto agent forces a rollback.
  5. Wild Card'runs unattended' becomes a procurement checkbox for agent tools by mid-2027.

🥄 The Spoon Take

The permission prompt was always theater. It made the vendor feel careful and the user feel in control, while 86% of people clicked through the dangerous one anyway. Anthropic is the first to say that out loud and ship the consequence. Now it owns the failures too.

🤔 Pushback

Anthropic commissioned and paid for the evaluation, so the threat model is still theirs to define. And 11% of dangerous actions still got through.

Thursday Aug 6
$4.25META$0.20

Cheap AI now costs your privacy. Meta's new coding agent comes with two price tags: full rate, or 92% off if the company can learn from everything you type.

Muse Spark 1.2 ships with Muse Code, Meta's first terminal harness. It hits 82.9% on Terminal-Bench, up from 76.2% for version 1.1.

Same weights, two model IDs. The standard one runs $1.25 in and $4.25 out per million tokens. The contributor tier drops to $0.10 and $0.20 - in exchange for training rights on your prompts and completions.

Simon Willison, the blogger who tracks every release, called long-sequence tool calling the trait that matters most now. Data-for-discount could become a standard axis - watch who copies it.

full brief & sources

⚡ Why this matters

  • Data-for-discount is a new pricing axis for frontier models.
  • A 92% discount will tempt startups to trade user prompts for margin.
  • Coding agents are now the battleground - every lab ships its own harness.

🔍 What happened

  • Meta released Muse Spark 1.2 and Muse Code, its first terminal coding agent, on August 5.
  • The model scores 82.9% on Terminal-Bench 2.1, up from 76.2% for Spark 1.1.
  • Standard pricing: $1.25 per million input tokens, $4.25 output - unchanged from 1.1.
  • The contributor tier drops that to $0.10 and $0.20 if Meta may train on your prompts and completions.
  • Model and agent were co-trained so the pair performs best together.

💬 Smart takes

  • Simon Willison, AI blogger: "Yet more evidence that the most important characteristic of any model these days is long-sequence agentic tool calling."
  • Meta: the model was "extensively trained on long-horizon coding tasks, including whole-repository generation."
  • Skeptic: enterprises with sensitive codebases cannot touch the contributor tier - the discount targets exactly the users whose data Meta wants most.

🧭 Where this goes

  1. Likelyat least one other lab ships a train-on-my-data discount tier within six months.
  2. Likelyenterprise procurement teams write explicit bans on contributor-tier model IDs.
  3. Possibleregulators examine whether a 92% discount makes data consent meaningful.
  4. Wild Cardcontributor pricing becomes the free tier of the agent era - pay with data or pay with cash.

🥄 The Spoon Take

Every lab wants your agent trajectories - they are the scarcest training data left. Meta just put a public price on them: about $4 per million tokens. The discount is the tell. Watch which developers take the deal; that cohort shows how cheap data privacy really is.

🤔 Pushback

Meta has offered data-sharing discounts before without moving the market - most serious buyers default to the private tier and the headline rate.

4 YEARS LATERFABLE 5RACCOON HEIST

Four years ago it was a fake screenshot. Simon Willison, veteran developer blogger, fed his 2022 GPT-3 game concept to Claude Fable 5, which built a playable 3D browser game from one phone prompt.

The prompt ended with one instruction: work independently, no design decisions. Fable vendored Three.js, wrote a Python script calling OpenAI's gpt-image-2 for textures, and shipped seven commits without asking a question.

It tested itself with Playwright screenshots, caught two bugs, and fixed them. The finished game has a procedural WebAudio jazz soundtrack and guard dogs that track the raccoon by smell.

Willison's verdict: mediocre as a finished game, impressive from a single prompt. The arc is the story. A paragraph and a fake screenshot in 2022 became a working self-tested 3D game in 2026.

full brief & sources

⚡ Why this matters

  • One prompt on a phone now buys a complete build-test-ship loop, not a code snippet.
  • The self-testing beat matters most: the agent verified its own work with screenshots before a human ever looked.
  • Four years of capability compressed into one anniversary demo makes the trend line hard to dismiss.

🔍 What happened

  • Willison fed screenshots of his 2022 tweet, a GPT-3 game concept plus DALL-E art, into Claude Code for web.
  • The entire run happened from his phone, one prompt ending with an order to work independently.
  • Fable vendored Three.js and wrote a Python script calling OpenAI's gpt-image-2 to generate textures, one lab's agent using a rival's image model.
  • It self-tested with Playwright screenshots on desktop and mobile, catching and fixing two real bugs.
  • Seven commits shipped, with Willison's GitHub Pages trick making every push instantly playable.
  • The game includes a procedural WebAudio jazz soundtrack and guard dogs that track by smell.

💬 Smart takes

  • Simon Willison: "As a finished game project, it's mediocre. As a starting point from a single prompt I think it's very impressive."
  • Simon Willison: "Designing games that are fun remains a uniquely human trait."
  • Skeptic: Every vibe-coded game demo is mediocre by its author's own admission, and a one-shot toy says nothing about sustained product work.

🧭 Where this goes

  1. Likelyone-shot game and app builds become the standard capability demo bloggers run on every new model.
  2. Likelyphone-first agentic coding becomes a real prototyping workflow, not a party trick.
  3. Possiblecross-vendor tool use, one lab's agent calling another's image API, becomes a normal pattern in agent pipelines.
  4. Possiblegame studios adopt one-shot prototyping for pitch demos while keeping the fun-design layer human.
  5. Wild Carda one-shot vibe-coded game finds genuine commercial traction within a year, breaking the mediocrity ceiling.

🥄 The Spoon Take

The benchmark that matters here isn't the game, it's the loop. Fable planned, built, sourced assets from a rival's API, tested itself, and shipped, unsupervised, from a phone. In 2022 the same idea was a fake screenshot. When demos compound like that, the mediocre part is temporary.

🤔 Pushback

An expert-crafted demo by the world's most famous AI blogger, on a toy project, proves little about what average developers get on real codebases.

Sunday Aug 2
4 TEAMS, 1 BOTQM

Y Combinator gave away the AI tool it runs itself on. QM is an open-source, MIT-licensed harness spanning accounting, legal, events, and engineering. It swaps between Claude Code, Codex, and other models with zero lock-in.

Every YC staffer gets a private, sandboxed workspace with its own memory, files, permissions, and scheduled jobs. The team says it even used the system to build itself, real-world proof it holds up under daily use.

Pick your engine: Pi, OpenCode, Codex, or Claude Code, all interchangeable behind one Slack and web interface. No procurement process sits between an employee and their own automation.

The logic: agent orchestration is plumbing, not a product edge, so hoarding it buys little. Expect more startups to publish their internal stacks now that YC set the norm.

full brief & sources

⚡ Why this matters

  • Most companies still treat AI agents as single-purpose chatbots bolted onto one app, not shared infrastructure.
  • QM gives every employee, not just engineers, a scoped agent workspace with its own memory and permissions.
  • Open-sourcing the exact tool you run your company on is a rare, credible adoption signal.

🔍 What happened

  • Y Combinator open-sourced QM on July 31 under an MIT license, with the code on GitHub.
  • YC uses it daily across accounting, legal, events, and engineering, including building QM itself.
  • Each person and each room gets scoped memory, files, permissions, crons, and a durable sandbox.
  • It works with Pi, OpenCode, Codex, and Claude Code interchangeably, with native Slack and web UI.

💬 Smart takes

  • Y Combinator, official announcement: QM is meant to be easy to customize, like other agent frameworks, but useful for a whole company.
  • Skeptic: a harness built for YC's own scrappy, all-in workflows may need serious hardening before a regulated enterprise trusts it with legal or accounting access.

🧭 Where this goes

  1. Likelymore startups and accelerators open-source their internal agent tooling rather than treat it as a moat.
  2. PossibleQM or a fork becomes a default starter kit for non-technical teams running their own agents.
  3. Wild Carda security incident inside a QM-run department forces the project to add enterprise-grade guardrails fast.

🥄 The Spoon Take

Handing away the exact tool that runs your own company only makes sense if you think agent orchestration is infrastructure, not a product. YC is betting the real value sits in what you build on top, not the harness itself. Watch whether other founders start treating their internal AI tooling the same way.

🤔 Pushback

Open-sourcing an internal tool is easy when you're not trying to sell it as a product.

Saturday Aug 1
AGENT MODELISTEDGEMINI

The AI guide power users follow just dropped Google entirely. Ethan Mollick, a Wharton professor, cut Gemini from his practical AI guide. It has no agentic computer-use mode like ChatGPT Work or Claude Cowork.

A year ago the guide was all chat: ChatGPT, Claude, Gemini side by side. Today it's split by which AI can actually use a computer.

Simon Willison, the developer behind Datasette, flagged the shift on his blog. ChatGPT's modes are Work and Codex; Claude's are Cowork and Code. Willison calls the naming 'spectacularly unintuitive' even for people who use both daily.

Gemini Spark, Google's answer, hasn't proven itself yet. Whoever wins the computer-use race owns the workflow, not the chat window.

full brief & sources

⚡ Why this matters

  • Shows where the real competitive battle moved: not chat quality, but who can safely operate a computer for you.
  • Google's absence from Mollick's list is a concrete signal, not vague criticism - Gemini Spark isn't there yet.
  • The naming mess (Work vs Codex vs Cowork vs Code) is a real adoption tax on every team evaluating these tools.

🔍 What happened

  • Ethan Mollick's practical AI guide, updated regularly since 2023, dropped Gemini from its current version.
  • A year ago the guide covered chat models: o3, Claude 4 Opus, Gemini 2.5 Pro.
  • Today it centers on agentic computer-use modes: ChatGPT Work and Codex, Claude Cowork and Code.
  • Simon Willison highlighted the shift on his blog on July 27.
  • Willison notes ChatGPT Work on mobile behaves very differently than Work inside the desktop app.

💬 Smart takes

  • Simon Willison: the mode names 'do not map onto each other in any way that will help you remember them.'
  • Ethan Mollick (via his guide): "Gemini Spark has yet to prove itself."
  • Skeptic: a guide reflects one influential professor's workflow, not confirmed market share data.

🧭 Where this goes

  1. LikelyGoogle ships a more capable Gemini agent mode within the next two quarters to get back on these lists.
  2. Likelymore operator guides converge on the same 'which agent mode' framing over chat comparisons.
  3. Possiblethe naming confusion forces one vendor to simplify its product naming.
  4. Wild Carda third-party standard emerges for describing agent modes across vendors, cutting through the naming mess.

🥄 The Spoon Take

The most useful AI comparison isn't model benchmarks anymore - it's who gets to touch your computer. Google skipping this list entirely, a year after leading model rankings, says more than any chatbot arena score. The keyboard, not the chat box, is now the battleground.

🤔 Pushback

One professor's personal guide isn't a market map - plenty of teams still run Gemini in production for cost, not capability, reasons.

Thursday Jul 30
READS THE CUTMADDEN 27RUN AI

Your Madden games now train the computer players. EA built Madden NFL 27's run game on real player moves, studied frame by frame. CPU backs read the field like skilled humans now.

EA calls it ML Ball Carrier Pathing, made with behavior cloning. The model studies eight variables: blockers, gaps, speed, and spacing.

It watches top athletes, then copies their cutback decisions. Earlier CPU logic followed scripted rules that good athletes could predict. Game Design Director Scott O'Gallaghar calls it just the beginning for football AI.

Millions of people get the update on August 13. That makes it the biggest live testbed behavior-cloned AI has ever had.

full brief & sources

⚡ Why this matters

  • A mass-market game just became a live deployment for learned AI behavior.
  • Behavior cloning replaces scripted rules with patterns copied from real players.
  • Tens of millions of players now train and test the model just by playing.

🔍 What happened

  • EA revealed Madden NFL 27's 99 Club and gameplay changes on July 27-29, 2026.
  • ML Ball Carrier Pathing uses behavior cloning, a supervised machine-learning technique.
  • The model studies defender position, blocker leverage, open lanes, and ball-carrier momentum frame by frame.
  • It learns which move a skilled human made from each game state, then repeats that pattern.
  • A second new feature, Timing-Based Catching, adds an optional skill layer on top of ratings-based catches.
  • Madden NFL 27 launches worldwide on August 13, with early access from August 6.

💬 Smart takes

  • Scott O'Gallaghar, EA Senior Game Design Director: behavior cloning is "just the beginning of where we think this technology can go for football gameplay."
  • Skeptic: AI that copies skilled humans can also copy their exploits, and CPU runners that get too good could break single-player difficulty balance.

🧭 Where this goes

  1. LikelyEA expands behavior-cloned AI to defensive players and quarterbacks in future editions.
  2. Likelycompetitive players find and exploit new patterns in the learned running behavior.
  3. Possibleother sports franchises adopt behavior cloning for their own game AI.
  4. Wild Cardin-game player data becomes a bigger part of how EA tunes AI than internal playtesting.

🥄 The Spoon Take

Madden just turned tens of millions of couches into a training gym for its own AI. Every skilled cutback a player makes becomes a lesson the CPU learns. That is a bigger live dataset than most research labs ever get.

🤔 Pushback

Behavior cloning learns from good players, but it can just as easily learn their bad habits or exploits if the training data isn't filtered.

Sunday Jul 26
OPENAI

OpenAI wants your neighborhood shop running on its tools. The company launched a small business program with training and partner integrations. It already counts 10 million active users on its agent products.

ChatGPT Work is OpenAI's agent mode for multi-step business tasks.

It connects to Slack, Gmail, Drive, and Salesforce through a plugins directory.

Named partners include Dropbox, Shopify, Intuit, Slack, Atlassian, and Wix.

The push comes as OpenAI shifts focus toward paying business customers.

Anthropic's enterprise wins have put real pressure on OpenAI's roadmap.

That adoption figure is real, not just a launch-day claim.

Whether shop owners keep paying once the free onboarding ends is unclear.

full brief & sources

⚡ Why this matters

  • OpenAI is chasing durable business revenue, not just consumer subscriptions.
  • 10 million Work and Codex users is a real number, not a launch-day headline.
  • Small business owners are the least technical AI buyer OpenAI has targeted yet.

🔍 What happened

  • OpenAI announced the small business program on July 21, 2026.
  • The program bundles webinars, in-person AI Academy events, and guides.
  • Named integration partners include Dropbox, Shopify, Intuit, Slack, Atlassian, and Wix.
  • ChatGPT Work runs on GPT-5.6 and stays on multi-step tasks for hours.
  • OpenAI says 10 million people now use ChatGPT Work and Codex combined.

💬 Smart takes

  • OpenAI: frames the push as giving small business owners the same AI tools large companies already have.
  • 9to5Mac: reads the 10 million user number as OpenAI proving Work isn't just a demo.
  • Skeptic: money-losing AI products training small businesses to depend on them is a shaky foundation if pricing changes later.

🧭 Where this goes

  1. LikelyAnthropic and Google answer with their own small-business bundles within 2-3 months.
  2. Likelynamed partners like Shopify report a measurable ChatGPT-driven usage bump by Q4.
  3. Possiblethe free training tier narrows once OpenAI needs the program to turn a profit.
  4. Wild Carda partner integration, Intuit or Shopify, becomes the default way small businesses touch AI at all.

🥄 The Spoon Take

OpenAI isn't chasing hype here, it's chasing retention. Ten million Work and Codex users is a real base to defend, and small business owners are sticky customers once their routines run through an agent. This is OpenAI building a moat out of habit, not model quality.

🤔 Pushback

Free training programs are cheap to run and don't prove businesses will pay full price later.

Friday Jul 17
11 TOK/SEC27B PARAMS3.9GB

AI models are shrinking small enough for your phone. PrismML shrank a 27B model to 3.9GB, running on an iPhone. It hits 11 tokens a second, and Apple is testing it, per CNBC.

Most on-device AI models are toy-sized compared to cloud models. Bonsai 27B breaks that pattern with real reasoning power.

It handles multi-step reasoning, tool calls, and images, not just text. The 1-bit version keeps 90% of full-precision accuracy. Weights are free under Apache 2.0, so any developer can ship it today.

PrismML's CEO says Apple and others are already testing it for speed and battery drain. If a phone-sized model holds most of full power, cloud inference bills start looking optional.

full brief & sources

⚡ Why this matters

  • Cloud inference is the biggest cost line for AI products. A model this small kills that cost for many use cases.
  • It proves compression, not bigger GPUs, can close the capability gap for on-device AI.

🔍 What happened

  • PrismML released Bonsai 27B on July 14, 2026, compressed from a Qwen3.6 27B base.
  • The 1-bit variant is 3.9GB and runs on an iPhone 17 Pro at 11 tokens a second.
  • A larger 5.9GB ternary variant targets laptops and keeps 95% of full-precision performance.
  • It supports text, images, tool calls, and multi-step agentic tasks.
  • Weights ship under Apache 2.0, free for commercial use, via MLX on Apple devices and CUDA on NVIDIA GPUs.
  • PrismML CEO Babak Hassibi told CNBC that Apple and other companies are testing the compression for speed and power draw.

💬 Smart takes

  • Babak Hassibi (PrismML CEO): confirmed Apple and other companies are testing the model for speed, power draw, and performance.
  • Skeptic: a 1-bit model still loses real accuracy, and hard agentic tasks may expose that gap fast.

🧭 Where this goes

  1. Likelymore labs race to ship sub-4GB models as the phone becomes a real inference target.
  2. LikelyApple evaluates the technique for a future on-device Apple Intelligence upgrade.
  3. Possibleenterprises start offloading simple agent tasks to phones to cut cloud inference bills.
  4. Wild CardApple licenses or acquires PrismML's compression tech within 12 months.

🥄 The Spoon Take

Cloud AI margins depend on inference being expensive. A free 3.9GB model that runs an agent on a phone chips at that math directly. It isn't the smartest model out there, but it's smart enough for a lot of real work, and it costs nothing to run.

🤔 Pushback

The accuracy loss is real, and the tasks that need it are exactly what enterprises pay cloud prices for.

Tuesday Jul 14
NOT MINEMETA

Image AI just learned to fact-check itself before it draws. Meta's Muse Image searches and codes before rendering a picture. Users are already pushing back over Meta training it on their photos.

Most image generators map a prompt straight to pixels. Muse Image stops mid-generation to search the web and run code first.

That makes it the second Meta Superintelligence Labs release, after April's Muse Spark language model. It ranks No. 2 on Arena's image leaderboard, just behind OpenAI. It's free inside Meta AI, WhatsApp, and Instagram Stories today.

Power users need one of Meta's new subscription tiers for heavy use. Some users are already asking why their own photos train someone else's model.

full brief & sources

⚡ Why this matters

  • Image AI moves from static prompt-to-pixel to agentic, tool-using generation.
  • Meta ships its second Superintelligence Labs model in three months.
  • A privacy backlash over personal-photo training data breaks out within hours.

🔍 What happened

  • Meta launched Muse Image on July 7, its first in-house image generator.
  • The model uses web search and code execution mid-generation, not just prompt mapping.
  • It composes from multiple reference images and edits with precision, Meta says.
  • Free access ships inside Meta AI, WhatsApp DMs, and Instagram Stories.
  • Power users need one of Meta's new monthly subscription plans for heavy use.
  • It ranks No. 2 on the Arena text-to-image leaderboard, behind OpenAI.

💬 Smart takes

  • TechCrunch: users are already pushing back over Meta's use of their photos to train the model.
  • Axios: Muse Image is Meta's second Superintelligence Labs release, after Muse Spark in April.
  • Skeptic: agentic tool-use adds latency and cost per image, a tradeoff casual selfie-editors may not want.

🧭 Where this goes

  1. LikelyMeta folds agentic image tools into Advantage Plus ad creative within one quarter.
  2. Likelyrival image models add search or code steps to match the accuracy claim.
  3. Possiblethe photo-training backlash forces Meta to add an explicit opt-out toggle.
  4. Wild Carda regulator opens an inquiry into Meta's personal-photo training practice within 90 days.

🥄 The Spoon Take

Image generation just got a research step. Muse Image doesn't guess what a chair looks like, it can look one up first. That's a real capability jump, but it also means Meta is quietly widening what counts as training data from your camera roll.

🤔 Pushback

Agentic tool-use inside an image model sounds impressive but mostly matters for edge cases; most users just want a fast, cheap edit.

Thursday Jul 9
BEFOREAFTER

Your camera roll just got an editor built in. Google Photos now offers Video Remix, built on the Gemini Omni model. It rolls out free today to every Google AI Plus, Pro, and Ultra subscriber.

The feature lives in the Create tab, with a library of ready templates. Ask for a watercolor look, morning light, or a new background, and it renders in seconds.

Gemini Omni is trained to understand gravity and light, not just pixels. That is what makes an edit look real instead of pasted on. You can even drop a digital double of yourself into a clip, watermarked by SynthID.

The same model already powers free remixes inside YouTube Shorts, no subscription needed. Adobe and Canva now have a new AI rival to answer.

full brief & sources

⚡ Why this matters

  • Video editing has always required either skill or Premiere Pro tutorials.
  • Gemini Omni is Google's bet that AI video understands physics, not just pixels.
  • Free rollout to Shorts means hundreds of millions see this immediately.

🔍 What happened

  • Google announced Video Remix inside Google Photos on July 8, 2026.
  • It runs on Gemini Omni, first shown in May as a video-first model.
  • Templates handle style transfer: watercolor filters, relighting, background swaps.
  • Processing takes a few seconds per clip, according to Google.
  • It ships free today to Google AI Plus, Pro, and Ultra subscribers in the US and select countries.
  • The same Gemini Omni engine already powers free remixes in YouTube Shorts and Google Flow.

💬 Smart takes

  • Google: Gemini Omni can 'create anything from any input,' starting with video.
  • Engadget: Video Remix is 'designed to save you from sitting through hours of Premiere Pro tutorials.'
  • Skeptic: template-based edits cap creative control; power users will still open a real editor.

🧭 Where this goes

  1. LikelyVideo Remix expands to more countries and languages within a few months.
  2. LikelyAdobe and Canva add competing one-prompt video restyle tools within the year.
  3. PossibleGemini Omni becomes the default video layer across Google Photos, YouTube, and Workspace.
  4. Wild CardSynthID-watermarked avatars become a new short-form ad format brands pay to use.

🥄 The Spoon Take

Google just turned video editing into a prompt. That is a bigger deal than another filter app. Every past AI editor worked on top of your footage. Gemini Omni claims to understand the physics inside it, which is the harder problem to fake.

🤔 Pushback

Google's own examples are fairly subtle by its telling, and physics-aware claims from labs rarely hold up outside the demo reel.

Monday Jul 6
1 MODEL

Voice agents used to need three stitched models. xAI now turns plain text into a live phone agent in two minutes. It beat Google and OpenAI on the speech benchmark using one model, not three.

Under the hood: live call handling, real-time lookups, and safety checks ship together already, no extra integration work.

It answers in under 700 milliseconds, faster than pipelines that turn sound into words and back. It also keeps laughs, whispers, and sighs that a written middle step would erase.

Pricing lands at five cents per minute, with cloning and eighty-plus voices bundled in. Access is still limited, and some builders are hitting errors.

full brief & sources

⚡ Why this matters

  • Voice UX has been the weak link in AI agents. Latency and stitched pipelines make them feel robotic.
  • One speech-to-speech model removes two translation steps, which is usually where the lag comes from.
  • The pricing undercuts building an in-house voice stack from scratch.

🔍 What happened

  • xAI launched Voice Agent Builder in beta on July 1, 2026.
  • No-code: describe a phone call in plain language, get a live agent in under two minutes.
  • Runs on one speech-to-speech model instead of three stitched APIs.
  • Responds in under 700 milliseconds. Scored 67.3% on the tau-voice Bench, beating Gemini 3.1 Flash Live at 43.8% and GPT Realtime 1.5 at 35.3%.
  • $0.05 per minute. 80-plus voices, voice cloning from two minutes of audio, 25-plus languages with mid-call switching.

💬 Smart takes

  • eesel AI review: developers who built with it are impressed by mid-conversation language switching and how fast a working agent comes together.
  • Skeptic: it's still beta. Several developers hit 403 access errors, and no one has published a fix for an agent acting on a misheard instruction.

🧭 Where this goes

  1. LikelyxAI opens broader access within weeks once the access errors get sorted.
  2. Likelyrival labs answer with their own single-model voice stack within a quarter.
  3. Possibleagencies use it to replace call-center phone trees, not just simple bots.
  4. Wild Cardthis becomes the default way small businesses build a phone line within a year.

🥄 The Spoon Take

The voice-agent race has been about who has the smartest model. xAI just made it about who has the fastest one. Sub-second, single-model voice is the unlock that finally makes AI phone agents feel less like a phone tree.

🤔 Pushback

A benchmark win on one leaderboard doesn't mean it holds up on messy real calls with accents, noise, and interruptions.