Wednesday Jul 8

Meta Claims Watermelon Matches GPT-5.5

2JUL
METAGPT-5.5

Meta says its next model finally caught up to OpenAI's best. Alexandr Wang, Meta's AI chief, reportedly told staff Watermelon matches GPT-5.5. No named benchmarks, no outside check yet.

Watermelon is the successor to Avocado, the model behind Muse Spark.

It reportedly uses 10 times more compute than its predecessor.

Wang made the claim at an internal town hall, not a launch.

He didn't name which benchmarks Watermelon supposedly matches.

Meta also teased a faster Muse Spark coding update coming soon.

This is the fourth Meta AI claim this year without independent verification.

If true, it ends OpenAI's run as the undisputed benchmark leader.

full brief & sources

Why this matters

  • An unverified benchmark claim from a lab chief can move competitive narratives overnight.
  • Meta has a track record of teasing capability before shipping it.
  • If real, it changes who enterprises trust for their next model migration.

🔍 What happened

  • Alexandr Wang told an internal Meta town hall on July 2 that Watermelon matches GPT-5.5.
  • Watermelon is Meta's next flagship model, successor to Avocado (Muse Spark's codename).
  • Wang said Watermelon uses roughly 10x the compute of its predecessor.
  • No specific benchmark names or scores were disclosed.
  • Business Insider sourced the claim from people familiar with the town hall.
  • OpenAI has not commented on the claim.

💬 Smart takes

  • Business Insider: the claim came from 'one executive, in one internal room, on one round of unnamed benchmarks.'
  • Alexandr Wang: a Muse Spark coding and agent update is coming 'pretty soon.'
  • Skeptic: Meta has claimed benchmark parity before and shipped models that fell short in independent testing.

🧭 Where this goes

  1. LikelyMeta publishes at least partial Watermelon benchmarks within the next 60 days.
  2. Possibleindependent evaluators test Watermelon and find a smaller gap than claimed.
  3. PossibleOpenAI responds with its own updated GPT-5.5 benchmark refresh.
  4. Wild CardWatermelon ships and actually beats GPT-5.5 on reasoning, not just matches it.

🥄 The Spoon Take

An internal town hall claim isn't a launch, but it's a signal Meta wants told. Every 'we caught up' leak before an actual release buys goodwill cheaply. The real test is the benchmark table Meta eventually has to publish.

🤔 Pushback

This is a single executive's claim in a private room, with no named benchmarks and no outside verification. Meta has made similar catch-up claims before that didn't hold up once independent labs ran the numbers.