Thursday Jun 11

Nvidia Ships America's Best Open Model

11JUN
NEMOTRON 3

Nvidia released the strongest open-weight model from a US lab. Nemotron 3 Ultra has 550 billion parameters and is built to run agents for hundreds of steps. Jensen Huang unveiled it at Computex.

Nemotron 3 Ultra scores highest among US open models on the Artificial Analysis index. It beats prior open releases on reasoning.

It uses a hybrid Mamba-Transformer design with 55 billion active parameters per token. Context runs to 1 million tokens. It burns fewer tokens than rival open models on long agent runs.

Early adopters include Cursor, Perplexity, ServiceNow, and Deloitte. This is a real US answer to China's DeepSeek and Qwen.

full brief & sources

Why this matters

  • First US open model to clearly top the open-weight leaderboard in 2026.
  • Tuned for agents: plans, calls tools, recovers from errors across hundreds of turns.
  • Direct challenge to Chinese open models that led on cost and openness.

🔍 What happened

  • Announced at Computex June 1; full release June 4.
  • 550B mixture-of-experts, roughly 55B active per token.
  • Hybrid Mamba-Transformer; up to 1M token context.
  • Scores 48 on the Artificial Analysis Intelligence Index.
  • 300-plus tokens per second on BF16; 5x faster with NVFP4 on Blackwell.
  • Adopters: Accenture, CrowdStrike, Cursor, Perplexity, ServiceNow, Siemens, Zoom.

💬 Smart takes

  • Artificial Analysis: the most capable open model from a US lab to date.
  • Jensen Huang, Nvidia CEO: built to orchestrate agents that plan, delegate, and recover.
  • Skeptic: open weights from a chipmaker also sell more Nvidia GPUs to run them.

🧭 Where this goes

  1. LikelyUS enterprises wary of Chinese models adopt Nemotron for agents.
  2. LikelyNvidia keeps shipping open models to drive GPU demand.
  3. Possible'tokens per task' becomes the key agent cost metric.
  4. Wild Cardan open model tops a closed frontier model on agent benchmarks within a year.

🥄 The Spoon Take

Nvidia is not just selling shovels anymore. It is handing out a best-in-class open model that happens to run fastest on its own chips. Open weights win developer trust. The agent focus wins the next workload. Every token it saves is a token that still bills on a Blackwell.

🤔 Pushback

Benchmark wins fade fast, and a chipmaker's open model is a marketing engine for GPUs as much as a research milestone.