Tuesday Sep 22

TypeSafe's Jev Returns Odds, Not Text

22SEP
OUTPUT: FREECHATJEV

A ChatGPT co-inventor shipped a model that never writes words. Jev reads text and returns typed probabilities: yes or no, pick one, score it. It cannot hallucinate. Input costs 4 cents per million tokens.

Diogo Almeida, TypeSafe co-founder and ex-OpenAI, helped invent RLHF. His new model answers only in odds. Trained on synthetic data with a method he calls reinforcement learning from calibrated decisions.

Vercel swapped an OpenAI classifier for Jev and got 5 to 18 times faster with better accuracy. Output is free. Open-weight clones and a JevBench appeared within days.

Simon Willison, independent developer, is uneasy. A black box that ranks things is a bias machine. His line: he really hopes nobody uses Jev to rank job applicants.

full brief & sources

⚡ Why this matters

  • Most production LLM calls are classification in disguise. Jev makes that a product category with its own price point.
  • Almeida is saying the quiet part: frontier labs sell fear or hype, and most of the capability is not useful yet.
  • If typed outputs win the routing and moderation layer, chat models lose their cheapest and highest-volume traffic.

🔍 What happened

  • TypeSafe AI launched Jev on September 15. TechCrunch covered the developer reaction on September 18, Simon Willison wrote it up on September 21.
  • Jev takes text and returns a typed distribution: a yes or no probability, a choice among options, or a score.
  • Pricing: $0.042 per million input tokens, output free. GPT-5 Nano costs $0.05 for input.
  • Vercel's Pranit Sharma replaced an OpenAI Luna 5.6 classifier and reported 5 to 18 times lower latency with higher accuracy.
  • Bryo AI CTO Nikhil Mudholkar found Gemini slightly more accurate but 10 to 20 times more expensive.
  • Community shipped a Qwen 3.5 based clone called Kev and a benchmark called JevBench. The API was briefly overloaded.

💬 Smart takes

  • Diogo Almeida, TypeSafe CEO: "We have lightning in a bottle, and yet it is not useful." He says the main product of frontier labs is fear or hype.
  • Armin Ronacher, Earendil CTO: Jev "delegates the hallucination problem a little bit to the user." He expects competitors to copy the shape.
  • Skeptic, Simon Willison: a probability with no explanation is a regression in debuggability. Calibrated is not the same as fair.

🧭 Where this goes

  1. Likelyevery major lab ships a typed-output or decision model tier within six months.
  2. Likelyrouting, moderation, and ranking calls move off chat models first, because that is where the cost gap is 10x.
  3. PossibleJev-style scores end up in hiring, credit, and content pipelines with no audit trail.
  4. Wild Cardregulators treat opaque decision models as scoring systems under existing credit and employment law.

🥄 The Spoon Take

Chat was the demo. Decisions are the business. Most of what companies pay LLMs for is yes or no, this or that, how likely. Jev just priced that at almost nothing and made it fast. The trap is obvious. A model that can't hallucinate can still be wrong, and nobody can see why.

🤔 Pushback

Jev only wins on tasks you can phrase as a choice. The moment you need a reason, you are back to a chat model.