Thursday Sep 24

OpenAI's Raters Got Fired For Using AI

24SEP
10,000 RATERSHIDDEN AITHE RATER

Contractors grading ChatGPT answers were dismissed after vendors caught them leaning on language models and Grammarly. The tell was em dashes and speed. Human judgment is the input nobody can fake.

404 Media's Joseph Cox reports multiple workers on OpenAI rating projects lost their gigs. The projects run through firms like Mercor and span 10,000 people. Project Lily has hundreds scoring real chats for sycophancy.

An internal guide tells reviewers to spot repetitive words, quick completions and dashes, and warns: do not tell evaluators why you suspect AI. Mercor says its contracts ban LLMs and it enforces that.

Meanwhile the labeling business is booming. Snorkel AI raised $350 million at $3.5 billion with ARR up 18x. Micro1 is worth $4 billion. The product they sell is unautomated human opinion.

full brief & sources

⚡ Why this matters

  • The frontier labs are paying a premium for one thing: judgment that did not come from a model. When the graders use models, the signal collapses into the thing it was meant to correct.
  • This is the model-collapse problem showing up as an HR policy. Training on your own output looks like progress until it does not.
  • Ten thousand contractors is a workforce. The rules they work under will set the template for every AI evaluation job.

🔍 What happened

  • 404 Media reported on September 22 that several contractors rating ChatGPT responses were fired for using AI tools, including LLMs, GPTZero, Grammarly and AI translation.
  • The rating programs span more than 10,000 contractors through vendors such as Mercor. Project Lily assigns hundreds of people to read real user conversations and score responses from 1 to 7 on sycophancy and anthropomorphizing.
  • An internal document instructs reviewers not to use AI detection tools or AI themselves, and not to tell evaluators why they are suspected. Red flags listed: repetitive wording, em dashes, and completing tasks too fast.
  • One contractor told 404 Media they had deliberately picked the worst outputs as a form of sabotage. Mercor said its contracts strictly prohibit LLM use and it enforces that. OpenAI declined to comment.
  • Separately, Snorkel AI announced a $350 million Series E at a $3.5 billion valuation led by Insight Partners and S32, with ARR up 18x to $375 million on the back of expert data services.

💬 Smart takes

  • Mercor spokesperson: "Our contracts strictly prohibit the use of LLMs to complete projects and we enforce that." The vendor is the enforcement layer, not OpenAI.
  • Joseph Cox, 404 Media: the people training the AI were fired for using the AI. The irony is the story, but the mechanism is the lesson: the labs can detect their own fingerprints.
  • Skeptic: firing gig workers over a grammar checker is a labor story as much as a data story. If the pay assumed AI-speed throughput, the incentive to cheat was built in.

🧭 Where this goes

  1. Likelyrating vendors add keystroke and screen monitoring, and the rate cards rise to compensate.
  2. Possiblea fired contractor sues over the no-explanation dismissal policy, and the internal guidance becomes an exhibit.
  3. Wild Carda lab publishes a study showing how much AI-assisted ratings degraded a model, and the whole industry reprices human data.

🥄 The Spoon Take

Here is the tell: the labs can detect AI writing well enough to fire people for it, but cannot use AI to grade AI. That asymmetry is the market. Snorkel's 18x ARR is the price of verified human judgment. If your product depends on evaluation data, budget for humans and for policing them. Both costs just went up.

🤔 Pushback

This rests on one outlet's reporting and anonymous workers. OpenAI has not confirmed the firings or the scale.