Cybersecurity & Privacy

Contractors Are Reading Real ChatGPT Conversations, Report Finds

OpenAI pays hundreds of contractors to read real ChatGPT conversations and rate the chatbot’s replies, a practice most users likely don’t know exists.

By Marcus Lee Edited by Maria Konash Published: Updated:
Contractors Are Reading Real ChatGPT Conversations, Report Finds

Key Notes

  • OpenAI runs "Project Lily," paying hundreds of contractors (via recruiter Crossing Hurdles and payer Mercor, $50+/hour) to read real ChatGPT conversations and rate responses on a scale, aimed at reducing anthropomorphizing and sycophancy.
  • Reviewers don't see usernames, but do see a "user memories summary" that can reveal prior usage and rough location.
  • Anthropic confirmed it also uses human review to improve Claude, but only for users who've opted in via the "Help improve our AI models" setting, removing account identifiers first — a more clearly opt-in structure than OpenAI's practice, which applies by default on consumer plans.

OpenAI is paying hundreds of contractors to read real ChatGPT conversations and rate the chatbot’s responses, according to a 404 Media investigation based on leaked internal documents and real prompts reviewed by the outlet.

The internal codename for the effort is Project Lily. Contractors are recruited through a staffing firm called Crossing Hurdles and paid via Mercor, an AI-training company; one worker told 404 Media the pay exceeds $50 an hour.

The reviewers see whole conversations, not just isolated prompts, and rate ChatGPT’s replies to help OpenAI improve response quality. Internal documents show contractors specifically training the model to stop anthropomorphizing itself and to reduce sycophancy, excessive flattery or validation of a user’s stated views regardless of accuracy.

Sycophancy became a serious concern for OpenAI after its GPT-4o model’s tendency toward this behavior was cited as a contributing factor in multiple wrongful-death lawsuits involving user suicides, according to the suits.

OpenAI says reviewers cannot see usernames. But the dashboard contractors use reportedly includes a “user memories summary” above each prompt, which can reveal what a person has previously used ChatGPT for and, in some cases, roughly where they live.

The company said conversations pass through an automated Privacy Filter meant to strip personal details before reaching human reviewers. OpenAI’s own documentation for that system acknowledges it can miss uncommon identifiers or fail to adequately redact information when context is limited, meaning some sensitive material likely does reach contractors despite the safeguard.

When 404 Media asked OpenAI where it discloses to users that humans may read their chats, the company did not directly answer, and after publication pointed only to a general help page describing human review in the context of content moderation and model improvement.

How Anthropic’s Practice Differs

Anthropic confirmed to 404 Media that it also uses human review to improve its Claude models. The company said this applies specifically to users who have opted in through a “Help improve our AI models” setting, and that it removes account identifiers from conversations before human reviewers see them.

That distinction matters: Anthropic’s process is described as opt-in, requiring users to actively enable data sharing, while OpenAI’s Project Lily reviews conversations from ChatGPT’s consumer plans, where model-training use is turned on by default rather than requiring explicit opt-in. Google’s Gemini reportedly uses a similar human-review process as well, with its own distinct opt-in defaults and disclosure language.

Why the Disclosure Gap Matters

The core tension the report surfaces isn’t that human review happens, since most major AI companies use some form of it to refine model behavior, but whether users meaningfully understand that it’s happening to their specific, personal conversations.

One person who works with the prompts told 404 Media plainly that they do not believe ChatGPT users know humans are reading their chats. Many of the prompts reviewers see reportedly include users explicitly asking ChatGPT to keep the conversation’s contents confidential, a request the human-review pipeline does not honor.

With more than 900 million weekly users, ChatGPT has become a place people turn to for deeply personal conversations, from mental health struggles to relationship troubles to medical questions, often without realizing that a real person, not just an algorithm, may eventually read what they wrote.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, Cybersecurity & Privacy, News