OpenAI has hired hundreds of human contractors under an internal initiative dubbed ‘Project Lily’ to review and rate a continuous stream of real user ChatGPT prompts, according to internal documents obtained by 404 Media. The reviewers evaluate chatbot outputs to reduce sycophancy, tone down anthropomorphism, and refine response accuracy.

Although OpenAI strips usernames and attempts to scrub personally identifiable data prior to review, sensitive personal, medical, and professional details shared by users regularly bypass filters. The revelation highlights privacy risks for consumer users who input confidential information into conversational AI systems.

Competitors like Anthropic confirmed they employ similar human review methods to refine model behavior. The practice underlines the continuous dependence on human feedback networks (RLHF) to optimize frontier models beyond automated web scraping.

Why it matters

  • Enterprise operators must enforce strict data-handling policies to prevent employee prompts containing sensitive IP from being reviewed by third-party human contractors.

  • Consumer AI platforms face potential regulatory scrutiny and public backlash regarding data scrubbing practices and consent transparency.

  • Highlights the ongoing necessity of human-in-the-loop data annotation for maintaining model safety and conversational quality.

Source: 404media.co