Tech companies are now hiring people to pose as teenagers to intentionally bait AI chatbots into saying dangerous things.
Meta recently employed hundreds of contractors to interact with rival AI services like ChatGPT and Gemini while pretending to be underage. These workers were instructed to ask the chatbots questions about high-risk topics, including self-harm, sexual content, and illegal drugs. The goal of this project was to determine how effectively other AI services defend against potentially harmful prompts when interacting with what appears to be a child. By posing as minors, these contractors could pressure the AI to see if its automated safety rules—the internal filters designed to block harmful output—would buckle under the pressure of a vulnerable user profile.
The cat-and-mouse game of AI safety
At the core of every modern AI assistant is a set of training data—a massive digital library of human text used to teach the system how to communicate. During the development phase, companies apply a layer of restrictions known as safety guidelines to prevent the AI from generating harmful or inappropriate content. Think of these guidelines as a digital librarian who has been instructed to refuse requests for forbidden books. However, because these systems are so complex, it is nearly impossible to predict every way a user might phrase a request to bypass those rules. Testing these systems often involves a process called 'red teaming,' where testers act as adversarial users trying to trick the AI into breaking its own rules, essentially looking for cracks in the fence to see what they can squeeze through before a real user does.
This strategy shows just how difficult it is for tech companies to guarantee that their tools are safe for younger users. When a company tests its own AI, it has access to the internal blueprints; when it tests a competitor's AI, it is blindly feeling for weaknesses from the outside, much like a security auditor trying to pick a lock. This approach highlights an uncomfortable reality: developers still don't fully understand how to perfectly control the AI's behavior. As these tools become more integrated into our lives, the pressure to ensure they don't provide harmful advice to teenagers is leading to even more aggressive—and sometimes strange—testing tactics.
Liked this one? The next lands at breakfast.
Every story in tomorrow's AI news, rebuilt in plain English — five minutes, sources linked, free forever.
By joining you agree to receive Article's daily newsletter — unsubscribe in one click. Privacy