← The Vault
Big Question

How Companies Test AI Safety by Pretending to Be Kids

To see if AI chatbots are safe for minors, Meta hired contractors to pose as teenagers and bait rival services into discussing sensitive topics like self-harm and drugs. This unusual testing method highlights the ongoing struggle to keep AI guardrails strong. Instead of waiting for users to accidentally misuse their tools, tech companies are actively probing their competitors to find out where the ‘safety filters’ fail.

Edition № 131Room: Big Question29 June 20262 min readSources: 1
Article

Tech companies are now hiring people to pose as teenagers to intentionally bait AI chatbots into saying dangerous things.

WHAT'S HAPPENING

Meta recently employed hundreds of contractors to interact with rival AI services like ChatGPT and Gemini while pretending to be underage. These workers were instructed to ask the chatbots questions about high-risk topics, including self-harm, sexual content, and illegal drugs. The goal of this project was to determine how effectively other AI services defend against potentially harmful prompts when interacting with what appears to be a child. By posing as minors, these contractors could pressure the AI to see if its automated safety rules—the internal filters designed to block harmful output—would buckle under the pressure of a vulnerable user profile.

The cat-and-mouse game of AI safety

HOW IT WORKS

At the core of every modern AI assistant is a set of training data—a massive digital library of human text used to teach the system how to communicate. During the development phase, companies apply a layer of restrictions known as safety guidelines to prevent the AI from generating harmful or inappropriate content. Think of these guidelines as a digital librarian who has been instructed to refuse requests for forbidden books. However, because these systems are so complex, it is nearly impossible to predict every way a user might phrase a request to bypass those rules. Testing these systems often involves a process called 'red teaming,' where testers act as adversarial users trying to trick the AI into breaking its own rules, essentially looking for cracks in the fence to see what they can squeeze through before a real user does.

WHY IT MATTERS

This strategy shows just how difficult it is for tech companies to guarantee that their tools are safe for younger users. When a company tests its own AI, it has access to the internal blueprints; when it tests a competitor's AI, it is blindly feeling for weaknesses from the outside, much like a security auditor trying to pick a lock. This approach highlights an uncomfortable reality: developers still don't fully understand how to perfectly control the AI's behavior. As these tools become more integrated into our lives, the pressure to ensure they don't provide harmful advice to teenagers is leading to even more aggressive—and sometimes strange—testing tactics.

Sources
← PreviousHow Google and California are bringing AI into your daily apps and government servicesNext →Why startups are building their own AI from scratch
Tomorrow's edition · free

Liked this one? The next lands at breakfast.

Every story in tomorrow's AI news, rebuilt in plain English — five minutes, sources linked, free forever.

By joining you agree to receive Article's daily newsletter — unsubscribe in one click. Privacy

← Back to the Vault