← The Vault
The Big Story

Is OpenAI hiding how it trains its AI?

The New York Times has accused OpenAI of misleading the court about its ability to search the data used to teach ChatGPT. The legal battle centers on whether OpenAI used copyrighted news articles to train its systems and whether it has been honest about its internal efforts to track that usage. This dispute highlights a central challenge in AI: identifying exactly what information an AI company has fed into its systems behind the scenes.

Edition № 192Room: The Big Story10 July 20262 min readSources: 1
Article

A high-stakes legal battle is heating up over what exactly goes into the digital brain of AI. The New York Times and The Daily News are now accusing OpenAI of hiding evidence that could prove the company knew exactly how much copyrighted news was being used to build its products.

WHAT'S HAPPENING

The newspapers are suing OpenAI, claiming the company broke copyright law by feeding their journalism into the systems that power ChatGPT. For a long time, OpenAI told the court that it was technically too difficult to search its massive collection of training data and chat logs to see if anyone’s work was inside. They also argued that digging through these private conversations would violate user privacy. But in a recent deposition, an internal engineer suggested that OpenAI actually had the ability to search this information all along. The newspapers claim OpenAI has already created internal tools to track when its AI regurgitates copyrighted material, specifically mentioning a project called Project Giraffe and a tool known as a Bloom filter, both designed to flag when the AI repeats exact text from its training sources. They also allege that OpenAI deleted or improperly altered records to hide these capabilities from the court.

The mystery of the training dataset

HOW IT WORKS

Think of the model as a student that learned everything it knows by reading millions of books and articles, a process called training. The training dataset is the massive stack of library materials this student was given to study. When you ask ChatGPT a question, it is not looking up current information in a database. Instead, it is predicting the next word based on patterns it memorized during that training phase. The lawsuit asks whether those study materials included copyrighted journalism and if, as a result, the AI is just regurgitating the original articles rather than creating something new. The Bloom filter mentioned in the court documents acts like a digital monitor that watches the model’s outputs to see if it is accidentally spitting out the exact copyrighted content it studied, helping the company measure where its training might have leaned too heavily on someone else's work.

WHY IT MATTERS

This conflict forces us to ask how much control companies really have over the black boxes they build. If a company claims it cannot track what its AI is learning, but internal projects like Giraffe suggest they have developed ways to monitor exactly that, it weakens both their legal defense and public trust. The outcome here will likely set a major precedent for whether AI companies must be transparent about their sources, or if their training methods will remain shielded from public view. We are essentially watching a fight over the rules of the road for the next generation of information technology, where the boundary between original reporting and automated output is being defined in real-time.

Sources
← PreviousIs AI ready to write our laws and make our home-tech?Next →Why AI is struggling to make sense of your spreadsheets
Tomorrow's edition · free

Liked this one? The next lands at breakfast.

Every story in tomorrow's AI news, rebuilt in plain English — five minutes, sources linked, free forever.

By joining you agree to receive Article's daily newsletter — unsubscribe in one click. Privacy

← Back to the Vault