A high-stakes legal battle is heating up over what exactly goes into the digital brain of AI. The New York Times and The Daily News are now accusing OpenAI of hiding evidence that could prove the company knew exactly how much copyrighted news was being used to build its products.
The newspapers are suing OpenAI, claiming the company broke copyright law by feeding their journalism into the systems that power ChatGPT. For a long time, OpenAI told the court that it was technically too difficult to search its massive collection of training data and chat logs to see if anyone’s work was inside. They also argued that digging through these private conversations would violate user privacy. But in a recent deposition, an internal engineer suggested that OpenAI actually had the ability to search this information all along. The newspapers claim OpenAI has already created internal tools to track when its AI regurgitates copyrighted material, specifically mentioning a project called Project Giraffe and a tool known as a Bloom filter, both designed to flag when the AI repeats exact text from its training sources. They also allege that OpenAI deleted or improperly altered records to hide these capabilities from the court.
The mystery of the training dataset
Think of the model as a student that learned everything it knows by reading millions of books and articles, a process called training. The training dataset is the massive stack of library materials this student was given to study. When you ask ChatGPT a question, it is not looking up current information in a database. Instead, it is predicting the next word based on patterns it memorized during that training phase. The lawsuit asks whether those study materials included copyrighted journalism and if, as a result, the AI is just regurgitating the original articles rather than creating something new. The Bloom filter mentioned in the court documents acts like a digital monitor that watches the model’s outputs to see if it is accidentally spitting out the exact copyrighted content it studied, helping the company measure where its training might have leaned too heavily on someone else's work.
This conflict forces us to ask how much control companies really have over the black boxes they build. If a company claims it cannot track what its AI is learning, but internal projects like Giraffe suggest they have developed ways to monitor exactly that, it weakens both their legal defense and public trust. The outcome here will likely set a major precedent for whether AI companies must be transparent about their sources, or if their training methods will remain shielded from public view. We are essentially watching a fight over the rules of the road for the next generation of information technology, where the boundary between original reporting and automated output is being defined in real-time.
Liked this one? The next lands at breakfast.
Every story in tomorrow's AI news, rebuilt in plain English — five minutes, sources linked, free forever.
By joining you agree to receive Article's daily newsletter — unsubscribe in one click. Privacy