Monday, August 3, 2026

Policy & Regulation

NYT accuses OpenAI of hiding evidence in ChatGPT copyright suit

The New York Times and The Daily News have accused OpenAI of hiding evidence and deleting billions of ChatGPT outputs in their ongoing copyright lawsuit.

The New York Times and The Daily News have asked a federal judge to discipline OpenAI, alleging the company has been lying about its ability to search customer chat log data and training datasets for their copyrighted works. It’s the latest escalation in a two-year lawsuit accusing OpenAI of violating copyright law by training its generative AI models on the Times’ content and reproducing that journalism in user outputs. The plaintiffs also claim OpenAI deleted billions of ChatGPT outputs after they filed suit in direct violation of the court’s preservation order, and substituted millions of logs in the requested sample — conduct they say amounts to withholding evidence and messing with the discovery process (the pre-trial phase in which parties exchange evidence). Crosby, lead counsel for the plaintiffs, argued that if OpenAI genuinely believed its use of the outlets’ journalism was fair and legal, it would not have concealed having done it.

An April court-ordered deposition of OpenAI data privacy engineer Vinnie Monaco allegedly revealed that OpenAI had already conducted internal searches and evaluations of its training corpus to look for copyrighted journalism works. Monaco’s testimony also allegedly showed that, before the lawsuit was even filed, OpenAI had amassed a database of about 78 million de-identified ChatGPT conversations, which it used internally to gauge how much it was infringing on others’ works.

Monaco’s deposition further allegedly revealed that OpenAI implemented a Bloom filter (a data structure used to detect matching text) as part of an internal set of tools called Project Giraffe, which detected and recorded regurgitation — the reproduction of training text in outputs — shortly after the lawsuit was filed. The plaintiffs say this undercuts OpenAI’s argument that searching its logs was too technically burdensome, especially since OpenAI negotiated a requested sample of 120 million chat logs down to just 20 million, which it submitted to the courts last December, and which the plaintiffs say included millions of substituted logs.

OpenAI spokesperson Drew Pusateri denied the allegations. “As the Times’ case weakens and they’ve been forced to drop claims against us, they’re persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations,” Pusateri said. He added that OpenAI will continue defending its users’ privacy and the long-established principles of fair use.

Why it matters

The dispute highlights the growing legal friction over AI training transparency, as publishers try to prove copyright infringement while AI developers guard their training data and user privacy.