Unredacted NYT v. OpenAI Filings: 91,692 Article Copies in Mid-Training Data and a 93% Click-Through Collapse
TechCrunch reported on 17 September 2026 that newly unredacted filings in the New York Times copyright suit against OpenAI and Microsoft show OpenAI mid-training datasets containing more than 91,692 copies of NYT, Daily News and investigative journalism works, and internal data putting NYT click-through rates down as much as 93% under Microsoft Copilot versus traditional Bing search. The filings include a January 2023 internal memo from Microsoft's director of Applied Science Brent Hecht calling AI scraping 'the largest theft of labor in human history' and 'an astonishing theft of unprecedented proportions.' Filings also allege employees bypassed paywalls and stripped copyright notices from training data; neither company responded to comment requests.
↳ Follow the thread