Legal Battle Intensifies: Reddit’s AI Scraping Case Against Anthropic Advances
The ongoing legal friction between Reddit and AI developer Anthropic has reached a significant turning point. A San Francisco Superior Court judge has issued a ruling that allows the majority of Reddit’s lawsuit-which alleges that the company’s Claude chatbot was trained using unauthorized scrapes of user-generated content-to move toward discovery.
Court Ruling Narrows the Scope of Litigation
In a formal decision issued on September 17 and officially recorded the following day, Judge Harold Kahn addressed a demurrer filed by Anthropic. The AI firm had attempted to have the case dismissed entirely, arguing that Reddit’s grievances failed to establish a viable legal foundation.
While Judge Kahn did not grant a total victory to either side, he effectively pruned the lawsuit. The court dismissed the claims regarding “unjust enrichment” and “trespass to chattels.” However, the core of Reddit’s argument remains intact, as the judge ruled that the following three claims possess enough merit to proceed:
* Breach of Contract: Allegations that Anthropic violated the terms governing how data can be harvested from the platform.
* Interference with Contract: Claims that the scraping activities disrupted Reddit’s existing business relationships and data usage agreements.
* Violation of California’s Unfair Competition Law: Assertions that Anthropic’s methods of data acquisition constitute predatory or deceptive business practices under state law.
The Broader Context of AI Data Training
This case is part of a growing wave of litigation as content platforms grapple with the rapid expansion of Large Language Models (LLMs). As of late 2024, industry analysts note that the “data gold rush” has led to a surge in copyright and terms-of-service disputes. Much like the high-profile legal battles between news publishers and AI labs, this conflict highlights the tension between the open web and the proprietary needs of generative AI companies.
By allowing these specific claims to move forward, the court is signaling that the methods used to “feed” AI models are subject to rigorous legal scrutiny. As the case progresses, it may set a critical precedent for how AI companies interact with user-generated content, potentially forcing a shift toward more transparent licensing models for training data.
