Reddit Can Proceed With Scraping Claims Against Perplexity, SerpApi

Reddit can proceed with a lawsuit alleging that Perplexity and data scraper SerpApi wrongly obtained copyrighted Reddit posts from Google's search results, a federal judge said Friday.

The ruling, issued by U.S. District Court Judge Paul Engelmayer in New York, comes in a lawsuit filed by Reddit in October, when it claimed that SerpApi and artificial intelligence company Perplexity violated the Digital Millennium Copyright Act's anti-circumvention provisions, which prohibit bypassing technological restrictions on copying.

Reddit's initial complaint included allegations that SerpApi used a workaround to bypass Google's SearchGuard -- an anti-scraping system that relies on CAPTCHAs -- in order to provide Perplexity with Reddit posts.

An amended complaint filed by Reddit in February includes allegations that Perplexity itself used tools provided by SerpApi to obtain Reddit posts.

advertisement

advertisement

"Perplexity is doing more than simply receiving data from SerpApi," Reddit alleged in that amended complaint. "Perplexity circumvents Google’s security protocols and gains access, through unauthorized and automated processes, to Reddit data by using (at least) SerpApi’s tools or directing others to use those tools."

Reddit also alleged that content posted to its site includes posts by users as well as "tens of thousands of posts, comments, and other original works authored by Reddit itself."

For instance, one post allegedly scraped by SerpApi was an "Admin-authored 700-word history of the origin of the subreddit."

Reddit also claimed that it suffered "reputational harm" with regard to privacy. 

Specifically, Reddit said it promises users they can delete their posts, and requires licensing partners (like Google) to honor those deletions, but that these data scrapers deprive Reddit of the ability "to give effect to subsequent user-deletion requests."

SerpApi and Perplexity sought an early dismissal for several reasons.

Among other arguments, SerpApi said Reddit lacked "standing" to sue because it did not own a copyright interest in users' posts, adding that the allegations in its complaint, even if proven true, would not show that it had been harmed by the alleged scraping.

Perplexity separately argued that the allegation that it used SerpApi's "tools" to directly scrape data was too conclusory to warrant further proceedings.

Engelmayer rejected those arguments for now, essentially ruling that if Reddit's allegations were proven true, they could support a conclusion that the company violated the Digital Millennium Copyright Act.

He also said Reddit's privacy-related allegations -- essentially that data scrapers would not necessarily shed posts that users deleted -- could support the claim that scraping injured its reputation.

While Engelmayer dismissed some of Reddit's claims, the ruling allows the company to move forward with its main contention.

Reddit praised the ruling. "Redditors create some of the most valuable human conversations on the internet," the spokesperson stated. "We intend to protect them."

The ruling comes less than two weeks after a different federal judge dismissed a complaint by Google against SerpApi over alleged scraping of the search results.

In that matter, U.S. District Court Judge Yvonne Gonzalez Rogers in the Northern District of California said Google could not pursue a claim that SerpApi circumvented restrictions on copying material that Google had not licensed and didn't otherwise own a copyright interest in.

Next story loading loading..