Reddit can proceed with a lawsuit alleging that Perplexity and data scraper SerpApi wrongly obtained copyrighted Reddit posts from Google's search results, a federal judge said
Friday.
The ruling, issued by U.S. District Court Judge Paul Engelmayer in New York, comes in a lawsuit filed by Reddit in October, when it claimed that SerpApi and artificial
intelligence company Perplexity violated the Digital Millennium Copyright Act's anti-circumvention provisions, which prohibit bypassing technological restrictions on copying.
Reddit's initial complaint included allegations that SerpApi used a workaround to bypass Google's SearchGuard -- an anti-scraping system that relies on CAPTCHAs -- in order to provide
Perplexity with Reddit posts.
An amended complaint filed by Reddit in February includes allegations that Perplexity itself used tools provided by SerpApi to obtain Reddit
posts.
advertisement
advertisement
"Perplexity is doing more than simply receiving data from SerpApi," Reddit alleged in that amended complaint. "Perplexity circumvents Google’s security protocols
and gains access, through unauthorized and automated processes, to Reddit data by using (at least) SerpApi’s tools or directing others to use those tools."
Reddit also
alleged that content posted to its site includes posts by users as well as "tens of thousands of posts, comments, and other original works authored by Reddit itself."
For
instance, one post allegedly scraped by SerpApi was an "Admin-authored 700-word history of the origin of the subreddit."
Reddit also claimed that it suffered "reputational
harm" with regard to privacy.
Specifically, Reddit said it promises users they can delete their posts, and requires licensing partners (like Google) to honor those deletions, but that
these data scrapers deprive Reddit of the ability "to give effect to subsequent user-deletion requests."
SerpApi and Perplexity sought an early dismissal for several
reasons.
Among other arguments, SerpApi said Reddit lacked "standing" to sue because it did not own a copyright interest in users' posts, adding that the allegations in its
complaint, even if proven true, would not show that it had been harmed by the alleged scraping.
Perplexity separately argued that the allegation that it used SerpApi's "tools"
to directly scrape data was too conclusory to warrant further proceedings.
Engelmayer rejected those arguments for now, essentially ruling that if Reddit's allegations were
proven true, they could support a conclusion that the company violated the Digital Millennium Copyright Act.
He also said Reddit's privacy-related allegations -- essentially
that data scrapers would not necessarily shed posts that users deleted -- could support the claim that scraping injured its reputation.
While Engelmayer dismissed some of
Reddit's claims, the ruling allows the company to move forward with its main contention.
Reddit praised the ruling. "Redditors create some of the most valuable human
conversations on the internet," the spokesperson stated. "We intend to protect them."
The ruling comes less than two weeks after a different federal judge dismissed a complaint by Google against SerpApi over alleged scraping of the
search results.
In that matter, U.S. District Court Judge Yvonne Gonzalez Rogers in the Northern District of California said Google could not pursue a claim that SerpApi circumvented
restrictions on copying material that Google had not licensed and didn't otherwise own a copyright interest in.