The AI company Perplexity is complaining their bots can't bypass Cloudflare's firewall

Davriellelouna@lemmy.world · edit-2 1 day ago

The AI company Perplexity is complaining their bots can't bypass Cloudflare's firewall

kreskin@lemmy.world · edit-2 3 hours ago

they cant get their ai to check a box that says “I am not a robot”? I’d think thatd be a first year comp sci student level task. And robots.txt files were basically always voluntary compliance anyway.

TheGrandNagus@lemmy.world · 5 hours ago

Can’t believe I’ve lived to see Cloudflare be the good guys

poopkins@lemmy.world · edit-2 2 hours ago

I’ve developed my own agent for assisting me with researching a topic I’m passionate about, and I ran into the exact same barrier: Cloudflare intercepts my request and is clearly checking if I’m a human using a web browser. (For my network requests, I’ve defined my own user agent.)

So I use that as a signal that the website doesn’t want automated tools scraping their data. That’s fine with me: my agent just tells me that there might be interesting content on the site and gives me a deep link. I can extract the data and carry on my research on my own.

I completely understand where Perplexity is coming from, but at scale, implementations like this are awful for the web.

Wispy2891@lemmy.world · edit-2 6 hours ago

Here comes the ridiculous offer to buy Google chrome with money they don’t have: easy delicious scraping directly from the user source

kittenzrulz123@lemmy.blahaj.zone · 9 hours ago

tibi@lemmy.world · 15 hours ago

You could say they are… Perplexed.

Kissaki@feddit.org · edit-2 17 hours ago

Perplexity argues that a platform’s inability to differentiate between helpful AI assistants and harmful bots causes misclassification of legitimate web traffic.

So, I assume Perplexity uses appropriate identifiable user-agent headers, to allow hosters to decide whether to serve them one way or another?

lime!@feddit.nu · 14 hours ago

yeah it’s almost like there as already a system for this in place

WolfLink@sh.itjust.works · 18 hours ago

This is a nice CloudFlare ad

pyre@lemmy.world · 15 hours ago

yeah. still not worth dealing with fucking cloudflare. fuck cloudflare.

oppy1984@lemdro.id · 36 minutes ago

I’m out of the loop, what’s wrong with cloud flare?

int32@lemmy.dbzer0.com · 15 hours ago

DEATH TO CLOUDFLARE!

pressanykeynow@lemmy.world · 6 hours ago

That would be terrible for a lot of people as they are the only company providing such services that doesn’t charge for traffic.

int32@lemmy.dbzer0.com · edit-2 4 hours ago

They can use web.archive.org as a cdn(I do that to cloudflare websites). But honestly, cloudflare or not, the internet is broken.

pressanykeynow@lemmy.world · 1 hour ago

Can you explain please? How can I use archive.org as a cdn for my website?

NotASharkInAManSuit@lemmy.world · 17 hours ago

That’s the entire point, dipshit. I wish we got one of the cool techno dystopias rather than this boring corporate idiot one.

Leon@pawb.social · 17 hours ago

I’m still holding out for Stephen Hawking to mail out Demon Summoning programs.

Frezik@lemmy.blahaj.zone · 19 hours ago

Traveling snake oil salesman complains he can’t pick people’s locks.

Glitchvid@lemmy.world · 24 hours ago

When a firm outright admits to bypassing or trying to bypass measures taken to keep them out, you think that would be a slam dunk case of unauthorized access under the CFAA with felony enhancements.

GamingChairModel@lemmy.world · 23 hours ago

Fuck that. I don’t need prosecutors and the courts to rule that accessing publicly available information in a way that the website owner doesn’t want is literally a crime. That logic would extend to ad blockers and editing HTML/js in an “inspect element” tag.

kibiz0r@midwest.social · 20 hours ago

They already prosecute people under the unauthorized access provision. They just don’t prosecute rich people under it.

GamingChairModel@lemmy.world · 17 hours ago

They prosecuted and convicted a guy under the CFAA for figuring out the URL schema for an AT&T website designed to be accessed by the iPad when it first launched, and then just visiting that site by trying every URL in a script. And then his lawyer (the foremost expert on the CFAA) got his conviction overturned:

https://www.eff.org/cases/us-v-auernheimer

We have to maintain that fight, to make sure that the legal system doesn’t criminalize normal computer tinkering, like using scripts or even browser settings in ways that site owners don’t approve of.

Encrypt-Keeper@lemmy.world · 23 hours ago

That logic would not extend to ad blockers, as the point of concern is gaining unauthorized access to a computer system or asset. Blocking ads would not be considered gaining unauthorized access to anything. In fact it would be the opposite of that.

GamingChairModel@lemmy.world · 22 hours ago

gaining unauthorized access to a computer system

And my point is that defining “unauthorized” to include visitors using unauthorized tools/methods to access a publicly visible resource would be a policy disaster.

If I put a banner on my site that says “by visiting my site you agree not to modify the scripts or ads displayed on the site,” does that make my visit with an ad blocker “unauthorized” under the CFAA? I think the answer should obviously be “no,” and that the way to define “authorization” is whether the website puts up some kind of login/authentication mechanism to block or allow specific users, not to put a simple request to the visiting public to please respect the rules of the site.

To me, a robots.txt is more like a friendly request to unauthenticated visitors than it is a technical implementation of some kind of authentication mechanism.

Scraping isn’t hacking. I agree with the Third Circuit and the EFF: If the website owner makes a resource available to visitors without authentication, then accessing those resources isn’t a crime, even if the website owner didn’t intend for site visitors to use that specific method.

Glitchvid@lemmy.world · edit-2 22 hours ago

When sites put challenges like Anubis or other measures to authenticate that the viewer isn’t a robot, and scrapers then employ measures to thwart that authentication (via spoofing or other means) I think that’s a reasonable violation of the CFAA in spirit — especially since these mass scraping activities are getting attention for the damage they are causing to site operators (another factor in the CFAA, and one that would promote this to felony activity.)

The fact is these laws are already on the books, we may as well utilize them to shut down this objectively harmful activity AI scrapers are doing.

ubergeek@lemmy.today · 21 hours ago

The fact is these laws are already on the books, we may as well utilize them to shut down this objectively harmful activity AI scrapers are doing.

Silly plebe! Those laws are there to target the working class, not to be used against corporations. See: Copyright.

RangerAndTheCat@lemmy.dbzer0.com · 20 hours ago

Aatube@lemmy.dbzer0.com · 16 hours ago

That same logic is how Aaron Swartz was cornered into suicide for scraping JSTOR, something widely agreed to be a bad idea by a wide range of lawspeople including SCOTUS in its 2021 decision Van Buren v. US that struck this interpretation off the books.

tomalley8342@lemmy.world · 20 hours ago

Nah, that would also mean using Newpipe, YoutubeDL, Revanced, and Tachiyomi would be a crime, and it would only take the re-introduction of WEI to extend that criminalization to the rest of the web ecosystem. It would be extremely shortsighted and foolish of me to cheer on the criminalization of user spoofing and browser automation because of this.

Glitchvid@lemmy.world · edit-2 14 hours ago

Do you think DoS/DDoS activities should be criminal?

If you’re a site operator and the mass AI scraping is genuinely causing operational problems (not hard to imagine, I’ve seen what it does to my hosted repositories pages) should there be recourse? Especially if you’re actively trying to prevent that activity (revoking consent in cookies, authorization captchas).

In general I think the idea of “your right to swing your fists ends at my face” applies reasonably well here — these AI scraping companies are giving lots of admins bloody noses and need to be held accountable.

I really am amenable to arguments wrt the right to an open web, but look at how many sites are hiding behind CF and other portals, or outright becoming hostile to any scraping at all; we’re already seeing the rapid death of the ideal because of these malicious scrapers, and we should be using all available recourse to stop this bleeding.

tomalley8342@lemmy.world · 13 hours ago

DoS attacks are already a crime, so of course the need for some kind of solution is clear. But any proposal that gatekeeps the internet and restricts the freedoms with which the user can interact with it is no solution at all. To me, the openness of the web shouldn’t be something that people just consider, or are amenable to. It should be the foundation in which all reasonable proposals should consider as a principle truth.

Encrypt-Keeper@lemmy.world · 20 hours ago

If I put a banner on my site that says “by visiting my site you agree not to modify the scripts or ads displayed on the site,” does that make my visit with an ad blocker “unauthorized” under the CFAA?

How would you “authorize” a user to access assets served by your systems based on what they do with them after they’ve accessed them? That doesn’t logically follow so no, that would not make an ad blocker unauthorized under the CFAA. Especially because you’re not actually taking any steps to deny these people access either.

AI scrapers on the other hand are a type of users that you’re not authorizing to begin with, and if you’re using CloudFlares bot protection you’re putting into place a system to deny them access. To purposefully circumvent that access would be considered unauthorized.

GamingChairModel@lemmy.world · 17 hours ago

That doesn’t logically follow so no, that would not make an ad blocker unauthorized under the CFAA.

The CFAA also criminalizes “exceeding authorized access” in every place it criminalizes accessing without authorization. My position is that mere permission (in a colloquial sense, not necessarily technical IT permissions) isn’t enough to define authorization. Social expectations and even contractual restrictions shouldn’t be enough to define “authorization” in this criminal statute.

To purposefully circumvent that access would be considered unauthorized.

Even as a normal non-bot user who sees the cloudflare landing page because they’re on a VPN or happen to share an IP address with someone who was abusing the network? No, circumventing those gatekeeping functions is no different than circumventing a paywall on a newspaper website by deleting cookies or something. Or using a VPN or relay to get around rate limiting.

The idea of criminalizing scrapers or scripts would be a policy disaster.

cm0002@piefed.world · 20 hours ago

You say, just as news breaks that the top German court has over turned a decision that declared “AD blocking isn’t piracy”

Encrypt-Keeper@lemmy.world · 20 hours ago

Unauthorized access into a computer system and “Piracy” are two very different things.

cm0002@piefed.world · 20 hours ago

Please instruct me on how I go to the timeline where the legal system always makes decisions based on logic, reasoning, evidence and fairness and not…the opposite…of all those things

You have a lot of trust placed in the courts to actually do the right thing

Encrypt-Keeper@lemmy.world · edit-2 19 hours ago

I’m not saying courts couldn’t pass a new law saying whatever they want. But the laws we have today would not allow for ad blocking to be considered unauthorized access. Not under the CFAA as mentioned.

I said “The logic would not extend to that” not that a legal system could not act illogically.

cm0002@piefed.world · 18 hours ago

The original comment reply to you was all about how the legal system would act, that’s the primary concern. All it would take is a Trump loyalist judge, a Trump leaning appeals court and the right-wing Supreme Court and boom suddenly the CFAA covers a whole lot more than what was “logical”

Demdaru@lemmy.world · 22 hours ago

Ehhhh, you are gaining access to content due to assumption you are going to interact with ads and thus, bring revenue to the person and/or company producing said content. If you block ads, you remove authorisation brought to you by ads.

gian @lemmy.grys.it · 3 hours ago

Carefull, this way even not looking at an ads positioned at the bottom of the page (or anyway not visible without scrolling) would mean to remove authorisation brought to you by ads.

Encrypt-Keeper@lemmy.world · edit-2 19 hours ago

That doesn’t make any logical sense. You cant tie legal authorization to an unsaid implicit assumption, especially when that is in turn based on what you do with the content you’ve retrieved from a system after you’ve accessed and retrieved it.

When you access a system, are you authorized to do so, or aren’t you? If you are, that authorization can’t be retroactively revoked. If that were the case, you could be arrested for having used a computer at a job, once you’ve quit. Because even though you were authorized to use it and your corporate network while you worked there, now that you’ve quit and are no longer authorized that would apply retroactively back to when you DID work there.

ℍ𝕂-𝟞𝟝@sopuli.xyz · 21 hours ago

There was no header on the request saying I want ads though

jve@lemmy.world · edit-2 18 hours ago

Right? Isn’t this a textbook DMCA violation, too?

WhyJiffie@sh.itjust.works · 10 hours ago

for us, not for them. wait until they argue in court that actually its us at fault and we need to provide access or else

EtherWhack@lemmy.world · 20 hours ago

floquant@lemmy.dbzer0.com · 1 day ago

It’s difficult to be a shittier company than OpenAI, but Perplexity seems to be trying hard.

BigFig@lemmy.world · 22 hours ago

Step 1, SOMEHOW find a more punchable face than Altman

Tollana1234567@lemmy.today · edit-2 5 hours ago

put META android zuckerberg on or mechahitler musk.

SugarCatDestroyer@lemmy.world · edit-2 18 hours ago

It seems like it’s some kind of distraction to make people think things aren’t as bad as they really are, it just sounds too far-fetched to me.

It’s like a bear that has eaten too much and starts whining because a small rabbit is running away from him, even though the bear has already eaten almost all the rabbits and is clearly full.

StocktonCrushed@sh.itjust.works · 18 hours ago

deleted by creator

SugarCatDestroyer@lemmy.world · 17 hours ago

So that he doesn’t have to run after the rabbits, he will learn to raise them and manage them with a fake smile, providing them with a stable life lol.

Well, I think the thing is that we still live by the law: the strong do what they want, and the weak just whine and complain.

ubergeek@lemmy.today · 21 hours ago

Good. I went through my CF panel, and blocked some of those “AI Assistants” that by default were open, including Perplexity’s.

_g_be@lemmy.world · edit-2 18 hours ago

CF panel? Your light bulb??

ubergeek@lemmy.today · 17 hours ago

CF == Cloudflare :)

The AI company Perplexity is complaining their bots can't bypass Cloudflare's firewall

The AI company Perplexity is complaining their bots can't bypass Cloudflare's firewall

Perplexity Says Cloudflare Is Blocking Legitimate AI Assistants