Pay-per-output? AI firms blindsided by beefed up robots.txt instructions.
https://arstechnica.com/tech-policy/2025/09/pay-per-output-ai-firms-blindsided-by-beefed-up-robots-txt-instructions/?utm_social-type=owned
138 points · 27 comments · view on lemmy.world
27 Comments
underline960@sh.itjust.works · 66 pts · 357d
This article tries to slip in the idea that creators will benefit from this arrangement. Just like with Spotify and Getty Images, it's the publisher that's getting paid.
Then they decide how much they'll let trickle down to creators.
ICastFist@programming.dev · 19 pts · 357d
Cue an even greater influx of AI slop pages in hopes of getting crawled for that juicy trickled down money
ccunning@lemmy.world · 7 pts · 357d
I would assume creators and published would agree to those terms in advance (moving forward of course).
Kissaki@feddit.org · 31 pts · 356d
robots.txt - the well known technology to block bad-intention bots /s
What's automated about the licensing layer? At some point, I started skimming the article. They didn't seem clear about it. The AI can "automatically" parse it?
Yeah, this is as useless as I thought it would be. Nothing here is actively blocking.
I love that the XML then points to a text/html content website. I guess nothing for machine parsing, maybe for AI parsing.
I don't remember which AI company, but they argued they're not crawlers but agents acting on the users behalf for their specific request/action, ignoring robots.txt. Who knows how they will react. But their incentives and history is ignoring robots.txt.
Why
am Iis this comment so negative. Oh well.FaceDeer@fedia.io · 25 pts · 357d
And suddenly the Internet is gung-ho in favor of EULAs being enforceable simply by reading the content the website has already provided.
Recent major court cases have held that the training of an AI model is fair use and doesn't involve copyright violation, so I don't think licensing actually matters in this case. They'd have to put the content behind a paywall to stop the trainer from seeing it in the first place.
ccunning@lemmy.world · 8 pts · 357d
I guess that’s a different court case than the one where Anthropic offered to pay $1.5 billion?
NewNewAugustEast@lemmy.zip · 6 pts · 356d
Totally different. Anthropic could have bought all the books and trained on them. Pirating is a different topic.
corsicanguppy@lemmy.ca · 3 pts · 356d
You think buying the books would let them plagiarize ? That doesn't seem to be normal in the "book buying" process.
NewNewAugustEast@lemmy.zip · 6 pts · 356d
Doesn't really matter what I think, its a different concept than pirating. Hence a different thing than what was getting ruled on.
I mean AI or not look at it this way: if a company wanted to train their workers and pirated all the training manuals, piracy is the issue, not the training.
Womble@piefed.world · 2 pts · 356d
Given the judege in that case flat out rejected the claim that there was any infringement for works they had legally aquired, yes.
FaceDeer@fedia.io · 5 pts · 356d
Nope, this was one of them. The case had two parts, one about the training and one about the downloading of pirated books. The judge issued a preliminary judgment about the training part, that was declared fair use without any further need to address it in trial. The downloading was what was proceeding to trial and what the settlement offer was about.
tabular@lemmy.world · 0 pts · 356d
Is it hypocrisy to be for EULA enforcement on reading when it's machines, but not when it's humans? Crawlers "read" on a massive scale that doesn't compare to humans.
WhyJiffie@sh.itjust.works · 1 pts · 356d
I don't think so, or not always. humans need to find the EULA on the website by first loading the main page or another they found a link to. but if the path of that document was standardized, it could be enforced that way for robots
GissaMittJobb@lemmy.ml · 24 pts · 357d
I have no idea what they think this will accomplish, to be honest. It has the legal value of posting on Facebook that you don't allow them to use your photos.
ccunning@lemmy.world · 7 pts · 357d
I think the idea is that all parties would find it beneficial:
ricecake@sh.itjust.works · 16 pts · 357d
The thing is a robots.txt file doesn't work as licensing. There's no legal requirement to fetch the file, and no mechanism to consent or track consent.
This is putting up a sign that says everyone must pay, and then giving it to anyone who asks for free.
ccunning@lemmy.world · 1 pts · 357d
The thing is if all parties find the terms agreeable it doesn’t matter if it’s legally binding.
It’s more like putting a price on the shelf at the grocery store. Not every one will agree the price is agreeable and you might still get shoplifters but it doesn’t mean it’s a waste of time to list the price.
ricecake@sh.itjust.works · 8 pts · 357d
It really does matter if it's legally binding if you're talking about content licensing. That's the whole thing with a licensing agreement: it's a legal agreement.
The store analogy isn't quite right. Leaving a store with something you haven't purchased with the consent of the store is explicitly illegal.
With a website, it's more like if the "shoplifter" walked in, didn't request a price sheet, picked up what they wanted and went to the cashier who explicitly gave it to them without payment.
The crux of the issue is that the website is still providing the information even if the requester never agreed or was even presented with the terms.
If your site wants to make access to something conditional then it needs to actually enforce that restriction.
It's why the current AI training situation is unlikely to be resolved without laws to address it explicitly.
Telorand@reddthat.com · 0 pts · 356d
I think the analogy is apt. If you post a price on goods, and somebody walks into a store, picks up the item, and walks out without paying, they can't simply say, "Well, I didn't care to read the price, and nobody presented me with a contract, so I just took it," as a valid defense. There's sometimes an explicit agreement upon terms, sure, but there are times where that agreement is implicit: they put a price on a thing, I pay it, else it's stealing. I don't need to sign a contract every time I get groceries.
I do, however, agree that this will only have teeth once it's argued and upheld in court the first (few) time(s). If nothing else, it's good to see people trying to solve the problem, rather than just throwing up their hands and letting billionaires run amok with virtual impunity. Maybe this won't work to reign in AI tech bros, but maybe it will inspire the things that do.
ricecake@sh.itjust.works · 2 pts · 356d
Except that with the website example it's not that they're ignoring the price or just walking out with the item. It's that the item was not labeled with a price, nor were they informed of the price. Then, rather than just walking out, they requested the item and it was delivered to them with no attempt to collect payment.
The key part of a website is that the user cannot take something. The site has to give it to them.
A more apt retail analogy might be you go to a website. You see a scooter you like, so you click "I want it!". The site then asks for your address and a few days later you get a scooter in the mail.
That's not theft, it's a free scooter. If the site accused you of theft because you didn't navigate to an unlinked page they didn't tell you about to find the prices, or try to figure out payment before requesting, you'd rightly be pretty miffed.
The shoplifting analogy doesn't work because it's not shoplifting if the vendor gives it to you knowingly and you never misrepresented the cost or tried to avoid paying. Additionally, taking someone's property without their permission is explicitly illegal, and we have a subcategory that explicitly spells out how retail fraud works and is illegal.
Under our current system the way to prevent someone from having your thing without paying or meeting some other criteria first is to collect payment or check that criteria before giving it to them.
To allow people to have things on their website freely available to humans but to prevent grabbing and using it for training will require a new law of some sort.
billwashere@lemmy.world · 19 pts · 356d
The issue is the line that says “compensate creators”. Reddit still thinks it’s the creator, not the individual users.
trailee@sh.itjust.works · 10 pts · 357d
Neither the article nor the RSL website makes clear how pricing or payment works, which seems like a huge miss. It’s not obvious if a publisher can price-differentiate among content, or even choose their own prices at all.
RSL makes an analogy:
I’d like to get excited about this because AI companies suck, but if the best example they have is that ASCAP helps “musicians get paid fairly” I’m afraid this isn’t a solution that most content creators will celebrate.
zrst@lemmy.cif.su · 9 pts · 357d
Does AI cost advertisers money?
I'd be cool with it if that's the case.
tchambers@crust.piefed.social · 9 pts · 356d
Wonder if this could work for Fediverse servers too.
rimu@piefed.social · 3 pts · 356d
Hold my beer
BrianTheeBiscuiteer@lemmy.world · 7 pts · 357d
Not a bad idea but the biggest challenge will probably be determining who needs to be sued for non-compliance. Google might not be hiding the origin of its bots now but that could easily change.
tchambers@crust.piefed.social · 3 pts · 357d
Interesting.