if a C&D doesn't accuse you of anything illegal, then it absolutely cannot compel you under any circumstance:
This doesn't have to mean criminal penalties. If WMF simply tells the scrapers that they are no longer authorized to access their systems, they can litigate against companies who continue to breach their request to discontinue scraping. That can be a civil action.
I can't imagine that Wikimedia is somehow required to feed the LLMs, and that they cannot simply request that they stop. You and I may not get a lot of traction there, but $263M definitely ought to make it possible to hire a lawyer to defend against (now) unauthorized access.
We don't know that because WMF hasn't bothered to try defending contributors.
Very false. Bot detection and blocking efforts have always been pretty documented
I don't see how bot detection shows that WMF is trying to defend contributors' IP.
Instead, they created a glide path for the pirates taking advantage of them.
Before Wikimedia Enterprise it was way easier. Again, API Access was, until very recently, unlimited and free. I'm not sure how Enterprise is the glide path here.
Before Wikimedia Enterprise they weren't getting paid and the LLM companies may have had to depend on residential proxies -- if the bot detection you mention was successful, for example. The glide path allows Enterprise customers to not experience rate limits or rely on shoddy infrastructure. Clearly, it's better to not be in the shadows, if you are trying to evade detection.
PS: Were they using the API or were they scraping? If they were using the API, why couldn't WMF just revoke their keys?
To have legal standing for a letter, what you're asking to cease and desist needs to be illegal.
FWIW, that isn't true - I didn't post the letter, but I got a C&D from SoFi for this post. Clearly I had done nothing illegal.
WMF isn't Nintendo or Hatchette; you need way more than $296M to pursue all that, not to mention $200M of that is the standard practice of keeping a 12 months' rain fund in case something massive happens to current revenue.
They can't defend the contributors, but they can hire a union busting law firm to keep their staff in check. Got it.
we know that these violations exploiting the commons are going to happen regardless of whether Enterprise exists
We don't know that because WMF hasn't bothered to try defending contributors. Instead, they created a glide path for the pirates taking advantage of them.
so they might as well make the exploiters pay for it to help maintain the commons' infrastructure.
I think it is interesting you say that WMF is being "exploited" yet you believe that they don't have any standing for the damages they have experienced.
Send one to whom? A WHOIS of the Brazilian IP that turns out to be a residential proxy? Anonymous scraping for LLMs is a problem all over the Internet without a solution (save Anubis). There's a reason all the big companies have went for Cloudflare instead of any lawsuits, which you'd hear in the news.
I know individuals that have been unmasked in torrent swarms and have had their ISP cancel their service due to that. The idea that Wikimedia's hands are powerless to send a Cease and Desist to ISPs to warn and ban their customers for scraping is hilarious.
Again, they have $296M in reserves. WMF can send a letter.
They're not even sure if they have the standing; they've looked at that question
But that isn't what you linked to says; the word standing doesn't appear in the text, nor do they seem to explore that. Thanks for the reference, but it doesn't actually support your argument.
I am simply representing the perspective of contributors as a signatory myself, and as a developer who makes API calls to Wikipedia. "Our content is always free to use, but our infrastructure is not" sums it up nicely and Wikimedia Enterprise keeps the free infrastructure good.
As I responded to another commenter:
I also don't know how much pointing out that this access has been granted for other users matters - the current CEO very clearly states that this was built to support the scraping use cases; Wikimedia knows that the scrapers are violating the licenses with every derivative work created not licensed reciprocally, and designed the feature to make that happen faster.
Enabling non-violating use cases don't erase the violating ones.
This omits how scrapers and crawlers would have been getting the corpus for free without that, profiteering and clogging up the tubes that bring you Wikipedia and tons of amazing applications that rely on it.
I didn't know that I had to preemptively defend the Foundation.
The reason I don't think that that is particularly relevant is because these companies are violating the license that Wikimedia projects are distributed under. WMF has $296M in reserves. They couldn't send a Cease and Desist for the server load?
BTW it's already established that training AI models on copyrighted materials is legal without approval of the copyright holder, and IMO it's unikely WP's material would be an exception.
That isn't accurate - this is a highly unsettled question, and there are multiple cases in litigation today.
WMF would need some better argument if they'd want to sue successfully, especially aginst companies that are sitting on 100x more money than them.
A better argument than what - that they are openly violating the licenses under which the encyclopedia is distributed? What more do they need?
My actual sources for statements of fact were already listed. It's a shorthand for "This is my personal position as an experienced editor that this is an overreaction" that you're warping in bad-faith.
I have no idea why you think I am "warping" my reaction in bad faith - my reaction is based on the license text and what is written in the post. I also think it is ironic that you ask me to extend you grace in accepting your shorthand, and you clearly don't bother to accept mine (that your statement was a reference to your own authority as an experienced editor).
Material published to Wikimedia is CC BY-SA 4.0; thus, those current Enterprise customers have every right to use the material basically however they see fit regardless of Enterprise.
You don't actually tell us why this is the case - I argue that the companies are violating the license by not licensing their derivative works reciprocally - you don't even bother to respond to that and just posit that they have "every right to use the material however they see fit". Do you really believe that? Are the CC-BY-SA and GFDL licenses just completely worthless?
Wikimedia isn't selling access to the material, because it's literally nobody's to sell; they're selling access to stream the data on their servers which they host.
I'm not sure how much that matters. Would Warner Brothers not have an issue with me "streaming" access to their movies via my home server for payment?
I also don't know how much pointing out that this access has been granted for other users matters - the current CEO very clearly states that this was built to support the scraping use cases; Wikimedia knows that the scrapers are violating the licenses with every derivative work created not licensed reciprocally, and designed the feature to make that happen faster.
Enabling non-violating use cases don't erase the violating ones.
All contributors agree to release their content in perpetuity under CC-BY-SA (or similar licences that preceded it), meaning that all content is free for anyone to use, including commercial uses and derivatives, with the only restrictions being requiring attribution (credit the original contributor) and releasing under an equivalent license. The WMF can't restrict access to any use complying with that license.
The Wikipedia community (its editors) could decide not to allow its content to be used by AI, but it has not. It would be very legally complicated, anyway, given the content's license.
Why do you think the license allows for big tech to produce derivative works that are not licensed CC-BY-SA?
The WMF negotiating paid privileged access to data streams for these large clients is a win for everyone. The purpose of Wikipedia is to disseminate information, not gatekeep it, and the WMF has literally no right to decide who can access it and who cannot.
If WMF has no right to decide who can access it and who cannot, how can they sell privileged access to who can access it? Frankly, that assertion fails on its face.
Clearly, everyone realizes that the big tech AI pirates will scrape and stream the data - what I am objecting to is the response. When Google began to pirate Disney's IP, Disney didn't immediately offer them a data sharing deal (that also somehow doesn't provide access to the IP) - they sued.
The WMF has $296M in assets and they are grubbing for the pocket change that big tech throws at them for the fruits of the unpaid labor in Wikipedia. Why aren't they suing for us? They are the stewards of the corpus.
Has WMF even threatened the big tech owners with a good time, or were they simply salivating for the addition to their bottom line? Did they tell them to download the database? Did they attempt to ban their servers? Or did they simply provide big tech with privileged access to data they do not own?
While clearly commentary from the community will be more interesting than from it outside of it, I don't begrudge analysis or reporting from traditional media outlets - or do we want this all to be a private matter that 404 Media and the like don't cover, allowing the theft (and union busting) to continue apace?
But any source code leak is also open sourcing in that world.
I don't see how that helps free software, though. Those programmers got paid. Volunteers didn't.
It ends up with a weird reverse robin hood situation. LLM vendors steal from the poor, sell that to the rich. Do the rich give back? Only if it is stolen from them.
Can I legally reverse engineer AI generated software?
If you have the source, why would you need to?
Can you even put terms and conditions on this supposed public domain copyright free compiled software product?
You can put terms on anything, but you can't protect the underlying asset if someone breaks your terms. Think of the code produced by Grsecruity that they put behind a paywall -- people were free to release the code (since it was licensed as open source as a derivative work), but obviously Grsecruity was able to discontinue their agreement with their clients who would do so.
Is the compiled version even different than the raw AI generated source code in its ability to be licensed?
People aren't generally licensing compiled binaries as open source, since you can't produce derivative works from them. But I think that if there is no copyright protection for the work, compiling it doesn't change the copyrightability. Curious what you think.
What rights does one have to AI generated code? Be it compiled or source. It’s surely not just communal.
Why is that surely the case? It is public domain - that is the most "communal" you can get for copyright.
I have seen this sentiment, but I don't know what the world looks like without copyright protections for creative works.
Does open source exist in your vision? How?
My imagination for this topic may not be as expansive as yours, but my interpretation is that if people contribute code to the commons, it will immediately available for any use - including for use by massive corporations.
So it ends up looking like people working for big companies for free.
as soon as it's modified by a human in nontrivial ways
is doing a lot of heavy lifting here.
We know that people are using coding LLMs as slot machines - pull the handle and see if it solves your problem. Where is the human modifying anything? That is a "straight dump" of AI output without modifications.
Honestly, if AI destroys copyright, it's the best thing it can do.
I have seen this being said, but I really don't understand it. Just because copyright can be abused doesn't mean (to me) that we ought to throw the baby out with the bathwater.
If copyright no longer exists, what incentive do people have to share copyleft code at all? It clearly would no longer exist, so can you help me understand how both copyright can be dead and open source exist? Or are you simply accepting that rather than copyright, we are using trade secrets (like the KFC chicken recipe) to protect works?
How does this apply to software made by, say, Anthropic? They proudly say Claude Code is written by AI. If it can’t be copywritten, or licensed, then it’s just a matter of figuring out how to acquire a copy of the source code, and you could do whatever with it. Right?
If you were on Mastodon last week when the Claude source code was released (by Claude, accidentally), people were joking about how Anthropic was trying to use the DMCA to get the source removed from websites -- even though clearly, copyrights don't apply, since the code is clearly in the public domain.
All works created by a person are copyright by default, so people need to release their works to allow others to build on it or use it (except for the limited uses allowed by fair use). Like-minded people have come up with various licenses that allow people to release their works in ways that people prefer.
This doesn't have to mean criminal penalties. If WMF simply tells the scrapers that they are no longer authorized to access their systems, they can litigate against companies who continue to breach their request to discontinue scraping. That can be a civil action.
I can't imagine that Wikimedia is somehow required to feed the LLMs, and that they cannot simply request that they stop. You and I may not get a lot of traction there, but $263M definitely ought to make it possible to hire a lawyer to defend against (now) unauthorized access.
I don't see how bot detection shows that WMF is trying to defend contributors' IP.
Before Wikimedia Enterprise they weren't getting paid and the LLM companies may have had to depend on residential proxies -- if the bot detection you mention was successful, for example. The glide path allows Enterprise customers to not experience rate limits or rely on shoddy infrastructure. Clearly, it's better to not be in the shadows, if you are trying to evade detection.
PS: Were they using the API or were they scraping? If they were using the API, why couldn't WMF just revoke their keys?
FWIW, that isn't true - I didn't post the letter, but I got a C&D from SoFi for this post. Clearly I had done nothing illegal.
They can't defend the contributors, but they can hire a union busting law firm to keep their staff in check. Got it.
We don't know that because WMF hasn't bothered to try defending contributors. Instead, they created a glide path for the pirates taking advantage of them.
I think it is interesting you say that WMF is being "exploited" yet you believe that they don't have any standing for the damages they have experienced.
I know individuals that have been unmasked in torrent swarms and have had their ISP cancel their service due to that. The idea that Wikimedia's hands are powerless to send a Cease and Desist to ISPs to warn and ban their customers for scraping is hilarious.
Again, they have $296M in reserves. WMF can send a letter.
But that isn't what you linked to says; the word standing doesn't appear in the text, nor do they seem to explore that. Thanks for the reference, but it doesn't actually support your argument.
As I responded to another commenter:
I didn't know that I had to preemptively defend the Foundation.
The reason I don't think that that is particularly relevant is because these companies are violating the license that Wikimedia projects are distributed under. WMF has $296M in reserves. They couldn't send a Cease and Desist for the server load?
That isn't accurate - this is a highly unsettled question, and there are multiple cases in litigation today.
A better argument than what - that they are openly violating the licenses under which the encyclopedia is distributed? What more do they need?
They are sitting on $296M. Why aren't they suing the AI companies to defend the contributors?
I have no idea why you think I am "warping" my reaction in bad faith - my reaction is based on the license text and what is written in the post. I also think it is ironic that you ask me to extend you grace in accepting your shorthand, and you clearly don't bother to accept mine (that your statement was a reference to your own authority as an experienced editor).
You don't actually tell us why this is the case - I argue that the companies are violating the license by not licensing their derivative works reciprocally - you don't even bother to respond to that and just posit that they have "every right to use the material however they see fit". Do you really believe that? Are the CC-BY-SA and GFDL licenses just completely worthless?
I'm not sure how much that matters. Would Warner Brothers not have an issue with me "streaming" access to their movies via my home server for payment?
I also don't know how much pointing out that this access has been granted for other users matters - the current CEO very clearly states that this was built to support the scraping use cases; Wikimedia knows that the scrapers are violating the licenses with every derivative work created not licensed reciprocally, and designed the feature to make that happen faster.
Enabling non-violating use cases don't erase the violating ones.
Why do you think the license allows for big tech to produce derivative works that are not licensed CC-BY-SA?
If WMF has no right to decide who can access it and who cannot, how can they sell privileged access to who can access it? Frankly, that assertion fails on its face.
She runs the foundation and is firing union workers. I'm just writing commentary. We are not the same.
Volunteers' work is stolen no matter if it is my own or others - me being an editor has no bearing on that.
Hmm: "Source: ~50,000+ contributions to Wikimedia projects."
Clearly, everyone realizes that the big tech AI pirates will scrape and stream the data - what I am objecting to is the response. When Google began to pirate Disney's IP, Disney didn't immediately offer them a data sharing deal (that also somehow doesn't provide access to the IP) - they sued.
The WMF has $296M in assets and they are grubbing for the pocket change that big tech throws at them for the fruits of the unpaid labor in Wikipedia. Why aren't they suing for us? They are the stewards of the corpus.
Has WMF even threatened the big tech owners with a good time, or were they simply salivating for the addition to their bottom line? Did they tell them to download the database? Did they attempt to ban their servers? Or did they simply provide big tech with privileged access to data they do not own?
While clearly commentary from the community will be more interesting than from it outside of it, I don't begrudge analysis or reporting from traditional media outlets - or do we want this all to be a private matter that 404 Media and the like don't cover, allowing the theft (and union busting) to continue apace?
In any case, I'll show you mine.
[Removed a reference to a different comment that I should have investigated more deeply.]
I don't see how that helps free software, though. Those programmers got paid. Volunteers didn't.
It ends up with a weird reverse robin hood situation. LLM vendors steal from the poor, sell that to the rich. Do the rich give back? Only if it is stolen from them.
If you have the source, why would you need to?
You can put terms on anything, but you can't protect the underlying asset if someone breaks your terms. Think of the code produced by Grsecruity that they put behind a paywall -- people were free to release the code (since it was licensed as open source as a derivative work), but obviously Grsecruity was able to discontinue their agreement with their clients who would do so.
People aren't generally licensing compiled binaries as open source, since you can't produce derivative works from them. But I think that if there is no copyright protection for the work, compiling it doesn't change the copyrightability. Curious what you think.
Why is that surely the case? It is public domain - that is the most "communal" you can get for copyright.
I have seen this sentiment, but I don't know what the world looks like without copyright protections for creative works.
Does open source exist in your vision? How?
My imagination for this topic may not be as expansive as yours, but my interpretation is that if people contribute code to the commons, it will immediately available for any use - including for use by massive corporations.
So it ends up looking like people working for big companies for free.
is doing a lot of heavy lifting here.
We know that people are using coding LLMs as slot machines - pull the handle and see if it solves your problem. Where is the human modifying anything? That is a "straight dump" of AI output without modifications.
I have seen this being said, but I really don't understand it. Just because copyright can be abused doesn't mean (to me) that we ought to throw the baby out with the bathwater.
If copyright no longer exists, what incentive do people have to share copyleft code at all? It clearly would no longer exist, so can you help me understand how both copyright can be dead and open source exist? Or are you simply accepting that rather than copyright, we are using trade secrets (like the KFC chicken recipe) to protect works?
If you were on Mastodon last week when the Claude source code was released (by Claude, accidentally), people were joking about how Anthropic was trying to use the DMCA to get the source removed from websites -- even though clearly, copyrights don't apply, since the code is clearly in the public domain.
If the LLM wrote the code, it is uncopyrightable.
All works created by a person are copyright by default, so people need to release their works to allow others to build on it or use it (except for the limited uses allowed by fair use). Like-minded people have come up with various licenses that allow people to release their works in ways that people prefer.
Except for the fact that it is public domain and not protected by the open source license that the code is ostensibly submitted under.