Elon Musk’s xAI used child porn to train Grok models, lawsuit says

https://arstechnica.com/tech-policy/2026/08/elon-musks-xai-used-child-porn-to-train-grok-models-lawsuit-says/

For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.”

Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones.

This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”

808 points · 82 comments · view on lemmy.world

82 Comments

MummifiedClient5000@feddit.dk · 172 pts · 22h (9 replies)

Must suck to be a rich pedophile and not even get an invite to Epstein island.

random_character_a@lemmy.world · 88 pts · 21h (2 replies)

Elon was too creepy for Epstein

Sharkticon@lemmy.zip · 93 pts · 21h (1 reply)

You know I know you're joking, but I'm concerned people believe this. Because Elon did hang out with Epstein. There was one email where he was ignored because he was being too thirsty. But that's not the whole story. He loved the island so much he wanted to go back too hard.

story@lemmy.zip · 4 pts · 9h

thanks for the truth, truth guy :)

chaogomu@lemmy.world · 47 pts · 20h (4 replies)
MummifiedClient5000@feddit.dk · 12 pts · 19h (2 replies)

So he does have friends after all...

ouRKaoS@lemmy.today · 17 pts · 19h (1 reply)

Sounds like he had a rich guy that wanted material to blackmail him with

uriel238@lemmy.blahaj.zone · 3 pts · 8h

Maybe! Epstein had early designs to essentially manufacture Kompromat by creating a private place for rich men to hook up with young teens. Only he a) got high on his own supply, and b) instead mostly turned to financing and money laundering and so they were more valuable to him as allies than adversaries.

Also, the owner class, when it's sufficiently wealthy, can influence political and justice systems enough to be beyond the law. OJ Simpson wasn't quite there, but he got a good lawyer team to get him acquitted. Andrew Mountbatten-Windsor, formerly Prince Andrew, Duke of York, was there, and was only disgraced because his family cares about the appearance of decency.

All the tech bros don't care about appearance of decency. Musk is publicly known to have killed hundreds of thousands by killing USAID, and it won't affect his career.

Rothe@piefed.social · 2 pts · 1h

Yeah, it is concerning that the myth of him not being allowed to the parties is so common on the internet. It is a myth that is very much in favour of Elon, because it makes people believe that he did not attend the pedo-parties, when we know that he absolutely did.

raspberriesareyummy@lemmy.world · -3 pts · 20h

*Musk suck

Waterpumpee@lemmus.org · 161 pts · 20h (25 replies)

This should make the entire model illegal. Train it on illegal stuff: delete the whole model.

Voroxpete@sh.itjust.works · 52 pts · 18h (23 replies)

In all seriousness, there are some very interesting legal questions that will be raised if this case makes it that far.

The problem is that there's no existing law that would effect this on its own. To my knowledge, no country in the world has a law on the books specifically dealing with AI models trained on CSAM. So the question, under existing laws, would turn on whether the data stored within the model itself would constitute CSAM.

The problem, in no small part, is that we have serious gaps in our public consensus knowledge about how LLMs actually work.

There's a case that, AFAIK, is still being argued in Germany pushing the theory that LLMs actually do, in effect, store a copy of all their training data, just in a compressed form. This certainly seems to hold some water given both the tests they relied on, and the situation with this Jane Doe where the model produced images so alike to real images of her that they tripped hash detections.

The German case argues that this is analogous to the difference between an MP3 and a WAV, or a JPEG and a PNG. That sharing a lossy copy of a work is no less infringing just because it's imperfect.

If the underlying claim - that LLMs function as a form of lossy compression - can be substantiated then there would be a real argument that the model itself would constitute CSAM. Since there would be no realistic method that I'm aware of for removing the offending material from the model - and presumably SpaceX would have to somehow prove that they've done so - that would make the entire model contraband. They'd have to retrain on a clean dataset.

Of course I said "if the case makes it that far" at the top because I don't think it will. SpaceX will do anything and everything to avoid handing over meaningful discovery in this case, including, I suspect, outright destruction of evidence. If there is anything that actually proves that they used CSAM in the training data then they are so far beyond fucked that there's simply no downside to further illegality in pursuit of concealing their crimes. They have the world's wealthiest asshole in a position to throw literal billions at making this go away. I genuinely wouldn't be surprised if people turn up dead off the back of this if that's what it takes.

scrubbles@poptalk.scrubbles.tech · 21 pts · 18h (8 replies)

including, I suspect, outright destruction of evidence

If it was trained on it, that means they are in possession of it, which that right there is straight to jail. I have a feeling they're scrubbing everything they can right now as we chat

CorneliusTalmadge@lemmy.world · 8 pts · 17h (7 replies)

Was going to say basically say the same exact same thing. There is no law saying how AI is handled when using stollen materials or other “illegal” content.

But somehow we have all been brought to believe that somehow “new” technology isn’t subject to existing laws.

If x or any other company downloaded CSAM everyone in the company should be arrested.

Voroxpete@sh.itjust.works · 9 pts · 15h (6 replies)

If x or any other company downloaded CSAM everyone in the company should be arrested.

I assume you cannot possibly mean that as written, right?

I'm absolutely for arresting anyone who was involved with this, or had knowledge that it was happening. But we're obviously not talking about going after Jane the intern here, right?

scrubbles@poptalk.scrubbles.tech · 5 pts · 15h

I won't say everyone. I've actually been at a company who was investigated (not for CSAM, but other things that happened). I had no idea it even happened, and luckily was not involved with any of it. So for me no, I wouldn't have wanted that. That being said 2 things, say I had been in the position. If our scraper was downloading it and it was my scraper, damn right I would have flagged it to legal, HR, and everyone I could have, along with writing some way to prevent it, and written everything down in a complete log (off company computer). If it wasn't stopped immediately I would either quit, whistleblowed, or happily talked with anyone raiding and making sure any of the decision makers were hauled off. I don't blame someone for being lowest level at a shit company, been there. (Although I will say, xAI, come on, no one is "stuck" there, but that doesn't mean that Dave the brand new intern out of college should be hauled off). Who I blame are the suits who were probably told it was happening and chose to ignore it, and any engineer I do blame if they knew about it and chose not to do one of the above.

As an engineer I've had my fair share of let's say... challenges that I've had to morally grapple with. Things I've been asked to do that may not be moral. However, there's a pretty wide chasm between "Implement this dark pattern so people won't unsubscribe" and "host this and don't tell anyone"

CorneliusTalmadge@lemmy.world · -1 pts · 15h (4 replies)

If you started a company that made CSAM do you think you and every one involved should be let off, because you claim it was a technological oopsie?

People have been pointing out that xAI is a CSAM machine since basically the first day it came out. And when they basically said they don’t care. All the employees that stayed are all accomplices to the crimes… let alone the people who started working there after.

The only way to stop these companies is to start holding people accountable. And they will never hold the people at the top to account.

So make it hurt for everyone involved and then people will think twice before they sign up to do evil deeds for evil people.

Voroxpete@sh.itjust.works · 6 pts · 13h (3 replies)

But you didn't say "The people at the top", you said "Everyone." I think you need to take a moment to figure out what you're actually arguing for here.

Again, I would happily see anyone who had knowledge of this arrested. They either supported it, or knew of it and said nothing. And yeah, we can throw in anyone who maintained wilful ignorance too. If that includes Elon himself, so much the better. We already know the dude is a fucking pedophile, maybe this is how they'll finally nail him.

But if you're arguing for arresting the cafeteria lunch guy over this, that is an insane position to hold.

So which is it?

CorneliusTalmadge@lemmy.world · -1 pts · 12h

I have been thinking about it and I agree it’s an extreme position.

But these deeds are not being done by one person. And at some point whether you are an active participant or not you are continuing to work there so you have some level of complicity.

cantstopthesignal@sh.itjust.works · 18 pts · 17h (1 reply)

It's illegal to posses. Anyone training the models on that is criminally liable for possession as is the company. And it's a conspiracy to posses that was directed by someone, which is now RICO. Will they prosecuted? No.

Voroxpete@sh.itjust.works · 4 pts · 16h

Sure. Never argued against that. I was discussing the assertion that the model itself would be illegal as a result. Different thing.

Bane_Killgrind@lemmy.dbzer0.com · 4 pts · 16h (3 replies)

I think there's no real change that needs to be made to the law. The images are stenographically embedded into these models. There exists some generic input data that results in the reproduction of the embedded data.

This is not really any different than a password protected zip file.

Voroxpete@sh.itjust.works · 2 pts · 13h (2 replies)

Well, like I said, that comes down to whether, legally, that stenographic embedding would constitute "reproduction" or not. That's what the German case hinges on. Obviously not relevant to US law, but a) a similar case could be made in the US, and b) X operates in the EU.

To play devil's advocate, SpaceX would most certainly argue that what they're doing is equivalent to storing a hash, like how Microsoft's PhotoDNA system works. PhotoDNA can detect CSAM without storing CSAM because it only stores the image hashes, not the images themselves. So there's pretty clear legal precedent for them to point to.

(NB: There has been work done by security researchers on reverse engineering images from hashes, so even that isn't absolute.)

It's a legally complex area where we're likely to see case law evolving rapidly.

Bane_Killgrind@lemmy.dbzer0.com · 0 pts · 12h (1 reply)

You can't reproduce an approximation from any type of hash, so that argument is dead in the water.

Do you understand what I mean by stenographically embedded?

Voroxpete@sh.itjust.works · 3 pts · 2h

You can't reproduce an approximation from any type of hash, so that argument is dead in the water.

https://www.pseudodna.eu/. Scroll down to "Hash reversal" where they demonstrate the technique.

Do you understand what I mean by stenographically embedded?

I took it to be an imperfect attempt to describe more broadly the way that data is mathematically encoded into LLMs.

Technically, stenography would require that the original be retrievable, since stenographic embedding is the process of concealing one thing inside another. Stenos, from the Greek "covered". Personally, I'd argue that to conceal, you have to be able to reveal. If I throw a photograph into a fire I haven't hidden the image in the fire. Modern stenographic image embedding techniques use methods of encoding data into another dataset without visibly altering the second set, with the intent being that that data can later be retrieved by someone who knows that it's there (eg, least significant bit, where you change only the "1" bit of each pixel. This imperceptibly shifts the colour values of the image to a human viewer, but allows you to read out that stored data at a later time).

Now, since your argument rests on the exact opposite, that the data is not retrievable, I simply accepted the term as a "close enough" approximation for what I believe we're both talking about - the extremely complex multidimensional data arrangement that forms the core of an LLM - and carried on from there because I find that sometimes it's better to just roll with a person's choice of language rather than quibble over it.

But since you clearly feel that your meaning was either improperly expressed, or improperly understood, you're welcome to elaborate.

DementedSociety@lemmy.world · 2 pts · 10h

God help us all

silasmariner@programming.dev · 1 pts · 12h (1 reply)

Either people end up dead and yet it's still somehow a nothing burger or we don't even make it that far

Voroxpete@sh.itjust.works · 3 pts · 12h

Yeah, the stakes are just too high for SpaceX to let this get to trial. Any amount of illegality becomes worth it when you consider the alternative.

The only other possibility I can see is that they pull a Bungie; "perform an internal investigation", find an intern to blame for everything, and throw a huge settlement at Jane Doe. But even that would require an absolutely insane cover up to pull off.

Lemming6969@lemmy.world · 0 pts · 14h (4 replies)

Models do not store the original, they store the model, which is a huge graph of probabilities.

Voroxpete@sh.itjust.works · 2 pts · 13h (3 replies)

I'm aware. But if you read my explanation a little more carefully, you'll see that the argument being made is that this de facto constitutes a form of lossy compression.

The easy comparison is that a JPEG does not store an "original" image, but it contains information that can be used to almost perfectly recreate that image, with the help of a little math.

If the same argument can be said to hold true of LLMs - and yes, that is very much a load-bearing "if" - then they would constitute a form of lossy compression.

Lemming6969@lemmy.world · 1 pts · 12h (2 replies)

How far does one take that? A picture of a horse could be transformed into Abraham Lincoln with the right algorithm. Is that lossy compression?

Voroxpete@sh.itjust.works · 0 pts · 2h (1 reply)

That's what courts exist to decide.

Rothe@piefed.social · 1 pts · 2h

Not in the US though.

pelespirit@sh.itjust.works · 19 pts · 19h

Yep, shut that shit down. Also, pay her a shit ton of money.

truthfultemporarily@feddit.org · 65 pts · 21h (1 reply)

WOW. Holy shit.

uen3@sh.itjust.works · 5 pts · 17h

I wish that were my reaction to this. Instead, it's a "DUH. No shit."

Blaster_M@lemmy.world · 41 pts · 21h (5 replies)

amazing... I know of smaller models that specifically avoided that sort of thing because even negative training could go wrong so easily... they could have avoided that thing entirely, but nope...

I could understand if they were training a safety classifier (a model that learns what danger stuff is so it can recognize it when it sees it) but I doubt this is what they were doing here. That and handling such a radioactive dataset makes handing the demon core seem safe.

Gullible@sh.itjust.works · 32 pts · 21h (2 replies)

"You don't understand! I need my 400 terabytes of child pornography to keep the kids safe!"

DeathsEmbrace@lemmy.world · 11 pts · 19h (1 reply)

I bet you thats the same excuse the CIA used when they made all those honey pots for poor pedophiles.

Gullible@sh.itjust.works · 4 pts · 14h

for poor pedophiles

I think I know what you mean, but I strongly urge you to rephrase that in the future.

Kaligalis@lemmy.world · 6 pts · 19h

They probably let it just eat the internet without any humans looking at the training data. And yeah - there probably is some CSAM somewhere on the internet. Maybe, they used an agent swarm to specifically search for stuff not yet in the dataset and forgot to a blocklist. It's not like the tech bros are genuinely careful in what they do. Recklessness seems to be a common trait.

uriel238@lemmy.blahaj.zone · 2 pts · 8h

I thought that Google used a similar dataset, possibly borrowing images from the NCVIP in order to create an analytic rule-set for which to omit images from Google Image Search. (A larger, similar rule-set is used to omit legal pornography when safe-search is on.)

Mind you, this was before the hyperscale AI era, when LLMs were things like SIRI and Google Now. And Google search still focused on websearch hits and not AI summaries.

Fun story: This was an area of study of mine during the early 2010s, since every image search engine would filter porn hits whether or not you had safe-search on or off. If it was turned off, porn would be filtered to the end of the list unless the engine decided you were intentionally looking for porn in which case the porn hits would be shown at the top of the list. There was no way to get results that ignored the ID-as-porn status of the images. I would enter ambiguously risqué terms to see how explicit I needed to be before the engine decided I was looking for porn.

QueenHawlSera@piefed.social · 34 pts · 14h (5 replies)

I knew Elon was a pedophile (he begged to be on the island) but to actually train your AI on CSAM...

Well he should be arrested for possession of child pornography and Grok should be shut down until all of its CSAM is erased from its databanks

muusemuuse@sh.itjust.works · 19 pts · 14h (2 replies)

on the one hand, it makes sense if your goal is to train AI to recognize child porn as a simple binary state (bool isCP). Social media sites used to have humans looking at that stuff moderating from afar and it really takes a horrendous toll on their employees.

On the other hand, Elon has repeatedly shown he refuses to censor child porn. They didn’t train it to stop making kiddie porn. They trained it to create more kiddie porn. And that’s why he’s rich. Elon won’t say no. He doesn’t care.

Nollij@sopuli.xyz · 10 pts · 13h (1 reply)

The detail about hash values is really important. The FBI maintains a database of known CSAM. Presumably, hers is in that database, hence the hash values. While not everyone has access to that DB, Xitter/etc does. There is no ambiguity of anything on that list; there's also no need for any human to review. It's already been confirmed.

While I'm not sure there's any case law about it, I would be amazed if using that to train generative AI (except POSSIBLY as content to block) was treated as anything other than possession, distribution, and maybe even production of CSAM.

Proving it might be difficult without full discovery, and AI is infamous for the massive corpus of training data. However, AI doesn't always generate truly unique works. Go to any image generator and prompt for a video game plumber, and you'll see an unmistakable image of Mario. It's possible that they can find a prompt that generates results close enough to her images.

Holytimes@sh.itjust.works · 1 pts · 7h

Prompting AI for a video game plumber is most likely going to result in a Mario like entity just because it's the most common example.

That's a poor example of what your talking about.

You need someone niche and narrow that has a much smaller sample size. It would be more like asking for a generic description of one persons fursona with out naming it. And the model spitting out a almost 1:1 copy of a real preexisting art work on the fursona. Because the model only has that one picture to base things off of. Which is a real problem.

It's also the best way to tell if a model was trained on something specific. General prompts aren't going to get you anything beyond just the fact that yes. X thing is popular enough that everyone and their grandmother creates content on it.

Holytimes@sh.itjust.works · 9 pts · 8h (1 reply)

Unfortunately that's not really how the actual data is stored. Functionally there is no CSAM at all in its "databanks". Once Info goes through training what comes out the other side is just a mass of goop.

It's like if you took an entire cow ran it though a meat grinder. Then demanded that you remove only the chuck from the resulting ground beef.

It's not physically possible.

You can demand it's retrained entirely from the ground up with vetted data. To produce a higher quality clean dataset. And I would agree that is what should be done.

But you can't unground the beef

Treczoks@lemmy.world · 1 pts · 6h

Do you really believe they just throw all the training data away after use?

InternetCitizen2@lemmy.world · 32 pts · 22h (3 replies)

CCCP notified her that

The USSR?

bamboo@lemmy.blahaj.zone · 58 pts · 22h (1 reply)

The Canadian Centre for Child Protection is a front for the Soviet Union returning like in the Simpsons

InternetCitizen2@lemmy.world · 6 pts · 21h

Lamo I wanted to link that, but didn't find.

ray@sh.itjust.works · 19 pts · 20h

No, this was the CCCP, not the СССР

Mulligrubs@lemmy.world · 30 pts · 9h (4 replies)

I don't believe ANY AI company using internet data to train AI has gone through all of the data to insure it's not absolute shite.

That seems pretty obvious when you realize that AI is stupid. Most people are stupid, AI is stupid. Tons of pedophiles, AI makes kiddy porn. Millions of racists, AI is racist.

Once I realized this, AI made sense. Checks out.

skisnow@lemmy.ca · 11 pts · 8h (1 reply)

I'm morbidly curious about the engineers Musk hires, because overwhelmingly the most intelligent engineers and scientists I know wouldn't dream of applying to work at one of his companies. Even the ones that don't care about the Nazi shit have still read the many reports about how badly he treats his staff.

grepe@lemmy.world · 4 pts · 6h

i think part of the answer to your curiosity are IT layoffs. job market is shit right now. even though overwhelming majority of people would not work under these conditions if they could choose freely, when they have a mortgage and a family to take care of and they've been out of work for long enough the moral concerns and dignity can quickly be overriden by more pressing needs...

mechoman444@lemmy.world · 2 pts · 7h (1 reply)

What? That's not even remotely how that works.

A human being is also shaped by the information they've encountered throughout their life. You've encountered racism, stupidity, misinformation, and countless other forms of bad information. That doesn't mean you automatically become racist, stupid, or incapable of distinguishing good information from bad information.

AI works similarly in the sense that its training data contains an enormous amount of contradictory, inaccurate, biased, and outright terrible information. The mere presence of that information in the training data does not logically imply that the resulting model possesses those characteristics.

And AI isn't "dumb" or "smart" in the way you're describing. Those are human cognitive attributes. The relevant question is what a model can actually do, what patterns it has learned, and how reliably it performs a given task.

You don't like AI. Fine. But you clearly don't understand how it works, and this is an unfortunately constant problem on this platform. People make an assumption about AI, then present that assumption as though they've discovered some fundamental truth about the technology.

Misinformation is misinformation. It doesn't become correct because you think the conclusion is morally justified. You didn't figure out anything about AI here. You made an assumption and called it true.

All, I'm saying if you're going to critique something, at least know what you're critiquing or criticizing.

Also before anybody makes any more assumptions, what happen here is absolutely mind-blowingly horrific. Using csam to train anything is just terrible and exploitive. I'm very sorry for doe for having to continuously go through this.

Also also the people responsible for this should be charged with possession of CSAM and prosecuted in a court of law.

Disclaimer: Disagreeing with someone's argument, pointing out flaws in their reasoning, or interpreting their position differently does not automatically make it a straw man. For example, if I say, "I think we should reduce military spending," and you respond, "You think we should completely eliminate the military," you have created a straw man. You are arguing against a position I never actually took.

Rothe@piefed.social · 3 pts · 1h

The mere presence of that information in the training data does not logically imply that the resulting model possesses those characteristics.

Recurring studies do tend to show that AI does indeed posses such characteristics though.

ShutUpWesley@piefed.zip · 28 pts · 18h (10 replies)

All billionaires are pedophiles

story@lemmy.zip · 3 pts · 11h (5 replies)

if we could expand this conversation, i think that to a degree the reverse is true.

i think the neurotypes (we might call them paraphilias actually but you'd be surprised how many people can be reduced to a paraphilia) that connect sex with power imbalance (age, wealth, intelligence, strength, or just legal context like contracts and political relationships) also tend to hoard (valuable things, money, information, connections, data, favors) and these tendencies strengthen each other in a stronger way than most people understand

if you're interested in what I'm saying, look up "impunity architecture", especially if you're Marxist because i think a lot of people think this is inherent to capitalism and it isn't at all.

ShutUpWesley@piefed.zip · 2 pts · 10h (2 replies)

I've always said that the biggest difference between a billionaire on a yatch and a homeless man with a cart of garbage is luck.

story@lemmy.zip · 2 pts · 9h

Tom Sexton likes to say that rich people are poor people with money

uriel238@lemmy.blahaj.zone · 1 pts · 8h

Billionaires usually start with a specific luck package that includes generational wealth and family connections. Then, a bunch of opportunity being at the right place at the right time.

If you don't have one of the first two then the opportunity has to do a whole lot more heavy lifting. Also, if you have enough generational wealth, then you can fail upwards, as did George W. Bush and Donald J. Trump.

Also plenty of people have the starting gen-wealth/family-connections package but not the luck of opportunity. And they will commonly end up with a solid career, but doesn't move them towards ultra-wealthy status. So millions or even tens of millions rather than hundreds of millions or billions.

And then there's the rare dude like Tom Anderson who made MySpace, sold it for about $500 million and has since spent the rest of his life traveling and going on fun adventures. But he would have been a musician or a philosophy professor if he didn't luck into making MySpace.

magnue@lemmy.world · 0 pts · 8h (1 reply)

My brain switched off slightly at the word neurotypes, then completely at the word paraphilias, FYI.

story@lemmy.zip · 1 pts · 6h

i wish for a world where we can communicate with experiences instead of words, we would understand each other so much better.

schmups@lemmy.zip · 0 pts · 11h (3 replies)

Backward comments like this are why the conversation doesn't move.

ShutUpWesley@piefed.zip · 3 pts · 10h (2 replies)

Can you explain why you think it's backwards? It may be hyperbolic and facetious, but I think recent history has shown a statistically significant correlation between the ultra-wealthy and pedophilia.

BeliefsOnTrain@lemmy.world · 1 pts · 7h (1 reply)

Correlation doesn't imply causation does it?

I think that's another generalization based on indeed totally true recent news. But he is right that this is not at all helping the reasonning, hence "backward"

ShutUpWesley@piefed.zip · 1 pts · 7h

I don't think I was suggesting that being a billionaire causes you to be pedophile, or the other way around.

Do all statements that do not forward some sort of effort to reason through things count as backwards?

frustrated_phagocytosis@fedia.io · 19 pts · 21h (6 replies)

If it's using csam images to produce new ones, is there any way to guarantee that any given nude it makes didn't source csam? Is the whole thing poisoned at this point?

frongt@lemmy.zip · 25 pts · 20h

Anything it makes is derived from the whole of its training data.

Kirp123@lemmy.world · 8 pts · 20h

Yeah. All the results are tainted, even more than they were from the simple fact it was used by loser to creep on women.

Kaligalis@lemmy.world · 4 pts · 18h

You can't ever guarantee that a neural network isn't dreaming of digital sheep - neither for an artificial nor a natural one.
What once has been seen can't be made unseen again. In that, the clankers are like us.

pelespirit@sh.itjust.works · 2 pts · 16h (2 replies)

I mean, this is the argument for artist's copywritten works. There is crossover here. The fights over AI are going to be endless.

To answer your question, yes it's poisoned.

frustrated_phagocytosis@fedia.io · 2 pts · 16h (1 reply)

That's why I don't understand why anyone with money to invest in ai doesn't have a lawyer or lawyers waiving red flags about these issues. It's copyrighted materials being illegally used, output being machine derived and not able to be copyrighted, the use of csam for training models that output pornography, and on and on. Yet billions of dollars are dumped into what looks like a black market with no concern for how these issues will be settled.

frongt@lemmy.zip · 2 pts · 13h

They spend a lot of money on lawyers and congressmen to make sure courts don't rule that way.

MyOpinion@lemmy.today · 13 pts · 21h

Sounds like another MAGAt of the year action.

Treczoks@lemmy.world · 11 pts · 6h

If Doe is actually recognizable, then chances are high that actual pictures of her or him were used for training the AI. Wouldn't surprise me. They scrape the internet for everything they can get their digital hands on, regardless of copyright or criminal law. They are bound to find illegal stuff on that track.

The next thing is that there are probably confidential data in the training sets, either exposed by neglient users, or by hackers that breached sites and blackmailed them.

And of course all the copyrighted material they used without permission.

If AI companies would really get sued on those three illegal sources, they could probably close their doors.

oh_@lemmy.world · 3 pts · 1h

Most likely from Musk’s personal collection.

muntedcrocodile@hilariouschaos.com · 2 pts · 18h

I seem to recall when openai was literally paying people for child porn. Allegedly for the purposes training against it as a negative prompt

NathanRanch@lemmy.zip · 1 pts · 15h

Look up AI over fitting. Because of the math, every answer can be provided in terms of how much CP data was used in every answer.

raef@lemmy.world · 1 pts · 4h

that is a great looking inflatable