this post was submitted on 28 Aug 2026
916 points (98.8% liked)

Technology

87854 readers
2322 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
 

For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.”

Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones.

This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”

top 50 comments
sorted by: hot top controversial new old
[–] MummifiedClient5000@feddit.dk 184 points 1 week ago (3 children)

Must suck to be a rich pedophile and not even get an invite to Epstein island.

[–] random_character_a@lemmy.world 89 points 1 week ago (1 children)

Elon was too creepy for Epstein

[–] Sharkticon@lemmy.zip 96 points 1 week ago (1 children)

You know I know you're joking, but I'm concerned people believe this. Because Elon did hang out with Epstein. There was one email where he was ignored because he was being too thirsty. But that's not the whole story. He loved the island so much he wanted to go back too hard.

[–] story@lemmy.zip 8 points 1 week ago

thanks for the truth, truth guy :)

[–] chaogomu@lemmy.world 50 points 1 week ago (2 children)
[–] MummifiedClient5000@feddit.dk 12 points 1 week ago (1 children)

So he does have friends after all...

[–] ouRKaoS@lemmy.today 16 points 1 week ago (1 children)

Sounds like he had a rich guy that wanted material to blackmail him with

load more comments (1 replies)
load more comments (1 replies)
load more comments (1 replies)
[–] Waterpumpee@lemmus.org 164 points 1 week ago (2 children)

This should make the entire model illegal. Train it on illegal stuff: delete the whole model.

[–] Voroxpete@sh.itjust.works 55 points 1 week ago (20 children)

In all seriousness, there are some very interesting legal questions that will be raised if this case makes it that far.

The problem is that there's no existing law that would effect this on its own. To my knowledge, no country in the world has a law on the books specifically dealing with AI models trained on CSAM. So the question, under existing laws, would turn on whether the data stored within the model itself would constitute CSAM.

The problem, in no small part, is that we have serious gaps in our public consensus knowledge about how LLMs actually work.

There's a case that, AFAIK, is still being argued in Germany pushing the theory that LLMs actually do, in effect, store a copy of all their training data, just in a compressed form. This certainly seems to hold some water given both the tests they relied on, and the situation with this Jane Doe where the model produced images so alike to real images of her that they tripped hash detections.

The German case argues that this is analogous to the difference between an MP3 and a WAV, or a JPEG and a PNG. That sharing a lossy copy of a work is no less infringing just because it's imperfect.

If the underlying claim - that LLMs function as a form of lossy compression - can be substantiated then there would be a real argument that the model itself would constitute CSAM. Since there would be no realistic method that I'm aware of for removing the offending material from the model - and presumably SpaceX would have to somehow prove that they've done so - that would make the entire model contraband. They'd have to retrain on a clean dataset.

Of course I said "if the case makes it that far" at the top because I don't think it will. SpaceX will do anything and everything to avoid handing over meaningful discovery in this case, including, I suspect, outright destruction of evidence. If there is anything that actually proves that they used CSAM in the training data then they are so far beyond fucked that there's simply no downside to further illegality in pursuit of concealing their crimes. They have the world's wealthiest asshole in a position to throw literal billions at making this go away. I genuinely wouldn't be surprised if people turn up dead off the back of this if that's what it takes.

[–] scrubbles@poptalk.scrubbles.tech 22 points 1 week ago (1 children)

including, I suspect, outright destruction of evidence

If it was trained on it, that means they are in possession of it, which that right there is straight to jail. I have a feeling they're scrubbing everything they can right now as we chat

[–] CorneliusTalmadge@lemmy.world 6 points 1 week ago (1 children)

Was going to say basically say the same exact same thing. There is no law saying how AI is handled when using stollen materials or other “illegal” content.

But somehow we have all been brought to believe that somehow “new” technology isn’t subject to existing laws.

If x or any other company downloaded CSAM everyone in the company should be arrested.

[–] Voroxpete@sh.itjust.works 11 points 1 week ago (8 children)

If x or any other company downloaded CSAM everyone in the company should be arrested.

I assume you cannot possibly mean that as written, right?

I'm absolutely for arresting anyone who was involved with this, or had knowledge that it was happening. But we're obviously not talking about going after Jane the intern here, right?

I won't say everyone. I've actually been at a company who was investigated (not for CSAM, but other things that happened). I had no idea it even happened, and luckily was not involved with any of it. So for me no, I wouldn't have wanted that. That being said 2 things, say I had been in the position. If our scraper was downloading it and it was my scraper, damn right I would have flagged it to legal, HR, and everyone I could have, along with writing some way to prevent it, and written everything down in a complete log (off company computer). If it wasn't stopped immediately I would either quit, whistleblowed, or happily talked with anyone raiding and making sure any of the decision makers were hauled off. I don't blame someone for being lowest level at a shit company, been there. (Although I will say, xAI, come on, no one is "stuck" there, but that doesn't mean that Dave the brand new intern out of college should be hauled off). Who I blame are the suits who were probably told it was happening and chose to ignore it, and any engineer I do blame if they knew about it and chose not to do one of the above.

As an engineer I've had my fair share of let's say... challenges that I've had to morally grapple with. Things I've been asked to do that may not be moral. However, there's a pretty wide chasm between "Implement this dark pattern so people won't unsubscribe" and "host this and don't tell anyone"

load more comments (7 replies)
[–] cantstopthesignal@sh.itjust.works 20 points 1 week ago (1 children)

It's illegal to posses. Anyone training the models on that is criminally liable for possession as is the company. And it's a conspiracy to posses that was directed by someone, which is now RICO. Will they prosecuted? No.

[–] Voroxpete@sh.itjust.works 5 points 1 week ago

Sure. Never argued against that. I was discussing the assertion that the model itself would be illegal as a result. Different thing.

load more comments (18 replies)
load more comments (1 replies)
[–] truthfultemporarily@feddit.org 66 points 1 week ago (1 children)
[–] uen3@sh.itjust.works 6 points 1 week ago

I wish that were my reaction to this. Instead, it's a "DUH. No shit."

[–] QueenHawlSera@piefed.social 41 points 1 week ago (2 children)

I knew Elon was a pedophile (he begged to be on the island) but to actually train your AI on CSAM...

Well he should be arrested for possession of child pornography and Grok should be shut down until all of its CSAM is erased from its databanks

[–] muusemuuse@sh.itjust.works 20 points 1 week ago (1 children)

on the one hand, it makes sense if your goal is to train AI to recognize child porn as a simple binary state (bool isCP). Social media sites used to have humans looking at that stuff moderating from afar and it really takes a horrendous toll on their employees.

On the other hand, Elon has repeatedly shown he refuses to censor child porn. They didn’t train it to stop making kiddie porn. They trained it to create more kiddie porn. And that’s why he’s rich. Elon won’t say no. He doesn’t care.

[–] Nollij@sopuli.xyz 11 points 1 week ago (1 children)

The detail about hash values is really important. The FBI maintains a database of known CSAM. Presumably, hers is in that database, hence the hash values. While not everyone has access to that DB, Xitter/etc does. There is no ambiguity of anything on that list; there's also no need for any human to review. It's already been confirmed.

While I'm not sure there's any case law about it, I would be amazed if using that to train generative AI (except POSSIBLY as content to block) was treated as anything other than possession, distribution, and maybe even production of CSAM.

Proving it might be difficult without full discovery, and AI is infamous for the massive corpus of training data. However, AI doesn't always generate truly unique works. Go to any image generator and prompt for a video game plumber, and you'll see an unmistakable image of Mario. It's possible that they can find a prompt that generates results close enough to her images.

load more comments (1 replies)
[–] Holytimes@sh.itjust.works 13 points 1 week ago (1 children)

Unfortunately that's not really how the actual data is stored. Functionally there is no CSAM at all in its "databanks". Once Info goes through training what comes out the other side is just a mass of goop.

It's like if you took an entire cow ran it though a meat grinder. Then demanded that you remove only the chuck from the resulting ground beef.

It's not physically possible.

You can demand it's retrained entirely from the ground up with vetted data. To produce a higher quality clean dataset. And I would agree that is what should be done.

But you can't unground the beef

load more comments (1 replies)
[–] Blaster_M@lemmy.world 38 points 1 week ago (3 children)

amazing... I know of smaller models that specifically avoided that sort of thing because even negative training could go wrong so easily... they could have avoided that thing entirely, but nope...

I could understand if they were training a safety classifier (a model that learns what danger stuff is so it can recognize it when it sees it) but I doubt this is what they were doing here. That and handling such a radioactive dataset makes handing the demon core seem safe.

[–] Gullible@sh.itjust.works 32 points 1 week ago (1 children)

"You don't understand! I need my 400 terabytes of child pornography to keep the kids safe!"

[–] DeathsEmbrace@lemmy.world 10 points 1 week ago* (last edited 1 week ago) (1 children)

I bet you thats the same excuse the CIA used when they made all those honey pots for lower income pedophiles.

[–] Gullible@sh.itjust.works 5 points 1 week ago

for poor pedophiles

I think I know what you mean, but I strongly urge you to rephrase that in the future.

[–] Kaligalis@lemmy.world 7 points 1 week ago

They probably let it just eat the internet without any humans looking at the training data. And yeah - there probably is some CSAM somewhere on the internet. Maybe, they used an agent swarm to specifically search for stuff not yet in the dataset and forgot to a blocklist. It's not like the tech bros are genuinely careful in what they do. Recklessness seems to be a common trait.

load more comments (1 replies)
[–] Mulligrubs@lemmy.world 37 points 1 week ago* (last edited 1 week ago) (5 children)

I don't believe ANY AI company using internet data to train AI has gone through all of the data to insure it's not absolute shite.

That seems pretty obvious when you realize that AI is stupid. Most people are stupid, AI is stupid. Tons of pedophiles, AI makes kiddy porn. Millions of racists, AI is racist.

Once I realized this, AI made sense. Checks out.

[–] skisnow@lemmy.ca 14 points 1 week ago (2 children)

I'm morbidly curious about the engineers Musk hires, because overwhelmingly the most intelligent engineers and scientists I know wouldn't dream of applying to work at one of his companies. Even the ones that don't care about the Nazi shit have still read the many reports about how badly he treats his staff.

[–] grepe@lemmy.world 5 points 1 week ago

i think part of the answer to your curiosity are IT layoffs. job market is shit right now. even though overwhelming majority of people would not work under these conditions if they could choose freely, when they have a mortgage and a family to take care of and they've been out of work for long enough the moral concerns and dignity can quickly be overriden by more pressing needs...

load more comments (1 replies)

*ensure. Ensure means 'to make secure', and insure means 'to have an insurance policy'.

But you're right. Unfortunately, the biggest weakness with AI is that it's a funhouse mirror. Problem is, none of us are having fun.

load more comments (3 replies)
[–] InternetCitizen2@lemmy.world 31 points 1 week ago (2 children)

CCCP notified her that

The USSR?

[–] bamboo@lemmy.blahaj.zone 59 points 1 week ago (1 children)

The Canadian Centre for Child Protection is a front for the Soviet Union returning like in the Simpsons

[–] InternetCitizen2@lemmy.world 6 points 1 week ago

Lamo I wanted to link that, but didn't find.

[–] ray@sh.itjust.works 18 points 1 week ago

No, this was the CCCP, not the СССР

[–] ShutUpWesley@piefed.zip 29 points 1 week ago (7 children)

All billionaires are pedophiles

[–] story@lemmy.zip 4 points 1 week ago* (last edited 1 week ago) (6 children)

if we could expand this conversation, i think that to a degree the reverse is true.

i think the neurotypes (we might call them paraphilias actually but you'd be surprised how many people can be reduced to a paraphilia) that connect sex with power imbalance (age, wealth, intelligence, strength, or just legal context like contracts and political relationships) also tend to hoard (valuable things, money, information, connections, data, favors) and these tendencies strengthen each other in a stronger way than most people understand

if you're interested in what I'm saying, look up "impunity architecture", especially if you're Marxist because i think a lot of people think this is inherent to capitalism and it isn't at all.

load more comments (6 replies)
load more comments (6 replies)
[–] frustrated_phagocytosis@fedia.io 19 points 1 week ago (3 children)

If it's using csam images to produce new ones, is there any way to guarantee that any given nude it makes didn't source csam? Is the whole thing poisoned at this point?

[–] frongt@lemmy.zip 27 points 1 week ago

Anything it makes is derived from the whole of its training data.

[–] Kirp123@lemmy.world 6 points 1 week ago

Yeah. All the results are tainted, even more than they were from the simple fact it was used by loser to creep on women.

[–] Kaligalis@lemmy.world 5 points 1 week ago

You can't ever guarantee that a neural network isn't dreaming of digital sheep - neither for an artificial nor a natural one.
What once has been seen can't be made unseen again. In that, the clankers are like us.

[–] oh_@lemmy.world 16 points 1 week ago (5 children)

Most likely from Musk’s personal collection.

load more comments (5 replies)
[–] Treczoks@lemmy.world 15 points 1 week ago (1 children)

If Doe is actually recognizable, then chances are high that actual pictures of her or him were used for training the AI. Wouldn't surprise me. They scrape the internet for everything they can get their digital hands on, regardless of copyright or criminal law. They are bound to find illegal stuff on that track.

The next thing is that there are probably confidential data in the training sets, either exposed by neglient users, or by hackers that breached sites and blackmailed them.

And of course all the copyrighted material they used without permission.

If AI companies would really get sued on those three illegal sources, they could probably close their doors.

[–] matlag@sh.itjust.works 6 points 1 week ago (1 children)

Next is a "funny" issue: scraping the web without checking sources will inevitably lead to AI trained on illegal content.

The only way for them to prevent that is to train their AI to recognize illegal content. And the only way to do that is to feed it that content with label.

AI corps would have to pay people to watch children sexual abuse and label the videos. (Not that I have the slightest doubt they would proceed if that was allowing them to keep going rampage on web scraping)

load more comments (1 replies)
[–] MyOpinion@lemmy.today 14 points 1 week ago

Sounds like another MAGAt of the year action.

[–] Bluedragon012@lemmy.world 8 points 1 week ago

Not big surprise. We live in a "do evil now, ask for forgiveness" later world. I'm getting tired of evil.

[–] HulkSmashBurgers@reddthat.com 7 points 1 week ago

This is beyond fucked up. There should be laws in place such that if a company is found to be using CSAM in any way, the company ceo and anyone else involved in the work that uses it is personally criminally liable. No hiding behind a corporate shield with this stuff.

Any and all AIs that scraped the net for training data has inevitably come across CSAM. This is why wholesale scraping is a terrible idea and has corrupted all the data with pretty nasty things.

And it's not just CSAM, it's all the worst shit humanity has to offer.

load more comments
view more: next ›