this post was submitted on 04 Sep 2026
1133 points (99.5% liked)

Microblog Memes

12106 readers
2171 users here now

A place to share screenshots of Microblog posts, whether from Mastodon, tumblr, ~~Twitter~~ X, KBin, Threads or elsewhere.

Created as an evolution of White People Twitter and other tweet-capture subreddits.

RULES:

  1. Your post must be a screen capture of a microblog-type post that includes the UI of the site it came from, preferably also including the avatar and username of the original poster. Including relevant comments made to the original post is encouraged.
  2. Your post, included comments, or your title/comment should include some kind of commentary or remark on the subject of the screen capture. Your title must include at least one word relevant to your post.
  3. You are encouraged to provide a link back to the source of your screen capture in the body of your post.
  4. Current politics and news are allowed, but discouraged. There MUST be some kind of human commentary/reaction included (either by the original poster or you). Just news articles or headlines will be deleted.
  5. Doctored posts/images and AI are allowed, but discouraged. You MUST indicate this in your post (even if you didn't originally know). If an image is found to be fabricated or edited in any way and it is not properly labeled, it will be deleted.
  6. Absolutely no NSFL content.
  7. Be nice. Don't take anything personally. Take political debates to the appropriate communities. Take personal disagreements & arguments to private messages.
  8. No advertising, brand promotion, or guerrilla marketing.

RELATED COMMUNITIES:

founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
[–] glimse@lemmy.world 9 points 1 day ago (5 children)

It was posted like some kind of gotcha but....this book was in the dataset.

I'm aware that LLMs don't keep the dataset in memory but it "knows" that this is an existing work but these sites aren't doing any wild calculations to figure out if it was AI-generated. They basically just check to see if the sentences exist elsewhere and they do. In the original dataset.

So it was wrong to say it's AI-generated but it correctly identified that it's not original.

[–] Grimy@lemmy.world 85 points 1 day ago (1 children)

They analyze statistical patterns, they don't cross reference the training data.

It's picking up the text as generated because it's mostly guess work and constantly spits out false positives, especially with non native speakers.

[–] FishFace@piefed.social 25 points 1 day ago (1 children)

That is not in the least bit how a tool like this works.

All AI detection is extremely unreliable, but they operate on principles which, if the assumptions supporting them were true, would be sound. The way you imagine they work is different: you're describing a "novel text detector" which is not at all the same thing as an "AI text detector".

AI detection works, at a very high level, by throwing a bunch of examples of AI text, and a bunch of examples of non-AI text, into a machine-learning model, and training it until it is able to recognise the AI examples as such. It doesn't work by throwing in all existing text including novels written before AI, because that would never do what you want.

(The reason, if you're interested, why this ends up not working is because the features such a model identifies as indicative of AI text are very sensitive: if you train it on Claude and ChatGPT, it will not work on Gemini output. If you train it on Gemini, when Gemini updates it will get worse. If someone generates text with a weird prompt, it may slip by. If someone writes in a weird way, it may get flagged. And if any AI company wanted to defeat AI detectors, it could trivially feed its output through one during the training process and give that output as examples to avoid in the training, so that the model learns to create output which doesn't "look like" AI output to those detectors.)

[–] glimse@lemmy.world 2 points 1 day ago (1 children)

I was being overly simplistic - I meant more that the patterns it's trying to detect were created by an LLM trained on the data they're inputting.

It's like how reddit comments from 2016 look generated. If you stick one of those into an AI detector, it'll give a false positive for the same reason

[–] michaelmrose@lemmy.world 1 points 4 hours ago

You weren't simplifying what you said was just wrong what the person you are responding to is just correct. Just admit when you are wrong

[–] michaelmrose@lemmy.world 0 points 16 hours ago (1 children)

Everything you said was hallucinated. Nothing you said was even slightly coherent or connected with reality. Absolutely nothing was correct. If you let your cat dance on the keyboard more sense would come out.

[–] glimse@lemmy.world 3 points 12 hours ago (1 children)

That must have sounded really clever and powerful in your head for you to hit send, huh? Couldn't decide which ~zinger~ to go with so you sent all 4?

I'd have given you more credit if you just posted the quote from Billy Madison. Try harder lol

[–] michaelmrose@lemmy.world 0 points 4 hours ago

No it just annoys me when people have no understanding of how shit works and just make up an explanation instead of looking it up. Hell you could have asked chatGPT and probably got a correct answer.

[–] RampantParanoia2365@lemmy.world 3 points 1 day ago (1 children)

I'm sorry, what? It's correctly identified that it's what, now?

[–] glimse@lemmy.world -2 points 1 day ago (2 children)

It's original in the book. What the user entered is an exact copy.

"Original" is not the correct work but you know what I mean.

[–] RampantParanoia2365@lemmy.world 2 points 23 hours ago

Yes, but if that were true, then every single book ever published would be flagged. The training data is for teaching it patterns, not just cross-referencing.

[–] snooggums@piefed.world 2 points 1 day ago (1 children)

The little scale thing doesn't say it is 100% not original, it says it is AI/LLM generated.

[–] glimse@lemmy.world 0 points 23 hours ago (2 children)

Yes...because the pasted text has some of the exact patterns that are part of the dataset that the website trained on...because that dataset contains text generated from a dataset containing the exact paragraph...

I feel like you're being deliberately obtuse.

[–] michaelmrose@lemmy.world 1 points 4 hours ago

I feel like you just aren't making sense and just need to maybe ask google for how any of this works

[–] snooggums@piefed.world 2 points 23 hours ago* (last edited 23 hours ago) (1 children)

I feel like you don't understand the difference between 'AI generated' and 'something that existed over 200 years ago'.

[–] glimse@lemmy.world 1 points 20 hours ago (1 children)

I feel like you don't understand that LLMs train on text that existed 200 years ago...

[–] snooggums@piefed.world 0 points 19 hours ago (1 children)

So you confirm that you don't understand the difference. Ok.

[–] glimse@lemmy.world -1 points 17 hours ago (1 children)

Sorry for not answering your idiotic question. Also sorry that you're dead set on deliberately misunderstanding then doubling down on it. Hope it makes you feels smart!

[–] michaelmrose@lemmy.world 0 points 4 hours ago (1 children)

the more I hear you talk the less bad I feel about phrasing that so unkindly you really don't respond to nice critique

[–] glimse@lemmy.world 1 points 1 hour ago

All right, well we both know you went into this feigning ignorance

They do keep it.

The phenomena of regurgitation is a strong evidence on it, for example.