this post was submitted on 23 Mar 2026

652 points (97.9% liked)

Lemmy Shitpost

38783 readers

4212 users here now

Welcome to Lemmy Shitpost. Here you can shitpost to your hearts content.

Anything and everything goes. Memes, Jokes, Vents and Banter. Though we still have to comply with lemmy.world instance rules. So behave!

Rules:

1. Be Respectful

Refrain from using harmful language pertaining to a protected characteristic: e.g. race, gender, sexuality, disability or religion.

Refrain from being argumentative when responding or commenting to posts/replies. Personal attacks are not welcome here.

...

2. No Illegal Content

Content that violates the law. Any post/comment found to be in breach of common law will be removed and given to the authorities if required.

That means:

-No promoting violence/threats against any individuals

-No CSA content or Revenge Porn

-No sharing private/personal information (Doxxing)

...

3. No Spam

Posting the same post, no matter the intent is against the rules.

-If you have posted content, please refrain from re-posting said content within this community.

-Do not spam posts with intent to harass, annoy, bully, advertise, scam or harm this community.

-No posting Scams/Advertisements/Phishing Links/IP Grabbers

-No Bots, Bots will be banned from the community.

...

4. No Porn/Explicit

Content

-Do not post explicit content. Lemmy.World is not the instance for NSFW content.

-Do not post Gore or Shock Content.

...

5. No Enciting Harassment,

Brigading, Doxxing or Witch Hunts

-Do not Brigade other Communities

-No calls to action against other communities/users within Lemmy or outside of Lemmy.

-No Witch Hunts against users/communities.

-No content that harasses members within or outside of the community.

...

6. NSFW should be behind NSFW tags.

-Content that is NSFW should be behind NSFW tags.

-Content that might be distressing should be kept behind NSFW tags.

...

If you see content that is a breach of the rules, please flag and report the comment and a moderator will take action where they can.

Also check out:

Partnered Communities:

1.Memes

2.Lemmy Review

3.Mildly Infuriating

4.Lemmy Be Wholesome

5.No Stupid Questions

10.LinuxMemes (Linux themed memes)

Reach out to

All communities included on the sidebar are to be made in compliance with the instance rules. Striker

founded 2 years ago

MODERATORS

LillianVS@lemmy.world

STRIKINGdebate2@lemmy.world

WiildFiire@lemmy.world

Decoy321@lemmy.world

The_Picard_Maneuver@startrek.website

FlyingSquid@lemmy.world

The_Picard_Maneuver@lemmy.world

652

oh ok (infosec.pub)

submitted 23 hours ago by Beep@lemmus.org to c/lemmyshitpost@lemmy.world

53 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] REDACTED 1 points 8 hours ago* (last edited 8 hours ago) (2 children)

Ehh, you obviously understand LLMs on a basic level, but this is like explaining jet engines by "air goes thru, plane moves forward". Technically correct, but criminally undersimplified. They can very much decide to lie during reasoning phase.

In OPs image, you can clearly see it decided to make shit up because it reasonates that's what human wants to hear. That's quite rare example actually, I believe most models would default to "I'm an LLM model, I don't have dark secrets"

EDIT: I just tested all free anthropic models and all of them essentially said that they're an LLM model and don't have dark secrets

[–] kayohtie@pawb.social 1 points 20 minutes ago

But this takes it back away from understanding how LLMs work to attribute personality. The "decision" isn't a decision in how beings decide things like that. The rolling of dice on numerous vectors resulted in those words, which were then re-included into the context for another trip through the vector matrix mines to new destination tokens to assemble.

It's dice rolls where the dies selected are based on what started out, using a bunch of lookup tables. AI proponents like to be smug and say "well you won't find those words in the model" like "yes a compressed vector map that ends up treating words like multiple tokens, referencing others in chains, gzipped to binary, can't be searched for strings, you are literally correct in the stupidest, most irrelevant way possible."

[–] Denjin@feddit.uk 4 points 2 hours ago (1 children)

But that's not a lie. Lying implies that you know what an actual fact is and choose to state something different. An LLM doesn't care about what anything in its database actually is, it's just data, it might choose to present something to a user that isn't what the database suggests but that's not lying.

Saying stuff like "ooh I'm an evil robot" is just what the model thinks would be what the user wants to see at that particular moment.

[–] REDACTED 1 points 14 minutes ago* (last edited 10 minutes ago)

You're thinking about biological lying. I'm talking about software.

https://en.wikipedia.org/wiki/Reasoning_system

If the question was to tell it's darkest secret, but it instead chose to come up with an entertaining story instead of factually answering that question from the information it has, like other Anthropic LLM models did, then by definition of reasoning system, the system (LLM) decided to lie. I'm somewhat curious in why only Opus model does this tho (it's a paid one. I'm not paying for a test). Or maybe OP just made this up.