Technology

83069 readers

4404 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws

Nvidia CEO Jensen Huang says ‘I think we’ve achieved AGI’ (www.theverge.com)

submitted 2 days ago by return2ozma@lemmy.world to c/technology@lemmy.world

64 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] Technus@lemmy.zip 1 points 20 hours ago

Not sure what this internal state you are referring to is. Are you talking about all the values that come out of each step of the computations?

It would need to be able to form memories like real brains do, by creating new connections between neurons and adjusting their weights in real time in response to stimuli, and having those connections persist. I think that's a prerequisite to models that are capable of higher-level reasoning and understanding. But then you would need to store those changes to the model for each user, which would be tens or hundreds of gigabytes.

These current once-through LLMs don't have time to properly digest what they're looking at, because they essentially forget everything once they output a token. I don't think you can make up for that by spitting some tokens out to a file and reading them back in, because it still has to be human-readable and coherent. That transformation is inherently lossy.

This is basically what I'm talking about:

But for every single token the LLM outputs. The fact that it's allowed to take notes is a mitigation for this context loss, not a silver bullet.