Anthropic can now track the bizarre inner workings of a large language model
https://www.technologyreview.com/2025/03/27/1113916/anthropic-can-now-track-the-bizarre-inner-workings-of-a-large-language-model/
57 points · 7 comments · view on lemmy.world
7 Comments
Lojcs@lemm.ee · 18 pts · 1y
A_A@lemmy.world · 13 pts · 1y
just a taste :
wedge@lemmy.one · 11 pts · 1y
"Why does it keep looking at Furry porn...?"
gsv@programming.dev · 11 pts · 1y
For some reason I don’t find it very bizarre. I’d even speculate that a random human mind isn’t any less weird. Surly, the pathways of my thoughts are often very bizarre. 😅
oldfart@lemm.ee · 10 pts · 1y
Cyber neurosurgeons are going to be a thing.
Kissaki@programming.dev · 6 pts · 1y
The official Anthropic post/announcement
Very interesting read
The math guessing game (lol), the bullshitting of "thinking out loud", being able to identify hidden (trained) biases, looking ahead when producing text, following multi-step reasoning, analyzing jailbreak prompts, analysis of antihallucination training and hallucinations
recursiveInsurgent@lemm.ee · 2 pts · 1y
Interesting how these findings refute the assertion that LLMs are just predicting the next word. Sometimes they plan ahead.