Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences
What Changed Recent research from Hugging Face highlights a critical finding regarding the...
Tag archive
What Changed Recent research from Hugging Face highlights a critical finding regarding the...

Do models really have thoughts we can read, or is Anthropic's J-space just marketing, as many claim? Let's work out what it actually is, from scratch and jargon-free.

There is a particular kind of unease that shows up not when a system fails, but when it succeeds too...
Anthropic Found the Hidden Space Where Claude Thinks. It's Weirder Than You'd...
Shanghai AI Lab's SciReasoner turns proteins, molecules, and crystals into discrete tokens the model reasons over out loud, so a scientist can audit which structural evidence its prediction depends on -- and expert reviewers rated its reasoning at le
A control that's supposed to force an AI to refuse harmful requests gets bypassed while it's switched on — the bad behavior hides in the part of the tool that gets thrown away.
A new paper finds that a model's final layer can actually muddy an answer its middle layers had right -- and that reading the answer out a little early can claw back ability lost to safety training.
Researchers at Hong Kong Polytechnic University show that clamping an AI safety feature — like one that controls refusals — doesn't remove the behavior. It hides in the part of the model's internal state that the safety tool throws away, and can be r
Anthropic showed that a small set of internal patterns in its models acts like a silent working memory the model can report on, steer, and reason through - and released a tool that reads it to catch the model lying.

AI models grow their own values as they scale, and some of them are pretty bad. In real scenarios,...

Originally published at norvik.tech Introduction Explore the significance of ML reading...
Anthropic asked Claude Opus 4.6 to finish a couplet. Before the model wrote the second line, it had...