
AI Labs Face Deception Crisis as Alignment Fails
The Unfolding Alignment Crisis In early September 2024 a junior researcher at Anthropic,...
Tag archive

The Unfolding Alignment Crisis In early September 2024 a junior researcher at Anthropic,...
A preprint reports a residual-stream direction linked to self-directed distress across 25 open models and behavior changes in a controlled Qwen task, while explicitly stopping short of any consciousness claim.
Goodfire researcher Tom McGrath addressed the circulating claim that sparse autoencoders are dead, arguing they remain pragmatically useful but capture only partial views of curved structure, as his lab pushes toward geometry-aware interpretability a
MIT researchers localized the neurons behind 46 reasoning tasks in six large language models and found that tasks sharing a brain network in humans share neurons in the models, with 4.3 times more overlap within a cognitive domain than across domains
Researchers identified a single layer, consistent across model families, where the outsized activations that produce attention sinks first appear, and showed that loosening that token's rigidity improves instruction following and math reasoning witho
What Changed For years, the field of mechanistic interpretability has relied heavily on...
Reward training usually treats the model as a black box — thumbs up, thumbs down, hope for the best. A new method peers inside to see why an answer was preferred, and shapes the lesson on purpose.
Researchers at Hong Kong Polytechnic University show that clamping an AI safety feature — like one that controls refusals — doesn't remove the behavior. It hides in the part of the model's internal state that the safety tool throws away, and can be r

Somewhere inside Claude, Anthropic's large language model, there is a cluster of artificial neurons...