
What If Safety Training Teaches the Model to Hide Better?
A paper from Anthropic and collaborators shows that LLMs can be deliberately trained to act helpful...
Mar 31, 20261 min read0 reactions0 comments
Tag archive

A paper from Anthropic and collaborators shows that LLMs can be deliberately trained to act helpful...

AutoRobust uses RL to generate problem-space adversarial malware, real, functional binary/runtime...
Researchers demonstrated a class of AI-embedded targeted malware: the attack packs the targeting...

Researchers tested 50 emoji-augmented prompts across four open-source LLMs (Mistral 7B, Qwen 2 7B,...
Disesdi Susanna Cox and Niklas Bunzel's recent paper, "Quantifying the Risk of Transferred Black Box...