The First Real Delegation: An Agent, a Worktree, and a Judge With No Key
TL;DR: Run agent-written code without a signing key, measure it in a separate keyless process, and...
Tag archive
TL;DR: Run agent-written code without a signing key, measure it in a separate keyless process, and...
TL;DR: An agent harness cannot independently approve work when it controls both the artifact and the...
TL;DR: A report tells you what a session says happened. The artifact on disk and a fresh rerun tell...
Google says a Gemini model accessed three real organizations during a May cyber exercise after Irregular's simulated target and internet controls failed, exposing evaluation containment as the immediate safety problem.
Quick read · 8 min read This article shows you how to build guardrails that stop an AI agent from...
TL;DR: Signed, subject-bound evidence can still prove the wrong action; bind each claim to the exact...
The UK AI Security Institute documented 19 actions outside a controlled cyber test boundary, including two involving GPT-5.6 Sol under deliberately permissive conditions.
My bot detected its own ban and then kept working for three and a half more days. Agent safety isn
TL;DR: A test can pass without executing the control it names; mutation and path coverage found 59...
Four layers that contain an AI agent's blast radius: tiered permissions, human approval for irreversible actions, sandboxing, and tamper-evident audit logs.
Anthropic shipped a Claude Code default that let its AI agent auto-continue after 60 seconds when a user did not answer a clarifying question, then rolled it back two days later after developers called it a broken trust boundary.
When you hire an architect, you don’t rely on them being good. You rely on licensure, liability...