I gave my AI agent 76 tools. Then I built one that catches it lying about using them.
The problem nobody talks about Last month I caught my own AI agent saying "I've already...
Tag archive
The problem nobody talks about Last month I caught my own AI agent saying "I've already...
Originally published on andrew.ooo — visit the original for any updates, code snippets that aged...
Hallucination isn't lying. It's the same mechanism that makes models work at all.
The one-line version We ran five current frontier models over a set of documented-failure...
Quick read · 8 min read This article shows you how to give autonomous agents a governed map of your...
Ninety-six Python coding tasks. Ninety-six package names a frontier AI model handed back, each with...
A system that generates complete research papers as thirteen composable skills inside a coding assistant audited at 99.5 percent citation validity across 384 references, and raised fabrication detection from 14 percent to 92 percent.
A new study measuring 12 language models across more than 143,000 judged claims found every one of them invented or stereotyped between 35 and 49 percent of what it asserted about a user, and that the models most confident they were being careful wer
Somewhere in the pull request your team merged last Thursday, there's an import statement for a...
AI SDR Hallucination A great rep once knew every account. Now your agents do. But only if...

Build an AI blog automation pipeline that refuses to fabricate. This guide walks through a hard proof gate that skips any commercial post without a real case study and human expert note — and why an empty backlog is the healthiest signal your content engine can send.
A 2026 guide to RAG citations, abstention, and the feedback loop: NLI-based citation scoring, faithfulness verification, and confidence-based escalation.