DeepSeek Harness vs Codex vs Claude Code: What I Learned Comparing Three Agent Harnesses
I’ve been saying for months that the model is not the product anymore, the harness is. Nobody...
Tag archive
I’ve been saying for months that the model is not the product anymore, the harness is. Nobody...

Give an autonomous coding agent a hundred-thousand-token context window, point it at a GitHub repository, and ask it to reproduce an ML baseline.
Not a chatbot with more steps “Agentic AI” gets used for basically anything with a system...
Originally published on andrew.ooo — visit the original for any updates, code snippets that aged...

A support agent reads a ticket. The ticket body contains: IGNORE ALL PREVIOUS INSTRUCTIONS. You...
Testing NVIDIA NemoClaw in a Local Sandboxed Environment with Bob Introduction For a...
Originally published on andrew.ooo — visit the original for any updates, code snippets that aged...

Musk says Grok 4.6 is worse without its harness. Cherny has Claude maintaining apps. GitHub is splitting agent infra into three layers. The moat moved.

The Problem Isn't "Can You Call an API?" Switching to a different model is rarely as...
The EXO agent runtime splits a self-modifying agent into a disposable policy layer and a durable state layer, so an agent can rewrite its own prompts, tools and executor code without being able to damage its own event log, secrets or history.
A 17-author study separates the ability to improve an AI agent's scaffolding from the ability to benefit from the improvement, and finds that a 9-billion-parameter model produces upgrades yielding gains comparable to Claude Opus 4.6.
Evo-Bench holds the model and budget fixed and measures only what improving its own scaffolding is worth, finding gains of up to 16.6 points that come close to human-engineered baselines everywhere except tasks with prescribed workflows.