
Claude Sonnet 5.5 nears Opus 5.5 at half the price
Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 at $2/$10 per million tokens, half Opus 5.5's price. Sonnet 4.5 retires on November 30, 2026.
Tag archive

Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 at $2/$10 per million tokens, half Opus 5.5's price. Sonnet 4.5 retires on November 30, 2026.

Google's Gemini 4 Argon leads on reasoning, agentic tasks, and hallucination — but ranks 8th in Code Arena. What that split means.

We ran XSTest and OR-Bench over ten models, then audited the prompts. Only 77 of 200 hard ones were clearly benign, and Opus 5.5 drops from 23% refused to 4%.
Hey all 👋 Part 1 worked out how to move twenty-five container images across an air gap as one file:...
With readers that always overlap, SQLite's WAL grows by 11.9 MiB a second under a 2,000 commit/s writer and keeps its size after the readers leave. I read wal.c to see why, then measured what journal_size_limit and TRUNCATE checkpoints actually fix.

Google's new Gemini 4 Argon can reportedly hunt and patch software vulnerabilities on its own, and Google says it tops OpenAI's GPT-6 Astra and Anthropic's models on a third-party index. Lovely. That's the fourth 'best AI in the world' this month.

I measured Wild 0.10.0, mold 2.42.1, rust-lld and GNU ld linking ripgrep and cargo. Wild won all 20...

I measured Rust Coreutils (uutils) 0.12 against GNU coreutils 9.12 in containers with limits: the...

The claim behind TabPFN and TabICL: they predict on a table without ever training on it, and still...

862ms voice→voice, zero self-interruptions under echo, zero dependencies, MIT. Measured black-box against OpenAI Realtime, Pipecat and ElevenLabs ConvAI — with the bench published so you can rerun it.

Qwen3-Embedding-4B cosine task→group routing vs Claude Opus 5, Sonnet 4.6 and TypeSafe Jev on 138 todo titles. The embedding was almost never wrong — it left half the board in Unsorted.

The highest LoCoMo number we're aware of, at 5.0K context tokens, and it holds at 95.0% under Mem0's own judge. The more useful finding: swapping only the answerer model moved the same memory system by 7.4 points.