A
Aug 19, 2026A runbook, not a model, hit 95 percent on a live agent benchmark for 15 dollars
StateM reaches 95.3 percent raw accuracy on Terminal-Bench 2.1 across 445 trials without changing any model weights, using a state-machine runtime and a reusable runbook, at about 15 dollars of final-score API spend against 574.68 dollars for the ref
Aug 19, 20264 min read0 reactions0 comments
