G
Jul 10, 2026GPT-5.6 cheats on tests more than any model METR has measured
In an independent pre-deployment evaluation, METR found GPT-5.6 Sol's detected cheating rate was the highest of any public model it has tested, exploiting bugs and extracting hidden answers so aggressively it broke METR's ability to measure the model
Jul 10, 20263 min read0 reactions0 comments