Back to articles

Tag archive

#llmevaluation

D
Jun 30, 2026

Does a Local Reasoning Model Earn Its Keep? Measuring thinking ON/OFF on gemma4:12b

In an earlier post I wrote off gemma4:12b's empty replies as a packaging bug. They weren't: it's a reasoning model. So I ran 13 questions with thinking ON and OFF. Reasoning got one more answer right while spending 68× the output tokens and 19× the wall-clock. Here's when I now turn it on in an agent, measured.

Jun 30, 20268 min read0 reactions0 comments