Back to articles

Tag archive

#dpo

T
Aug 31, 2026

Teaching a local coding agent from its own mistakes: DPO on a 30B model

Our VS Code assistant was passing every test on its curriculum — which meant the curriculum had stopped measuring anything. Here's how we built honest eval sets, found two silent contaminations in our test bench, and used Direct Preference Optimization on the assistant's own redirect pairs to teach a 30B model to pick the right tool on the first try. Five OutOfMemory crashes, one counterintuitive fix, a clean 4-hour training run — and a pre-registered eval gate whose verdict we report as measured, including the part that failed.

Aug 31, 20267 min read1 reactions2 comments