DPO vs supervised fine-tuning: do student projects need preference tuning at all?
Direct Preference Optimization is having a moment. For most student projects, plain supervised fine-tuning on good examples still gets you further, fa
Sep 7, 20262 min read0 reactions0 comments
