T
Aug 23, 2026Train a 125M LLM on Your Laptop in <10 Minutes with NanoGPT
Train a 125‑M‑Parameter LLM on Your Laptop in Under 10 Minutes NanoGPT speedrun – the...
Aug 23, 20264 min read0 reactions0 comments
Tag archive
Train a 125‑M‑Parameter LLM on Your Laptop in Under 10 Minutes NanoGPT speedrun – the...
What Changed A new paper from Hugging Face, titled "GRPO, Dr. GRPO, and DAPO Are Three...
Researchers from Tsinghua and ByteDance show that a small model's reinforcement learning gains can be distilled into a larger, already-stronger model by transferring the change in the teacher's policy rather than its outputs, letting a 1.5B teacher i
DOPD and MOPD advance on-policy distillation -- training a student on its own outputs -- with DOPD routing supervision to avoid a 'privilege illusion' and MOPD merging multiple specialist RL teachers into one model without cross-domain interference.