G
Jul 24, 2026Giving experts no memory cuts optimizer state from 50 gigabytes to 1.3
A single-author preprint shows that a mixture-of-experts model can drop momentum entirely for its expert layers, shrinking persistent optimizer state from 50.55 gigabytes to 1.29 with almost no effect on final quality.
Jul 24, 20264 min read0 reactions0 comments