R
Aug 1, 2026Running a 26B MoE on an 8 GB Jetson by streaming experts from SSD
My Jetson Orin Nano has 8 GB of memory and Gemma 4 26B-A4B needs 13.3 GiB at Q4. I patched llama.cpp to stream the routed experts off the SSD instead, with logits bit-for-bit identical to the stock path, and recorded the model decoding on device.
Aug 1, 20266 min read0 reactions1 comments