A 2.78T Model in 8 GB of RAM, and the Gate Moves to the Disk
kimi-k3-in-c runs a 2.78-trillion-parameter Kimi K3 on one CPU in 8.24 GB of RAM, with byte-identical output from 8 GB up to 224 GB, so memo
Tag archive
kimi-k3-in-c runs a 2.78-trillion-parameter Kimi K3 on one CPU in 8.24 GB of RAM, with byte-identical output from 8 GB up to 224 GB, so memo

For everyday mixed enterprise workloads with bursty traffic and short prompts, CPU inference is usually cheaper per query on-premise. A GPU only pays for itself once one model runs at sustained high u
Originally published on andrew.ooo — visit the original for any updates, code snippets that aged...
VIDRAFT의 POCKET-35B-GGUF, GPU 없이 CPU만으로 350억 파라미터 LLM 실행. Q2_K(13GB)로 미니PC 일상 사용, IQ1_M(8.2GB)으로 16GB RAM 기기까지 지원. 로컬 AI 추론의 새 기준.
A $300 Xeon from 2013 runs Gemma 4 26B at 5 tok/s with no GPU. Here's the memory bandwidth math, quantization tradeoffs, and production decision framework nobody else is covering.

Successfully deploying llama.cpp and Ollama on a 64-core RISC-V workstation to run a 70-billion parameter language model without GPU acceleration - a
Learn how to set up, optimize, and execute popular AI models on a CPU‑only machine in just a few hours