Running 100B+ MoE LLMs on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp
Tuần này trên Hacker News có bài được hơn 300 điểm: chạy model MoE cỡ 125B trên một con RTX 4090, tốc...
Tag archive
Tuần này trên Hacker News có bài được hơn 300 điểm: chạy model MoE cỡ 125B trên một con RTX 4090, tốc...

TL;DR: llama-server --sleep-idle-seconds N unloads the model after N idle seconds and is documented...
A model running locally loads cleanly and answers correctly, then keeps generating past its answer,...

The live demo, step by step: Gemma 4 E2B answering at 76 tok/s from a GTX 1650 Ti with 4 GB of memory, and the exact re-pack of Google's QAT weights on Hugging Face that makes it fit, run faster and stay close to bf16.

A practical breakdown of my llama.cpp configuration for running Qwen 3.8 27B locally on an M5 Mac with 128 GB RAM — from MTP speculative decoding and 512K context to batching, KV cache, Flash Attention, and GPU utilization.
Discover 10 advanced Apple Silicon hacks to run local LLMs 300% faster. Master unified memory, Metal quantization, and hardware acceleration on Mac in 2026.
Ollama and llama.cpp are often compared as if they were rival inference engines. The real choice is...
Run your existing .gguf models in KoboldCpp in under 15 minutes. I’ll show the exact settings for context, GPU layers, streaming, and when Ollama still wins.

Originally published at norvik.tech Introduction Descubre cómo ejecutar modelos GGUF...
A reproducible 16GB recipe for “Qwen 35B”: which Qwen2.5-32B quants fit, how to budget KV cache, and the exact serving flags that stop OOMs.

Gemma 4 E2B q4_0 served by llama.cpp on one laptop, CPU-only and on a GTX 1650 Ti, rebuilt on CUDA 13.4 and re-measured in CPU, GPU, GPU, CPU order with a temperature gate. The card takes decode by 4.14x, and run order moves the answer by about 2%.
ROCm and Vulkan both accelerate AMD GPUs for local LLM hosting, but they are not interchangeable. The...