Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default
A serving setup for Qwen3.8-27B on a single RTX 3090 with vLLM whose README measures when `DFLASH_TOKENS=15` pays and what it costs in reque
Tag archive
A serving setup for Qwen3.8-27B on a single RTX 3090 with vLLM whose README measures when `DFLASH_TOKENS=15` pays and what it costs in reque

FP16, INT8, and binary. Pick your spot on the accuracy-cost dial before Amazon OpenSearch Service...

When I started this project, I had a question in mind: if a charging network keeps adding stations...
Hot take For many production features, defaulting to a cloud LLM is the lazy opt-in. But...
Developers download Q4_K_M models daily but few understand how floating-point weights become 4-bit integers. Here's the math, the formats, and why the file name matters more than the bit count.
Prism's ternary Bonsai 2 compresses its smallest 27B language-weight pack to 5.95 GB, but the headline is a disk-size claim rather than a full runtime or VRAM requirement.
A reproducible 16GB recipe for “Qwen 35B”: which Qwen2.5-32B quants fit, how to budget KV cache, and the exact serving flags that stop OOMs.

I ran Prism's ternary-quantized Bonsai 2 27B, a 27B model squeezed under 6GB, on a GPU-less VPS to test its near-lossless claim. Here's what happened.

Self-hosting an open-source LLM is all about control—until the GPU bill lands. This no-nonsense guide walks through VRAM sizing, quantisation trade-offs, serving stacks, and the budget line item that catches every team off guard.
A from-scratch model scored 44% on ARC-AGI-1 in 67 cents of compute. The lesson for small models: the training bill is optional.
A per-block scale that cuts FP4 gradient-quantization error on 45 of 45 tensors, 14% against the published state of the art, two rented-GPU training runs that landed within noise anyway, and the measurement that explains both: the gap only shows at million-token batches, 35 to 643 times larger than anything I ran.
In this video: 0:00 The Crash Nobody Can Explain 0:18 It Loads, It Answers... Then Dies 1:36...