
vLLM's weight cache can serve another checkpoint's weights when the tensor layout matches
TL;DR: vLLM 0.30.0 can keep a model's weights in a long-running daemon so that engines restart...
Tag archive

TL;DR: vLLM 0.30.0 can keep a model's weights in a long-running daemon so that engines restart...

TL;DR: A PEFT LoRA adapter can give individual modules their own rank and alpha through rank_pattern...
A serving setup for Qwen3.8-27B on a single RTX 3090 with vLLM whose README measures when `DFLASH_TOKENS=15` pays and what it costs in reque
An engineering evaluation of speculative decoding methods in vLLM on AMD Instinct MI300X hardware, analyzing throughput, acceptance rates, and tradeoffs.

Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.

Repacking Gemma 4's QAT weights five ways and serving each on the same SageMaker NVIDIA L4 endpoint: int4 linears, int4 embeddings and lm_head, FP8 and int8, across E2B, E4B, 12B, 26B A4B and 31B.

The same Gemma 4 build, vLLM version and GPU served from a SageMaker endpoint and from a plain EC2 instance, on a T4 and an L4: identical decode and answers, a different call path, and what the managed endpoint's 1.40x buys.

Google's QAT Gemma 4 E2B keeps its embedding tables in bf16, and on a Tesla T4 they are most of the model. Packing them to int4 on the grid QAT trained them onto cuts model loading from 6.33 to 2.86 GiB, with every greedy test output token-identical, and raises vLLM's output throughput 11-37% over Google's own W4A16 export.

A short background on SageMaker real-time endpoints, then a measured comparison of Gemma 4 E2B's QAT w4a16 checkpoint against the full-size bf16 release on the same NVIDIA L4 endpoint: decode speed, parallel throughput, answers and cost.
We examine NVIDIA Dynamo's Kubernetes prefill/decode architecture, build a release-pinned vLLM deployment workflow, and identify the validation gates required before publishing performance or cost claims.

Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.
A production-first checklist for self-hosting an OpenAI-compatible vLLM endpoint: auth, per-tenant quotas, streaming SSE, queueing, and redaction-safe logs.