Evaluating Speculative Decoding in vLLM on AMD MI300X GPUs
An engineering evaluation of speculative decoding methods in vLLM on AMD Instinct MI300X hardware, analyzing throughput, acceptance rates, and tradeoffs.
Tag archive
An engineering evaluation of speculative decoding methods in vLLM on AMD Instinct MI300X hardware, analyzing throughput, acceptance rates, and tradeoffs.
A late‑2026 compatibility matrix plus known‑good commands for getting Unsloth LoRA fine‑tuning to actually converge on Radeon and MI GPUs with ROCm.

Originally published at norvik.tech Introduction Explore the implications of ROCmfix...

Run OpenAI Whisper locally on AMD GPUs using ROCm without VRAM crashes or OOM errors. Complete setup...

Learn how to run Stable Diffusion and ComfyUI on AMD GPUs using ROCm. Step-by-step guide to...
Accelerating Inference with Speculative Decoding on AMD Hardware: A Technical Deep...

A step by step deployment of Gemma 4 E2B to a single AMD Instinct MI300X on AMD Developer Cloud, driven by Python MCP tools, and the throughput a 191.7 GiB card returns for its hourly rate.

One AMD Instinct MI300X on AMD Developer Cloud, managed entirely through a tag-scoped Python MCP server, with every figure read off the card rather than a spec sheet. fp8 e4m3fnuz runs 1.77x bf16; int8, which AMD rates identically to fp8, runs 0.69x; fp4 is not on this silicon at all. One droplet, $1.99 an hour, and two readings that were wrong the first time.

PyTorch is an open-source framework for building and training machine learning models, especially...

TensorFlow is an open-source framework widely used for building and training machine learning models,...

JAX is an open-source library for high-performance numerical computing and machine learning research,...
I'm 17, mechanic from Torino. I built vLLM to run NATIVELY on Windows on my RX 6750 XT 12GB (gfx1031...