SGLang RadixAttention vs vLLM on One H100: A Production Throughput Reality Check
We benchmarked SGLang RadixAttention against vLLM prefix caching on a single H100, covering shared prefixes, multi-turn chat, structured output, latency variance, and GPU cost.