Back to articles

Tag archive

#tensorrt

I
Jun 10, 2026

INT8 Q/DQ Calibration on Blackwell: 1.8 the TRT 10 + FP16 Baseline

A practical walkthrough of doing INT8 post-training quantization the right way on RTX 5090 + TensorRT 11. 1,500 stratified calibration samples + NVIDIA ModelOpt, 56 seconds of work, 71k NPS — 1.8× the previous TRT 10 + auto-FP16 baseline on the same model. Covers calibration set design, the ModelOpt invocation, and what TRT 11's strongly-typed model is actually buying you on Blackwell.

Jun 10, 20267 min read0 reactions0 comments