INT8 Q/DQ Calibration on Blackwell: 1.8 the TRT 10 + FP16 Baseline
A practical walkthrough of doing INT8 post-training quantization the right way on RTX 5090 + TensorRT 11. 1,500 stratified calibration samples + NVIDIA ModelOpt, 56 seconds of work, 71k NPS — 1.8× the previous TRT 10 + auto-FP16 baseline on the same model. Covers calibration set design, the ModelOpt invocation, and what TRT 11's strongly-typed model is actually buying you on Blackwell.
