Qwen 2.5 7B
+51.8% higher generation rate
Mean time to first token was 23.8% lower in the Geodd run.
Test configuration
Baseline: vLLM with FP8 dynamic on H200. Geodd precision, batch settings, workload lengths, and runtime versions were not recorded for this pair.