Qwen 2.5 7B · NVIDIA H200
51.8% higher generation rate. 23.8% lower time to first token.
In the recorded H200 comparison, Qwen 2.5 7B reached 195.16 tokens per second, up from 128.57 with the vLLM baseline. Mean time to first token fell from 789.2 ms to 601.4 ms.
Baseline: vLLM with FP8 dynamic on H200. Geodd precision, batch settings, workload lengths, and runtime versions were not recorded for this pair. Results describe these runs; they do not isolate the contribution of a particular kernel change.
Trinity Mini · NVIDIA H200
From 61.04 to 114.51 tokens per second.
Across four recorded versions, Trinity Mini’s mean generation rate increased by 87.6% from the first version to the final recorded version.
First version61.04 tokens/s
Second version64.67 tokens/s
Third version75.70 tokens/s
Final version114.51 tokens/s
This measures progress from our first recorded optimization version. Against the separately recorded vLLM baseline of 78.00 tokens/s, the final run’s generation rate was 46.8% higher. Its mean time to first token was also higher: 2.9265 s, compared with 0.7055 s for vLLM.