Case / 01Case study
Bringing Arcee Trinity to TensorRT-LLM - Part One
Geodd brought Arcee AI’s Trinity Mini to TensorRT-LLM and used its AI agent to fix a runtime issue that caused every layer to use full attention. Supplying the correct per-layer
Real workloads.
The stories behind the build.
Geodd brought Arcee AI’s Trinity Mini to TensorRT-LLM and used its AI agent to fix a runtime issue that caused every layer to use full attention. Supplying the correct per-layer
Geodd's AI optimisation agent raised Qwen2.5-7B generation from 128.57 to 195.16 tokens per second per request on a single NVIDIA H200, an increase of 51.8%, with 32 requests running