The work, up close

Case studies

Real workloads.
The stories behind the build.

The collection

2 studies
Case / 01Case study

Bringing Arcee Trinity to TensorRT-LLM - Part One

Geodd brought Arcee AI’s Trinity Mini to TensorRT-LLM and used its AI agent to fix a runtime issue that caused every layer to use full attention. Supplying the correct per-layer

Case / 02Case study

Qwen2.5-7B Inference Optimisation with Phala

Geodd's AI optimisation agent raised Qwen2.5-7B generation from 128.57 to 195.16 tokens per second per request on a single NVIDIA H200, an increase of 51.8%, with 32 requests running

Have a workload in mind?Let's talk