Mosaic
NVIDIA / CUDA
View benchmarksGeodd is an AI inference company. We serve models through one API and develop the hardware kernels, the code that runs on accelerators, that help those models execute more efficiently.
Our goal is to turn the performance problems we find in real workloads into improvements people can use through our service.
A short conversation, a long document, and many requests arriving together place different demands on a model. The work needed to improve one can be different from the work needed to improve another.
We bring inference operations and kernel development together so the problems we encounter while serving models can guide what we build next.
Our AI agents use hardware-specific LLMs to develop kernels for the bottlenecks we identify. Changes are tested in staging for correctness and performance under the relevant workload conditions.
Improvements that pass verification are deployed automatically into serving. We measure their effect and use the results to guide the next experiments and further development of our specialist models.
Explore the Model EngineVerification passes?
Return to kernel development
Revise and retestDeploy automatically and measure
Back to serving observationsWe develop specialist models around the accelerator and software stack each one targets.
NVIDIA / CUDA
View benchmarksAMD / ROCm
Coming soon
Tenstorrent / TT-Metalium
Coming soon
Our engineering work spans model serving, accelerator performance, and production operations. Keeping those responsibilities connected helps us investigate problems with the context of the system that produced them.
Developers can speak with our team about their workload, deployment needs, and the performance they are seeing.
Meet the team through a conversationRead our benchmark results and engineering updates, or run inference runtimes with model-specific optimized kernels on your own hardware.
Compare recorded results and inspect the available test details.
View benchmarksFollow the work behind our models, infrastructure, and service.
Read articlesTest and use free inference runtimes with model-specific kernels developed by our Mosaic-powered AI agents.
Browse local releasesThese local runtime releases are separate from the specialist kernel-development LLMs we use internally. Their availability does not mean the specialist model weights are publicly released.
Run a model through our API, or talk with our engineers about the workload you want to deploy.