About Geodd

We run inference. We keep improving it.

Geodd is an AI inference company. We serve models through one API and develop the hardware kernels, the code that runs on accelerators, that help those models execute more efficiently.

Our goal is to turn the performance problems we find in real workloads into improvements people can use through our service.

Why we do this

Performance changes with the workload.

A short conversation, a long document, and many requests arriving together place different demands on a model. The work needed to improve one can be different from the work needed to improve another.

We bring inference operations and kernel development together so the problems we encounter while serving models can guide what we build next.

Operate inference
Understand where time and compute are being spent.
Develop improvements
Address the bottlenecks that measurements reveal.
Our approach

Measure the problem. Verify the improvement.

Our AI agents use hardware-specific LLMs to develop kernels for the bottlenecks we identify. Changes are tested in staging for correctness and performance under the relevant workload conditions.

Improvements that pass verification are deployed automatically into serving. We measure their effect and use the results to guide the next experiments and further development of our specialist models.

Explore the Model Engine
From serving to release
  1. Observe serving performance
  2. Agents develop kernels
  3. Test in staging
Release gate

Verification passes?

No

Return to kernel development

Revise and retest
Yes

Deploy automatically and measure

Back to serving observations

We develop specialist models around the accelerator and software stack each one targets.

Druze

AMD / ROCm

Coming soon

Strata

Tenstorrent / TT-Metalium

Coming soon

The people behind Geodd

Built and operated by the same team.

Our engineering work spans model serving, accelerator performance, and production operations. Keeping those responsibilities connected helps us investigate problems with the context of the system that produced them.

Developers can speak with our team about their workload, deployment needs, and the performance they are seeing.

Meet the team through a conversation
Our work

Explore the results. Try optimized runtimes.

Read our benchmark results and engineering updates, or run inference runtimes with model-specific optimized kernels on your own hardware.

Benchmarks

Compare recorded results and inspect the available test details.

View benchmarks

Engineering updates

Follow the work behind our models, infrastructure, and service.

Read articles

Run locally

Test and use free inference runtimes with model-specific kernels developed by our Mosaic-powered AI agents.

Browse local releases

These local runtime releases are separate from the specialist kernel-development LLMs we use internally. Their availability does not mean the specialist model weights are publicly released.

Build with Geodd.

Run a model through our API, or talk with our engineers about the workload you want to deploy.