AI Inference

Fast, steady inference for AI agents.

With our AI agents that continuously tune how models run on the hardware.

Long runs that finish. Agents complete long tasks without one failure breaking the whole run, and latency stays flat as runs get longer.

Consistency you can rely on. Steady performance on every call, so your team builds product instead of retries and backup providers.

How It Works02

Production traffic guides every improvement.

  1. 01

    Observe real workloads

    Execution graphs, context lengths, and changing batch sizes show where inference spends time.

  2. 02

    Write better kernels

    Our AI agents use hardware-specific LLMs to generate and refine code for those bottlenecks.

  3. 03

    Test against the workload

    Check correctness and measure whether the changes improve inference under the conditions that exposed the problem.

  4. 04

    Deploy and measure

    Release verified improvements into serving, then monitor their effect on production performance.

  5. 05

    Feed the next cycle

    Successful kernels and measured results inform the next experiments and further development of our hardware-specific LLMs.

Our Kernel Development Models03

The LLMs powering
this process.

Meridian, Helix, and Stride power the agents in this process. Each is fine-tuned for the hardware it optimizes.

01 / CUDA

Meridian · NVIDIA

Kernel development for NVIDIA GPUs and CUDA.

View benchmarks
Parallel / OrderedCUDA
02 / ROCm

Helix · AMD

Kernel development for AMD GPUs and ROCm.

Coming soon
Branching / AdaptiveROCm
03 / TT-Metalium

Stride · Tenstorrent

Kernel development for Tenstorrent accelerators and TT-Metalium.

Coming soon
Mesh / DistributedTT-Metalium
Pricing & Deployment

Serverless Inference

GLM 5.3 Flash

Model Type
Text
Quantization
fp8

GLM-5.3 Flash brings coding, reasoning, and image understanding to AI applications. Use it to power coding assistants, debug software, analyze documents and screenshots, or build agents that call tools and complete multistep tasks. Its long-context capabilities make it useful for working across large codebases and detailed research material.

zai-org/glm-5.3-flash
$0.110
per 1M tokens

DeepSeek V4.1 Flash

Model Type
Text
Quantization
fp8

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively.

deepseek-ai/DeepSeek-V4.1-Flash
$0.300
per 1M tokens

openai/gpt-oss-120b

Model Type
Text
Quantization
bf16

GPT-OSS-120B is OpenAI's open-weight reasoning model designed for high-capability reasoning, tool use and agentic workloads. Geodd serves GPT-OSS-120B through an OpenAI-compatible API with serverless inference in supported Geodd regions.

openai/gpt-oss-120b
$0.039
per 1M tokens

Multi-Regional

Deploy in our US and EU regions. Deployment in APAC is underway.

GDPR Ready & SOC 2 Pending

Enterprise-grade security and data isolation for all workloads.

Unified API

One SDK for both serverless inference and dedicated compute.

Locations Topology

Built for Scale

Select a region

Select a location dot or region to inspect capacity

Select a region
Deployment footprint

Global infrastructure

Regional compute capacity positioned for production inference workloads.

3
Global
regions
2
Active
regions
Active
Expansion
Blog

Latest Updates

Dedicated GPU vs Dedicated AI Inference

Dedicated GPU vs Dedicated AI Inference

A dedicated GPU gives your team reserved GPU compute. Dedicated AI inference gives your team a dedicated or isolated inference environment that may in...

Total Cost of Self-Hosted Inference

Total Cost of Self-Hosted Inference

The total cost of self-hosted inference is not only the hourly GPU price. It includes GPU capacity, supporting compute, storage, networking, orchestra...

Integration

Developer-First
Control

Fully compatible with the OpenAI SDK. Switch providers with a single line of code. No migration headaches, just immediate performance gains. This example uses the GPT-OSS-120B API.

  • Direct OpenAI SDK compatibility
  • Real-time token usage and observability
  • Privacy first with Zero Data Retention (ZDR) and logging policy
deploy_inference.py
from openai import OpenAI

# Switch to Geodd by changing base_url
client = OpenAI(
  api_key="GEODD_API_KEY",
  base_url="https://api.geodd.io/inference/v1"
)

completion = client.chat.completions.create(
  model="openai/gpt-oss-120b",
  messages=[
    {"role": "user", "content": "What is machine learning?"}
  ]
)

print(completion.choices[0].message)

Ready to Scale?

Explore Geodd
Today.

Get instant access to our Model APIs and dedicated GPUs. Precision engineered for the most demanding production workloads.

deploy — mistral-large
Provisioning GPUdone
Preparing Serverdone
Configuring Accessdone
Downloading Modelrunning…
Starting Model Serverpending
Finalizing Servicepending
$awaiting pipeline