Inference as a Service

Production
Inference Stack

End-to-end inference with automatic scaling, streaming tokens, and real-time monitoring.

Infrastructure

Deployment Modes

Choose the optimal runtime for your workload. From elastic API endpoints to bare-metal isolated instances.

Serverless and dedicated deployment comparison
Compare by01 / Serverless

Serverless Inferencing

AUTO_SCALING
02 / Dedicated

Dedicated Deployment

ISOLATED_RUNTIME
Workload fit

Serverless endpoints are optimized for rapid deployment and elastic workloads. Fully abstracted infrastructure with automatic scaling.

Used when workload predictability, isolation, or sustained throughput becomes critical. Single-tenant GPU allocation.

ExecutionMulti-tenant optimized runtimeSingle-tenant isolated runtime
InfrastructureFully abstractedDedicated GPU allocation
ControlParameter-level tuningInfra + runtime control
ScalingAutomatic, workload-drivenCluster-level scaling
Use CaseDynamic workloads, API productsStable high-throughput systems
Get started
Available Models

Run the models you need.

01 / Model Type: Text

GLM 5.2

We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include:

zai-org/glm-5.2
$0.900per 1M tokens
02 / Model Type: Text

DeepSeek V4 Flash

DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.

deepseek-ai/DeepSeek-V4-Flash
$0.140per 1M tokens
03 / Model Type: Text

openai/gpt-oss-120b

GPT-OSS-120B is OpenAI's open-weight reasoning model designed for high-capability reasoning, tool use and agentic workloads. Geodd serves GPT-OSS-120B through an OpenAI-compatible API with serverless inference in supported Geodd regions.

openai/gpt-oss-120b
$0.039per 1M tokens

How it works

Better kernels.
Better inference.

Our hardware-specific LLMs help agents write better kernels. We test the changes and return verified improvements to the inference service.

  1. 01 / Runtime

    Measure performance.

    Runtime performance measurements show where execution can improve. Our optimization agents do not inspect customer prompts or completions.

  2. 02 / Model Engine

    Write and test kernels.

    Meridian, our NVIDIA-specific LLM, powers AI agents that write and refine CUDA kernels. The fully automated workflow tests correctness and execution performance, then deploys verified improvements to our inference service.

  3. 03 / Inference service

    Deploy and improve.

    Verified improvements return to serving. Kernel code and measured results guide the next cycle, not customer content.

Infrastructure

Built for continuous operation.

Regional GPU capacity, redundant infrastructure, and direct engineering ownership support production inference under sustained load.

NVIDIA GPUs
500+
Inference capacity across multiple regions.
Observed uptime
99.95%
Across multi-location deployment with failover.
Datacenter infrastructure
Tier III
Redundant power, network, and hardware.
Failure handling

Resilience in the stack.
Engineers in the loop.

Infrastructure and MLOps work together, from failure detection to recovery. Incidents go directly to the engineers responsible for the runtime.

  1. Reduce single points of failure

    Redundant power, network, and hardware underpin the serving infrastructure.

  2. Detect and recover

    Failure detection and automated failover support continuous operation across locations.

  3. Respond directly

    Engineers receive incident alerts directly. Infrastructure and MLOps coordinate recovery and runtime tuning.

Compliance

Enterprise AI Inference,
Built for GDPR.

Deploy production-ready AI models without the compliance headache. Geodd provides high-performance, GDPR-ready inference designed to eliminate unnecessary data exposure for products serving the EU, EEA, and UK.

  • Zero-Data Retention

    We never store your API prompts, completions, request bodies, or customer datasets.

  • No Training or Human Review

    Your data is strictly yours. It is never used for model training or human-in-the-loop screening.

  • EU Data Sovereignty

    Route and process your inference workloads entirely within EU-based data center infrastructure.

  • Enterprise Security

    Hardened with encryption at rest and in transit (TLS), role-based access controls, and a comprehensive DPA framework.

  • The Geodd Standard

    We act strictly as a Data Processor for your API data, keeping your internal workflows, product logic, and customer messages safe by default.

Ready to Scale?

Join the next generation of AI

Build on Geodd's hyper-optimized inference stack. Get instant API access to the world's most capable open-source models or talk to our team for custom deployments.