All articlesThe Chronicle / Geodd

GPU Cloud Pricing: 7 Providers Compared by Hourly Rate

Renting a GPU looks simple until you compare invoices. One provider quotes an H100 at a low hourly rate but bills storage, egress, and idle time separately. Another bundles everything but only sells eight-GPU nodes. If you are trying to understand GPU cloud pricing, the headline number rarely tells you what you will actually pay.

Here is the short answer. Hourly rates for an NVIDIA H100 vary by a factor of two or more across providers, and the cheapest listed rate is often not the cheapest total cost. Specialist GPU clouds usually beat the big hyperscalers on price, while reserved capacity beats on-demand once your usage is steady. For cloud GPU pricing, the real variables are commitment length, minimum cluster size, and how much of the hour your GPU does useful work.

Below, we compare seven providers by hourly rate for H100, H200, B200, and A100 GPUs, and note what each one charges extra for. At Geodd, we run serverless and dedicated inference every day, so we also flag when paying per token beats renting a GPU at all.

1. Geodd

Pricing model and what you pay for

Geodd sells a managed inference service and dedicated GPU servers, so its pricing works differently from the rest of this list. Serverless inference is billed by usage, and real-time token usage observability shows you spend as it happens. Dedicated GPUs are single-tenant bare-metal servers (H200, H100, and RTX Pro 6000) with server-level SSH access, for when you want reserved capacity or full control. Check Geodd's GPU pricing for current quotes.

Per-token billing means you never pay for an idle GPU. For spiky traffic, that usually beats any hourly rate. For steady, high-volume load, a dedicated deployment is the closer match to an NVIDIA GPU cluster price.

Pay per token when traffic is spiky, and pay for dedicated GPUs when it is steady.

GPUs and models available

Geodd runs 500+ GPUs in North America, with active regions in US-EAST and EU-NORTH (Norway). APAC-SOUTH in Colombo is being built out. You reach everything through one API and SDK, across these model families:

  • Text: GLM-5.2, DeepSeek V4 Flash, GPT-OSS-120B, Gemma 4 31B IT
  • Image and video: ByteDance Seed, Seedream, and Seedance

Under the hood, hardware-specific LLMs (Meridian for NVIDIA/CUDA, Helix for AMD/ROCm, Stride for Tenstorrent) write and refine kernel code from your production workloads. The goal is steady latency on long agentic runs, which means fewer failed runs and retries. Retries cost money, so that lowers your effective price.

Best for

Teams shipping agents and AI features who want predictable performance without running GPUs themselves fit best. The API is OpenAI SDK compatible, so switching providers is a one-line change. EU companies also get GDPR-ready handling and a Zero Data Retention policy.

  • Long-running, agentic workloads
  • Teams migrating off OpenAI or self-hosted GPUs
    • Startups that want to avoid GPU operations overhead
      • Teams that need dedicated bare-metal GPUs for training, custom runtimes, or fixed capacity

Watch out for

Dedicated GPUs are bare metal, so you own the OS, drivers, and runtime, and no orchestration layer is included. If you would rather not manage that, the managed inference service is the better fit. Also note that SOC 2 is still pending and APAC capacity is not live yet, so confirm both against your compliance and latency requirements.

2. Lambda

Pricing model and what you pay for

Lambda bills on-demand GPUs by the minute with no egress fees. That removes one of the most common surprise lines on a GPU invoice. Persistent storage is billed separately per GB per month, and the lowest rates require a reserved contract. The figures below are approximate list prices per GPU hour, so confirm them before you budget.

GPUApprox. on-demand rate
A100 80GB~$1.79
H100 SXM~$2.99
B200~$4.99

Per-minute billing and zero egress fees keep Lambda's listed rate close to what you actually pay.

GPUs and models available

Lambda offers H100, H200, B200, and A100 GPUs, plus GH200 superchips. You can rent single-GPU instances or eight-GPU nodes. For multi-node training, 1-Click Clusters give you interconnected GPUs without building the network yourself.

Instances ship with a preinstalled software stack covering CUDA drivers, PyTorch, and common ML libraries. Lambda also runs a hosted inference API for open models, but its core product is the GPU rental.

Best for

Researchers and ML teams who train or fine-tune models and want simple, predictable billing will like Lambda. It also suits anyone who needs a ready environment on a raw GPU.

  • Fine-tuning and training runs
  • Small teams comparing cloud GPU pricing without hidden fees
  • Multi-node jobs on clusters

Watch out for

Popular GPUs, especially H100 and B200, often sell out, so the cheapest rate does you no good if you can't get an instance. Lambda has fewer regions than the hyperscalers, which matters if you have data residency or latency requirements. Idle instances also keep billing until you shut them down.

3. RunPod

Pricing model and what you pay for

RunPod bills Pods by the second and charges no ingress or egress fees. You choose Secure Cloud, which runs in vetted data centers, or Community Cloud, which costs less and runs on third-party hosts. Storage is billed separately, and you keep paying for volume disks while a Pod is stopped. These rates are approximate per GPU hour, so confirm them before you budget.

GPUApprox. on-demand rate
A100 80GB~$1.64
H100 SXM~$2.69
H200~$3.59
B200~$5.98

Per-second billing makes short experiments cheap, as long as you remember to shut Pods down.

GPUs and models available

The catalog covers A100, H100, H200, and B200, plus cheaper RTX 4090 and L40S cards for smaller jobs. Instant Clusters link multiple nodes for training runs.

Serverless endpoints scale workers down to zero, so you pay only while requests run. Templates for PyTorch, ComfyUI, and vLLM let you launch in minutes, and you bring your own model.

Best for

Indie developers, startups, and researchers who want low hourly rates and fast experimentation fit RunPod well.

  • Prototyping and fine-tuning on a tight budget
  • Bursty serverless endpoints
  • Image generation and other GPU-heavy side projects

Watch out for

Community Cloud reliability varies by host, so keep it away from jobs that can't restart from a checkpoint. Popular GPUs also sell out in some regions.

Serverless cold starts can add latency when workers scale from zero. If your agents need steady response times, test under real load before you commit.

4. Vast.ai

Pricing model and what you pay for

Vast.ai is a marketplace, not a single fleet. Independent hosts set their own prices, so rates move with supply and demand. Billing runs per second, and storage and bandwidth are charged separately on each listing. Interruptible instances bid lower but can be paused when someone outbids you. In a GPU cloud pricing comparison, Vast.ai usually posts the lowest numbers. These ranges are approximate per GPU hour, so confirm them live.

GPUApprox. marketplace range
A100 80GB~$0.80 to $1.50
H100~$1.50 to $2.50
H200~$2.50 to $3.50
B200~$3.50 to $5.50

On a marketplace, the lowest price is a range, so check the host before you check the rate.

GPUs and models available

Listings cover A100, H100, H200, and B200 cards, plus consumer RTX 4090 and 5090 GPUs. You can filter by GPU, reliability score, bandwidth, and location. You launch your own Docker image or a template, and you bring your own model, because Vast.ai does not host one for you.

Best for

Cost-driven users who can tolerate some risk get the most from Vast.ai. It suits short experiments and fault-tolerant jobs where the lowest hourly rate matters more than uptime.

  • Fine-tuning with regular checkpoints
  • Batch rendering and hyperparameter sweeps
  • Students and hobbyists on a tight budget

Watch out for

Host quality varies widely. Two listings with the same GPU can differ in network speed, disk performance, and uptime, so read the reliability score before you rent. Interruptible instances can stop mid-run. Community hosts also run your workload on hardware you don't control, so keep sensitive or regulated data off them. That makes Vast.ai a poor fit for GDPR-bound production work.

5. CoreWeave

Pricing model and what you pay for

CoreWeave prices per GPU hour, but you rent whole nodes, usually eight GPUs at a time. Storage and networking are billed separately, and committed contracts cut the on-demand rate substantially. The figures below are approximate per GPU hour, so confirm them before you budget.

GPUApprox. on-demand rate
A100 80GB~$2.21
H100 HGX~$6.16
H200 HGX~$6.31
B200 HGX~$8.60

CoreWeave's on-demand rates sit above marketplace prices, so its value shows up in reserved capacity and cluster performance.

GPUs and models available

Expect H100, H200, B200, and A100 GPUs, plus GB200 NVL72 rack-scale systems. Nodes connect over InfiniBand networking, which matters for large distributed training. The platform is Kubernetes-native, and you bring your own models and frameworks.

Best for

Large training teams and AI labs that need thousands of interconnected GPUs get the most from CoreWeave. It also suits companies that want enterprise support and capacity commitments. In a cloud GPU pricing comparison, it wins on cluster quality, not on the lowest hourly number.

Watch out for

Smaller teams may find the eight-GPU minimum and sales-led onboarding heavy. You also need Kubernetes skills to run workloads well, and quotas for newer GPUs can take time to unlock. If you only need a single GPU for a weekend experiment, a cheaper provider above will serve you better.

6. Google Cloud

Pricing model and what you pay for

Google Cloud bills GPU VMs by the second, and the GPUs come bundled into fixed machine types such as A2 and A3. You can cut the bill with Spot VMs, which cost far less but can be preempted, or with one- and three-year committed use discounts. Disk and network egress are billed on top. These are approximate on-demand rates per GPU hour in a US region, so check the official GPU pricing page before you budget.

GPUApprox. on-demand rate
A100 80GB~$5.07
H100 80GB~$11.06
H200~$10.60

On demand, Google Cloud costs several times more than specialist clouds, so discounts decide whether it makes sense.

GPUs and models available

The catalog spans A100 (A2), H100 and H200 (A3), and B200 (A4) machines, plus cheaper L4 and T4 cards for light inference. Availability differs by region and zone.

Vertex AI adds managed endpoints and a Model Garden of hosted open and proprietary models. On raw VMs, you bring your own models and frameworks.

Best for

Enterprises already on Google Cloud get the most value, because GPUs sit next to BigQuery, Cloud Storage, IAM, and compliance tooling you already use. It also suits teams that need global regions and formal support contracts.

  • Existing Google Cloud customers
  • Regulated workloads that need audited infrastructure
  • Teams using Vertex AI or TPUs alongside GPUs

Watch out for

Quotas are the first obstacle. New projects often start with zero GPU quota, and approval for H100 or B200 capacity can take days. Spot capacity can also vanish when demand spikes.

Egress is the second. Moving large datasets or checkpoints out of Google Cloud adds per-GB network charges that specialist clouds like Lambda and RunPod waive. If you only need cheap training hours, this is rarely the best deal.

7. AWS

Pricing model and what you pay for

AWS bills EC2 GPU instances by the second, but you buy fixed eight-GPU instance types such as P5. Savings Plans, Spot Instances, and Capacity Blocks for ML (short reservations for a set window) cut the on-demand rate. Storage and data transfer out cost extra. These approximate per-GPU rates come from instance prices in a US region, so confirm them on the EC2 pricing page.

GPUApprox. on-demand rate
A100 80GB~$3.43
H100 80GB~$6.88
H200~$7.90
B200~$14.24

On AWS, your commitment and instance size matter more than the listed per-GPU rate.

GPUs and models available

The lineup covers A100 (P4d, P4de), H100 (P5), H200 (P5e, P5en), and B200 (P6) instances. AWS also sells its own Trainium and Inferentia chips.

Amazon Bedrock hosts managed foundation models, and SageMaker handles managed training. On raw EC2, you bring your own models and frameworks.

Best for

Enterprises already running on AWS get the most value, since GPUs sit beside S3, IAM, VPC, and compliance tooling you already use. It also suits teams that need global regions and formal support.

  • Existing AWS customers
  • Regulated production workloads
  • Teams using Bedrock or SageMaker

Watch out for

Expect the highest on-demand prices in this comparison. Because you rent eight GPUs at a time, a single-GPU experiment makes little sense here, and a cheaper provider above will serve you better.

Quotas and egress also bite. New accounts often need quota increases before they can launch P5 or P6 instances, and data transfer out adds per-GB charges that Lambda and RunPod waive.

Choosing the right GPU provider for your budget

The best deal in GPU cloud pricing depends on your workload, not the lowest number in a table. Vast.ai and RunPod win on raw hourly rate for experiments and fault-tolerant jobs. Lambda keeps billing simple with no egress fees. CoreWeave earns its premium on large clusters, while Google Cloud and AWS make sense mainly when you already live inside their ecosystems.

Before you commit, check three things: minimum cluster size, egress and storage charges, and whether your GPU sits idle for much of the day. A cheap rate on hardware you use 30% of the time costs more than a fair rate you use fully.

If your goal is shipping an agent or AI feature rather than operating GPUs, per-token pricing may beat any hourly rental. You can estimate your usage-based inference costs from request volume and token counts, then compare that number against the rentals above.

The Chronicle / Bartosz Neuman
Keep reading

More from Geodd.

All articles