Model-specific inference kernels

Free optimized runtimes.

Our AI agents use Mosaic to build faster inference kernels. Run them free in Docker with vLLM, SGLang, or TensorRT-LLM.

5 runtime releases / Free to use

  • Qwen 2.5 7B

    geodd/qwen-2.5-7b:latest
  • Llama 3.3 70B

    geodd/llama-3.3-70b:latest
  • Mistral Nemo

    geodd/mistral-nemo:latest
  • Qwen3 32B

    geodd/qwen3-32b:latest
  • Trinity Mini

    geodd/trinity-mini:latest

Before you run

Requires an NVIDIA GPU, driver, and NVIDIA Container Toolkit. GPU Docker setup

Performance depends on your hardware and workload. A public image tag alone does not guarantee reproduction of a recorded benchmark.

These images package inference runtimes with model-specific kernels, not downloads of Mosaic or our other internal kernel-development LLMs.