Model-specific inference kernels
Free optimized runtimes.
Our AI agents use Mosaic to build faster inference kernels. Run them free in Docker with vLLM, SGLang, or TensorRT-LLM.
Qwen 2.5 7B
geodd/qwen-2.5-7b:latestLlama 3.3 70B
geodd/llama-3.3-70b:latestMistral Nemo
geodd/mistral-nemo:latestQwen3 32B
geodd/qwen3-32b:latestTrinity Mini
geodd/trinity-mini:latest
Before you run
Requires an NVIDIA GPU, driver, and NVIDIA Container Toolkit. GPU Docker setup
Performance depends on your hardware and workload. A public image tag alone does not guarantee reproduction of a recorded benchmark.
These images package inference runtimes with model-specific kernels, not downloads of Mosaic or our other internal kernel-development LLMs.