All articlesThe Chronicle / Geodd

12 LLM Benchmarks & Leaderboards to Compare Models in 2026

Picking a model off a single leaderboard score is how teams end up with an expensive, slow model that fails on their actual workload. LLM benchmarks measure narrow skills under controlled conditions. Your production traffic is neither narrow nor controlled. A model that tops a reasoning test can still stall on a long agentic task or blow your latency budget.

Here is the short answer. Use no single benchmark as your decision. Pair one general knowledge test, one coding benchmark such as SWE-bench, and one live human-preference leaderboard. Then add price and speed data from an independent tracker. That mix gives you a fair view of quality, cost, and throughput.

Below, you get 12 benchmarks and leaderboards worth checking in 2026. For each one, we explain what it measures, where it misleads, and when to trust it. At Geodd, we run models like GLM, DeepSeek, and GPT-OSS in production, so we also flag which scores carry over to real inference and which ones you should verify on your own prompts.

1. Geodd model catalog and live inference metrics

What it measures

Unlike the academic tests below, the Geodd catalog shows how models behave as running services, not how they score on a quiz. You get the models we host, including GLM-5.2, DeepSeek V4 Flash, GPT-OSS-120B, Gemma 4 31B IT, and the ByteDance Seed, Seedream, and Seedance families. All sit behind one OpenAI-compatible API. You can see:

  • Real-time token usage per model
  • Which regions serve each model (US-EAST and EU-NORTH today)
  • Serverless and dedicated GPU options

Who it is for

This fits engineers who already shortlisted models from public llm benchmarks and now need to test them on real prompts. It also suits teams with EU data requirements, since our handling is GDPR-ready and we apply Zero Data Retention.

If you are still deciding which model family to evaluate, start with the leaderboards below and come back here to validate.

Coding and agentic strengths

The catalog leans toward models people use for coding and agents. Our hardware-specific LLMs, Meridian for NVIDIA, Helix for AMD, and Stride for Tenstorrent, write and refine kernel code from observed production workloads. The goal is steady latency on long agentic runs, so you see fewer failed runs and retries.

That matters for any llm coding benchmark you reproduce. A score only helps if the model performs the same way at hour three of an agent loop as it does in the first minute.

Price and speed data

Pricing and usage show up per model, so you can run an llm model comparison with your own traffic instead of a vendor's demo. Because the API matches the OpenAI SDK, switching providers takes a single-line change. That lets you send the same prompt to several models and compare LLM models on cost per task, not just cost per token.

Speed numbers come from your own workloads, which makes them more relevant than a generic average.

Limitations

This is not a neutral scoreboard. We host a curated set of models, so you will not find every competitor here, and we do not publish an independent quality ranking.

A catalog tells you how a model runs, not how well it reasons, so pair it with a real benchmark.

Treat it as the last step of your evaluation. Use public leaderboards to narrow the field, then use live metrics to confirm the winner holds up in production.

2. LiveBench

What it measures

LiveBench is one of the latest LLM benchmarks built to resist training-data contamination. The maintainers refresh questions regularly from recent sources, and answers are scored against objective ground truth, with no LLM judge in the loop. Scores roll up from these categories:

  • Reasoning and math
  • Coding and agentic coding
  • Data analysis
  • Language
  • Instruction following

Who it is for

Pick it if you want a current LLM benchmarks view that models are unlikely to have memorized. It suits ML leads who distrust static test sets, and anyone who needs a fast, broad first filter before running deeper evaluations on a shortlist.

Coding and agentic strengths

The coding and agentic coding categories work as a lightweight llm coding benchmark. They show whether a model can write and fix code on problems it has not seen before. They are quicker to read than full repository tests like SWE-bench, but they are also shallower, so treat them as a signal and not a verdict.

Price and speed data

LiveBench reports quality scores only. It does not publish token prices, latency, or throughput. Pair it with an independent tracker, or measure cost and speed on your own traffic, before you commit to a model.

Limitations

Because questions rotate, positions on this llm benchmarks leaderboard shift over time, and older results are not directly comparable to newer ones. Tasks are also short, so they say little about how a model holds up across an hour-long agent run.

A fresh test beats a famous test, but neither replaces your own prompts.

Use LiveBench to narrow the field, then confirm the finalists on your real workload.

3. Artificial Analysis LLM leaderboard

What it measures

Artificial Analysis puts quality, price, and speed on one page. Its core number is a composite Intelligence Index that blends scores from several evaluations, and the team runs those tests itself instead of repeating vendor claims. The result is an ai benchmarks ranking you can sort by more than one axis.

Who it is for

Choose it when you need to compare LLM models across providers with a budget in mind. It suits CTOs and platform leads who must defend a pick with numbers. Engineers also use it for a quick llm compare pass before they test anything.

Coding and agentic strengths

The index folds in coding and agentic evaluations, so a high score reflects more than trivia recall. Coding is still one slice of an average, though. If your workload is code, open the coding breakdown and ignore the headline number.

Price and speed data

This is where the site beats most rivals. Each model lists the following:

MetricWhat it tells you
Price per million tokensInput and output cost
Output tokens per secondGeneration throughput
Time to first tokenResponsiveness
Context windowMaximum input size

These llm performance benchmarks come from public provider endpoints, so the same model can show different speeds depending on the host.

Limitations

Composite scores hide weak spots. A model can average well and still fail at the one task you care about. The index methodology also changes between versions, so only compare scores within the same version.

A blended score ranks models in general, not for your workload.

Speed figures reflect standard test prompts and the time of day they were measured. Treat them as a baseline to beat, then confirm on your own traffic.

4. BenchLM leaderboard

What it measures

BenchLM is an aggregator. It does not write its own tests. It pulls published results from public evaluations and rolls them into category scores and an overall rank. Typical categories include:

  • Agentic tasks
  • Coding
  • Reasoning and math
  • Knowledge
  • Multimodal

Who it is for

Use it when you want one page that summarizes many AI LLM benchmarks without opening ten sites. It suits product managers and engineers who need a quick LLM model benchmarks overview before a deeper review.

Coding and agentic strengths

Coding and agentic categories sit side by side, so you can check whether a model's coding score matches its agent score. That matters because coding LLM benchmarks often disagree with each other. Click through to the underlying tests, such as SWE-bench, before you trust a category average.

Price and speed data

Models list pricing next to scores, which makes a rough cost-versus-quality cut easy. Speed coverage is thinner than on a dedicated tracker, so confirm throughput elsewhere or on your own traffic.

Limitations

Aggregation inherits every flaw of its sources. Coverage can be uneven across models, and weighting choices decide the final rank, so read the methodology page before you cite a position. Vendor-reported numbers can also slip in, which makes self-reported scores worth a second look.

An aggregated rank is only as honest as the tests underneath it.

Treat BenchLM as a map of the field, then drill into the source benchmarks for the models you actually plan to run.

5. Humanity's Last Exam

What it measures

Humanity's Last Exam, or HLE, was built by the Center for AI Safety and Scale AI to replace tests that top models had already saturated. It holds roughly 2,500 expert-written questions across math, science, and the humanities. Key traits:

  • Closed-ended answers, either multiple choice or short exact match
  • Questions screened against frontier models, so easy ones were dropped
  • A small share that requires reading an image

Who it is for

Look at it if you track frontier reasoning ability or evaluate models for research-heavy work. It is a poor first filter for a typical product team, because the questions look nothing like support tickets or pull requests from your own repo.

Coding and agentic strengths

HLE is not a coding benchmark. Few questions involve code, so it says almost nothing about the coding AI benchmarks you care about, such as SWE-bench. Some leaderboards report scores with tools enabled, which hints at agentic skill, but check which setting a number uses before comparing.

Price and speed data

The public leaderboards report accuracy only. Cost and latency are missing, so pull them from a tracker or measure them yourself. Reasoning-heavy models that score well here often burn many output tokens, which raises your cost per answer.

Limitations

Scores near the top are hard to interpret. Public questions can leak into training data, and audits have flagged wrong answer keys in some items. Obscure-fact questions also reward memorization over useful work.

A high HLE score proves a model is good at expert quizzes, not that it is good at your job.

Use it as a ceiling check on reasoning, never as the number that decides your pick.

6. MMLU-Pro

What it measures

MMLU-Pro is a harder rebuild of the classic MMLU test, released by TIGER-Lab in 2024. It keeps multiple-choice knowledge questions but raises the difficulty:

  • About 12,000 questions across 14 subjects
  • Ten answer options instead of four, so guessing pays far less
  • Noisy and trivial MMLU items removed, with more reasoning-heavy problems added

Who it is for

Choose it for a quick llm model comparison on general knowledge and reasoning, especially among open-weight models. It suits teams that want a cheap, widely reported number that appears on almost every model card.

Coding and agentic strengths

Weak. A computer science category exists, but it asks multiple-choice questions about code, not writing or fixing it. For an llm coding benchmark, use SWE-bench or LiveCodeBench instead. Nothing here tests tool use or multi-step agent behavior.

Price and speed data

MMLU-Pro reports accuracy only. Cost and latency have to come from a tracker or your own traffic. Scores often use chain-of-thought prompting, so a strong result can mean many output tokens per question and a bigger bill.

Limitations

Top models now cluster in a narrow band, so gaps of a point or two are within noise. Prompt format and answer parsing also shift results, and vendor-reported numbers rarely use identical settings. Public questions can leak into training data too.

MMLU-Pro separates weak models from strong ones, but it no longer separates strong from stronger.

Use it as a sanity filter at the start of your shortlist, then move on to tests that resemble your workload.

7. SWE-bench Pro and SWE-bench-Live

What it measures

Both benchmarks give a model a real GitHub issue and ask for a patch that passes the project's tests. SWE-bench Pro, from Scale AI, uses longer multi-file tasks drawn from public, commercial, and held-out repositories. SWE-bench-Live, from Microsoft, adds fresh issues on a rolling basis to limit contamination.

  • Pro: harder patches, private code the model is unlikely to have seen
  • Live: regular refreshes, isolated environment per task

Who it is for

Start here if you are building code agents or pull request automation. It is the closest public proxy for repo-level work, and it belongs at the top of any shortlist built from llm coding benchmarks. Teams doing autocomplete or chat help will find it overkill.

Coding and agentic strengths

This is the strongest agentic signal on the list. Models must read a codebase, edit several files, and run tests. Scores depend heavily on the agent harness, meaning the tools, prompts, and retries wrapped around the model, so check which one a result used. Long tasks also reward models that stay consistent across many steps.

Price and speed data

Most listings report resolve rate only. A few add cost per task or step counts, but not consistently. Agent runs burn many tokens, so estimate cost per resolved issue on your own traffic before you compare two models on score alone.

Limitations

Tasks are mostly bug fixes and small features, not architecture or design work. Vendors also publish numbers from custom harnesses, which makes cross-vendor comparisons shaky.

A SWE-bench score describes a model plus its harness, not the model alone.

Reproduce the setup on a slice of your own repo before you trust a ten-point gap.

8. LiveCodeBench Pro

What it measures

LiveCodeBench Pro scores competitive programming, not repository work. Problems come from Codeforces, ICPC, and IOI contests and are added continuously to limit contamination. Olympiad medalists tag each one by skill and difficulty:

  • Easy, medium, and hard tiers
  • Tags such as knowledge-heavy, logic-heavy, and observation-heavy

Who it is for

Pick it if you train or evaluate models for algorithmic reasoning, or you want a harder check than saturated tests. Teams building CRUD features or review bots can skip it and spend the time elsewhere.

Coding and agentic strengths

It is the sharpest llm coding benchmark for pure problem solving. At launch, top models handled medium problems but solved none of the hard tier, so it still separates strong from stronger. It does not test tool use, file edits, or multi-step agent behavior.

Price and speed data

Leaderboards report pass rates only. Reasoning models often spend thousands of thinking tokens per problem, so cost per solved problem can dwarf the sticker price. Pull token counts from a tracker or your own runs.

Limitations

Contest puzzles are not software engineering. A model can ace them and still fail a messy bug fix in your codebase. Scores also shift with attempt count and whether tools are allowed.

Contest skill predicts puzzle solving, not shipping code.

Pair it with SWE-bench for repo-level evidence, and treat both as inputs, not verdicts.

9. ARC-AGI and FrontierMath

What it measures

These two LLM benchmarks test novel reasoning, not recall. ARC-AGI, created by François Chollet, shows a model a few colored grid puzzles and asks it to infer the hidden rule, then apply it to a new grid. FrontierMath, from Epoch AI, holds original, unpublished math problems with exact answers that can be checked automatically.

Who it is for

Reach for them if you follow frontier reasoning progress or build research tools. A team shipping a support bot or a summarizer can skip both, because the scores say little about everyday task quality.

Coding and agentic strengths

Neither is a coding benchmark. Newer ARC-AGI versions add interactive tasks that hint at agentic adaptation, but FrontierMath has no repository work or tool use. For an llm coding benchmark, stay with SWE-bench.

Price and speed data

ARC-AGI stands out here. Its leaderboard plots score against cost per task, which exposes models that buy accuracy with huge reasoning budgets. FrontierMath reports accuracy only, so pull token counts and latency from a tracker or your own runs.

Limitations

Both are narrow and costly to run, so fewer models get tested and new results arrive late. Most FrontierMath items are private, so you cannot inspect or reproduce them. OpenAI funded its creation, which Epoch has disclosed, so read vendor claims with that in mind.

Abstract reasoning scores show a model's ceiling, not its usefulness.

10. Chatbot Arena

What it measures

Chatbot Arena, now run as LMArena, ranks models by human preference. You see two anonymous answers to the same prompt, pick the better one, and the votes feed an Elo-style rating. Among LLM benchmarks, it is the hardest to game directly, because the prompts are whatever real users decide to type.

Who it is for

Use it when you build chat products, assistants, or writing tools where tone and helpfulness matter. It also works as a check on static test rankings, since it shows whether benchmark winners feel better to actual people.

Coding and agentic strengths

Separate coding and web-development categories let you filter votes to programming prompts. They capture what developers prefer to read, not whether the code passes tests. Agentic work is barely covered, so do not use it in place of SWE-bench.

Price and speed data

Ratings are the product here. The leaderboard offers little on cost, latency, or throughput, so pull those from an independent tracker or your own traffic. Verbose, reasoning-heavy models can win votes while staying slow and expensive in production.

Limitations

Voters tend to reward longer, well-formatted answers, and style-control adjustments only partly correct that. The crowd is self-selected and the prompts are mostly short, so results say little about long documents or multi-step jobs. Vendors can also test private variants before launch, so confirm the listed model is the one you can actually call.

A vote measures what people like, not what is correct.

Treat it as a taste signal for chat quality, then verify accuracy with a task-based test.

11. Berkeley Function-Calling Leaderboard

What it measures

The Berkeley Function-Calling Leaderboard, or BFCL, comes from UC Berkeley's Gorilla project. It tests whether a model calls tools correctly. It checks function names, arguments, and types against a schema instead of judging prose, so scoring is objective. Test categories include:

  • Single, multiple, and parallel function calls
  • Irrelevance detection, where the right move is to call nothing
  • Multi-turn conversations that carry state
  • Agentic tasks such as web search and memory

Who it is for

Pick it if your product depends on structured output and tool use, such as support agents, workflow automation, or API-driven assistants. Among LLM benchmarks, it is the one teams skip and then regret skipping. Teams building pure chat or summarization can safely ignore it.

Coding and agentic strengths

It is not a code-writing test, so do not treat it as an LLM coding benchmark. Its value is agentic. It shows whether a model picks the right tool, fills arguments without inventing values, and keeps state across turns. Those failures break many agent runs, which makes BFCL a good complement to SWE-bench.

Price and speed data

The leaderboard lists cost and latency columns next to accuracy, which most academic tests skip. Those figures come from the maintainers' own setup, so use them for relative ranking only. Agents chain many calls, so measure latency per tool call on your own endpoint.

Limitations

Scores shift with prompt-based versus native tool-calling mode, and the benchmark's clean schemas are tidier than yours. Real APIs have vague descriptions, overlapping tools, and auth errors. Top models also cluster tightly on the simple categories.

A model that handles clean function schemas can still fumble your messy ones.

Test finalists with your actual tool definitions, and log every malformed call so you can see where they fail.

12. Vellum LLM leaderboard

What it measures

Vellum publishes a curated comparison table of recent frontier models. It pulls results from tests such as GPQA Diamond, SWE-bench, and Humanity's Last Exam, then adds tool-use and agentic evaluations. It favors new releases and drops older models, so the list stays short and current.

Who it is for

Use it when you want a fast llm model comparison across proprietary and open-weight models on one page. It suits teams that need a pre-meeting overview without opening a dozen sites.

Coding and agentic strengths

Coding and agentic scores sit in their own columns. You can read llm coding benchmarks like SWE-bench next to tool-use results for the same model. The numbers come from other sources, so click through to the original test and check which harness produced each score.

Price and speed data

Unlike most academic tests, each model carries cost and speed figures you can scan:

ColumnWhat you use it for
Cost per million tokensBudget estimates
Output tokens per secondThroughput
Time to first tokenResponsiveness
Context windowLong-input fit

Speed values depend on the provider that served the test, so treat them as a baseline and re-measure on your own traffic.

Limitations

Much of the data is reported by model vendors, and settings differ between them. Coverage also stops at recent models, so the older model you run in production may be missing entirely.

A curated table saves you time, but it cannot replace a test on your own prompts.

Use Vellum to build the shortlist, then validate the finalists yourself.

Choosing the right benchmark for your model

No leaderboard picks a model for you. The best LLM benchmarks each answer one narrow question, so match the test to your workload. Building code agents? Start with SWE-bench and the Berkeley Function-Calling Leaderboard. Shipping a chat product? Check Chatbot Arena. Watching your budget? Use Artificial Analysis for price and speed.

Then do the part no public score can do. Take your top two or three models, run your own prompts, and measure cost per task and latency on real traffic. Public numbers narrow the field. Your workload decides the winner.

When you are ready to test your shortlist, browse the Geodd model catalog and send the same prompt to several models through one OpenAI-compatible API.

The Chronicle / Bartosz Neuman
Keep reading

More from Geodd.

All articles