AI Glossary term

Benchmark

A fixed set of tasks used to score models against each other. MMLU, SWE-bench, GPQA and friends.

Also called: eval, leaderboard

A fixed set of tasks used to score models against each other. MMLU, SWE-bench, GPQA and friends.

Why it mattersUseful for rough ranking. A benchmark win rarely predicts whether a model is good at your job.

See also Evals

Explore more AI terms

Browse the full glossary for plain-English definitions across models, agents, data, and safety.

Back to all terms