LLMs
1–25 of 435
Rank reflects Intelligence Index
| RankPosition based on the Intelligence Index. Tied scores share the same rank. | ModelModel name and publisher matched to the published benchmark result. | Composite score across reasoning, knowledge, mathematics, science, and coding evaluations. | Composite score for programming and software-engineering capability. | Composite score for tool use, planning, and autonomous task completion. | Maximum number of tokens the model can process in one request. | Median output generation throughput, measured in tokens per second. | Blended USD cost per 1 million tokens using a 3:1 input-to-output ratio. | Public release month and year for this model. | Whether downloadable model weights are open or access is closed. |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 53.4 | 81.6 | 58.0 | 1M | 69 tokens/s | $20 | Sep 2026 | Closed source | |
| 2 | 52.8 | 77.1 | 51.5 | 1M | 56 tokens/s | $20 | Sep 2026 | Closed source | |
| 3 | 50.7 | 78.0 | 56.2 | 1M | 54 tokens/s | $10 | Jul 2026 | Closed source | |
| 4 | 49.7 | 76.5 | 51.0 | 1M | 64 tokens/s | $20 | Jun 2026 | Closed source | |
| 5 | 48.2 | 76.5 | 55.7 | 1M | 232 tokens/s | $2 | Sep 2026 | Closed source | |
| 6 | 47.1 | 78.3 | 50.5 | 1M | 68 tokens/s | $8 | Jul 2026 | Closed source | |
| 7 | 44.9 | 74.8 | 53.4 | 1M | 70 tokens/s | $2.15 | Aug 2026 | Open source | |
| 8 | 44.4 | 76.8 | 53.4 | 500K | 58 tokens/s | $3 | Aug 2026 | Closed source | |
| 9 | 43.8 | 76.2 | 50.6 | 1M | 41 tokens/s | $6 | Jul 2026 | Open source | |
| 10 | 42.3 | 76.7 | 43.7 | 1M | 116 tokens/s | $4.5 | Jul 2026 | Closed source | |
| 11 | 42.2 | 73.1 | — | 256K | 72 tokens/s | $0.23 | Aug 2026 | Open source | |
| 12 | 42.0 | 74.3 | 42.6 | 1M | 59 tokens/s | $10 | May 2026 | Closed source | |
| 13 | 41.9 | 71.5 | 51.2 | 1M | 71 tokens/s | $0.24 | Aug 2026 | Open source | |
| 14 | 41.2 | 76.3 | 41.1 | 1M | 281 tokens/s | $1.5 | Sep 2026 | Closed source | |
| 15 | 40.7 | 73.6 | 39.5 | 1M | 50 tokens/s | $10 | Apr 2026 | Closed source | |
| 16 | 40.3 | 71.8 | 49.6 | 1M | 40 tokens/s | $3 | Aug 2026 | Closed source | |
| 17 | 40.0 | 71.9 | 50.4 | 983.6K | 41 tokens/s | $3 | Aug 2026 | Open source | |
| 18 | 39.8 | 72.2 | 44.0 | 1M | 238 tokens/s | $2 | Aug 2026 | Closed source | |
| 19 | 39.6 | 76.1 | 36.4 | 1M | 311 tokens/s | $1.5 | Aug 2026 | Closed source | |
| 20 | 39.1 | 72.4 | 42.1 | 500K | 59 tokens/s | $3 | Jul 2026 | Closed source | |
| 21 | 39.0 | 71.1 | — | 1.1M | 158 tokens/s | $5.63 | Mar 2026 | Closed source | |
| 22 | 38.6 | 68.8 | 39.4 | 1M | 63 tokens/s | $2.15 | Jun 2026 | Open source | |
| 23 | 38.6 | 74.9 | 37.3 | 922K | 90 tokens/s | $11.25 | Apr 2026 | Closed source | |
| 24 | 38.4 | 71.5 | 44.3 | 1M | 79 tokens/s | $4 | Jun 2026 | Closed source | |
| 25 | 37.5 | 71.4 | 42.7 | 1M | 120 tokens/s | $0.45 | Jul 2026 | Closed source | |
Rank uses Intelligence Index. Sort any index, context, speed, pricing, release date, or license; missing data always appears last. | |||||||||
Explore every published snapshot across coding, agentic search, reasoning, instruction following, long context, and safety evaluations.
Browse by category
31 of 31 benchmarks
A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.
73 models
A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.
44 models
An independently evaluated Tau3 banking benchmark from Artificial Analysis.
15 models
How important are skills for agents?
35 models
A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.
16 models
A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.
45 models
MRCR v2 slice focused on very long contexts at 128K-256K lengths.
2 models
Human-validated software engineering issues from real-world Python repositories.
33 models
Factoid questions spanning politics, science, technology, art, sports, geography, and music.
80 models
Abstract reasoning and pattern generalization on grid-based tasks.
203 models
Exceptionally difficult research-level mathematics problems.
62 models
Real-world multi-step tool-use evaluation through the Model Context Protocol.
35 models
Real-world cybersecurity evaluation of AI agents reproducing vulnerabilities with working proof-of-concept tests.
70 models
The official leaderboard for Terminal-Bench 3.0.
12 models
Expert-level questions across mathematics, science, and humanities, as published by the CAIS AI Dashboard.
60 models
Average of the text capability benchmarks published by the CAIS AI Dashboard.
60 models
Average of the vision capability benchmarks published by the CAIS AI Dashboard.
48 models
Average of the risk and safety benchmarks published by the CAIS AI Dashboard; lower is better.
9 models
Remote Labor Index automation rates published by the CAIS AI Dashboard.
13 models
Composite Artificial Analysis index across mathematics, science, coding, and reasoning evaluations.
633 models
Artificial Analysis composite index for programming and software-engineering capability.
256 models
Artificial Analysis composite index for tool use, planning, autonomy, and complex agentic workflows.
151 models
Artificial Analysis factual-knowledge and hallucination-resistance evaluation.
99 models
Graduate-level, expert-written science questions evaluated by Artificial Analysis.
612 models
Broad expert-level reasoning and knowledge evaluation published by Artificial Analysis.
607 models
Instruction-following benchmark results published by Artificial Analysis.
450 models
Scientific programming benchmark results published by Artificial Analysis.
167 models
Agentic terminal and software-engineering results published by Artificial Analysis.
236 models
Critical-points programming evaluation published by Artificial Analysis.
520 models
Long-context reasoning results published by Artificial Analysis.
516 models
Tool-agent interaction benchmark results published by Artificial Analysis.
440 models
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.