Comprehensive performance comparison across major AI labs · Data sourced from official model cards and benchmark publications
Last updated: February 19, 2026 · Click column headers to sort
Lab
Category
Model | ReasoningGPQA Diamond | KnowledgeMMLU | KnowledgeMMMLU | CodingHumanEval | CodingSWE-bench Ve… | MathAIME 2024 | MathMATH-500 | ReasoningHLE | ReasoningARC-AGI-2 | CodingLiveCodeBench | MultimodalMMMU | Instruction FollowingIFEval | ReasoningARC-AGI-1 | EQEQ-Bench 3 (… |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Claude 3.7 Sonnet Anthropic · 2025-02 | 68 | 88 | — | 93.7 | 62.3 | — | 80 | — | — | — | — | 86.5 | 32 | 1083.7 |
Claude Opus 4.1 Anthropic · 2025-04 | 79.2 | — | 89.5 | 92 | 72.8 | — | 94.8 | — | — | — | — | 91.4 | 60 | 1407.9 |
Claude Opus 4.5 Anthropic · 2025-10 | 87 | — | 90.8 | — | 80.9 | 100 | 97.5 | 32 | 37.6 | — | 78.5 | 93.1 | 76 | 1683.1 |
Claude Opus 4.6 Anthropic · 2026-02 | 91.3 | — | — | — | 80.8 | — | 97.6 | 40 | 75.2 | — | — | — | 85 | 1961 |
Claude Sonnet 4.5 Anthropic · 2025-10 | 83.4 | — | 89.1 | — | 82 | 87 | 96.2 | 19.8 | — | 64 | 75.4 | 92.8 | 64 | 1501.4 |
Claude Sonnet 4.6 Anthropic · 2026-02 | 74.1 | 79.1 | — | — | 79.6 | — | 97.8 | 19.1 | 58.3 | — | — | — | — | — |
DeepSeek V3 DeepSeek · 2024-12 | 59.1 | 88.5 | — | 82.6 | 42 | — | 90.2 | — | — | — | — | 87.5 | — | 1048.7 |
DeepSeek-R1 DeepSeek · 2025-01 | 71.5 | 90.8 | — | — | 49.2 | 79.8 | 97.3 | — | — | 65.9 | — | 83.3 | 45 | 1185.6 |
Gemini 2.5 Flash Google · 2025-04 | 70.5 | — | 85.1 | — | 49.6 | 83 | 91.2 | — | — | — | 68 | 87.5 | 42 | 1074.2 |
Gemini 2.5 Pro Google · 2025-03 | 84 | — | 89.2 | — | 63.8 | 92 | 95.2 | 21.6 | — | 70.4 | 75.8 | 89.5 | 63 | 1347.8 |
Gemini 3 Flash Google · 2025-12 | 82.1 | — | 87.5 | — | 62 | 90 | 93.5 | — | — | — | 74.2 | 89 | 55 | — |
Gemini 3 Pro Google · 2025-11 | 91.9 | — | 91.8 | — | 76.2 | 100 | 97 | 45.8 | 31.1 | 79.5 | 81 | 91.5 | 80 | 1629.5 |
Gemini 3.1 Pro Google · 2026-02 | 94.3 | — | 92.6 | — | 80.6 | — | — | 44.4 | 77.1 | — | — | — | — | — |
Gemini Deep Think Google · 2026-02 | 93.8 | — | — | — | — | — | — | 48.4 | 84.6 | — | — | — | — | — |
GPT-4.1 OpenAI · 2025-04 | 66.3 | 90.2 | — | 91.5 | 54.6 | — | 90.2 | — | — | — | — | 88.4 | 52 | 1137.7 |
GPT-4o OpenAI · 2024-05 | 53.6 | 88.7 | — | 90.2 | 38.4 | — | 76.6 | — | — | — | 69.1 | 86.1 | 21 | 1321.9 |
GPT-5 OpenAI · 2025-08 | 88.4 | — | — | — | 74.9 | 94.6 | 97.3 | 35.2 | 18 | — | 84.2 | 92.5 | 72 | 1456.9 |
GPT-5.1 OpenAI · 2025-08 | 88.1 | — | 91 | — | 76.3 | 100 | 97.8 | — | 17.6 | — | 85.4 | — | 75 | 1727.6 |
GPT-5.2 OpenAI · 2025-12 | 92.4 | — | 91 | — | 80 | 100 | 98.5 | — | 52.9 | — | 86.5 | — | 82 | 1637 |
Grok 3 xAI · 2025-02 | 68.2 | 88.5 | — | 89.3 | 48.5 | 86.7 | 93 | — | — | — | — | — | 34 | 1180.6 |
Grok 4 xAI · 2025-07 | 87.5 | — | 86.6 | — | 75 | 95 | 97 | 25.4 | — | — | 76.5 | — | 71 | 1131.7 |
Grok 4.20 xAI · 2026-02 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
Kimi K2 Thinking Moonshot · 2025-07 | 84.5 | — | — | — | 71.3 | 99.1 | 97 | 44.9 | — | 83.1 | — | — | 68 | 1622.5 |
Llama 4 Maverick Meta · 2025-04 | 69.8 | 88.4 | — | 85.5 | — | — | 86 | — | — | — | 73.4 | 90 | — | 833.2 |
Llama 4 Scout Meta · 2025-04 | 57.2 | 85.8 | — | 81.2 | — | — | 79.6 | — | — | — | 69.4 | 87.6 | — | 626.8 |
o3 OpenAI · 2025-04 | 83.3 | — | — | — | 69.1 | 98.4 | 96.7 | — | — | — | 82.9 | 91.8 | 87.5 | 1500 |
o4-mini OpenAI · 2025-04 | 81.4 | — | — | — | 68.1 | 99.5 | 96.3 | — | — | — | 79.6 | 90.2 | 72 | 1210.6 |
Qwen 3 Alibaba · 2025-04 | 71.1 | 89.5 | — | 88.4 | — | 87.5 | 95 | — | — | 62.5 | — | 88.2 | — | 1167.9 |
Data sourced from official model cards, blog posts, arXiv papers, and independent evaluations. Curated by Natural20 and BuseyBench.