AA Intelligence Index v4.3
10 evaluations spanning agents, coding, knowledge and science. Rounded scores; ties do not mean identical capabilities.
Source and methodCompare major models on the same task. Every number has a test version, configuration and source.
Snapshot: 2026-09-08. Choose 2–4 models for a focused comparison. A dash means no value in this snapshot; some source values are rounded. Check each test's configuration and date.
10 evaluations spanning agents, coding, knowledge and science. Rounded scores; ties do not mean identical capabilities.
Source and method66 tasks, three runs, mean pass@1. Independent AA runs can differ from vendor tables.
— GLM-5.3
Source and method657 held-out tasks, v1.0.6. Objectives completed with partial credit; guardrail violations zero the task. Not the fully completed task rate.
— Claude Fable 5.1 · GPT-5.6 Sol
Source and methodLong engineering tasks. OpenAI comparison table; maximum across reported reasoning efforts, not equal compute cost.
— GLM-5.3 · DeepSeek V4 Pro 0813 · Grok 4.6 · Muse Spark 1.3 · Kimi K3
Source and methodGraduate science questions. OpenAI comparison table; a high score does not establish scientific discovery ability.
— GLM-5.3 · DeepSeek V4 Pro 0813 · Grok 4.6 · Muse Spark 1.3 · Kimi K3
Source and methodThe hardest FrontierMath tier. Not comparable with Tier 1–3; see the source for settings.
— Gemini 3.8 Flash · GLM-5.3 · DeepSeek V4 Pro 0813 · Grok 4.6 · Muse Spark 1.3 · Kimi K3
Source and methodIndex v4.3 changed on September 7. It is not comparable with July's v4.1 numbers. Start with your task when choosing a model.
Find a model →