综合能力榜 Overall Intelligence

融合多平台综合评测的模型综合能力榜(AISOTA 加权)

Data as of 2026-08-14 48 models
# Model Artificial AnalysisLiveBench AISOTA Score
1 SOTA Claude Fable 5 (with fallback) 2 sources 90.0 97.7 93.8
2 gpt-5.5-xhigh 93.0 93.0
3 Claude Opus 5 (max) 2 sources 95.0 90.7 92.8
4 GPT-5.6 Sol (max) 2 sources 85.0 95.3 90.2
5 smaug-agentic 86.0 86.0
6 Grok 4.6 (high) 2 sources 80.0 79.1 79.5
7 Kimi K3 (max) 2 sources 75.0 81.4 78.2
8 Qwen3.8 Max 2 sources 70.0 83.7 76.8
9 gpt-5.4-xhigh 74.4 74.4
10 Gemini 3.7 Flash (high) 2 sources 55.0 88.4 71.7
11 Muse Spark 1.2 (xhigh) 2 sources 65.0 76.7 70.8
12 gemini-3.1-pro-preview-high 67.4 67.4
13 GPT-5.6 Terra (max) 2 sources 60.0 72.1 66.0
14 claude-opus-4-8-max-effort 65.1 65.1
15 grok-4.5 62.8 62.8
16 claude-opus-4-7-xhigh-effort 60.5 60.5
17 DeepSeek V4 Pro 0813 (max) 2 sources 50.0 69.8 59.9
18 claude-sonnet-5-xhigh-effort 58.1 58.1
19 muse-spark-1.1-xhigh 55.8 55.8
20 gemini-3.5-flash-high 53.5 53.5
21 gpt-5.2-2025-12-11-high 51.2 51.2
22 claude-opus-4-6-thinking-auto-high-effort 48.8 48.8
23 deepseek-v4-flash-0731 46.5 46.5
24 gemini-3.6-flash-high 44.2 44.2
25 qwen3.7-max 41.9 41.9
26 gpt-5.2-codex 39.5 39.5
27 GLM-5.2 (max) 2 sources 45.0 32.6 38.8
28 GPT-5.6 Luna (max) 2 sources 40.0 37.2 38.6
29 Motif 3 35.0 35.0
30 claude-sonnet-4-6-thinking-auto-medium-effort 34.9 34.9
31 claude-opus-4-5-20251101-thinking-64k-high-effort 30.2 30.2
32 Inkling 2 sources 25.0 27.9 26.4
33 deepseek-v4-pro 25.6 25.6
34 kimi-k2.6-thinking 23.3 23.3
35 gpt-5.4-nano-xhigh 20.9 20.9
36 MiniMax-M3 2 sources 30.0 11.6 20.8
37 Nemotron 3 Ultra 20.0 20.0
38 qwen3.6-plus 18.6 18.6
39 kimi-k2.7-code 16.3 16.3
40 grok-build-0.1 14.0 14.0
41 Solar Open2 250B 10.0 10.0
42 gpt-5.4-mini-xhigh 9.3 9.3
43 Gemini 3.5 Flash-Lite 2 sources 15.0 2.3 8.7
44 deepseek-v4-flash 7.0 7.0
45 Muse Glimmer (high) 5.0 5.0
46 qwen3.6-27b 4.7 4.7
47 A.X-K2 0.0 0.0
48 grok-4.3 0.0 0.0