编程能力榜 Coding Ability

融合多平台编程评测的模型编程能力榜(AISOTA 加权)

Data as of 2026-08-14 48 models
# Model Artificial Analysis AgenticLiveBench Coding AISOTA Score
1 SOTA smaug-agentic 97.7 97.7
2 Claude Opus 5 (max) 2 sources 95.0 93.0 94.0
3 Qwen3.8 Max 2 sources 85.0 88.4 86.7
4 claude-sonnet-5-xhigh-effort 86.0 86.0
5 Claude Fable 5 (with fallback) 2 sources 75.0 95.3 85.2
6 GPT-5.6 Sol (max) 2 sources 80.0 83.7 81.8
7 Grok 4.6 (high) 2 sources 90.0 72.1 81.0
8 Kimi K3 (max) 2 sources 70.0 90.7 80.3
9 muse-spark-1.1-xhigh 79.1 79.1
10 gpt-5.5-xhigh 74.4 74.4
11 GPT-5.6 Terra (max) 2 sources 65.0 69.8 67.4
12 Muse Spark 1.2 (xhigh) 2 sources 55.0 76.7 65.8
13 gpt-5.4-xhigh 65.1 65.1
14 DeepSeek V4 Pro 0813 (max) 2 sources 60.0 67.4 63.7
15 claude-opus-4-7-xhigh-effort 62.8 62.8
16 Gemini 3.7 Flash (high) 2 sources 40.0 81.4 60.7
17 gpt-5.2-codex 60.5 60.5
18 claude-opus-4-8-max-effort 58.1 58.1
19 GPT-5.6 Luna (max) 2 sources 50.0 53.5 51.8
20 grok-4.5 51.2 51.2
21 GLM-5.2 (max) 2 sources 45.0 55.8 50.4
22 claude-opus-4-6-thinking-auto-high-effort 48.8 48.8
23 gemini-3.5-flash-high 46.5 46.5
24 gpt-5.2-2025-12-11-high 44.2 44.2
25 kimi-k2.6-thinking 41.9 41.9
26 deepseek-v4-flash-0731 39.5 39.5
27 Motif 3 35.0 35.0
28 claude-sonnet-4-6-thinking-auto-medium-effort 32.6 32.6
29 Inkling 2 sources 25.0 37.2 31.1
30 gemini-3.6-flash-high 30.2 30.2
31 gemini-3.1-pro-preview-high 27.9 27.9
32 kimi-k2.7-code 25.6 25.6
33 gpt-5.4-nano-xhigh 23.3 23.3
34 Gemini 3.5 Flash-Lite 2 sources 10.0 34.9 22.4
35 qwen3.6-plus 20.9 20.9
36 Solar Open2 250B 20.0 20.0
37 qwen3.7-max 18.6 18.6
38 MiniMax-M3 2 sources 30.0 4.7 17.4
39 claude-opus-4-5-20251101-thinking-64k-high-effort 16.3 16.3
40 Nemotron 3 Ultra 15.0 15.0
41 gpt-5.4-mini-xhigh 14.0 14.0
42 grok-build-0.1 11.6 11.6
43 deepseek-v4-pro 9.3 9.3
44 qwen3.6-27b 7.0 7.0
45 A.X-K2 5.0 5.0
46 deepseek-v4-flash 2.3 2.3
47 Muse Glimmer (high) 0.0 0.0
48 grok-4.3 0.0 0.0