· State of the Art

每日 AI 决策情报

多来源榜单融合为一个分数,每天告诉你哪个模型才是真正的 SOTA。

数据截止 2026-08-14更新于 8月14日 23:27

综合能力榜

融合多平台综合评测,汇成一个 AISOTA 综合分。

排名模型Artificial AnalysisLiveBenchAISOTA 分
1SOTAClaude Fable 5 (with fallback)2 sources90.097.793.8
22gpt-5.5-xhigh93.093.0
33Claude Opus 5 (max)2 sources95.090.792.8
4GPT-5.6 Sol (max)2 sources85.095.390.2
5smaug-agentic86.086.0
6Grok 4.6 (high)2 sources80.079.179.5
7Kimi K3 (max)2 sources75.081.478.2
8Qwen3.8 Max2 sources70.083.776.8
9gpt-5.4-xhigh74.474.4
10Gemini 3.7 Flash (high)2 sources55.088.471.7

编程能力榜

融合多平台编程评测,汇成一个 AISOTA 综合分。

排名模型Artificial Analysis AgenticLiveBench CodingAISOTA 分
1SOTAsmaug-agentic97.797.7
22Claude Opus 5 (max)2 sources95.093.094.0
33Qwen3.8 Max2 sources85.088.486.7
4claude-sonnet-5-xhigh-effort86.086.0
5Claude Fable 5 (with fallback)2 sources75.095.385.2
6GPT-5.6 Sol (max)2 sources80.083.781.8
7Grok 4.6 (high)2 sources90.072.181.0
8Kimi K3 (max)2 sources70.090.780.3
9muse-spark-1.1-xhigh79.179.1
10gpt-5.5-xhigh74.474.4

性价比榜

单位成本能买到的智能,基于输入+输出价格中位数。

排名模型AA IntelligenceOpenRouter PriceAISOTA 分
1SOTAGemini 3.7 Flash (high)2 sources55.065.060.0
22GPT-5.6 Luna (max)2 sources40.074.857.4
33DeepSeek V4 Pro 0813 (max)2 sources50.059.254.6
4Claude Opus 5 (max)2 sources95.012.253.6
5MiniMax-M32 sources30.072.851.4
6Grok 4.6 (high)2 sources80.020.450.2
7Claude Fable 5 (with fallback)2 sources90.06.148.0
8GPT-5.6 Sol (max)2 sources85.08.846.9
9Muse Spark 1.2 (xhigh)2 sources65.026.245.6
10Qwen3.8 Max2 sources70.020.145.0

调用量榜

近 7 天各模型 API 请求量,数据来自 OpenRouter 公开排行榜。

数据来源:OpenRouter Usage
排名模型OpenRouter UsageAISOTA 分
1SOTAdeepseek-v4-flash843.17M99.8
22gemini-2-5-flash-lite254.84M99.5
33gpt-5-6-luna237.25M99.3
4gemini-2-5-flash152.47M99.0
5hy3132.91M98.8
6gpt-4o-mini119.03M98.6
7gemini-3-1-flash-lite117.7M98.3
8deepseek-v4-pro107.73M98.1
9gemini-3-flash104.59M97.9
10gemma-4-31b-it104.07M97.6

AISOTA 评分机制

每个模型按榜单的数据来源逐一打分。在每个来源内,模型按得分排名并换算成 0–100 的百分位。模型的 AISOTA 分就是它在各来源百分位的加权平均。SOTA 标记当天第一名。

数据每天自动从公开 API 抓取。空白单元格表示该来源没有收录这个模型。