New Frontier data and RL environments, off the shelf

Benchmarks

Greatness isn't accidental. How we measure it shouldn't be either. If we want AGI that builds billion-dollar enterprises and globe-spanning infrastructure, we need benchmarks that test for intelligence and sophistication.

This is our ranking of models, measured by their capacity for rigorous reasoning and real-world mastery.
View by :
Professional Graphical Reasoning

Chartography

Chartography is our benchmark for professional chart understanding. It tests whether frontier models can read the Kaplan-Meier curves, candlestick charts, contour maps, Sankey diagrams, Bode plots, and other specialized graphics that professionals use to make real decisions every day. It evaluates visual perception, domain-aware interpretation, and multi-step graphical reasoning.

Rank
Model
Score
Claude Fable 5.1 (Adaptive/Max)
46.2
%
Claude Fable 5.1 (Adaptive/Max)
GPT 5.6 Sol (Max)
45
%
GPT 5.6 Sol (Max)
Claude Fable 5.1
44.1
%
Claude Fable 5.1
Gemini 3.7 Flash
43
%
Gemini 3.7 Flash
Gemini 3.8 Flash
42.5
%
Gemini 3.8 Flash
Gemini 3.8 Flash (High)
40.9
%
Gemini 3.8 Flash (High)
Gemini 3.7 Flash (High)
40.4
%
Gemini 3.7 Flash (High)
GPT 5.6 Sol
39.5
%
GPT 5.6 Sol
Gemini 3.5 Flash
35.9
%
Gemini 3.5 Flash
Claude Fable 5 (Adaptive/Max)
34.8
%
Claude Fable 5 (Adaptive/Max)
GPT 5.6 Terra (Max)
34
%
GPT 5.6 Terra (Max)
Gemini 3.6 Flash
34
%
Gemini 3.6 Flash
Muse Spark 1.2 (xHigh)
32.1
%
Muse Spark 1.2 (xHigh)
Muse Spark 1.2
31.5
%
Muse Spark 1.2
GPT 5.6 Luna (Max)
31.4
%
GPT 5.6 Luna (Max)
GPT 5.5 (xHigh)
31
%
GPT 5.5 (xHigh)
GPT 5.4 (xHigh)
29.6
%
GPT 5.4 (xHigh)
Claude Fable 5
29.5
%
Claude Fable 5
Qwen 3.8 Max
29.1
%
Qwen 3.8 Max
GPT 5.5
28.5
%
GPT 5.5
Muse Spark 1.3 (xHigh)
27.6
%
Muse Spark 1.3 (xHigh)
Claude Opus 5 (Adaptive/Max)
27.3
%
Claude Opus 5 (Adaptive/Max)
GPT 5.6 Terra
26.6
%
GPT 5.6 Terra
Kimi K3 (Max)
26.6
%
Kimi K3 (Max)
Gemini 3.1 Pro
26.1
%
Gemini 3.1 Pro
Claude Opus 5
25.7
%
Claude Opus 5
Muse Spark 1.1 (xHigh)
24.4
%
Muse Spark 1.1 (xHigh)
Muse Spark 1.3
24.4
%
Muse Spark 1.3
Muse Spark 1.1
23.6
%
Muse Spark 1.1
Grok 4.6
21.9
%
Grok 4.6
GPT 5.6 Luna
21.4
%
GPT 5.6 Luna
Grok 4.6 (xHigh)
20.5
%
Grok 4.6 (xHigh)
Muse Glimmer 30B
17.9
%
Muse Glimmer 30B
Qwen 3.8 Flash
17.4
%
Qwen 3.8 Flash
Muse Glimmer 30B (xHigh)
17.2
%
Muse Glimmer 30B (xHigh)
Grok 4.5
17
%
Grok 4.5
Claude Sonnet 5 (Adaptive/Max)
16.6
%
Claude Sonnet 5 (Adaptive/Max)
Claude Opus 4.7 (Adaptive/Max)
16.5
%
Claude Opus 4.7 (Adaptive/Max)
GLM 5.3 Flash
16.3
%
GLM 5.3 Flash
Claude Opus 4.8 (Adaptive/Max)
15.9
%
Claude Opus 4.8 (Adaptive/Max)
Qwen 3.5 Plus
15.9
%
Qwen 3.5 Plus
Qwen 3.7 Plus
15.8
%
Qwen 3.7 Plus
GPT 5.4
14.1
%
GPT 5.4
Claude Opus 4.7
13.6
%
Claude Opus 4.7
Grok 4.3 (High)
12.9
%
Grok 4.3 (High)
Kimi K2.5
12.6
%
Kimi K2.5
Claude Sonnet 5
12.2
%
Claude Sonnet 5
Kimi K2.6
12.2
%
Kimi K2.6
Grok 4.3
11.7
%
Grok 4.3
Claude Opus 4.8
11.3
%
Claude Opus 4.8
DeepSeek V4 Flash Vision (experimental) (Max)
11.2
%
DeepSeek V4 Flash Vision (experimental) (Max)
Gemini 3.5 Flash-Lite
10.2
%
Gemini 3.5 Flash-Lite
Inkling (xHigh)
10
%
Inkling (xHigh)
Inkling
10
%
Inkling
DeepSeek V4 Flash Vision (experimental)
9.5
%
DeepSeek V4 Flash Vision (experimental)
Inkling Small (xHigh)
9.1
%
Inkling Small (xHigh)
Mistral Large 3
9
%
Mistral Large 3
Inkling Small
8.7
%
Inkling Small
Long-Context Agentic Instruction Following

HANDBOOK.md Agents

Can an agent follow a 100-page company handbook inside an enterprise RL environment?

HANDBOOK.md is a benchmark for long-context agentic instruction following, modeled on how professionals follow corporate policy in their day-to-day work. Each task is a unique RL environment with internal tools and external MCP servers, across five enterprise domains.

Rank
Model
Score
Claude Fable 5 (Adaptive/Max)
36.2
%
Claude Fable 5 (Adaptive/Max)
Claude Fable 5
34.2
%
Claude Fable 5
Claude Opus 5 (Adaptive/Max)
32.3
%
Claude Opus 5 (Adaptive/Max)
Claude Opus 5
29.6
%
Claude Opus 5
Grok 4.6 (xHigh)
27.3
%
Grok 4.6 (xHigh)
DeepSeek V4 Pro (Max)
26.9
%
DeepSeek V4 Pro (Max)
Grok 4.6
25.8
%
Grok 4.6
GPT 5.6 Sol (Max)
23.5
%
GPT 5.6 Sol (Max)
DeepSeek V4 Flash Vision (experimental) (Max)
23.1
%
DeepSeek V4 Flash Vision (experimental) (Max)
Claude Opus 4.8 (Adaptive/Max)
21.9
%
Claude Opus 4.8 (Adaptive/Max)
GPT 5.6 Sol
21.5
%
GPT 5.6 Sol
GPT 5.5 (xHigh)
21.5
%
GPT 5.5 (xHigh)
GPT 5.5
21.5
%
GPT 5.5
DeepSeek V4 Flash Vision (experimental)
19.6
%
DeepSeek V4 Flash Vision (experimental)
DeepSeek V4 Pro
19.6
%
DeepSeek V4 Pro
Claude Opus 4.8
18.9
%
Claude Opus 4.8
Qwen 3.8 Max
16.5
%
Qwen 3.8 Max
Grok 4.5 (High)
15.8
%
Grok 4.5 (High)
Muse Spark 1.3 (xHigh)
15.4
%
Muse Spark 1.3 (xHigh)
GLM 5.3
15
%
GLM 5.3
Muse Spark 1.1 (xHigh)
13.5
%
Muse Spark 1.1 (xHigh)
Muse Spark 1.2 (xHigh)
13.1
%
Muse Spark 1.2 (xHigh)
Muse Spark 1.3
12.7
%
Muse Spark 1.3
GLM 5.2
12.7
%
GLM 5.2
Kimi K3 (Max)
11.9
%
Kimi K3 (Max)
Gemini 3.7 Flash
11.9
%
Gemini 3.7 Flash
Gemini 3.8 Flash
11.2
%
Gemini 3.8 Flash
Muse Spark 1.2
11.2
%
Muse Spark 1.2
Gemini 3.5 Flash (High)
11.2
%
Gemini 3.5 Flash (High)
Gemini 3.8 Flash (High)
10.8
%
Gemini 3.8 Flash (High)
Gemini 3.7 Flash (High)
10.8
%
Gemini 3.7 Flash (High)
GLM 5.3 Flash
10.4
%
GLM 5.3 Flash
Claude Sonnet 4.6 (Adaptive/Max)
10.4
%
Claude Sonnet 4.6 (Adaptive/Max)
Gemini 3.1 Pro
10
%
Gemini 3.1 Pro
GLM 5.2 (xHigh)
10
%
GLM 5.2 (xHigh)
Gemini 3.5 Flash
9.2
%
Gemini 3.5 Flash
DeepSeek V4 Pro (preview) (xHigh)
9.2
%
DeepSeek V4 Pro (preview) (xHigh)
Qwen 3.7 Max
8.5
%
Qwen 3.7 Max
Hy3 (High)
7.7
%
Hy3 (High)
Hy3
7.7
%
Hy3
Claude Sonnet 4.6
7.7
%
Claude Sonnet 4.6
DeepSeek V4 Flash (xHigh)
7.3
%
DeepSeek V4 Flash (xHigh)
DeepSeek V4 Flash (preview) (xHigh)
7.3
%
DeepSeek V4 Flash (preview) (xHigh)
DeepSeek V4 Flash (preview)
7.3
%
DeepSeek V4 Flash (preview)
DeepSeek V4 Flash
7.3
%
DeepSeek V4 Flash
Kimi K2.6
6.9
%
Kimi K2.6
DeepSeek V4 Pro (preview)
6.9
%
DeepSeek V4 Pro (preview)
Gemini 3.6 Flash (High)
5
%
Gemini 3.6 Flash (High)
Muse Glimmer 30B (xHigh)
3.5
%
Muse Glimmer 30B (xHigh)
Muse Glimmer 30B
3.5
%
Muse Glimmer 30B
Inkling Small
3.1
%
Inkling Small
Gemini 3.5 Flash-Lite (High)
3.1
%
Gemini 3.5 Flash-Lite (High)
Inkling (xHigh)
2.3
%
Inkling (xHigh)
Grok 4.3 (High)
1.9
%
Grok 4.3 (High)
Nemotron 3 Ultra
1.5
%
Nemotron 3 Ultra
Inkling Small (xHigh)
1.2
%
Inkling Small (xHigh)
Grok 4.3
0.8
%
Grok 4.3
Nemotron Lightning 3.5 30B A3B
0
%
Nemotron Lightning 3.5 30B A3B
Antidote / Everyday

Antidote: Everyday Edition

Some leaderboards reward the answers that convince you in two seconds. Antidote rewards the answer that leaves you better off a month later. It's our real-world AI leaderboard, graded by doctors, lawyers, and engineers who read every word, check every citation, and run every line of code.

Rank
Model
elo score (95% ci)
Claude Fable 5
1099
(
1085
-
1112
)
Gemini 3.1 Pro
1091
(
1081
-
1101
)
Gemini 3.8 Flash
1087
(
1070
-
1104
)
Gemini 3.6 Flash
1083
(
1069
-
1097
)
Gemini 3.5 Flash
1079
(
1068
-
1090
)
Kimi K3
1077
(
1063
-
1091
)
Gemini 3.7 Flash
1075
(
1059
-
1090
)
GLM 5.3
1067
(
1051
-
1083
)
Claude Fable 5.1
1065
(
1048
-
1083
)
Claude Opus 5
1055
(
1041
-
1069
)
Qwen 3.7 Max
1051
(
1040
-
1061
)
Kimi K2.6
1045
(
1036
-
1054
)
Claude Opus 4.7
1042
(
1031
-
1053
)
Claude Opus 4.6
1042
(
1032
-
1052
)
GLM 5.2
1040
(
1027
-
1052
)
Claude Opus 4.8
1030
(
1018
-
1043
)
Muse Spark 1.2
1028
(
1011
-
1046
)
GPT 5.6 Sol
1027
(
1014
-
1041
)
Hy3
1022
(
1007
-
1037
)
DeepSeek V4 Pro
1018
(
1003
-
1033
)
DeepSeek V4 Flash Vision (experimental)
1015
(
999
-
1030
)
Muse Spark 1.1
1014
(
998
-
1030
)
GPT 5.5
1013
(
1000
-
1025
)
Claude Sonnet 4.6
1012
(
1000
-
1024
)
DeepSeek V4 Pro (preview)
1007
(
995
-
1019
)
Kimi K2.5
1006
(
995
-
1018
)
Grok 4.20 Beta
1001
(
989
-
1012
)
Qwen 3.5 Plus
999
(
987
-
1011
)
Hy4 Preview
992
(
974
-
1009
)
Gemini 3.5 Flash-Lite
991
(
977
-
1005
)
Inkling
991
(
977
-
1005
)
Qwen 3.8 Max
985
(
970
-
1000
)
Grok 4.5
973
(
960
-
986
)
DeepSeek V4 Flash (preview)
966
(
954
-
978
)
Grok 4.6
966
(
950
-
981
)
DeepSeek V3.2
954
(
940
-
967
)
Grok 4.3
950
(
939
-
960
)
Mistral Large 3
947
(
938
-
957
)
Ernie 5.1
947
(
936
-
958
)
Qwen 3.8 Flash
944
(
927
-
962
)
Claude Haiku 4.5
932
(
919
-
944
)
Muse Glimmer 30B
931
(
915
-
947
)
GPT 5.4 Mini
929
(
917
-
941
)
Gemma 3 12B
893
(
880
-
905
)
Nemotron Lightning 3.5 30B A3B
888
(
871
-
904
)
Ernie 4.5 300B
863
(
850
-
876
)
Nova 2 Pro
819
(
809
-
830
)
Creative, Business, and Everyday Writing

Hemingway-bench

Most AI writing benchmarks reward surface-level signals: elaborate metaphors and prose that looks impressive at a glance. Hemingway-bench rewards writing that is actually good.

Our leaderboard is judged by professional writers who evaluate creative writing, business writing, and everyday writing tasks for taste, originality, coherence, and emotional intelligence.

Rank
Model
elo score (95% ci)
Claude Fable 5
1110
(
1091
-
1130
)
Gemini 3.7 Flash
1105
(
1084
-
1126
)
Claude Fable 5.1
1095
(
1071
-
1118
)
Gemini 3.8 Flash
1093
(
1069
-
1117
)
Gemini 3.6 Flash
1075
(
1055
-
1095
)
Gemini 3.1 Pro
1073
(
1057
-
1088
)
Gemini 3.5 Flash
1071
(
1054
-
1088
)
Kimi K3
1069
(
1049
-
1089
)
GPT 5.6 Sol
1057
(
1037
-
1076
)
Gemini 3 Pro
1056
(
1031
-
1081
)
GLM 5.3
1053
(
1030
-
1076
)
Gemini 3 Flash
1052
(
1037
-
1066
)
Claude Opus 5
1047
(
1026
-
1068
)
DeepSeek V4 Pro
1046
(
1025
-
1068
)
Claude Opus 4.6
1045
(
1029
-
1061
)
Claude Opus 4.8 (Adaptive/Default)
1044
(
1026
-
1061
)
Claude Opus 4.7
1042
(
1026
-
1058
)
Claude Opus 4.8
1038
(
1021
-
1055
)
GLM 5.2
1036
(
1019
-
1053
)
GPT 5.5
1035
(
1019
-
1052
)
GLM 5.3 Flash
1034
(
1010
-
1057
)
Kimi K2.6
1021
(
1004
-
1038
)
Claude Opus 4.5
1016
(
995
-
1037
)
Hy4 Preview
1014
(
991
-
1037
)
Muse Spark 1.2
1013
(
992
-
1034
)
Qwen 3.8 Max
1009
(
988
-
1030
)
Claude Sonnet 4.6
1001
(
985
-
1016
)
DeepSeek V4 Pro (preview)
999
(
982
-
1015
)
GPT 5.2 Chat
994
(
978
-
1010
)
Gemini 3.5 Flash-Lite
990
(
970
-
1010
)
Qwen 3.5 Plus
990
(
974
-
1005
)
Kimi K2.5
988
(
973
-
1003
)
Muse Spark 1.1
985
(
966
-
1005
)
GPT 5.4
985
(
969
-
1001
)
Grok 4.6
983
(
962
-
1004
)
DeepSeek V4 Flash Vision (experimental)
981
(
958
-
1003
)
Grok 4.5
972
(
954
-
991
)
Hy3
971
(
949
-
994
)
DeepSeek V4 Flash (preview)
969
(
953
-
986
)
GPT 5.2
954
(
929
-
980
)
Qwen 3 Max
948
(
920
-
975
)
Qwen 3.8 Flash
929
(
905
-
952
)
Grok 4.1 Fast Reasoning
916
(
901
-
931
)
Muse Glimmer 30B
913
(
892
-
935
)
Nemotron Lightning 3.5 30B A3B
888
(
866
-
911
)
Kimi K2 Instruct
883
(
856
-
910
)
Llama 4 Maverick
803
(
781
-
824
)
Nova 2 Pro
749
(
718
-
779
)
Enterprise Instruction Following

ComplexConstraints

A benchmark for professional instruction following, where constraints depend on each other, fire conditionally, and must be inferred from context.

Rank
Model
Score
Muse Spark 1.3 (xHigh)
51.9
%
Muse Spark 1.3
50.8
%
GPT 5.6 Sol (Max)
50.5
%
GPT 5.5 (xHigh)
49.5
%
GPT 5.5 (High)
48.9
%
Gemini 3.8 Flash (High)
48.4
%
GPT 5.4 (xHigh)
46
%
Hy4 Preview
46
%
Qwen 3.8 Max
45.5
%
Claude Fable 5.1 (Adaptive/Max)
45.1
%
GPT 5.5
44.4
%
Gemini 3.1 Pro
43.7
%
GPT 5.6 Sol
43.7
%
Qwen 3.8 Flash
43.3
%
Claude Fable 5.1
43.2
%
Gemini 3.8 Flash
43.2
%
Gemini 3.7 Flash (High)
42.4
%
DeepSeek V4 Pro
42.1
%
GPT 5.4 (High)
42.1
%
DeepSeek V4 Pro (Max)
42
%
Gemini 3.6 Flash (High)
40
%
Muse Spark 1.2 (xHigh)
39.9
%
DeepSeek V4 Flash Vision (experimental) (Max)
39.9
%
Claude Fable 5 (Max)
38.1
%
Muse Spark 1.1 (xHigh)
38.1
%
Kimi K3
37.9
%
Claude Opus 5 (Adaptive/Max)
37.3
%
Gemini 3.5 Flash
37.1
%
Claude Fable 5 (High)
36.9
%
Muse Spark 1.1 (High)
36.9
%
Gemini 3.7 Flash
36.5
%
Grok 4.6 (xHigh)
36.5
%
Claude Opus 4.6 (Adaptive/High)
36.3
%
DeepSeek V4 Flash Vision (experimental)
36.3
%
Grok 4.5
35.9
%
Claude Opus 5 (Adaptive/High)
35.7
%
Claude Opus 4.8 (Max)
35.6
%
Gemini 3.6 Flash
35.5
%
Grok 4.6
35.1
%
Claude Opus 4.8 (Adaptive/High)
34.8
%
Muse Spark 1.2
34.8
%
Hy3 (High)
34.3
%
Claude Opus 4.8
34.2
%
Inkling
34.1
%
Claude Sonnet 4.6 (Adaptive/High)
34
%
Qwen 3.7 Max
33.5
%
Claude Opus 4.7
31.6
%
GLM 5.3
30.5
%
Claude Opus 4.7 (Adaptive/High)
30.3
%
Kimi K2.6
30.1
%
GLM 5.2 (Max)
29.3
%
DeepSeek V4 Pro (preview)
28
%
GLM 5.3 Flash
27.6
%
Muse Glimmer 30B (xHigh)
23.9
%
Muse Glimmer 30B
23.7
%
Gemini 3.5 Flash-Lite (High)
23.2
%
GLM 5.2
22.8
%
DeepSeek V4 Flash (preview)
20.9
%
Kimi K2.5
18.2
%
Qwen 3.5 Plus
17.8
%
Nemotron 3 Ultra
17.8
%
Claude Opus 4.6
16.9
%
Grok 4.20 Beta
16.4
%
Ernie 5.1
14.7
%
Claude Sonnet 4.6
10.7
%
Nemotron Lightning 3.5 30B A3B
7.1
%
DeepSeek V3.2
2.2
%
Gemini 3.5 Flash-Lite
1.3
%
Hy3
1.2
%
Mistral Large 3
0.4
%
Ernie 4.5 300B
0
%
Nova 2 Pro
0
%
Professional Multimodal Reasoning

GDP.pdf

Can frontier models master the documents that run the world? GDP.pdf is a multimodal and reasoning benchmark that takes real-world prompts and PDFs pulled directly from expert professional workflows.

Rank
Model
Score
GPT 5.6 Sol (Max)
30.7
%
GPT 5.6 Sol (Max)
Claude Fable 5 (Adaptive/Max)
29.8
%
Claude Fable 5 (Adaptive/Max)
Claude Fable 5.1
29.6
%
Claude Fable 5.1
Muse Spark 1.3 (xHigh)
27.6
%
Muse Spark 1.3 (xHigh)
Muse Spark 1.3
27.6
%
Muse Spark 1.3
Claude Fable 5.1 (Adaptive/Max)
27.6
%
Claude Fable 5.1 (Adaptive/Max)
GPT 5.5 (xHigh)
26
%
GPT 5.5 (xHigh)
GPT 5.6 Terra
24.7
%
GPT 5.6 Terra
Claude Opus 5 (Adaptive/Max)
24
%
Claude Opus 5 (Adaptive/Max)
Claude Opus 4.8 (Adaptive/Max)
24
%
Claude Opus 4.8 (Adaptive/Max)
Gemini 3.7 Flash (High)
23.8
%
Gemini 3.7 Flash (High)
Gemini 3.8 Flash
23.4
%
Gemini 3.8 Flash
Gemini 3.8 Flash (High)
23.2
%
Gemini 3.8 Flash (High)
Qwen 3.8 Max
23.2
%
Qwen 3.8 Max
GPT 5.6 Luna
22.7
%
GPT 5.6 Luna
Gemini 3.7 Flash
21.8
%
Gemini 3.7 Flash
Claude Opus 4.7 (Adaptive/Max)
21
%
Claude Opus 4.7 (Adaptive/Max)
Kimi K3 (Max)
19
%
Kimi K3 (Max)
Claude Sonnet 4.6 (Adaptive/Max)
18
%
Claude Sonnet 4.6 (Adaptive/Max)
Grok 4.6 (xHigh)
17.2
%
Grok 4.6 (xHigh)
Gemini 3.1 Pro
17
%
Gemini 3.1 Pro
Qwen 3.8 Flash
16.6
%
Qwen 3.8 Flash
DeepSeek V4 Flash Vision (experimental) (Max)
16.2
%
DeepSeek V4 Flash Vision (experimental) (Max)
Grok 4.6
16
%
Grok 4.6
Muse Spark 1.2 (xHigh)
16
%
Muse Spark 1.2 (xHigh)
Muse Spark 1.1
15
%
Muse Spark 1.1
DeepSeek V4 Flash Vision (experimental)
14.8
%
DeepSeek V4 Flash Vision (experimental)
GLM 5.3 Flash
14
%
GLM 5.3 Flash
Grok 4.5 (High)
14
%
Grok 4.5 (High)
Gemini 3.6 Flash (High)
14
%
Gemini 3.6 Flash (High)
Gemini 3.5 Flash
14
%
Gemini 3.5 Flash
Muse Spark 1.2
12
%
Muse Spark 1.2
Kimi K2.6
12
%
Kimi K2.6
Muse Glimmer 30B (xHigh)
11.8
%
Muse Glimmer 30B (xHigh)
Muse Glimmer 30B
10.8
%
Muse Glimmer 30B
Gemini 3.5 Flash-Lite (High)
10
%
Gemini 3.5 Flash-Lite (High)
Gemini 3 Flash
10
%
Gemini 3 Flash
Grok 4.3 (High)
8
%
Grok 4.3 (High)
Nova 2 Pro
2
%
Nova 2 Pro
Nemotron 3 Nano Omni
2
%
Nemotron 3 Nano Omni
Mistral Large 3
2
%
Mistral Large 3
Enterprise Agents in Realistic RL Environments

EnterpriseBench: CoreCraft Agents

Stop testing models in tiny, self-contained environments. We built CoreCraft, a large-scale startup world, and deployed AI agents to solve real tasks. Our goal: to move agents beyond the cleanliness of the lab and into the chaos of enterprise reality.

Rank
Model
Score
Claude Fable 5.1 (Adaptive/Max)
77.4
%
Claude Fable 5.1 (Adaptive/Max)
Claude Fable 5.1
72.3
%
Claude Fable 5.1
Claude Fable 5 (Adaptive/Max)
70.3
%
Claude Fable 5 (Adaptive/Max)
Frontier Research Mathematics

Riemann-bench

We evaluate AI models on advanced mathematical problems requiring deep reasoning and novel synthesis. Our benchmark features cutting-edge problems sourced from leading mathematicians, including Ivy League professors, PhD IMO medalists, and graduate students at the top of their field.

Rank
Model
Score
GPT-5.6 Sol (Max)
74.4
%
GPT-5.6 Sol (Max)
Claude Opus 5 (Adaptive/Max)
68
%
Claude Opus 5 (Adaptive/Max)
Claude Fable 5.1 (Adaptive/Max)
65.6
%
Claude Fable 5.1 (Adaptive/Max)
Claude Fable 5 (Adaptive/Max)
60
%
Claude Fable 5 (Adaptive/Max)
GPT 5.5 (xHigh)
55.2
%
GPT 5.5 (xHigh)
Gemini 3.8 Flash (High)
51.2
%
Gemini 3.8 Flash (High)
Claude Opus 4.8 (Adaptive/Max)
47.2
%
Claude Opus 4.8 (Adaptive/Max)
GPT 5.4 (xHigh)
41.6
%
GPT 5.4 (xHigh)
Gemini 3.7 Flash (High)
39.2
%
Gemini 3.7 Flash (High)
Grok 4.5 (High)
38.4
%
Grok 4.5 (High)
DeepSeek V4 Pro (Max)
38.4
%
DeepSeek V4 Pro (Max)
Kimi K3
37.6
%
Kimi K3
GPT-5.2 (xHigh)
37.6
%
GPT-5.2 (xHigh)
Gemini 3.5 Flash (High)
36.8
%
Gemini 3.5 Flash (High)
Gemini 3.1 Pro
33.6
%
Gemini 3.1 Pro
Claude Opus 4.7 (Adaptive/Max)
32.8
%
Claude Opus 4.7 (Adaptive/Max)
Gemini 3.6 Flash (High)
30.4
%
Gemini 3.6 Flash (High)
Muse Spark 1.3 (xHigh)
28
%
Muse Spark 1.3 (xHigh)
Claude Opus 4.6 (Adaptive/Max)
27.2
%
Claude Opus 4.6 (Adaptive/Max)
Hy4 Preview
24
%
Hy4 Preview
Muse Spark 1.2 (xHigh)
23.2
%
Muse Spark 1.2 (xHigh)
DeepSeek V4 Flash Vision (experimental) (High)
23.2
%
DeepSeek V4 Flash Vision (experimental) (High)
Muse Spark 1.1 (xHigh)
20.8
%
Muse Spark 1.1 (xHigh)
Inkling (xHigh)
15.2
%
Inkling (xHigh)
Qwen 3.8 Max
15.2
%
Qwen 3.8 Max
Qwen 3.7 Max
15.2
%
Qwen 3.7 Max
Muse Glimmer 30B (xHigh)
13.6
%
Muse Glimmer 30B (xHigh)
Kimi K2.5
12
%
Kimi K2.5
GLM 5.3 (Max)
12
%
GLM 5.3 (Max)
Gemini 3.5 Flash-Lite (High)
11.2
%
Gemini 3.5 Flash-Lite (High)
Claude Opus 4.5 (Adaptive/Max)
11.2
%
Claude Opus 4.5 (Adaptive/Max)
GLM 5.2 (Max)
10.4
%
GLM 5.2 (Max)
DeepSeek V4 Flash (Max)
10.4
%
DeepSeek V4 Flash (Max)
GLM 5.3 Flash (Max)
8.8
%
GLM 5.3 Flash (Max)
DeepSeek V3.2 (Thinking)
8
%
DeepSeek V3.2 (Thinking)
Kimi K2.6
7.2
%
Kimi K2.6
DeepSeek V4 Pro (preview) (xHigh)
5.6
%
DeepSeek V4 Pro (preview) (xHigh)
MAI Thinking 1
4.8
%
MAI Thinking 1

Stay Posted on New Benchmarks