벤치마크 상세

MMLU-Pro

폭넓은 전문 지식: 넓은 일반 지식, 시험형 문제, 기본 reasoning 비교

공식 리더보드 보기 공식 리더보드 CSV의 Overall 값입니다. TIGER-Lab 측정과 Self-Reported 제출을 표에서 구분합니다. 제출별 모델·프롬프트 조건 차이를 확인해야 합니다.

모든 공개 결과

모델 검색 · 점수순
확인 2026. 09. 07.TIGER-Lab public leaderboard · Overall% overall accuracy
261개 결과원문 ↗
순서모델 / 실행 구성점수평가 조건·근거
1Gemini-3.1-Pro91.16TIGER-Lab
2Gemini-3-Pro(11/25)90.1Self-Reported
3GPT-o189.3Self-Reported
4Claude-4.6-Opus(Thinking)89.1Self-Reported
5Gemini-3-Flash(12/25)88.6Self-Reported
6MiniMax-M2.188Self-Reported
7Qwen3.5-397B-A17B87.8Self-Reported
8Seed2.0-Lite87.7Self-Reported
9GPT-5.487.5Self-Reported
10GPT-5.287.4Self-Reported
11Claude-4.5-Sonnet(Thinking)87.4Self-Reported
12Claude-4-Opus-Thinking87.3Self-Reported
13Claude-4.5-Opus(Thinking)87.3Self-Reported
14Claude-4.6-Sonnet(Thinking)87.3Self-Reported
15Hunyuan-T187.2Self-Reported
16GPT-5(high)87.1Self-Reported
17K2.5-1T-A32B87.1Self-Reported
18Seed-Thinking-v1.587Self-Reported
19Grok-487Self-Reported
20Seed2.0-Pro87Self-Reported
21Qwen3.5-122B-A10B86.7Self-Reported
22Seed1.6-Thinking86.6Self-Reported
23Seed1.6-Base86.6Self-Reported
24Seed1.6-Ada-Thinking86.4Self-Reported
25GPT-5.186.4Self-Reported
26Gemini-3.1-Flash-Lite-Preview86.2Self-Reported
27GPT-4.586.1Self-Reported
28Qwen3.5-27B86.1Self-Reported
29Gemini-2.5-Pro86Self-Reported
30GLM-586Self-Reported
31Qwen3-Max-Thinking85.7Self-Reported
32Qwen3.5-35B-A3B85.3Self-Reported
33GPT-o3-high85Self-Reported
34DeepSeek-V3.2-Thinking85Self-Reported
35DeepSeek-V3.1-Thinking84.8Self-Reported
36GLM-4.584.6Self-Reported
37Gemini-2.5-Pro-Exp-03-2584.52TIGER-Lab
38Qwen3-235B-A22B-Thinking-250784.5Self-Reported
39Grok-4.1-Fast(Reasoning)84.2Self-Reported
40DeepSeek-R184Self-Reported
41Claude-3.7-Sonnet-Thinking84Self-Reported
42Claude-4-Sonnet83.7Self-Reported
43DeepSeek-V3.1-NonThinking83.7Self-Reported
44Seed2.0-Mini83.6Self-Reported
45Intern-S183.5Self-Reported
46DeepSeek-R1-052883.4Self-Reported
47Grok-3-mini83Self-Reported
48GPT-4-mini (high)83Self-Reported
49Qwen3-235B-A22B-Instruct-250783Self-Reported
50Llama4-Behemoth82.8Self-Reported
51Seed-OSS-36B-Instruct82.7Self-Reported
52LongCat-Flash-Chat82.7Self-Reported
53Qwen3.5-9B82.5Self-Reported
54MiniMax-M282Self-Reported
55GPT-4.181.8Self-Reported
56GLM-4.5-Air81.4Self-Reported
57Deepseek-V3-032481.3Self-Reported
58MiniMax-M181.1Self-Reported
59Kimi-K2-Instruct81Self-Reported
60Qwen3-30B-A3B-Thinking-250780.9Self-Reported
61GPT-oss-120B(high)80.8Self-Reported
62Llama4-Maverick80.5Self-Reported
63GPT-o1-mini80.3Self-Reported
64Doubao-1.5-Pro80.1Self-Reported
65MiniMax-M2.580.1Self-Reported
66Grok3-Beta79.9Self-Reported
67GPT-o3-mini79.4Self-Reported
68Gemini-2.0-Pro79.1Self-Reported
69Qwen3.5-4B79.1Self-Reported
70HunyuanTurboS79Self-Reported
71Grok3-mini-Beta78.9Self-Reported
72Qwen3-30B-A3B-Thinking78.5Self-Reported
73ERNIE-4.5-300B-A47B78.4Self-Reported
74Nemotron-3-Nano-30B-A3B(BF16)78.3Self-Reported
75Nemotron-3-Nano-30B-A3B(FP8)78.1Self-Reported
76Claude-3.5-Sonnet (2024-10-22)78Self-Reported
77GPT-4o (2024-11-20)77.9Self-Reported
78Claude-3.5-Sonnet (2024-10-22)77.64TIGER-LAb
79Gemini-2.0-Flash77.6Self-Reported
80Gemini-2.0-Flash-exp76.24TIGER-Lab
81Claude-3.5-Sonnet (2024-06-20)76.12TIGER-Lab
82Qwen2.5-Max76.1Self-Reported
83Phi-4-reasoning-plus76Self-Reported
84Deepseek-V375.87Self-Reported
85MiniMax-Text-0175.7Self-Reported
86Grok-275.46Self-Reported
87Grok-4.1-Fast(Non-Reasoning)75.2Self-Reported
88GPT-4o (2024-08-06)74.68TIGER-Lab
89Llama4-Scout74.3Self-Reported
90Phi-4-reasoning74.3Self-Reported
91GPT-oss-20B(high)73.6Self-Reported
92Llama-3.1-405B-Instruct73.3Self-Reported
93GPT-oss-20B(medium)73.14TIGER-Lab
94Athene-V2-Chat (0-shot)73.11TIGER-Lab
95GPT-4o (2024-05-13)72.55TIGER-Lab
96Grok-2-mini71.85Self-Reported
97Gemini-2.0-Flash-Lite71.6Self-Reported
98Qwen2.5-72B71.59Self-Reported
99ECHO_Ego_v2_14B71.24Self-Reported
100QwQ-32B-Preview70.97Self-Reported
101Phi-470.4Self-Reported
102Gemini-1.5-Pro-00270.25TIGER-Lab
103Athene-V2-Chat70.21TIGER-Lab
104ERNIE-4.5-300B-A47B-Base69.5Self-Reported
105Qwen2.5-32B69.23Self-Reported
106SkyThought-T169.2Self-Reported
107QwQ-32B69.07TIGER-Lab
108Gemini-1.5-Pro69.03Self-Reported
109Claude-3-Opus68.45TIGER-Lab
110Qwen3-235B-A22B68.18Self-Reported
111Mistral-Large-Instruct-241167.94TIGER-Lab
112Gemma-3-27B-it67.5Self-Reported
113Hunyuan-A13B67.3Self-Reported
114Mistral-3.1-Small66.8Sefl-Reported
115General-Reasoner-14B66.6Self-Reported
116Mistral-Small-instruct66.3Self-Reported
117Llama-3.3-70B-Instruct65.92TIGER-Lab
118Mistral-Large-Instruct-240765.91TIGER-Lab
119DeepSeek-Chat-V2_565.83TIGER-Lab
120Seed-OSS-36B-Base(w/ syn.)65.1Self-Reported
121Nemotron-3-Nano-30B-A3B-Base65.1Self-Reported
122Reka 365Self-Reported
123Qwen2-72B-Chat64.38TIGER-Lab
124Gemini-1.5-Flash-00264.09TIGER-Lab
125magnum-72b-v163.93TIGER-Lab
126GPT-4-Turbo63.71TIGER-Lab
127Qwen2.5-14B63.69Self-Reported
128DeepSeek-Coder-V2-Instruct63.63TIGER-Lab
129Higgs-Llama-3-70B63.16Self-Reported
130GPT-4o-mini63.09TIGER-Lab
131azerogpt63.07Self-Reported
132Llama-3.1-70B-Instruct62.84TIGER-Lab
133Llama-3.1-Nemotron-70B-Instruct-HF62.78TIGER-Lab
134Yi-Lightning62.38TIGER-Lab
135Claude-3-5-Haiku-2024102262.12TIGER-Lab
136RRD2.5-9B61.84Self-Reported
137Qwen3-30B-A3B-Base61.7Self-Reported
138Llama-3.1-405B61.6Self-Reported
139Gemma-3-12B-it60.6Self-Reported
140Nemotron-H-56B-Base60.5Self-Reported
141Seed-OSS-36B-Base(w/o syn.)60.4Self-Reported
142Reflection-Llama-3.1-70B60.35TIGER-Lab
143Hunyuan-Large60.2Self-Reported
144Gemini-1.5-Flash59.12TIGER-Lab
145EXAONE-3.5-32B-Instruct58.91TIGER-Lab
146General-Reasoner-7B58.9Self-Reported
147Mimo-7B-RL58.6Self-Reported
148Yi-large58.09TIGER-Lab
149NewenAI/Phi4-sft57.7Self-Reported
150Internlm3-8B-Instruct57.6Self-Reported
151Claude-3-Sonnet56.8TIGER-Lab
152ERNIE-4.5-21B-A3B-Base56.7Self-Reported
153Gemma-2-27B-it56.54TIGER-Lab
154Mixtral-8x22B-Instruct-v0.156.33TIGER-Lab
155Llama-3-70B-Instruct56.2TIGER-Lab
156Phi3-medium-4k55.7TIGER-Lab
157Qwen2.5-Turbo55.6Self-Reported
158Qwen2-72B-32k55.59TIGER-Lab
159Qwen3.5-2B55.3Self-Reported
160Deepseek-V2-Chat54.81TIGER-Lab
161Mistral-Small-base54.4Self-Reported
162Phi-4-mini52.8Self-Reported
163Llama-3-70B52.78TIGER-Lab
164Qwen1.5-72B-Chat52.64TIGER-Lab
165Llama-3.1-70B52.47TIGER-Lab
166Yi-1.5-34B-Chat52.29TIGER-Lab
167Gemma-2-9B-it52.08TIGER-Lab
168Phi3-medium-128k51.91TIGER-Lab
169MAmmoTH2-8x7B-Plus50.4TIGER-Lab
170Qwen1.5-110B49.93TIGER-Lab
171Jamba-1.5-Large49.46TIGER-Lab
172Mistral-Small-Instruct-240948.4TIGER-Lab
173GLM-4-9B-Chat48.01TIGER-Lab
174GLM-4-9B47.92TIGER-Lab
175Phi-3.5-mini-instruct47.87TIGER-Lab
176Qwen2-7B-Instruct47.24TIGER-Lab
177Cohere-Aya-Vision47.2Self-Reported
178EXAONE-3.5-7.8B-Instruct46.24TIGER-Lab
179Yi-1.5-9B-Chat45.95TIGER-Lab
180Phi3-mini-4k45.66TIGER-Lab
181Aya-Expanse-32B45.41TIGER-Lab
182Gemma-2-9B45.1TIGER-Lab
183Qwen2.5-7B45Self-Reported
184Mistral-Nemo-Instruct-240744.81TIGER-Lab
185Llama-3.1-8B-Instruct44.25TIGER-Lab
186Nemotron-H-8B-Base44Self-Reported
187Phi3-mini-128k43.86TIGER-Lab
188Qwen2.5-3B43.73Self-Reported
189Gemma-3-4B-it43.6Self-Reported
190MAmmoTH2-8B-Plus43.35TIGER-Lab
191Mixtral-8x7B-Instruct-v0.143.27TIGER-Lab
192Yi-34B43.03TIGER-Lab
193Claude-3-Haiku-2024030742.29TIGER-Lab
194Mathstral-7B-v0.142TIGER-Lab
195Mimo-7B-base41.9Self-Reported
196DeepSeek-Coder-V2-Lite-Instruct41.57TIGER-Lab
197Mixtral-8x7B-v0.141.03TIGER-Lab
198Granite-3.1-8B-Instruct41.03TIGER-Lab
199Llama-3-8B-Instruct40.98TIGER-Lab
200MAmmoTH2-7B-Plus40.85TIGER-Lab
201Qwen2-7B40.73TIGER-Lab
202Mistral-Nemo-Base-240739.77TIGER-Lab
203WizardLM-2-8x22B39.24TIGER-Lab
204EXAONE-3.5-2.4B-Instruct39.1TIGER-Lab
205Yi-1.5-6B-Chat38.23TIGER-Lab
206Qwen1.5-14B-Chat38.02TIGER-Lab
207Ministral-8B-Instruct-241037.93TIGER-Lab
208c4ai-command-r-v0137.9TIGER-Lab
209Staring-7B37.9TIGER-Lab
210Llama-2-70B37.53TIGER-Lab
211OpenChat-3.5-8B37.24TIGER-Lab
212InternMath-20B-Plus37.1TIGER-Lab
213LLaDA37Self-Reported
214Llama3-Smaug-8B36.93TIGER-Lab
215Llama-3.1-8B36.6TIGER-Lab
216Llama-3-8B35.36TIGER-Lab
217DeepseekMath-7B-Instruct35.3TIGER-Lab
218DeepSeek-Coder-V2-Lite-Base34.37TIGER-Lab
219Aya-Expanse-8B33.74TIGER-Lab
220Gemma-7B33.73TIGER-Lab
221InternMath-7B-Plus33.5TIGER-Lab
222Granite-3.1-8B-Base33.08TIGER-Lab
223Zephyr-7B-Beta32.97TIGER-Lab
224Qwen2.5-1.5B32.1Self-Reported
225Granite-3.1-2B-Instruct31.97TIGER-Lab
226Granite-3.0-8B-Base31.03TIGER-Lab
227Mistral-7B-v0.130.88TIGER-Lab
228Mistral-7B-Instruct-v0.230.84TIGER-Lab
229Mistral-7B-v0.230.43TIGER-Lab
230Qwen3.5-0.8B29.7Self-Reported
231Qwen1.5-7B-Chat29.06TIGER-Lab
232Yi-6B-Chat28.84TIGER-Lab
233Neo-7B-Instruct28.74TIGER-Lab
234Yi-6B26.51TIGER-Lab
235Neo-7B25.85TIGER-Lab
236Mistral-7B-Instruct-v0.125.75TIGER-Lab
237Granite-3.1-3B-A800M-Instruct25.42TIGER-Lab
238Llama-2-13B25.34TIGER-Lab
239Granite-3.1-2B-Base23.89TIGER-Lab
240Llemma-7B23.45TIGER-Lab
241Qwen2-1.5B-Instruct22.62TIGER-Lab
242Qwen2-1.5B22.56TIGER-Lab
243Llama-3.2-3B22.17TIGER-Lab
244Granite-3.0-2B-Base21.72TIGER-Lab
245Granite-3.1-3B-A800M-Base20.39TIGER-Lab
246Llama-2-7B20.32TIGER-Lab
247SmolLM2-1.7B18.31TIGER-Lab
248Qwen2-0.5B-Instruct15.93TIGER-Lab
249Gemma-2B15.85TIGER-Lab
250Gemma-2-2B-it15.6Self-Reported
251Qwen2-0.5B14.97TIGER-Lab
252Qwen2.5-0.5B14.92Self-Reported
253Gemma-3-1B-it14.7Self-Reported
254Granite-3.1-1B-A400M-Instruct13.27TIGER-Lab
255Granite-3.1-1B-A400M-Base12.34TIGER-Lab
256Llama-3.2-1B11.95TIGER-Lab
257SmolLM-1.7B11.93TIGER-Lab
258SmolLM2-360M11.38TIGER-Lab
259SmolLM-135M11.22TIGER-Lab
260SmolLM-360M10.95TIGER-Lab
261SmolLM2-135M10.85TIGER-Lab

공개된 숫자만 높은 순으로 표시합니다. 순서가 신뢰구간을 고려한 확정 순위는 아닙니다. 공식 리더보드 CSV의 Overall 값입니다. TIGER-Lab 측정과 Self-Reported 제출을 표에서 구분합니다. 제출별 모델·프롬프트 조건 차이를 확인해야 합니다.

MMLU보다 어려운 다지선다 지식/추론 평가입니다.

범용 지식 비교의 기준점으로는 좋지만, 실제 업무 자동화 성능은 따로 봐야 합니다.