벤치마크 상세
MMLU-Pro
폭넓은 전문 지식: 넓은 일반 지식, 시험형 문제, 기본 reasoning 비교
공식 리더보드 보기
공식 리더보드 CSV의 Overall 값입니다. TIGER-Lab 측정과 Self-Reported 제출을 표에서 구분합니다. 제출별 모델·프롬프트 조건 차이를 확인해야 합니다.
모든 공개 결과
모델 검색 · 점수순확인 2026. 09. 07.TIGER-Lab public leaderboard · Overall% overall accuracy
261개 결과원문 ↗
| 순서 | 모델 / 실행 구성 | 점수 | 평가 조건·근거 |
|---|---|---|---|
| 1 | Gemini-3.1-Pro | 91.16 | TIGER-Lab |
| 2 | Gemini-3-Pro(11/25) | 90.1 | Self-Reported |
| 3 | GPT-o1 | 89.3 | Self-Reported |
| 4 | Claude-4.6-Opus(Thinking) | 89.1 | Self-Reported |
| 5 | Gemini-3-Flash(12/25) | 88.6 | Self-Reported |
| 6 | MiniMax-M2.1 | 88 | Self-Reported |
| 7 | Qwen3.5-397B-A17B | 87.8 | Self-Reported |
| 8 | Seed2.0-Lite | 87.7 | Self-Reported |
| 9 | GPT-5.4 | 87.5 | Self-Reported |
| 10 | GPT-5.2 | 87.4 | Self-Reported |
| 11 | Claude-4.5-Sonnet(Thinking) | 87.4 | Self-Reported |
| 12 | Claude-4-Opus-Thinking | 87.3 | Self-Reported |
| 13 | Claude-4.5-Opus(Thinking) | 87.3 | Self-Reported |
| 14 | Claude-4.6-Sonnet(Thinking) | 87.3 | Self-Reported |
| 15 | Hunyuan-T1 | 87.2 | Self-Reported |
| 16 | GPT-5(high) | 87.1 | Self-Reported |
| 17 | K2.5-1T-A32B | 87.1 | Self-Reported |
| 18 | Seed-Thinking-v1.5 | 87 | Self-Reported |
| 19 | Grok-4 | 87 | Self-Reported |
| 20 | Seed2.0-Pro | 87 | Self-Reported |
| 21 | Qwen3.5-122B-A10B | 86.7 | Self-Reported |
| 22 | Seed1.6-Thinking | 86.6 | Self-Reported |
| 23 | Seed1.6-Base | 86.6 | Self-Reported |
| 24 | Seed1.6-Ada-Thinking | 86.4 | Self-Reported |
| 25 | GPT-5.1 | 86.4 | Self-Reported |
| 26 | Gemini-3.1-Flash-Lite-Preview | 86.2 | Self-Reported |
| 27 | GPT-4.5 | 86.1 | Self-Reported |
| 28 | Qwen3.5-27B | 86.1 | Self-Reported |
| 29 | Gemini-2.5-Pro | 86 | Self-Reported |
| 30 | GLM-5 | 86 | Self-Reported |
| 31 | Qwen3-Max-Thinking | 85.7 | Self-Reported |
| 32 | Qwen3.5-35B-A3B | 85.3 | Self-Reported |
| 33 | GPT-o3-high | 85 | Self-Reported |
| 34 | DeepSeek-V3.2-Thinking | 85 | Self-Reported |
| 35 | DeepSeek-V3.1-Thinking | 84.8 | Self-Reported |
| 36 | GLM-4.5 | 84.6 | Self-Reported |
| 37 | Gemini-2.5-Pro-Exp-03-25 | 84.52 | TIGER-Lab |
| 38 | Qwen3-235B-A22B-Thinking-2507 | 84.5 | Self-Reported |
| 39 | Grok-4.1-Fast(Reasoning) | 84.2 | Self-Reported |
| 40 | DeepSeek-R1 | 84 | Self-Reported |
| 41 | Claude-3.7-Sonnet-Thinking | 84 | Self-Reported |
| 42 | Claude-4-Sonnet | 83.7 | Self-Reported |
| 43 | DeepSeek-V3.1-NonThinking | 83.7 | Self-Reported |
| 44 | Seed2.0-Mini | 83.6 | Self-Reported |
| 45 | Intern-S1 | 83.5 | Self-Reported |
| 46 | DeepSeek-R1-0528 | 83.4 | Self-Reported |
| 47 | Grok-3-mini | 83 | Self-Reported |
| 48 | GPT-4-mini (high) | 83 | Self-Reported |
| 49 | Qwen3-235B-A22B-Instruct-2507 | 83 | Self-Reported |
| 50 | Llama4-Behemoth | 82.8 | Self-Reported |
| 51 | Seed-OSS-36B-Instruct | 82.7 | Self-Reported |
| 52 | LongCat-Flash-Chat | 82.7 | Self-Reported |
| 53 | Qwen3.5-9B | 82.5 | Self-Reported |
| 54 | MiniMax-M2 | 82 | Self-Reported |
| 55 | GPT-4.1 | 81.8 | Self-Reported |
| 56 | GLM-4.5-Air | 81.4 | Self-Reported |
| 57 | Deepseek-V3-0324 | 81.3 | Self-Reported |
| 58 | MiniMax-M1 | 81.1 | Self-Reported |
| 59 | Kimi-K2-Instruct | 81 | Self-Reported |
| 60 | Qwen3-30B-A3B-Thinking-2507 | 80.9 | Self-Reported |
| 61 | GPT-oss-120B(high) | 80.8 | Self-Reported |
| 62 | Llama4-Maverick | 80.5 | Self-Reported |
| 63 | GPT-o1-mini | 80.3 | Self-Reported |
| 64 | Doubao-1.5-Pro | 80.1 | Self-Reported |
| 65 | MiniMax-M2.5 | 80.1 | Self-Reported |
| 66 | Grok3-Beta | 79.9 | Self-Reported |
| 67 | GPT-o3-mini | 79.4 | Self-Reported |
| 68 | Gemini-2.0-Pro | 79.1 | Self-Reported |
| 69 | Qwen3.5-4B | 79.1 | Self-Reported |
| 70 | HunyuanTurboS | 79 | Self-Reported |
| 71 | Grok3-mini-Beta | 78.9 | Self-Reported |
| 72 | Qwen3-30B-A3B-Thinking | 78.5 | Self-Reported |
| 73 | ERNIE-4.5-300B-A47B | 78.4 | Self-Reported |
| 74 | Nemotron-3-Nano-30B-A3B(BF16) | 78.3 | Self-Reported |
| 75 | Nemotron-3-Nano-30B-A3B(FP8) | 78.1 | Self-Reported |
| 76 | Claude-3.5-Sonnet (2024-10-22) | 78 | Self-Reported |
| 77 | GPT-4o (2024-11-20) | 77.9 | Self-Reported |
| 78 | Claude-3.5-Sonnet (2024-10-22) | 77.64 | TIGER-LAb |
| 79 | Gemini-2.0-Flash | 77.6 | Self-Reported |
| 80 | Gemini-2.0-Flash-exp | 76.24 | TIGER-Lab |
| 81 | Claude-3.5-Sonnet (2024-06-20) | 76.12 | TIGER-Lab |
| 82 | Qwen2.5-Max | 76.1 | Self-Reported |
| 83 | Phi-4-reasoning-plus | 76 | Self-Reported |
| 84 | Deepseek-V3 | 75.87 | Self-Reported |
| 85 | MiniMax-Text-01 | 75.7 | Self-Reported |
| 86 | Grok-2 | 75.46 | Self-Reported |
| 87 | Grok-4.1-Fast(Non-Reasoning) | 75.2 | Self-Reported |
| 88 | GPT-4o (2024-08-06) | 74.68 | TIGER-Lab |
| 89 | Llama4-Scout | 74.3 | Self-Reported |
| 90 | Phi-4-reasoning | 74.3 | Self-Reported |
| 91 | GPT-oss-20B(high) | 73.6 | Self-Reported |
| 92 | Llama-3.1-405B-Instruct | 73.3 | Self-Reported |
| 93 | GPT-oss-20B(medium) | 73.14 | TIGER-Lab |
| 94 | Athene-V2-Chat (0-shot) | 73.11 | TIGER-Lab |
| 95 | GPT-4o (2024-05-13) | 72.55 | TIGER-Lab |
| 96 | Grok-2-mini | 71.85 | Self-Reported |
| 97 | Gemini-2.0-Flash-Lite | 71.6 | Self-Reported |
| 98 | Qwen2.5-72B | 71.59 | Self-Reported |
| 99 | ECHO_Ego_v2_14B | 71.24 | Self-Reported |
| 100 | QwQ-32B-Preview | 70.97 | Self-Reported |
| 101 | Phi-4 | 70.4 | Self-Reported |
| 102 | Gemini-1.5-Pro-002 | 70.25 | TIGER-Lab |
| 103 | Athene-V2-Chat | 70.21 | TIGER-Lab |
| 104 | ERNIE-4.5-300B-A47B-Base | 69.5 | Self-Reported |
| 105 | Qwen2.5-32B | 69.23 | Self-Reported |
| 106 | SkyThought-T1 | 69.2 | Self-Reported |
| 107 | QwQ-32B | 69.07 | TIGER-Lab |
| 108 | Gemini-1.5-Pro | 69.03 | Self-Reported |
| 109 | Claude-3-Opus | 68.45 | TIGER-Lab |
| 110 | Qwen3-235B-A22B | 68.18 | Self-Reported |
| 111 | Mistral-Large-Instruct-2411 | 67.94 | TIGER-Lab |
| 112 | Gemma-3-27B-it | 67.5 | Self-Reported |
| 113 | Hunyuan-A13B | 67.3 | Self-Reported |
| 114 | Mistral-3.1-Small | 66.8 | Sefl-Reported |
| 115 | General-Reasoner-14B | 66.6 | Self-Reported |
| 116 | Mistral-Small-instruct | 66.3 | Self-Reported |
| 117 | Llama-3.3-70B-Instruct | 65.92 | TIGER-Lab |
| 118 | Mistral-Large-Instruct-2407 | 65.91 | TIGER-Lab |
| 119 | DeepSeek-Chat-V2_5 | 65.83 | TIGER-Lab |
| 120 | Seed-OSS-36B-Base(w/ syn.) | 65.1 | Self-Reported |
| 121 | Nemotron-3-Nano-30B-A3B-Base | 65.1 | Self-Reported |
| 122 | Reka 3 | 65 | Self-Reported |
| 123 | Qwen2-72B-Chat | 64.38 | TIGER-Lab |
| 124 | Gemini-1.5-Flash-002 | 64.09 | TIGER-Lab |
| 125 | magnum-72b-v1 | 63.93 | TIGER-Lab |
| 126 | GPT-4-Turbo | 63.71 | TIGER-Lab |
| 127 | Qwen2.5-14B | 63.69 | Self-Reported |
| 128 | DeepSeek-Coder-V2-Instruct | 63.63 | TIGER-Lab |
| 129 | Higgs-Llama-3-70B | 63.16 | Self-Reported |
| 130 | GPT-4o-mini | 63.09 | TIGER-Lab |
| 131 | azerogpt | 63.07 | Self-Reported |
| 132 | Llama-3.1-70B-Instruct | 62.84 | TIGER-Lab |
| 133 | Llama-3.1-Nemotron-70B-Instruct-HF | 62.78 | TIGER-Lab |
| 134 | Yi-Lightning | 62.38 | TIGER-Lab |
| 135 | Claude-3-5-Haiku-20241022 | 62.12 | TIGER-Lab |
| 136 | RRD2.5-9B | 61.84 | Self-Reported |
| 137 | Qwen3-30B-A3B-Base | 61.7 | Self-Reported |
| 138 | Llama-3.1-405B | 61.6 | Self-Reported |
| 139 | Gemma-3-12B-it | 60.6 | Self-Reported |
| 140 | Nemotron-H-56B-Base | 60.5 | Self-Reported |
| 141 | Seed-OSS-36B-Base(w/o syn.) | 60.4 | Self-Reported |
| 142 | Reflection-Llama-3.1-70B | 60.35 | TIGER-Lab |
| 143 | Hunyuan-Large | 60.2 | Self-Reported |
| 144 | Gemini-1.5-Flash | 59.12 | TIGER-Lab |
| 145 | EXAONE-3.5-32B-Instruct | 58.91 | TIGER-Lab |
| 146 | General-Reasoner-7B | 58.9 | Self-Reported |
| 147 | Mimo-7B-RL | 58.6 | Self-Reported |
| 148 | Yi-large | 58.09 | TIGER-Lab |
| 149 | NewenAI/Phi4-sft | 57.7 | Self-Reported |
| 150 | Internlm3-8B-Instruct | 57.6 | Self-Reported |
| 151 | Claude-3-Sonnet | 56.8 | TIGER-Lab |
| 152 | ERNIE-4.5-21B-A3B-Base | 56.7 | Self-Reported |
| 153 | Gemma-2-27B-it | 56.54 | TIGER-Lab |
| 154 | Mixtral-8x22B-Instruct-v0.1 | 56.33 | TIGER-Lab |
| 155 | Llama-3-70B-Instruct | 56.2 | TIGER-Lab |
| 156 | Phi3-medium-4k | 55.7 | TIGER-Lab |
| 157 | Qwen2.5-Turbo | 55.6 | Self-Reported |
| 158 | Qwen2-72B-32k | 55.59 | TIGER-Lab |
| 159 | Qwen3.5-2B | 55.3 | Self-Reported |
| 160 | Deepseek-V2-Chat | 54.81 | TIGER-Lab |
| 161 | Mistral-Small-base | 54.4 | Self-Reported |
| 162 | Phi-4-mini | 52.8 | Self-Reported |
| 163 | Llama-3-70B | 52.78 | TIGER-Lab |
| 164 | Qwen1.5-72B-Chat | 52.64 | TIGER-Lab |
| 165 | Llama-3.1-70B | 52.47 | TIGER-Lab |
| 166 | Yi-1.5-34B-Chat | 52.29 | TIGER-Lab |
| 167 | Gemma-2-9B-it | 52.08 | TIGER-Lab |
| 168 | Phi3-medium-128k | 51.91 | TIGER-Lab |
| 169 | MAmmoTH2-8x7B-Plus | 50.4 | TIGER-Lab |
| 170 | Qwen1.5-110B | 49.93 | TIGER-Lab |
| 171 | Jamba-1.5-Large | 49.46 | TIGER-Lab |
| 172 | Mistral-Small-Instruct-2409 | 48.4 | TIGER-Lab |
| 173 | GLM-4-9B-Chat | 48.01 | TIGER-Lab |
| 174 | GLM-4-9B | 47.92 | TIGER-Lab |
| 175 | Phi-3.5-mini-instruct | 47.87 | TIGER-Lab |
| 176 | Qwen2-7B-Instruct | 47.24 | TIGER-Lab |
| 177 | Cohere-Aya-Vision | 47.2 | Self-Reported |
| 178 | EXAONE-3.5-7.8B-Instruct | 46.24 | TIGER-Lab |
| 179 | Yi-1.5-9B-Chat | 45.95 | TIGER-Lab |
| 180 | Phi3-mini-4k | 45.66 | TIGER-Lab |
| 181 | Aya-Expanse-32B | 45.41 | TIGER-Lab |
| 182 | Gemma-2-9B | 45.1 | TIGER-Lab |
| 183 | Qwen2.5-7B | 45 | Self-Reported |
| 184 | Mistral-Nemo-Instruct-2407 | 44.81 | TIGER-Lab |
| 185 | Llama-3.1-8B-Instruct | 44.25 | TIGER-Lab |
| 186 | Nemotron-H-8B-Base | 44 | Self-Reported |
| 187 | Phi3-mini-128k | 43.86 | TIGER-Lab |
| 188 | Qwen2.5-3B | 43.73 | Self-Reported |
| 189 | Gemma-3-4B-it | 43.6 | Self-Reported |
| 190 | MAmmoTH2-8B-Plus | 43.35 | TIGER-Lab |
| 191 | Mixtral-8x7B-Instruct-v0.1 | 43.27 | TIGER-Lab |
| 192 | Yi-34B | 43.03 | TIGER-Lab |
| 193 | Claude-3-Haiku-20240307 | 42.29 | TIGER-Lab |
| 194 | Mathstral-7B-v0.1 | 42 | TIGER-Lab |
| 195 | Mimo-7B-base | 41.9 | Self-Reported |
| 196 | DeepSeek-Coder-V2-Lite-Instruct | 41.57 | TIGER-Lab |
| 197 | Mixtral-8x7B-v0.1 | 41.03 | TIGER-Lab |
| 198 | Granite-3.1-8B-Instruct | 41.03 | TIGER-Lab |
| 199 | Llama-3-8B-Instruct | 40.98 | TIGER-Lab |
| 200 | MAmmoTH2-7B-Plus | 40.85 | TIGER-Lab |
| 201 | Qwen2-7B | 40.73 | TIGER-Lab |
| 202 | Mistral-Nemo-Base-2407 | 39.77 | TIGER-Lab |
| 203 | WizardLM-2-8x22B | 39.24 | TIGER-Lab |
| 204 | EXAONE-3.5-2.4B-Instruct | 39.1 | TIGER-Lab |
| 205 | Yi-1.5-6B-Chat | 38.23 | TIGER-Lab |
| 206 | Qwen1.5-14B-Chat | 38.02 | TIGER-Lab |
| 207 | Ministral-8B-Instruct-2410 | 37.93 | TIGER-Lab |
| 208 | c4ai-command-r-v01 | 37.9 | TIGER-Lab |
| 209 | Staring-7B | 37.9 | TIGER-Lab |
| 210 | Llama-2-70B | 37.53 | TIGER-Lab |
| 211 | OpenChat-3.5-8B | 37.24 | TIGER-Lab |
| 212 | InternMath-20B-Plus | 37.1 | TIGER-Lab |
| 213 | LLaDA | 37 | Self-Reported |
| 214 | Llama3-Smaug-8B | 36.93 | TIGER-Lab |
| 215 | Llama-3.1-8B | 36.6 | TIGER-Lab |
| 216 | Llama-3-8B | 35.36 | TIGER-Lab |
| 217 | DeepseekMath-7B-Instruct | 35.3 | TIGER-Lab |
| 218 | DeepSeek-Coder-V2-Lite-Base | 34.37 | TIGER-Lab |
| 219 | Aya-Expanse-8B | 33.74 | TIGER-Lab |
| 220 | Gemma-7B | 33.73 | TIGER-Lab |
| 221 | InternMath-7B-Plus | 33.5 | TIGER-Lab |
| 222 | Granite-3.1-8B-Base | 33.08 | TIGER-Lab |
| 223 | Zephyr-7B-Beta | 32.97 | TIGER-Lab |
| 224 | Qwen2.5-1.5B | 32.1 | Self-Reported |
| 225 | Granite-3.1-2B-Instruct | 31.97 | TIGER-Lab |
| 226 | Granite-3.0-8B-Base | 31.03 | TIGER-Lab |
| 227 | Mistral-7B-v0.1 | 30.88 | TIGER-Lab |
| 228 | Mistral-7B-Instruct-v0.2 | 30.84 | TIGER-Lab |
| 229 | Mistral-7B-v0.2 | 30.43 | TIGER-Lab |
| 230 | Qwen3.5-0.8B | 29.7 | Self-Reported |
| 231 | Qwen1.5-7B-Chat | 29.06 | TIGER-Lab |
| 232 | Yi-6B-Chat | 28.84 | TIGER-Lab |
| 233 | Neo-7B-Instruct | 28.74 | TIGER-Lab |
| 234 | Yi-6B | 26.51 | TIGER-Lab |
| 235 | Neo-7B | 25.85 | TIGER-Lab |
| 236 | Mistral-7B-Instruct-v0.1 | 25.75 | TIGER-Lab |
| 237 | Granite-3.1-3B-A800M-Instruct | 25.42 | TIGER-Lab |
| 238 | Llama-2-13B | 25.34 | TIGER-Lab |
| 239 | Granite-3.1-2B-Base | 23.89 | TIGER-Lab |
| 240 | Llemma-7B | 23.45 | TIGER-Lab |
| 241 | Qwen2-1.5B-Instruct | 22.62 | TIGER-Lab |
| 242 | Qwen2-1.5B | 22.56 | TIGER-Lab |
| 243 | Llama-3.2-3B | 22.17 | TIGER-Lab |
| 244 | Granite-3.0-2B-Base | 21.72 | TIGER-Lab |
| 245 | Granite-3.1-3B-A800M-Base | 20.39 | TIGER-Lab |
| 246 | Llama-2-7B | 20.32 | TIGER-Lab |
| 247 | SmolLM2-1.7B | 18.31 | TIGER-Lab |
| 248 | Qwen2-0.5B-Instruct | 15.93 | TIGER-Lab |
| 249 | Gemma-2B | 15.85 | TIGER-Lab |
| 250 | Gemma-2-2B-it | 15.6 | Self-Reported |
| 251 | Qwen2-0.5B | 14.97 | TIGER-Lab |
| 252 | Qwen2.5-0.5B | 14.92 | Self-Reported |
| 253 | Gemma-3-1B-it | 14.7 | Self-Reported |
| 254 | Granite-3.1-1B-A400M-Instruct | 13.27 | TIGER-Lab |
| 255 | Granite-3.1-1B-A400M-Base | 12.34 | TIGER-Lab |
| 256 | Llama-3.2-1B | 11.95 | TIGER-Lab |
| 257 | SmolLM-1.7B | 11.93 | TIGER-Lab |
| 258 | SmolLM2-360M | 11.38 | TIGER-Lab |
| 259 | SmolLM-135M | 11.22 | TIGER-Lab |
| 260 | SmolLM-360M | 10.95 | TIGER-Lab |
| 261 | SmolLM2-135M | 10.85 | TIGER-Lab |
일치하는 모델이 없습니다. 검색어를 줄여 보세요.
공개된 숫자만 높은 순으로 표시합니다. 순서가 신뢰구간을 고려한 확정 순위는 아닙니다. 공식 리더보드 CSV의 Overall 값입니다. TIGER-Lab 측정과 Self-Reported 제출을 표에서 구분합니다. 제출별 모델·프롬프트 조건 차이를 확인해야 합니다.
읽는 법
전체 벤치마크MMLU보다 어려운 다지선다 지식/추론 평가입니다.
범용 지식 비교의 기준점으로는 좋지만, 실제 업무 자동화 성능은 따로 봐야 합니다.