OSAIM
Open Source AI Models

Leaderboards

Self-reported benchmark scores compiled from model cards and papers. Higher is better. Numbers should be treated as guidance, not gospel — labs use slightly different evaluation harnesses. See methodology for sources.

MMLU

50 models

Massive Multitask Language Understanding — 57 academic subjects, 5-shot. Saturating among frontier models; included for legacy comparison.

#ModelParamsMMLUPer B
1DeepSeek R1· deepseek671B90.800.14
2Kimi K2 Instruct· kimi1000B89.500.09
3DeepSeek V3· deepseek671B88.500.13
4Llama 3.1 405B Instruct· llama405B87.300.22
5Qwen 3 235B (A22B)· qwen235B87.100.37
6Qwen2.5 72B Instruct· qwen72B86.101.20
7Llama 3.3 70B Instruct· llama70B86.001.23
8DeepSeek R1 Distill Llama 70B· deepseek70B86.001.23
9Llama 3.2 90B Vision· llama90B86.000.96
10Llama 4 Maverick 17B (128E)· llama17B85.505.03
11Llama 3.1 Nemotron 70B Instruct· nemotron70B85.001.21
12Phi-4 14B· phi14B84.806.06
13Llama 3.1 70B Instruct· llama70B83.601.19
14Qwen 3 32B· qwen32B83.402.61
15Qwen2.5 32B Instruct· qwen32B83.302.60
16Jamba 1.5 Large· jamba398B81.200.20
17Nemotron-4 340B Instruct· nemotron340B81.100.24
18Mistral Small 3· mistral24B81.003.38
19Qwen2.5 14B Instruct· qwen14B79.705.69
20Hermes 3 Llama 3.1 70B· hermes70B79.601.14
21Llama 4 Scout 17B (16E)· llama17B79.604.68
22DeepSeek Coder V2· deepseek236B79.200.34
23Phi-3 Medium 14B· phi14B78.005.57
24Mixtral 8×22B Instruct· mistral141B77.750.55
25Qwen 3 8B· qwen8B76.909.61
26Yi 1.5 34B Chat· yi34B76.802.26
27Command R+· command104B75.700.73
28Gemma 2 27B· gemma27B75.202.79
29Qwen2.5 Coder 32B· qwen32B75.102.35
30QwQ 32B Preview· qwen32B75.002.34
31Qwen2.5 7B Instruct· qwen7B74.2010.60
32DBRX Instruct· dbrx132B73.700.56
33Llama 3.2 11B Vision· llama11B73.006.64
34Grok 1· grok314B73.000.23
35Gemma 2 9B· gemma9B71.307.92
36Mixtral 8×7B Instruct· mistral47B70.601.51
37Llama 3.1 8B Instruct· llama8B69.408.68
38Llama 2 70B Chat· llama70B68.900.98
39Phi-3 Mini 4K Instruct· phi4B68.8018.11
40Falcon 3 7B Instruct· falcon7B68.509.79
41Command R· command35B68.201.95
42Mistral Nemo 12B· mistral12B68.005.67
43OLMo 2 13B· olmo13B67.505.19
44Hermes 3 Llama 3.1 8B· hermes8B65.408.18
45OLMo 2 7B· olmo7B63.709.10
46Llama 3.2 3B· llama3B63.4021.13
47Falcon Mamba 7B· falcon7B62.008.86
48Stable LM 2 12B· stablelm12B61.005.08
49Mistral 7B v0.3· mistral7B60.108.59
50Llama 2 13B Chat· llama13B54.804.22
Click any column header to sort.

HumanEval

49 models

OpenAI's Python coding benchmark — pass@1 on function completion. Saturating around 90+ for frontier models.

#ModelParamsHumanEvalPer B
1Qwen2.5 Coder 32B· qwen32B92.702.90
2Qwen 3 235B (A22B)· qwen235B90.900.39
3DeepSeek Coder V2· deepseek236B90.200.38
4Qwen 3 32B· qwen32B89.602.80
5Llama 3.1 405B Instruct· llama405B89.000.22
6Llama 3.3 70B Instruct· llama70B88.401.26
7Qwen2.5 32B Instruct· qwen32B88.402.76
8Kimi K2 Instruct· kimi1000B88.400.09
9Qwen2.5 72B Instruct· qwen72B86.601.20
10DeepSeek R1 Distill Llama 70B· deepseek70B86.001.23
11Llama 4 Maverick 17B (128E)· llama17B85.505.03
12Mistral Small 3· mistral24B84.803.53
13Qwen 3 8B· qwen8B84.8010.60
14Qwen2.5 7B Instruct· qwen7B84.8012.11
15Llama 3.1 Nemotron 70B Instruct· nemotron70B84.001.20
16Qwen2.5 14B Instruct· qwen14B83.505.96
17DeepSeek V3· deepseek671B82.600.12
18Phi-4 14B· phi14B82.605.90
19Llama 3.1 70B Instruct· llama70B80.501.15
20Llama 4 Scout 17B (16E)· llama17B79.904.70
21Hermes 3 Llama 3.1 70B· hermes70B78.801.13
22Mixtral 8×22B Instruct· mistral141B76.000.54
23Yi 1.5 34B Chat· yi34B75.202.21
24Nemotron-4 340B Instruct· nemotron340B73.200.22
25Llama 3.1 8B Instruct· llama8B72.609.07
26Jamba 1.5 Large· jamba398B71.300.18
27Command R+· command104B70.700.68
28DBRX Instruct· dbrx132B70.100.53
29Mistral Nemo 12B· mistral12B64.405.37
30Grok 1· grok314B63.200.20
31Phi-3 Medium 14B· phi14B62.204.44
32Hermes 3 Llama 3.1 8B· hermes8B60.407.55
33Phi-3 Mini 4K Instruct· phi4B59.1015.55
34Falcon 3 7B Instruct· falcon7B56.708.10
35Command R· command35B53.701.53
36Gemma 2 27B· gemma27B51.801.92
37Llama 3.2 3B· llama3B51.5017.17
38Mixtral 8×7B Instruct· mistral47B40.200.86
39Gemma 2 9B· gemma9B40.204.47
40Llama 3.2 1B· llama1B37.2037.20
41Mistral 7B v0.3· mistral7B30.504.36
42Llama 2 70B Chat· llama70B29.900.43
43Falcon Mamba 7B· falcon7B29.904.27
44OLMo 2 13B· olmo13B28.702.21
45Stable LM 2 12B· stablelm12B27.402.28
46OLMo 2 7B· olmo7B22.603.23
47Llama 2 13B Chat· llama13B18.301.41
48Gemma 2 2B· gemma3B17.706.81
49Llama 2 7B Chat· llama7B12.801.83
Click any column header to sort.

MATH

41 models

Hendrycks competition mathematics. Exact-match grading. Reasoning models like DeepSeek R1 push 95+; non-reasoning frontier sits around 70–85.

#ModelParamsMATHPer B
1DeepSeek R1· deepseek671B97.300.15
2DeepSeek R1 Distill Llama 70B· deepseek70B94.501.35
3Qwen 3 235B (A22B)· qwen235B91.200.39
4QwQ 32B Preview· qwen32B90.602.83
5Kimi K2 Instruct· kimi1000B90.000.09
6Qwen 3 32B· qwen32B87.402.73
7DeepSeek V3· deepseek671B84.000.13
8Qwen2.5 72B Instruct· qwen72B83.101.15
9Qwen2.5 32B Instruct· qwen32B83.102.60
10Phi-4 14B· phi14B80.405.74
11Qwen 3 8B· qwen8B80.2010.03
12Qwen2.5 14B Instruct· qwen14B80.005.71
13Llama 3.3 70B Instruct· llama70B77.001.10
14DeepSeek Coder V2· deepseek236B75.700.32
15Qwen2.5 7B Instruct· qwen7B75.5010.79
16Llama 3.1 405B Instruct· llama405B73.800.18
17Mistral Small 3· mistral24B70.602.94
18Llama 3.1 70B Instruct· llama70B68.000.97
19Llama 3.1 Nemotron 70B Instruct· nemotron70B67.400.96
20Nemotron-4 340B Instruct· nemotron340B65.500.19
21Qwen2.5 Coder 32B· qwen32B65.002.03
22Llama 4 Maverick 17B (128E)· llama17B61.203.60
23Mistral Nemo 12B· mistral12B55.104.59
24Llama 3.1 8B Instruct· llama8B51.906.49
25Llama 4 Scout 17B (16E)· llama17B50.302.96
26Yi 1.5 34B Chat· yi34B50.101.47
27Llama 3.2 3B· llama3B48.0016.00
28Gemma 2 27B· gemma27B42.301.57
29Phi-3 Medium 14B· phi14B41.802.99
30Mixtral 8×22B Instruct· mistral141B41.800.30
31Falcon 3 7B Instruct· falcon7B39.305.61
32Command R+· command104B38.600.37
33Gemma 2 9B· gemma9B36.604.07
34DBRX Instruct· dbrx132B34.600.26
35Llama 3.2 1B· llama1B30.6030.60
36Mixtral 8×7B Instruct· mistral47B28.400.61
37Phi-3 Mini 4K Instruct· phi4B28.007.37
38Command R· command35B26.600.76
39Grok 1· grok314B23.900.08
40Mistral 7B v0.3· mistral7B13.101.87
41Gemma 2 2B· gemma3B11.804.54
Click any column header to sort.

What does "Per B" mean?

Score divided by parameters in billions — a rough efficiency metric. Models that punch above their weight on a benchmark (Phi-4 on reasoning, Qwen2.5 Coder 32B on code) climb this ranking. Not a perfect measure (training data quality matters more than headline parameter count) but useful for spotting capable small models.