OSAIM
Open Source AI Models

How we track benchmarks (and why some scores stay blank)

Every benchmark row on this site has a source and a verification date — here's the discipline behind that.

Every benchmark row on opensourceaimodels.net carries a source_type ('official_card', 'paper', 'provider_blog', 'independent') and a last_verified_at date. Some cells stay deliberately blank. This post explains the discipline behind those decisions.

Rule one: only what the lab publishes

For every benchmark score we display, someone with credibility (the lab, the paper authors, the provider) has published the number. If we can't find an official source, the cell stays NULL — not zero, not a guess.

This means the site under-reports. Llama 3.2 1B's HumanEval score is real; the fact that some benchmarks appear blank for smaller models is because those labs didn't run those evals. We don't fill the gaps.

Rule two: dates every time

Benchmarks age. HumanEval was frontier in 2022 and is now saturating for frontier models. MATH scores that looked impressive in 2024 look unremarkable in 2026. Every row has a last_verified_at so readers can weight recency.

Rule three: no aggregate scores across families

Different labs use different eval harnesses, different prompt formats, different few-shot counts. Comparing MMLU scores across families is at best directional. We show the numbers because they're useful signal; we don't average them or produce a composite 'quality' rank.

Rule four: the leaderboard reflects reality, not our opinions

SWE-bench Verified is more predictive of real coding-agent quality than HumanEval. IFEval matters more for chat workloads than MMLU. ArenaHard correlates with human preference. That's why we surface those alongside the classical academic set. When a model doesn't publish a score, its row is absent — silence, not zero.

For the full methodology see /methodology. For the raw benchmark data see /leaderboards. For corrections, open a PR at github.com/bryanflowers/opensourceaimodels.

methodologymeta