Size

70B+

21 open-source models in this size bucket.

Moonshot AI's 1-trillion-parameter mixture-of-experts (32B active per token). Trained on 15.5T tokens with a heavy emphasis on tool-use and agentic behaviour. Modified-MIT licence with an attribution clause for very-large deployments. Exceptional at long-horizon agent tasks; benchmarked well against Claude Sonnet on SWE-bench Verified.

Context: 128K
License: kimi
VRAM Q4: 600 GB

DeepSeek V3

671B

671B-parameter MoE model with 37B active per token. Trained for roughly $5.6M of compute — a landmark in cost-efficient frontier training. Frontier-class quality at a fraction of the cost of the closed proprietary frontier. The DeepSeek licence permits commercial use with limited restrictions on military and unlawful applications. Running V3 yourself requires serious hardware (8× H100 at fp8); most teams will use it via the DeepSeek API or providers like Together.

Context: 128K
License: deepseek
VRAM Q4: 402.6 GB

DeepSeek R1

671B

Reasoning model trained with reinforcement learning on top of DeepSeek V3-Base. MIT licence — even the weights are unrestricted, making R1 the most permissively-licensed frontier reasoning model. Generates long internal chains-of-thought before answering, trading latency for accuracy on math, code, and reasoning benchmarks. Distilled variants (e.g. R1 Distill Llama 70B) recover most of the quality at much smaller scales.

Context: 128K
License: mit
VRAM Q4: 402.6 GB

Llama 3.1 405B Instruct

405B

Meta's July 2024 flagship — the first open-weights model at 405B parameters. Trained on 15T tokens with 128K context. Rivals GPT-4o on many academic benchmarks and set the ceiling for open-weights quality for most of 2024. Running it self-hosted requires serious hardware (8× H100 at fp8 or multi-node at fp16); most users will run it via a hosted provider (Together, Groq, Fireworks). Llama 3.3 70B closed most of the practical gap at a fraction of the cost, so 405B is now most useful when 70B specifically hits its ceiling.

Context: 128K
License: llama-3
VRAM Q4: 243 GB

Jamba 1.5 Large

398B

Hybrid Mamba-Transformer-MoE model with native 256K context (effective beyond 140K). 94B active parameters out of 398B total. The state-space-model layers give it linear-time scaling with sequence length, making it interesting for very long contexts. Licensed under AI21's open model licence, which permits most commercial use.

Context: 256K
License: jamba-open
VRAM Q4: 238.8 GB

Nemotron-4 340B Instruct

340B

NVIDIA's reward-modelling research vehicle. Trained primarily to be a synthetic-data-generation specialist rather than a chat-first model. Useful for teams building instruction-tuning datasets at scale.

Context: 4K
License: llama-3
VRAM Q4: 204 GB

Grok 1

314B

xAI's first open-weights release: a 314B-parameter mixture-of-experts model. Apache 2.0 licensed. Largely a research artefact at this size — most users will run smaller models for production — but useful as a permissively-licensed reference for MoE research.

Context: 8K
License: apache-2-0
VRAM Q4: 188.4 GB

Grok 2

300B

xAI's second open-weights release, Apache 2.0. ~300B mixture-of-experts. xAI's pattern of open-sourcing the previous frontier when a new one ships continues from Grok 1. Competitive with GPT-4-class chat quality at release; today useful mainly as a research artefact given the compute needed to run it.

Context: 131K
License: apache-2-0
VRAM Q4: 180 GB

DeepSeek Coder V2

236B

Coding-focused MoE model with 21B active parameters out of 236B total. Supports 338 programming languages with strong performance across mainstream stacks (Python, TypeScript, Go, Rust, Java, C++) and competent results on niche languages where most open models falter. The DeepSeek licence applies — commercial use permitted with some application restrictions.

Context: 128K
License: deepseek
VRAM Q4: 141.6 GB

Qwen 3 235B (A22B)

235B

The flagship Qwen 3 release: a 235B-total MoE with 22B active parameters per token. Competitive with DeepSeek V3 and Llama 4 Maverick on reasoning benchmarks while being smaller total. Apache 2.0 — one of the most permissively licenced frontier-class models.

Context: 128K
License: apache-2-0
VRAM Q4: 141 GB

Mixtral 8×22B Instruct

141B

Scaled-up Mixtral with 22B-parameter experts. ~39B active parameters out of 141B total. Strong long-context performance and competitive coding scores. Apache 2.0 makes it attractive for self-hosting where the licence terms of Llama 3 are a non-starter.

Context: 66K
License: apache-2-0
VRAM Q4: 84.6 GB

DBRX Instruct

132B

Databricks' 132B mixture-of-experts — 16 experts, 4 active per token (36B active params). Trained on 12T tokens on Mosaic infrastructure and released under the Databricks Open Model Licence. DBRX was best-in-class on release; now beaten by Llama 3.3 70B and Qwen 2.5 72B on most benchmarks, but retains value as a well-documented MoE reference.

Context: 33K
License: dbrx-open
VRAM Q4: 79.2 GB

Command R+

104B

Cohere's flagship 104B model. RAG-focused with native multilingual support across ~10 high-resource languages. CC-BY-NC weights; commercial use via Cohere's hosted API.

Context: 128K
License: mrl
VRAM Q4: 62.4 GB

Llama 3.2 90B Vision

90B

Larger vision-language Llama variant, competitive with the proprietary multimodal frontier on standard image-understanding benchmarks. Drops in as a vision upgrade where 11B isn't sharp enough. Requires substantial GPU memory in fp16; most teams will run it quantized or on multi-GPU. A natural pairing with retrieval pipelines that fetch image-rich chunks alongside text.

Context: 128K
License: llama-3
VRAM Q4: 54 GB

Qwen2.5 72B Instruct

72B

The flagship Qwen 2.5 release. Competes with Llama 3.1 405B on many benchmarks at one-fifth the parameter count. Note the 72B specifically uses the Qwen License (commercial use up to 100M MAU) — the smaller Qwen2.5 sizes are Apache 2.0.

Context: 128K
License: qwen
VRAM Q4: 43.2 GB

Llama 2 70B Chat

70B

Flagship Llama 2 release. Fundamentally superseded by Llama 3 70B on every benchmark, but relevant historically: the model that made 'open-weights chat model at frontier scale' credible for enterprise workloads.

Context: 4K
License: llama-2
VRAM Q4: 42 GB

Llama 3.3 70B Instruct

70B

Meta's December 2024 refresh of Llama 3 70B that closes most of the gap with Llama 3.1 405B for chat workloads while remaining tractable on a single H100. Strong instruction following, robust tool-use behaviour, and a 128K context window make it the default choice for production chat at 70B scale. The 3.3 release was trained on a refreshed instruction-tuning data mix and benefits from Meta's most recent alignment work. It outperforms the much larger 3.1 405B on several reasoning benchmarks at a fraction of inference cost. The licence is the Llama 3 Community License, which permits commercial use unless your service exceeds 700M monthly active users. Good pick for: production chat at scale, RAG over long documents, agentic workflows where tool use matters, and any 70B-tier replacement for closed proprietary models.

Context: 128K
License: llama-3
VRAM Q4: 42 GB

Llama 3.1 70B Instruct

70B

The pre-3.3 70B workhorse. Same base architecture as Llama 3.3 70B but the earlier instruction-tuning recipe. Still widely referenced as a baseline in papers and provider docs, and still the default 70B on some hosted providers.

Context: 128K
License: llama-3
VRAM Q4: 42 GB

Llama 3.1 Nemotron 70B Instruct

70B

NVIDIA's RLHF-tuned Llama 3.1 70B. Tops several Arena-style human-preference leaderboards and shipped with NVIDIA's reward-model research. Inherits the Llama 3 community licence.

Context: 128K
License: llama-3
VRAM Q4: 42 GB

DeepSeek R1 Distill Llama 70B

70B

R1 reasoning capabilities distilled into a Llama 3.3 70B base. The most accessible way to run R1-class reasoning locally — fits on a single H100 in fp16 or on a 4090 at Q4. Inherits Llama 3's community licence (commercial use under 700M MAU). Great pick for production reasoning workloads where the full R1 is too expensive to host but o1/R1-style quality is required.

Context: 128K
License: llama-3
VRAM Q4: 42 GB

Hermes 3 Llama 3.1 70B

70B

Larger Hermes 3 variant on top of Llama 3.1 70B. Widely used in agent-heavy workloads that need strong tool use combined with reliable function-calling schemas.

Context: 128K
License: llama-3
VRAM Q4: 42 GB