Cost calculator
Estimate monthly hosted-inference cost across every model + provider we track. Numbers assume the pricing on our verified pricing table — always confirm on the provider's own page before committing.
| Provider | Input $/M | Output $/M | ||
|---|---|---|---|---|
| Llama 3.2 1B | groq | $0.040 | $0.040 | $30.00 |
| Llama 3.1 8B Instruct | groq | $0.050 | $0.080 | $42.00 |
| Llama 3.2 3B | groq | $0.060 | $0.060 | $45.00 |
| Phi-3 Mini 4K Instruct | deepinfra | $0.080 | $0.080 | $60.00 |
| Qwen2.5 7B Instruct | deepinfra | $0.080 | $0.300 | $93.00 |
| Mistral Nemo 12B | deepinfra | $0.130 | $0.130 | $97.50 |
| Qwen 3 32B | deepinfra | $0.100 | $0.300 | $105.00 |
| Llama 4 Scout 17B (16E) | groq | $0.110 | $0.340 | $117.00 |
| DeepSeek Coder V2 | deepinfra | $0.140 | $0.280 | $126.00 |
| Llama 3.1 8B Instruct | together | $0.180 | $0.180 | $135.00 |
| Llama 3.2 11B Vision | groq | $0.180 | $0.180 | $135.00 |
| Gemma 2 9B | groq | $0.200 | $0.200 | $150.00 |
| Mistral 7B v0.3 | together | $0.200 | $0.200 | $150.00 |
| Mixtral 8×7B Instruct | groq | $0.240 | $0.240 | $180.00 |
| Llama 4 Scout 17B (16E) | together | $0.180 | $0.590 | $196.50 |
| Llama 3.1 70B Instruct | deepinfra | $0.230 | $0.400 | $198.00 |
| Llama 3.3 70B Instruct | deepinfra | $0.230 | $0.400 | $198.00 |
| Llama 4 Maverick 17B (128E) | groq | $0.200 | $0.600 | $210.00 |
| Qwen2.5 32B Instruct | deepinfra | $0.250 | $0.400 | $210.00 |
| Gemma 2 27B | together | $0.300 | $0.300 | $225.00 |
| Phi-4 14B | together | $0.300 | $0.300 | $225.00 |
| Qwen2.5 14B Instruct | together | $0.300 | $0.300 | $225.00 |
| Llama 4 Maverick 17B (128E) | together | $0.270 | $0.850 | $289.50 |
| Phi-3 Medium 14B | together | $0.400 | $0.400 | $300.00 |
| Qwen 3 32B | together | $0.400 | $0.400 | $300.00 |
| DeepSeek V3 | deepinfra | $0.490 | $0.890 | $427.50 |
| Mixtral 8×7B Instruct | together | $0.600 | $0.600 | $450.00 |
| Qwen 3 235B (A22B) | together | $0.600 | $0.600 | $450.00 |
| Qwen 3 235B (A22B) | deepinfra | $0.400 | $1.400 | $450.00 |
| Llama 3.1 70B Instruct | groq | $0.590 | $0.790 | $472.50 |
| Llama 3.3 70B Instruct | groq | $0.590 | $0.790 | $472.50 |
| Command R | together | $0.500 | $1.500 | $525.00 |
| DeepSeek R1 Distill Llama 70B | groq | $0.750 | $0.990 | $598.50 |
| Llama 3.1 405B Instruct | deepinfra | $0.800 | $0.800 | $600.00 |
| Mistral Small 3 | together | $0.800 | $0.800 | $600.00 |
| Qwen2.5 Coder 32B | together | $0.800 | $0.800 | $600.00 |
| Yi 1.5 34B Chat | together | $0.800 | $0.800 | $600.00 |
| DeepSeek R1 | deepinfra | $0.550 | $2.190 | $658.50 |
| Llama 3.1 70B Instruct | together | $0.880 | $0.880 | $660.00 |
| Llama 3.3 70B Instruct | together | $0.880 | $0.880 | $660.00 |
| Llama 3.2 90B Vision | groq | $0.900 | $0.900 | $675.00 |
| Mixtral 8×22B Instruct | together | $1.200 | $1.200 | $900.00 |
| Qwen2.5 72B Instruct | together | $1.200 | $1.200 | $900.00 |
| QwQ 32B Preview | together | $1.200 | $1.200 | $900.00 |
| DeepSeek V3 | together | $1.250 | $1.250 | $937.50 |
| Kimi K2 Instruct | groq | $1.000 | $3.000 | $1050.00 |
| Kimi K2 Instruct | together | $1.000 | $3.000 | $1050.00 |
| Llama 3.1 405B Instruct | groq | $2.150 | $2.150 | $1612.50 |
| Jamba 1.5 Large | together | $2.000 | $8.000 | $2400.00 |
| Llama 3.1 405B Instruct | together | $3.500 | $3.500 | $2625.00 |
| DeepSeek R1 | together | $3.000 | $7.000 | $2850.00 |
| Command R+ | together | $3.000 | $15.000 | $4050.00 |
Self-hosting comparison: at typical utilisation, a single RTX 4090 (~$400/mo amortised) breaks even against hosted 7B pricing around 2B tokens/month, and against hosted 70B pricing around 200M tokens/month. See /hardware.