1602 model families, 155 providers, 3301 price changes in 30 days
Prices are re-read from every provider daily and appended to a ledger, never overwritten. 1863 of the recent changes were drops. Watch a family to get an email the morning after it moves.
Price changes, last 30 days
| When | Model | Provider | Side | $/M before → after | Change | |
|---|---|---|---|---|---|---|
| Sep 29 | claude-fable-5-1 | anthropicvia anthropic | output | 10.00 → 50.00 | +400.0% | Watch |
| Sep 29 | claude-opus-5-5 | anthropicvia anthropic | output | 5.00 → 20.00 | +300.0% | Watch |
| Sep 29 | glm-5v-turbo | zaivia models-dev | input | 1.20 → 5.00 | +316.7% | Watch |
| Sep 29 | glm-5v-turbo | zaivia models-dev | output | 4.00 → 22.00 | +450.0% | Watch |
| Sep 29 | qwen3.7-plus | dashscopevia models-dev | input | 0.500 → 0.400 | -20.0% | Watch |
| Sep 29 | qwen3.7-plus | dashscopevia models-dev | output | 3.00 → 1.60 | -46.7% | Watch |
| Sep 29 | vercel_ai_gateway/spacexai/grok-4.7 | vercel_ai_gatewayvia vercel-gateway | input | 1.20 → 2.00 | +66.7% | Watch |
| Sep 29 | vercel_ai_gateway/spacexai/grok-4.7 | vercel_ai_gatewayvia vercel-gateway | output | 3.60 → 6.00 | +66.7% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v3.0-t2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.224 → 0.336 | +50.0% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v3.0-t2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.308 → 0.224 | -27.3% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v3.0-t2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.252 → 0.308 | +22.2% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v3.0-t2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.168 → 0.252 | +50.0% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v3.0-t2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.392 → 0.168 | -57.1% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v3.0-motion-control@default | vercel_ai_gatewayvia vercel-gateway | output | 0.168 → 0.126 | -25.0% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v3.0-i2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.224 → 0.336 | +50.0% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v3.0-i2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.308 → 0.224 | -27.3% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v3.0-i2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.252 → 0.308 | +22.2% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v3.0-i2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.168 → 0.252 | +50.0% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v3.0-i2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.392 → 0.168 | -57.1% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v2.6-t2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.042 → 0.070 | +66.7% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v2.6-t2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.140 → 0.042 | -70.0% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v2.6-motion-control@default | vercel_ai_gatewayvia vercel-gateway | output | 0.112 → 0.070 | -37.5% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v2.6-i2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.042 → 0.070 | +66.7% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v2.6-i2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.140 → 0.042 | -70.0% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v2.5-turbo-t2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.070 → 0.042 | -40.0% | Watch |
| Sep 29 | vercel_ai_gateway/klingai/kling-v2.5-turbo-i2v@default | vercel_ai_gatewayvia vercel-gateway | output | 0.070 → 0.042 | -40.0% | Watch |
| Sep 29 | vercel_ai_gateway/google/veo-3.1-lite-generate-001@720p | vercel_ai_gatewayvia vercel-gateway | output | 0.050 → 0.030 | -40.0% | Watch |
| Sep 29 | vercel_ai_gateway/google/veo-3.1-lite-generate-001@1080p | vercel_ai_gatewayvia vercel-gateway | output | 0.080 → 0.050 | -37.5% | Watch |
| Sep 29 | vercel_ai_gateway/google/veo-3.1-generate-001@720p | vercel_ai_gatewayvia vercel-gateway | output | 0.400 → 0.200 | -50.0% | Watch |
| Sep 29 | vercel_ai_gateway/google/veo-3.1-generate-001@4k | vercel_ai_gatewayvia vercel-gateway | output | 0.600 → 0.400 | -33.3% | Watch |
| Sep 29 | vercel_ai_gateway/google/veo-3.1-generate-001@1080p | vercel_ai_gatewayvia vercel-gateway | output | 0.400 → 0.200 | -50.0% | Watch |
| Sep 29 | vercel_ai_gateway/google/veo-3.1-fast-generate-001@720p | vercel_ai_gatewayvia vercel-gateway | output | 0.150 → 0.100 | -33.3% | Watch |
| Sep 29 | vercel_ai_gateway/google/veo-3.1-fast-generate-001@4k | vercel_ai_gatewayvia vercel-gateway | output | 0.350 → 0.300 | -14.3% | Watch |
| Sep 29 | vercel_ai_gateway/google/veo-3.1-fast-generate-001@1080p | vercel_ai_gatewayvia vercel-gateway | output | 0.150 → 0.100 | -33.3% | Watch |
| Sep 29 | vercel_ai_gateway/bytedance/seedance-v1.5-pro@1080p | vercel_ai_gatewayvia vercel-gateway | output | 0.117 → 0.058 | -50.0% | Watch |
| Sep 29 | vercel_ai_gateway/bytedance/seedance-v1.5-pro@480p | vercel_ai_gatewayvia vercel-gateway | output | 0.024 → 0.012 | -49.8% | Watch |
| Sep 29 | vercel_ai_gateway/bytedance/seedance-v1.5-pro@720p | vercel_ai_gatewayvia vercel-gateway | output | 0.052 → 0.026 | -50.0% | Watch |
| Sep 29 | vercel_ai_gateway/google/veo-3.0-fast-generate-001@1080p | vercel_ai_gatewayvia vercel-gateway | output | 0.150 → 0.100 | -33.3% | Watch |
| Sep 29 | vercel_ai_gateway/google/veo-3.0-fast-generate-001@720p | vercel_ai_gatewayvia vercel-gateway | output | 0.150 → 0.100 | -33.3% | Watch |
| Sep 29 | vercel_ai_gateway/google/veo-3.0-generate-001@1080p | vercel_ai_gatewayvia vercel-gateway | output | 0.400 → 0.200 | -50.0% | Watch |
Showing the 40 most recent of 3301.
Price index, last 90 days
Median list price of every paid chat listing the catalog saw that day, one price per seller, in USD per million tokens. Only sources that polled that day count, so a missing day is a gap rather than a guess. The index is not held to a fixed basket: when a new source adds listings the median moves with them, and the listing count below says when that happened.
2026-09-28: 4,490 listings from 6 of 6 sources; cheapest paid input $0.003, output $0.002. Over 89 indexed days, 343 listing-cuts and 293 listing-raises; 41 days had a source missing. Series by provider, family and /decide task via /api/v1/trends.
Price catalog
One row per model family, ranked by how many providers compete on it. Cheapest is by input + output list price; free and preview tiers are excluded. Full list at /api/v1/families.
| Family | Providers | Cheapest input $/M | Cheapest output $/M | Via | |
|---|---|---|---|---|---|
| deepseek-v4 | 51 | 0.044 | 0.132 | streamlake | Watch |
| glm-5-3-flash | 42 | 0.045 | 0.140 | inferencenet | Watch |
| deepseek-v4-1 | 40 | 0.140 | 0.280 | amd | Watch |
| glm-5-2 | 40 | 0.490 | 1.54 | decart | Watch |
| glm-5-3 | 40 | 0.370 | 1.14 | reka | Watch |
| gpt-oss | 36 | 0.018 | 0.090 | darkbloom | Watch |
| qwen-3 | 31 | 0.010 | 0.000 | nebius | Watch |
| llama-3-3-70b-instruct | 30 | 0.200 | 0.200 | nscale | Watch |
| kimi-k2-6 | 29 | 0.475 | 2.45 | inceptron | Watch |
| deepseek-v3-2 | 28 | 0.134 | 0.202 | atlascloud | Watch |
| kimi-k3 | 28 | 0.679 | 8.79 | sail_research | Watch |
| kimi-k2-7-code | 27 | 0.713 | 3.00 | streamlake | Watch |
| qwen-3-235b | 27 | 0.087 | 0.350 | openrouter | Watch |
| qwen-3-8-27b | 26 | 0.150 | 1.88 | deepinfra | Watch |
| glm-5-1 | 25 | 0.966 | 3.04 | streamlake | Watch |
| deepseek-r1 | 24 | 0.025 | 0.025 | nscale | Watch |
| llama-3-1-8b-instruct | 23 | 0.025 | 0.025 | inference_net | Watch |
| gemma-4-26b | 22 | 0.060 | 0.200 | reka | Watch |
| gemma-4-31b | 21 | 0.080 | 0.300 | reka | Watch |
| minimax-m3 | 21 | 0.230 | 0.960 | coreweave | Watch |
| qwen-3-5-397b | 20 | 0.168 | 1.01 | baidu | Watch |
| qwen-3-6-35b | 19 | 0.050 | 0.700 | darkbloom | Watch |
| deepseek-v3 | 18 | 0.200 | 0.200 | hyperbolic | Watch |
| glm-4-7 | 17 | 0.400 | 1.75 | deepinfra | Watch |
| kimi-k2-5 | 17 | 0.450 | 2.25 | deepinfra | Watch |
| qwen-3-32b | 17 | 0.050 | 0.100 | lambda_ai | Watch |
| mistral-small | 16 | 0.050 | 0.080 | deepinfra | Watch |
| claude-sonnet-4-6 | 15 | 3.00 | 15.00 | bedrock_converse | Watch |
| deepseek-v3-1 | 15 | 0.250 | 0.950 | deepinfra | Watch |
| glm-5 | 15 | 0.600 | 1.92 | openrouter | Watch |
| qwen-3-coder | 15 | 0.120 | 0.800 | openrouter | Watch |
| claude-opus-4-7 | 14 | 4.50 | 22.50 | gmi | Watch |
| claude-opus-4-8 | 14 | 5.00 | 25.00 | bedrock | Watch |
| claude-sonnet-4-5 | 14 | 3.00 | 15.00 | bedrock | Watch |
| mistral-nemo | 14 | 0.018 | 0.030 | dekallm | Watch |
| qwen-3-6-27b | 14 | 0.300 | 2.00 | chutes | Watch |
| claude-fable-5 | 13 | 10.00 | 50.00 | bedrock | Watch |
| claude-haiku-4-5 | 13 | 1.00 | 5.00 | bedrock | Watch |
| claude-opus-4-5 | 13 | 5.00 | 25.00 | bedrock | Watch |
| claude-opus-4-6 | 13 | 5.00 | 25.00 | bedrock_converse | Watch |
| claude-opus-5 | 13 | 5.00 | 25.00 | bedrock | Watch |
| claude-sonnet-5 | 13 | 2.00 | 10.00 | bedrock | Watch |
| deepseek-chat | 13 | 0.280 | 0.420 | deepseek | Watch |
| llama-3-1-70b-instruct | 13 | 0.100 | 0.100 | fireworks_ai | Watch |
| llama-3-2-3b-instruct | 13 | 0.015 | 0.025 | lambda_ai | Watch |
| minimax-m2-7 | 13 | 0.210 | 0.840 | gmi | Watch |
| mistral-large | 13 | 0.500 | 1.50 | azure_ai | Watch |
| qwen-2 | 13 | 0.020 | 0.060 | nebius | Watch |
| qwen-3-30b | 13 | 0.048 | 0.193 | openrouter | Watch |
| qwen-3-5-35b | 13 | 0.056 | 0.448 | baidu | Watch |
| minimax-m2-5 | 12 | 0.270 | 0.950 | venice | Watch |
| qwen-3-coder-30b | 12 | 0.070 | 0.260 | huggingface | Watch |
| qwen-3-coder-480b | 12 | 0.250 | 1.00 | siliconflow | Watch |
| claude-fable-5-1 | 11 | 10.00 | 50.00 | bedrock | Watch |
| claude-opus-5-5 | 11 | 4.00 | 20.00 | databricks | Watch |
| claude-sonnet-5-5 | 11 | 2.00 | 10.00 | bedrock_converse | Watch |
| glm-4-6 | 11 | 0.400 | 1.75 | io_net | Watch |
| grok-4 | 11 | 0.200 | 0.500 | azure_ai | Watch |
| llama-3-1-405b-instruct | 11 | 0.100 | 0.100 | fireworks_ai | Watch |
| llama-4-maverick-128e | 11 | 0.050 | 0.100 | lambda_ai | Watch |
Showing the 60 most-contested of 1602 families.
2000 throughput numbers verified by ≥2 independent measurements
Every cost quote on this site sits on numbers that came from somewhere real. We don't average estimates and call it a benchmark. When two independent labs (NVIDIA, Dell, Cisco, Lenovo, AMD, ASUSTeK, Google, Oracle, Supermicro, GigaComputing) publish the same MLPerf result within 10%, we mark itverified-cross-checkedand surface it here.
Best deploy: Llama 3.1 70B, 1.2M tokens
The end-to-end optimizer pulls hosted-API pricing, GPU rental hourlies, and self-deploy electricity (with PUE and grid carbon) into one ranked list. Same workload, three honest answers: cheapest dollars, greenest emissions, and the closed-source alternative.
Full ranked list: POST /api/v1/cost/best-deploy with {family, inputTokens, outputTokens}. Returns hosted-API + hosted-GPU + self-deploy options ranked by cost; closed-source families return hosted-API only.
Headline agreements
Sorted by number of independent labs that published the same benchmark. Delta is the tightest pairwise agreement. Click through to MLPerf v5.1 for the raw submissions.
| Model | Hardware | Labs | Tightest delta | Throughput |
|---|---|---|---|---|
| Llama 3 1 8B | 8× B200 180GB SXM | 6 gigacomputing, dell, supermicro, nvidia +2 | 0.06% | 142,136tok/s |
| Llama 3 1 405B | 8× B200 180GB SXM | 6 nvidia, gigacomputing, dell, google +2 | 0.10% | 1,624tok/s |
| Llama 2 70B | 8× MI325X 256GB | 5 gigacomputing, amd, asustek, supermicro +1 | 0.00% | 34,483tok/s |
| Mixtral 8x7b | 8× MI325X 256GB | 4 amd, gigacomputing, asustek, nvidia | 0.00% | 68,781tok/s |
| Llama 2 70B | 8× H200 141GB SXM | 4 dell, asustek, google, nvidia | 0.00% | 35,317tok/s |
| Llama 2 70B | 8× B200 180GB SXM | 4 nvidia, dell, supermicro, gigacomputing | 0.17% | 101,527tok/s |
Cheapest self-host options
Cost per million tokens for cross-checked throughput rows joined against verified public hourly GPU pricing. Single-node configurations only (multi-node MLPerf rows excluded — not rentable at public per-hour rates). Math:($/hr × gpu_count × 1M) / (tok/s × 3600)
| Model | Hardware | Provider | Throughput | $/M tokens |
|---|---|---|---|---|
| Mixtral 8x7b | 8× MI325X 256GB | vultr | 69,692tok/s | $0.064 |
| Mixtral 8x7b | 8× MI300X 192GB | vultr | 53,334tok/s | $0.077 |
| Llama 2 70B | 8× MI325X 256GB | vultr | 34,555tok/s | $0.129 |
| Llama 2 70B | 8× MI300X 192GB | vultr | 27,804tok/s | $0.148 |
| Llama 2 70B | 32× H100 80GB SXM | coreweave | 124,879tok/s | $0.177 |
| Llama 2 70B | 8× H200 141GB | lambda | 31,267tok/s | $0.276 |
Electricity by major AI city
Industrial meter rate × PUE (cooling overhead) = what the GPU actually pays for power. Same hyperscale workload (8×H100 SXM running MLPerf Llama-2-70B at 5,500 tok/s, 5,600W IT draw, PUE-adjusted facility draw). 5.3× differential from cheapest to most expensive. Sources: EIA (US, verified), Eurostat (EU + UK, verified), DESNZ (UK direct, verified), training approximations marked.
| City | Meter | PUE | Effective $/kWh | gCO2/kWh | $/M (workload) | kg CO2 (workload) |
|---|---|---|---|---|---|---|
| Abu DhabiUAE | $0.0500 | 1.45 | $0.0725 | 468 | $0.0205 | 0.230 |
| Council BluffsUS | $0.0643 | 1.15 | $0.0739 | 417 | $0.0209 | 0.163 |
| QuincyUS | $0.0685 | 1.15 | $0.0788 | 287 | $0.0223 | 0.112 |
| DallasUS | $0.0626 | 1.32 | $0.0826 | 333 | $0.0234 | 0.149 |
| AtlantaUS | $0.0670 | 1.25 | $0.0838 | 382 | $0.0237 | 0.162 |
| MemphisUS | $0.0706 | 1.28 | $0.0904 | 407 | $0.0256 | 0.177 |
| RenoUS | $0.0766 | 1.25 | $0.0958 | 287 | $0.0271 | 0.121 |
| PhoenixUS | $0.0717 | 1.40 | $0.1004 | 319 | $0.0284 | 0.152 |
| StockholmSweden | $0.1033 | 1.10 | $0.1136 | 35 | $0.0321 | 0.013 |
| AshburnUS | $0.1025 | 1.25 | $0.1281 | 269 | $0.0362 | 0.114 |
| SeoulSouth Korea | $0.1100 | 1.30 | $0.1430 | 417 | $0.0404 | 0.184 |
| MumbaiIndia | $0.1000 | 1.45 | $0.1450 | 670 | $0.0410 | 0.330 |
| ParisFrance | $0.1388 | 1.22 | $0.1693 | 41 | $0.0479 | 0.017 |
| HobartAustralia | $0.1450 | 1.30 | $0.1885 | — | $0.0533 | — |
| AmsterdamNetherlands | $0.1722 | 1.20 | $0.2066 | 254 | $0.0584 | 0.103 |
| TokyoJapan | $0.1600 | 1.35 | $0.2160 | 477 | $0.0611 | 0.219 |
| MelbourneAustralia | $0.1600 | 1.40 | $0.2240 | — | $0.0634 | — |
| BrisbaneAustralia | $0.1550 | 1.48 | $0.2294 | — | $0.0649 | — |
| SydneyAustralia | $0.1650 | 1.42 | $0.2343 | — | $0.0663 | — |
| FrankfurtGermany | $0.1956 | 1.20 | $0.2347 | 330 | $0.0664 | 0.134 |
| SingaporeSingapore | $0.1800 | 1.45 | $0.2610 | 497 | $0.0738 | 0.245 |
| DublinIreland | $0.2651 | 1.12 | $0.2969 | 256 | $0.0840 | 0.098 |
| AdelaideAustralia | $0.2200 | 1.45 | $0.3190 | — | $0.0902 | — |
| LondonUK | $0.3150 | 1.22 | $0.3843 | 217 | $0.1087 | 0.090 |
Pass ?cityCode=us-memphis to /api/v1/cost/run-locally and the calculator uses that city's rate instead of the $0.12 default. Full roster: /api/v1/electricity.
How cross-check works
Two or more MLPerf labs published the same (model, hardware) result within 10%. The strongest signal. Headline rows above all use this.
Same AA provider scraped twice >5 days apart agreed within 10%. Confirms the number is stable, not a one-off blip. Starts populating a week after deploy.
AA per-stream throughput lies in the physically-plausible band given MLPerf hardware ceiling. Weaker than multi-submitter but rules out rows that exceed hardware capability.
Full JSON roster: GET /api/v1/health/cross-check. Filter by ?axis= or ?family=. Each row carries the full evidence used in the upgrade.