TokenPylon

1602 model families, 155 providers, 3301 price changes in 30 days

Prices are re-read from every provider daily and appended to a ledger, never overwritten. 1863 of the recent changes were drops. Watch a family to get an email the morning after it moves.

Price changes, last 30 days

WhenModelProviderSide$/M before → afterChange
Sep 29claude-fable-5-1anthropicvia anthropicoutput10.00 → 50.00+400.0%Watch
Sep 29claude-opus-5-5anthropicvia anthropicoutput5.00 → 20.00+300.0%Watch
Sep 29glm-5v-turbozaivia models-devinput1.20 → 5.00+316.7%Watch
Sep 29glm-5v-turbozaivia models-devoutput4.00 → 22.00+450.0%Watch
Sep 29qwen3.7-plusdashscopevia models-devinput0.500 → 0.400-20.0%Watch
Sep 29qwen3.7-plusdashscopevia models-devoutput3.00 → 1.60-46.7%Watch
Sep 29vercel_ai_gateway/spacexai/grok-4.7vercel_ai_gatewayvia vercel-gatewayinput1.20 → 2.00+66.7%Watch
Sep 29vercel_ai_gateway/spacexai/grok-4.7vercel_ai_gatewayvia vercel-gatewayoutput3.60 → 6.00+66.7%Watch
Sep 29vercel_ai_gateway/klingai/kling-v3.0-t2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.224 → 0.336+50.0%Watch
Sep 29vercel_ai_gateway/klingai/kling-v3.0-t2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.308 → 0.224-27.3%Watch
Sep 29vercel_ai_gateway/klingai/kling-v3.0-t2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.252 → 0.308+22.2%Watch
Sep 29vercel_ai_gateway/klingai/kling-v3.0-t2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.168 → 0.252+50.0%Watch
Sep 29vercel_ai_gateway/klingai/kling-v3.0-t2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.392 → 0.168-57.1%Watch
Sep 29vercel_ai_gateway/klingai/kling-v3.0-motion-control@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.168 → 0.126-25.0%Watch
Sep 29vercel_ai_gateway/klingai/kling-v3.0-i2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.224 → 0.336+50.0%Watch
Sep 29vercel_ai_gateway/klingai/kling-v3.0-i2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.308 → 0.224-27.3%Watch
Sep 29vercel_ai_gateway/klingai/kling-v3.0-i2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.252 → 0.308+22.2%Watch
Sep 29vercel_ai_gateway/klingai/kling-v3.0-i2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.168 → 0.252+50.0%Watch
Sep 29vercel_ai_gateway/klingai/kling-v3.0-i2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.392 → 0.168-57.1%Watch
Sep 29vercel_ai_gateway/klingai/kling-v2.6-t2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.042 → 0.070+66.7%Watch
Sep 29vercel_ai_gateway/klingai/kling-v2.6-t2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.140 → 0.042-70.0%Watch
Sep 29vercel_ai_gateway/klingai/kling-v2.6-motion-control@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.112 → 0.070-37.5%Watch
Sep 29vercel_ai_gateway/klingai/kling-v2.6-i2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.042 → 0.070+66.7%Watch
Sep 29vercel_ai_gateway/klingai/kling-v2.6-i2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.140 → 0.042-70.0%Watch
Sep 29vercel_ai_gateway/klingai/kling-v2.5-turbo-t2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.070 → 0.042-40.0%Watch
Sep 29vercel_ai_gateway/klingai/kling-v2.5-turbo-i2v@defaultvercel_ai_gatewayvia vercel-gatewayoutput0.070 → 0.042-40.0%Watch
Sep 29vercel_ai_gateway/google/veo-3.1-lite-generate-001@720pvercel_ai_gatewayvia vercel-gatewayoutput0.050 → 0.030-40.0%Watch
Sep 29vercel_ai_gateway/google/veo-3.1-lite-generate-001@1080pvercel_ai_gatewayvia vercel-gatewayoutput0.080 → 0.050-37.5%Watch
Sep 29vercel_ai_gateway/google/veo-3.1-generate-001@720pvercel_ai_gatewayvia vercel-gatewayoutput0.400 → 0.200-50.0%Watch
Sep 29vercel_ai_gateway/google/veo-3.1-generate-001@4kvercel_ai_gatewayvia vercel-gatewayoutput0.600 → 0.400-33.3%Watch
Sep 29vercel_ai_gateway/google/veo-3.1-generate-001@1080pvercel_ai_gatewayvia vercel-gatewayoutput0.400 → 0.200-50.0%Watch
Sep 29vercel_ai_gateway/google/veo-3.1-fast-generate-001@720pvercel_ai_gatewayvia vercel-gatewayoutput0.150 → 0.100-33.3%Watch
Sep 29vercel_ai_gateway/google/veo-3.1-fast-generate-001@4kvercel_ai_gatewayvia vercel-gatewayoutput0.350 → 0.300-14.3%Watch
Sep 29vercel_ai_gateway/google/veo-3.1-fast-generate-001@1080pvercel_ai_gatewayvia vercel-gatewayoutput0.150 → 0.100-33.3%Watch
Sep 29vercel_ai_gateway/bytedance/seedance-v1.5-pro@1080pvercel_ai_gatewayvia vercel-gatewayoutput0.117 → 0.058-50.0%Watch
Sep 29vercel_ai_gateway/bytedance/seedance-v1.5-pro@480pvercel_ai_gatewayvia vercel-gatewayoutput0.024 → 0.012-49.8%Watch
Sep 29vercel_ai_gateway/bytedance/seedance-v1.5-pro@720pvercel_ai_gatewayvia vercel-gatewayoutput0.052 → 0.026-50.0%Watch
Sep 29vercel_ai_gateway/google/veo-3.0-fast-generate-001@1080pvercel_ai_gatewayvia vercel-gatewayoutput0.150 → 0.100-33.3%Watch
Sep 29vercel_ai_gateway/google/veo-3.0-fast-generate-001@720pvercel_ai_gatewayvia vercel-gatewayoutput0.150 → 0.100-33.3%Watch
Sep 29vercel_ai_gateway/google/veo-3.0-generate-001@1080pvercel_ai_gatewayvia vercel-gatewayoutput0.400 → 0.200-50.0%Watch

Showing the 40 most recent of 3301.

Price index, last 90 days

Median list price of every paid chat listing the catalog saw that day, one price per seller, in USD per million tokens. Only sources that polled that day count, so a missing day is a gap rather than a guess. The index is not held to a fixed basket: when a new source adds listings the median moves with them, and the listing count below says when that happened.

Median input$0.57
$0.45$0.56$0.662026-07-022026-09-28
Median output$1.85
$1.30$1.67$2.042026-07-022026-09-28
Listings priced4490
0259251832026-07-022026-09-28

2026-09-28: 4,490 listings from 6 of 6 sources; cheapest paid input $0.003, output $0.002. Over 89 indexed days, 343 listing-cuts and 293 listing-raises; 41 days had a source missing. Series by provider, family and /decide task via /api/v1/trends.

Price catalog

One row per model family, ranked by how many providers compete on it. Cheapest is by input + output list price; free and preview tiers are excluded. Full list at /api/v1/families.

FamilyProvidersCheapest input $/MCheapest output $/MVia
deepseek-v4510.0440.132streamlakeWatch
glm-5-3-flash420.0450.140inferencenetWatch
deepseek-v4-1400.1400.280amdWatch
glm-5-2400.4901.54decartWatch
glm-5-3400.3701.14rekaWatch
gpt-oss360.0180.090darkbloomWatch
qwen-3310.0100.000nebiusWatch
llama-3-3-70b-instruct300.2000.200nscaleWatch
kimi-k2-6290.4752.45inceptronWatch
deepseek-v3-2280.1340.202atlascloudWatch
kimi-k3280.6798.79sail_researchWatch
kimi-k2-7-code270.7133.00streamlakeWatch
qwen-3-235b270.0870.350openrouterWatch
qwen-3-8-27b260.1501.88deepinfraWatch
glm-5-1250.9663.04streamlakeWatch
deepseek-r1240.0250.025nscaleWatch
llama-3-1-8b-instruct230.0250.025inference_netWatch
gemma-4-26b220.0600.200rekaWatch
gemma-4-31b210.0800.300rekaWatch
minimax-m3210.2300.960coreweaveWatch
qwen-3-5-397b200.1681.01baiduWatch
qwen-3-6-35b190.0500.700darkbloomWatch
deepseek-v3180.2000.200hyperbolicWatch
glm-4-7170.4001.75deepinfraWatch
kimi-k2-5170.4502.25deepinfraWatch
qwen-3-32b170.0500.100lambda_aiWatch
mistral-small160.0500.080deepinfraWatch
claude-sonnet-4-6153.0015.00bedrock_converseWatch
deepseek-v3-1150.2500.950deepinfraWatch
glm-5150.6001.92openrouterWatch
qwen-3-coder150.1200.800openrouterWatch
claude-opus-4-7144.5022.50gmiWatch
claude-opus-4-8145.0025.00bedrockWatch
claude-sonnet-4-5143.0015.00bedrockWatch
mistral-nemo140.0180.030dekallmWatch
qwen-3-6-27b140.3002.00chutesWatch
claude-fable-51310.0050.00bedrockWatch
claude-haiku-4-5131.005.00bedrockWatch
claude-opus-4-5135.0025.00bedrockWatch
claude-opus-4-6135.0025.00bedrock_converseWatch
claude-opus-5135.0025.00bedrockWatch
claude-sonnet-5132.0010.00bedrockWatch
deepseek-chat130.2800.420deepseekWatch
llama-3-1-70b-instruct130.1000.100fireworks_aiWatch
llama-3-2-3b-instruct130.0150.025lambda_aiWatch
minimax-m2-7130.2100.840gmiWatch
mistral-large130.5001.50azure_aiWatch
qwen-2130.0200.060nebiusWatch
qwen-3-30b130.0480.193openrouterWatch
qwen-3-5-35b130.0560.448baiduWatch
minimax-m2-5120.2700.950veniceWatch
qwen-3-coder-30b120.0700.260huggingfaceWatch
qwen-3-coder-480b120.2501.00siliconflowWatch
claude-fable-5-11110.0050.00bedrockWatch
claude-opus-5-5114.0020.00databricksWatch
claude-sonnet-5-5112.0010.00bedrock_converseWatch
glm-4-6110.4001.75io_netWatch
grok-4110.2000.500azure_aiWatch
llama-3-1-405b-instruct110.1000.100fireworks_aiWatch
llama-4-maverick-128e110.0500.100lambda_aiWatch

Showing the 60 most-contested of 1602 families.

2000 throughput numbers verified by ≥2 independent measurements

Every cost quote on this site sits on numbers that came from somewhere real. We don't average estimates and call it a benchmark. When two independent labs (NVIDIA, Dell, Cisco, Lenovo, AMD, ASUSTeK, Google, Oracle, Supermicro, GigaComputing) publish the same MLPerf result within 10%, we mark itverified-cross-checkedand surface it here.

1000
multi-lab MLPerf agreements
162
AA rows within hardware ceilings
5909
total throughput rows tracked

Best deploy: Llama 3.1 70B, 1.2M tokens

The end-to-end optimizer pulls hosted-API pricing, GPU rental hourlies, and self-deploy electricity (with PUE and grid carbon) into one ranked list. Same workload, three honest answers: cheapest dollars, greenest emissions, and the closed-source alternative.

Cheapest by $
$0.1000/M tokens
Mode: hosted-api
fireworks_ai
Greenest by gCO2
0.0060 kg CO2
Mode: hosted-gpu
1× h200-141gb via modal
$0.8407/M

Full ranked list: POST /api/v1/cost/best-deploy with {family, inputTokens, outputTokens}. Returns hosted-API + hosted-GPU + self-deploy options ranked by cost; closed-source families return hosted-API only.

Headline agreements

Sorted by number of independent labs that published the same benchmark. Delta is the tightest pairwise agreement. Click through to MLPerf v5.1 for the raw submissions.

ModelHardwareLabsTightest deltaThroughput
Llama 3 1 8B8× B200 180GB SXM6
gigacomputing, dell, supermicro, nvidia +2
0.06%142,136tok/s
Llama 3 1 405B8× B200 180GB SXM6
nvidia, gigacomputing, dell, google +2
0.10%1,624tok/s
Llama 2 70B8× MI325X 256GB5
gigacomputing, amd, asustek, supermicro +1
0.00%34,483tok/s
Mixtral 8x7b8× MI325X 256GB4
amd, gigacomputing, asustek, nvidia
0.00%68,781tok/s
Llama 2 70B8× H200 141GB SXM4
dell, asustek, google, nvidia
0.00%35,317tok/s
Llama 2 70B8× B200 180GB SXM4
nvidia, dell, supermicro, gigacomputing
0.17%101,527tok/s

Cheapest self-host options

Cost per million tokens for cross-checked throughput rows joined against verified public hourly GPU pricing. Single-node configurations only (multi-node MLPerf rows excluded — not rentable at public per-hour rates). Math:($/hr × gpu_count × 1M) / (tok/s × 3600)

ModelHardwareProviderThroughput$/M tokens
Mixtral 8x7b8× MI325X 256GBvultr69,692tok/s$0.064
Mixtral 8x7b8× MI300X 192GBvultr53,334tok/s$0.077
Llama 2 70B8× MI325X 256GBvultr34,555tok/s$0.129
Llama 2 70B8× MI300X 192GBvultr27,804tok/s$0.148
Llama 2 70B32× H100 80GB SXMcoreweave124,879tok/s$0.177
Llama 2 70B8× H200 141GBlambda31,267tok/s$0.276

Electricity by major AI city

Industrial meter rate × PUE (cooling overhead) = what the GPU actually pays for power. Same hyperscale workload (8×H100 SXM running MLPerf Llama-2-70B at 5,500 tok/s, 5,600W IT draw, PUE-adjusted facility draw). 5.3× differential from cheapest to most expensive. Sources: EIA (US, verified), Eurostat (EU + UK, verified), DESNZ (UK direct, verified), training approximations marked.

CityMeterPUEEffective $/kWhgCO2/kWh$/M (workload)kg CO2 (workload)
Abu DhabiUAE$0.05001.45$0.0725468$0.02050.230
Council BluffsUS$0.06431.15$0.0739417$0.02090.163
QuincyUS$0.06851.15$0.0788287$0.02230.112
DallasUS$0.06261.32$0.0826333$0.02340.149
AtlantaUS$0.06701.25$0.0838382$0.02370.162
MemphisUS$0.07061.28$0.0904407$0.02560.177
RenoUS$0.07661.25$0.0958287$0.02710.121
PhoenixUS$0.07171.40$0.1004319$0.02840.152
StockholmSweden$0.10331.10$0.113635$0.03210.013
AshburnUS$0.10251.25$0.1281269$0.03620.114
SeoulSouth Korea$0.11001.30$0.1430417$0.04040.184
MumbaiIndia$0.10001.45$0.1450670$0.04100.330
ParisFrance$0.13881.22$0.169341$0.04790.017
HobartAustralia$0.14501.30$0.1885—$0.0533—
AmsterdamNetherlands$0.17221.20$0.2066254$0.05840.103
TokyoJapan$0.16001.35$0.2160477$0.06110.219
MelbourneAustralia$0.16001.40$0.2240—$0.0634—
BrisbaneAustralia$0.15501.48$0.2294—$0.0649—
SydneyAustralia$0.16501.42$0.2343—$0.0663—
FrankfurtGermany$0.19561.20$0.2347330$0.06640.134
SingaporeSingapore$0.18001.45$0.2610497$0.07380.245
DublinIreland$0.26511.12$0.2969256$0.08400.098
AdelaideAustralia$0.22001.45$0.3190—$0.0902—
LondonUK$0.31501.22$0.3843217$0.10870.090

Pass ?cityCode=us-memphis to /api/v1/cost/run-locally and the calculator uses that city's rate instead of the $0.12 default. Full roster: /api/v1/electricity.

How cross-check works

multi-submitter

Two or more MLPerf labs published the same (model, hardware) result within 10%. The strongest signal. Headline rows above all use this.

temporal

Same AA provider scraped twice >5 days apart agreed within 10%. Confirms the number is stable, not a one-off blip. Starts populating a week after deploy.

cross-source

AA per-stream throughput lies in the physically-plausible band given MLPerf hardware ceiling. Weaker than multi-submitter but rules out rows that exceed hardware capability.

Full JSON roster: GET /api/v1/health/cross-check. Filter by ?axis= or ?family=. Each row carries the full evidence used in the upgrade.