LLM API pricing by model, provider and Region

What 94 language models cost per 1M tokens on the vendor APIs, AWS Bedrock, Azure AI Foundry and GCP Vertex AI: 213 offers, with Region availability and Azure deployment profiles beside each price, so one model can be compared across every place that sells it.

Prices as of 94 models 11 endpoints 80 Regions USD per 1M tokens JSON CSV Sources FAQ

AWS Bedrock

Of the 30 Bedrock models priced per Region, 23 cost more in some Regions. Claude carries one figure. Details

Azure AI Foundry

The price follows the deployment profile, not the country: Global Standard is the cheapest. Details

GCP Vertex AI

Google publishes endpoint scopes (global, us, eu) instead of a Region list. Details

94 of 94 models · 213 of 213 offers
Global
one worldwide endpoint, no Region to pick
cyan dot
cheapest provider in an open row
“from”
lowest of a model’s providers
$1.25 – $1.50
price moves by Region
“—”
not published
LLM prices in USD per 1M tokens, one row per model; expand a row for its providers
VendorProvidersRegions
Anthropic1Mfrom 4.00from 20.004 providersGlobal + 35 regions
Anthropic1Mfrom 5.00from 25.004 providersGlobal + 35 regions
Anthropic1Mfrom 5.00from 25.004 providersGlobal + 35 regions
Anthropic1Mfrom 5.00from 25.004 providersGlobal + 33 regions
Anthropic1Mfrom 5.00from 25.004 providersGlobal + 33 regions
Anthropic1Mfrom 10.00from 50.004 providersGlobal + 35 regions
Anthropic1Mfrom 10.00from 50.004 providersGlobal + 33 regions
Anthropic1Mfrom 2.00from 10.004 providersGlobal + 35 regions
Anthropic1Mfrom 2.00from 10.004 providersGlobal + 35 regions
Anthropic1Mfrom 3.00from 15.004 providersGlobal + 33 regions
Anthropic200Kfrom 3.00from 15.004 providersGlobal + 35 regions
Anthropic200Kfrom 5.00from 25.004 providersGlobal + 33 regions
Anthropic200Kfrom 1.00from 5.004 providersGlobal + 33 regions
OpenAI1.1M10.0050.002 providersGlobal + 27 regions
OpenAI1.1M2.0010.002 providersGlobal + 26 regions
OpenAI1.1M0.100.502 providersGlobal + 26 regions
OpenAI1.1M4.0020.002 providersGlobal + 26 regions
OpenAI1.1M2.0012.003 providersGlobal + 60 regions
OpenAI1.1M0.201.203 providersGlobal + 60 regions
OpenAI1.1M5.0030.002 providersGlobal + 31 regions
OpenAI1M30.00180.001 providerGlobal
OpenAI1.1M2.5015.003 providersGlobal + 37 regions
OpenAI272K0.754.502 providersGlobal + 33 regions
OpenAI272K0.201.252 providersGlobal + 33 regions
OpenAI1.1M30.00180.002 providersGlobal + 25 regions
OpenAI400K1.7514.002 providersGlobal + 30 regions
OpenAI272K1.7514.002 providersGlobal + 32 regions
OpenAI272K21.00168.002 providersGlobal + 24 regions
OpenAI272K1.2510.002 providersGlobal + 35 regions
OpenAI272K1.2510.002 providersGlobal + 27 regions
OpenAI272K0.252.002 providersGlobal + 27 regions
OpenAI272K0.050.402 providersGlobal + 27 regions
OpenAI400K15.00120.002 providersGlobal + 29 regions
OpenAI1.0M2.008.002 providersGlobal + 30 regions
OpenAI128K2.5010.002 providersGlobal + 29 regions
OpenAI128K0.150.602 providersGlobal + 31 regions
OpenAI200K2.008.002 providersGlobal + 27 regions
OpenAI200K1.104.402 providersGlobal + 30 regions
OpenAI8K0.13—2 providersGlobal + 32 regions
OpenAI8K0.02—2 providersGlobal + 32 regions
Google1.0M0.753.752 providersGlobal
Google1.0M0.753.752 providersGlobal
Google1.0M0.753.752 providersGlobal
Google1.0M1.509.002 providersGlobal
Google1.0M0.302.502 providersGlobal
Google1.0M0.251.502 providersGlobal
Google1.0M2.0012.002 providersGlobal
Google1.0M2.0012.002 providersGlobal
Google1.0M1.2510.002 providersGlobal
Google1.0M0.302.502 providersGlobal
Google8K0.20—2 providersGlobal
xAI500Kfrom 2.00from 6.002 providersGlobal + 28 regions
xAI500Kfrom 2.00from 6.003 providersGlobal + 34 regions
xAI500K2.006.001 providerGlobal
xAI1M1.252.502 providersGlobal + 4 regions
xAI2M0.200.503 providersGlobal + 41 regions
DeepSeek1Mfrom 0.435from 0.872 providersGlobal + 42 regions
DeepSeek1Mfrom 0.19from 0.512 providersGlobal + 42 regions
DeepSeek164Kfrom 0.2288from 0.34324 providersGlobal + 52 regions
Mistral262K0.501.503 providersGlobal + 49 regions
Mistral262K1.507.501 providerGlobal
Mistral40K2.005.001 providerGlobal
Mistral262K0.150.152 providersGlobal + 10 regions
Mistral128K0.200.202 providersGlobal + 10 regions
Mistral128K0.100.102 providersGlobal + 10 regions
Mistral256K0.300.901 providerGlobal
Mistral8K0.100.101 providerGlobal
Meta128Kfrom 0.20from 0.604 providersGlobal + 46 regions
Meta128Kfrom 0.11from 0.344 providersGlobal + 4 regions
Meta128Kfrom 0.59from 0.713 providersGlobal + 45 regions
Amazon1M1.37511.001 provider—
Amazon1M0.302.501 provider26 regions
Amazon1M2.5012.501 provider3 regions
Amazon300K0.803.201 provider20 regions
Amazon300K0.060.241 provider21 regions
Amazon128K0.0350.141 provider17 regions
Amazon8K0.02—1 provider21 regions
Alibaba262Kfrom 0.221.802 providers8 regions
Alibaba256K0.501.201 provider3 regions
Alibaba262Kfrom 0.22from 0.882 providers10 regions
Alibaba256K0.141.201 provider10 regions
Alibaba131Kfrom 0.15from 0.592 providersGlobal + 13 regions
OpenAI128Kfrom 0.09from 0.364 providersGlobal + 56 regions
OpenAI128Kfrom 0.07from 0.253 providersGlobal + 15 regions
Moonshot AI1.0Mfrom 3.00from 15.002 providers52 regions
Moonshot AI262K0.603.002 providers47 regions
Moonshot AI128K0.602.503 providers34 regions
Cohere256K2.5010.002 providersGlobal + 41 regions
Cohere128K0.12—3 providersGlobal + 64 regions
Z.ai1M1.755.501 provider36 regions
Z.ai200K1.003.202 providers10 regions
Z.ai203K0.602.201 provider10 regions
MiniMax1M0.331.321 provider22 regions
MiniMax1M0.301.201 provider13 regions
“from” is the cheapest among the providers matching your filters; narrow the Region or the Provider and the number moves“—” means the provider does not publish that price; “unpriced” means AWS lists the model but meters it in GovCloud only

How to read the price table

One row per model. Expand it and you get one line per place that sells it: the vendor’s API, AWS Bedrock, Azure AI Foundry and GCP Vertex AI. The same model is often priced differently in those places and almost always has a different Region story. That gap is what this page shows, and what an LLM gateway hides behind one API.

Model
Name, with the id you paste into a request underneath. An asterisk means our data has a known gap for this model; see Where the numbers come from.
Context
Maximum context window in tokens. The Cache & batch switch also shows the highest published output cap.
Input from · Output from
USD per 1M tokens, the lowest price among the providers that match your filters. “from” appears only when providers disagree; expand the row to see each one.
Cache read · Cache write · Batch in · Batch out
The same rule for prompt-cache and batch prices, behind the Cache & batch switch.
Providers · Regions
How many endpoints sell the model and where it is reachable. “Global” means one worldwide endpoint or routing tier with no Region to choose: a vendor API, Bedrock’s Global tier, Azure Global Standard or Vertex’s global scope. “Global + 35 regions” means both exist for that model.
mark
Inside an expanded row, the cheapest input or output price for that model among the providers shown.
“—” and “unpriced”
A dash: no source we read publishes that number. “unpriced”: AWS lists the model on Bedrock but meters it only in GovCloud, so there is no commercial price to print.

AWS Bedrock pricing by Region

AWS publishes a per-Region price for 30 of the 48 models sold on Bedrock, and 23 of those are billed at a different rate in different Regions: their rows show a range rather than one number. The other 18, every Claude model among them, are published as a single figure with no per-Region breakdown, so their Regions column says where the model runs, not what it costs there. Hover Regions for the routing tiers. In-Region and Geo keep a request inside a geography; Global does not.

Azure OpenAI pricing by deployment profile

Azure AI Foundry, which includes Azure OpenAI, prices by deployment profile, not by country. Global Standard is the cheapest and is the headline number. Data Zone Standard keeps traffic inside the EU or US zone and costs more; the two zones are not priced alike, which is why 39 models show a range. Regional Standard, where Microsoft meters it, is the dearest. The price cell’s tooltip names every profile the feed found.

GCP Vertex AI pricing and endpoint scopes

Google renders its availability matrix in the browser, so no feed can read it. Vertex rows show the endpoint scopes Google states in prose instead: global, us and eu. Vertex list prices match the Gemini API for the models here (see the Gemini rows); the choice between the two is billing, IAM and residency.

Where the numbers come from

Clouds that bill a model themselves are read directly: the AWS Price List for Bedrock and the Azure Retail Prices API for Azure. Vendor API prices come from pydantic’s genai-prices, an open hand-maintained table, cross-checked against the vendors’ own pricing pages; 34 figures matched on the last run and none disagreed. Where no feed names a model under the id it is sold at, the price is pinned by hand from the vendor’s page: 63 of the 213 offers here, each listed with its source in the feed’s dataQuality block. The same block lists every known gap; 27 of the 94 models carry one and wear an asterisk in the table.

Availability
Vertex AI locations
This dataset
/llm-pricing.json /llm-pricing.csv CORS-open, no key.

LLM pricing FAQ

How much does the Claude API cost per 1M tokens?
Claude Opus 5.5 costs $4 per 1M input tokens and $20 per 1M output tokens on the Anthropic API. Prompt caching is $0.20 to read and $5 to write. On AWS Bedrock the same model is $4.40 in and $22 out. Every Claude model and endpoint has its own row above. For seat plans rather than API tokens, see Claude Enterprise pricing. Prices as of October 1, 2026.
How much does the Gemini API cost per 1M tokens?
Gemini 3.1 Pro is $2 per 1M input tokens and $12 per 1M output tokens. Gemini 3.8 Flash is $0.75 in and $3.75 out. GCP Vertex AI lists both at the same numbers, so the choice between the two is billing, IAM and residency rather than cost. Prices as of October 1, 2026.
What is the cheapest LLM API?
By list price, the lowest input price in this table is Nova Micro at $0.035 per 1M tokens on AWS Bedrock, and the lowest output price is Ministral 3 3B at $0.10. Cheapest depends on what you send: cache reads and batch rates cut the bill by half or more on several models, so switch on Cache & batch and sort by the column that matches your traffic. Both figures are computed over chat models only. Prices as of October 1, 2026.
Is AWS Bedrock cheaper than the Anthropic API?
No. Bedrock lists Claude Opus 5.5 at $4.40 / $22 against $4 / $20 on the Anthropic API. What Bedrock adds is Region choice, an AWS invoice and IAM instead of a second set of API keys.
Is Azure OpenAI more expensive in Europe?
Azure prices by deployment profile, not by country. GPT-6 Astra is $10 per 1M input tokens on Global Standard and $11 – $12 on Data Zone Standard. The EU zone sits at the low end of that range and Asia Pacific at the high end.
Do AWS Bedrock prices differ by Region?
AWS publishes a per-Region price for 30 of the 48 models on Bedrock, and 23 of those are billed at a different rate in different Regions; their rows show a range. Seven cost the same in every Region AWS sells them in. For the remaining 18, every Claude model among them, AWS states one figure and no per-Region breakdown, so the table prints that number.
Which models are available in EU Regions?
79 of the 94 models here are offered in at least one European Region, against 82 in the United States. Set the Region filter to EU (Europe) to see them. Vendor APIs are a single global endpoint, so they stay in the list under any Region filter.
Why are prices in USD and not EUR?
Every provider here sets and bills these prices in US dollars. AWS, Azure and GCP can invoice you in your local currency, but they convert the USD list price at their own rate on their own date. A converted number would move every day without adding anything true.
How often is this page updated?
Last checked on October 1, 2026; the numbers last changed on October 1, 2026. The same data is at secondstack.ai/llm-pricing.json and secondstack.ai/llm-pricing.csv, one row per offer, CORS-open and without a key.

The same prices, enforced per request

We keep this table because SecondGate, the gateway inside SecondStack, has to know what a request costs before it sends it. Your apps call one OpenAI-compatible endpoint with virtual keys; the gateway routes to the vendor APIs and cloud endpoints above under your own provider keys and meters spend against these prices. Budgets per user, per team and per API key, threshold alerts and per-request usage logs come with it, self-hosted on your infrastructure.