# LLM API pricing by model, provider and Region

> What 94 LLMs cost per 1M tokens on vendor APIs, AWS Bedrock, Azure AI Foundry and Vertex AI with Region availability. Updated Oct 1, 2026. Free JSON, CSV.

- Canonical: https://secondstack.ai/llm-pricing/
- Updated: 2026-10-01

---

What 94 language models cost per 1M tokens on the vendor APIs, AWS Bedrock, Azure AI Foundry and GCP Vertex AI: 213 offers, with Region availability and Azure deployment profiles beside each price, so one model can be compared across every place that sells it.

- Prices as of: October 1, 2026 (last price change)
- Last checked: October 1, 2026
- Coverage: 94 models, 213 offers, 11 endpoints, 80 Regions
- Currency: USD per 1M tokens

## The data

The full table is served as data rather than repeated here:

- https://secondstack.ai/llm-pricing.json: every model, offer, price, Region list, per-offer source URL and a `dataQuality` block naming the known gaps.
- https://secondstack.ai/llm-pricing.csv: the same dataset flattened to one row per offer.

CORS-open, no key.

## How the price moves on each cloud

- **AWS Bedrock**: Of the 30 Bedrock models priced per Region, 23 cost more in some Regions. Claude carries one figure.
- **Azure AI Foundry**: The price follows the deployment profile, not the country: Global Standard is the cheapest.
- **GCP Vertex AI**: Google publishes endpoint scopes (`global`, `us`, `eu`) instead of a Region list.

## How to read the price table

One row per model. Expand it and you get one line per place that sells it: the vendor’s API, AWS Bedrock, Azure AI Foundry and GCP Vertex AI. The same model is often priced differently in those places and almost always has a different Region story. That gap is what this page shows, and what [an LLM gateway](/blog/what-is-an-llm-gateway/) hides behind one API.

- **Model**: Name, with the id you paste into a request underneath. An asterisk means our data has a known gap for this model; see [Where the numbers come from](/llm-pricing/#sources).
- **Context**: Maximum context window in tokens. The Cache & batch switch also shows the highest published output cap.
- **Input from · Output from**: USD per 1M tokens, the lowest price among the providers that match your filters. “from” appears only when providers disagree; expand the row to see each one.
- **Cache read · Cache write · Batch in · Batch out**: The same rule for prompt-cache and batch prices, behind the Cache & batch switch.
- **Providers · Regions**: How many endpoints sell the model and where it is reachable. “Global” means one worldwide endpoint or routing tier with no Region to choose: a vendor API, Bedrock’s Global tier, Azure Global Standard or Vertex’s `global` scope. “Global + 35 regions” means both exist for that model.
- **Cyan dot**: Inside an expanded row, the cheapest input or output price for that model among the providers shown.
- **“—” and “unpriced”**: A dash: no source we read publishes that number. “unpriced”: AWS lists the model on Bedrock but meters it only in GovCloud, so there is no commercial price to print.

## AWS Bedrock pricing by Region

AWS publishes a per-Region price for 30 of the 48 models sold on Bedrock, and 23 of those are billed at a different rate in different Regions: their rows show a range rather than one number. The other 18, every Claude model among them, are published as a single figure with no per-Region breakdown, so their Regions column says where the model runs, not what it costs there. Hover Regions for the routing tiers. In-Region and Geo keep a request inside a geography; Global does not.

## Azure OpenAI pricing by deployment profile

Azure AI Foundry, which includes Azure OpenAI, prices by deployment profile, not by country. Global Standard is the cheapest and is the headline number. Data Zone Standard keeps traffic inside the EU or US zone and costs more; the two zones are not priced alike, which is why 39 models show a range. Regional Standard, where Microsoft meters it, is the dearest. The price cell’s tooltip names every profile the feed found.

## GCP Vertex AI pricing and endpoint scopes

Google renders its availability matrix in the browser, so no feed can read it. Vertex rows show the endpoint scopes Google states in prose instead: `global`, `us` and `eu`. Vertex list prices match the Gemini API for the models here (see the Gemini rows); the choice between the two is billing, IAM and residency.

## Where the numbers come from

Clouds that bill a model themselves are read directly: the AWS Price List for Bedrock and the Azure Retail Prices API for Azure. Vendor API prices come from pydantic’s genai-prices, an open hand-maintained table, cross-checked against the vendors’ own pricing pages; 34 figures matched on the last run and none disagreed. Where no feed names a model under the id it is sold at, the price is pinned by hand from the vendor’s page: 63 of the 213 offers here, each listed with its source in the feed’s dataQuality block. The same block lists every known gap; 27 of the 94 models carry one and wear an asterisk in the table.

- Price feed: [AWS Price List (Amazon Bedrock)](https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/AmazonBedrock/current/index.json)
- Price feed: [Azure Retail Prices API (Foundry Models)](https://prices.azure.com/api/retail/prices?api-version=2023-01-01-preview&$filter=serviceName%20eq%20'Foundry%20Models')
- Price table: [pydantic/genai-prices (price backbone)](https://raw.githubusercontent.com/pydantic/genai-prices/main/prices/data.json)
- Availability: [AWS Bedrock Regional availability](https://docs.aws.amazon.com/bedrock/latest/userguide/models-region-compatibility.html)
- Availability: [Azure model availability matrix](https://raw.githubusercontent.com/MicrosoftDocs/azure-ai-docs/main/articles/foundry/openai/includes/model-matrix/standard-global.md)
- Availability: [Vertex AI locations](https://cloud.google.com/vertex-ai/generative-ai/docs/learn/locations?hl=en)

## LLM pricing FAQ

### How much does the Claude API cost per 1M tokens?

Claude Opus 5.5 costs $4 per 1M input tokens and $20 per 1M output tokens on the Anthropic API. Prompt caching is $0.20 to read and $5 to write. On AWS Bedrock the same model is $4.40 in and $22 out. Every Claude model and endpoint has its own row above. For seat plans rather than API tokens, see [Claude Enterprise pricing](/blog/claude-enterprise-pricing/). Prices as of October 1, 2026.

### How much does the Gemini API cost per 1M tokens?

Gemini 3.1 Pro is $2 per 1M input tokens and $12 per 1M output tokens. Gemini 3.8 Flash is $0.75 in and $3.75 out. GCP Vertex AI lists both at the same numbers, so the choice between the two is billing, IAM and residency rather than cost. Prices as of October 1, 2026.

### What is the cheapest LLM API?

By list price, the lowest input price in this table is Nova Micro at $0.035 per 1M tokens on AWS Bedrock, and the lowest output price is Ministral 3 3B at $0.10. Cheapest depends on what you send: cache reads and batch rates cut the bill by half or more on several models, so switch on Cache & batch and sort by the column that matches your traffic. Both figures are computed over chat models only. Prices as of October 1, 2026.

### Is AWS Bedrock cheaper than the Anthropic API?

No. Bedrock lists Claude Opus 5.5 at $4.40 / $22 against $4 / $20 on the Anthropic API. What Bedrock adds is Region choice, an AWS invoice and IAM instead of a second set of API keys.

### Is Azure OpenAI more expensive in Europe?

Azure prices by deployment profile, not by country. GPT-6 Astra is $10 per 1M input tokens on Global Standard and $11 – $12 on Data Zone Standard. The EU zone sits at the low end of that range and Asia Pacific at the high end.

### Do AWS Bedrock prices differ by Region?

AWS publishes a per-Region price for 30 of the 48 models on Bedrock, and 23 of those are billed at a different rate in different Regions; their rows show a range. Seven cost the same in every Region AWS sells them in. For the remaining 18, every Claude model among them, AWS states one figure and no per-Region breakdown, so the table prints that number.

### Which models are available in EU Regions?

79 of the 94 models here are offered in at least one European Region, against 82 in the United States. Set the Region filter to EU (Europe) to see them. Vendor APIs are a single global endpoint, so they stay in the list under any Region filter.

### Why are prices in USD and not EUR?

Every provider here sets and bills these prices in US dollars. AWS, Azure and GCP can invoice you in your local currency, but they convert the USD list price at their own rate on their own date. A converted number would move every day without adding anything true.

### How often is this page updated?

Last checked on October 1, 2026; the numbers last changed on October 1, 2026. The same data is at secondstack.ai/llm-pricing.json and secondstack.ai/llm-pricing.csv, one row per offer, CORS-open and without a key.

## The same prices, enforced per request

We keep this table because SecondGate, the gateway inside SecondStack, has to know what a request costs before it sends it. Your apps call one OpenAI-compatible endpoint with virtual keys; the gateway routes to the vendor APIs and cloud endpoints above under your own provider keys and meters spend against these prices. Budgets per user, per team and per API key, threshold alerts and per-request usage logs come with it, self-hosted on your infrastructure.
