Model Pricing
Per-token rates for every model, long-context and US-hosted rates, scheduled rate changes, and a worked monthly example.
How token costs are calculated
Every agent request is metered by the token counts the model provider reports: input tokens, cached input tokens (written and read), and output tokens. Each count is multiplied by the rate per one million tokens for the endpoint that served the request, and the results are added up for the month. Self-serve plans add a 20% platform fee on top; enterprise agreements set their own terms. The Usage page reflects what you pay. Seats, allowances, and usage limits are covered on Usage & Billing.
Base rates
All prices and rates are per one million tokens and reflect the base pricing published by the third-party model providers (Anthropic, OpenAI, Google, and Fireworks AI). For models served through AWS Bedrock's US-hosted endpoints, a 10% surcharge applies to these base rates and the rates you are billed are the ones in the US-hosted endpoints table. Self-serve plans add the 20% platform fee on top.
Base provider pricing
These figures make model costs easier to compare, but provider-specific conditions can change the final price. Review the linked provider pricing before making a cost-sensitive decision.
| Model | Input | Cache write | Cache read | Output |
|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $12.50 | $0.25 | $50 |
| Claude Fable 5 | $10 | $12.50 | $1 | $50 |
| Claude Opus 5 | $5 | $6.25 | $0.50 | $25 |
| Claude Sonnet 5 | $2 | $2.50 | $0.20 | $10 |
| Claude Opus 4.8 | $5 | $6.25 | $0.50 | $25 |
| Claude Opus 4.7 | $5 | $6.25 | $0.50 | $25 |
| Claude Opus 4.6 | $5 | $6.25 | $0.50 | $25 |
| Claude Opus 4.5 | $5 | $6.25 | $0.50 | $25 |
| Claude Sonnet 4.6 | $3 | $3.75 | $0.30 | $15 |
| Claude Sonnet 4.5 | $3 | $3.75 | $0.30 | $15 |
| Claude Haiku 4.5 | $1 | $1.25 | $0.10 | $5 |
| Gemini 3.7 Flash | $0.75 | — | $0.075 | $3.75 |
| Gemini 3.1 Pro | $2 | — | $0.20 | $12 |
| Gemini 3.5 Flash | $1.50 | — | $0.15 | $9 |
| Gemini 3 Flash | $0.50 | — | $0.05 | $3 |
| Gemini 3.1 Flash Lite | $0.25 | — | $0.025 | $1.50 |
| GPT-6 Astra | $10 | $12.50 | $1 | $50 |
| GPT-5.6 Sol | $4 | $5 | $0.40 | $20 |
| GPT-5.6 Terra | $2 | $2.50 | $0.20 | $12 |
| GPT-5.6 Luna | $0.20 | $0.25 | $0.02 | $1.20 |
| GPT-5.5 | $5 | — | $0.50 | $30 |
| GPT-5.4 | $2.50 | — | $0.25 | $15 |
| GPT-5.3 Codex | $1.75 | — | $0.175 | $14 |
| GPT-5.2 | $1.75 | — | $0.175 | $14 |
| GPT-5.4 Mini | $0.75 | — | $0.075 | $4.50 |
| Kimi K3 | $3 | — | $0.30 | $15 |
| Kimi K2.7 Code | $0.95 | — | $0.19 | $4 |
| Kimi K2.6 | $0.95 | — | $0.16 | $4 |
| GLM 5.3 | $1.40 | — | $0.26 | $4.40 |
| GLM 5.2 | $1.40 | — | $0.14 | $4.40 |
A dash means there is no separate charge for that category. Claude Fable 5 and Claude Fable 5.1 are off by default; an organization admin can enable them from Settings → Organization → Models.
Long-context requests
Some models cost more when a single request carries a very large context. When the request's context exceeds the threshold below, the whole request is billed at the long-context rates: 2x the standard rate for input and cached tokens, and 1.5x for output. The rates below are the ones billed, including the US-hosted surcharge where it applies. No other model is billed at a long-context tier.
| Model | Applies when the context exceeds | Input | Cache write | Cache read | Output |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | 200,000 tokens | $4 | — | $0.40 | $18 |
| GPT-6 Astra | 272,000 tokens | $20 | $25 | $2 | $75 |
| GPT-5.6 Sol | 272,000 tokens | $11 | $13.75 | $1.10 | $49.50 |
| GPT-5.6 Terra | 272,000 tokens | $4.40 | $5.50 | $0.44 | $19.80 |
| GPT-5.6 Luna | 272,000 tokens | $0.44 | $0.55 | $0.044 | $1.98 |
| GPT-5.5 | 272,000 tokens | $11 | — | $1.10 | $49.50 |
| GPT-5.4 | 272,000 tokens | $5.50 | — | $0.55 | $24.75 |
The tier is applied per request, based on the size of that request's context, and the context window setting caps how large that context can be. Gemini 3.1 Pro and GPT-6 Astra default to a 300,000-token window, and GPT-5.6 Sol, Terra, and Luna default to 272,000 tokens; a larger window can be selected where the model supports one. GPT-5.4 and GPT-5.5 default to their full window of about 1 million tokens, so their requests can cross the threshold without any setting being changed.
US-hosted endpoints
Claude models and GPT-5.4, GPT-5.5, and GPT-5.6 are normally served through AWS Bedrock's US-hosted endpoints. AWS prices these endpoints 10% above the model provider's list price, so every rate for these models is the base rate multiplied by 1.1; these are the rates you are billed. For GPT-5.6 Sol, AWS applies the surcharge to OpenAI's list price rather than the lower rate OpenAI currently publishes. Gemini, Kimi, and GLM models carry no surcharge.
| Model | Input | Cache write | Cache read | Output |
|---|---|---|---|---|
| Claude Fable 5.1 | $11 | $13.75 | $0.275 | $55 |
| Claude Fable 5 | $11 | $13.75 | $1.10 | $55 |
| Claude Opus 5 | $5.50 | $6.875 | $0.55 | $27.50 |
| Claude Sonnet 5 | $2.20 | $2.75 | $0.22 | $11 |
| Claude Opus 4.8 | $5.50 | $6.875 | $0.55 | $27.50 |
| Claude Opus 4.7 | $5.50 | $6.875 | $0.55 | $27.50 |
| Claude Opus 4.6 | $5.50 | $6.875 | $0.55 | $27.50 |
| Claude Opus 4.5 | $5.50 | $6.875 | $0.55 | $27.50 |
| Claude Sonnet 4.6 | $3.30 | $4.125 | $0.33 | $16.50 |
| Claude Sonnet 4.5 | $3.30 | $4.125 | $0.33 | $16.50 |
| Claude Haiku 4.5 | $1.10 | $1.375 | $0.11 | $5.50 |
| GPT-5.6 Sol | $5.50 | $6.875 | $0.55 | $33 |
| GPT-5.6 Terra | $2.20 | $2.75 | $0.22 | $13.20 |
| GPT-5.6 Luna | $0.22 | $0.275 | $0.022 | $1.32 |
| GPT-5.5 | $5.50 | — | $0.55 | $33 |
| GPT-5.4 | $2.75 | — | $0.275 | $16.50 |
Scheduled rate changes
Google has scheduled a rate change for Gemini 3.7 Flash. Each request is billed at the rate in effect on the day it runs.
| Model | Through December 31, 2026 | From January 1, 2027 |
|---|---|---|
| Gemini 3.7 Flash | $0.75 input, $0.075 cache read, $3.75 output | $1.50 input, $0.15 cache read, $7.50 output |
New models may be added over time, and this page is updated whenever the catalog or a rate changes. Consult each provider's current API pricing:
Example: a month of usage
For illustration only, here is how a monthly usage charge is calculated for an organization that used five models. The Claude Sonnet 5 line shows the US-hosted surcharge, the two Gemini 3.1 Pro lines show the long-context multiplier, and the GPT-5.5 line shows both applied to the same requests.
| Model | Tokens | Rate per 1M | Cost |
|---|---|---|---|
| Claude Sonnet 5, served through AWS Bedrock (US-hosted, +10%) | 40,000,000 input | $2.20 ($2.00 + 10%) | $88.00 |
| 10,000,000 cache read | $0.22 ($0.20 + 10%) | $2.20 | |
| 8,000,000 output | $11.00 ($10.00 + 10%) | $88.00 | |
| Gemini 3.1 Pro, requests up to 200,000 tokens of context | 20,000,000 input | $2.00 | $40.00 |
| 4,000,000 output | $12.00 | $48.00 | |
| Gemini 3.1 Pro, 10 requests with 250,000 tokens of context | 2,500,000 input | $4.00 (2x) | $10.00 |
| 100,000 output | $18.00 (1.5x) | $1.80 | |
| GPT-5.5, served through AWS Bedrock (US-hosted), 5 requests with 300,000 tokens of context | 1,500,000 input | $11.00 (2x on $5.50) | $16.50 |
| 100,000 output | $49.50 (1.5x on $33.00) | $4.95 | |
| Kimi K3 | 10,000,000 input | $3.00 | $30.00 |
| 2,000,000 output | $15.00 | $30.00 | |
| Gemini 3.7 Flash, rate through December 31, 2026 | 50,000,000 input | $0.75 | $37.50 |
| 10,000,000 output | $3.75 | $37.50 | |
| Usage total | $434.45 |
Each of the ten Gemini 3.1 Pro long-context requests carried 250,000 tokens of context, over the 200,000-token threshold, so every token in those requests was billed at the long-context rate: 250,000 tokens at $4.00 per million is $1.00 per request instead of $0.50. The other Gemini 3.1 Pro requests stayed under the threshold and were billed at the standard rate. The GPT-5.5 requests show the two adjustments stacking: the US-hosted rate ($5.50 input, $33 output) is doubled for input and multiplied by 1.5 for output because each request's context exceeded 272,000 tokens.
On a self-serve plan, the 20% platform fee is added to the usage total ($86.89 here, for $521.34). Enterprise agreements set their own terms.
Continue to the Changelog.