Per-token billing
Serverless requests are billed per token, in USD per million tokens:
| Token type | Billed at |
|---|---|
| Input tokens | The model’s input rate |
| Cached input tokens | The model’s cached rate, when a prompt prefix is reused |
| Output tokens | The model’s output rate |
| Reasoning tokens | The output rate |
The usage object in every response reports the tokens you were billed for.
Credits
Self-serve accounts use prepaid credits: usage is deducted as requests complete. Requests are rejected with 402 when the
balance runs out. Credit expiry, refunds, and taxes are covered by the
Billing, Credits & Refund Policy.
Enterprise invoicing
Enterprise customers can be invoiced monthly. Contact sales.
Dedicated endpoints
Dedicated endpoints are billed per GPU-hour rather than per token — see Dedicated endpoints.