AI pricing
Every model call is metered in the tokens the provider reports: input, output, cached input and cache writes. Each model has its own price, so you can choose the model that fits the task and see exactly what it costs.
- Agent conversations and tool use
- Pipeline and plan authoring
- Predictions and forecasting
- Audio transcription
Prices
Live prices are briefly unavailable. They will be back shortly; in the meantime, contact us for a quote.
How it's measured
- Prompt tokens sent to the model, not served from cache, as the provider reports them.
- Tokens the model generates, including reasoning tokens.
- Prompt tokens served from the provider's cache, usually far cheaper than fresh input.
- Prompt tokens written to the provider's cache so later calls can reuse them.
- Charged per call only for models priced as a whole call rather than per token.
- Audio transcribed to text, measured by the length of the audio.
Example workloads
These workloads illustrate what is measured. Cost calculations are available when live rates return.
An agent turn
20,000 input tokens of context, 15,000 of them from cache, and a 1,000-token answer.