Start with the model, route and billing unit
Check the model’s rate table before estimating a request. Text models use input and output token rates; other models may use audio minutes or generated images. Keep the units separate when comparing costs.
The catalogue shows From prices in EUR: the lowest price the dynamic multiplier can reach for each model. Depending on the hour, prices can be higher than the From price; some models are effectively fixed. Confirm the active rate and billing unit in the console before estimating production costs.
| Workload | What to check |
|---|---|
| Text and vision chat | For token-billed models, use the input and output rates for that model. Reasoning can count towards output tokens. |
| Document OCR and parsing | The Vision & OCR catalogue lists token rates for its models. The dedicated /v1/parse route documents billing per processed image. Do not apply a token price to a parsing request or infer its model support from the catalogue category. |
| Speech and audio | The catalogue quotes audio minutes. Confirm how the selected transcription or speech route meters the request; a duration returned in the response does not by itself establish the charged unit. |
| Image generation | Use the selected generation model’s per-image rate. This is separate from document parsing and vision chat. |
Apply the hourly multiplier
Eligible shared-pool models use a GPU-class pricing schedule. The multiplier in force when a request is submitted is used for that request. Moving flexible work to a lower-priced hour can reduce its cost; the schedule can also rise above the From price.
From prices already include the lowest multiplier a model can reach, so do not multiply them again. For a token-billed request:
cost at From prices = (input tokens ÷ 1,000,000 × From input price)
+ (output tokens ÷ 1,000,000 × From output price)
inference cost = list rate × multiplier recorded at submission
(within the model’s range, never below the cost at From prices)
Your console records the rate applied to each request. This formula covers the model inference portion. Optional server-tool calls can incur separate charges. Check your measured usage and billing details rather than treating a token estimate as an invoice.
A worked example
Example token volumes, using the From prices updated on 30 September 2026. The calculation below uses the same catalogue as the pricing page; the actual request cost depends on the multiplier in force.
Using the From prices for gemma-4-26b-a4b, assume 1,000,000 input tokens and 500,000 output tokens. Input: From €0.0398 / 1M tokens; output: From €0.149 / 1M tokens.
| Scenario | Calculation | Inference cost |
|---|---|---|
| Input tokens | 1 × €0.0398 | €0.0398 |
| Output tokens | 0.5 × €0.149 | €0.0745 |
| Total at the From prices | €0.0398 + €0.0745 | €0.1143 |
These totals cover inference in EUR at the From prices, the lowest the dynamic multiplier can reach, before any separately billed tool calls or applicable taxes. Changing the prompt length or generated output changes the total, even at the same hourly multiplier.
Decide what can wait
Batch OCR, offline evaluations and embeddings can often run within a flexible time window. Keep time-sensitive requests on the schedule your users need. Compare the published schedule for the relevant GPU class; cheaper hours are not a fixed daytime promise for every model.
The pricing page compares From prices across models. It does not show an hourly forecast. Check current rates and model eligibility in the console before scheduling work.
Check the actual request
Read the usage returned by the API and the charges recorded for your account in the console. Response formats vary by route: token counts are not the same as a final price, and speech generation returns audio rather than a JSON usage block. Reasoning can contribute to billed output; do not equate the visible answer’s length with every token generated by the model. Optional web search is charged separately from token inference.
Compare model From prices or make your first API call.
Rate source: the model catalogue updated 30 September 2026, also available as Markdown. The tables and worked example use that same catalogue. The API reference describes route parameters and response formats. Neither source replaces the active rates and charges for your account; no hourly multiplier is inferred from the catalogue From prices.