Open Models inference pricing
Same models.
Cheaper hours.
Flexible work has a timing advantage. Move batch OCR, embeddings or evals to quieter hours and use lower dynamic rates.
Compare From prices below: the lowest price the dynamic multiplier can reach for each model. Prices move up or down within each model’s range.
Explore model ratesCompare model
inference rates.
From prices in EUR, updated 30 September 2026.
Prices move within a per-model range.
No models match your search.
Chat, code & reasoning
| Model | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
qwen3.6-35b-a3bReasoning with a long context · 256k contextAWQ-INT4 · 131 tok/s · 419 ms TTFT | From €0.0808 | From €0.542 |
qwen3.8-27bDense reasoning model · 64k contextW4A16 (INT4) · 58 tok/s · 328 ms TTFT | From €0.0816 | From €1.09 |
gemma-4-26b-a4bLong-context text generation · 256k contextAWQ-INT4 · 140 tok/s · 370 ms TTFT | From €0.0398 | From €0.149 |
mistral-small-24bText generation · 32k contextAWQ-INT4 · 32 tok/s · 399 ms TTFT | From €0.034 | From €0.0547 |
laguna-xs-2.1Coding and reasoning · 128k contextINT4 · 140 tok/s · 396 ms TTFT | From €0.0409 | From €0.0827 |
gpt-oss-20bReasoning · 128k contextMXFP4 · 180 tok/s · 372 ms TTFT | From €0.0183 | From €0.0697 |
hypernova-60bReasoning · 128k contextMXFP4 · 174 tok/s · 436 ms TTFT | From €0.043 | From €0.215 |
nemotron3-nano-30bLong-context reasoning · 256k contextAWQ-INT4 · 215 tok/s · 390 ms TTFT | From €0.034 | From €0.136 |
Vision & OCR
| Model | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
qwen3-vl-30b-a3bImage understanding · 32k contextAWQ-INT4 · 179 tok/s · 515 ms TTFT | From €0.101 | From €0.402 |
nemotron3-omni-30bVision and audio understanding · 128k contextW4A16 (INT4) · 187 tok/s · 397 ms TTFT | From €0.06 | From €0.18 |
deepseek-ocr-2Document OCR · 8k contextBF16 · 144 tok/s · 900 ms TTFT | From €0.0232 | From €0.0232 |
Text embeddings
| Model | Input / 1M tokens |
|---|---|
qwen3-embedding-8bText embeddings · 7k contextBF16 | From €0.00774 |
Speech & audio
| Model | Per audio minute |
|---|---|
voxtral-mini-3bTranscription and audio understanding · 8k contextBF16 · 25 tok/s (chat) · 855 ms TTFT | From €0.000774 |
whisper-large-v3-turboSpeech to text · 30 s / segmentFP16 · 21.8× real-time (1 stream) | From €0.000155 |
kokoro-82mText to speechFP32 · ~1× real-time (serial processing) | From €0.00399 |
Image generation
| Model | Per image |
|---|---|
flux-2-kleinImage generationBF16 · 4.15 s / image | From €0.00132 |
These are catalogue rates for the listed models, not a tariff for every API route. For OCR and audio, check the route and billing unit before estimating your workload.
From prices in EUR, updated on 30 September 2026: the lowest price the dynamic multiplier can reach for each model. Confirm current rates and access in the console. Performance figures are indicative and shown per stream; TTFT means time to first token. Reasoning tokens count as output.
A rate you can
trace to the request.
For eligible models, the multiplier is set when you submit a request. Check recorded usage and charges in your console.
- 01
Start from the From price.
Each model has a From price for its billing unit.
- 02
Follow the hour’s rate.
Eligible models follow their GPU class’s hourly schedule, never below the From price.
- 03
Check your usage.
Review the request usage and charges recorded for your account.
A few useful
details.
Can dynamic prices go above list?
Yes. Prices follow demand and can rise above the list rate. Use the published GPU-class schedule when timing flexible work.
Does every model use dynamic pricing?
No. Dynamic pricing applies to eligible models. Other models use a fixed rate.
How is usage measured?
Start with the model and its API route. The catalogue uses tokens, audio minutes or generated images; dedicated parsing routes document per-image billing. Confirm the unit for your selected route in the console. Read the billing-unit guide.
What about Serverless pricing?
Serverless is opening gradually. GPU classes, scaling settings and billing terms are confirmed during onboarding. Join the waitlist with your workload; Open Models rates are not a Serverless GPU quote.
Explore Serverless