blade Get your API key

Open Models inference pricing

Same models.
Cheaper hours.

Flexible work has a timing advantage. Move batch OCR, embeddings or evals to quieter hours and use lower dynamic rates.

Compare From prices below: the lowest price the dynamic multiplier can reach for each model. Prices move up or down within each model’s range.

Explore model rates
No subscriptionNo minimum spendRate set at submissionTrack usage in the console

Compare model
inference rates.

From prices in EUR, updated 30 September 2026.
Prices move within a per-model range.

Chat, code & reasoning

Chat, code & reasoning From prices in euros: the lowest price the dynamic multiplier can reach for each model
ModelInput / 1M tokensOutput / 1M tokens
qwen3.6-35b-a3bReasoning with a long context · 256k contextAWQ-INT4 · 131 tok/s · 419 ms TTFTFrom €0.0808From €0.542
qwen3.8-27bDense reasoning model · 64k contextW4A16 (INT4) · 58 tok/s · 328 ms TTFTFrom €0.0816From €1.09
gemma-4-26b-a4bLong-context text generation · 256k contextAWQ-INT4 · 140 tok/s · 370 ms TTFTFrom €0.0398From €0.149
mistral-small-24bText generation · 32k contextAWQ-INT4 · 32 tok/s · 399 ms TTFTFrom €0.034From €0.0547
laguna-xs-2.1Coding and reasoning · 128k contextINT4 · 140 tok/s · 396 ms TTFTFrom €0.0409From €0.0827
gpt-oss-20bReasoning · 128k contextMXFP4 · 180 tok/s · 372 ms TTFTFrom €0.0183From €0.0697
hypernova-60bReasoning · 128k contextMXFP4 · 174 tok/s · 436 ms TTFTFrom €0.043From €0.215
nemotron3-nano-30bLong-context reasoning · 256k contextAWQ-INT4 · 215 tok/s · 390 ms TTFTFrom €0.034From €0.136

Vision & OCR

Vision & OCR From prices in euros: the lowest price the dynamic multiplier can reach for each model
ModelInput / 1M tokensOutput / 1M tokens
qwen3-vl-30b-a3bImage understanding · 32k contextAWQ-INT4 · 179 tok/s · 515 ms TTFTFrom €0.101From €0.402
nemotron3-omni-30bVision and audio understanding · 128k contextW4A16 (INT4) · 187 tok/s · 397 ms TTFTFrom €0.06From €0.18
deepseek-ocr-2Document OCR · 8k contextBF16 · 144 tok/s · 900 ms TTFTFrom €0.0232From €0.0232
Text embeddings From prices in euros: the lowest price the dynamic multiplier can reach for each model
ModelInput / 1M tokens
qwen3-embedding-8bText embeddings · 7k contextBF16From €0.00774

Speech & audio

Speech & audio From prices in euros: the lowest price the dynamic multiplier can reach for each model
ModelPer audio minute
voxtral-mini-3bTranscription and audio understanding · 8k contextBF16 · 25 tok/s (chat) · 855 ms TTFTFrom €0.000774
whisper-large-v3-turboSpeech to text · 30 s / segmentFP16 · 21.8× real-time (1 stream)From €0.000155
kokoro-82mText to speechFP32 · ~1× real-time (serial processing)From €0.00399

Image generation

Image generation From prices in euros: the lowest price the dynamic multiplier can reach for each model
ModelPer image
flux-2-kleinImage generationBF16 · 4.15 s / imageFrom €0.00132

These are catalogue rates for the listed models, not a tariff for every API route. For OCR and audio, check the route and billing unit before estimating your workload.

From prices in EUR, updated on 30 September 2026: the lowest price the dynamic multiplier can reach for each model. Confirm current rates and access in the console. Performance figures are indicative and shown per stream; TTFT means time to first token. Reasoning tokens count as output.

A rate you can
trace to the request.

For eligible models, the multiplier is set when you submit a request. Check recorded usage and charges in your console.

  1. 01

    Start from the From price.

    Each model has a From price for its billing unit.

  2. 02

    Follow the hour’s rate.

    Eligible models follow their GPU class’s hourly schedule, never below the From price.

  3. 03

    Check your usage.

    Review the request usage and charges recorded for your account.

A few useful
details.

Can dynamic prices go above list?

Yes. Prices follow demand and can rise above the list rate. Use the published GPU-class schedule when timing flexible work.

Does every model use dynamic pricing?

No. Dynamic pricing applies to eligible models. Other models use a fixed rate.

How is usage measured?

Start with the model and its API route. The catalogue uses tokens, audio minutes or generated images; dedicated parsing routes document per-image billing. Confirm the unit for your selected route in the console. Read the billing-unit guide.

What about Serverless pricing?

Serverless is opening gradually. GPU classes, scaling settings and billing terms are confirmed during onboarding. Join the waitlist with your workload; Open Models rates are not a Serverless GPU quote.

Explore Serverless

Contact the team

Loading the form…