blade Get your API key

Open Models Available now

Open models.
One familiar API.

Run open-weight LLMs through an OpenAI-compatible inference API. Keep your OpenAI SDK and use Blade’s European GPU infrastructure.

Start Building

Test a prompt.
Take the code.

Try prompts in LLM Studio, adjust generation settings, then open “View code” for a Python, JavaScript or cURL example. Explore dedicated studios for audio, images, OCR and embeddings.

For current sources and calculations, add Flash by Staan web search or the code interpreter. These tools require account activation and a model with tool calling.

Open the studio

“When should I use the median?”

Flash by Staan Web search
Python documentation

The median is the middle value. It is less sensitive to extreme values than the mean.

10, 12, 14, 16, 98
Mean 30Median 14
Illustrative answer · example data

Explore
open-weight models.

Find a model for your workload. Prices are shown in EUR as From prices: the lowest each model’s dynamic multiplier can reach.

16 models · From prices in EUR, updated 30 September 2026: the lowest price the dynamic multiplier can reach for each model. Prices move within a per-model range. Check the console for current dynamic rates. How dynamic pricing works →

Use the OpenAI SDK
with Blade.

Start in Get started: create your key, copy an example and make your first call. Keep the OpenAI client you already use.

  1. Create a key with inference access. Save it when it is shown.
  2. Choose an available chat model and copy its model ID.
  3. Set your key and the Blade base URL, then send a request.
  4. Track requests, tokens and estimated cost in Usage.
Read the API docs
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.models.blade.sh/v1",
    api_key=os.environ["BLADE_API_KEY"],
)

response = client.chat.completions.create(
    model=os.environ["BLADE_MODEL"],
    messages=[{"role": "user",
               "content": "Hello, Blade."}],
)
print(response.choices[0].message.content)

Know what’s supported.

OpenAI compatibility on Open Models
FeatureSupport on Blade
AuthenticationYour Blade key, the Blade base URL and a model ID available to your account.
Chat Completions/v1/chat/completions, with streaming and tool calling on supported models.
Other workloadsAudio, images and embeddings use their own API routes. Document parsing has dedicated endpoints. Start with the relevant studio’s code example.
ResponsesStateless requests. store: true and previous_response_id are not supported.
Server toolsWeb search, web fetch and code execution require account access and a model with native tool calling.

Read the full compatibility reference →

Contact the team

Loading the form…