blade Get your API key

Developer docs / Open Models / Getting started

Your SDK.
Your first model call.

Call a hosted model with Python, JavaScript or curl. Set up your key, make a request, then stream the answer into your app.

1. Get your key and model ID

  1. Create an account or sign in to the console.
  2. In Get started, create an API key with inference access. Save it securely: the full key is shown only once. Manage existing keys from your account’s Keys page.
  3. Open Models, choose a chat-capable model. Copy its exact model ID from View code.
Base URL
https://api.models.blade.sh/v1
Authentication
Authorization: Bearer YOUR_BLADE_API_KEY
First endpoint
POST /v1/chat/completions

Use your Blade key with this URL; an OpenAI key will not authenticate here.

2. Set up your project

Python

Use Python 3.10 or newer with pip. In a new project directory, create a virtual environment and install the OpenAI SDK. These commands use a macOS or Linux shell.

Terminal · Python setup
python3 -m venv .venv
. .venv/bin/activate
python3 -m pip install openai

Windows / PowerShell setup
PowerShell · Python setup
py -m venv .venv
.\.venv\Scripts\python.exe -m pip install openai

Use .\.venv\Scripts\python.exe hello.py to run the example without activating the environment.

JavaScript / Node.js

Use Node.js 22 or a newer supported LTS release with npm. Run this in a new project directory. In an existing npm project, only the install command is needed.

Terminal · Node.js setup
npm init -y
npm install openai

The examples run on your server with Node.js. The .mjs extension enables JavaScript module imports.

curl

Use a recent version of curl that supports --fail-with-body. The commands below use a macOS or Linux shell; no SDK is required.

Set the following environment variables in the terminal you will use to run the example. Replace the placeholders with your key and model ID.

Terminal · Environment
export BLADE_API_KEY="YOUR_BLADE_API_KEY"
export BLADE_MODEL="YOUR_AVAILABLE_MODEL_ID"

Windows / PowerShell environment variables
PowerShell · Environment
$env:BLADE_API_KEY="YOUR_BLADE_API_KEY"
$env:BLADE_MODEL="YOUR_AVAILABLE_MODEL_ID"

BLADE_API_KEY and BLADE_MODEL are local names used by this guide. The examples read them explicitly; creating a .env file alone will not load them.

Terminal · Environment
export BLADE_API_KEY="YOUR_BLADE_API_KEY"

You will put the model ID in request.json in the next step.

3. Send your first request

Python

Save this file as hello.py in your project directory.

hello.py
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.models.blade.sh/v1",
    api_key=os.environ["BLADE_API_KEY"],
)

response = client.chat.completions.create(
    model=os.environ["BLADE_MODEL"],
    messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(response.choices[0].message.content)

Terminal · Run the request
python3 hello.py

JavaScript / Node.js

Save this file as hello.mjs in your project directory.

hello.mjs
import OpenAI from "openai";

const { BLADE_API_KEY, BLADE_MODEL } = process.env;
if (!BLADE_API_KEY || !BLADE_MODEL) {
  throw new Error("Set BLADE_API_KEY and BLADE_MODEL first.");
}

const client = new OpenAI({
  baseURL: "https://api.models.blade.sh/v1",
  apiKey: BLADE_API_KEY,
});

const response = await client.chat.completions.create({
  model: BLADE_MODEL,
  messages: [{ role: "user", content: "Say hello in one sentence." }],
});
console.log(response.choices[0].message.content);

Terminal · Run the request
node hello.mjs

curl

Save this JSON as request.json. Replace YOUR_AVAILABLE_MODEL_ID with the model ID you copied from the console.

request.json
{
  "model": "YOUR_AVAILABLE_MODEL_ID",
  "messages": [
    { "role": "user", "content": "Say hello in one sentence." }
  ]
}

Run this command from the same directory as request.json.

Terminal · Send the request
curl --fail-with-body --silent --show-error \
  https://api.models.blade.sh/v1/chat/completions \
  -H "Authorization: Bearer $BLADE_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @request.json

Read the result

The SDK examples print the model’s answer. curl returns the full JSON response; the answer is in choices[0].message.content. The wording varies between runs.

choices[0].message.content
The generated answer.
choices[0].finish_reason
Why generation ended. A value of length means the output reached a limit; review the token budget.
usage
Token counts returned with the response. These are not a currency amount; see how inference pricing works.

If the request fails, use the status code and error message to identify the cause.

4. Stream the answer

Set stream to true to receive the output as it is generated. Check that the selected model supports streaming.

Python

Save this complete example as stream.py and run python3 stream.py in the same environment. On Windows, use .\.venv\Scripts\python.exe stream.py.

stream.py
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.models.blade.sh/v1",
    api_key=os.environ["BLADE_API_KEY"],
)

stream = client.chat.completions.create(
    model=os.environ["BLADE_MODEL"],
    messages=[{"role": "user", "content": "Say hello in one sentence."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="", flush=True)
print()

JavaScript / Node.js

Save this complete example as stream.mjs and run node stream.mjs in the same terminal.

stream.mjs
import OpenAI from "openai";

const { BLADE_API_KEY, BLADE_MODEL } = process.env;
if (!BLADE_API_KEY || !BLADE_MODEL) {
  throw new Error("Set BLADE_API_KEY and BLADE_MODEL first.");
}

const client = new OpenAI({
  baseURL: "https://api.models.blade.sh/v1",
  apiKey: BLADE_API_KEY,
});

const stream = await client.chat.completions.create({
  model: BLADE_MODEL,
  messages: [{ role: "user", content: "Say hello in one sentence." }],
  stream: true,
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta.content ?? "");
}
process.stdout.write("\n");

curl

Save this as stream.json and replace the model placeholder. --no-buffer lets curl display incoming events immediately.

stream.json
{
  "model": "YOUR_AVAILABLE_MODEL_ID",
  "stream": true,
  "messages": [
    { "role": "user", "content": "Say hello in one sentence." }
  ]
}

Terminal · Stream the request
curl --no-buffer --fail-with-body --silent --show-error \
  https://api.models.blade.sh/v1/chat/completions \
  -H "Authorization: Bearer $BLADE_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @stream.json

curl displays raw server-sent events: data: lines containing JSON chunks, followed by data: [DONE]. Your app should parse these events instead of displaying the raw stream.

Read text from choices[0].delta.content. Some chunks carry no text or have an empty choices array; the SDK examples handle both. For streaming token totals, set stream_options.include_usage to true and read usage from the final usage chunk.

Check support before adding features

OpenAI-compatible means you can reuse the client for supported routes. It does not mean every OpenAI feature or parameter is available.

FeatureWhat to check
Chat CompletionsUse /v1/chat/completions. Check each model’s capabilities, context window and supported_parameters in the model reference or GET /v1/models.
Conversation historySend the relevant earlier messages with each request. The client does not preserve your conversation automatically.
ResponsesThe current /v1/responses API is stateless. store: true and previous_response_id are not supported. Replay history in the request instead.
Function callingChoose a model with native tool calling. Your application executes client-defined functions and returns their results.
Staan web searchEnable web search for the account and choose a model with native tool calling. Search is opt-in per request and billed per executed search. Search queries go to Staan; see data handling and the API reference before enabling it.
Other server toolsTool access and model support are required. Check the console’s Tools page for activation and billing details.
List models from your terminal

With a key that includes models:read, this read-only request returns a JSON list in data. Inspect the model’s id and supported parameters before using it.

Terminal · List models
curl --fail-with-body --silent --show-error \
  https://api.models.blade.sh/v1/models \
  -H "Authorization: Bearer $BLADE_API_KEY"

When a request fails

Read the error body as well as the HTTP status. Retrying an invalid key, unsupported parameter or exhausted daily quota will not fix the request.

StatusWhat to do
400 · RequestCheck JSON syntax, accepted parameters, the model’s output-token cap and context limit. Remove or correct the field named in the error.
401 · KeyConfirm the environment variable is set in the same terminal. Check for a missing, malformed, expired or revoked key and use a Blade key with the Blade base URL.
403 · AccessCheck key permissions and the account’s model, external-provider or tool access. Ask an account administrator to enable the required access.
404 · Model / routeCopy the exact model ID from the console and check the route. Set the SDK base URL to https://api.models.blade.sh/v1; do not append /chat/completions to the SDK base URL.
413 · Request sizeReduce the request body to fit the endpoint’s documented size limit.
429 · LimitsFor rate_limit_exceeded, reduce concurrency and respect Retry-After. For insufficient_quota, check the daily token allowance; the documented reset is midnight UTC.
502–504 · UpstreamUse a limited number of retries with a delay for transient failures; investigate repeated failures rather than retrying indefinitely.
The example fails before making a request
  • Python cannot import openai: use the same interpreter for installation and execution. Activate .venv, or run its Python executable directly.
  • Node.js cannot find openai: install it in the project directory and keep the example’s .mjs extension.
  • A variable is missing: set it in the current terminal, then run the example there. A new terminal will not inherit variables exported in another session.
  • curl cannot read the file: run the command in the directory containing request.json or stream.json.

Check usage. Then build on it.

Open Usage in the console and select the relevant date range. Review requests and token counts by model and key, and export CSV when needed. Displayed costs are estimates; the invoice is authoritative.

For audio, image generation, OCR or embeddings, start from the corresponding console studio and the model’s API example. These workloads have their own inputs and routes; a text-chat request is not a universal template.

Checked on against the console and Blade’s quickstart and API reference. SDK syntax: official Python and JavaScript references. Model availability and enabled features depend on your account.

Contact the team

Loading the form…