1. Get your key and model ID
- Create an account or sign in to the console.
- In Get started, create an API key with inference access. Save it securely: the full key is shown only once. Manage existing keys from your account’s Keys page.
- Open Models, choose a chat-capable model. Copy its exact model ID from View code.
- Base URL
https://api.models.blade.sh/v1- Authentication
Authorization: Bearer YOUR_BLADE_API_KEY- First endpoint
POST /v1/chat/completions
Use your Blade key with this URL; an OpenAI key will not authenticate here.
2. Set up your project
Python
Use Python 3.10 or newer with pip. In a new project directory, create a virtual environment and install the OpenAI SDK. These commands use a macOS or Linux shell.
python3 -m venv .venv
. .venv/bin/activate
python3 -m pip install openai
Windows / PowerShell setup
py -m venv .venv
.\.venv\Scripts\python.exe -m pip install openai
Use .\.venv\Scripts\python.exe hello.py to run the example without activating the environment.
JavaScript / Node.js
Use Node.js 22 or a newer supported LTS release with npm. Run this in a new project directory. In an existing npm project, only the install command is needed.
npm init -y
npm install openai
The examples run on your server with Node.js. The .mjs extension enables JavaScript module imports.
curl
Use a recent version of curl that supports --fail-with-body. The commands below use a macOS or Linux shell; no SDK is required.
Set the following environment variables in the terminal you will use to run the example. Replace the placeholders with your key and model ID.
export BLADE_API_KEY="YOUR_BLADE_API_KEY"
export BLADE_MODEL="YOUR_AVAILABLE_MODEL_ID"
Windows / PowerShell environment variables
$env:BLADE_API_KEY="YOUR_BLADE_API_KEY"
$env:BLADE_MODEL="YOUR_AVAILABLE_MODEL_ID"
BLADE_API_KEY and BLADE_MODEL are local names used by this guide. The examples read them explicitly; creating a .env file alone will not load them.
export BLADE_API_KEY="YOUR_BLADE_API_KEY"
You will put the model ID in request.json in the next step.
3. Send your first request
Python
Save this file as hello.py in your project directory.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.models.blade.sh/v1",
api_key=os.environ["BLADE_API_KEY"],
)
response = client.chat.completions.create(
model=os.environ["BLADE_MODEL"],
messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(response.choices[0].message.content)
python3 hello.py
JavaScript / Node.js
Save this file as hello.mjs in your project directory.
import OpenAI from "openai";
const { BLADE_API_KEY, BLADE_MODEL } = process.env;
if (!BLADE_API_KEY || !BLADE_MODEL) {
throw new Error("Set BLADE_API_KEY and BLADE_MODEL first.");
}
const client = new OpenAI({
baseURL: "https://api.models.blade.sh/v1",
apiKey: BLADE_API_KEY,
});
const response = await client.chat.completions.create({
model: BLADE_MODEL,
messages: [{ role: "user", content: "Say hello in one sentence." }],
});
console.log(response.choices[0].message.content);
node hello.mjs
curl
Save this JSON as request.json. Replace YOUR_AVAILABLE_MODEL_ID with the model ID you copied from the console.
{
"model": "YOUR_AVAILABLE_MODEL_ID",
"messages": [
{ "role": "user", "content": "Say hello in one sentence." }
]
}
Run this command from the same directory as request.json.
curl --fail-with-body --silent --show-error \
https://api.models.blade.sh/v1/chat/completions \
-H "Authorization: Bearer $BLADE_API_KEY" \
-H "Content-Type: application/json" \
--data-binary @request.json
Read the result
The SDK examples print the model’s answer. curl returns the full JSON response; the answer is in choices[0].message.content. The wording varies between runs.
choices[0].message.content- The generated answer.
choices[0].finish_reason- Why generation ended. A value of
lengthmeans the output reached a limit; review the token budget. usage- Token counts returned with the response. These are not a currency amount; see how inference pricing works.
If the request fails, use the status code and error message to identify the cause.
4. Stream the answer
Set stream to true to receive the output as it is generated. Check that the selected model supports streaming.
Python
Save this complete example as stream.py and run python3 stream.py in the same environment. On Windows, use .\.venv\Scripts\python.exe stream.py.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.models.blade.sh/v1",
api_key=os.environ["BLADE_API_KEY"],
)
stream = client.chat.completions.create(
model=os.environ["BLADE_MODEL"],
messages=[{"role": "user", "content": "Say hello in one sentence."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="", flush=True)
print()
JavaScript / Node.js
Save this complete example as stream.mjs and run node stream.mjs in the same terminal.
import OpenAI from "openai";
const { BLADE_API_KEY, BLADE_MODEL } = process.env;
if (!BLADE_API_KEY || !BLADE_MODEL) {
throw new Error("Set BLADE_API_KEY and BLADE_MODEL first.");
}
const client = new OpenAI({
baseURL: "https://api.models.blade.sh/v1",
apiKey: BLADE_API_KEY,
});
const stream = await client.chat.completions.create({
model: BLADE_MODEL,
messages: [{ role: "user", content: "Say hello in one sentence." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta.content ?? "");
}
process.stdout.write("\n");
curl
Save this as stream.json and replace the model placeholder. --no-buffer lets curl display incoming events immediately.
{
"model": "YOUR_AVAILABLE_MODEL_ID",
"stream": true,
"messages": [
{ "role": "user", "content": "Say hello in one sentence." }
]
}
curl --no-buffer --fail-with-body --silent --show-error \
https://api.models.blade.sh/v1/chat/completions \
-H "Authorization: Bearer $BLADE_API_KEY" \
-H "Content-Type: application/json" \
--data-binary @stream.json
curl displays raw server-sent events: data: lines containing JSON chunks, followed by data: [DONE]. Your app should parse these events instead of displaying the raw stream.
Read text from choices[0].delta.content. Some chunks carry no text or have an empty choices array; the SDK examples handle both. For streaming token totals, set stream_options.include_usage to true and read usage from the final usage chunk.
Check support before adding features
OpenAI-compatible means you can reuse the client for supported routes. It does not mean every OpenAI feature or parameter is available.
| Feature | What to check |
|---|---|
| Chat Completions | Use /v1/chat/completions. Check each model’s capabilities, context window and supported_parameters in the model reference or GET /v1/models. |
| Conversation history | Send the relevant earlier messages with each request. The client does not preserve your conversation automatically. |
| Responses | The current /v1/responses API is stateless. store: true and previous_response_id are not supported. Replay history in the request instead. |
| Function calling | Choose a model with native tool calling. Your application executes client-defined functions and returns their results. |
| Staan web search | Enable web search for the account and choose a model with native tool calling. Search is opt-in per request and billed per executed search. Search queries go to Staan; see data handling and the API reference before enabling it. |
| Other server tools | Tool access and model support are required. Check the console’s Tools page for activation and billing details. |
List models from your terminal
With a key that includes models:read, this read-only request returns a JSON list in data. Inspect the model’s id and supported parameters before using it.
curl --fail-with-body --silent --show-error \
https://api.models.blade.sh/v1/models \
-H "Authorization: Bearer $BLADE_API_KEY"
When a request fails
Read the error body as well as the HTTP status. Retrying an invalid key, unsupported parameter or exhausted daily quota will not fix the request.
| Status | What to do |
|---|---|
400 · Request | Check JSON syntax, accepted parameters, the model’s output-token cap and context limit. Remove or correct the field named in the error. |
401 · Key | Confirm the environment variable is set in the same terminal. Check for a missing, malformed, expired or revoked key and use a Blade key with the Blade base URL. |
403 · Access | Check key permissions and the account’s model, external-provider or tool access. Ask an account administrator to enable the required access. |
404 · Model / route | Copy the exact model ID from the console and check the route. Set the SDK base URL to https://api.models.blade.sh/v1; do not append /chat/completions to the SDK base URL. |
413 · Request size | Reduce the request body to fit the endpoint’s documented size limit. |
429 · Limits | For rate_limit_exceeded, reduce concurrency and respect Retry-After. For insufficient_quota, check the daily token allowance; the documented reset is midnight UTC. |
502–504 · Upstream | Use a limited number of retries with a delay for transient failures; investigate repeated failures rather than retrying indefinitely. |
The example fails before making a request
- Python cannot import
openai: use the same interpreter for installation and execution. Activate.venv, or run its Python executable directly. - Node.js cannot find
openai: install it in the project directory and keep the example’s.mjsextension. - A variable is missing: set it in the current terminal, then run the example there. A new terminal will not inherit variables exported in another session.
- curl cannot read the file: run the command in the directory containing
request.jsonorstream.json.
Check usage. Then build on it.
Open Usage in the console and select the relevant date range. Review requests and token counts by model and key, and export CSV when needed. Displayed costs are estimates; the invoice is authoritative.
For audio, image generation, OCR or embeddings, start from the corresponding console studio and the model’s API example. These workloads have their own inputs and routes; a text-chat request is not a universal template.
Checked on against the console and Blade’s quickstart and API reference. SDK syntax: official Python and JavaScript references. Model availability and enabled features depend on your account.