Basics
Open Models is compatible with the OpenAI API. Every route below takes and returns JSON unless stated otherwise, and needs your API key in the Authorization header.
- Base URL
https://api.models.blade.sh/v1- Authentication
Authorization: Bearer $BLADE_API_KEY
Create a key in the console. An error always comes back as the same envelope, with a stable code to branch on:
{
"error": {
"message": "prompt_tokens + max_tokens (9000) exceeds the model's context length (8192).",
"type": "invalid_request_error",
"code": "context_length_exceeded"
}
}
Text
Chat completions
POST /v1/chat/completions
Send a conversation and get the model's next message, in one JSON document or streamed as it is written.
Compatible with the OpenAI Chat Completions API: the official SDKs work with the base URL changed.
A field the API does not list is passed to the model as is. GET /v1/models lists, per model, the parameters Blade supports, and a parameter the API had to drop is named in the x-shadow-ignored-params response header.
Parameters
| Name | Type | Description |
|---|---|---|
frequency_penalty | numberOptional | Lowers the chance of repeating tokens in proportion to how often they already appear. |
max_completion_tokens | integerOptional | OpenAI's newer name for max_tokens, with the same limits. If both are sent, this one wins. |
max_tokens | integerOptional | The most tokens to generate. Capped by the model's max_output_tokens, and the prompt plus the answer must fit in its context. Leave it out to let the model use the room it has left. |
messages | array of objectRequired | The conversation so far: a list of messages with a role (system, user, assistant or tool) and a content. Content can be a list of parts to add images. |
model | stringRequired | The model ID, from GET /v1/models. |
n | integerOptional | How many answers to generate. Each one counts as a request. More than one needs a temperature above zero. |
presence_penalty | numberOptional | Lowers the chance of repeating any token that already appears, to push towards new topics. |
seed | integerOptional | Makes sampling repeatable, as far as the model allows. |
stop | string or array of stringOptional | Up to four sequences where generation stops. A string or a list of strings. |
stream | booleanOptional | Send the answer as server-sent events while it is written. Default: false. |
stream_options | objectOptional | Streaming options. Set include_usage to receive the token counts at the end of the stream. |
temperature | numberOptional | Sampling temperature, from 0 to 2. Higher is more varied, lower is more focused. |
tools | array of objectOptional | Functions the model may call, in the OpenAI format, and Blade's built-in tools: {"type": "web_search"}, {"type": "web_fetch"} and {"type": "code_interpreter"}. Blade runs the built-in tools itself and returns the final answer. |
top_p | numberOptional | Nucleus sampling: only the most likely tokens whose probabilities add up to top_p are considered. |
web_search_options | objectOptional | Turns on web search: Blade runs the searches the model asks for and returns an answer with its sources. Searches are billed per search and leave Blade's infrastructure. |
Responses
| Status | Meaning |
|---|---|
200 | The completion. JSON, or a stream of server-sent events when stream is true. |
400 | Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed. |
401 | Missing, invalid or revoked API key. |
403 | Your account cannot use this: a tool is not enabled for it, or the model is not available to it. |
404 | Unknown model, or a job that does not exist or is not yours. |
413 | The request body is too large. |
429 | Rate limit reached: too many requests or tokens per minute, or the daily token quota is used up. Retry after the Retry-After header. |
502 | The model failed to answer. |
503 | The model is not available right now, for example while it starts. Retry later. |
504 | The model took too long to answer. |
curl -X POST https://api.models.blade.sh/v1/chat/completions \
-H "Authorization: Bearer $BLADE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{
"role": "user",
"content": "Hello!"
}
],
"model": "MODEL_ID"
}'
Text
Completions
POST /v1/completions
The legacy text completion API, for clients that send a raw prompt instead of messages.
Same limits and the same streaming as chat completions. Prefer chat completions for new code.
Parameters
| Name | Type | Description |
|---|---|---|
max_tokens | integerOptional | The most tokens to generate. Same limits as in chat completions. |
model | stringRequired | The model ID, from GET /v1/models. |
n | integerOptional | How many completions to generate. Each one counts as a request. |
prompt | string or array of stringRequired | The text to complete. A string, or a list of strings. |
seed | integerOptional | Makes sampling repeatable, as far as the model allows. |
stop | string or array of stringOptional | Sequences where generation stops. A string or a list of strings. |
stream | booleanOptional | Send the answer as server-sent events while it is written. Default: false. |
stream_options | objectOptional | Streaming options. Set include_usage to receive the token counts at the end of the stream. |
temperature | numberOptional | Sampling temperature, from 0 to 2. |
top_p | numberOptional | Nucleus sampling threshold. |
Responses
| Status | Meaning |
|---|---|
200 | The completion. JSON, or a stream of server-sent events when stream is true. |
400 | Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed. |
401 | Missing, invalid or revoked API key. |
403 | Your account cannot use this: a tool is not enabled for it, or the model is not available to it. |
404 | Unknown model, or a job that does not exist or is not yours. |
413 | The request body is too large. |
429 | Rate limit reached: too many requests or tokens per minute, or the daily token quota is used up. Retry after the Retry-After header. |
502 | The model failed to answer. |
503 | The model is not available right now, for example while it starts. Retry later. |
504 | The model took too long to answer. |
curl -X POST https://api.models.blade.sh/v1/completions \
-H "Authorization: Bearer $BLADE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "MODEL_ID",
"prompt": "Once upon a time"
}'
Text
Responses
POST /v1/responses
The OpenAI Responses API: item-based input and output, with Blade's built-in tools (web search, web fetch, code execution).
Stateless: nothing is stored between requests. store: true and previous_response_id are refused; send the earlier turns in input instead.
Only plain text output is available: a json_schema format in text is refused.
Parameters
| Name | Type | Description |
|---|---|---|
input | string or array of objectRequired | What the model answers: a string for a single user turn, or a list of items (message, function_call, function_call_output) to replay a conversation or return the result of your own function. |
instructions | stringOptional | A system message placed before the input. |
max_output_tokens | integerOptional | The most tokens to generate, with the same limits as max_tokens in chat completions. |
metadata | objectOptional | Your own key-value pairs, returned unchanged with the response. |
model | stringRequired | The model ID of a chat model, from GET /v1/models. |
parallel_tool_calls | booleanOptional | Lets the model call several tools in one turn. |
previous_response_id | stringOptional | Not supported. Send the earlier turns in input instead. |
reasoning | objectOptional | {"effort": "none"} turns the model's reasoning off; any other effort turns it on, on models that can reason. |
seed | integerOptional | Makes sampling repeatable, as far as the model allows. |
store | booleanOptional | Must be false or left out: responses are not stored. |
stream | booleanOptional | Send the response as Responses events while it is written. Default: false. |
temperature | numberOptional | Sampling temperature, from 0 to 2. |
text | objectOptional | Output format. Only {"format": {"type": "text"}} is accepted. |
tool_choice | string or objectOptional | Which tool the model must or may use. Cannot be combined with a built-in tool. |
tools | array of objectOptional | Functions the model may call, in the Responses format, and Blade's built-in tools. |
top_p | numberOptional | Nucleus sampling threshold. |
user | stringOptional | An identifier for your end user, for your own tracking. |
Responses
| Status | Meaning |
|---|---|
200 | The response. JSON, or a stream of Responses events when stream is true. |
400 | Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed. |
401 | Missing, invalid or revoked API key. |
403 | Your account cannot use this: a tool is not enabled for it, or the model is not available to it. |
404 | Unknown model, or a job that does not exist or is not yours. |
413 | The request body is too large. |
429 | Rate limit reached: too many requests or tokens per minute, or the daily token quota is used up. Retry after the Retry-After header. |
502 | The model failed to answer. |
503 | The model is not available right now, for example while it starts. Retry later. |
504 | The model took too long to answer. |
curl -X POST https://api.models.blade.sh/v1/responses \
-H "Authorization: Bearer $BLADE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": "Hello!",
"model": "MODEL_ID"
}'
Audio
Speech
POST /v1/audio/speech
Turn text into spoken audio.
Voices: tara, leah, jess, leo, dan, mia, zac and zoe. Emotion tags such as <laugh> or <sigh> can be written inline in the text.
Parameters
| Name | Type | Description |
|---|---|---|
input | stringRequired | The text to speak, up to 4,096 characters. |
model | stringRequired | The ID of a speech model, from GET /v1/models. |
voice | stringOptional | The voice. Default: tara. |
Responses
| Status | Meaning |
|---|---|
200 | WAV audio, mono, 24 kHz. |
400 | Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed. |
401 | Missing, invalid or revoked API key. |
404 | Unknown model, or a job that does not exist or is not yours. |
429 | Rate limit reached: too many requests or tokens per minute, or the daily token quota is used up. Retry after the Retry-After header. |
502 | The model failed to answer. |
503 | The model is not available right now, for example while it starts. Retry later. |
504 | The model took too long to answer. |
curl -X POST https://api.models.blade.sh/v1/audio/speech \
-H "Authorization: Bearer $BLADE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": "Hello!",
"model": "MODEL_ID"
}' \
--output speech.wav
Audio
Transcriptions
POST /v1/audio/transcriptions
Turn an audio file into text.
Send the file as multipart form data: WAV, MP3, M4A, FLAC or OGG, up to 25 MB.
Parameters
| Name | Type | Description |
|---|---|---|
file | fileRequired | The audio file: WAV, MP3, M4A, FLAC or OGG, up to 25 MB. |
model | stringRequired | The ID of a transcription model, from GET /v1/models. |
Responses
| Status | Meaning |
|---|---|
200 | The transcribed text. |
400 | Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed. |
401 | Missing, invalid or revoked API key. |
404 | Unknown model, or a job that does not exist or is not yours. |
413 | The request body is too large. |
429 | Rate limit reached: too many requests or tokens per minute, or the daily token quota is used up. Retry after the Retry-After header. |
502 | The model failed to answer. |
503 | The model is not available right now, for example while it starts. Retry later. |
504 | The model took too long to answer. |
curl -X POST https://api.models.blade.sh/v1/audio/transcriptions \
-H "Authorization: Bearer $BLADE_API_KEY" \
-F file=@recording.mp3 \
-F model=MODEL_ID
Video
Create a video
POST /v1/video/generations
Start generating a video from a text prompt. Generation takes one to three minutes, so the request returns a job right away.
Poll the job with GET /v1/video/generations/{job_id} until it completes. You can have two unfinished jobs at a time; a third returns 429 video_jobs_quota_exceeded.
Parameters
| Name | Type | Description |
|---|---|---|
fps | integerOptional | Frames per second, from 8 to 30. |
height | integerOptional | Height in pixels, up to 480. |
model | stringRequired | The ID of a video model, from GET /v1/models. |
negative_prompt | stringOptional | What the video should avoid. |
num_frames | integerOptional | Number of frames, up to 161. Rounded to a multiple of 8 plus 1. |
num_inference_steps | integerOptional | Generation steps, up to 50. More steps take longer and usually look better. |
prompt | stringRequired | What the video shows, up to 2,000 characters. |
seed | integerOptional | Makes the generation repeatable. |
width | integerOptional | Width in pixels, up to 704. |
Responses
| Status | Meaning |
|---|---|
202 | The job, created. Keep its id to follow it. |
400 | Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed. |
401 | Missing, invalid or revoked API key. |
404 | Unknown model, or a job that does not exist or is not yours. |
429 | Rate limit reached: too many requests or tokens per minute, or the daily token quota is used up. Retry after the Retry-After header. |
curl -X POST https://api.models.blade.sh/v1/video/generations \
-H "Authorization: Bearer $BLADE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "MODEL_ID",
"prompt": "A red kite over the sea at dawn"
}'
Video
List videos
GET /v1/video/generations
Your 20 most recent video jobs.
Responses
| Status | Meaning |
|---|---|
200 | The jobs, most recent first. |
curl https://api.models.blade.sh/v1/video/generations \
-H "Authorization: Bearer $BLADE_API_KEY"
Video
Get a video
GET /v1/video/generations/{job_id}
The state of a video job, and the link to the video once it is ready.
The download link expires: read the job again to get a fresh one. A job that does not exist or is not yours returns 404.
Parameters
| Name | Type | Description |
|---|---|---|
job_id | stringRequired | The id returned when the job was created. |
Responses
| Status | Meaning |
|---|---|
200 | The job. |
401 | Missing, invalid or revoked API key. |
404 | Unknown model, or a job that does not exist or is not yours. |
curl https://api.models.blade.sh/v1/video/generations/JOB_ID \
-H "Authorization: Bearer $BLADE_API_KEY"
Video
Cancel a video
POST /v1/video/generations/{job_id}/cancel
Stop a video job that has not finished.
Cancellation is best effort: a job that is about to finish may complete anyway.
Parameters
| Name | Type | Description |
|---|---|---|
job_id | stringRequired | The id returned when the job was created. |
Responses
| Status | Meaning |
|---|---|
200 | The job, cancelled if the cancellation went through. |
401 | Missing, invalid or revoked API key. |
404 | Unknown model, or a job that does not exist or is not yours. |
curl -X POST https://api.models.blade.sh/v1/video/generations/JOB_ID/cancel \
-H "Authorization: Bearer $BLADE_API_KEY"
Documents
Detect
POST /v1/detect
Find regions in an image, such as the tables on a document page.
One JSON document, no streaming. Name the model explicitly; there is no default. Billed per processed image.
Parameters
| Name | Type | Description |
|---|---|---|
image | stringRequired | The image, in base64 or as a data:image/...;base64,... URI. |
model | stringRequired | The ID of a detection model, from GET /v1/models. |
Responses
| Status | Meaning |
|---|---|
200 | The regions found. |
400 | Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed. |
401 | Missing, invalid or revoked API key. |
404 | Unknown model, or a job that does not exist or is not yours. |
502 | The model failed to answer. |
503 | The model is not available right now, for example while it starts. Retry later. |
504 | The model took too long to answer. |
curl -X POST https://api.models.blade.sh/v1/detect \
-H "Authorization: Bearer $BLADE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"image": "BASE64_IMAGE",
"model": "MODEL_ID"
}'
Documents
Parse
POST /v1/parse
Read the content of an image as structured text: handwriting, general OCR, a chart as a data table, a formula as LaTeX, or the structure of a table.
One JSON document, no streaming. The shape of the result depends on the model. Billed per processed image.
The table structure model needs the boxes of the tables in table_bboxes: get them from POST /v1/detect, or pass the whole page.
Parameters
| Name | Type | Description |
|---|---|---|
image | stringRequired | The image, in base64 or as a data:image/...;base64,... URI. |
max_new_tokens | integerOptional | The most tokens to generate, for the models that write text. Capped by the model's own limit. |
model | stringRequired | The ID of a parsing model, from GET /v1/models. |
table_bboxes | array of array of numberOptional | Required by the table structure model: the pixel boxes of the tables to read. Ignored by the other models. |
Responses
| Status | Meaning |
|---|---|
200 | The parsed content. Its shape depends on the model. |
400 | Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed. |
401 | Missing, invalid or revoked API key. |
404 | Unknown model, or a job that does not exist or is not yours. |
502 | The model failed to answer. |
503 | The model is not available right now, for example while it starts. Retry later. |
504 | The model took too long to answer. |
curl -X POST https://api.models.blade.sh/v1/parse \
-H "Authorization: Bearer $BLADE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"image": "BASE64_IMAGE",
"model": "MODEL_ID"
}'
Catalogue
List models
GET /v1/models
The models you can call, with their limits, capabilities, supported parameters and prices.
Use the id of a model as the model of your requests.
Responses
| Status | Meaning |
|---|---|
200 | The model list, in the OpenAI format. |
401 | Missing, invalid or revoked API key. |
curl https://api.models.blade.sh/v1/models \
-H "Authorization: Bearer $BLADE_API_KEY"
For agents and tools
The same reference in machine-readable formats, for code generators and AI agents.
- OpenAPI contract (YAML)
- API reference (Markdown)
- Documentation index (llms.txt)
Generated from the API contract of . Model availability and enabled features depend on your account: check them in the console before integrating.