blade Get your API key

Developer docs / Open Models / API reference

Every route.
Every parameter.

The Open Models API, route by route: what to send, what comes back, and the errors to handle.

Basics

Open Models is compatible with the OpenAI API. Every route below takes and returns JSON unless stated otherwise, and needs your API key in the Authorization header.

Base URL
https://api.models.blade.sh/v1
Authentication
Authorization: Bearer $BLADE_API_KEY

Create a key in the console. An error always comes back as the same envelope, with a stable code to branch on:

JSON · Error
{
  "error": {
    "message": "prompt_tokens + max_tokens (9000) exceeds the model's context length (8192).",
    "type": "invalid_request_error",
    "code": "context_length_exceeded"
  }
}

Text

Chat completions

POST /v1/chat/completions

Send a conversation and get the model's next message, in one JSON document or streamed as it is written.

Compatible with the OpenAI Chat Completions API: the official SDKs work with the base URL changed.

A field the API does not list is passed to the model as is. GET /v1/models lists, per model, the parameters Blade supports, and a parameter the API had to drop is named in the x-shadow-ignored-params response header.

Parameters

NameTypeDescription
frequency_penaltynumber
Optional
Lowers the chance of repeating tokens in proportion to how often they already appear.
max_completion_tokensinteger
Optional
OpenAI's newer name for max_tokens, with the same limits. If both are sent, this one wins.
max_tokensinteger
Optional
The most tokens to generate. Capped by the model's max_output_tokens, and the prompt plus the answer must fit in its context. Leave it out to let the model use the room it has left.
messagesarray of object
Required
The conversation so far: a list of messages with a role (system, user, assistant or tool) and a content. Content can be a list of parts to add images.
modelstring
Required
The model ID, from GET /v1/models.
ninteger
Optional
How many answers to generate. Each one counts as a request. More than one needs a temperature above zero.
presence_penaltynumber
Optional
Lowers the chance of repeating any token that already appears, to push towards new topics.
seedinteger
Optional
Makes sampling repeatable, as far as the model allows.
stopstring or array of string
Optional
Up to four sequences where generation stops. A string or a list of strings.
streamboolean
Optional
Send the answer as server-sent events while it is written. Default: false.
stream_optionsobject
Optional
Streaming options. Set include_usage to receive the token counts at the end of the stream.
temperaturenumber
Optional
Sampling temperature, from 0 to 2. Higher is more varied, lower is more focused.
toolsarray of object
Optional
Functions the model may call, in the OpenAI format, and Blade's built-in tools: {"type": "web_search"}, {"type": "web_fetch"} and {"type": "code_interpreter"}. Blade runs the built-in tools itself and returns the final answer.
top_pnumber
Optional
Nucleus sampling: only the most likely tokens whose probabilities add up to top_p are considered.
web_search_optionsobject
Optional
Turns on web search: Blade runs the searches the model asks for and returns an answer with its sources. Searches are billed per search and leave Blade's infrastructure.

Responses

StatusMeaning
200The completion. JSON, or a stream of server-sent events when stream is true.
400Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed.
401Missing, invalid or revoked API key.
403Your account cannot use this: a tool is not enabled for it, or the model is not available to it.
404Unknown model, or a job that does not exist or is not yours.
413The request body is too large.
429Rate limit reached: too many requests or tokens per minute, or the daily token quota is used up. Retry after the Retry-After header.
502The model failed to answer.
503The model is not available right now, for example while it starts. Retry later.
504The model took too long to answer.
curl · Chat completions
curl -X POST https://api.models.blade.sh/v1/chat/completions \
  -H "Authorization: Bearer $BLADE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ],
    "model": "MODEL_ID"
  }'

Text

Completions

POST /v1/completions

The legacy text completion API, for clients that send a raw prompt instead of messages.

Same limits and the same streaming as chat completions. Prefer chat completions for new code.

Parameters

NameTypeDescription
max_tokensinteger
Optional
The most tokens to generate. Same limits as in chat completions.
modelstring
Required
The model ID, from GET /v1/models.
ninteger
Optional
How many completions to generate. Each one counts as a request.
promptstring or array of string
Required
The text to complete. A string, or a list of strings.
seedinteger
Optional
Makes sampling repeatable, as far as the model allows.
stopstring or array of string
Optional
Sequences where generation stops. A string or a list of strings.
streamboolean
Optional
Send the answer as server-sent events while it is written. Default: false.
stream_optionsobject
Optional
Streaming options. Set include_usage to receive the token counts at the end of the stream.
temperaturenumber
Optional
Sampling temperature, from 0 to 2.
top_pnumber
Optional
Nucleus sampling threshold.

Responses

StatusMeaning
200The completion. JSON, or a stream of server-sent events when stream is true.
400Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed.
401Missing, invalid or revoked API key.
403Your account cannot use this: a tool is not enabled for it, or the model is not available to it.
404Unknown model, or a job that does not exist or is not yours.
413The request body is too large.
429Rate limit reached: too many requests or tokens per minute, or the daily token quota is used up. Retry after the Retry-After header.
502The model failed to answer.
503The model is not available right now, for example while it starts. Retry later.
504The model took too long to answer.
curl · Completions
curl -X POST https://api.models.blade.sh/v1/completions \
  -H "Authorization: Bearer $BLADE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MODEL_ID",
    "prompt": "Once upon a time"
  }'

Text

Responses

POST /v1/responses

The OpenAI Responses API: item-based input and output, with Blade's built-in tools (web search, web fetch, code execution).

Stateless: nothing is stored between requests. store: true and previous_response_id are refused; send the earlier turns in input instead.

Only plain text output is available: a json_schema format in text is refused.

Parameters

NameTypeDescription
inputstring or array of object
Required
What the model answers: a string for a single user turn, or a list of items (message, function_call, function_call_output) to replay a conversation or return the result of your own function.
instructionsstring
Optional
A system message placed before the input.
max_output_tokensinteger
Optional
The most tokens to generate, with the same limits as max_tokens in chat completions.
metadataobject
Optional
Your own key-value pairs, returned unchanged with the response.
modelstring
Required
The model ID of a chat model, from GET /v1/models.
parallel_tool_callsboolean
Optional
Lets the model call several tools in one turn.
previous_response_idstring
Optional
Not supported. Send the earlier turns in input instead.
reasoningobject
Optional
{"effort": "none"} turns the model's reasoning off; any other effort turns it on, on models that can reason.
seedinteger
Optional
Makes sampling repeatable, as far as the model allows.
storeboolean
Optional
Must be false or left out: responses are not stored.
streamboolean
Optional
Send the response as Responses events while it is written. Default: false.
temperaturenumber
Optional
Sampling temperature, from 0 to 2.
textobject
Optional
Output format. Only {"format": {"type": "text"}} is accepted.
tool_choicestring or object
Optional
Which tool the model must or may use. Cannot be combined with a built-in tool.
toolsarray of object
Optional
Functions the model may call, in the Responses format, and Blade's built-in tools.
top_pnumber
Optional
Nucleus sampling threshold.
userstring
Optional
An identifier for your end user, for your own tracking.

Responses

StatusMeaning
200The response. JSON, or a stream of Responses events when stream is true.
400Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed.
401Missing, invalid or revoked API key.
403Your account cannot use this: a tool is not enabled for it, or the model is not available to it.
404Unknown model, or a job that does not exist or is not yours.
413The request body is too large.
429Rate limit reached: too many requests or tokens per minute, or the daily token quota is used up. Retry after the Retry-After header.
502The model failed to answer.
503The model is not available right now, for example while it starts. Retry later.
504The model took too long to answer.
curl · Responses
curl -X POST https://api.models.blade.sh/v1/responses \
  -H "Authorization: Bearer $BLADE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "input": "Hello!",
    "model": "MODEL_ID"
  }'

Audio

Speech

POST /v1/audio/speech

Turn text into spoken audio.

Voices: tara, leah, jess, leo, dan, mia, zac and zoe. Emotion tags such as <laugh> or <sigh> can be written inline in the text.

Parameters

NameTypeDescription
inputstring
Required
The text to speak, up to 4,096 characters.
modelstring
Required
The ID of a speech model, from GET /v1/models.
voicestring
Optional
The voice. Default: tara.

Responses

StatusMeaning
200WAV audio, mono, 24 kHz.
400Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed.
401Missing, invalid or revoked API key.
404Unknown model, or a job that does not exist or is not yours.
429Rate limit reached: too many requests or tokens per minute, or the daily token quota is used up. Retry after the Retry-After header.
502The model failed to answer.
503The model is not available right now, for example while it starts. Retry later.
504The model took too long to answer.
curl · Speech
curl -X POST https://api.models.blade.sh/v1/audio/speech \
  -H "Authorization: Bearer $BLADE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "input": "Hello!",
    "model": "MODEL_ID"
  }' \
  --output speech.wav

Audio

Transcriptions

POST /v1/audio/transcriptions

Turn an audio file into text.

Send the file as multipart form data: WAV, MP3, M4A, FLAC or OGG, up to 25 MB.

Parameters

NameTypeDescription
filefile
Required
The audio file: WAV, MP3, M4A, FLAC or OGG, up to 25 MB.
modelstring
Required
The ID of a transcription model, from GET /v1/models.

Responses

StatusMeaning
200The transcribed text.
400Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed.
401Missing, invalid or revoked API key.
404Unknown model, or a job that does not exist or is not yours.
413The request body is too large.
429Rate limit reached: too many requests or tokens per minute, or the daily token quota is used up. Retry after the Retry-After header.
502The model failed to answer.
503The model is not available right now, for example while it starts. Retry later.
504The model took too long to answer.
curl · Transcriptions
curl -X POST https://api.models.blade.sh/v1/audio/transcriptions \
  -H "Authorization: Bearer $BLADE_API_KEY" \
  -F file=@recording.mp3 \
  -F model=MODEL_ID

Video

Create a video

POST /v1/video/generations

Start generating a video from a text prompt. Generation takes one to three minutes, so the request returns a job right away.

Poll the job with GET /v1/video/generations/{job_id} until it completes. You can have two unfinished jobs at a time; a third returns 429 video_jobs_quota_exceeded.

Parameters

NameTypeDescription
fpsinteger
Optional
Frames per second, from 8 to 30.
heightinteger
Optional
Height in pixels, up to 480.
modelstring
Required
The ID of a video model, from GET /v1/models.
negative_promptstring
Optional
What the video should avoid.
num_framesinteger
Optional
Number of frames, up to 161. Rounded to a multiple of 8 plus 1.
num_inference_stepsinteger
Optional
Generation steps, up to 50. More steps take longer and usually look better.
promptstring
Required
What the video shows, up to 2,000 characters.
seedinteger
Optional
Makes the generation repeatable.
widthinteger
Optional
Width in pixels, up to 704.

Responses

StatusMeaning
202The job, created. Keep its id to follow it.
400Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed.
401Missing, invalid or revoked API key.
404Unknown model, or a job that does not exist or is not yours.
429Rate limit reached: too many requests or tokens per minute, or the daily token quota is used up. Retry after the Retry-After header.
curl · Create a video
curl -X POST https://api.models.blade.sh/v1/video/generations \
  -H "Authorization: Bearer $BLADE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MODEL_ID",
    "prompt": "A red kite over the sea at dawn"
  }'

Video

List videos

GET /v1/video/generations

Your 20 most recent video jobs.

Responses

StatusMeaning
200The jobs, most recent first.
curl · List videos
curl https://api.models.blade.sh/v1/video/generations \
  -H "Authorization: Bearer $BLADE_API_KEY"

Video

Get a video

GET /v1/video/generations/{job_id}

The state of a video job, and the link to the video once it is ready.

The download link expires: read the job again to get a fresh one. A job that does not exist or is not yours returns 404.

Parameters

NameTypeDescription
job_idstring
Required
The id returned when the job was created.

Responses

StatusMeaning
200The job.
401Missing, invalid or revoked API key.
404Unknown model, or a job that does not exist or is not yours.
curl · Get a video
curl https://api.models.blade.sh/v1/video/generations/JOB_ID \
  -H "Authorization: Bearer $BLADE_API_KEY"

Video

Cancel a video

POST /v1/video/generations/{job_id}/cancel

Stop a video job that has not finished.

Cancellation is best effort: a job that is about to finish may complete anyway.

Parameters

NameTypeDescription
job_idstring
Required
The id returned when the job was created.

Responses

StatusMeaning
200The job, cancelled if the cancellation went through.
401Missing, invalid or revoked API key.
404Unknown model, or a job that does not exist or is not yours.
curl · Cancel a video
curl -X POST https://api.models.blade.sh/v1/video/generations/JOB_ID/cancel \
  -H "Authorization: Bearer $BLADE_API_KEY"

Documents

Detect

POST /v1/detect

Find regions in an image, such as the tables on a document page.

One JSON document, no streaming. Name the model explicitly; there is no default. Billed per processed image.

Parameters

NameTypeDescription
imagestring
Required
The image, in base64 or as a data:image/...;base64,... URI.
modelstring
Required
The ID of a detection model, from GET /v1/models.

Responses

StatusMeaning
200The regions found.
400Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed.
401Missing, invalid or revoked API key.
404Unknown model, or a job that does not exist or is not yours.
502The model failed to answer.
503The model is not available right now, for example while it starts. Retry later.
504The model took too long to answer.
curl · Detect
curl -X POST https://api.models.blade.sh/v1/detect \
  -H "Authorization: Bearer $BLADE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "image": "BASE64_IMAGE",
    "model": "MODEL_ID"
  }'

Documents

Parse

POST /v1/parse

Read the content of an image as structured text: handwriting, general OCR, a chart as a data table, a formula as LaTeX, or the structure of a table.

One JSON document, no streaming. The shape of the result depends on the model. Billed per processed image.

The table structure model needs the boxes of the tables in table_bboxes: get them from POST /v1/detect, or pass the whole page.

Parameters

NameTypeDescription
imagestring
Required
The image, in base64 or as a data:image/...;base64,... URI.
max_new_tokensinteger
Optional
The most tokens to generate, for the models that write text. Capped by the model's own limit.
modelstring
Required
The ID of a parsing model, from GET /v1/models.
table_bboxesarray of array of number
Optional
Required by the table structure model: the pixel boxes of the tables to read. Ignored by the other models.

Responses

StatusMeaning
200The parsed content. Its shape depends on the model.
400Invalid request: a parameter is out of range, the prompt does not fit in the context, or the body is malformed.
401Missing, invalid or revoked API key.
404Unknown model, or a job that does not exist or is not yours.
502The model failed to answer.
503The model is not available right now, for example while it starts. Retry later.
504The model took too long to answer.
curl · Parse
curl -X POST https://api.models.blade.sh/v1/parse \
  -H "Authorization: Bearer $BLADE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "image": "BASE64_IMAGE",
    "model": "MODEL_ID"
  }'

Catalogue

List models

GET /v1/models

The models you can call, with their limits, capabilities, supported parameters and prices.

Use the id of a model as the model of your requests.

Responses

StatusMeaning
200The model list, in the OpenAI format.
401Missing, invalid or revoked API key.
curl · List models
curl https://api.models.blade.sh/v1/models \
  -H "Authorization: Bearer $BLADE_API_KEY"

For agents and tools

The same reference in machine-readable formats, for code generators and AI agents.

Generated from the API contract of . Model availability and enabled features depend on your account: check them in the console before integrating.

Contact the team

Loading the form…