Chat that helps.
Guide onboarding, explain a product or answer support questions using context supplied by your app. Stream the answer into your interface.
Customer support / in-app assistantsOpen Models / LLM inference
Answer a customer. Research a question. Write the next function. Connect open models to your product or coding workflow through an OpenAI-compatible inference API.
Start BuildingOpenAI-compatible API. Hosted on European GPUs.
“What changed in this library?
Find the release notes.”
web_searchA summary of the changes.
Links to the original sources.
Start with one job your users need done. Choose the model and tools for that job.
Guide onboarding, explain a product or answer support questions using context supplied by your app. Stream the answer into your interface.
Customer support / in-app assistantsDraft functions, propose a refactor or generate test cases. Give the model the relevant files and constraints, then check the result in your project.
Implementation / tests / documentationCompare migration approaches, diagnose an issue or plan a change. Add tool calling when the answer needs fresh information from the web or your own service.
Research / analysis / connected workflowsReasoning controls, streaming and tool calling depend on the selected model. Check its capabilities in the console or GET /v1/models.
In your product
Keep your website or service. Call Blade from your backend with the OpenAI SDK, and return the model’s answer to your users.
pip install openai, then run this Python example.import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.models.blade.sh/v1",
api_key=os.environ["BLADE_API_KEY"],
timeout=60.0,
)
response = client.chat.completions.create(
model=os.environ["BLADE_MODEL"],
messages=[{
"role": "user",
"content": "Explain this API to a new user.",
}],
max_tokens=512,
)
print(response.choices[0].message.content)Set BLADE_MODEL to a model ID from your account. OpenAI compatibility applies to supported routes and parameters.
Bring your context. Send the conversation and relevant product data with each request.
Keep actions in your control. Validate function arguments and permissions before running your own tools.
Measure real usage. Handle timeouts and rate limits. Track tokens, latency and cost before rolling out.
Flash by Staan
Connect your assistant to current web information. With Staan web search enabled, Blade handles the search loop and returns an answer with source annotations.
Requires web search access on your account and a model with native tool calling. Searches are billed separately and sent to the search provider.
How external tools handle dataAdd search to the same client
response = client.chat.completions.create(
model=os.environ["BLADE_MODEL"],
messages=[{
"role": "user",
"content": "Find the latest Python release notes.",
}],
extra_body={
"tools": [{"type": "web_search"}],
},
)
message = response.choices[0].message
print(message.content)
print(message.model_dump().get("annotations", []))Blade’s built-in tool extension. Render the returned annotations as source links in your interface.
Blade’s built-in toolsDeclare web search; Blade executes it and returns the answer.
Your own functionsDeclare a function; your backend executes the requested call, sends back its result, then asks the model to continue.
In your daily workflow
Spend less time on repetitive implementation. Let your primary assistant scope and review the work; send focused coding tasks to an open model through Blade.
Scope, files, acceptance checks.
A bounded task through the API.
Inspect the diff. Run the tests.
Download the Python example, install openai and set the same key and model as above. Your assistant can call this script from a skill or a tool. It returns a proposal; you test and apply it.
python delegate.py brief.txt \
--context src/validator.py \
> proposal.txtCreate brief.txt with your task and replace the context path with a file you want to send. Each run makes a billed API call. No automatic file edits.
Create a project skill in .claude/skills/blade-delegate/SKILL.md that tells Claude when to call the runner, which files to include and which checks to run. Configure the runner’s key and model, then ask Claude to delegate a bounded task. Claude keeps the project context and reviews the proposed changes.
Start with tests, documentation or an isolated function. The runner returns code or a diff for review.
Claude Code skillsIn Codex, create .agents/skills/blade-delegate/SKILL.md with instructions to call the runner and review its output. In a GPT-powered application, expose the same API call as a tool in your own backend. Pass the brief and selected context, then return the result to the primary assistant.
This is a delegation workflow, not a change to the default model in a standard ChatGPT conversation.
Build a Codex skillRun hermes model, choose Custom endpoint, and enter https://api.models.blade.sh/v1, your Blade key and an available model ID. Choose a model that supports native tool calling for agent workflows.
This configures Blade as the inference provider. Delegating only selected jobs requires a separate tool or runner in your workflow.
Hermes provider setupNo. Select a model for the features you need. Check the console and the model’s supported_parameters in GET /v1/models. Reasoning controls and supported values vary by deployment.
A custom function call is a request for your application to execute a function. Your backend validates it, runs the allowed operation and returns the result. Blade’s built-in server tools follow a separate, managed execution path.
Both workflows call the Open Models API. Use separate keys for separate projects where possible, monitor usage, and compare the selected model’s rates on the pricing page.