Make your first API call.
Set up your key and model. Follow complete Python, JavaScript or curl examples, then add streaming and troubleshoot errors.
Get started with Open Models → PricingKnow what a request costs.
Read the billing unit and hourly multiplier. Walk through a worked example.
Understand dynamic pricing → Serverless · WaitlistBring your own model.
Prepare the runtime, model weights and workload requirements for onboarding.
Prepare a GPU deployment →API reference and model details
Open Models is Blade’s hosted inference API. Use https://api.models.blade.sh/v1 with your API key and an available model ID. Start with the guides and reference below.
- API routes, accepted parameters and errors
- Model catalogue
- Model pricing
- Hosting, external tools and data handling
For agents and tools
Use the catalogue below for this site’s published model rates. The gateway documentation index covers the API protocol; some of those documents still use the legacy name Shadow Inference.
- Model catalogue (Markdown)
- Gateway documentation index (llms.txt)
The catalogue contains 16 models with From prices in EUR: the lowest price the dynamic multiplier can reach for each model. Check model access, supported parameters and the current rate in your console before integrating.