blade Get your API key

Developer docs Preparation guide · waitlist

Your model.
Ready for a GPU.

Prepare the workload now. Complete deployment with the runtime and access provided during onboarding.

Start with the workload

Serverless is for your model or Python code. Open Models is for calling models that are already hosted behind an API. Choose the former when you need control over model weights, dependencies or the work around inference.

Serverless is opening gradually. This guide helps you prepare for onboarding; the install commands, GPU inventory and deployment interface will be provided with your access.

Package the model and its dependencies

Keep the model identifier or weights location, Python dependencies and inference entry point together. Record the model licence and any access requirements for its weights. Pin the dependencies you have actually tested.

A useful project structure is:

my-model/
  requirements.txt
  predict.py
  sample-request.json
  README.md

This is a suggested project layout, not a Serverless SDK contract. Your predict.py should load the model and accept the input shape your application will send.

Size the workload before choosing a GPU

Prepare Why it matters
Model size, precision and expected input size Establish the GPU memory requirement.
A representative sample request Measure latency and output quality on the real workload.
Expected request rate or batch volume Decide how much concurrent work the service must handle.
Model loading time Understand how startup affects the first request after idle.
Runtime dependencies and weights access Make deployment reproducible.
Required timeout and retry behaviour Define how the application handles long or interrupted work.

Check the deployment during onboarding

Confirm the supported Python runtime, GPU class, dependency installation and secrets configuration with the team. Deploy a small test workload first, then exercise its input and output contract before connecting it to production traffic.

For an interactive endpoint, measure cold and warm latency separately. For a batch pipeline, test a representative group of inputs and verify that retrying a job does not duplicate downstream work.

Confirm scaling and costs

Discuss the concurrency limit, startup behaviour, idle policy and billing unit for your workload. The Serverless product is designed around autoscaling and scale to zero; the available settings and commercial terms are confirmed when access opens.

Open Models token prices and their dynamic pricing are not a Serverless GPU quote. Dedicated model endpoints described in the inference API documentation are also a separate billing surface.

Ready to bring your model?

Join the Serverless waitlist with your model, memory requirement, representative input and expected traffic. If you need a hosted model API now, start with Open Models.

Contact the team

Loading the form…