blade Get your API key

Open Models / Speech & audio

Let your app
listen and speak.

Add speech-to-text and text-to-speech to your app. Transcribe recordings, give written answers a voice and connect both with the same Blade API key.

Start Building

Speech to text. Text to speech. Open models on European GPUs.

audio → text → audioExample workflow
recording.mp3

“How do I invite my team?”

Transcribe the question

How do I invite my team?

Speak your app’s reply

“Open your workspace settings…”

Your app provides the reply, or calls a chat model between transcription and speech generation. Illustrative content.

Transcribe recordings.
Generate speech.

Start with the audio your users already have. Add the feature that saves them the next manual step.

Voice notes, ready to use.

Turn recorded feedback or a voice memo into a transcript. Pass the text to your app to create a draft note, support ticket or task.

Content people can find.

Transcribe interviews, lessons and podcasts. Add search or summarisation in your application so a useful detail is easier to find.

A voice for your interface.

Read an onboarding message, an article excerpt or an assistant’s answer aloud. Generate a speech file from text your app controls.

Speech to text

An audio file in.
Useful text out.

Send the recording from your backend. Blade returns a transcript you can display, store or pass to a chat model.

  1. Create an API key and select a transcription model in the console.
  2. Set BLADE_API_KEY and BLADE_STT_MODEL on your server. Install the client with pip install openai.
  3. Save a sample as recording.mp3 and run the example.
Transcription API reference
transcribe.pyPython
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.models.blade.sh/v1",
    api_key=os.environ["BLADE_API_KEY"],
)

with open("recording.mp3", "rb") as audio:
    transcript = client.audio.transcriptions.create(
        model=os.environ["BLADE_STT_MODEL"],
        file=audio,
    )

print(transcript.text)

Documented inputs: WAV, MP3, M4A, FLAC or OGG, up to 25 MB. Model-specific limits may also apply.

Start with real recordings. Test the languages, background noise and vocabulary your users actually bring.

Keep the original. Let a user check and correct a transcript before it drives an important action.

Build on the text. Search, summarisation and task extraction are separate steps in your application.

Connect the steps.
Keep your product logic.

A recorded question can become a spoken answer. Your backend coordinates each request and decides what happens between them.

01 / TranscribeListen

Audio file → written question.

02 / Your applicationRespond

Use your logic or a chat model.

03 / Generate speechSpeak

Written reply → audio file.

Need reasoning or web search in the middle?

Build an assistant

Before you add voice.

Is this a real-time voice agent API?

These examples use file transcription and text-to-speech requests. They do not provide a live microphone session, turn detection or interruption handling. Build and validate those behaviours separately if your product needs them.

Can I transcribe a long recording?

The documented upload limit is 25 MB. Check the selected model’s limits, split large recordings in your application when needed and preserve the order of the segments.

How do I choose a model and estimate cost?

Compare the audio models and their billing units in the pricing catalogue, then test representative audio. Use the console for current rates.

Contact the team

Loading the form…