Voice notes, ready to use.
Turn recorded feedback or a voice memo into a transcript. Pass the text to your app to create a draft note, support ticket or task.
Open Models / Speech & audio
Add speech-to-text and text-to-speech to your app. Transcribe recordings, give written answers a voice and connect both with the same Blade API key.
Start BuildingSpeech to text. Text to speech. Open models on European GPUs.
“How do I invite my team?”
How do I invite my team?
“Open your workspace settings…”
Start with the audio your users already have. Add the feature that saves them the next manual step.
Turn recorded feedback or a voice memo into a transcript. Pass the text to your app to create a draft note, support ticket or task.
Transcribe interviews, lessons and podcasts. Add search or summarisation in your application so a useful detail is easier to find.
Read an onboarding message, an article excerpt or an assistant’s answer aloud. Generate a speech file from text your app controls.
Speech to text
Send the recording from your backend. Blade returns a transcript you can display, store or pass to a chat model.
BLADE_API_KEY and BLADE_STT_MODEL on your server. Install the client with pip install openai.recording.mp3 and run the example.import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.models.blade.sh/v1",
api_key=os.environ["BLADE_API_KEY"],
)
with open("recording.mp3", "rb") as audio:
transcript = client.audio.transcriptions.create(
model=os.environ["BLADE_STT_MODEL"],
file=audio,
)
print(transcript.text)Documented inputs: WAV, MP3, M4A, FLAC or OGG, up to 25 MB. Model-specific limits may also apply.
Start with real recordings. Test the languages, background noise and vocabulary your users actually bring.
Keep the original. Let a user check and correct a transcript before it drives an important action.
Build on the text. Search, summarisation and task extraction are separate steps in your application.
Text to speech
Send a welcome message, product guidance or a generated response to the speech endpoint. Return the audio to your app and let the user press play.
The documented speech route accepts up to 4,096 characters and returns a WAV file. This example uses its tara voice. Select a model and voice supported by your account.
# Reuse the client configured above.
with client.audio.speech.with_streaming_response.create(
model=os.environ["BLADE_TTS_MODEL"],
voice="tara",
input="Your workspace is ready. Let's get started.",
) as audio:
audio.stream_to_file("welcome.wav")Set BLADE_TTS_MODEL to the model ID for speech generation. The SDK writes the returned audio to a local file.
A recorded question can become a spoken answer. Your backend coordinates each request and decides what happens between them.
Audio file → written question.
Use your logic or a chat model.
Written reply → audio file.
Need reasoning or web search in the middle?
Build an assistantThese examples use file transcription and text-to-speech requests. They do not provide a live microphone session, turn detection or interruption handling. Build and validate those behaviours separately if your product needs them.
The documented upload limit is 25 MB. Check the selected model’s limits, split large recordings in your application when needed and preserve the order of the segments.
Compare the audio models and their billing units in the pricing catalogue, then test representative audio. Use the console for current rates.