Transcribe Audio
Transcribe speech to text using Whisper on GPU — 99+ languages, word timestamps, subtitle output.
/v1/audio/transcribe
curl "https://audio.toolkitapi.io/v1/audio/transcribe?audio_url=https://example.com/audio.mp3&language=en&model=base"
import httpx
resp = httpx.get(
"https://audio.toolkitapi.io/v1/audio/transcribe?audio_url=https://example.com/audio.mp3&language=en&model=base",
)
print(resp.json())
const resp = await fetch("https://audio.toolkitapi.io/v1/audio/transcribe?audio_url=https://example.com/audio.mp3&language=en&model=base", {
});
const data = await resp.json();
console.log(data);
# See curl example
{
"status": "ok",
"text": "Hello, this is a sample transcription.",
"segments": [{"start": 0.0, "end": 2.5, "text": "Hello, this is a sample transcription."}],
"language": "en",
"duration_seconds": 2.5,
"task_id": "abc123",
"processing_time_seconds": 1.2
}
Try It Live
Description
How to Use
1. Provide a publicly accessible audio URL (WAV, MP3, FLAC, etc.).
2. Set the `language` parameter to the language code (e.g., `en`, `fr`, `de`) or `auto` for automatic detection.
3. Choose a Whisper model size with the `model` parameter. `base` is a good default balance of speed and accuracy.
4. Set `response_format` to `json` for structured segments, `srt` or `vtt` for subtitles, or `text` for plain text.
About This Tool
Transcribe audio files to text using OpenAI's Whisper model running on GPU. Supports 99+ languages, automatic language detection, and optional word-level timestamps. Output formats include plain text, JSON with segments, SRT, and VTT subtitles.
Multiple Whisper model sizes are available: `tiny` (fastest), `base`, `small`, `medium`, and `large-v3` (most accurate).
Why Use This Tool
- Podcast transcription — Generate text from podcast episodes for search and accessibility
- Meeting notes — Transcribe recorded meetings into searchable text
- Subtitle generation — Create SRT/VTT subtitles for videos
- Multi-language support — Transcribe audio in 99+ languages
Frequently Asked Questions
What audio formats are supported?
How long can the audio be?
Can I transcribe multiple languages in one file?
Start using Transcribe Audio now
Get your free API key and make your first request in under a minute.