🤖

Transcribe Audio

Transcribe speech to text using Whisper on GPU — 99+ languages, word timestamps, subtitle output.

GET 1 credit /v1/audio/transcribe
curl "https://audio.toolkitapi.io/v1/audio/transcribe?audio_url=https://example.com/audio.mp3&language=en&model=base"
import httpx

resp = httpx.get(
    "https://audio.toolkitapi.io/v1/audio/transcribe?audio_url=https://example.com/audio.mp3&language=en&model=base",
)
print(resp.json())
const resp = await fetch("https://audio.toolkitapi.io/v1/audio/transcribe?audio_url=https://example.com/audio.mp3&language=en&model=base", {
});
const data = await resp.json();
console.log(data);
# See curl example
Response 200 OK
{
  "status": "ok",
  "text": "Hello, this is a sample transcription.",
  "segments": [{"start": 0.0, "end": 2.5, "text": "Hello, this is a sample transcription."}],
  "language": "en",
  "duration_seconds": 2.5,
  "task_id": "abc123",
  "processing_time_seconds": 1.2
}

Try It Live

Live Demo

Description

Transcribe speech to text using Whisper on GPU — 99+ languages, word timestamps, subtitle output.

How to Use

1

1. Provide a publicly accessible audio URL (WAV, MP3, FLAC, etc.).

2

2. Set the `language` parameter to the language code (e.g., `en`, `fr`, `de`) or `auto` for automatic detection.

3

3. Choose a Whisper model size with the `model` parameter. `base` is a good default balance of speed and accuracy.

4

4. Set `response_format` to `json` for structured segments, `srt` or `vtt` for subtitles, or `text` for plain text.

About This Tool

Transcribe audio files to text using OpenAI's Whisper model running on GPU. Supports 99+ languages, automatic language detection, and optional word-level timestamps. Output formats include plain text, JSON with segments, SRT, and VTT subtitles.

Multiple Whisper model sizes are available: `tiny` (fastest), `base`, `small`, `medium`, and `large-v3` (most accurate).

Why Use This Tool

Frequently Asked Questions

What audio formats are supported?
WAV, MP3, FLAC, OGG, AAC, M4A, WMA, and most other common audio formats.
How long can the audio be?
There's no hard limit, but longer files take proportionally more time. For files over 30 minutes, consider using the `tiny` or `base` model for faster results.
Can I transcribe multiple languages in one file?
Set `language=auto` and Whisper will detect the language automatically. However, it works best when the audio is predominantly in one language.

Start using Transcribe Audio now

Get your free API key and make your first request in under a minute.