🤖

Text-to-Speech

Convert text to natural-sounding speech using Kokoro TTS on GPU — multiple voices, adjustable speed.

GET 1 credit /v1/audio/synthesize
curl "https://audio.toolkitapi.io/v1/audio/synthesize?text=Hello%20world&voice=en_female&speed=1.0&output_format=mp3"
import httpx

resp = httpx.get(
    "https://audio.toolkitapi.io/v1/audio/synthesize?text=Hello%20world&voice=en_female&speed=1.0&output_format=mp3",
)
print(resp.json())
const resp = await fetch("https://audio.toolkitapi.io/v1/audio/synthesize?text=Hello%20world&voice=en_female&speed=1.0&output_format=mp3", {
});
const data = await resp.json();
console.log(data);
# See curl example
Response 200 OK
{
  "status": "ok",
  "audio_url": "https://storage.example.com/audio/abc123.mp3",
  "format": "mp3",
  "sample_rate": 24000,
  "duration_seconds": 1.5,
  "task_id": "abc123"
}

Try It Live

Live Demo

Description

Convert text to natural-sounding speech using Kokoro TTS on GPU — multiple voices, adjustable speed.

How to Use

1

1. Provide the text to synthesize in the `text` parameter (up to 5000 characters).

2

2. Select a voice with the `voice` parameter (e.g., `en_female`, `en_male`).

3

3. Optionally adjust playback speed with the `speed` parameter (0.5× to 2.0×).

4

4. Choose an output format: `mp3` (default, best compatibility), `wav` (lossless), or `ogg`.

About This Tool

Convert text into natural-sounding speech using Kokoro TTS running on GPU. Choose from multiple voice options and adjust playback speed. Output formats include WAV, MP3, and OGG.

Kokoro TTS produces high-quality, natural-sounding speech suitable for voiceovers, accessibility, and conversational AI applications.

Why Use This Tool

Frequently Asked Questions

How long does generation take?
Most short text (under 500 characters) generates in under 2 seconds. Longer text scales linearly.
Can I use custom voices?
Currently, the available voices are `en_female` and `en_male`. Additional voices may be added in future releases.
What's the maximum text length?
Up to 5000 characters per request. For longer content, split into multiple requests and concatenate the audio files.

Start using Text-to-Speech now

Get your free API key and make your first request in under a minute.