Skip to main content

Create speech

Generates audio from the input text.

Parameters

string
required
The text to generate audio for. The maximum length is 4096 characters.
string
required
One of the available TTS models: tts-1, tts-1-hd, gpt-4o-mini-tts, or gpt-4o-mini-tts-2025-12-15.
  • tts-1 - Standard quality, optimized for speed
  • tts-1-hd - High definition quality
  • gpt-4o-mini-tts - GPT-4 optimized TTS with voice instructions support
  • gpt-4o-mini-tts-2025-12-15 - Latest GPT-4 optimized TTS model
string
required
The voice to use when generating the audio. Supported built-in voices are alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse, marin, and cedar.Previews of the voices are available in the Text to speech guide.
string
Control the voice of your generated audio with additional instructions. Does not work with tts-1 or tts-1-hd.
string
default:"mp3"
The format to audio in. Supported formats are mp3, opus, aac, flac, wav, and pcm.
float
default:"1.0"
The speed of the generated audio. Select a value from 0.25 to 4.0. 1.0 is the default.
string
The format to stream the audio in. Supported formats are sse and audio. sse is not supported for tts-1 or tts-1-hd.

Response

Returns the audio file content as binary data.

Examples

Generate speech with different voices

Adjust speech speed

Use GPT-4 TTS with instructions

Generate in different formats

Async usage

Supported audio formats

  • mp3 - MPEG audio format (default)
  • opus - Opus audio format
  • aac - AAC audio format
  • flac - FLAC lossless audio format
  • wav - Waveform audio format
  • pcm - Raw PCM audio at 24kHz (16-bit signed, low-endian)

Available voices

The following voices are available for speech generation:
  • alloy - Neutral and balanced
  • ash - Clear and articulate
  • ballad - Smooth and melodic
  • coral - Warm and friendly
  • echo - Resonant and clear
  • fable - Expressive and storytelling
  • nova - Energetic and engaging
  • onyx - Deep and authoritative
  • sage - Calm and measured
  • shimmer - Light and cheerful
  • verse - Poetic and rhythmic
  • marin - Natural and conversational
  • cedar - Warm and grounded
Preview audio samples for each voice in the OpenAI Text-to-Speech guide.