GPT Proto

GPTProto

  • Dashboard
  • Text

    • claude
      Claude Opus 5New
    • google
      Gemini 3.6 Flash
    • google
      Gemini 3.5 Flash Lite
    • moonshotai
      Kimi K3
    • openai
      GPT 5.6 Luna

    Image

    • bytedance
      Dola Seedream 5.0 Pro 260628New
    • google
      Gemini 3.1 Flash Lite Image
    • google
      Gemini 3.1 Flash Image
    • openai
      GPT Image 2
    • google
      Gemini 3.1 Flash Image Preview

    Video

    • kling
      Kling v3.0 4kNew
    • bytedance
      Dreamina Seedance 2.0 Mini 260615
    • kling
      Kling v3 Omni 4k
    • bytedance
      Dreamina Seedance 2.0 Fast 260128
    • bytedance
      Dreamina Seedance 2.0 260128
    Explore 214+ Models >
  • Generator

    • Create Image
    • Create Video
    • Edit in Canvas

    Features

    • Anime to Real Life AINew
    • Anime AI Art Generator
    • AI Object Remover
    • AI Image Editor
    • Unrestricted AI Image Generator
    • AI Motion Transfer
    • AI Clothes Remover
    • AI Watermark Remover
    • AI Image Enhancer Online
    • Online Background Remover Tool
    • AI Face Swap Image
    • AI Passport Photo Maker
    • MS Paint AI Generator

    Prompts

    • Seedance 2.0 PromptsNew
    • GPT Image 2 Prompts
    • Nano Banana Pro Prompts
  • AI Blog

    • GLM 5.2 vs MiniMax M3: Which Is Better for Coding and Frontend Work?
    • How to Create Your Own AI Character With an API—No Coding Required
    • Kimi K3 vs Claude Opus 5: Which Is Better for Coding and AI Agents?
    • 20 Free Seedream 5.0 Pro Packaging Design Prompts for Products and E-commerce
    • GLM-5.2 vs Kimi K3 for Coding: Which Is Better for Developers in 2026?
    Explore All >

    AI Insight

    • What Is Seedance 2.5? What It Can Make and How It Upgrades Seedance 2.0
    • MiniMax H3 Is Here: What Its Video Editing Upgrade Actually Changes
    • What Is Emochi AI—and Why Is It Growing So Fast? (2026)
    • What Is Kimi K3—and Is It Really Close to GPT-5.6 and Fable 5?
    • 12 Best AI Video Generation Tools in 2026 for YouTube, TikTok, Text and Images
    Explore All >

    AI Docs

    • gpt-image-2
    • gpt-5.4
    • kimi-k2.5
    • claude-opus-4-6
    • kling-v3.0-pro
    Explore All >

    AI Skills

    • document-illustrator
    • video-wrapper
    • claude-to-im
    • openclaw-gptproto-config
    • openclaw-installer
    Explore All >
PricingGet Started Now
  1. Home
  2. /Model
  3. /OpenAI
  4. /gpt-4o-mini-tts
OpenAI
gpt-4o-mini-tts
Documentation
Document attachment
Turn up to 2,000 input tokens into steerable spoken audio with 13 built-in voices and MP3, Opus, AAC, FLAC, WAV, or PCM output. Access gpt-4o-mini-tts on GPTProto at $0.42 per 1M text-input tokens and $8.40 per 1M audio-output tokens—30% below OpenAI's listed rates.

$ 0.42
$ 0.6

$ 8.4
$ 12

text

audio

$ 0.42
$ 0.6

text

$ 8.4
$ 12

audio

Related Models
All Models
Google
Google
gemini-2.5-flash-preview-tts
$ 6
$ 10
Google
Google
gemini-2.5-pro-preview-tts
$ 12
$ 20
MiniMax
MiniMax
speech-2.6-hd
$ 60
$ 100
MiniMax
MiniMax
speech-2.5-turbo-preview
$ 36
$ 60
MiniMax
MiniMax
speech-2.5-turbo-preview-voice-clone
$ 0.5003
$ 0.8338
MiniMax
MiniMax
speech-02-turbo
$ 0.002
$ 0.0034

GPT-4o-mini-TTS API

Generate natural-sounding speech with plain-language control over accent, emotion, intonation, pace, tone, and whispering. The affordable GPT-4o-mini-TTS API on GPTProto uses one API key and the same account balance as 200+ other AI models.

Steerable Speech Delivery

Describe how a line should sound in natural language. Direct the accent, emotional range, intonation, pace, tone, or whispering through the model's instructions field.

gpt tts emotional

13 Built-In Voices

Choose from 13 documented voices, including alloy, coral, marin, and cedar. OpenAI currently recommends marin or cedar when output quality is the priority.

gpt api low latency

Streaming and Six Formats

Return MP3, Opus, AAC, FLAC, WAV, or PCM audio. Streaming lets playback begin before the complete file is available; WAV and PCM avoid decoding overhead in latency-sensitive workflows.

gpt 4o mini cost

30% Lower API Rates

GPTProto lists text input at $0.42 and audio output at $8.40 per 1M tokens—30% below OpenAI's $0.60 and $12.00 published rates.

gpt 4o mini tts api

What Is GPT-4o-mini-TTS?

GPT-4o-mini-TTS is OpenAI's text-to-speech model for applications that already have written text and need controllable spoken audio. It accepts text only and returns audio only. Unlike a general chat model, it is not designed to analyze images, transcribe incoming speech, maintain a long conversation history, or return a text answer.

The model's main advantage over older preset-style TTS workflows is its instructions field. Instead of selecting only a voice and speed, developers can describe the intended delivery in ordinary language: calm or excited, conversational or formal, softly spoken or energetic, with guidance for accent, pacing, intonation, and emphasis. This makes the OpenAI GPT-4o-mini-TTS API useful for narration, product onboarding, accessibility audio, game dialogue, and spoken responses generated from an application's text output.

Specification GPT-4o-mini-TTS
Model string gpt-4o-mini-tts
Input modality Text
Output modality Audio
Maximum input 2,000 input tokens per request
Built-in voices 13
Output formats MP3, Opus, AAC, FLAC, WAV, PCM
Delivery Complete audio response or streamed audio
Adjustable speed 0.25 to 4.0; default 1.0
GPTProto pricing $0.42 text input / $8.40 audio output per 1M tokens

Choose the Right Voice and Audio Format

GPT-4o-mini-TTS currently documents the built-in voices alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, and cedar. OpenAI recommends trying marin or cedar first for quality-focused output. The voices can generate speech in multiple languages, but they are optimized for English, so test names, acronyms, numbers, and local pronunciation with real scripts before release.

Voice and format solve different problems. The voice determines the base sound; the instruction controls delivery; the response format determines how the audio fits the product.

Format Best fit
MP3 General web playback, downloads, and compact files
Opus Internet streaming and communication applications
AAC Mobile apps and video publishing workflows
FLAC Lossless storage, editing, and archiving
WAV Low-latency playback and audio editing without compressed-audio decoding
PCM Raw 24 kHz, 16-bit little-endian audio for custom playback pipelines

Use MP3 as a practical default. Choose Opus when bandwidth matters, FLAC when lossless storage matters, and WAV or PCM when avoiding decoding overhead is more important than file size.

How to Write a GPT-4o-mini-TTS Prompt

A useful GPT-4o-mini-TTS prompt describes four things: the speaker's role, the listener, the emotional intent, and the delivery. Keep the spoken script in the input field and put voice direction in the instructions field. This separation makes prompts easier to reuse and test.

Application Example instructions prompt
Customer support Speak calmly and reassuringly. Use a warm conversational tone, moderate pace, and gentle emphasis on the next step. Do not sound overly cheerful.
Product onboarding Sound friendly, clear, and confident. Keep the pace brisk, pause briefly after each action, and emphasize button names without sounding like an advertisement.
Short-form narration Use an intimate documentary style with controlled energy. Begin softly, build curiosity through the middle, and slow down for the final sentence.
Accessibility reading Read clearly at a measured pace. Pronounce every number and abbreviation distinctly, avoid dramatic emotion, and pause between headings and body text.

Change one variable at a time when testing a GPT-4o-mini-TTS API voice. First select the base voice, then refine tone and pacing, and finally test pronunciation. This makes it easier to identify whether a problem comes from the voice choice, the instruction, or the source text.

Common GPT-4o-mini-TTS API Applications

  • Narration and voiceovers: Convert articles, product videos, lessons, and short scripts into consistent spoken audio.

  • App guidance: Read onboarding steps, alerts, navigation cues, or generated summaries aloud.

  • Voice-agent output: Turn a text response from an assistant into streamed speech. Use a Realtime API workflow instead when the product needs full two-way, speech-to-speech interaction.

  • Accessibility: Add spoken versions of messages, documents, and interface content for users who prefer or require audio.

  • Multilingual delivery: Generate speech from supported-language text, while testing each target language because the built-in voices are optimized for English.

Plan Long Scripts Before Sending Requests

The maximum input is 2,000 tokens per request, not a 128K context window. Split longer material at sentence or paragraph boundaries. Reuse the same voice, instruction, speed, and response format for every segment, and avoid cutting a sentence, number, or proper name across two requests. After generation, listen to the joins for repeated pauses, inconsistent emphasis, or pronunciation changes.

For interactive playback, request streaming so the application can start playing audio before generation finishes. For prerecorded narration, generate complete files and run a quality check before publishing. In either case, clearly disclose to end users that the voice is AI-generated rather than human.

How to Get a gpt-4o-mini-tts API Key

Getting a gpt-4o-mini-tts API key takes four steps and a few minutes. Create a free GPTProto account, add credits, generate your key, and make your first call — at $0.42 / $8.4 it's a cheaper gpt-4o-mini-tts API key than going direct, and one key works across every model on the platform. Full gpt-4o-mini-tts Documentation is in the docs.

Sign up

Sign up

Create your free GPT Proto account to begin. You can set up an organization for your team at any time.

Top up

Top up

Your balance can be used across all models on the platform, including gpt-4o-mini-tts, giving you the flexibility to experiment and scale as needed.

Generate your API key

Generate your API key

In your dashboard, create an API key — you'll need it to authenticate when making requests to gpt-4o-mini-tts.

Make your first API call

Make your first API call

Use your API key with our sample code to send a request to gpt-4o-mini-tts via GPT Proto and see instant AI-powered results.

Get API Key

FAQ

How much does the GPT-4o-mini-TTS API cost on GPTProto?

GPTProto lists GPT-4o-mini-TTS API pricing at $0.42 per 1M text-input tokens and $8.40 per 1M audio-output tokens. OpenAI currently lists $0.60 and $12.00 respectively, so both GPTProto rates are 30% lower. Pricing can change; use the live price panel as the final billing reference.

What is the GPT-4o-mini-TTS input limit?

The model accepts a maximum of 2,000 input tokens per request. It does not have a 128K context window. Split longer scripts at natural sentence or paragraph boundaries and keep the same voice and instructions across segments.

Can I control the GPT-4o-mini-TTS voice with a prompt?

Yes. Use natural-language instructions to guide accent, emotional range, intonation, speed, tone, and whispering. Keep the words to be spoken in the input and the performance direction in the instructions so each part can be edited independently.

Which voices and audio formats are available?

OpenAI documents 13 built-in voices: alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, and cedar. Supported formats are MP3, Opus, AAC, FLAC, WAV, and PCM.

Does the GPT-4o-mini-TTS API support streaming?

The model's Speech API supports streamed audio, allowing playback to begin before the full output is complete. OpenAI documents both direct audio streaming and an SSE stream format. Before publishing the SSE claim for GPTProto, verify that the current GPTProto endpoint exposes stream_format: "sse".

Does GPT-4o-mini-TTS support custom voice cloning?

OpenAI now offers custom voices only to eligible customers and requires consent and sample recordings. That availability does not automatically extend to a third-party API route. Unless GPTProto has verified custom voice-ID support, present the 13 built-in voices as the supported option on this page.

Can GPT-4o-mini-TTS generate speech in languages other than English?

Yes. The TTS guide lists many supported languages, including Chinese, Japanese, Korean, Spanish, French, German, Portuguese, Arabic, and Hindi. OpenAI notes that its voices are optimized for English, so test pronunciation and naturalness with representative text for every target language.

Related Articles

More Blogs
7 Best AI Text to Speech Tools in 2026 for TikTok and YouTube

7 Best AI Text to Speech Tools in 2026 for TikTok and YouTube

Compare the best AI text to speech tools for TikTok and YouTube, including free options, voice quality, commercial rights, video workflows, and API automation.

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite Explained: Which One Should You Use?

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite Explained: Which One Should You Use?

Compare Gemini 3.6 Flash and Gemini 3.5 Flash-Lite pricing, speed, benchmarks and use cases. See which new Google model fits your AI workload.

Which Text to Speech AI API Is Actually Best in 2026?

Which Text to Speech AI API Is Actually Best in 2026?

Compare the best text to speech AI APIs for voice quality, real-time agents, multilingual audio, pricing, free tiers, voice cloning, and production use.

5 Best Chinese LLM Models in 2026: Which One Is Best for Coding?

5 Best Chinese LLM Models in 2026: Which One Is Best for Coding?

Compare Kimi K3, GLM 5.2, Qwen3.7 Max, MiniMax M3 and DeepSeek V4 Pro for coding, cost, speed, 1M context and open-weight access.

GPT Proto

Empowering AI Innovation with Global Scale and Stability:

With our flagship product GPT Proto, we offer a unified interface to access and combine APIs from the world's leading AI providers—spanning text, vision, speech, and beyond. We empower developers and enterprises to simplify integration and accelerate innovation without limits.

Global Infrastructure, Local Compliance:

To ensure enterprise-grade reliability and compliance, Talent Tech Global Limited operates specifically as our global Billing and Contracting Entity. Meanwhile, our core technical infrastructure and R&D teams are strategically distributed across global innovation hubs, including Silicon Valley, Singapore, and Hong Kong.

Built to Scale:

We understand that stability is paramount. Our platform is built on a robust, decentralized architecture supporting dynamic Auto-scaling. Whether you are running a pilot or handling millions of concurrent requests, our system expands instantly to meet demand—guaranteeing that your business never outgrows our infrastructure.

Navigation

  • Dashboard
  • Models
  • Create Image
  • AI Image Upscale
  • AI Background Remover
  • Create Video
  • Edit in Canvas
  • Features
  • Pricing
  • AI Docs
  • AI Blog
  • AI Insight
  • AI Skills

Features

  • Anime to Real Life AI
  • Anime AI Art Generator
  • AI Object Remover
  • AI Image Editor
  • Unrestricted AI Image Generator
  • AI Motion Transfer
  • AI Clothes Remover
  • AI Watermark Remover
  • AI Image Enhancer Online
  • Online Background Remover Tool
  • AI Face Swap Image
  • AI Passport Photo Maker
  • MS Paint AI Generator
Explore all features >

Text

  • Claude Opus 5
  • Gemini 3.6 Flash
  • Gemini 3.5 Flash Lite
  • Kimi K3
  • GPT 5.6 Luna
  • GPT 5.6 Terra
  • GPT 5.6 Sol
  • Grok 4.5
  • Claude Sonnet 5
  • Minimax M3
  • GLM 5.2
  • GPT 5.1 Chat Latest
  • Claude Fable 5
  • Qwen3.7 Max
  • Claude Opus 4.8 Thinking
  • Claude Opus 4.8
  • Gemini 3.5 Flash
  • DeepSeek v4 Flash
  • DeepSeek v4 Pro
  • Grok 4.3
Explore all models >

Image

  • Dola Seedream 5.0 Pro 260628
  • Gemini 3.1 Flash Lite Image
  • Gemini 3.1 Flash Image
  • GPT Image 2
  • Gemini 3.1 Flash Image Preview
  • Seedream 5.0 260128
  • Doubao Seedream 5.0 260128
  • Viduq2
  • Grok Imagine Image
  • Kling Image O1
  • GPT Image 1.5
  • Seedream 4.5 251128
  • Doubao Seedream 4.5 251128
  • Grok Imagine 0.9
  • Gemini 3 Pro Image Preview
  • Qwen Image Lora
  • Qwen Image Plus Lora
  • Qwen Image Plus
  • Grok 4 Image
  • GPT Image 1 Mini
Explore all models >

Video

  • Kling v3.0 4k
  • Dreamina Seedance 2.0 Mini 260615
  • Kling v3 Omni 4k
  • Dreamina Seedance 2.0 Fast 260128
  • Dreamina Seedance 2.0 260128
  • Vidu2.0
  • Doubao Seedance 2.0 260128
  • Doubao Seedance 2.0 Fast 260128
  • Kling v3 Omni Pro
  • Kling v3 Omni Std
  • Viduq3 Turbo
  • Kling v3.0 Pro
  • Kling v3.0 Std
  • Viduq3 Pro
  • Kling v2.6 Std
  • Viduq2 Pro
  • Viduq2 Turbo
  • Viduq2 Pro Fast
  • Viduq2
  • Kling Video O1 Pro
Explore all models >

© 2026 Talent Tech Global Limited (Hong Kong) / Talent Tech Global LLC (US). All rights reserved.

  • About Us
  • Privacy Policy
  • Terms of Service
  • Sitemap