48kHz HD Audio Output
Experience superior clarity with 48kHz sampling. This high-definition text to speech output eliminates aliasing, making it ideal for professional dubbing and media.
Input
Output
curl --request POST "https://gptproto.com/api/v3/minimax/speech-02-hd/text-to-audio" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"text": "Welcome to our advanced text-to-speech system! Experience high-quality voice synthesis with natural pronunciation and clear articulation.",
"voice_id": "Wise_Woman",
"speed": 1,
"volume": 1,
"pitch": "0",
"emotion": "happy",
"english_normalization": false,
"sample_rate": 8000,
"bitrate": 32000,
"channel": 1,
"format": "mp3",
"language_boost": "English",
"enable_sync_mode": ""
}'Start from the cost of a single sample and pick a testing budget. GPTProto rates are 40% below list price.
GPTProto · Price est.
Estimated from the rate card. Final charges may vary.No extra settings
Top up
GPTProto vs official pricing.Save$66.66 (40%)vs MiniMax official
Advanced technical capabilities that make speech-02-hd a leader in the audio generation market.
Experience superior clarity with 48kHz sampling. This high-definition text to speech output eliminates aliasing, making it ideal for professional dubbing and media.
Achieve near-instant response times. The speech-02-hd model is optimized for low-latency text processing, ensuring your conversational AI feels responsive and human.
The 02 model adds realistic breaths, laughter, and sighs. It transforms flat text into emotionally nuanced speech, achieving a industry-leading 4.62 MOS score.
Replicate any voice with just a 3-second sample. Text Speech 02 maintains the original speaker's timbre and rhythm across all supported languages and text inputs.
Expert answers regarding the Text Speech 02 (speech-02-hd) model capabilities, integration, and technical performance for your audio projects.
Guides, comparisons, and updates related to this model.
All Articles
Learn about GPT-4o Mini TTS, OpenAI's text-to-speech model that provides natural-sounding voices, emotional expression, and fast response times.

Master high-fidelity voice synthesis with minimax speech 02. Learn to build low-latency, emotional AI audio applications today.

Instantly convert audio to text with GPT-4o transcribe. Learn how to access this game-changing AI, its practical uses, and its affordable pricing.

Claude Mythos is a step change in AI performance. Learn why its reasoning and cyber capabilities have the industry on alert. Get the full breakdown.