# MiniMax Speech 2.5 HD Preview Voice Clone — Text To Audio > Harness speech 2.5 voice cloning with 48kHz HD audio and low latency. Build real-time AI apps with zero-shot synthesis through the GPTProto.com unified API. ## Overview - **Endpoint**: `POST https://gptproto.com/api/v3/minimax/speech-2.5-hd-preview-voice-clone/text-to-audio` - **Result URL**: `GET https://gptproto.com/api/v3/predictions/{result_id}/result` — the submit response also returns the authoritative URL in `data.urls.get` - **Model ID**: `speech-2.5-hd-preview-voice-clone` - **Vendor**: MiniMax - **Scene**: `text-to-audio` - **Category**: text-to-audio - **Modalities**: input text → output audio - **Playground**: https://gptproto.com/model/minimax/speech-2.5-hd-preview-voice-clone - **API documentation**: https://docs.gptproto.com - **Other scenes of this model**: `voice-clone` — same auth, different endpoint path and input schema ## Authentication Every request needs a GPTProto API key in the `Authorization` header. Create one at https://gptproto.com/dashboard/api-key, then export it: ```bash export GPTPROTO_API_KEY="your-api-key" ``` Header: `Authorization: Bearer $GPTPROTO_API_KEY` (plus `Content-Type: application/json` on POST). ## Pricing Platform price by tier (USD, already includes the GPTProto discount): - **Per run** — $0.5003 per run Price range: $0.5003 – $0.5003 per generation. Prices may change. The model page always shows the live price: https://gptproto.com/model/minimax/speech-2.5-hd-preview-voice-clone ## API Information The API is asynchronous: POST the endpoint to create a prediction, then poll its result URL until `status` is `completed` or `failed`. ### Input Schema The endpoint accepts the following JSON body parameters: - **`audio`** (`string`, _required_): The uploaded file is cloned and supports formats such as MP3, M4A, and WAV. - **`custom_voice_id`** (`string`, _required_): Custom user-defined ID. Minimum 8 characters; must include letters and numbers and start with a letter. Duplicate voice-ids will throw an error. - **`need_noise_reduction`** (`boolean`, _optional_): Enable noise reduction. Default is false (no noise reduction). - Default: `false` - Options: true, false - **`need_volume_normalization`** (`boolean`, _optional_): Specify whether to enable volume normalization. If not provided, the default value is false. - Default: `false` - Options: true, false - **`accuracy`** (`range`, _optional_): Uploading this parameter will set the text validation accuracy threshold, with a value range of [0,1]. If not provided, the default value for this parameter is 0.7. - Default: `0.7` - Range: 0–1 - **`text`** (`string`, _optional_): Text for audio preview. Limited to 2000 characters. - Default: `Hello! Welcome to GPT Proto! This is a preview of your cloned voice. I hope you enjoy it!` **Required Parameters Example**: ```json { "audio": "", "custom_voice_id": "" } ``` **Full Example**: ```json { "audio": "", "custom_voice_id": "", "need_noise_reduction": false, "need_volume_normalization": false, "accuracy": 0.7, "text": "Hello! Welcome to GPT Proto! This is a preview of your cloned voice. I hope you enjoy it!" } ``` ### Output Schema Both submit and poll return the same envelope: - **`data.id`** (`string`): Prediction id. Use it as `result_id` when polling for the result. - **`data.status`** (`string`): One of `created`, `running`, `completed`, `failed`. See Status Values below. - **`data.outputs`** (`array of string`): Generated files. Empty until `status` is `completed`; then it holds the output URLs (or base64 strings when the model exposes a base64 option). - **`data.urls.get`** (`string`): Authoritative result URL for this prediction. Prefer it over building the poll URL yourself. - **`data.error`** (`string | null`): Failure reason when `status` is `failed`, otherwise `null`. - **`data.executionTime`** (`integer`): Total processing time in milliseconds. - **`data.timings.inference`** (`integer`): Model inference time in milliseconds. - **`data.hasNsfwContents`** (`array of boolean`): Per-output moderation flags. - **`message`** (`string`): `success` on a normal response, otherwise the error message. - **`code`** (`integer`): Business status code. `200` means the request was accepted. **Example Response — submit (task accepted)**: ```json { "data": { "id": "pred_example_01", "model": "speech-2.5-hd-preview-voice-clone", "outputs": [], "urls": { "get": "https://gptproto.com/api/v3/predictions/pred_example_01/result" }, "hasNsfwContents": [], "status": "created", "createdAt": "2026-01-01T12:00:00Z", "executionTime": 0, "timings": { "inference": 0 } }, "message": "success", "code": 200 } ``` **Example Response — poll (completed)**: ```json { "data": { "id": "pred_example_01", "model": "speech-2.5-hd-preview-voice-clone", "outputs": [ "https://oss-us.gptproto.com/example/output.mp3" ], "urls": { "get": "https://gptproto.com/api/v3/predictions/pred_example_01/result" }, "status": "completed", "error": null, "executionTime": 12345, "timings": { "inference": 12000 }, "has_nsfw_contents": [], "created_at": "2026-01-01T12:00:00Z" }, "message": "success", "code": 200 } ``` ### Status Values - `created` — Task accepted. Use `data.id` as `result_id` for polling. - `running` — Generation is in progress. Keep polling. - `completed` — Finished successfully. Generated files are in `data.outputs`. - `failed` — Generation failed. Read `data.error` for the reason. ## Usage Examples ### 1. Submit a request ```bash curl --request POST "https://gptproto.com/api/v3/minimax/speech-2.5-hd-preview-voice-clone/text-to-audio" \ --header "Authorization: Bearer $GPTPROTO_API_KEY" \ --header "Content-Type: application/json" \ --data '{ "audio": "", "custom_voice_id": "", "need_noise_reduction": false, "need_volume_normalization": false, "accuracy": 0.7, "text": "Hello! Welcome to GPT Proto! This is a preview of your cloned voice. I hope you enjoy it!" }' ``` ```python import os import requests url = "https://gptproto.com/api/v3/minimax/speech-2.5-hd-preview-voice-clone/text-to-audio" headers = { "Authorization": f"Bearer {os.environ['GPTPROTO_API_KEY']}", "Content-Type": "application/json" } payload = { "audio": "", "custom_voice_id": "", "need_noise_reduction": False, "need_volume_normalization": False, "accuracy": 0.7, "text": "Hello! Welcome to GPT Proto! This is a preview of your cloned voice. I hope you enjoy it!" } response = requests.request("POST", url, headers=headers, json=payload) print(response.json()) ``` ```typescript const response = await fetch("https://gptproto.com/api/v3/minimax/speech-2.5-hd-preview-voice-clone/text-to-audio", { method: "POST", headers: { "Authorization": `Bearer ${process.env.GPTPROTO_API_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ audio: "", custom_voice_id: "", need_noise_reduction: false, need_volume_normalization: false, accuracy: 0.7, text: "Hello! Welcome to GPT Proto! This is a preview of your cloned voice. I hope you enjoy it!", }), }); const data = await response.json(); console.log(data); ``` ### 2. Poll until the task finishes Replace `YOUR_RESULT_ID` with `data.id` from the submit response (or call `data.urls.get` directly), and keep polling every 1–3 seconds while `status` is `created` or `running`. ```bash result_id="YOUR_RESULT_ID" curl --request GET "https://gptproto.com/api/v3/predictions/$result_id/result" \ --header "Authorization: Bearer $GPTPROTO_API_KEY" ``` ```python import os import requests result_id = "YOUR_RESULT_ID" url = f"https://gptproto.com/api/v3/predictions/{result_id}/result" headers = { "Authorization": f"Bearer {os.environ['GPTPROTO_API_KEY']}" } response = requests.request("GET", url, headers=headers) print(response.json()) ``` ```typescript const result_id = "YOUR_RESULT_ID"; const response = await fetch(`https://gptproto.com/api/v3/predictions/${result_id}/result`, { method: "GET", headers: { "Authorization": `Bearer ${process.env.GPTPROTO_API_KEY}`, }, }); const data = await response.json(); console.log(data); ``` ## Example Prompts - Hello everyone. We're thrilled to introduce the next generation of our voice model: MiniMax Speech 2.5. Building on its predecessor, Speech 2.0, this new version is more powerful than ever. But where it truly shines is in its incredible realism. The model masterfully captures the… ## HTTP Status Codes - `400` — malformed body, a parameter or value this model does not accept, or input blocked by the provider's content moderation - `401` — missing or invalid API key - `403` — insufficient credits - `413` — request body too large - `429` — rate limited; retry with backoff - `500` — An internal server error occurred - `502` — An internal server error occurred - `504` — Gateway timeout — upstream service did not respond in time; retry later ## Additional Resources - [Model playground](https://gptproto.com/model/minimax/speech-2.5-hd-preview-voice-clone) - [API documentation](https://docs.gptproto.com) - [All models](https://gptproto.com/model) - [API keys](https://gptproto.com/dashboard/api-key) - [Platform overview for LLMs](https://gptproto.com/llm-full.txt) - Any other model: `https://gptproto.com/model/{vendor}/{model}/{scene}/llms.txt`