Skip to main content
Use the MiniMax text-to-speech (TTS) capability with both synchronous and asynchronous calling styles. All requests are forwarded to MiniMax by GregAPI — no self-signing required, just pass the GregAPI API Key in the header.
Standard headers:

Endpoints

Supported Models

Request Parameters

Emotion enum

happy, sad, angry, fearful, disgusted, surprised, calm, fluent, whisper Support depends on the model and voice combination; not all voice_id values support every emotion.

Sound effects enum

spacious_echo, auditorium_echo, lofi_telephone, robotic

Voice aliases

For OpenAI-style consistency, the platform maps common aliases to MiniMax voices: When you use one of these aliases in voice_id, the platform converts it to the actual MiniMax voice, and fills in the default emotion automatically when the voice has one and no emotion is specified.

Return format (audio_mode)

The audio return mode can be set via channel custom parameters:
  • audio_mode = json (default): the response body is JSON, with data.audio as hex or URL (determined by output_format).
  • audio_mode = hex: returns a raw audio stream when output_format=hex; still returns JSON when output_format=url.

Synchronous example

With output_format=hex, the response is a raw hex audio stream that can be written directly to an .mp3 file. With output_format=url, the response is JSON; read the audio URL from data.audio.

Asynchronous example

Query the result:
Async synthesis preserves your voice_id and parameters in the JSON payload; if you used an alias, the platform converts it and fills in the emotion before the upstream request.

Billing

  • Billed by the audio successfully synthesized; see the MiniMax speech entries under “Model Pricing” in the console for exact prices.
  • Actual amounts follow the upstream consumption logs.