Endpoints
Supported Models
Request Parameters
Emotion enum
happy, sad, angry, fearful, disgusted, surprised, calm, fluent, whisper
Support depends on the model and voice combination; not all voice_id values support every emotion.
Sound effects enum
spacious_echo, auditorium_echo, lofi_telephone, robotic
Voice aliases
For OpenAI-style consistency, the platform maps common aliases to MiniMax voices:
When you use one of these aliases in
voice_id, the platform converts it to the actual MiniMax voice, and fills in the default emotion automatically when the voice has one and no emotion is specified.
Return format (audio_mode)
The audio return mode can be set via channel custom parameters:audio_mode = json(default): the response body is JSON, withdata.audioas hex or URL (determined byoutput_format).audio_mode = hex: returns a raw audio stream whenoutput_format=hex; still returns JSON whenoutput_format=url.
Synchronous example
output_format=hex, the response is a raw hex audio stream that can be written directly to an .mp3 file.
With output_format=url, the response is JSON; read the audio URL from data.audio.
Asynchronous example
voice_id and parameters in the JSON payload; if you used an alias, the platform converts it and fills in the emotion before the upstream request.
Billing
- Billed by the audio successfully synthesized; see the MiniMax speech entries under “Model Pricing” in the console for exact prices.
- Actual amounts follow the upstream consumption logs.