How to create an audio file
The creation process is simple:
- Enter text : Type or paste your text into the input box.
Note : Text input is crucial. High-quality input leads to high-quality output.
- Select Voice : Choose the voice you wish to use from your list of voices.
Tip : If you want a happy voice, choose one cloned from happy samples; if you want a specific accent, choose a voice trained on that accent.
- Adjusting Settings : Configure settings such as speed or the creativity factor (detailed under Sound and Performance Control).
- Create : Click the " Generate " button to create your audio file.
Audio generation modes and scenarios
The most important step is deciding whether to check the box for enabling emotion and dialect modes, as this will alter how the AI interprets your text.
Scenario A: Expressive and personal content (with emotional enablement)
Ideal for games, animation, and audiobooks, where an emotional range and interaction between characters are required.
- Capabilities : Generating natural and realistic dialogue with a wide emotional range (laughter, whispering, etc.).
- Languages : Supports over 70 languages, including: Afrikaans, Arabic, Armenian, Assamese, Azerbaijani, Belarusian, Bengali, Bosnian, Bulgarian, Catalan, Cebuano, Chichewa, Croatian, Czech, Danish, Dutch, English, Estonian, Filipino, Finnish, French, Galician, Georgian, German, Greek, Gujarati, Hausa, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Irish, Italian, Japanese, Javanese, Kannada, Kazakh, Kyrgyz, Korean, Latvian, Lingala, Lithuanian, Luxembourgish, Macedonian, Malay, Malayalam, Mandarin Chinese, Marathi, Nepali, Norwegian, Pashto, Persian, Polish, Portuguese, Punjabi, Romanian, Russian, Serbian, Sindhi, Slovak, Slovenian, Somali, Spanish, Swahili, Swedish, Tamil, Telugu, Thai, Turkish, Ukrainian, Urdu, Vietnamese, Welsh.
- Limitations :
- Maximum number of characters per request: 3,000
- Approximate duration: 3 minutes
Scenario B: Professional and stable content (emotions disabled)
Perfectly suited for corporate videos, e-learning materials, and scenarios requiring consistent audio quality.
- Capabilities : Delivers consistent voice quality and persona while preserving the speaker's unique characteristics.
- Languages : Supports 29 languages: English (USA, UK, Australia, Canada), Japanese, Chinese, German, Hindi, French (France, Canada), Korean, Portuguese (Brazil, Portugal), Italian, Spanish (Spain, Mexico), Indonesian, Dutch, Turkish, Filipino, Polish, Swedish, Bulgarian, Romanian, Arabic (Saudi Arabia, UAE), Czech, Greek, Finnish, Croatian, Malay, Slovak, Danish, Tamil, Ukrainian, Russian.
- Limitations :
- Maximum characters per request: 1,500
- Approximate duration: ~1.5 minutes
Go to the next article in the Makina guide.