Best Practices for Input
Numbers
- Numbers : Text-to-speech models may struggle with formats such as phone numbers or currencies.
- Recommendation : Write numbers as words to avoid ambiguity (for example, write "one" instead of "1").
- Phone numbers : 123-456-7890 ← one, two, three; four, five, six...
- Currencies : $45.67 ← Forty-five dollars and sixty-seven cents.
- Abbreviations : Expand them and write them out in full for clarity.
- Dr. ← Doctor, St. ← Street.
- Links/Symbols :
- moknah.io/docs ← moknah dot io slash docs (type it exactly as
is pronounced).
- 100% ← One hundred percent.
Overview
Generating your first audio file using text-to-audio or Text to Audio Studio is very simple. However, to get the most out of these features, there are a few things to keep in mind.
The Guide
1
Type or paste your text into the input box on the text-to-speech page.
You can always include expressive tags in your text to achieve superior speech synthesis performance.
- If emotions are enabled , use tags enclosed in brackets, such as [shouting] or [breathing]. These tags are highly effective (refer to the Voice Tags section for the full list) .
- If emotions are not enabled , you can still describe the desired performance (such as writing "excited") or use tags enclosed in , although these tags are experimental and may not work as consistently.
- Note : All tags must be in English, as using them in other languages is experimental and may not work reliably. We are working on simplifying this process by using emojis and buttons in future updates.
- Character limits and long-form content: The standard audio generation box is designed for short clips and accommodates only 1,500 characters (we recommend ~200 characters per generation for optimal quality).
- For large volumes: If you want to generate continuous audio as a single piece—such as for audiobooks or long articles—do not use this input box. Instead, use the Text to Audio Studio. The Studio allows you to upload entire documents and efficiently automate the production of long-form content.
2
Voice selection
Select the sound you wish to use from your list.
Choosing the right voice is important; for instance, if you want a voice that speaks Modern Standard Arabic, it is best to select one designed for that purpose from the start, rather than one tailored to a colloquial dialect. However, this does not mean that a voice trained on a specific dialect cannot read text in Standard Arabic; it certainly can, and may even deliver it with great liveliness. That said, if you increase the emotion parameter or the temperature, you might encounter unexpected results, as the voice will tend to drift toward the dialect on which it was originally trained.
For instance, suppose you want a performance in a specific dialect. In this case, if you choose not to activate the emotion parameter but select a voice with a colloquial style, increasing expressiveness and warmth will bring the performance closer to your desired dialect. Ultimately, the vocal performance is shaped by the input context; if the model recognizes the input dialect, it will attempt to emulate and deliver it, even if the selected voice is not specialized in dialects. Always remember that the more you practice, the better you will understand the balance achievable through your knowledge of the settings and how to manage them, leading to impressive results.
3
Adjusting settings (optional)
Adjust the audio settings to obtain the desired output.
The recommended default setting is shown in the image below:
Settings :
- Creativity Parameter (Temperature):
- Higher = more flexibility
- Less = more rigid
- Increasing the "temperature" gives the AI model greater freedom to control the audio, but raising the level too high risks producing strange results; conversely, setting it to zero makes the vocal performance stiff and lifeless. Therefore, finding a suitable middle ground that meets your needs is the best approach.
- Similarity: Similarity
- Highest = closest to the original sound
- Less = less similar to the original sound
- Increasing the matching factor enhances the resemblance to the original source audio; however, setting it to the maximum level when the reference sample is not clean or of high quality risks introducing audio artifacts, such as crackling. You can set it to the maximum if you require an exact match and the sample is clean; otherwise, lower the setting to achieve a cleaner, superior result.
- Speed: Speed
- Higher = Faster
- Less = slower
- Increasing the speech speed to the maximum or reducing it to the minimum may compromise the quality of the generated audio; therefore, we recommend using this option only when necessary.
- Expressiveness : Expressiveness
- Highest = Emotional delivery
- Less = neutral tone
- Increasing the emotion coefficient brings the output closer to a natural performance, but raising it excessively risks introducing unusual behavior; therefore, when generating audio in large volumes—particularly in Voice Studio—we recommend keeping it at zero and increasing it experimentally for specific sentences.
Emotion Settings (Emotion Setting slides):
The recommended default setting is shown in the adjacent image:
If you choose to enable emotion (Enable Emotion), as shown in the image above, the recommended setting is to increase the temperature (Temperature).
4
Select the desired normalization method.
- Basic : Suitable for Arabic without any additions, and suitable for languages other than Arabic.
- AI- enhanced : Specifically designed for the Arabic language, it understands the context, segments sentences or paragraphs, and applies the necessary diacritics for proper pronunciation.
5
Generate
Click the " Run " button on the Text to Audio page or the " Generate " button in Text to Audio Studio to create your audio file. In the Studio, you can generate thousands of lines at once by clicking the "Merge" button.
Go to the next article in the Makina guide.