Sound and Performance Control
Directing and controlling emotions
When you select the emotion-enabling box, you can control the mode of expression using the context and specified tags.
- Contextual Guidance : The model interprets the context directly from the text. Adding descriptive text—such as "she said enthusiastically"—or using exclamation marks will influence the emotion of the speech.
- Audio markers (related to sound): Place these markers within the text to guide the audio.
Note : The effectiveness depends on the specific sound.
- [laughs], [laughs harder], [starts laughing], [wheezing]
- [whispers]
- [sighs], [exhales]
- [sarcastic], [curious], [excited], [crying], [snorts], [mischievously]
- Sound effects :
- [gunshot], [applause], [clapping], [explosion]
- [swallows], [gulps]
- Experimental marks :
- [strong X accent] (replace X with the desired accent, e.g., [strong Lebanese accent])
- [sings], [woo], [fart]
- Punctuation Marks and Text Structure :
- The ellipsis (…): to add pauses and weight.
- Uppercase letters: for added emphasis.
- Standard punctuation: to provide a natural speech rhythm.
Important note : The emotional model is in the alpha stage. Very short prompts are likely to result in inconsistent output. We encourage you to try prompts longer than 250 characters. When the emotion-enabling option is unchecked, the system prioritizes stability and consistency.
- How to issue instructions :
- Context is key : rely on the natural flow of the sentence. The AI will read the text exactly as written.
- Separators : Use marks like these.
- Syntax:
- Note : Do not use too many markers, as this may cause instability.
- Note : It may or may not work, depending on the context. Try increasing the expressiveness and temperature when using it.
- Punctuation : Use dashes (-) for short pauses or ellipses (...) for hesitant tones.
- No emotional markers : Do not use curly braces {} or vocal cues here; the AI is likely to read them aloud or ignore them.
Audio Settings (Advanced Controls)
Adjust these parameters to fine-tune the output.
- Reliability coefficient.
- Range : More rigid (lower) ↔ More flexible (higher).
- Effect : A higher temperature gives the AI more freedom and expressiveness but risks instability (hallucinations). A lower temperature makes the output more rigid and consistent.
- Recommendation : The default value (around 25–50) is a good starting point.
- Resemblance
- Range : Less like a voiceover (lower) ↔ Like a voiceover (higher).
- Effect : Increasing this enhances similarity to the original version. However, if the reference audio contains noise, high similarity may result in the reproduction of distortions or crackling sounds.
- Speed
- Range : 0.7 (slower) ↔ 1.2 (faster). The default value is 1.0.
- Effect : Controls the pace. Extreme values may affect quality.
- Expressionism
- Range : Neutral tone (lower) ↔ Emotional performance (higher).
- Effect : Increasing this brings the output closer to natural performance. Raising it excessively risks erratic behavior.
Go to the next article in the Makina guide.
خ