Text-to-Audio Studio
Designed specifically for large-scale projects such as audiobooks, podcasts, and long-form voiceovers. Unlike standard text-to-speech tools that handle short inputs, the studio allows you to manage, edit, and generate entire documents within a single workspace.
Key features
Bulk Import : Uploading entire books or texts directly as Word documents (.doc, .docx)
- Smart Segmentation : The system automatically divides the uploaded document into manageable lines or sections. This allows you to treat each sentence or paragraph as a separate unit and apply different settings or edits as needed.
Integrated Translation : You can enable the "Translation" option during project setup to automatically translate your entire uploaded document into over 100 languages before generating the audio.
Automated Production : Once your project is set up, you can generate audio for the entire project at once with a single click using the "Convert" button.
How to create a project in the studio
- Project Information : Give your project a title (e.g., My First Audiobook).
- Content source : Drag and drop your Word file.
- Text processing : Choose between:
- Basic : Fast processing of standard text.
- AI -Enhanced: Advanced processing (optimized for Arabic) to ensure correct pronunciation and diacritics throughout the entire document.
- Translation (optional): Enable the "Enable Translation" option if you want to convert the source text into another language.
- Create : Click on "Create Project." You will then be taken to the workspace, where you can review the clips and begin the batch generation process.
Learn more about creating an audiobook in a few quick steps by reading this article:
https://moknah.io/blog/2025/7/27/convert-book-to-audiobook-in-15-minutes/
The Professionals' Guide: Mastering AI-Generated Arabic Speech
Skills, tricks, and best practices
The ultimate goal of an AI voice generator is to produce the highest audio quality and the most natural performance while minimizing credit consumption. At Maknah, we provide you with advanced tools to create professional voiceovers. This guide covers the essential skills you need.
- The Art of Segmentation (Text Preparation)
AI models follow the structure of your text literally. The way you segment your text determines performance quality, pacing, and cost.
The Problem: The Text Block
If you paste a long paragraph without splitting it, the form treats it as a single line.
- The result : AI often rushes to finish long sentences, resulting in an unnatural speed.
- Cost : If a single word is incorrect, you must regenerate the entire paragraph, wasting credit.
Error: Random partitioning
If you press Enter randomly in the middle of a sentence, the AI follows your instructions and forces...
Pausing (using *sukūn*) at the end of each line.
- The consequence : This creates a disjointed, incoherent sound and increases grammatical errors, as the AI determines the vocalization based on the structure of the entire sentence.
The Professional Solution: Conscious Segmentation
- Rule No. 1 : Divide the text based on punctuation marks (commas, periods, semicolons).
- Rule No. 2: If the source text lacks punctuation marks, add them yourself, especially before conjunctions (such as "and" or "but").
- The golden rule : Read the text aloud. Wherever you naturally pause to catch your breath, start a new line.
- For long-form content : Use the text-to-speech studio; it is designed to handle full Word documents and automatically manages the segmentation for you.
- Forming and Processing
Question : Do I need to manually add diacritics to the entire text? Answer: No. Excessive manual diacritics place too many constraints on the model, weakening its ability to utilize self-diacritization (which it inherits from the original cloned human voice).
Solution A : AI-Enhanced Processing (Magic Button)
This feature (available in the interface) is your secret weapon. When selected, the system...
Automatically with:
- Contextual vocalization of Arabic text.
- Adding the necessary punctuation marks.
- Correction of typographical errors.
- Result : It resolves 99% of pronunciation issues without manual effort.
- Solution B : Strategic Manual Shaping
If you are using basic normalization, follow a minimal approach:
- What do you vocalize ? Only the words that are difficult for you to read, or words with double meanings.
- How : You do not need to add diacritics to the entire word. Simply add the vowel mark to the specific letter causing ambiguity.
- Example : To distinguish between *kutiba* (passive voice) and *kataba* (past tense verb), you simply need to add a *damma* to the first letter: *kutiba*.
- Text formatting rules
The template prefers clean, simple text. Follow these formatting rules:
- ✅ Convert numbers to words: The model may struggle with complex numbers (such as dates or currencies) or their grammatical parsing. Write "two thousand and twenty" instead of "2020".
- ✅ Use dashes (-): Acceptable for parenthetical clauses, although commas are preferred.
- ❌ No list points: Do not use bullet points.
- ❌ No character elongation: Avoid elongation/kashida (e.g., مـرحـبـاً).
- ❌ No slashes/underscores: avoid / or _
- ❌ Avoid ellipses (...): Do not use multiple dots to separate sentences unless you specifically intend to create a hesitant or trailing-off tone.
The foreign letters trick (P, G, V, CH)
Arabic lacks certain sounds such as P, G, or V; to ensure AI pronounces foreign names correctly within Arabic text, use Urdu/Persian characters:
| The required sound | The Letter | Writing in Arabic | Urdu/Persian – The Letter Trick |
| P | P | I drank a can of Pepsi. | I drank a can of Pepsi. |
| G | G | I searched for the information on Google. | I searched for the information on Google. |
| CH | Ch | This is a cheese sandwich. | This is a cheese sandwich. |
These letters are as follows:
🟢 The letter Pe (پ)
🔸 The sentence:
I traveled to Pakistan with my friend.
Replacement: Pakistan
🔹 Explanation: The letter “پ” is pronounced like "P," which is the correct pronunciation of the word "Pakistan ."
🟢 The letter Che (چ)
🔸 The sentence:
I saw delicious chai in India.
Instead of: tea (accusative) or tea (nominative/genitive)
🔹 Clarification: The word “چاي” is written this way in Urdu and pronounced “چاي,” like “chai.”
🟢 The letter ٹ (ṭe)
🔸 The sentence:
I worked in a tin factory.
Substitute: *tin*, to indicate the velarized sound in certain foreign names.
🔹 Explanation: “ٹ” is a retroflex ‘t’, commonly used in names such as “Saeed Teacher” (Mr. Saeed).
🟢 Letter ڊ (Dāl)
🔸 The sentence:
I met the doctor at the hospital.
Instead of: Dr.
🔹 Explanation: The letter “ڈ” is pronounced like a heavy (retroflex) ‘D’, and it is common in Urdu words adapted from English.
🟢 The letter ڑ (ṛe)
🔸 The sentence:
I lived in a village near the mountain.
Pahār = The mountain
🔹 Explanation: “ڑ” is a retroflex ‘r’ sound, used in words such as “پہاڑ” (mountain).
🟢 The letter Gāf (گ)
🔸 The sentence:
We went to Gujarat on a short trip.
Alternative: Measures
🔹 Clarification: The letter “گ” is pronounced like “G”; this applies to the pronunciation of the state name “Gujarat”.
🟢 The letter ے (Bari Ye)
🔸 The sentence:
I met Ali at the mosque.
Substitute: Ali
🔹 Explanation: “ے” is used at the end of nouns to indicate an *imāla* (inclined) or long *yā’*.
🟢 The letter ں (Nūn Ghunnah)
🔸 The sentence:
This is Ahmed from the city.
Representation of the nasal 'n' at the end of a word
🔹 Explanation: “ں” is pronounced as a light nasal ‘n’; it is common in Urdu, especially at the end of nouns.
- Performance and Emotional Control
- You can direct the AI to act like an actor rather than a news anchor. This is done through modes and settings.
- Level 1 : Mode Selection
- In the generation interface, you will see a checkbox: Enable Emotions & Dialects Mode.
- a. Standard mode (checkbox unchecked)
- Best for: professional narration, audiobooks, e-learning.
- How to control :
- Pauses: To enforce a silence, add this tag on a new line:
- Mild emotion: Use context or descriptive tags at the end of the line (experimental): I looked at the flowers...
b. Emotion and accent settings (checkbox selected)
- Best for : Drama, games, character dialogue.
- How to control :
- Emotion tags : Curly brackets { } must be used, with the emotion written in English at the beginning of the sentence.
- Format : {Emotion} Your Arabic text here.
- Examples: {Shouting} Shouting // {Whispering} Whispering // {Laughing} Laughing // {Sighs} Sighing.
Level 2: Audio Settings
Fine-tune the output using the sliders:
- Temperature :
- Less (rigid): stable, steady, slightly mechanical. Good for news.
- Higher (flexible): More creative and emotional. Warning: Raising it too high may cause instability or auditory hallucinations.
- Similarity :
- Controls how closely the audio resembles the original cloned version. Keep it high for accuracy.
- Expressiveness :
- It enhances emotional range. Use it with caution; for long texts, keep it low to maintain stability.
- Mastering dialects
- The golden rule: For Arabic dialects (Saudi, Egyptian, Jordanian, etc.), the audio must match the text.
- If you are writing in Egyptian colloquial Arabic, you should choose a voice trained in the Egyptian dialect.
- Do not try to force a voice that speaks Standard Arabic to speak in a colloquial dialect; the result will sound unnatural.
- The Priming Trick :
- To help the AI instantly recognize the dialect, start your text with a strong word from that dialect.
- Example: Instead of starting with a neutral phrase, begin with "How are you? How are things going?" This sets the model up to maintain an Egyptian tone for the rest of the paragraph.
Go to the next article in the Makina guide.