Over the past two years, the "Listen to this article" button has evolved from a rarity into a near-standard feature on major news websites—both in the Arab world and globally. Yet, a question rarely asked with sufficient seriousness remains: Is this investment truly justified? Or is it simply another tech trend that publishers have jumped on because everyone else is doing it?
The honest answer: It relies on questions that most of those who raise them have not answered.
The Case for Audio: Why Major Sites Are Turning to It
The commercial rationale is not an illusion; a number of major global media organizations have developed real-world use cases:
The Economist doubled its monthly podcast audience from approximately 2.5 million to 5 million listeners between 2022 and 2025, prompting the launch of an audio-specific subscription service; meanwhile, Bloomberg recorded that its app users listen to an average of six stories per session—a level of engagement difficult to achieve through text alone. Additionally, The Times Ireland turned audio listening into a direct gateway for converting free readers into subscribers by requiring them to log in upon clicking the listening icon.
The most prominent Arab example is Al Jazeera Net , which set a regional precedent by implementing audio listening features on a large scale across its articles, thereby establishing a benchmark against which the rest of the Arab market is measured today.
But... beware of the figures being circulated.
Here, any serious consultant must pause at a point often obscured by marketing: most of the "legendary" figures cited by the TTS industry are outdated.
A recent critical review (July 2026) of the evidence base underpinning the journalism audio industry concluded that virtually all prominent figures—such as threefold higher engagement, 80% audio consumption, and listening completion rates nearing 50%—date back to the 2018–2022 period, with no documented updates found for 2025 or 2026. More concerning still is that the primary promoters of these figures are the TTS technology vendors themselves, who have a direct commercial interest in inflating them.
Furthermore, the famous Washington Post case—which saw three times higher engagement—dates back to the era before AI-driven, near-free voice synthesis; the content was produced using paid human voices or more costly synthesis methods. This introduces a "selection effect," making the prospect of replicating those results with today’s cheap, automated generation tools seem far from guaranteed to decision-makers at news or cultural outlets.
The professional takeaway here is that voice technology is worth exploring; however, any investment decision should be based on direct measurements of your own site's performance, rather than on figures from use cases in completely different industries and markets.
The insufficiently discussed challenge: Arabic is not English.
This is the most important point overlooked in most superficial discussions about TTS in the Arab market: it is a matter that is purely technical rather than marketing-related.
The issue of diacritics lies at the very heart of the problem, rather than being a marginal detail. Arabic text in the daily press overwhelmingly lacks diacritics, and a decade’s worth of academic research documents that this absence renders the text susceptible to genuine ambiguity regarding both meaning and pronunciation—a factor that directly impacts the quality of any text-to-speech engine lacking an accurate automatic diacritization layer.
In practical terms , a word without diacritics can be read in various ways—differing in both meaning and pronunciation—and a Text-to-Speech (TTS) engine that fails to resolve this ambiguity with genuine linguistic intelligence will produce audio that appears superficially "correct" yet is grammatically or semantically flawed to a trained Arabic ear.
This remains an active area of research: modern automatic diacritization models published in peer-reviewed studies—spanning 2024 through early 2025 and 2026—continue to compete to reduce error rates, with some even outperforming general-purpose large language models. This indicates that accurate automatic diacritization for Arabic is a specialized challenge that cannot be solved by a generic Text-to-Speech (TTS) engine originally built for English and subsequently "patched" with Arabic support.
Additional challenges specific to Arabic that are worth mentioning in any serious assessment:
- The Modern Standard Arabic/Colloquial duality: Arab readers expect a specific news-style tone in Standard Arabic, whereas content in a dialect (whether marketing, cultural, or interactive) requires a completely different engine—one that does not conflate Standard and colloquial pronunciations.
- The *tā’ marbūṭah* and contextual pausing: pronunciation rules shift depending on the word’s position within the sentence (mid-sentence versus at the end, before a punctuation mark). This is a minor detail, yet it instantly exposes any engine not calibrated to the Arab ear.
- Foreign names and embedded terms: Arabic news texts are replete with proper names and Latin-script terms interspersed within Arabic sentences; this requires an engine capable of transitioning smoothly between two phonetic systems without disrupting the rhythm.
But what if this particular challenge were solved?
None of the above is a reason to abandon audio; rather, it explains exactly why some early attempts in the Arab market failed to yield the expected results. The issue did not lie with the concept of audio itself, but with an engine that was never built to master such details.
Let us pose the question differently: What if there were an engine that resolved the challenge of diacritics with genuine linguistic precision rather than mere approximation? One that distinguishes between Modern Standard Arabic (as used in news) and colloquial speech when necessary, transitions seamlessly between foreign names and Arabic text without disrupting the rhythm, and automatically handles the pronunciation of the *ta’ marbuta* and contextual pauses correctly? And most importantly: What if its vocal delivery avoided a cold, robotic feel, instead adopting a natural, engaging news-reporting tone that made the listener feel they were hearing a real broadcast rather than a machine reading a text?
At that point, the equation changes completely, and this is precisely what practical experience on the ground demonstrates.
What actual usage data reveals
There is an important point worth clarifying here: the common assumption is that user enthusiasm for listening wanes over time following the initial impression. Our experience with clients at Makna suggests quite the opposite— the number of users listening to articles rises cumulatively over time , rather than declining. When the audio performance is precise and pleasing to the Arab ear from the very first encounter, there is no "try-and-drop" scenario; instead, a listening habit is formed that grows stronger week by week—much like what has occurred with websites that made listening an integral part of the reading experience, rather than a mere peripheral add-on.
This makes sense when you think about it: a reader doesn’t simply "try" listening once and then make a decision—they gradually build trust in the audio quality. With every successful experience (accurate pronunciation, natural delivery), the likelihood of them returning increases, until listening becomes a habitual daily routine rather than a conscious choice requiring reconsideration each time.
Executive Summary for Decision-Makers
The real debate isn't whether voice technology is useful—global evidence indicates it is effective for retention and loyalty. The real question is: Is the engine you are using truly built to handle the challenges of the Arabic language, or is it a general-purpose engine simply presented with an Arabic interface?
The difference between an engine that translates Arabic and one that understands it is the difference between a feature that is added and forgotten, and one that becomes a daily habit the reader relies on. With an engine that is both highly accurate in news delivery and engaging, the question is not "Is it worth the investment?" but rather "How much time might we lose before we start?"
If you want to be sure, try one machine and then another, and listen to the difference yourself.