Real-time Text to Speech (TTS)
Stream text progressively and receive raw PCM or MP3 audio bytes instantly
Establishes a bidirectional WebSocket connection for real-time text-to-speech streaming. Send text chunks progressively (ideal for LLM outputs) and receive binary audio streams continuously. Credits are calculated identically to the HTTP endpoint.
To achieve ultra-low latency (~200ms Time-To-First-Byte), ensure your text chunks end with hard punctuation (periods ., exclamation marks !, or question marks ؟). The AI naturally buffers text until it detects a sentence boundary to ensure perfect human intonation.
Authentication
Authenticate using the Authorization header during the WebSocket handshake:
Authorization: Bearer YOUR_API_KEY
Connection Setup
After connecting, send an initialization message with your desired configuration:
{
"type": "init",
"voice_id": 1710,
"prerecording": 1,
"output_format": "pcm_24000",
"voice_settings": {
"temperature": 25,
"similarity": 100,
"expressiveness": 0,
"speed": 1.0,
"fast_mode": 2
}
}
Init Parameters
| Parameter | Required | Type | Description |
|---|---|---|---|
| type | required | string | Must be "init" for initialization message. |
| voice_id | required | integer | Voice ID from the voice list. |
| output_format | optional | string | The desired audio format. Defaults to pcm_24000. Options include:
|
| prerecording | optional | integer | 1 (normal, default) or 2 (AI enhancement). Affects credit cost. |
| voice_settings | optional | object | Voice customization options. See Voice Settings section below. |
Voice Settings Object
| Parameter | Type | Range | Default | Description |
|---|---|---|---|---|
| temperature | integer | 0 – 100 | 25 | Lower = stable and predictable, Higher = creative and varied. |
| similarity | integer | 0 – 100 | 100 | How closely the speech matches the original voice clone. |
| speed | float | 0.5 – 2.0 | 1.0 | Playback speed multiplier. |
| fast_mode | integer | 0, 1, or 2 | 0 |
Reduces latency significantly.
|
Streaming Pipeline
1. Send Text Chunks
Send text progressively as your LLM generates it:
{"type": "text", "data": "مرحبا بكم."}
The server will instantly process the chunk, deduct the specific credits for that length, and begin streaming audio back.
2. Signal Completion (Stop)
When your text generation is completely finished, send the stop signal:
{"type": "stop"}
The server will finish synthesizing the remaining text in its buffer, flush the final audio bytes to you, and send the End of Stream signal.
Receiving Messages
The server pushes two types of frames over the WebSocket:
1. Binary Audio Chunks (PCM / MP3)
Raw audio bytes streamed as fast as they are generated. The format depends on your requested output_format (defaults to 16-bit PCM at 24kHz). If you requested MP3, you can write these bytes directly to a file.
2. JSON Control Messages
Init Ready:
{"status": "ready", "message": "Initialized and ready to receive text."}
End of Stream (Safe Close Signal):
{"type": "end_of_stream"}
When you receive this message, it mathematically guarantees that 100% of the audio chunks have been delivered. You can now safely close the WebSocket connection on the client side.
Mid-Stream Error:
{"error": "Insufficient credits during stream"}
Production-Ready Examples
# pip install websockets
import asyncio
import websockets
import json
API_KEY = "YOUR_API_KEY"
url = "wss://api.moknah.io/api/v1/tts/ws/"
async def stream_tts():
headers = {"Authorization": f"Bearer {API_KEY}"}
async with websockets.connect(url, additional_headers=headers) as ws:
# 1. Initialize session (Requesting MP3 so we can save to file)
await ws.send(json.dumps({
"type": "init",
"voice_id": 1710,
"output_format": "mp3_44100_128",
"voice_settings": {"fast_mode": 2}
}))
print(f"Server: {await ws.recv()}")
# 2. Receiver Task
async def receive_audio():
with open("audio.mp3", "wb") as f:
async for msg in ws:
if isinstance(msg, bytes):
f.write(msg)
else:
data = json.loads(msg)
if data.get("type") == "end_of_stream":
print("✅ Stream finished safely.")
await ws.close()
break
receiver = asyncio.create_task(receive_audio())
# 3. Send Text
for chunk in ["مرحبا بكم.", "في منصة مُكنة."]:
await ws.send(json.dumps({"type": "text", "data": chunk}))
await asyncio.sleep(0.5)
# 4. Signal Stop
await ws.send(json.dumps({"type": "stop"}))
await receiver
asyncio.run(stream_tts())
// npm install ws
const WebSocket = require('ws');
const fs = require('fs');
const API_KEY = "YOUR_API_KEY";
const ws = new WebSocket('wss://api.moknah.io/api/v1/tts/ws/', {
headers: { "Authorization": `Bearer ${API_KEY}` }
});
const audioStream = fs.createWriteStream('audio.mp3');
ws.on('open', function open() {
// Requesting MP3 to safely write to an MP3 file
ws.send(JSON.stringify({
type: "init",
voice_id: 1710,
output_format: "mp3_44100_128",
voice_settings: { fast_mode: 2 }
}));
});
ws.on('message', function incoming(data, isBinary) {
if (isBinary) {
audioStream.write(data); // Save MP3 chunks
} else {
const msg = JSON.parse(data.toString());
if (msg.status === 'ready') {
console.log("Ready! Sending text...");
ws.send(JSON.stringify({ type: "text", data: "مرحبا بكم." }));
setTimeout(() => ws.send(JSON.stringify({ type: "text", data: "في منصة مُكة." })), 500);
setTimeout(() => ws.send(JSON.stringify({ type: "stop" })), 1000);
} else if (msg.type === 'end_of_stream') {
console.log("✅ All audio received. Closing connection.");
ws.close();
}
}
});
<?php
// composer require textalk/websocket
require 'vendor/autoload.php';
use WebSocket\Client;
$apiKey = "YOUR_API_KEY";
$client = new Client("wss://api.moknah.io/api/v1/tts/ws/", [
'headers' => ['Authorization' => "Bearer $apiKey"],
'timeout' => 60
]);
// 1. Initialize (Requesting MP3 format)
$client->text(json_encode([
"type" => "init",
"voice_id" => 1710,
"output_format" => "mp3_44100_128",
"voice_settings" => ["fast_mode" => 2]
]));
echo "Server: " . $client->receive() . "\n";
// 2. Send Text Chunks
$client->text(json_encode(["type" => "text", "data" => "مرحبا بكم."]));
$client->text(json_encode(["type" => "text", "data" => "في منصة مُكنة."]));
$client->text(json_encode(["type" => "stop"]));
// 3. Receive Audio & Control Messages
$audioFile = fopen("audio.mp3", "wb");
while (true) {
$message = $client->receive();
$decoded = json_decode($message, true);
if ($decoded && isset($decoded['type']) && $decoded['type'] == 'end_of_stream') {
echo "✅ Stream finished successfully.\n";
break;
} elseif (!$decoded) {
// If it can't be decoded as JSON, it's our binary MP3 audio chunk
fwrite($audioFile, $message);
}
}
fclose($audioFile);
$client->close();
?>
# Standard cURL does not support bidirectional WebSockets easily.
# Use 'websocat' to test WebSockets from the command line:
websocat wss://api.moknah.io/api/v1/tts/ws/ \
-H "Authorization: Bearer YOUR_API_KEY"
# Once connected, paste the JSON payloads interactively:
{"type": "init", "voice_id": 1710, "output_format": "mp3_44100_128"}
{"type": "text", "data": "مرحبا بكم."}
{"type": "stop"}
Close Codes & Errors
The WebSocket connection may close with one of these standard codes:
| Code | Type | Description |
|---|---|---|
1000 |
SUCCESS | Stream completed naturally. |
3011 |
GENERATION_ERROR | Audio generation failed on the upstream API. Retry the request. |
4001 |
MISSING_API_KEY | No API key provided in the Authorization header. |
4002 |
AUTH_FAILED | Invalid key, inactive account, insufficient initial credits, or IP restriction. |
4290 |
CONCURRENT_LIMIT | User already has an active TTS WebSocket connection. The Limit is 1 per user. |
For API-related questions or issues, contact us at api@moknah.io.