User Guide API Reference

Real-time Text to Speech (TTS)

Stream text progressively and receive raw PCM or MP3 audio bytes instantly

WSS /api/v1/tts/ws/

Establishes a bidirectional WebSocket connection for real-time text-to-speech streaming. Send text chunks progressively (ideal for LLM outputs) and receive binary audio streams continuously. Credits are calculated identically to the HTTP endpoint.

⚡ Latency Optimization (TTFB)

To achieve ultra-low latency (~200ms Time-To-First-Byte), ensure your text chunks end with hard punctuation (periods ., exclamation marks !, or question marks ؟). The AI naturally buffers text until it detects a sentence boundary to ensure perfect human intonation.

Authentication

Authenticate using the Authorization header during the WebSocket handshake:

Authorization: Bearer YOUR_API_KEY

Connection Setup

After connecting, send an initialization message with your desired configuration:

{
  "type": "init",
  "voice_id": 1710,
  "prerecording": 1,
  "output_format": "pcm_24000",
  "voice_settings": {
    "temperature": 25,
    "similarity": 100,
    "expressiveness": 0,
    "speed": 1.0,
    "fast_mode": 2
  }
}

Init Parameters

Parameter Required Type Description
type required string Must be "init" for initialization message.
voice_id required integer Voice ID from the voice list.
output_format optional string The desired audio format. Defaults to pcm_24000. Options include:
  • MP3: mp3_22050_32, mp3_44100_32, mp3_44100_64, mp3_44100_96, mp3_44100_128, mp3_44100_192
  • PCM: pcm_8000, pcm_16000, pcm_22050, pcm_24000, pcm_44100
For zero-latency app streaming, use PCM. For saving files, request MP3.
prerecording optional integer 1 (normal, default) or 2 (AI enhancement). Affects credit cost.
voice_settings optional object Voice customization options. See Voice Settings section below.

Voice Settings Object

Parameter Type Range Default Description
temperature integer 0 – 100 25 Lower = stable and predictable, Higher = creative and varied.
similarity integer 0 – 100 100 How closely the speech matches the original voice clone.
speed float 0.5 – 2.0 1.0 Playback speed multiplier.
fast_mode integer 0, 1, or 2 0 Reduces latency significantly.
  • 0: Disabled (highest quality)
  • 1: Flash mode (medium latency reduction)
  • 2: Turbo mode (maximum latency reduction)

Streaming Pipeline

1. Send Text Chunks

Send text progressively as your LLM generates it:

{"type": "text", "data": "مرحبا بكم."}

The server will instantly process the chunk, deduct the specific credits for that length, and begin streaming audio back.

2. Signal Completion (Stop)

When your text generation is completely finished, send the stop signal:

{"type": "stop"}

The server will finish synthesizing the remaining text in its buffer, flush the final audio bytes to you, and send the End of Stream signal.

Receiving Messages

The server pushes two types of frames over the WebSocket:

1. Binary Audio Chunks (PCM / MP3)

Raw audio bytes streamed as fast as they are generated. The format depends on your requested output_format (defaults to 16-bit PCM at 24kHz). If you requested MP3, you can write these bytes directly to a file.

2. JSON Control Messages

Init Ready:

{"status": "ready", "message": "Initialized and ready to receive text."}

End of Stream (Safe Close Signal):

{"type": "end_of_stream"}

When you receive this message, it mathematically guarantees that 100% of the audio chunks have been delivered. You can now safely close the WebSocket connection on the client side.

Mid-Stream Error:

{"error": "Insufficient credits during stream"}

Production-Ready Examples

# pip install websockets
import asyncio
import websockets
import json

API_KEY = "YOUR_API_KEY"
url = "wss://api.moknah.io/api/v1/tts/ws/"

async def stream_tts():
    headers = {"Authorization": f"Bearer {API_KEY}"}
    async with websockets.connect(url, additional_headers=headers) as ws:
        # 1. Initialize session (Requesting MP3 so we can save to file)
        await ws.send(json.dumps({
            "type": "init",
            "voice_id": 1710,
            "output_format": "mp3_44100_128",
            "voice_settings": {"fast_mode": 2}
        }))
        print(f"Server: {await ws.recv()}")

        # 2. Receiver Task
        async def receive_audio():
            with open("audio.mp3", "wb") as f:
                async for msg in ws:
                    if isinstance(msg, bytes):
                        f.write(msg)
                    else:
                        data = json.loads(msg)
                        if data.get("type") == "end_of_stream":
                            print("✅ Stream finished safely.")
                            await ws.close()
                            break

        receiver = asyncio.create_task(receive_audio())

        # 3. Send Text
        for chunk in ["مرحبا بكم.", "في منصة مُكنة."]:
            await ws.send(json.dumps({"type": "text", "data": chunk}))
            await asyncio.sleep(0.5)

        # 4. Signal Stop
        await ws.send(json.dumps({"type": "stop"}))
        await receiver

asyncio.run(stream_tts())
// npm install ws
const WebSocket = require('ws');
const fs = require('fs');

const API_KEY = "YOUR_API_KEY";
const ws = new WebSocket('wss://api.moknah.io/api/v1/tts/ws/', {
    headers: { "Authorization": `Bearer ${API_KEY}` }
});

const audioStream = fs.createWriteStream('audio.mp3');

ws.on('open', function open() {
    // Requesting MP3 to safely write to an MP3 file
    ws.send(JSON.stringify({
        type: "init",
        voice_id: 1710,
        output_format: "mp3_44100_128",
        voice_settings: { fast_mode: 2 }
    }));
});

ws.on('message', function incoming(data, isBinary) {
    if (isBinary) {
        audioStream.write(data); // Save MP3 chunks
    } else {
        const msg = JSON.parse(data.toString());
        if (msg.status === 'ready') {
            console.log("Ready! Sending text...");
            ws.send(JSON.stringify({ type: "text", data: "مرحبا بكم." }));
            setTimeout(() => ws.send(JSON.stringify({ type: "text", data: "في منصة مُكة." })), 500);
            setTimeout(() => ws.send(JSON.stringify({ type: "stop" })), 1000);
        } else if (msg.type === 'end_of_stream') {
            console.log("✅ All audio received. Closing connection.");
            ws.close();
        }
    }
});
<?php
// composer require textalk/websocket
require 'vendor/autoload.php';
use WebSocket\Client;

$apiKey = "YOUR_API_KEY";
$client = new Client("wss://api.moknah.io/api/v1/tts/ws/", [
    'headers' => ['Authorization' => "Bearer $apiKey"],
    'timeout' => 60
]);

// 1. Initialize (Requesting MP3 format)
$client->text(json_encode([
    "type" => "init",
    "voice_id" => 1710,
    "output_format" => "mp3_44100_128",
    "voice_settings" => ["fast_mode" => 2]
]));
echo "Server: " . $client->receive() . "\n";

// 2. Send Text Chunks
$client->text(json_encode(["type" => "text", "data" => "مرحبا بكم."]));
$client->text(json_encode(["type" => "text", "data" => "في منصة مُكنة."]));
$client->text(json_encode(["type" => "stop"]));

// 3. Receive Audio & Control Messages
$audioFile = fopen("audio.mp3", "wb");
while (true) {
    $message = $client->receive();
    $decoded = json_decode($message, true);

    if ($decoded && isset($decoded['type']) && $decoded['type'] == 'end_of_stream') {
        echo "✅ Stream finished successfully.\n";
        break;
    } elseif (!$decoded) {
        // If it can't be decoded as JSON, it's our binary MP3 audio chunk
        fwrite($audioFile, $message);
    }
}
fclose($audioFile);
$client->close();
?>
# Standard cURL does not support bidirectional WebSockets easily.
# Use 'websocat' to test WebSockets from the command line:

websocat wss://api.moknah.io/api/v1/tts/ws/ \
  -H "Authorization: Bearer YOUR_API_KEY"

# Once connected, paste the JSON payloads interactively:
{"type": "init", "voice_id": 1710, "output_format": "mp3_44100_128"}

{"type": "text", "data": "مرحبا بكم."}

{"type": "stop"}

Close Codes & Errors

The WebSocket connection may close with one of these standard codes:

Code Type Description
1000 SUCCESS Stream completed naturally.
3011 GENERATION_ERROR Audio generation failed on the upstream API. Retry the request.
4001 MISSING_API_KEY No API key provided in the Authorization header.
4002 AUTH_FAILED Invalid key, inactive account, insufficient initial credits, or IP restriction.
4290 CONCURRENT_LIMIT User already has an active TTS WebSocket connection. The Limit is 1 per user.
API Support

For API-related questions or issues, contact us at api@moknah.io.