Real-time Speech to Text (STT)
Transcribe streaming live audio to text in real-time using WebSockets
Stream raw audio to our servers and receive bidirectional, real-time transcriptions. The server responds with partial results (live text as the user speaks) and final results (completed sentences with word-level timestamps).
Because streams are live and unpredictable, we use a prepaid meter. Upon connection, we immediately reserve 83 credits (paying for the first 60 seconds). When you disconnect, we calculate your exact usage to the second and automatically refund the unused credits back to your account!
Authentication & Connection
You must authenticate your WebSocket connection. The method depends on your client environment:
- Frontend Browsers: Standard browser WebSockets do not support custom HTTP headers. You must include your API key as a query parameter (?api_key=YOUR_KEY).
- Backend / Python / Node.js: You may use the standard HTTP header Authorization: Bearer YOUR_API_KEY.
| Query Parameter | Required | Type | Description |
|---|---|---|---|
| api_key | Conditional | string | Your Moknah API key. (Required here if not passed via the Authorization header). |
| language | Optional | string | The language code to transcribe. Default is ar-JO. Supported codes include ar-SA, ar-EG, en-US, etc. |
Audio Specifications
Do NOT send compressed audio files (like .mp3, .m4a) or standard .wav files with headers over the WebSocket. The AI will output random noise or empty strings. The WebSocket strictly accepts Raw Binary PCM frames.
Before sending binary frames over the socket, ensure your audio stream matches exactly:
- Sample Rate: 16,000 Hz (16 kHz)
- Bit Depth: 16-bit integer (Int16)
- Channels: 1 (Mono)
- Chunk Size: Send in small chunks (e.g., 100ms / 3200 bytes). Chunks larger than 256KB will be rejected.
Server Responses (JSON Events)
As you stream binary audio to the server, the server will reply asynchronously with JSON text frames.
| Event Type | Payload Example |
|---|---|
|
Connection Success Sent immediately upon successful authentication. |
{"status": "connected", "message": "Ready to stream audio..."} |
|
Partial Result Sent continuously. Used to render live, typing subtitles. |
{"type": "partial", "text": "مرحبا بك في منصة"} |
|
Final Result Sent when a sentence is completed. Includes timestamps. |
{
"type": "final",
"text": "مرحبا بك في منصة مكنة.",
"words": [
{"word": "مرحبا", "start_time_ms": 150, "duration_ms": 300},
{"word": "بك", "start_time_ms": 460, "duration_ms": 200}
]
}
|
Ending the stream: When the user stops speaking, send a JSON text frame: {"command": "stop"} to safely close the connection and calculate the final refund.
Example Integration
# pip install websockets pyaudio
import asyncio
import websockets
import json
import pyaudio
import sys
API_KEY = "YOUR_API_KEY"
URL = "wss://moknah.io/api/v1/stt/ws/?language=ar-SA"
# 16kHz, 16-bit, Mono (100ms chunks)
CHUNK = 1600
async def stream_live_microphone(ws):
audio = pyaudio.PyAudio()
stream = audio.open(format=pyaudio.paInt16, channels=1, rate=16000, input=True, frames_per_buffer=CHUNK)
print("🎤 Microphone is LIVE. Start speaking!")
loop = asyncio.get_running_loop()
try:
while True:
data = await loop.run_in_executor(None, stream.read, CHUNK, False)
await ws.send(data)
await asyncio.sleep(0.001)
except asyncio.CancelledError:
await ws.send(json.dumps({"command": "stop"}))
finally:
stream.stop_stream()
stream.close()
audio.terminate()
async def receive_transcriptions(ws):
async for message in ws:
data = json.loads(message)
if data.get('type') == 'partial':
sys.stdout.write(f"\r⏳ {data['text']}")
sys.stdout.flush()
elif data.get('type') == 'final':
sys.stdout.write(f"\r✅ {data['text']}\n")
sys.stdout.flush()
async def main():
headers = {"Authorization": f"Bearer {API_KEY}"}
async with websockets.connect(URL, additional_headers=headers) as ws:
await asyncio.gather(
asyncio.create_task(stream_live_microphone(ws)),
asyncio.create_task(receive_transcriptions(ws))
)
asyncio.run(main())
// 1. Connect to WebSocket
const ws = new WebSocket('wss://moknah.io/api/v1/stt/ws/?api_key=YOUR_API_KEY&language=ar-SA');
ws.onmessage = (event) => {
const data = JSON.parse(event.data);
if (data.type === 'partial') {
console.log("Live:", data.text);
} else if (data.type === 'final') {
console.log("Final:", data.text);
}
};
// 2. Capture Microphone Audio on Connection
ws.onopen = async () => {
const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
// Force 16kHz sample rate for the AudioContext
const audioCtx = new AudioContext({ sampleRate: 16000 });
const source = audioCtx.createMediaStreamSource(stream);
const processor = audioCtx.createScriptProcessor(4096, 1, 1);
processor.onaudioprocess = (e) => {
if (ws.readyState === WebSocket.OPEN) {
const float32Data = e.inputBuffer.getChannelData(0);
// Convert Browser Float32 to Int16 PCM (Required by Server)
const pcm16 = new Int16Array(float32Data.length);
for (let i = 0; i < float32Data.length; i++) {
pcm16[i] = Math.max(-1, Math.min(1, float32Data[i])) * 0x7FFF;
}
// Send binary PCM audio chunk to server
ws.send(pcm16.buffer);
}
};
source.connect(processor);
processor.connect(audioCtx.destination);
};
// 3. Handle graceful shutdown
function stopRecording() {
ws.send(JSON.stringify({ command: "stop" }));
}
Disconnect & Error Codes
If the connection is rejected or dropped, the WebSocket will close with one of these specific codes:
| Close Code | Name | Description |
|---|---|---|
| 1000 | NORMAL_CLOSURE |
Stream ended successfully. Usage logged and refunded appropriately. |
| 4001 | MISSING_API_KEY |
The ?api_key= parameter was not provided or is invalid. |
| 4002 | UNAUTHORIZED |
Invalid API Key, restricted IP, or insufficient credits to start/continue the stream. |
| 4009 | PAYLOAD_TOO_LARGE |
You sent an audio chunk larger than 256KB. Stream in smaller buffers. |
| 4290 | CONCURRENT_LIMIT |
Too many active WebSocket connections for your account. Limit is 1 active stream. |
For API-related questions or issues, contact us at api@moknah.io.