Rate Limits
Understanding API rate limits, concurrency, and how to work within them
Overview
Rate limits protect the API from abuse and ensure fair usage for all users. Moknah enforces limits on request frequency, credit consumption, and concurrent connections across both our HTTP and WebSocket endpoints.
Moknah processes exactly one action per user at a time.
• HTTP: Wait for the current request to complete before sending another. Concurrent requests receive a 409 Conflict or 429 Too Many Requests.
• WebSocket: You may only have 1 active WebSocket connection open at a time. Opening a second connection will instantly close with code 4290.
Default Limits
| Limit Type | Value | Applies To | Description |
|---|---|---|---|
| Requests per Minute (RPM) | 60 |
HTTP Only | Maximum number of standard API POST requests per minute. |
| Credits per Minute (CPM) | 10,000 |
HTTP & WebSocket | Maximum credits consumed per minute. |
| Active Connections | 1 |
WebSocket Only | Maximum number of simultaneous WSS streams per user. |
WebSocket Limits & Close Codes
Because WebSockets maintain a persistent connection, they do not use HTTP headers for rate limiting. Instead, limits are enforced via connection drops and JSON control messages:
| Limit Reached | Behavior / Output |
|---|---|
| Concurrency Limit | Connection instantly closes with code 4290. |
| Credits Exhausted | Receives JSON error and closes. |
| Chunk Size Limit | Receives JSON error if a single text chunk exceeds 2,000 characters. |
| Message Rate Limit | Receives JSON error if you send text chunks faster than the queue can process them. |
Note: If your WebSocket disconnects unexpectedly, wait at least 15 seconds before reconnecting to ensure the server-side lock has fully cleared.
HTTP Rate Limit Headers
For standard REST endpoints, every API response includes headers to help you track your status:
| Header | Description | Example |
|---|---|---|
RateLimit-Limit |
Your RPM limit | 60 |
RateLimit-Remaining |
Requests remaining this minute | 45 |
RateLimit-Reset |
Unix timestamp when limits reset | 1699574460 |
Moknah-Credits-Remaining |
Credits remain this minute | 10000 |
Moknah-Credits-Used |
Credits used this minute | 5000 |
When HTTP Rate Limited
If you exceed the HTTP rate limit, you'll receive a 429 Too Many Requests response with these additional headers:
| Header | Description | Example |
|---|---|---|
Retry-After |
Seconds to wait before retrying | 45 |
RateLimit-Reset |
Unix timestamp when you can retry | 1699574460 |
Credit Calculation
Credits are calculated based on text length and processing options. For WebSockets, credits are deducted dynamically per-chunk as they are streamed.
| Factor | Multiplier | Description |
|---|---|---|
| Base cost | 1x |
1 credit per character |
| AI-Enhanced Normalization | 2x |
Advanced Arabic processing with diacritics |
| Premium Voice | +% |
Additional percentage for premium voices |
Example: A 500-character text with AI-Enhanced normalization:
500 characters × 2 (AI-Enhanced) = 1,000 credits
Best Practices
1. Use WebSockets for High-Volume Sequential Text
If you are generating conversational AI responses or parsing long documents, do not make 50 rapid HTTP POST requests. Open a single WebSocket connection and stream the chunks over the persistent connection. This bypasses the 60 RPM HTTP limit entirely (though credit limits still apply).
2. Implement Request Queuing (HTTP)
Since Moknah doesn't support concurrent requests, queue your HTTP requests and process them sequentially:
import queue
import threading
import time
class TTSQueue:
def __init__(self, api_key):
self.api_key = api_key
self.queue = queue.Queue()
self.worker = threading.Thread(target=self._process_queue, daemon=True)
self.worker.start()
def _process_queue(self):
while True:
text, voice_id, callback = self.queue.get()
try:
result = self._generate(text, voice_id)
callback(result, None)
except Exception as e:
callback(None, e)
finally:
self.queue.task_done()
def _generate(self, text, voice_id):
# Your API call here
pass
def add(self, text, voice_id, callback):
self.queue.put((text, voice_id, callback))
# Usage
tts = TTSQueue("your_api_key")
tts.add("Hello world", "voice_123", lambda r, e: print(r or e))
class TTSQueue {
constructor(apiKey) {
this.apiKey = apiKey;
this.queue = [];
this.processing = false;
}
async add(text, voiceId) {
return new Promise((resolve, reject) => {
this.queue.push({ text, voiceId, resolve, reject });
this.processNext();
});
}
async processNext() {
if (this.processing || this.queue.length === 0) return;
this.processing = true;
const { text, voiceId, resolve, reject } = this.queue.shift();
try {
const result = await this.generate(text, voiceId);
resolve(result);
} catch (error) {
reject(error);
} finally {
this.processing = false;
this.processNext(); // Process next in queue
}
}
async generate(text, voiceId) {
// Your API call here
}
}
// Usage
const tts = new TTSQueue("your_api_key");
const audio = await tts.add("Hello world", "voice_123");
2. Implement Request Queuing (HTTP)
Since Moknah doesn't support concurrent requests, queue your HTTP requests and process them sequentially:
import time
import random
def request_with_backoff(func, max_retries=5):
retries = 0
while retries < max_retries:
try:
response = func()
if response.status_code == 429:
retry_after = int(response.headers.get("Retry-After", 60))
# Add jitter to prevent thundering herd
wait_time = retry_after + random.uniform(0, 1)
print(f"Rate limited. Waiting {wait_time:.1f}s...")
time.sleep(wait_time)
retries += 1
continue
return response
except Exception as e:
# Exponential backoff for other errors
wait_time = (2 ** retries) + random.uniform(0, 1)
print(f"Error: {e}. Retrying in {wait_time:.1f}s...")
time.sleep(wait_time)
retries += 1
raise Exception("Max retries exceeded")
async function requestWithBackoff(func, maxRetries = 5) {
let retries = 0;
while (retries < maxRetries) {
try {
const response = await func();
if (response.status === 429) {
const retryAfter = parseInt(response.headers.get("Retry-After") || 60);
// Add jitter to prevent thundering herd
const waitTime = retryAfter + Math.random();
console.log(`Rate limited. Waiting ${waitTime.toFixed(1)}s...`);
await new Promise(r => setTimeout(r, waitTime * 1000));
retries++;
continue;
}
return response;
} catch (error) {
// Exponential backoff for other errors
const waitTime = Math.pow(2, retries) + Math.random();
console.log(`Error: ${error.message}. Retrying in ${waitTime.toFixed(1)}s...`);
await new Promise(r => setTimeout(r, waitTime * 1000));
retries++;
}
}
throw new Error("Max retries exceeded");
}
3. Implement Exponential Backoff
When rate limited (HTTP 429) or connection limited (WebSocket 4290), use exponential backoff to retry securely without spamming the server:
Summary
HTTP limits: 60 RPM, 1 concurrent request.
WebSocket limits: 1 concurrent connection, 2,000 chars per chunk.
Global limits: 10,000 CPM (Credits Per Minute).
Need more capacity? Email sales@moknah.io
For API-related questions or issues, contact us at api@moknah.io.