User Guide API Reference

Rate Limits

Understanding API rate limits, concurrency, and how to work within them

Overview

Rate limits protect the API from abuse and ensure fair usage for all users. Moknah enforces limits on request frequency, credit consumption, and concurrent connections across both our HTTP and WebSocket endpoints.

Strict No-Concurrency Policy

Moknah processes exactly one action per user at a time.
• HTTP: Wait for the current request to complete before sending another. Concurrent requests receive a 409 Conflict or 429 Too Many Requests.
• WebSocket: You may only have 1 active WebSocket connection open at a time. Opening a second connection will instantly close with code 4290.

Default Limits

Limit Type Value Applies To Description
Requests per Minute (RPM) 60 HTTP Only Maximum number of standard API POST requests per minute.
Credits per Minute (CPM) 10,000 HTTP & WebSocket Maximum credits consumed per minute.
Active Connections 1 WebSocket Only Maximum number of simultaneous WSS streams per user.

WebSocket Limits & Close Codes

Because WebSockets maintain a persistent connection, they do not use HTTP headers for rate limiting. Instead, limits are enforced via connection drops and JSON control messages:

Limit Reached Behavior / Output
Concurrency Limit Connection instantly closes with code 4290.
Credits Exhausted Receives JSON error and closes.
Chunk Size Limit Receives JSON error if a single text chunk exceeds 2,000 characters.
Message Rate Limit Receives JSON error if you send text chunks faster than the queue can process them.

Note: If your WebSocket disconnects unexpectedly, wait at least 15 seconds before reconnecting to ensure the server-side lock has fully cleared.

HTTP Rate Limit Headers

For standard REST endpoints, every API response includes headers to help you track your status:

Header Description Example
RateLimit-Limit Your RPM limit 60
RateLimit-Remaining Requests remaining this minute 45
RateLimit-Reset Unix timestamp when limits reset 1699574460
Moknah-Credits-Remaining Credits remain this minute 10000
Moknah-Credits-Used Credits used this minute 5000

When HTTP Rate Limited

If you exceed the HTTP rate limit, you'll receive a 429 Too Many Requests response with these additional headers:

Header Description Example
Retry-After Seconds to wait before retrying 45
RateLimit-Reset Unix timestamp when you can retry 1699574460

Credit Calculation

Credits are calculated based on text length and processing options. For WebSockets, credits are deducted dynamically per-chunk as they are streamed.

Factor Multiplier Description
Base cost 1x 1 credit per character
AI-Enhanced Normalization 2x Advanced Arabic processing with diacritics
Premium Voice +% Additional percentage for premium voices

Example: A 500-character text with AI-Enhanced normalization:

500 characters × 2 (AI-Enhanced) = 1,000 credits

Best Practices

1. Use WebSockets for High-Volume Sequential Text

If you are generating conversational AI responses or parsing long documents, do not make 50 rapid HTTP POST requests. Open a single WebSocket connection and stream the chunks over the persistent connection. This bypasses the 60 RPM HTTP limit entirely (though credit limits still apply).

2. Implement Request Queuing (HTTP)

Since Moknah doesn't support concurrent requests, queue your HTTP requests and process them sequentially:

import queue
import threading
import time

class TTSQueue:
    def __init__(self, api_key):
        self.api_key = api_key
        self.queue = queue.Queue()
        self.worker = threading.Thread(target=self._process_queue, daemon=True)
        self.worker.start()

    def _process_queue(self):
        while True:
            text, voice_id, callback = self.queue.get()
            try:
                result = self._generate(text, voice_id)
                callback(result, None)
            except Exception as e:
                callback(None, e)
            finally:
                self.queue.task_done()

    def _generate(self, text, voice_id):
        # Your API call here
        pass

    def add(self, text, voice_id, callback):
        self.queue.put((text, voice_id, callback))

# Usage
tts = TTSQueue("your_api_key")
tts.add("Hello world", "voice_123", lambda r, e: print(r or e))
class TTSQueue {
  constructor(apiKey) {
    this.apiKey = apiKey;
    this.queue = [];
    this.processing = false;
  }

  async add(text, voiceId) {
    return new Promise((resolve, reject) => {
      this.queue.push({ text, voiceId, resolve, reject });
      this.processNext();
    });
  }

  async processNext() {
    if (this.processing || this.queue.length === 0) return;

    this.processing = true;
    const { text, voiceId, resolve, reject } = this.queue.shift();

    try {
      const result = await this.generate(text, voiceId);
      resolve(result);
    } catch (error) {
      reject(error);
    } finally {
      this.processing = false;
      this.processNext(); // Process next in queue
    }
  }

  async generate(text, voiceId) {
    // Your API call here
  }
}

// Usage
const tts = new TTSQueue("your_api_key");
const audio = await tts.add("Hello world", "voice_123");

2. Implement Request Queuing (HTTP)

Since Moknah doesn't support concurrent requests, queue your HTTP requests and process them sequentially:

import time
import random

def request_with_backoff(func, max_retries=5):
    retries = 0

    while retries < max_retries:
        try:
            response = func()

            if response.status_code == 429:
                retry_after = int(response.headers.get("Retry-After", 60))
                # Add jitter to prevent thundering herd
                wait_time = retry_after + random.uniform(0, 1)
                print(f"Rate limited. Waiting {wait_time:.1f}s...")
                time.sleep(wait_time)
                retries += 1
                continue

            return response

        except Exception as e:
            # Exponential backoff for other errors
            wait_time = (2 ** retries) + random.uniform(0, 1)
            print(f"Error: {e}. Retrying in {wait_time:.1f}s...")
            time.sleep(wait_time)
            retries += 1

    raise Exception("Max retries exceeded")
async function requestWithBackoff(func, maxRetries = 5) {
  let retries = 0;

  while (retries < maxRetries) {
    try {
      const response = await func();

      if (response.status === 429) {
        const retryAfter = parseInt(response.headers.get("Retry-After") || 60);
        // Add jitter to prevent thundering herd
        const waitTime = retryAfter + Math.random();
        console.log(`Rate limited. Waiting ${waitTime.toFixed(1)}s...`);
        await new Promise(r => setTimeout(r, waitTime * 1000));
        retries++;
        continue;
      }

      return response;

    } catch (error) {
      // Exponential backoff for other errors
      const waitTime = Math.pow(2, retries) + Math.random();
      console.log(`Error: ${error.message}. Retrying in ${waitTime.toFixed(1)}s...`);
      await new Promise(r => setTimeout(r, waitTime * 1000));
      retries++;
    }
  }

  throw new Error("Max retries exceeded");
}

3. Implement Exponential Backoff

When rate limited (HTTP 429) or connection limited (WebSocket 4290), use exponential backoff to retry securely without spamming the server:


Summary

Quick Reference

HTTP limits: 60 RPM, 1 concurrent request.
WebSocket limits: 1 concurrent connection, 2,000 chars per chunk.
Global limits: 10,000 CPM (Credits Per Minute).
Need more capacity? Email sales@moknah.io

API Support

For API-related questions or issues, contact us at api@moknah.io.