Kurrens applies per-key limits on requests per minute and tokens per minute. Limits depend on your account and the model. Contact sales for higher limits or a dedicated endpoint.
We fail fast instead of queueing
When a model is at capacity in its primary region, the request moves to the next region
instead of waiting in a queue. If no region has capacity — or you’re over your own limit — Kurrens returns
429 Too Many Requests immediately. That keeps latency predictable and lets your client retry or fall back right away.
Handling 429
- Retry with exponential backoff and jitter, starting around one second.
- Respect the
Retry-Afterheader when present. - Spread large jobs over time instead of sending bursts.
import random, time
from openai import OpenAI, RateLimitError
client = OpenAI(base_url="https://api.kurrens.ai/v1")
def complete(messages, attempts=5):
for attempt in range(attempts):
try:
return client.chat.completions.create(model="deepseek-ai/DeepSeek-V4.1-Flash", messages=messages)
except RateLimitError:
time.sleep(min(30, 2 ** attempt) + random.random())
raise RuntimeError("rate limited")
The OpenAI SDKs also retry 429 responses automatically (max_retries).