POST /v1/chat/completions is the main inference endpoint. It follows the OpenAI Chat Completions format, so existing
OpenAI clients work by changing only the base URL, the key, and the model id.
from openai import OpenAI
client = OpenAI(base_url="https://api.kurrens.ai/v1", api_key="<KURRENS_API_KEY>")
completion = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=[
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Explain KV caching in two sentences."},
],
temperature=0.6,
max_tokens=512,
)
print(completion.choices[0].message.content)
print(completion.usage)import OpenAI from 'openai';
const client = new OpenAI({ baseURL: 'https://api.kurrens.ai/v1', apiKey: process.env.KURRENS_API_KEY });
const completion = await client.chat.completions.create({
model: 'Qwen/Qwen3.8-27B',
messages: [
{ role: 'system', content: 'You are a concise assistant.' },
{ role: 'user', content: 'Explain KV caching in two sentences.' },
],
temperature: 0.6,
max_tokens: 512,
});
console.log(completion.choices[0].message.content, completion.usage);curl https://api.kurrens.ai/v1/chat/completions \
-H "Authorization: Bearer $KURRENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3.8-27B",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Explain KV caching in two sentences."}
],
"temperature": 0.6,
"max_tokens": 512
}'Messages
| Role | Purpose |
|---|---|
system / developer |
Instructions that shape the model’s behavior |
user |
Your input — text, or text plus images for vision models |
assistant |
Earlier model turns, including tool_calls |
tool |
Results of a tool call, with tool_call_id |
Common parameters
| Parameter | Notes |
|---|---|
model |
Required. A model id from the Model catalog |
max_tokens |
Upper bound on generated tokens, including reasoning tokens |
temperature, top_p |
Sampling controls |
stop |
Up to 4 stop sequences |
stream |
Stream tokens as server-sent events — see Streaming |
tools, tool_choice |
Tool calling |
response_format |
Structured outputs |
seed |
Best-effort reproducibility |
service_tier |
priority or flex where offered |
Supported parameters and their ranges vary by model; the full per-model list is published in
models.json. Unsupported parameters return 400.
Usage
Every response includes a usage object with prompt_tokens, completion_tokens, and — where applicable —
cached and reasoning token counts. This is what you’re billed for.
Data handling
The request and response are processed in memory and discarded when the response completes. See Zero data retention.