Pre-launchAPI access is by invitation. Prices shown are estimates.Request access →
Contact sales
ModelsPricingDedicatedDocsEnterpriseContact sales

Fast, private inference for open models.

Run DeepSeek, Kimi, GLM, Qwen, MiniMax, and more through one OpenAI-compatible API.Your prompts are processed in memory and never stored.

OpenAI-compatible.
Works with
OpenAI SDKVercel AI SDKLangChainLlamaIndexLiteLLMCursorDify
Every request, from prompt
to discard
Request traceIllustrative
StreamZDRfp8OpenAI API
POST/v1/chat/completions
Modeldeepseek-ai/DeepSeek-V4.1-Flash
User

Summarize this clause in one sentence: “Customer Content is processed in memory and is not stored.”

Assistant

Trace
  • Key authenticatedEdge
  • Scheduled on a dedicated GPU nodefp8
  • Prompt held in memory0 B on disk
  • Streaming tokens0 tokens
  • Content discarded from memoryZDR
Tokens streamedStored: 0 B
Serverless

Open models, one API

Call the latest open-weight models with the OpenAI SDK you already use. Streaming, tool calling, structured outputs, and prompt caching — billed per token, no minimums.

Every model lists its Hugging Face weights, precision, context window, and price.

Browse models
Catalog
AllDeepSeekQwen
ModelContextEst. in $/MEst. out $/MPrecision
DeepSeek V4.1 Flash1M$0.22$0.84fp8 · ZDR
DeepSeek V4 Pro 08131M$1.31$3.96fp8 · ZDR
DeepSeek V4 Flash 07311.3M$0.14$0.37fp8 · ZDR
GLM 5.31.3M$1.19$4.40fp8 · ZDR
GLM 5.3 Flash1.3M$0.15$0.50fp8 · ZDR
Copied model iddeepseek-ai/DeepSeek-V4.1-Flash
OpenAI SDK · base_url setapi.kurrens.ai/v1
Data path
Zero data retention
KURRENS · OPERATED INFRASTRUCTUREYOUR APPprompt← 0 tokensEDGEauthmeterGPU NODEHBMKV CACHE0 BLOCKSUSAGE LOGprompt_tokens42completion_tokens27status200DISKno contentPrompt in memory0 B on diskDisk write refusedZDRMetadata loggedno contentMemory freed0 B retained
Privacy

Zero data retention, in writing

Requests are authenticated and metered at our edge, then run on GPU nodes we operate. The prompt lives in memory only for as long as the model needs it.

We keep token counts, latency, and status codes for billing — never the content.

What happens to your data
Dedicated

Your capacity, your latency

Reserve GPUs for a single model and get a private base URL on the same API — consistent latency, custom rate limits, no noisy neighbors.

Under load we return an early 429 instead of queueing, so your client can retry elsewhere.

Talk to us
Node 01
DedicatedSingle tenant
NODE 018 × GPU · TENSOR PARALLEL · FP8GPU 0HBMSMGPU 1HBMSMGPU 2HBMSMGPU 3HBMSMGPU 4HBMSMGPU 5HBMSMGPU 6HBMSMGPU 7HBMSMINGRESS · PRIVATE BASE URLEndpoint readyprivate base URLCustom rate limitsRPM · TPMNo noisy neighborsisolated
Privacy commitments

Your prompts are your product.
We don't keep them.

Not stored

Inputs and outputs are processed in memory and discarded when the response completes.

Not trained on

We don't use your data to train, fine-tune, or evaluate any model — ours or anyone else's.

Not shared

We run every model ourselves. Your content never goes to a third-party model provider.

In writing

Zero data retention is a clause in our Terms of Service, and it overrides every other policy we publish.

Open models, served in flowDeepSeek · Kimi · GLM · Qwen · MiniMax
Get started with Kurrens