Fast, private inference for open models.
Run DeepSeek, Kimi, GLM, Qwen, MiniMax, and more through one OpenAI-compatible API.Your prompts are processed in memory and never stored.
Works with
to discard
/v1/chat/completionsdeepseek-ai/DeepSeek-V4.1-FlashSummarize this clause in one sentence: “Customer Content is processed in memory and is not stored.”
- Key authenticatedEdge
- Scheduled on a dedicated GPU nodefp8
- Prompt held in memory0 B on disk
- Streaming tokens0 tokens
- Content discarded from memoryZDR
Open models, one API
Call the latest open-weight models with the OpenAI SDK you already use. Streaming, tool calling, structured outputs, and prompt caching — billed per token, no minimums.
Every model lists its Hugging Face weights, precision, context window, and price.
Browse modelsZero data retention, in writing
Requests are authenticated and metered at our edge, then run on GPU nodes we operate. The prompt lives in memory only for as long as the model needs it.
We keep token counts, latency, and status codes for billing — never the content.
What happens to your dataYour capacity, your latency
Reserve GPUs for a single model and get a private base URL on the same API — consistent latency, custom rate limits, no noisy neighbors.
Under load we return an early 429 instead of queueing, so your client can retry elsewhere.
Talk to usYour prompts are your product.
We don't keep them.
Not stored
Inputs and outputs are processed in memory and discarded when the response completes.
Not trained on
We don't use your data to train, fine-tune, or evaluate any model — ours or anyone else's.
Not shared
We run every model ourselves. Your content never goes to a third-party model provider.
In writing
Zero data retention is a clause in our Terms of Service, and it overrides every other policy we publish.