# Reasoning

> Use models that think before answering, and control reasoning effort.

Models marked **Reasoning** produce internal reasoning before the final answer.

```python
completion = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V4.1-Flash",
    messages=[{"role": "user", "content": "How many primes are below 100?"}],
    reasoning_effort="medium",
)
message = completion.choices[0].message
print(message.content)
```

## What to know

- **Billing.** Reasoning tokens are billed at the output rate and reported in `usage.completion_tokens_details.reasoning_tokens`.
- **`max_tokens` includes reasoning.** Leave enough headroom, or the answer may be cut off before it starts.
- **Effort.** Where supported, `reasoning_effort` (`low`, `medium`, `high`) trades latency and cost for depth.
- **Reasoning content.** Some models return their reasoning in a `reasoning_content` field alongside `content`.
  It is processed like any other output — [not stored](/docs/data-security/zero-data-retention).
- **Keep-alives.** While a model reasons, streaming responses may include SSE comment lines; ignore them.
