Pre-launchAPI access is by invitation. Prices shown are estimates.Request access →
Contact sales
ModelsPricingDedicatedDocsEnterpriseContact sales
Inference

Docs / Inference

Vision

Send images alongside text to multimodal models.

View .md

Models whose input modalities include image accept images in user messages, in the OpenAI content-parts format.

completion = client.chat.completions.create(
    model="Qwen/Qwen3.8-27B",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What does this chart show?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}},
        ],
    }],
)

Sending images

  • URL — a publicly reachable https:// URL. We fetch it once to process the request.
  • Base64 — a data URL such as data:image/png;base64,.... Use this for private images.

Supported formats are PNG, JPEG, WebP, and non-animated GIF. Images are counted as input tokens; larger images cost more.

Video

Models whose input modalities include video — for example Kimi K3, MiniMax M3, and GLM 5.3 Flash — accept a video content part in the same way:

{"type": "video_url", "video_url": {"url": "https://example.com/clip.mp4"}}

Video is sampled into frames and counted as input tokens. Filter the Model catalog by “Image / video input” to see which models accept images or video.

Images, like all inputs, are processed in memory and not stored — see Zero data retention.

    Type to search titles, headings, and page text.

    ↑↓ to move · ↵ to open