image input LLM API, vision model API, multimodal
Models on this page accept images natively: send an image_url content block and the model reads the picture itself. Models that lack native vision are also usable with images here — the gateway reads the image and injects a description into the prompt — so you can send the same request body to any model in the catalogue and still get a sensible answer. Native multimodal models are preferable when the task depends on fine visual detail; the bridged path is preferable when you want one cheap model for everything.
Prices are per 1M tokens in USD · ¥6.6 = $1