reasoning LLM API, chain of thought model
Reasoning models spend extra tokens before answering. That is a real cost and a real latency, so they earn their keep only when the task has a verifiable answer — maths, code, multi-step extraction, debugging. For plain summarisation or rewriting they are usually the wrong tool: a cheaper chat model reaches the same output at a fraction of the price. Note that thinking tokens are billed as output tokens, so the effective price of a reasoning call is often several times the headline output rate.
Prices are per 1M tokens in USD · ¥6.6 = $1