long context LLM API, 1M token context, large document analysis
Long context lets you skip a retrieval pipeline, but it does not make long context free: input tokens are billed whether they are cached or not, and a 200K-token prompt costs the same as the output of hundreds of normal requests. Repeated prefixes are the case where the platform's cache-aware routing pays off — keep the stable part of the prompt (system instructions, documents) at the front and the volatile part at the end, and repeated calls against the same prefix cost materially less.
Prices are per 1M tokens in USD · ¥6.6 = $1