| qwen3.8-flash | qwen3.8-max | |
|---|---|---|
| Providers | Qwen | Qwen |
| Input $/M | $0.12 | $1.85 |
| Output $/M | $0.42 | $5.54 |
| Output multiplier | ×3.38 | ×3.00 |
| Context window | 1.05M tokens | 1.05M tokens |
| Intelligence index | — | 53.4 |
| Coding index | — | 68.9 |
qwen3.8-flash is roughly 15.0× cheaper on input than qwen3.8-max. If your workload passes the quality bar on the cheaper model, that gap compounds across every request; route to qwen3.8-max deliberately rather than by default.
Tags: qwen3.8-flash — Chat, Reasoning, Coding, Vision · qwen3.8-max — Chat, Reasoning, Coding, Vision
from openai import OpenAI
client = OpenAI(base_url="https://starseaapi.com/v1", api_key="sk-...")
# qwen3.8-flash
client.chat.completions.create(model="qwen3.8-flash", messages=[{"role":"user","content":"Hi"}])
# qwen3.8-max
client.chat.completions.create(model="qwen3.8-max", messages=[{"role": "user", "content": "Hi"}])