glm-5.2
40% off@cf/zai-org/glm-5.2Z.ai's flagship model. Strongest at multi-step reasoning, math and long-context analysis, with tool calls and strict JSON output.
Reasoning tokens are billed at the output rate — never twice.
Seven production models, priced in USD per 1M tokens. Upstream IDs public.
| Model | Capabilities | Context | Input | Cached input | Output |
|---|---|---|---|---|---|
| glm-5.240% off@cf/zai-org/glm-5.2 | toolsreasoningjson | 262K | |||
| kimi-k2.7-code40% off@cf/moonshotai/kimi-k2.7-code | toolsvisionreasoning | 262K | |||
| kimi-k2.640% off@cf/moonshotai/kimi-k2.6 | toolsvisionreasoning | 262K | |||
| glm-4.7-flash40% off@cf/zai-org/glm-4.7-flash | toolsreasoningfast | 131K | — | ||
| deepseek-v4-flash40% off@cf/deepseek-ai/deepseek-v4-flash-0731 | toolsreasoningfast | 1M | |||
| deepseek-v4-pro40% off@cf/deepseek-ai/deepseek-v4-pro-0813 | toolsreasoningjson | 1M | |||
| qwen3.8-27b40% off@cf/qwen/qwen3.8-27b | toolsreasoningvision | 262K |
Full cache hits are billed at 10% of the normal price. Prices as of 2026-08-18 — dashboard shows live rates.
Z.ai's flagship model. Strongest at multi-step reasoning, math and long-context analysis, with tool calls and strict JSON output.
Reasoning tokens are billed at the output rate — never twice.
Moonshot AI's coding specialist. Tuned for repository-scale code work and agentic tool use, with vision input for screenshots and diagrams.
Reasoning tokens are billed at the output rate — never twice.
General-purpose Kimi model balancing quality and cost. Strong tool calling and vision — a solid default for assistants and agents.
Reasoning tokens are billed at the output rate — never twice.
The fast, low-cost workhorse. Median time to first token is about 0.65 s in our own measurements — built for high-volume classification, extraction and chat.
Reasoning tokens are billed at the output rate — never twice.
DeepSeek's official V4 Flash release with a 1M-token context window. Fast in our tests (~1 s to first token, ~100 tok/s) with strong agentic tool calling — the value pick for long-context work.
Reasoning tokens are billed at the output rate — never twice.
DeepSeek's flagship V4 model: 1M-token context window, deep reasoning and reliable function calling. Built for repository-scale analysis and long-document workloads.
Reasoning tokens are billed at the output rate — never twice.
Alibaba's Qwen 3.8 27B — image input verified by our own probe, not copied from a spec sheet. Send images alongside text in the same request, with a 262K-token context window, tool calling and reasoning.
Reasoning tokens are billed at the output rate — never twice.
About reasoning tokens: reasoning models stream their thinking as part of the completion output. Those tokens are a subset of output tokens, billed once at the output rate, and shown separately in your usage logs — they are never billed on top.