Skip to content

Models & pricing

Nine production models, priced in USD per 1M tokens. Upstream IDs public.

ModelCapabilitiesContextInputCached inputOutput
glm-5.260% off@cf/zai-org/glm-5.2toolsreasoningjson262K$1.40$0.56$0.26$0.104$4.40$1.76
kimi-k2.7-code60% off@cf/moonshotai/kimi-k2.7-codetoolsvisionreasoning262K$0.95$0.38$0.19$0.076$4.00$1.60
kimi-k2.660% off@cf/moonshotai/kimi-k2.6toolsvisionreasoning262K$0.95$0.38$0.16$0.064$4.00$1.60
glm-4.7-flash60% off@cf/zai-org/glm-4.7-flashtoolsreasoningfast131K$0.06$0.024—$0.40$0.16
deepseek-v4-flash60% off@cf/deepseek-ai/deepseek-v4-flash-0731toolsreasoningfast1M$0.44$0.176$0.014$0.005$1.32$0.528
deepseek-v4-pro60% off@cf/deepseek-ai/deepseek-v4-pro-0813toolsreasoningjson1M$1.32$0.528$0.044$0.017$3.96$1.584
qwen3.8-27b60% off@cf/qwen/qwen3.8-27btoolsreasoningvision262K$0.45$0.18$0.05$0.02$3.20$1.28
glm-5.3-flash60% off@cf/zai-org/glm-5.3-flashtoolsvisionreasoningjson1M$0.15$0.06$0.03$0.012$0.50$0.20
glm-5.360% off@cf/zai-org/glm-5.3toolsreasoningjson1M$1.40$0.56$0.26$0.104$4.40$1.76

Full cache hits are billed at 10% of the normal price. Prices as of 2026-09-26 — dashboard shows live rates.

glm-5.2

60% off@cf/zai-org/glm-5.2

Z.ai’s proven reasoning workhorse. Strong at multi-step reasoning, math and long-context analysis, with tool calls and strict JSON output.

Context window262K
Capabilitiestoolsreasoningjson

Reasoning tokens are billed at the output rate — never twice.

Input$1.40$0.56
Cached input$0.26$0.104
Output$4.40$1.76

kimi-k2.7-code

60% off@cf/moonshotai/kimi-k2.7-code

Moonshot AI's coding specialist. Tuned for repository-scale code work and agentic tool use, with vision input for screenshots and diagrams.

Context window262K
Capabilitiestoolsvisionreasoning

Reasoning tokens are billed at the output rate — never twice.

Input$0.95$0.38
Cached input$0.19$0.076
Output$4.00$1.60

kimi-k2.6

60% off@cf/moonshotai/kimi-k2.6

General-purpose Kimi model balancing quality and cost. Strong tool calling and vision — a solid default for assistants and agents.

Context window262K
Capabilitiestoolsvisionreasoning

Reasoning tokens are billed at the output rate — never twice.

Input$0.95$0.38
Cached input$0.16$0.064
Output$4.00$1.60

glm-4.7-flash

60% off@cf/zai-org/glm-4.7-flash

The fast, low-cost workhorse. Median time to first token is about 0.65 s in our own measurements — built for high-volume classification, extraction and chat.

Context window131K
Capabilitiestoolsreasoningfast

Reasoning tokens are billed at the output rate — never twice.

Input$0.06$0.024
Cached input—
Output$0.40$0.16

deepseek-v4-flash

60% off@cf/deepseek-ai/deepseek-v4-flash-0731

DeepSeek's official V4 Flash release with a 1M-token context window. Fast in our tests (~1 s to first token, ~100 tok/s) with strong agentic tool calling — the value pick for long-context work.

Context window1M
Capabilitiestoolsreasoningfast

Reasoning tokens are billed at the output rate — never twice.

Input$0.44$0.176
Cached input$0.014$0.005
Output$1.32$0.528

deepseek-v4-pro

60% off@cf/deepseek-ai/deepseek-v4-pro-0813

DeepSeek's flagship V4 model: 1M-token context window, deep reasoning and reliable function calling. Built for repository-scale analysis and long-document workloads.

Context window1M
Capabilitiestoolsreasoningjson

Reasoning tokens are billed at the output rate — never twice.

Input$1.32$0.528
Cached input$0.044$0.017
Output$3.96$1.584

qwen3.8-27b

60% off@cf/qwen/qwen3.8-27b

Alibaba's Qwen 3.8 27B — image input verified by our own probe, not copied from a spec sheet. Send images alongside text in the same request, with a 262K-token context window, tool calling and reasoning.

Context window262K
Capabilitiestoolsreasoningvision

Reasoning tokens are billed at the output rate — never twice.

Input$0.45$0.18
Cached input$0.05$0.02
Output$3.20$1.28

glm-5.3-flash

60% off@cf/zai-org/glm-5.3-flash

Z.ai's GLM 5.3 Flash: a 1M-token context window and image input verified by our own probe, in one model — at the lowest input price of any vision model in the catalog. Tool calls, reasoning and strict JSON included.

Context window1M
Capabilitiestoolsvisionreasoningjson

Reasoning tokens are billed at the output rate — never twice.

Input$0.15$0.06
Cached input$0.03$0.012
Output$0.50$0.20

glm-5.3

60% off@cf/zai-org/glm-5.3

Z.ai’s flagship agentic coding model: a 1M-token context window at the same list price as GLM-5.2, with tool calls, five reasoning-effort levels and strict JSON output. Text-only — image input is refused up front, at no cost to you.

Context window1M
Capabilitiestoolsreasoningjson

Reasoning tokens are billed at the output rate — never twice.

Input$1.40$0.56
Cached input$0.26$0.104
Output$4.40$1.76

About reasoning tokens: reasoning models stream their thinking as part of the completion output. Those tokens are a subset of output tokens, billed once at the output rate, and shown separately in your usage logs — they are never billed on top.