Skip to content

Models & pricing

Seven production models, priced in USD per 1M tokens. Upstream IDs public.

ModelCapabilitiesContextInputCached inputOutput
glm-5.240% off@cf/zai-org/glm-5.2toolsreasoningjson262K$1.40$0.84$0.26$0.156$4.40$2.64
kimi-k2.7-code40% off@cf/moonshotai/kimi-k2.7-codetoolsvisionreasoning262K$0.95$0.57$0.19$0.114$4.00$2.40
kimi-k2.640% off@cf/moonshotai/kimi-k2.6toolsvisionreasoning262K$0.95$0.57$0.16$0.096$4.00$2.40
glm-4.7-flash40% off@cf/zai-org/glm-4.7-flashtoolsreasoningfast131K$0.06$0.036$0.40$0.24
deepseek-v4-flash40% off@cf/deepseek-ai/deepseek-v4-flash-0731toolsreasoningfast1M$0.44$0.264$0.014$0.008$1.32$0.792
deepseek-v4-pro40% off@cf/deepseek-ai/deepseek-v4-pro-0813toolsreasoningjson1M$1.32$0.792$0.044$0.026$3.96$2.376
qwen3.8-27b40% off@cf/qwen/qwen3.8-27btoolsreasoningvision262K$0.45$0.27$0.45$0.27$3.20$1.92

Full cache hits are billed at 10% of the normal price. Prices as of 2026-08-18 — dashboard shows live rates.

glm-5.2

40% off@cf/zai-org/glm-5.2

Z.ai's flagship model. Strongest at multi-step reasoning, math and long-context analysis, with tool calls and strict JSON output.

Context window262K
Capabilitiestoolsreasoningjson

Reasoning tokens are billed at the output rate — never twice.

Input$1.40$0.84
Cached input$0.26$0.156
Output$4.40$2.64

kimi-k2.7-code

40% off@cf/moonshotai/kimi-k2.7-code

Moonshot AI's coding specialist. Tuned for repository-scale code work and agentic tool use, with vision input for screenshots and diagrams.

Context window262K
Capabilitiestoolsvisionreasoning

Reasoning tokens are billed at the output rate — never twice.

Input$0.95$0.57
Cached input$0.19$0.114
Output$4.00$2.40

kimi-k2.6

40% off@cf/moonshotai/kimi-k2.6

General-purpose Kimi model balancing quality and cost. Strong tool calling and vision — a solid default for assistants and agents.

Context window262K
Capabilitiestoolsvisionreasoning

Reasoning tokens are billed at the output rate — never twice.

Input$0.95$0.57
Cached input$0.16$0.096
Output$4.00$2.40

glm-4.7-flash

40% off@cf/zai-org/glm-4.7-flash

The fast, low-cost workhorse. Median time to first token is about 0.65 s in our own measurements — built for high-volume classification, extraction and chat.

Context window131K
Capabilitiestoolsreasoningfast

Reasoning tokens are billed at the output rate — never twice.

Input$0.06$0.036
Cached input
Output$0.40$0.24

deepseek-v4-flash

40% off@cf/deepseek-ai/deepseek-v4-flash-0731

DeepSeek's official V4 Flash release with a 1M-token context window. Fast in our tests (~1 s to first token, ~100 tok/s) with strong agentic tool calling — the value pick for long-context work.

Context window1M
Capabilitiestoolsreasoningfast

Reasoning tokens are billed at the output rate — never twice.

Input$0.44$0.264
Cached input$0.014$0.008
Output$1.32$0.792

deepseek-v4-pro

40% off@cf/deepseek-ai/deepseek-v4-pro-0813

DeepSeek's flagship V4 model: 1M-token context window, deep reasoning and reliable function calling. Built for repository-scale analysis and long-document workloads.

Context window1M
Capabilitiestoolsreasoningjson

Reasoning tokens are billed at the output rate — never twice.

Input$1.32$0.792
Cached input$0.044$0.026
Output$3.96$2.376

qwen3.8-27b

40% off@cf/qwen/qwen3.8-27b

Alibaba's Qwen 3.8 27B — image input verified by our own probe, not copied from a spec sheet. Send images alongside text in the same request, with a 262K-token context window, tool calling and reasoning.

Context window262K
Capabilitiestoolsreasoningvision

Reasoning tokens are billed at the output rate — never twice.

Input$0.45$0.27
Cached input$0.45$0.27
Output$3.20$1.92

About reasoning tokens: reasoning models stream their thinking as part of the completion output. Those tokens are a subset of output tokens, billed once at the output rate, and shown separately in your usage logs — they are never billed on top.