AI model API costs compared: GPT, Claude and DeepSeek prices per million tokens (2026)
The recommendation: On a per-million-token basis through a gateway, DeepSeek is the budget leader (deepseek-v4-flash input from $0.14). Claude and GPT flagships sit around $0.5–0.6 input, while the premium claude-fable-5 costs $3.5. A quick reference and provider tables explain the main rates below.
Introduction
Your product manager gives you three days, but pricing is the first obstacle. GPT and Claude can differ severalfold for the same request, DeepSeek seems unusually cheap and Kimi specializes in long documents. Provider pricing pages use different currencies, units and context tiers, making comparisons exhausting.
This article explains three billing concepts, gives input/output prices per million tokens by provider, then recommends models for chat, coding, long documents and batches. Prices come from Yomi API Model Explorer data; one key lets you test them yourself.
Why not simply compare provider websites? Units differ between tokens, characters and plans; currencies and taxes complicate comparisons; and the same model can vary greatly by gateway route. We use USD per million tokens throughout for a common basis.
| Tier | Example model | Input $/1M | Output $/1M | Positioning |
|---|---|---|---|---|
| Value | deepseek-v4-flash |
0.14 | 0.28 | Fast and lightweight for frequent calls |
| Balanced | claude-sonnet-5 |
0.24 | 1.20 | Mainstay for coding and everyday chat |
| Flagship | claude-opus-5 |
0.60 | 3.00 | Complex reasoning and long-running agent tasks |
| Premium | claude-fable-5 |
3.50 | 17.50 | Peak performance when cost is less important |
Key takeaways
- Input and output tokens are priced separately. Output usually costs several times more: 2× for DeepSeek and 5–6× for most flagships.
- Cache hits cost 10–20% of normal input rates, reducing repeated-context costs substantially.
- Routes can price the same model differently. Tables use the cheapest currently available platform route.
- RMB top-ups through WeChat Pay / Alipay, precise token metering and itemized console records are supported.
Three concepts: tokens, separate input/output prices, and caching/batches
API billing multiplies token counts by rates, separately for input and output, rather than charging per message.
First, counting tokens. A token is roughly 0.7–1 Chinese characters or a little over half an English word. One million input tokens represents about 700,000–900,000 Chinese characters. Most products use less than this monthly, so actual calls often cost only a few cents to a few tenths of a yuan.
Second, separate input/output prices. Prompts, conversation history and documents count as input; generated answers count as output. Output rates are typically 2× input for DeepSeek, 5× for Claude/Kimi and 6× for GPT flagships. Long-prompt, short-answer tasks are input-heavy; reasoning with long answers makes output price the main cost driver.
Third, cache and batch prices. Most models support context caching, charging cache hits at 10–20% of input rates: around 10% for Claude/GPT, 19% for GLM and, for DeepSeek's deepseek-v4-pro , as low as 0.8%. Cache writes cost 1.25× the input rate when needed. Repeated questions over the same documents can reduce unit costs by an order of magnitude.
Remember this general approach:DeepSeek for frequent low-cost tasks, Sonnet for balance, Opus/GPT flagships for demanding work, and Kimi for long documents. The tables add the details.
Provider prices per million tokens: Anthropic, OpenAI, DeepSeek, Zhipu and Kimi
These input/output prices for available models come from a live pricing snapshot. Model names follow the current Model Explorer listing as the reference.
Anthropic (Claude series)
| Models | Input $/1M | Output $/1M | Positioning |
|---|---|---|---|
claude-sonnet-4-6 | 0.36 | 1.80 | Balanced, efficient everyday model |
claude-sonnet-5 | 0.24 | 1.20 | Cost-effective, balanced coding assistant |
claude-opus-4-6 | 0.60 | 3.00 | High-end programming reasoning |
claude-opus-4-7 | 0.60 | 3.00 | Advanced Professional Flagship |
claude-opus-4-8 | 0.60 | 3.00 | Flagship reasoning and long-running agents |
claude-opus-5 | 0.60 | 3.00 | Deep Reasoning King |
claude-fable-5 | 3.50 | 17.50 | Premium flagship and peak performance |
Claude has the widest internal price range. Sonnet handles everyday work; all four Opus entries cost $0.6 / $3, letting you choose by task difficulty. Only strict performance requirements justify claude-fable-5, whose output rate is nearly six times the Opus tier.
For teams on a budget:claude-sonnet-5 has an input rate of $0.24, less than half the Opus tier, while covering most everyday coding and chat needs. Reserve Opus for the roughly 10% of requests needing deeper reasoning.
OpenAI (GPT series)
| Models | Input $/1M | Output $/1M | Positioning |
|---|---|---|---|
gpt-5.4 | 0.25 | 1.50 | Previous flagship, mature and stable |
gpt-5.5 | 0.50 | 3.00 | Deep Thinking Workhorse |
gpt-5.6-terra | 0.20 | 1.20 | Coding specialist for frequent lightweight tasks |
gpt-5.6-sol | 0.50 | 3.00 | Premium flagship with 1M context |
Within GPT, gpt-5.6-terra stands out for coding at prices close to DeepSeek, especially completion and lightweight agents.
For products sensitive to ecosystem compatibility, such as extensive OpenAI-based code, GPT is a choice requiring no structural code changes.gpt-5.4 handles established workloads, while gpt-5.6-sol handles new tasks requiring 1M context, using the same integration code.
DeepSeek (the value benchmark)
| Models | Input $/1M | Output $/1M | Positioning |
|---|---|---|---|
deepseek-v4-flash | 0.14 | 0.28 | Fast and lightweight for frequent calls |
deepseek-v4-pro | 0.43 | 0.87 | Deep reasoning at excellent value |
Both DeepSeek models have a 2× output multiplier, the lowest among the main providers. For reasoning answers with thousands of tokens, this can reduce answer costs to about one third of comparable competitors.
DeepSeek also has low cache rates (deepseek-v4-pro cache hits are around 8% of input rates). Combined with 2× output pricing, this suits document Q&A and long answers. For truly demanding multi-step reasoning, Opus or GPT flagships may still be needed.
Zhipu (GLM series)
| Models | Input $/1M | Output $/1M | Positioning |
|---|---|---|---|
glm-5.2 | 0.84 | 2.64 | Reliable Chinese-model workhorse |
glm-5.2 | 0.84 | 2.64 | Chinese flagship with a comprehensive upgrade |
Two GLM generations share the same price, so upgrades need not change budgets. This suits teams requiring Chinese models and predictable costs.
Kimi (long-document benchmark)
| Models | Input $/1M | Output $/1M | Positioning |
|---|---|---|---|
kimi-k3 | 1.80 | 9.00 | Very long context for long documents |
Kimi's rates are higher, but long context is its core advantage. For million-token tasks, other models may need repeated chunks that multiply input costs. Reading the full document once can make Kimi cheaper per task.
All-model price comparison
Compare input/output rates and workloads in one table:
| Models | Input $/1M | Output $/1M | Use cases |
|---|---|---|---|
deepseek-v4-flash | 0.14 | 0.28 | Chatbots, batch classification and frequent small tasks |
gpt-5.6-terra | 0.20 | 1.20 | Code completion, lightweight agents and structured output |
claude-sonnet-5 | 0.24 | 1.20 | Everyday chat, general coding and tool calls |
gpt-5.4 | 0.25 | 1.50 | Stable production and compatibility-first workloads |
claude-sonnet-4-6 | 0.36 | 1.80 | Medium/long-context chat and everyday work |
deepseek-v4-pro | 0.43 | 0.87 | Deep reasoning, complex analysis and long outputs |
gpt-5.5 | 0.50 | 3.00 | Deep thinking and complex problem decomposition |
gpt-5.6-sol | 0.50 | 3.00 | Premium reasoning and 1M context |
claude-opus-4-6 | 0.60 | 3.00 | Difficult coding and architecture planning |
claude-opus-4-7 | 0.60 | 3.00 | Professional analysis and complex reasoning |
claude-opus-4-8 | 0.60 | 3.00 | Flagship reasoning and long-running agents |
claude-opus-5 | 0.60 | 3.00 | Leading deep reasoning for demanding tasks |
glm-5.2 | 0.84 | 2.64 | Chinese-model solutions and stable workloads |
glm-5.2 | 0.84 | 2.64 | Chinese flagship for general-purpose tasks |
kimi-k3 | 1.80 | 9.00 | Million-token documents and report analysis |
claude-fable-5 | 3.50 | 17.50 | Peak-performance, quality-first premium tasks |
Snapshot: 2026-08-08. Rates change with provider adjustments; live prices in Model Explorer take precedence. Three bands emerge: $0.14–0.25 for lightweight models, $0.24–0.6 for mainstream models, and $0.84+ for long-context and premium tiers.
Estimate costs for all 16 models using actual projected monthly input and output. For hundreds of thousands of tokens, the difference may be small; at tens of millions, lightweight and premium tiers can differ by more than 20×.
Choose by workload: chat, coding, documents and batches
The same million-token monthly workload can cost ten times more with an unsuitable model. Here are four common workloads, recommendations and approximate cost levels:
| Workload | Recommended models | Approximate cost | Selection rationale |
|---|---|---|---|
| Chat / customer support | deepseek-v4-flash, claude-sonnet-5 |
Input-heavy; about $0.14–0.24 per 1M tokens | Chat input often exceeds output; prioritize lower input rates |
| Coding assistant | gpt-5.6-terra, claude-sonnet-5, claude-opus-4-8 |
Moderate usage; about $0.2–0.6 per 1M tokens | Use terra for completion, Sonnet for everyday coding and Opus for demanding work |
| Long-document analysis | kimi-k3, gpt-5.6-sol |
Very large input, potentially a million tokens per task | Read once with a long-context model to avoid repeated chunk costs |
| Batch tasks | deepseek-v4-pro, deepseek-v4-flash |
Output-heavy; about $0.28–0.87 per 1M tokens | Output rates dominate; DeepSeek's lower multiplier offers clear savings |
For new projects, validate the workflow with a cheaper tier, then upgrade only bottleneck tasks. Switching through a gateway changes one
modelfield, keeping experimentation inexpensive.
Prices change: use Model Explorer for live rates
Providers adjust prices every quarter. Bookmark a live pricing source rather than relying only on this article.Yomi API Model Explorer shows current input/output prices for 200+ models from 30+ providers, updated with provider changes.
Price is only part of selection; data safety is essential. The June 2026 security advisory named exposed data, downgraded models, malicious implants and outbound transfers; the Cyberspace Administration is pursuing its AI application campaign, and operators have faced detention over noncompliant services. Itemized records, no content storage and TLS 1.3 matter more than low prices alone. Legitimate platforms show balances and every charge transparently.
The goal is useful models within budget. Start with a small top-up and a week of real requests, then evaluate latency and stability before scaling. Yomi supports RMB payments through WeChat Pay/Alipay and itemized records for controlled testing costs.
Yomi API