AI model API costs compared: GPT, Claude and DeepSeek prices per million tokens (2026)

The recommendation: On a per-million-token basis through a gateway, DeepSeek is the budget leader (deepseek-v4-flash input from $0.14). Claude and GPT flagships sit around $0.5–0.6 input, while the premium claude-fable-5 costs $3.5. A quick reference and provider tables explain the main rates below.

Introduction

Your product manager gives you three days, but pricing is the first obstacle. GPT and Claude can differ severalfold for the same request, DeepSeek seems unusually cheap and Kimi specializes in long documents. Provider pricing pages use different currencies, units and context tiers, making comparisons exhausting.

This article explains three billing concepts, gives input/output prices per million tokens by provider, then recommends models for chat, coding, long documents and batches. Prices come from Yomi API Model Explorer data; one key lets you test them yourself.

Why not simply compare provider websites? Units differ between tokens, characters and plans; currencies and taxes complicate comparisons; and the same model can vary greatly by gateway route. We use USD per million tokens throughout for a common basis.

TierExample modelInput $/1MOutput $/1MPositioning
Value deepseek-v4-flash 0.14 0.28 Fast and lightweight for frequent calls
Balanced claude-sonnet-5 0.24 1.20 Mainstay for coding and everyday chat
Flagship claude-opus-5 0.60 3.00 Complex reasoning and long-running agent tasks
Premium claude-fable-5 3.50 17.50 Peak performance when cost is less important

Key takeaways

  • Input and output tokens are priced separately. Output usually costs several times more: 2× for DeepSeek and 5–6× for most flagships.
  • Cache hits cost 10–20% of normal input rates, reducing repeated-context costs substantially.
  • Routes can price the same model differently. Tables use the cheapest currently available platform route.
  • RMB top-ups through WeChat Pay / Alipay, precise token metering and itemized console records are supported.

Three concepts: tokens, separate input/output prices, and caching/batches

API billing multiplies token counts by rates, separately for input and output, rather than charging per message.

First, counting tokens. A token is roughly 0.7–1 Chinese characters or a little over half an English word. One million input tokens represents about 700,000–900,000 Chinese characters. Most products use less than this monthly, so actual calls often cost only a few cents to a few tenths of a yuan.

Second, separate input/output prices. Prompts, conversation history and documents count as input; generated answers count as output. Output rates are typically 2× input for DeepSeek, 5× for Claude/Kimi and 6× for GPT flagships. Long-prompt, short-answer tasks are input-heavy; reasoning with long answers makes output price the main cost driver.

Third, cache and batch prices. Most models support context caching, charging cache hits at 10–20% of input rates: around 10% for Claude/GPT, 19% for GLM and, for DeepSeek's deepseek-v4-pro , as low as 0.8%. Cache writes cost 1.25× the input rate when needed. Repeated questions over the same documents can reduce unit costs by an order of magnitude.

Remember this general approach:DeepSeek for frequent low-cost tasks, Sonnet for balance, Opus/GPT flagships for demanding work, and Kimi for long documents. The tables add the details.

Provider prices per million tokens: Anthropic, OpenAI, DeepSeek, Zhipu and Kimi

These input/output prices for available models come from a live pricing snapshot. Model names follow the current Model Explorer listing as the reference.

Anthropic (Claude series)

ModelsInput $/1MOutput $/1MPositioning
claude-sonnet-4-60.361.80Balanced, efficient everyday model
claude-sonnet-50.241.20Cost-effective, balanced coding assistant
claude-opus-4-60.603.00High-end programming reasoning
claude-opus-4-70.603.00Advanced Professional Flagship
claude-opus-4-80.603.00Flagship reasoning and long-running agents
claude-opus-50.603.00Deep Reasoning King
claude-fable-53.5017.50Premium flagship and peak performance

Claude has the widest internal price range. Sonnet handles everyday work; all four Opus entries cost $0.6 / $3, letting you choose by task difficulty. Only strict performance requirements justify claude-fable-5, whose output rate is nearly six times the Opus tier.

For teams on a budget:claude-sonnet-5 has an input rate of $0.24, less than half the Opus tier, while covering most everyday coding and chat needs. Reserve Opus for the roughly 10% of requests needing deeper reasoning.

OpenAI (GPT series)

ModelsInput $/1MOutput $/1MPositioning
gpt-5.40.251.50Previous flagship, mature and stable
gpt-5.50.503.00Deep Thinking Workhorse
gpt-5.6-terra0.201.20Coding specialist for frequent lightweight tasks
gpt-5.6-sol0.503.00Premium flagship with 1M context

Within GPT, gpt-5.6-terra stands out for coding at prices close to DeepSeek, especially completion and lightweight agents.

For products sensitive to ecosystem compatibility, such as extensive OpenAI-based code, GPT is a choice requiring no structural code changes.gpt-5.4 handles established workloads, while gpt-5.6-sol handles new tasks requiring 1M context, using the same integration code.

DeepSeek (the value benchmark)

ModelsInput $/1MOutput $/1MPositioning
deepseek-v4-flash0.140.28Fast and lightweight for frequent calls
deepseek-v4-pro0.430.87Deep reasoning at excellent value

Both DeepSeek models have a 2× output multiplier, the lowest among the main providers. For reasoning answers with thousands of tokens, this can reduce answer costs to about one third of comparable competitors.

DeepSeek also has low cache rates (deepseek-v4-pro cache hits are around 8% of input rates). Combined with 2× output pricing, this suits document Q&A and long answers. For truly demanding multi-step reasoning, Opus or GPT flagships may still be needed.

Zhipu (GLM series)

ModelsInput $/1MOutput $/1MPositioning
glm-5.20.842.64Reliable Chinese-model workhorse
glm-5.20.842.64Chinese flagship with a comprehensive upgrade

Two GLM generations share the same price, so upgrades need not change budgets. This suits teams requiring Chinese models and predictable costs.

Kimi (long-document benchmark)

ModelsInput $/1MOutput $/1MPositioning
kimi-k31.809.00Very long context for long documents

Kimi's rates are higher, but long context is its core advantage. For million-token tasks, other models may need repeated chunks that multiply input costs. Reading the full document once can make Kimi cheaper per task.

All-model price comparison

Compare input/output rates and workloads in one table:

ModelsInput $/1MOutput $/1MUse cases
deepseek-v4-flash0.140.28Chatbots, batch classification and frequent small tasks
gpt-5.6-terra0.201.20Code completion, lightweight agents and structured output
claude-sonnet-50.241.20Everyday chat, general coding and tool calls
gpt-5.40.251.50Stable production and compatibility-first workloads
claude-sonnet-4-60.361.80Medium/long-context chat and everyday work
deepseek-v4-pro0.430.87Deep reasoning, complex analysis and long outputs
gpt-5.50.503.00Deep thinking and complex problem decomposition
gpt-5.6-sol0.503.00Premium reasoning and 1M context
claude-opus-4-60.603.00Difficult coding and architecture planning
claude-opus-4-70.603.00Professional analysis and complex reasoning
claude-opus-4-80.603.00Flagship reasoning and long-running agents
claude-opus-50.603.00Leading deep reasoning for demanding tasks
glm-5.20.842.64Chinese-model solutions and stable workloads
glm-5.20.842.64Chinese flagship for general-purpose tasks
kimi-k31.809.00Million-token documents and report analysis
claude-fable-53.5017.50Peak-performance, quality-first premium tasks

Snapshot: 2026-08-08. Rates change with provider adjustments; live prices in Model Explorer take precedence. Three bands emerge: $0.14–0.25 for lightweight models, $0.24–0.6 for mainstream models, and $0.84+ for long-context and premium tiers.

Estimate costs for all 16 models using actual projected monthly input and output. For hundreds of thousands of tokens, the difference may be small; at tens of millions, lightweight and premium tiers can differ by more than 20×.

Choose by workload: chat, coding, documents and batches

The same million-token monthly workload can cost ten times more with an unsuitable model. Here are four common workloads, recommendations and approximate cost levels:

WorkloadRecommended modelsApproximate costSelection rationale
Chat / customer support deepseek-v4-flash, claude-sonnet-5 Input-heavy; about $0.14–0.24 per 1M tokens Chat input often exceeds output; prioritize lower input rates
Coding assistant gpt-5.6-terra, claude-sonnet-5, claude-opus-4-8 Moderate usage; about $0.2–0.6 per 1M tokens Use terra for completion, Sonnet for everyday coding and Opus for demanding work
Long-document analysis kimi-k3, gpt-5.6-sol Very large input, potentially a million tokens per task Read once with a long-context model to avoid repeated chunk costs
Batch tasks deepseek-v4-pro, deepseek-v4-flash Output-heavy; about $0.28–0.87 per 1M tokens Output rates dominate; DeepSeek's lower multiplier offers clear savings

For new projects, validate the workflow with a cheaper tier, then upgrade only bottleneck tasks. Switching through a gateway changes one model field, keeping experimentation inexpensive.

Prices change: use Model Explorer for live rates

Providers adjust prices every quarter. Bookmark a live pricing source rather than relying only on this article.Yomi API Model Explorer shows current input/output prices for 200+ models from 30+ providers, updated with provider changes.

Price is only part of selection; data safety is essential. The June 2026 security advisory named exposed data, downgraded models, malicious implants and outbound transfers; the Cyberspace Administration is pursuing its AI application campaign, and operators have faced detention over noncompliant services. Itemized records, no content storage and TLS 1.3 matter more than low prices alone. Legitimate platforms show balances and every charge transparently.

The goal is useful models within budget. Start with a small top-up and a week of real requests, then evaluate latency and stability before scaling. Yomi supports RMB payments through WeChat Pay/Alipay and itemized records for controlled testing costs.

View live prices in Model Explorer →