Calling DeepSeek API: pricing, integration and model selection (2026)

The short answer: Call DeepSeek through the OpenAI-compatible API. Set base_url to https://api.yomiapi.com/v1 and model to deepseek-v4-pro or deepseek-v4-flash, then register for a key and make your first request in minutes. Pricing, selection, code and cost examples follow.

Introduction

Last month, developer Lin's GPT bill approached $200 for a batch tool primarily summarizing Chinese content. Most requests did not need a premium model. We switched the pipeline to DeepSeek, reducing costs by an order of magnitude with little change in summary quality.

DeepSeek's Chinese-language performance and separate input/output token pricing make it a popular value option in 2026. This article covers choosing between two models, pricing, integration, advanced use and cost estimates.

DeepSeek's open-source approach means rapid iteration, active communities and accessible troubleshooting. Its Chinese training supports stable copywriting, summaries and translation. Budget-conscious teams can start there and introduce stronger flagships as needed.

Key takeaways

  • DeepSeek uses the OpenAI-compatible protocol; change base_url without restructuring code.
  • Two common models:deepseek-v4-pro for flagship reasoning/coding and deepseek-v4-flash for lightweight high concurrency, at roughly one third the price.
  • Price snapshot 2026-08-08: v4-pro input $0.43 and output $0.87 per 1M tokens; v4-flash input $0.14 and output $0.28.
  • One key accesses all models; change model to A/B test Claude, GPT and DeepSeek.
  • Separate input/output token billing, itemized records and WeChat Pay/Alipay RMB top-ups are supported.

Choosing deepseek-v4-pro or deepseek-v4-flash

Two frequently used DeepSeek models are deepseek-v4-pro and deepseek-v4-flash. They are distinct product lines, not merely larger and smaller versions.

deepseek-v4-pro is the reasoning flagship, with 1M context and deep thinking for code generation, architecture and long documents. Speed is not its primary selling point; multi-step reasoning can justify the wait.

deepseek-v4-flash is lightweight, also with 1M context, emphasizing fast responses, concurrency and low cost for translation, summaries, classification and support. Its batch throughput is steadier than pro.

Actual prices, snapshot 2026-08-08, in USD per 1M tokens:

ModelInput priceOutput pricePositioning
deepseek-v4-pro $0.43 $0.87 Flagship reasoning/coding for complex tasks
deepseek-v4-flash $0.14 $0.28 Fast, lightweight, high-concurrency real-time tasks

Rates change with provider adjustments. Refer to the Model Explorer current listing. One recommendation:Start with flash for latency and cost; use pro for coding, deep reasoning and long documents. Mixing the two is much cheaper than using a flagship for everything.

For example, an e-commerce review analysis tool assigns classification and sentiment scoring to deepseek-v4-flash for tens of thousands of daily calls, then uses deepseek-v4-pro for weekly competitor research summaries. Both share a key, while most costs stay in the lower tier.

Three-step integration: your first call with the OpenAI SDK

Existing OpenAI SDK projects need change only base_url and model to use DeepSeek. The three steps:

  1. Register and create a key. Visit the Yomi API Register, then open the console's Keys page to generate a key. Create multiple project keys with separate permissions and quotas.
  2. Copy base_url. The OpenAI-compatible address is https://api.yomiapi.com/v1, including the trailing /v1.
  3. Write your first API call. Complete Python example:
# pip install openai
from openai import OpenAI

client = OpenAI(
    api_key="sk-YOUR-YOMI-API-KEY",        # Key created in the console
    base_url="https://api.yomiapi.com/v1" # OpenAI-compatible address including /v1
)

resp = client.chat.completions.create(
    model="deepseek-v4-pro",
    messages=[{"role": "user", "content": "Explain an LLM context window in three sentences"}],
)
print(resp.choices[0].message.content)

After it works, change model to deepseek-v4-flash and compare speed. Response formats remain OpenAI-compatible. Without an SDK, use curl to verify connectivity:

curl https://api.yomiapi.com/v1/chat/completions \
  -H "Authorization: Bearer sk-YOUR-YOMI-API-KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

If the response contains choices[0].message.content , the connection is working.

Advanced use: streaming and A/B testing with one key

For long answers, use stream=True to show the first token sooner and reduce perceived waiting:

stream = client.chat.completions.create(
    model="deepseek-v4-pro",
    messages=[{"role": "user", "content": "Write a 300-word product announcement"}],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="")

Switching models under one key for A/B tests is especially useful. One Yomi key accesses all 200+ models. Change only model:deepseek-v4-pro ↔ claude-sonnet-5 ↔ gpt-5.5. Save outputs and billing costs for the same task and compare them before choosing a primary model. Check the Model Explorer listing.

Two tips: control client concurrency and use retries for transient failures; for long-context sessions, trim and reorganize conversation history to reduce repeated input billing.

For complete API details, see the OpenAI API integration tutorial; both use the same protocol.

Cost estimates: how much is one million tokens?

Yomi prices input and output separately, with output costing more. Include both when budgeting. This example uses the 2026-08-08 snapshot:

Usagedeepseek-v4-prodeepseek-v4-flash
1M input + 1M output $0.43 + $0.87 = $1.30 $0.14 + $0.28 = $0.42
10,000 summaries (about 10,000 input + 2,000 output tokens each) About $604 About $196

For the same tasks, flash costs about one third as much as pro. In the example, 10,000 summaries save about $400, and the difference grows over time. RMB payments through WeChat Pay/Alipay and itemized records make spending clear.

Be cautious of implausibly cheap APIs. The June 8, 2026 advisory highlighted exposed data, model substitution, malicious implants and outbound transfers. Verify baseline protections such as TLS 1.3, no content storage and automatic failover.

FAQ: context, Chinese-language tasks and tool calls

Final thoughts

For Lin, the question was which tasks belong to which model. Assign batches and frequent calls to deepseek-v4-flash, complex reasoning and coding to deepseek-v4-pro, and A/B test with the same key when unsure. Lower bills let Lin fund more frequent monitoring and improve the product.

Start now: register → create a key → run the code above. Your first DeepSeek response can arrive in minutes. For comparisons, see the AI model API price comparison covering rates across 200+ models.

Get a free API key →