DeepSeek raised the cost of its flagship V4‑Pro and V4‑Flash large‑language models on Aug. 16, sending peak prices to $3.96 and $1.32 per million output tokens. The hike represents roughly a fourfold increase from the previous $0.87 and $0.28 rates. Off‑peak pricing, which applies during quieter compute periods, sits at half those figures—$1.98 for V4‑Pro and $0.66 for V4‑Flash—still more than double the old peak rates.
At the same time, the Beijing‑based firm rolled out a developer preview of DeepSeek Harness v0.1. The harness acts as a scaffolding layer around a model, enabling an autonomous agent to read and edit files, browse the web and continue working until a task is completed. DeepSeek markets the harness as an open‑architecture platform, allowing users to plug in any component, including models from rival providers. The company highlighted this openness as a contrast to competitors it says hard‑code their products.
Developer reaction and performance metrics
Early feedback from the AI community was mixed. Researchers praised the model’s capabilities in niche areas such as cybersecurity, but many developers found the overall performance underwhelming. Benchmark data released by DeepSeek places V4‑Pro at 80.6% on the SWE‑bench Verified test, roughly on par with Google’s Gemini 3.1 Pro and a hair behind Anthropic’s Claude Opus 4.6. On Terminal Bench 2.0 the model scores 67.9%, trailing OpenAI’s GPT‑5.4 at 75.1%, and it records 37.7% on Humanity’s Last Exam compared with Gemini 3.1 Pro’s 44.4%.
Despite the modest gains, the price jump has drawn criticism. Developers note that the new off‑peak rates exceed what they paid during peak hours before the hike, effectively doubling their compute costs even when they shift workloads to quieter times.
Strategic timing and market outlook
Bloomberg ties the pricing decision to DeepSeek’s upcoming initial public offering. The firm, valued at roughly $71 billion, is reportedly preparing to list as early as this year. Founder Liang Wenfeng must balance investor expectations, rapid expansion and the rising expense of compute resources. The company’s engineering advances—its V4‑Pro model uses a mixture‑of‑experts architecture with 1.6 trillion parameters and an attention design that cuts per‑token compute to 27% of the previous generation—have historically enabled low‑cost pricing. The recent surge suggests a shift toward higher margins rather than a reflection of increased operating costs.
Industry observers have coined a “death zone” around DeepSeek’s pricing, warning that models that become too expensive or underperform may be sidelined in favor of cheaper alternatives like Moonshot’s Kimi K3 ($15 per million tokens) or Anthropic’s Fable 5 ($50 per million tokens). Yet even at $3.96, DeepSeek remains cheaper than those rivals.
In a brief statement posted on its website, DeepSeek claimed the new V4‑Pro‑0813 build offered “significantly enhanced agent capabilities.” The message vanished later that afternoon, and the company has not commented on the removal. The disappearance adds another layer of uncertainty as developers weigh the cost increase against promised functionality.
Cet article a été rédigé avec l'assistance de l'IA.
News Factory APP - actualités agentiques pour booster votre SEO et AEO.