· 10 min read

DeepSeek Quietly Shipped V4 Pro's GA Release With No Blog Post. The Benchmarks Jumped More Than the Silence Suggests.

I've written about DeepSeek V4 Pro three times already on this blog: the April launch, the permanent price cut in May, and V4 Flash's GA in July beating its own flagship on agent benchmarks. This is the fourth chapter, and it's a different kind of story. On August 12, DeepSeek swapped the deepseek-v4-pro API endpoint to a new build called DeepSeek-V4-Pro-0813, ending a preview period that ran nearly four months. No blog post. No changelog entry. No press release. Just a version string on a pricing page, and a set of benchmark numbers that, if you take them at face value, look like the biggest single jump this model line has posted yet.

What actually shipped

DeepSeek's own API docs confirm it in one dry sentence: the deepseek-v4-pro model name "has been updated to DeepSeek-V4-Pro-0813," and the calling method is unchanged, so if you were already pointing at that endpoint, you're already running the new build whether you noticed or not. The architecture itself didn't change. It's still a mixture-of-experts model with 1.6 trillion total parameters and 49 billion active per token, still a 1 million token context window, still a 384K max output. What changed is the post-training, the same playbook DeepSeek ran on V4 Flash back in July: keep the base model, redo the fine-tuning, ship it as a version bump instead of a new model name.

The launch mechanics are the part I keep coming back to. Simon Willison flagged that the benchmark comparison table circulated first through DeepSeek's official WeChat group, then got copied into a Reddit post that was deleted for being low-effort, then got reproduced as an ASCII table on Hacker News. There's still no official DeepSeek announcement page for this release as of this writing. For a lab that's now shipped three major model updates in four months, treating a flagship GA release like a footnote is either supreme confidence or a sign they'd rather not draw attention to a pricing page that also contains a warning about a rate hike. Possibly both.

The benchmark claims, and why I'm not taking them at face value yet

Here's what DeepSeek is claiming, according to its own benchmark table: DeepSWE jumped from 12.8 to 62.7, CyberGym from 52.7 to 83.3, and Terminal Bench 2.1 from 72.1 to 87.9, all measured against the April preview build. If that Terminal Bench number holds up, DeepSeek would be putting V4 Pro within a tenth of a point of Claude Fable 5's 88.0 on a benchmark that actually grades whether an agent finishes a real task in a real shell, not multiple choice.

I want to be direct about the caveat here because it matters more than the headline number: no third-party evaluator has independently replicated any of these figures. The benchmark-tracking site benchable.ai shows an empty results section for V4 Pro 0813 as of publication, and DeepSeek hasn't published the evaluation harness it used to generate the agentic scores, which means nobody outside the company can currently reproduce them. That's not a small gap. The April preview build scored 80.6 on DeepSeek's own vendor-controlled SWE-bench Verified, a benchmark with a reported 8.5% false-positive rate, and 8% pass@1 on the independently run DeepSWE evaluation from yage.ai, which uses a stricter verifier with a 0.3% false-positive rate. That's not a rounding difference, it's an order-of-magnitude gap between the number the vendor picked to headline and the number an independent party actually measured. So when DeepSeek claims the 0813 build jumped to 62.7 on that same tighter DeepSWE benchmark, my honest read is: that would be a genuinely big deal if true, and "if true" is doing real work in that sentence until someone outside DeepSeek runs the eval.

The pricing cliff nobody's talking about enough

Pricing at GA carried straight over from preview: $0.435 per million input tokens on a cache miss, $0.003625 per million on a cache hit, and $0.87 per million output. Those are the numbers that made the May price cut a big deal, and if that's all that was happening here, this would be a shrug of a post.

It isn't all that's happening. Buried in the same pricing page that confirmed the 0813 GA is a notice that DeepSeek plans a significant price increase, and multiple outlets covering the story since have reported the specifics: starting at 00:00 Beijing time on August 17, DeepSeek is introducing peak and off-peak pricing, with peak hours running 9am to noon and 2pm to 6pm Beijing time. For V4 Pro specifically, cache-hit input rises 6x off-peak and 12x at peak. Cache-miss input rises 1.5x off-peak and 3x at peak, which in dollar terms takes that $0.435 input rate up toward roughly $1.32 per million at peak. Output pricing rises 2.25x off-peak and 4.5x at peak.

I check DeepSeek's pricing page more often than I'd like to admit, mostly because I got burned once by a provider that changed a rate mid-billing-cycle with a two-line notice I missed for a week. That habit is why this one stood out to me: DeepSeek gave developers a five-day window between confirming the GA build and flipping the pricing structure, and the actual multiplier depends heavily on which token type dominates your workload. If your usage is cache-hit-heavy (which a lot of agentic workflows are, since repeated system prompts and tool definitions cache well), you're looking at the 12x end of that range during peak hours, not the more modest cache-miss number people are quoting when they round this down to "prices are going up a bit."

What I'd actually do

If you're already running production traffic through deepseek-v4-pro, go pull your token breakdown by cache-hit versus cache-miss versus output before August 17, not after. A single blended cost estimate hides the difference between a workload that gets 50% more expensive and one that gets 12 times more expensive, and those are very different conversations to have with yourself about whether this model still pencils out. If a meaningful chunk of your load can shift to off-peak hours (roughly 6pm to 9am and noon to 2pm Beijing time), that's a real lever: off-peak rates are half of peak across the board.

On the benchmark claims, treat the 0813 numbers the same way I treated V4 Flash's numbers in July: real enough to justify running your own eval against your own tasks, not real enough to make a production routing decision on the vendor's table alone. Run three or four of your actual agent workloads against the 0813 build this week, while pricing is still at the preview rate, and compare failure rate and total cost, not just the headline benchmark gap. You'll get a truer read on whether the post-training gains are real for your use case than waiting for benchable.ai to catch up.

The honest counter-take

I could be overweighting the silent-launch angle. DeepSeek has now shipped three model updates this year (April, July, August) without much ceremony each time, and it's possible that's just how this lab operates rather than a sign of anything specific about this release. Plenty of legitimate, well-tested model updates ship as quiet version bumps; not every quiet launch is hiding something. And the pricing increase, while real and worth planning around, isn't unprecedented either: DeepSeek flagged a peak/off-peak scheme for V4 Flash back in July that still hasn't taken effect as far as I can tell, so "announced" and "actually enforced on schedule" aren't always the same thing with this vendor. It's possible August 17 slips too.

But even granting all of that, the combination here is specific: unreplicated benchmark claims on the exact model you'd be building agentic workloads on, landing four days before a confirmed price schedule that could raise your bill by an order of magnitude depending on your cache-hit ratio. That's not a reason to avoid V4 Pro. It's a reason to know your own numbers before August 17 instead of after.

Author

Sources

Stay in the Loop

Get new posts delivered to your inbox. No spam, unsubscribe anytime.

Newsletter coming soon. Set PUBLIC_CONVERTKIT_FORM_ID in .env to activate.

Related Posts