DeepSeek announced it was retiring its $1.98 model. Then, quietly, it didn't.
On September 10, 2026, DeepSeek cut its Flash pricing to $0.15 per million input tokens and $0.60 per million output tokens off-peak, and told everyone still calling deepseek-v4-pro that starting September 14, every one of those requests would get silently rerouted to Flash and billed at Flash rates. A request that used to cost $1.98 on output would cost $0.60. Four days is not much runway to notice you've hardcoded a model name into your cost model. Then, sometime before today, DeepSeek reversed that part of the plan, and the only place I found that admitted it was a footnote on the pricing page.
I build small AI tools for a living, and DeepSeek is one of three providers I keep an eye on for anything latency-tolerant. This week is a decent case study in what it actually feels like to run a business on someone else's pricing tier, because two contradictory things were both true depending on which day you checked.
What actually shipped on September 10
DeepSeek-V4.1-Flash launched at $0.15 per million input tokens and $0.60 per million output tokens off-peak, doubling to $0.30 and $1.20 at peak. Cached input hits cost $0.003 per million off-peak and $0.006 at peak. Peak hours are 01:00 to 04:00 UTC and 06:00 to 10:00 UTC, Monday through Friday; everything else counts as off-peak, so a solo operator running batch jobs overnight UTC time is already paying half price without changing anything. Flash also ships with a concurrency limit of 2,500 versus 500 for V4 Pro, which matters more than the headline price if you're running anything with real throughput. All of this is confirmed on DeepSeek's own Models & Pricing page and in the September 10 release notes.
The retirement that didn't happen
The same September 10 announcement said, in plain language: "Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches." That's a real quote from DeepSeek's own release notes, and it's what the eesel.ai and techbriefly writeups both picked up and ran with in the days after launch. Independent technical writers were still treating it as settled fact as late as September 13, one day before the cutover was supposed to happen.
Except it didn't happen, or at least not the way it was announced. As of today, the DeepSeek pricing page still lists deepseek-v4-pro as its own model with its own unchanged rates: $1.98 per million output tokens off-peak, $3.96 at peak, cache hits at $0.022 off-peak and cache misses at $0.66. The footnote reads: "In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." No blog post, no press mention, just a parenthetical on a pricing table that most people check once and never again.
So the actual sequence was: announce a forced migration, let it sit as the plan of record for four days, then quietly keep the old model and old pricing running anyway. If you'd already rewritten your code to call deepseek-flash explicitly because you assumed the migration was happening regardless, you're paying Flash rates now. If you left your code calling deepseek-v4-pro and did nothing, you're still paying $1.98. Neither choice was wrong, but neither was something you controlled.
Cache-hit ratio is the number that actually decides your bill
The headline price everyone quotes is the $0.15 input rate, but that's the cache miss rate, the price you pay the first time a token enters the model's context. Cached input on Flash is $0.003 off-peak, fifty times cheaper. For an agent that's re-sending the same system prompt, tool definitions, and conversation history on every turn, most of your input tokens should be cache hits by the third or fourth turn in a loop. If they're not, you're either restructuring your prompt in a way that busts the cache (reordering messages, injecting timestamps near the top, anything that changes the prefix) or you're on a provider or routing setup that isn't preserving cache state between calls.
I check this on every agent I ship by logging the cache-hit token count DeepSeek returns in the usage object, not by estimating it from the input price. A tool that looks cheap on paper at $0.15/million can cost three or four times more per completed task than a competitor with a higher headline rate, purely because its cache hit ratio is worse. Most solo builders price off the number in the pricing table because it's the number that's easy to find, not the number that determines what they'll actually pay at the end of the month.
552B parameters isn't "flash" anymore
The previous-generation V4-Flash ran at 284 billion parameters. V4.1-Flash is 552 billion, roughly double, using a new causal encoder-decoder split that activates about 8 billion parameters on input and 16 billion on output. That's still cheap to run relative to a dense model that size, but "Flash" now describes a model bigger than most companies' flagship products from two years ago.
The tradeoff shows up in DeepSeek's own technical report. To shrink the KV cache, the architecture uses a technique called bounded replay, which reconstructs local attention state from a short cached window instead of storing it. Replaying 128 tokens doesn't reconstruct every layer's exact original state, and DeepSeek says it saw negligible quality impact in its own testing, but it also acknowledges possible edge cases without publishing a numerical ablation for how often that approximation actually breaks something. That's a company telling you, in its own documentation, that it hasn't fully mapped the failure modes of the exact mechanism keeping the price this low.
What I'd actually do
I wouldn't hardcode a bare model name like deepseek-v4-pro or deepseek-flash into anything I bill a customer against. I'd pin to whatever dated version string the provider exposes, treat the changelog as a thing I check weekly rather than on launch day, and build a small script that pulls the current pricing page and diffs it against last week's, because that's apparently the only reliable record of what actually changed. I'd also actually log cache-hit ratios per task type instead of assuming they're high, because that number is what your margin is built on, not the sticker price.
Where I could be wrong: this time, DeepSeek's reversal was good for the people who did nothing. They kept V4 Pro's old pricing rather than forcing a migration, which is the opposite of the horror story. A vendor collapsing your pricing tier out from under you can also mean a vendor deciding not to, and you don't get to choose which one you'll get next time either. The lesson isn't that DeepSeek is untrustworthy, it's that a four-day-old public announcement from a frontier lab wasn't a reliable enough source to build a pricing model on, and the ground truth lived in a footnote instead.
Author
Lukas
@lukcombinator