Abstract:
From the official release to the time it was replaced by Flash, DeepSeek only left 32 days for the official version of V4 Pro.
On September 10, DeepSeek released V4.1 Flash and announced that starting from 12 noon on September 14, Beijing time, all requests sent to V4 Pro in the official API will be taken over by Flash and charged at a lower Flash price. This arrangement will continue until V4.1 Pro is launched.

V4 Pro will be available as a preview version from April 24th and will be officially released on August 13th.
Less than a month after the official version was launched, the company has already decided its "exit time".
What’s more worth asking is, why didn’t DeepSeek reduce the price of the Pro and keep it, or simply wait until the next generation of Pro is ready, and then let the two take over normally?
My judgment is that V4 Pro still has strengths, but it is no longer worth DeepSeek’s continued investment in computing power alone.
Taking over V4.1 Flash, which makes the architecture more efficient and has surpassed many intelligent agent tests, can not only reduce the company's operating expenses, but also lower the user's bills.01 Flash has caught up with the direction that Pro focused on last month
One month ago, the official version of V4 Pro was launched. Its main direction is very clear: intelligent agents, code programming and tool invocation. DeepSeek opened up different inference intensities for it at that time, and also specially adapted the Codex, hoping that it could take on more complex and longer-process work.

But something went wrong on the day the official version went online: the official website notification and open platform announcement were removed for a time and were restored that night. During this period, the API call entrance was always retained. According to media reports, some developers also reported that the model's thinking time was too short, long tasks ended prematurely, and the results may be very different if the same question is run a few hours later.
The subsequent price increase raised the standard for its fulfillment. Starting from August 17, the output price of Pro during peak hours has increased from 6 yuan to 27 yuan per million Tokens, which is 4.5 times the original price.
For developers who already feel that their performance is unstable, it is difficult to say that this extra expenditure has been exchanged for corresponding improvements.
Now, the newly released Flash has just caught up in these key directions.
Official data shows that in the software engineering test DeepSWE v1.1 and the office automation test AutomationBench, Flash scored 74.2 points and 54.8 points respectively, significantly ahead of V4Pro's 62.7 points and 43.2 points.
What the new model wins is precisely the task that Pro was promoted last month.

Pro doesn’t lose at every turn. In the HLE plain text subset that does not call tools, it still leads Flash with 42.7 points and 39.1 points; but in the complete HLE test that allows the use of tools, Flash scored 63.9 points, 60 points higher than Pro. In addition, in the comparison of basic models before post-training, Pro still maintains the lead in SimpleQA-Verified knowledge question and answer and LongBench-V2 long text tests.
Therefore, the "comprehensive surpass" in DeepSeek's official announcement cannot simply be understood as winning in every sub-category.
There are still things the Pro does well, and it's still valuable to users who happen to need those specific abilities. What is really shaken is the original commercial positioning of the flagship product.In the past, developers and enterprises easily understood the division of the two models: Flash was chosen if the budget was limited, and complex and long tasks were left to the more expensive Pro. Now that Flash has overtaken many mainstream smartphone tests, it is difficult for Pro to convince the market with the label of "more expensive, but generally stronger".
A broad lead is enough to support a high-priced flagship for the general public; but if it only dominates some segmented tasks, the company must recalculate: how many people cannot do without these capabilities, how much premium they are willing to pay for it, and whether it is worth allocating computing power clusters to support it.
Public information does not yet provide specific cost and user breakdown accounts. But judging from DeepSeek's product actions, it has obviously given up on the idea of retaining an independent and exclusive gear for the old Pro.
02 Reduce the price of Pro and save the cost of running it
If the problem is just that Flash is cheap, Pro can naturally follow the price reduction promotion. The difficulty is that price cuts can only change the numbers on the bill, but not the physical consumption underlying the model.
According to the data disclosed by the official model card, the number of parameters activated by each Token of V4 Pro is as high as 49 billion. V4.1Flash decouples processing input from generating answers, activating only 8 billion and 16 billion parameters respectively. Although the parameter ratio cannot be directly equated to the cost reduction, the scale of calculations called for each forward inference between the two is completely different in the same order of magnitude.
The new architecture also significantly reduces the storage overhead of long context tasks. Compared with V4 Flash, the global KV cache occupied by V4.1 Flash has been reduced to about one-quarter, and the SSD space occupied by the persistent KV cache has been reduced to about one-eighth. The running of the agent needs to repeatedly read the context code, reference documents and data returned by external tools. The cache is specially used to store the historical state calculated by the large model.
Once the task cycle is lengthened and the amount of concurrency increases, the storage memory will quickly bottom out. What Flash has to solve is not just answering a question correctly, but also the resource squeeze when massive concurrent tasks run at the same time.

DeepSeek directly mentioned in the announcement that higher architectural efficiency allows the company to carry more concurrency at a lower cost, so it is logical to lower the price.
As of September 11, calculated on a per million Token basis in Gushi, the input price of the V4 Pro miss cache is 4.5 yuan and the output price is 13.5 yuan; the corresponding V4.1 Flash prices are only 1 yuan and 4 yuan (the peak prices of both models are twice that of Gushi). With the implementation of the switch, the charges for the originally expensive Pro interface will be directly reduced to the level of Flash.
Although the pricing difference between the two is not equivalent to DeepSeek's real material cost, the strategy is very clear: use a highly efficient new architecture to fully accept requests for old interfaces, and then transfer the saved hardware costs to the market.
If you reduce the price of Pro, the platform will still have to bear the heavy inference and maintenance costs alone; switching to Flash will have the opportunity to reduce user bills and server load at the same time.
After the big model finished training, the bottomless pit of money was not closed. Every time it responds to a call, it occupies computing resources in real time, and operation and maintenance cannot stop for a moment.
"Just spent a lot of money to train" cannot be a death-free gold medal to keep the old Pro. Sunk costs have become a reality, and maintaining the old model requires continuous consumption of real money.
Even if the old Pro can continue to make money, it is obviously more cost-effective to use the same hardware computing power to run a new model with higher efficiency.03 The delivery of 160,000 Ascend chips may take more than a year
Combined with DeepSeek’s recent key actions, we can better understand the sense of urgency behind this resource transfer.
In August this year, DeepSeek launched peak and valley floating pricing to guide non-real-time batch processing tasks to run during idle times.
This is essentially to squeeze out all the existing computing power from internal scheduling when external hardware is limited.
Hardware procurement is also progressing, but water from afar cannot quench the thirst for nearness. Bloomberg reported on September 4, citing people familiar with the matter, that DeepSeek plans to deploy at least 160,000 Huawei Ascend 950DT chips in data centers in Inner Mongolia, which are currently mainly planned for model inference rather than training. However, the report also mentioned that Huawei's production capacity needs to take care of multiple customers, and due to supply chain bottlenecks, the entire order delivery cycle may take more than a year.

DeepSeek financing is also accelerating. Reuters reported on September 9, citing people familiar with the matter, that DeepSeek has hired CITIC Securities to promote the IPO process of the Science and Technology Innovation Board. The fund-raising is mainly to fill the funding gap for computing infrastructure, cutting-edge model research and development, and top talent recruitment. However, the listing timetable, fundraising amount and valuation target have not yet been finalized, and neither DeepSeek nor CITIC Securities commented on this.

Putting these pieces of the puzzle together, what emerges is a startup company that desires rapid expansion but is always constrained by the boundaries of physical resources.
Purchasing hardware from outside and seeking IPO financing solves the problem of long-term resource increment; while striving for efficiency in algorithm architecture is to enable the existing computing power to withstand a greater call peak.
The former requires a long delivery cycle and capital operation, while the latter can immediately release capacity in existing businesses as long as the technology matures.Under such tight resource constraints, it is feasible to continue to maintain the old flagship, but the price/performance ratio is extremely low.
When lighter and cheaper models are already capable of most core tasks, the few single advantages left by the old Pro are really hard to convince the team to continue to allocate precious clusters to support it.Founder Liang Wenfeng had made this trade-off clear when communicating with investors before. He clearly puts long-term AGI exploration above short-term profit maximization, believes that the open source model and commercialization are completely self-consistent, and bluntly states that computing power is the core bottleneck of current business expansion.
Looking back at this "flash model change" from this perspective, I prefer a strategic explanation: DeepSeek's goal is to push the scale of existing services to the extreme as soon as possible while promoting the next generation of frontier exploration.
Now that new technologies have completely restructured the computing power cost and performance curves, the product matrix must be reorganized vigorously. There is no need to force a long premium protection period for the previous generation flagships.04 DeepSeek is not waiting for the next Pro
The most eye-catching feature of this replacement is its decisiveness: DeepSeek did not even wait for V4.1 Pro to take shape.
According to the practice of conventional commercial software, it can completely retain the independent channel of the old Pro to appease those enterprise customers who rely on the old model, and at the same time launch Flash as a downgraded replacement; wait until the next generation Pro is fully mature, and then complete the step-by-step iteration step by step to make the product line look smooth and tidy.

But DeepSeek announced another arrangement: the old Flash will be taken over by V4.1 Flash, and the API traffic of the old Pro will also be forcibly merged into Flash after September 14. Although the original interface name is retained and billing is directly discounted, the underlying model core has been completely replaced.
The official still mentioned the future V4.1 Pro, but they did not wait for it and continue to keep the position of the old Pro.
DeepSeek is willing to let its new technology eliminate its old flagship first. The amount of investment in research and development and the fact that it has only been online for a month have not become a reason to protect it.However, just because the interface name remains unchanged does not mean that developers can sit back and relax.
Many companies have previously fine-tuned the Prompt framework, system settings and multi-step tool calling links around the old Pro, and most likely need to re-run regression tests on the new model.
If the performance of certain tasks that heavily rely on Pro long text or plain text reasoning declines, the calls from the community and corporate customers to retain the old channel also have practical rationality.This poses a problem for the next-generation Pro that has not yet been unveiled: if it wants to provide independent services at a higher price again, it must come up with capabilities that are difficult to accomplish with Flash and worth the extra payment for users, not just better than the old Pro.
Comments