The Hangzhou-headquartered lab that’s roiled the industry with inventive techniques said its latest model offered “more efficient architecture” at lower prices. (Photo: Bloomberg)
DeepSeek rolled out an AI model that charges as little as a fraction of a cent per million tokens, ramping up the pressure on rivals from Anthropic to Z.AI Co.
The Chinese startup unveiled the V4.1 Flash on Thursday, a slimmed-down platform it claims outperformed mainstays such as Moonshot’s Kimi K3, yet offers a steep discount to the competition. Shares in MiniMax and Z.ai plunged more than 8% in Hong Kong. Alibaba, the e-commerce giant that’s pivoting into artificial intelligence, slid more than 2%.
DeepSeek’s latest move highlights the intensifying price battle between Chinese open-weight models and their US counterparts, at a time Anthropic and ChatGPT-developer OpenAI are preparing to go public. The Chinese firm is betting that good-enough yet ultra-cheap models can beat top-tier models to drive the next phase of AI adoption. The Hangzhou-headquartered lab that’s roiled the industry with inventive techniques said its latest model offered “more efficient architecture” at lower prices.
DeepSeek’s discounts are reshaping how developers, startups and enterprises budget for AI agents that can run hours at a stretch, completing tasks autonomously while working toward a targeted goal. The industry’s focus has shifted toward AI agents that can write code over long sessions, summon tools or browse the web in search of information — all with little or no human intervention.
The advent of cheap yet high-powered Chinese models has created what Artificial Analysis terms a DeepSeek “death zone.” To compete, AI contenders must now either beat DeepSeek and its Chinese peers on price, or soundly surpass them on capability.
DeepSeek said it’s using a “causal encoder-decoder” design that activates only a tiny part of the capacity of its 552-billion-parameter model at a time. That’s a fraction of what the larger V4-Pro required. In benchmarks, the new model beats the V4-Pro on coding and agentic tasks, though it trails the flagship models of Anthropic and OpenAI.
Starting Sept. 14, the startup is retiring the V4-Pro by automatically rerouting all inference tasks to the V4.1 Flash, which will in turn be billed at the cheaper rates.
Tags: