Same architecture, same 13B active parameters, same price, and a 10-point jump from post-training alone!
DeepSeek V4 Flash 0731 now beats the 49B-active Pro preview on agentic work while using 12% fewer tokens.
DeepSeek found a remarkable amount of untapped capability in weights it already had.
The cache price is truly absurd. One million cached input tokens costs only $0.0028. This makes recurring agents with large, stable system prompts, codebases, or documents extremely cheap (reminds me of the DeepSeek 4 tech paper!)
DeepSeek thus delivers intelligence nearly on par with Luna at around 60% lower cost per task. However, Luna is faster, multimodal, and significantly less wasteful with tokens; following the price cut, the premium is now only about 2.4 times the cost, down from the previous 12-fold difference.
This makes me super freaking excited for the next DeepSeek 4 (4.5?) pro!
The whale is back, though todays its more a dolphin.