Chinese AI company DeepSeek has quietly released the API version of DeepSeek-V4-Flash, with a focus on improving its agent capabilities. The main upgrade is giving the model stronger abilities to use tools, operate software environments and complete multi-step tasks. DeepSeek's latest model approaches the level of leading frontier models on several agent and coding benchmarks, while keeping inference costs dramatically lower. V4-Flash is priced at just ¥1 per million input tokens and ¥2 per million output tokens. For repeated agent workflows using cached context, the cost drops to only ¥0.02 per million input tokens (around $0.003), making long-running AI agent tasks far cheaper to operate. For developers, the bigger change is ecosystem integration. V4-Flash now supports OpenAI-compatible Responses API and can work with the Codex ecosystem, allowing developers to call the model through Codex CLI, ChatGPT desktop and VS Code plugins. Currently, only V4-Flash supports Responses API integration. V4-Pro users will have to wait until early August for similar support. There is one important limitation: this update is currently available only through the V4-Flash API. The DeepSeek app, web interface, and V4-Pro API experience remain unchanged.
DeepSeek 悄然发布 V4-Flash API 版本,重点提升智能体能力,包括工具调用、软件环境操作和多步任务完成。该模型在多项智能体和编程基准上接近前沿水平,推理成本大幅降低,定价仅 ¥1/百万输入 token、¥2/百万输出 token,缓存上下文输入低至 ¥0.02/百万 token。
Chinese AI company DeepSeek has quietly released the API version of DeepSeek-V4-Flash, with a focus on improving its agent capabilities. The main upgrade is giving the model stronger abilities to use tools, operate software environments and complete multi-step tasks. DeepSeek's latest model approaches the level of leading frontier models on several agent and coding benchmarks, while keeping inference costs dramatically lower. V4-Flash is priced at just ¥1 per million input tokens and ¥2 per million output tokens. For repeated agent workflows using cached context, the cost drops to only ¥0.02 per million input tokens (around $0.003), making long-running AI agent tasks far cheaper to operate. For developers, the bigger change is ecosystem integration. V4-Flash now supports OpenAI-compatible Responses API and can work with the Codex ecosystem, allowing developers to call the model through Codex CLI, ChatGPT desktop and VS Code plugins. Currently, only V4-Flash supports Responses API integration. V4-Pro users will have to wait until early August for similar support. There is one important limitation: this update is currently available only through the V4-Flash API. The DeepSeek app, web interface, and V4-Pro API experience remain unchanged.