Rohan Paul@rohanpaul_ai
46AI 编辑部评分,满分 100
2026-08-04 09:54· 30分钟前
跳到正文
AI 摘要

atomic.chat 发布 DeepSeek V4 Flash 0731 的 14 款压缩量化构建,从无损 BF16 到 1-bit,提供 GGUF 格式供本地推理。其用 KL 散度对比未压缩权重,结果显示 3-bit 以上接近原版,以下则快速劣化。面向 128GB 硬件推荐的 AD-IQ2_M,token 匹配率达 83.6%。

atomic【.】chat just released 14 compressed quantized builds of DeepSeek V4 Flash 0731.

From lossless BF16 to 1-bit, GGUF versions for local inference runtimes.

This release measures KL divergence against the uncompressed weights instead, which asks how far the whole probability distribution drifts at every token.

The verdict is that everything above 3 bits is close to the original, and everything below falls apart fast.

The option it recommends for 128GB hardware is AD-IQ2_M. It matches the original's token choice 83.6% of the time, measured against all other V4 Flash GGUFs in the community.

@atomic_chat_hq is a desktop app that runs LLMs locally.

atomic.chatRun DeepSeek V4 Flash 0731 locally 🐳 We released 14 quants on Hugging Face, from lossless BF16 to 1-bit AD-IQ2_M is the best fit for 128GB hardware. It matches...
Rohan Paul · @rohanpaul_ai · X·2026-08-04 09:54·30分钟前
在 X 看原推· x.com(在新标签页打开)
AI 摘要

atomic.chat 发布 DeepSeek V4 Flash 0731 的 14 款压缩量化构建,从无损 BF16 到 1-bit,提供 GGUF 格式供本地推理。其用 KL 散度对比未压缩权重,结果显示 3-bit 以上接近原版,以下则快速劣化。面向 128GB 硬件推荐的 AD-IQ2_M,token 匹配率达 83.6%。

atomic【.】chat just released 14 compressed quantized builds of DeepSeek V4 Flash 0731.

From lossless BF16 to 1-bit, GGUF versions for local inference runtimes.

This release measures KL divergence against the uncompressed weights instead, which asks how far the whole probability distribution drifts at every token.

The verdict is that everything above 3 bits is close to the original, and everything below falls apart fast.

The option it recommends for 128GB hardware is AD-IQ2_M. It matches the original's token choice 83.6% of the time, measured against all other V4 Flash GGUFs in the community.

@atomic_chat_hq is a desktop app that runs LLMs locally.

atomic.chatRun DeepSeek V4 Flash 0731 locally 🐳 We released 14 quants on Hugging Face, from lossless BF16 to 1-bit AD-IQ2_M is the best fit for 128GB hardware. It matches...
在 X 查看原推x.com(在新标签页打开)