Atomic released 14 quants of "DeepSeek V4 Flash 0731" on Huggingface with the best "quality vs size" performance.
DeepSeek-V4-Flash is a 284B parameter mixture-of-experts model. It is quantization-aware-trained; the official checkpoint already stores its routed experts in MXFP4 and everything else in FP8 or BF16.
The AD-IQ2_M version is the best for testing on 128GB hardware!