🚀 Today, we're releasing INT4 and FP4 (MXFP4) variants of Ling-3.0-flash.
Both run end to end on a single NVIDIA DGX Spark via our Spark-adapted SGLang path.
For FP4, W4A16 is the stable default, while W4A8 is tuned for higher throughput.
The efficiency and accuracy of the quantized models are still evolving. We will continue updating the inference implementation and quantized weights over time.