One DGX Spark, 38.7 tok/s, full 256K context, and working software in one shot-huge thanks to @sudoingX for a week of rigorous testing that shows what Ling-3.0-flash can deliver as a truly local 124B-class model. 🙌
AI 导读
蚂蚁百灵 Ling-3.0-flash 124B 在单台 DGX Spark 上实现 38.7 tok/s(官方 int4 量化),社区 GGUF 为 35.2 tok/s,速度达 DeepSeek V4 Flash 同硬件成绩的 2.4 倍。该模型支持完整 256K 上下文,并能在智能体循环中一次性写出可运行代码,被评测者视为单机最快的 124B 级本地模型。
One DGX Spark, 38.7 tok/s, full 256K context, and working software in one shot-huge thanks to @sudoingX for a week of rigorous testing that shows what Ling-3.0-flash can deliver as a truly local 124B-class model. 🙌
ive been playing with @AntLingAGI's ling 3.0 flash 124b for a week now, and i need to publish some numbers and performance i found valuable from this lab. a wee...
来源:Ant Ling· x.com