OpenBMB · @OpenBMB · X·2026-08-18 15:34·18天前
AI 导读

面壁智能点赞社区贡献:@VukRosic99 开源了 MiniCPM5-1B 的 FP8 服务优化,不改官方 checkpoint 即在 RTX 3060 上实现显著加速。其自主研究系统将推理速度从约 165 tokens/秒提升至 287 tokens/秒,提速近 80%,并公开了可复现基准。面壁期待更多 GPU 与负载下的测试结果。

OpenBMB@OpenBMB
26AI 编辑部评分,满分 100
2026-08-18 15:34· 18天前
AI 导读

面壁智能点赞社区贡献:@VukRosic99 开源了 MiniCPM5-1B 的 FP8 服务优化,不改官方 checkpoint 即在 RTX 3060 上实现显著加速。其自主研究系统将推理速度从约 165 tokens/秒提升至 287 tokens/秒,提速近 80%,并公开了可复现基准。面壁期待更多 GPU 与负载下的测试结果。

Great to see the community pushing MiniCPM5-1B inference further!🥳

@VukRosic99 has open-sourced an FP8 serving optimization that achieves significant speedups on RTX 3060 without modifying the official checkpoint. The reproducible benchmarks and practical exploration of efficient local inference are great contributions to the community.👍

We would love to see more results across different NVIDIA GPUs and workloads.😍

Vuk Rosić 武克My AI startup made an LLM nearly 80% faster. An autonomous research system took MiniCPM5-1B from about 165 to 287 tokens/second on an RTX 3060. I show what work...