Great to see the community pushing MiniCPM5-1B inference further!🥳
@VukRosic99 has open-sourced an FP8 serving optimization that achieves significant speedups on RTX 3060 without modifying the official checkpoint. The reproducible benchmarks and practical exploration of efficient local inference are great contributions to the community.👍
We would love to see more results across different NVIDIA GPUs and workloads.😍