Somebody is running Kimi K3 full on 16x GB10 cluster at 20+tps average (llama-benchy coherent corpus) 38tps peak, 750tps prefill.
on a MikroTik Switch CRS804-4DDQ with 4x 400-to-4x100gbit breakout cables running with dspark
----
reddit .com/r/LocalLLaMA/comments/1vfl525/kimi_k3_full_model_running_on_16x_gb10_cluster_at/