TPU🚨 is working with the popular OSS inference optimization library Mooncake on integrating TPU with Mooncake Store. Similar to NVL72, KVCache DRAM P2P pooling will initially happen on the scale-out network via TENT instead of using ICI/NVLink. 🔥 Mooncake basically improves performance per TCO of production inference!
AI 导读
TPU🚨 正与流行的开源推理优化库 Mooncake 合作,将 TPU 与 Mooncake Store 集成。与 NVL72 类似,KVCache DRAM P2P 池化最初将通过 TENT 在横向扩展网络上进行,而非使用 ICI/NVLink。🔥 Mooncake 从根本上提升了生产推理的每 TCO 性能!
47
AI 编辑部评分,满分 100TPU🚨 正与流行的开源推理优化库 Mooncake 合作,将 TPU 与 Mooncake Store 集成。与 NVL72 类似,KVCache DRAM P2P 池化最初将通过 TENT 在横向扩展网络上进行,而非使用 ICI/NVLink。🔥 Mooncake 从根本上提升了生产推理的每 TCO 性能!
来源:SemiAnalysis· x.com