Google has open-sourced their TPU Raiden inference optimization library. This is the equivalent layer of the stack to NVIDIA NIXL, where it provides KVCache transfer between prefill & decode instances & has primitives for KVCache offloading movements! It is great to see Google open-source & externalize more and more of their TPU stack!
AI 导读
谷歌已开源其 TPU Raiden 推理优化库。这是与 NVIDIA NIXL 对等的栈层,提供预填充与解码实例之间的 KVCache 传输,并包含 KVCache 卸载移动的原语!很高兴看到谷歌将越来越多的 TPU 技术栈开源并外部化!
Google has open-sourced their TPU Raiden inference optimization library. This is the equivalent layer of the stack to NVIDIA NIXL, where it provides KVCache transfer between prefill & decode instances & has primitives for KVCache offloading movements! It is great to see Google open-source & externalize more and more of their TPU stack!
来源:SemiAnalysis· x.com