# Kimi 开源 FlashKDA 注意力核实现

- 来源：Kimi.ai (@Kimi_Moonshot)
- 发布时间：2026-07-27 23:25
- AIHOT 分数：62
- AIHOT 链接：https://aihot.virxact.com/items/cms3dxit00besro3fhbbott26
- 原文链接：https://x.com/Kimi_Moonshot/status/2081762799202746420

## AI 摘要

我们已开源 FlashKDA，这是我们基于 CUTLASS 的高性能 Kimi Delta Attention 核实现。
它在 H20 上相比 flash-linear-attention 基线实现了 1.72 倍至 2.22 倍的预填充加速，并可作为 flash-linear-attention 的即插即用后端。
在 GitHub 上探索：http://github.com/MoonshotAI/FlashKDA

## 正文

We've open-sourced FlashKDA， our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels.

It delivers 1.72×-2.22× prefill speedup over the flash-linear-attention baseline on H20， and works as a drop-in backend for flash-linear-attention.

Explore on GitHub：
http://github.com/MoonshotAI/FlashKDA
