Sky Computing Lab 在 ICML 2026 展示多项加速与推理研究

Hao AI Lab · @haoailab · X·2026-07-06 09:26·57天前
AI 导读

Sky Computing Lab 在 ICML 2026 首尔发布五项研究:MADQA 基准(2250 道人工问题、800 份 PDF)评估多模态智能体;d3LLM 通过伪轨迹蒸馏实现扩散 LLM 最高 10 倍速、自回归模型 5 倍速;Jacobi Forcing 将预训练自回归模型转为因果并行解码器,编码/数学任务达 4 倍真实加速;OnlineSpec 利用验证反馈在线调整草稿模型,平均提速 24%;Attn-QAT 实现 4-bit 注意力量化,在 RTX 5090 上加速 1.5 倍。

Hao AI Lab@haoailab
37AI 编辑部评分,满分 100

Sky Computing Lab 在 ICML 2026 展示多项加速与推理研究

2026-07-06 09:26· 57天前
AI 导读

Sky Computing Lab 在 ICML 2026 首尔发布五项研究:MADQA 基准(2250 道人工问题、800 份 PDF)评估多模态智能体;d3LLM 通过伪轨迹蒸馏实现扩散 LLM 最高 10 倍速、自回归模型 5 倍速;Jacobi Forcing 将预训练自回归模型转为因果并行解码器,编码/数学任务达 4 倍真实加速;OnlineSpec 利用验证反馈在线调整草稿模型,平均提速 24%;Attn-QAT 实现 4-bit 注意力量化,在 RTX 5090 上加速 1.5 倍。

Thrilled to share our poster sessions at #ICML2026 in Seoul! ☕️🇰🇷

Stop by to chat with our lab members and collaborators about parallel decoding, diffusion LLMs, speculative decoding, video sparse attention, quantization, and more. We’re excited to connect!

1️⃣ Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections 📜 TL;DR: This work introduces MADQA, a benchmark with 2,250 human-authored questions grounded in 800 heterogeneous PDF documents, to test whether multimodal agents truly reason strategically or rely on brute-force search. The study finds that even the best agents can match human searchers in raw accuracy but still fail to close a nearly 20% gap to oracle performance. 📍Time and Location: Tue, Jul 7, 2026 • 10:30 AM – 10:45 AM KST, HALL B2

2️⃣ d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation 📜 TL;DR: d3LLM improves the accuracy–parallelism trade-off in diffusion LLMs through pseudo-trajectory distillation with multi-block decoding, achieving up to 10x speedup over dLLMs and 5x speedup over AR models with little accuracy drop. 📍Time and Location: Tue, Jul 7, 2026 • 2:00 PM – 3:45 PM KST, HALL A #2505

3️⃣ Fast and Accurate Causal Parallel Decoding using Jacobi Forcing 📜 TL;DR: Jacobi Forcing turns pretrained AR models into causal parallel decoders while preserving AR-level quality, achieving up to 4.0x wall-clock speedup on coding and math benchmarks, by leveraging higher-quality drafts as block sizes scale. 📍Time and Location: Wed, Jul 8, 2026 • 2:30 PM – 4:15 PM KST, HALL A #2207

4️⃣ When Drafts Evolve: Speculative Decoding Meets Online Learning 📜 TL;DR: OnlineSpec uses the verification feedback already produced by speculative decoding to continuously adapt draft models online. Grounded in dynamic regret minimization, it improves speculative decoding acceleration and achieves up to 24% speedup across 7 benchmarks and 3 foundation models. 📍Time and Location: Wed, Jul 8, 2026 • 2:30 PM – 4:15 PM KST, HALL A #1600

5️⃣ Attn-QAT: 4-Bit Attention With Quantization-Aware Training 📜 TL;DR: Attn-QAT makes 4-bit attention practical by matching low-precision attention-score recomputation during training. Across diffusion and language models, it recovers FP4 attention quality drops and delivers up to 1.5× speedup on an RTX 5090. 📍Time and Location: Thu, Jul 9, 2026 • 10:30 AM – 12:15 PM KST, HALL A #2800