Thrilled to share our poster sessions at #ICML2026 in Seoul! ☕️🇰🇷
Stop by to chat with our lab members and collaborators about parallel decoding, diffusion LLMs, speculative decoding, video sparse attention, quantization, and more. We’re excited to connect!
1️⃣ Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections 📜 TL;DR: This work introduces MADQA, a benchmark with 2,250 human-authored questions grounded in 800 heterogeneous PDF documents, to test whether multimodal agents truly reason strategically or rely on brute-force search. The study finds that even the best agents can match human searchers in raw accuracy but still fail to close a nearly 20% gap to oracle performance. 📍Time and Location: Tue, Jul 7, 2026 • 10:30 AM – 10:45 AM KST, HALL B2
2️⃣ d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation 📜 TL;DR: d3LLM improves the accuracy–parallelism trade-off in diffusion LLMs through pseudo-trajectory distillation with multi-block decoding, achieving up to 10x speedup over dLLMs and 5x speedup over AR models with little accuracy drop. 📍Time and Location: Tue, Jul 7, 2026 • 2:00 PM – 3:45 PM KST, HALL A #2505
3️⃣ Fast and Accurate Causal Parallel Decoding using Jacobi Forcing 📜 TL;DR: Jacobi Forcing turns pretrained AR models into causal parallel decoders while preserving AR-level quality, achieving up to 4.0x wall-clock speedup on coding and math benchmarks, by leveraging higher-quality drafts as block sizes scale. 📍Time and Location: Wed, Jul 8, 2026 • 2:30 PM – 4:15 PM KST, HALL A #2207
4️⃣ When Drafts Evolve: Speculative Decoding Meets Online Learning 📜 TL;DR: OnlineSpec uses the verification feedback already produced by speculative decoding to continuously adapt draft models online. Grounded in dynamic regret minimization, it improves speculative decoding acceleration and achieves up to 24% speedup across 7 benchmarks and 3 foundation models. 📍Time and Location: Wed, Jul 8, 2026 • 2:30 PM – 4:15 PM KST, HALL A #1600
5️⃣ Attn-QAT: 4-Bit Attention With Quantization-Aware Training 📜 TL;DR: Attn-QAT makes 4-bit attention practical by matching low-precision attention-score recomputation during training. Across diffusion and language models, it recovers FP4 attention quality drops and delivers up to 1.5× speedup on an RTX 5090. 📍Time and Location: Thu, Jul 9, 2026 • 10:30 AM – 12:15 PM KST, HALL A #2800