# 语言模型需要睡眠：长时任务新解法

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-08-13 22:30
- AIHOT 分数：36
- AIHOT 链接：https://aihot.virxact.com/items/cmsrn020i003grosrncj4wi4a
- 原文链接：https://x.com/rohanpaul_ai/status/2087909542571172089

## AI 摘要

论文提出为语言智能体增加“睡眠阶段”，通过离线重读上下文并将有用信息写入固定大小记忆层，再清空注意力缓存，使长时运行更高效。该方法在元胞自动机、图查找和GSM-Infinite等任务上验证，睡眠越长性能越好，尤其利于深度推理。核心意义在于长时智能体无需无限扩大上下文，可整合关键信息并安全遗忘原始token。

## 正文

Long-running language agents may work better if they periodically stop to consolidate memory.

The problem is that today's transformer agents get slower and more expensive as their context grows, because attention has to keep checking more past tokens.

The usual fix for long context is to keep more tokens nearby, but that turns every next-token prediction into a larger search through the past.

The sharper idea here is that memory is not only storage.

Sometimes the hard part is converting a messy stretch of experience into a state that can actually be used later.

So the paper's idea is to add a sleep phase, where the model pauses, rereads recent context several times, writes the useful information into fixed-size memory layers, and then clears the short-term attention cache.

During sleep, the model runs several offline passes over recent context, writes the result into fast weights inside its state-space blocks, then clears the attention cache.

This means the model pays extra compute while sleeping, not while answering, so normal prediction can still happen with 1 forward pass.

The authors test this on cellular automata, graph lookup, and GSM-Infinite math problems, where the model must use old information that is no longer sitting in its attention cache.

The main result is that longer sleep improves performance, especially on harder cases that need deeper reasoning rather than just remembering a fact.

The big deal is that long-horizon agents may not need to carry bigger and bigger raw context forever, because they can consolidate the important parts and safely forget the raw tokens.

- arxiv. org/abs/2605.26099

Title: "Language Models Need Sleep"
