# LLM强化学习：单调推理策略是真正目标

- 来源：AK (@_akhaliq)
- 发布时间：2026-07-07 01:58
- AIHOT 分数：48
- AIHOT 链接：https://aihot.virxact.com/items/cmr9jd9tw008xihe8jyge0fc1
- 原文链接：https://x.com/_akhaliq/status/2074191328590671928

## AI 摘要

优化训练策略的幻象

单调推理策略才是 LLM 强化学习的真正目标

## 正文

The Mirage of Optimizing Training Policies

Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
