Objectively one of the coolest papers in LLM jailbreak world but most people didn't read the actual paper to know that the encrypted reasoning trace hack doesn't affect Fable and doesn't explain any purported distillation of that model.
AI 导读
Deedy Das 指出,LLM 越狱领域最酷的论文之一常被误读:利用 API 漏洞提取前沿模型隐藏推理的加密推理追踪攻击,对 Fable 模型无效,也无法解释该模型的所谓蒸馏。此前研究者称已通过该漏洞实现推理 token 计数与 API 计费 thinking tokens 1:1 匹配。
Objectively one of the coolest papers in LLM jailbreak world but most people didn't read the actual paper to know that the encrypted reasoning trace hack doesn't affect Fable and doesn't explain any purported distillation of that model.
We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We v...
来源:Deedy· x.com