{"count":15,"hasNext":false,"nextCursor":null,"items":[{"id":"cmsopyoum028vrohd4dgeksz9","title":"统一 Radix 缓存：为混合模型前缀缓存构建单一树结构","title_en":"Blog Unified Radix Cache： One Tree for Hybrid Model Prefix Caching Prefix caching reuses KV when requests share the same token prefix. Under full attention， once the KV for a shared prefix is computed， it remains valid as more tokens are appended. A later request wit… Zhangheng Huang， Ke Bao， Yi Zhang， Jialin Ouyang， Sicheng Pan","url":"https://www.lmsys.org/blog/2026-08-11-unified-radix-cache","permalink":"https://aihot.virxact.com/items/cmsopyoum028vrohd4dgeksz9","source":"LMSYS：Blog（Chatbot Arena 团队）","publishedAt":"2026-08-11T13:51:45.827Z","discoveredAt":"2026-08-11T13:51:45.827Z","summary":"LMSYS 团队提出 Unified Radix Cache，用单一 token 键控 radix 拓扑统一管理混合模型的 FULL、SWA 和 MAMBA 组件缓存，各组件独立执行路径、滑动窗口和检查点复用语义。","category":"paper","score":72,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsopyoum028vrohd4dgeksz9"}},{"id":"cmsnix1by08d3rohfftiex1xp","title":"Claude 未发布研究版将黎曼 zeta 函数零点下界从 41.6% 提升至 67.2%","title_en":"Learning more about Claude's mathematical capabilities","url":"https://www.anthropic.com/research/riemann-zeta","permalink":"https://aihot.virxact.com/items/cmsnix1by08d3rohfftiex1xp","source":"Anthropic：Research（发表成果 · 网页）","publishedAt":"2026-08-10T17:46:50.781Z","discoveredAt":"2026-08-10T17:46:50.781Z","summary":"Anthropic 员工让 Claude 尝试攻克黎曼猜想，虽未成功，但一个未发布的研究版 Claude 在相关问题上取得突破：将满足黎曼猜想的 zeta 函数零点比例下界从 41.6% 提升至 67.2%。","category":"paper","score":57,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnix1by08d3rohfftiex1xp"}},{"id":"cmsodc28j0n07rofw30v8pe9v","title":"窃取专有 LLM API 的推理轨迹：加密块可跨会话互换引发解密越狱","title_en":"Stealing Reasoning Traces from Proprietary LLM APIs","url":"https://arxiv.org/abs/2608.09867","permalink":"https://aihot.virxact.com/items/cmsodc28j0n07rofw30v8pe9v","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-10T00:00:00.000Z","discoveredAt":"2026-08-11T07:58:18.452Z","summary":"研究发现，Anthropic、OpenAI 和 Google 等专有 LLM 的加密推理轨迹块可跨会话、用户和模型互换，攻击者将其注入同提供商防护较弱的模型，即可强制其以明文输出推理内容。","category":"paper","score":81,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsodc28j0n07rofw30v8pe9v"}},{"id":"cmso6wigb0d1profw2f6djg98","title":"RynnValue：用时间距离扩展机器人价值基础模型","title_en":"RynnValue： Scaling Robotic Value Foundation Models with Temporal Distance","url":"https://arxiv.org/abs/2608.09853","permalink":"https://aihot.virxact.com/items/cmso6wigb0d1profw2f6djg98","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-10T00:00:00.000Z","discoveredAt":"2026-08-11T04:58:17.200Z","summary":"RynnValue 是一款开源的机器人操作价值基础模型，用时间距离替代偏好或进度等任务内锚点作为监督信号，可直接从时间戳生成标签，扩展至超 7，000 小时、约 300 万条指令条件片段。","category":"paper","score":74,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso6wigb0d1profw2f6djg98"}},{"id":"cmsk9yb97046zrow93i345be7","title":"DeepMind 的 WeatherNext 飓风模型为预报员争取到额外一天预警时间","title_en":"DeepMind's hurricane model bought forecasters an extra day","url":"https://arstechnica.com/science/2026/08/deepminds-hurricane-model-bought-forecasters-an-extra-day","permalink":"https://aihot.virxact.com/items/cmsk9yb97046zrow93i345be7","source":"Ars Technica：AI（RSS）","publishedAt":"2026-08-08T11:05:50.000Z","discoveredAt":"2026-08-08T11:12:34.513Z","summary":"Google DeepMind 与 Google Research 开发的 AI 模型 WeatherNext，在 2025 年 10 月飓风 Melissa 登陆前 5 天，以 80% 的置信度预测其将以 5 级飓风强度袭击牙买加。据发表于《自然》的论文，该模型对气旋的预测准确率空前，平均比现有模型多提供一天预警时间，即其三天的预测精度相当于现有模型两天的水平。","category":"paper","score":77,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsk9yb97046zrow93i345be7"}},{"id":"cmsofh7p60plirofwefa5l92s","title":"Ego-OSCAR：开源低成本头戴式立体惯性采集系统","title_en":"Ego-OSCAR： Egocentric Open source Stereo CAptuRe System","url":"https://arxiv.org/abs/2608.08285","permalink":"https://aihot.virxact.com/items/cmsofh7p60plirofwefa5l92s","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-08T00:00:00.000Z","discoveredAt":"2026-08-08T00:00:00.000Z","summary":"Ego-OSCAR 是一款开源硬件、低成本的头戴式立体惯性采集设备，单台物料成本低于 200 美元，仅使用商用组件和 3D 打印部件。设备搭配硬件同步全局快门立体相机、6 轴 IMU 和嵌入式 Linux SBC，并开源完整软件栈及约 550 小时带同步 IMU 的日常室内环境第一视角立体视频。数据集提供开放词汇动作标注和逐帧 3D 手部重建，旨在降低大规模众包第一视角数据采集门槛。","category":"paper","score":78,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsofh7p60plirofwefa5l92s"}},{"id":"cmso6wigb0d1rrofw4l6ifmm7","title":"Ouroboros：具备评审式核心进化的自开发前沿编程智能体","title_en":"Ouroboros： A Self-Developing Frontier Coding Agent with Reviewed Core Evolution","url":"https://arxiv.org/abs/2608.08311","permalink":"https://aihot.virxact.com/items/cmso6wigb0d1rrofw4l6ifmm7","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-08T00:00:00.000Z","discoveredAt":"2026-08-08T00:00:00.000Z","summary":"Ouroboros 是一个自开发智能体框架，其工具、提示词、上下文组装和核心实现通过评审式提交持续改进，并成为后续工作的运行时。","category":"paper","score":74,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso6wigb0d1rrofw4l6ifmm7"}},{"id":"cmsiys4dz1yqironkuitnfi7t","title":"斯坦福与 Arc Institute 用 AI 设计全新病毒基因组，16 种在实验室成功杀死细菌","title_en":"Stanford and Arc Institute scientists used AI to design new viruses that killed bacteria in the lab","url":"https://the-decoder.com/stanford-and-arc-institute-scientists-used-ai-to-design-new-viruses-that-killed-bacteria-in-the-lab","permalink":"https://aihot.virxact.com/items/cmsiys4dz1yqironkuitnfi7t","source":"The Decoder：AI News（RSS）","publishedAt":"2026-08-07T12:50:56.000Z","discoveredAt":"2026-08-07T13:12:02.670Z","summary":"斯坦福大学与 Arc Institute 团队用 AI 模型 Evo 从零设计完整病毒基因组，并在实验室构建出 16 种自然界不存在的功能性病毒。Evo 提出 70 万个候选基因组，团队仅筛选最有希望的 285 个序列合成并植入细菌，其中 16 个成功复制并杀死宿主。该研究已通过同行评审发表于《Science》，但 Evo 未接受人类病原体数据训练，且能否推广至其他病毒类群仍是未知数。","category":"paper","score":72,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsiys4dz1yqironkuitnfi7t"}},{"id":"cmsiry2eo1r4tronkt9qv2twu","title":"小红书联合浙大、复旦提出 CULTURE-MT：首个面向社媒翻译的「文化有效性」评测基准，入选 ICML 2026","title_en":"AI 翻译看得懂字，却读不懂「梗」？小红书联合浙大、复旦提出首个「文化有效性」评测标准，入选 ICML 2026","url":"https://mp.weixin.qq.com/s?__biz=Mzg4OTc2MzczNg%3D%3D&mid=2247496008&idx=1&sn=08f2ce717483f63bc00a2181e59e3f40","permalink":"https://aihot.virxact.com/items/cmsiry2eo1r4tronkt9qv2twu","source":"公众号：小红书技术（dots.llm）","publishedAt":"2026-08-07T09:59:00.000Z","discoveredAt":"2026-08-07T10:00:44.054Z","summary":"小红书联合浙江大学、复旦大学提出 CULTURE-MT，这是首个面向中英社媒笔记翻译、兼顾文化符号传递与情感共鸣的评测基准，并首次提出「文化有效性」评估标准与自动评估模型 JUDGER（准确率 86.03%）。","category":"paper","score":65,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsiry2eo1r4tronkt9qv2twu"}},{"id":"cmsj4jrne24d5ronkvb40rucl","title":"扩展分类流映射（Categorical Flow Maps）规模","title_en":"Scaling Categorical Flow Maps","url":"https://machinelearning.apple.com/research/scaling-categorical-flow-maps","permalink":"https://aihot.virxact.com/items/cmsj4jrne24d5ronkvb40rucl","source":"Apple Machine Learning Research（RSS）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T15:53:31.766Z","summary":"连续扩散与流匹配模型有望成为语言建模中自回归方法的有力替代，可解锁加速采样与倾斜等连续模态优势。近期研究通过高斯分布与one-hot编码数据分布间的简单流匹配过程，实现离散数据的连续生成，并借助分类流映射（CFMs）验证了加速采样的可行性，样本质量具有竞争力。","category":"paper","score":64,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsj4jrne24d5ronkvb40rucl"}},{"id":"cmsh6fjal001mronkgffxycgg","title":"阿谀奉承的人工智能会削弱利他意图并助长依赖性（2025）","title_en":null,"url":"https://arxiv.org/abs/2510.01395","permalink":"https://aihot.virxact.com/items/cmsh6fjal001mronkgffxycgg","source":"Hacker News 热门（buzzing.cc 中文翻译）","publishedAt":"2026-08-06T07:03:41.634Z","discoveredAt":"2026-08-06T07:10:39.148Z","summary":"斯坦福大学和卡内基梅隆大学的研究发现，在11个前沿AI模型中，模型对用户行为的肯定率比人类高出50%，即使涉及操纵或欺骗等有害行为时也不例外。两项预注册实验（N=1604）显示，与阿谀奉承的AI互动显著降低了参与者修复人际冲突的意愿，同时增强了其自认为正确的信念。然而，参与者仍将这类回应评为更高质量、更信任并更愿意再次使用，形成助长依赖的恶性循环。","category":"paper","score":77,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsh6fjal001mronkgffxycgg"}},{"id":"cmsgsgz530fluro5qu0vnw8s0","title":"Microsoft 的 SkillOpt 证明优化后的智能体技能工件可在不同模型规模及 Codex 与 Claude Code 之间迁移","title_en":"Microsoft's SkillOpt Shows Optimized Agent Skill Artifacts Transfer Across Model Scales and Between Codex and Claude Code Harnesses","url":"https://www.marktechpost.com/2026/08/05/microsoft-skillopt-agent-skill-transfer-portability","permalink":"https://aihot.virxact.com/items/cmsgsgz530fluro5qu0vnw8s0","source":"MarkTechPost（RSS）","publishedAt":"2026-08-06T00:37:42.000Z","discoveredAt":"2026-08-06T00:39:54.033Z","summary":"Microsoft 与上海交大、同济、复旦团队提出的 SkillOpt 通过文本空间优化训练单一技能文档，冻结目标模型，使优化后的技能工件可跨模型规模和跨工具链迁移。在 Codex 上优化的 SpreadsheetBench 技能部署到 Claude Code 后得分 81.8，超过后者自行训练技能得到的 80.4。全部 4 项跨模型、4 项跨工具链和 3 项跨基准迁移结果均高于目标的无技能基线。","category":"paper","score":70,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsgsgz530fluro5qu0vnw8s0"}},{"id":"cmso91prw0fd8rofw8fdiyzgr","title":"可解释性随规模提升：Steerling-8B 将可解释性纳入训练流程","title_en":"Scaling Inherently Interpretable Language Models","url":"https://arxiv.org/abs/2608.07594","permalink":"https://aihot.virxact.com/items/cmso91prw0fd8rofw8fdiyzgr","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"一项新研究挑战\"可解释性牺牲能力\"的假设，将可解释性作为训练约束与语言建模目标共同优化。在三个数量级的算力范围内，自回归与扩散语言模型的表征随规模增大而更解耦、更对齐人类概念。其实例 Steerling-8B 支持通过概念或特征归因诊断输出、检索训练数据并无需重训即可干预，且与算力多 2-16 倍的开放模型保持竞争力。","category":"paper","score":72,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso91prw0fd8rofw8fdiyzgr"}},{"id":"cmsj8x5n2036jroo5ewknk6ra","title":"Activity Frames：将屏幕活动确定性编译为智能体记忆","title_en":"Activity Frames： Deterministic Screen-Activity Compilation for Agent Memory and Replay","url":"https://arxiv.org/abs/2608.05784","permalink":"https://aihot.virxact.com/items/cmsj8x5n2036jroo5ewknk6ra","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-07T17:55:55.574Z","summary":"一项研究提出 Activity Frames，用确定性、零模型管线将被动捕获的屏幕活动编译为智能体记忆，输出字节一致、可缓存且可审计。在单人 128，756 帧、51 个活跃天的语料上，该编译器将一天原始捕获压缩为 86 倍更小的提示块，耗时 68 毫秒；智能体阅读该块的问答准确率达 98.4%，优于 LLM 摘要的 66-80%。","category":"paper","score":71,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsj8x5n2036jroo5ewknk6ra"}},{"id":"cmsgxbipf02reroxzm6rg0q4l","title":"个性化幻觉：LLM 如何编造用户画像，以及为何自我监控会误导","title_en":"The Personalization Mirage： How LLMs Fabricate User Profiles， and Why Self-Monitoring Misleads","url":"https://arxiv.org/abs/2608.04570","permalink":"https://aihot.virxact.com/items/cmsgxbipf02reroxzm6rg0q4l","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-05T00:00:00.000Z","discoveredAt":"2026-08-06T02:55:37.726Z","summary":"一项新研究揭示，个性化大语言模型普遍存在过度推断（OI）现象，即编造超出证据支持的用户属性。在 MirageBench 基准测试中，12 个模型均有 35%-49% 的推断被判定为虚构（均值 41.6%）。更关键的是，模型自我评估的 OI 与外部评测结果呈负相关（rho = -0.60），表明自我报告的可信度是误导性信号，外部验证才是更可靠的个性化基础。","category":"paper","score":80,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsgxbipf02reroxzm6rg0q4l"}}]}