{"count":100,"hasNext":true,"nextCursor":"eyJhIjoxNzg1OTc0NDAwMDAwLCJpIjoiY21zaXJyc29wMXFvcnJvbmtna2p1dzByZyJ9","items":[{"id":"cmsopyoum028vrohd4dgeksz9","title":"统一 Radix 缓存：为混合模型前缀缓存构建单一树结构","title_en":"Blog Unified Radix Cache： One Tree for Hybrid Model Prefix Caching Prefix caching reuses KV when requests share the same token prefix. Under full attention， once the KV for a shared prefix is computed， it remains valid as more tokens are appended. A later request wit… Zhangheng Huang， Ke Bao， Yi Zhang， Jialin Ouyang， Sicheng Pan","url":"https://www.lmsys.org/blog/2026-08-11-unified-radix-cache","permalink":"https://aihot.virxact.com/items/cmsopyoum028vrohd4dgeksz9","source":"LMSYS：Blog（Chatbot Arena 团队）","publishedAt":"2026-08-11T13:51:45.827Z","discoveredAt":"2026-08-11T13:51:45.827Z","summary":"LMSYS 团队提出 Unified Radix Cache，用单一 token 键控 radix 拓扑统一管理混合模型的 FULL、SWA 和 MAMBA 组件缓存，各组件独立执行路径、滑动窗口和检查点复用语义。","category":"paper","score":72,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsopyoum028vrohd4dgeksz9"}},{"id":"cmsooqkc107ogrop2qux03yh4","title":"Google ScientistOne 论文：用\"证据链\"解决 AI 生成研究的信任问题","title_en":"Google's ScientistOne paper tackles a basic problem with AI-generated research： The result can loo…","url":"https://x.com/rohanpaul_ai/status/2087161012554510600","permalink":"https://aihot.virxact.com/items/cmsooqkc107ogrop2qux03yh4","source":"X：Rohan Paul (@rohanpaul_ai)","publishedAt":"2026-08-11T12:55:37.000Z","discoveredAt":"2026-08-11T13:17:32.640Z","summary":"Google Cloud AI Research 审计五个自主研究系统的 75 篇论文，发现每个基线都至少出现一种系统性证据失败，如编造引用、分数无法复现。ScientistOne 通过\"Chain-of-Evidence\"机制，要求引用、数值和方法主张分别追溯至检索论文、评估日志和实现产物，方可定稿。","category":"paper","score":45,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsooqkc107ogrop2qux03yh4"}},{"id":"cmsok8wbs0vfgrofw0w5ak273","title":"我国锂一次电池实现\"双高\"突破：25安时级软包金属锂一次电池研制成功","title_en":"既有\"耐力\"又有\"爆发力\"，我国锂一次电池实现\"双高\"突破","url":"https://www.ithome.com/0/988/459.htm","permalink":"https://aihot.virxact.com/items/cmsok8wbs0vfgrofw0w5ak273","source":"IT之家（RSS）","publishedAt":"2026-08-11T10:53:20.000Z","discoveredAt":"2026-08-11T11:11:47.881Z","summary":"中国科学院大连化物所团队研制出25安时级软包金属锂一次电池，低倍率下能量密度超750Wh/kg，放电倍率可提高百倍，实现10C短时大电流放电。该电池常温比功率达4487W/kg，零下40摄氏度下1C能量密度保持447Wh/kg，年自放电低于1%，正加速验证于无人机、机器人等场景。","category":"paper","score":5,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsok8wbs0vfgrofw0w5ak273"}},{"id":"cmsok8wbt0vflrofwe09cqlil","title":"Anthropic 宣布研究版 Claude 攻克黎曼猜想取得重大突破","title_en":"未公开的新模型\"立功\"，Anthropic 宣布 Claude 攻克黎曼猜想取得重大突破","url":"https://www.ithome.com/0/988/453.htm","permalink":"https://aihot.virxact.com/items/cmsok8wbt0vflrofwe09cqlil","source":"IT之家（RSS）","publishedAt":"2026-08-11T10:35:17.000Z","discoveredAt":"2026-08-11T11:11:47.881Z","summary":"Anthropic 8 月 11 日宣布，未公开的研究版 Claude 在攻克黎曼猜想中取得重要进展，将黎曼 ζ 函数临界线上零点比例的长期下界从 41.6% 提高至 67.2%。","category":"paper","score":88,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsok8wbt0vflrofwe09cqlil"}},{"id":"cmsohmy4e0si0rofwg0gmdbro","title":"新数据有限，CD缩放定律揭示算力数据交换率","title_en":"Fresh data is finite. The next scaling law finds what truly limits learning. Paper： Bridging Comp…","url":"https://x.com/dongxi_nlp/status/2087113474115514467","permalink":"https://aihot.virxact.com/items/cmsohmy4e0si0rofwg0gmdbro","source":"X：马东锡 NLP (@dongxi_nlp)","publishedAt":"2026-08-11T09:46:43.000Z","discoveredAt":"2026-08-11T09:58:46.565Z","summary":"新鲜数据是有限的。\n\n下一个缩放定律将找出真正限制学习的因素。\n\n论文：\n\n《弥合计算最优与数据最优的预训练》","category":"paper","score":33,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsohmy4e0si0rofwg0gmdbro"}},{"id":"cmsobsdw00l14rofwbm182edx","title":"\"认知公共领域\"的悲剧：AI 如何侵蚀职业专长的再生基础","title_en":"\"认知公共领域\"的悲剧","url":"https://arxiv.org/abs/2607.29380","permalink":"https://aihot.virxact.com/items/cmsobsdw00l14rofwbm182edx","source":"Hacker News 热门（buzzing.cc 中文翻译）","publishedAt":"2026-08-11T06:50:12.715Z","discoveredAt":"2026-08-11T07:15:00.980Z","summary":"一篇概念性论文提出\"认知公共领域\"框架，指出理性的 AI 采用决策可能耗尽职业更新所需的共享专长池。该框架区分\"内化精通\"与\"分布式精通\"，并提出\"验证锚链\"：有效的 AI 监督依赖于 AI 采用本身可能削弱的人类专长。早期劳动力市场与临床证据显示，高 AI 暴露行业的专长再生路径或受干扰，但采用尚属近期，最强信号来自领先行业而非所有职业。","category":"paper","score":42,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsobsdw00l14rofwbm182edx"}},{"id":"cmso5fvd50bfmrofw7cvjjaw3","title":"SWE-Bench ProMax：大规模多语言代码重构基准","title_en":"SWE-Bench ProMax Benchmarking Agents on Large-Scale Multilingual Code Refactoring paper： https://h...","url":"https://x.com/_akhaliq/status/2087026792972308750","permalink":"https://aihot.virxact.com/items/cmso5fvd50bfmrofw7cvjjaw3","source":"X：AK (@_akhaliq)","publishedAt":"2026-08-11T04:02:17.000Z","discoveredAt":"2026-08-11T04:17:21.105Z","summary":"SWE-Bench ProMax\n\n在大规模多语言代码重构上对智能体进行基准测试\n\n论文：https://huggingface.co/papers/2608.09802","category":"paper","score":51,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso5fvd50bfmrofw7cvjjaw3"}},{"id":"cmsnprq5x0gbxrohfkzj6miqn","title":"小红书发布StreamArena长时视频理解研究","title_en":"Research from XiaoHongShu Understanding video means staying present over time. StreamArena tests m…","url":"https://x.com/dongxi_nlp/status/2086916939486744671","permalink":"https://aihot.virxact.com/items/cmsnprq5x0gbxrohfkzj6miqn","source":"X：马东锡 NLP (@dongxi_nlp)","publishedAt":"2026-08-10T20:45:46.000Z","discoveredAt":"2026-08-10T20:58:40.311Z","summary":"小红书的研究\n\n理解视频意味着在时间中保持在场。\n\nStreamArena 在单一实时流中测试记忆、主动性和工具使用。StreamMind 让快速行动与深层记忆共同成长。\n\n论文：\n\nStreamArena：迈向连续、交互、长时程的智能体流式视频理解","category":"paper","score":25,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnprq5x0gbxrohfkzj6miqn"}},{"id":"cmsno789d0eubrohfw7oczyn4","title":"Claude 在黎曼猜想相关问题上取得新进展：下界从 41.6% 提升至 67.2%","title_en":"进一步了解克劳德的数学能力","url":"https://www.anthropic.com/research/riemann-zeta","permalink":"https://aihot.virxact.com/items/cmsno789d0eubrohfw7oczyn4","source":"Hacker News 热门（buzzing.cc 中文翻译）","publishedAt":"2026-08-10T19:54:08.696Z","discoveredAt":"2026-08-10T20:14:44.112Z","summary":"Anthropic 让未发布的研究版 Claude 尝试攻克黎曼猜想，虽未成功，却在相关问题上取得突破：将满足黎曼猜想的 zeta 函数零点比例下界从 41.6% 提升至 67.2%。","category":"paper","score":80,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsno789d0eubrohfw7oczyn4"}},{"id":"cmsnn4noq0dwyrohf4ep1xipi","title":"Claude 研究版突破黎曼猜想相关下界","title_en":"Anybody who still says \"AI is just a tool\" is just embarrassing themselves at this point \"My hammer…","url":"https://x.com/AISafetyMemes/status/2086900500608471302","permalink":"https://aihot.virxact.com/items/cmsnn4noq0dwyrohf4ep1xipi","source":"X：AI Safety Memes (@AISafetyMemes)","publishedAt":"2026-08-10T19:40:26.000Z","discoveredAt":"2026-08-10T19:44:44.778Z","summary":"Anthropic 让未发布的研究版 Claude 尝试黎曼猜想，虽未解决原问题，但将满足猜想的黎曼 zeta 函数零点占比下界从 41.6% 提升至 67.2%。AI 已非单纯工具，而是能独立取得数学突破的智能体。","category":"paper","score":33,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnn4noq0dwyrohf4ep1xipi"}},{"id":"cmsnn7e730e0lrohfuchoprmc","title":"MatrAIx：83亿人格智能体模拟世界","title_en":"MatrAIx Simulating the World with 8.3 Billion Persona Agents paper： https://huggingface.co/papers/...","url":"https://x.com/_akhaliq/status/2086894806932865475","permalink":"https://aihot.virxact.com/items/cmsnn7e730e0lrohfuchoprmc","source":"X：AK (@_akhaliq)","publishedAt":"2026-08-10T19:17:49.000Z","discoveredAt":"2026-08-10T19:46:52.484Z","summary":"MatrAIx\n\n用83亿人格智能体模拟世界\n\n论文：https://huggingface.co/papers/2608.04205","category":"paper","score":34,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnn7e730e0lrohfuchoprmc"}},{"id":"cmsnm4p6q0d70rohfcsvww6az","title":"鼓励性话语也能提升大语言模型表现，Claude 意外改进黎曼猜想零点下界","title_en":"It's sweet that motivational sayings like \"believe in yourself！\" not only motivate people， especiall…","url":"https://x.com/kimmonismus/status/2086889910930215413","permalink":"https://aihot.virxact.com/items/cmsnm4p6q0d70rohfcsvww6az","source":"X：Kim (@kimmonismus)","publishedAt":"2026-08-10T18:58:21.000Z","discoveredAt":"2026-08-10T19:16:46.649Z","summary":"研究发现，\"相信自己！\"这类鼓励性话语不仅能激励人，也能帮助大语言模型提升表现。Anthropic 让未发布版 Claude 尝试攻克黎曼猜想，虽未证明，却意外将满足假设的 zeta 函数零点比例下界从 41.6% 提升至 67.2%，期间协调了 60 个子智能体并形式化验证了结果。","category":"paper","score":46,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnm4p6q0d70rohfcsvww6az"}},{"id":"cmsnlheoe0cyvrohfpslr4ood","title":"Claude研究版提升黎曼零点下界","title_en":"an unreleased research version = Nothing","url":"https://x.com/dongxi_nlp/status/2086881714710679851","permalink":"https://aihot.virxact.com/items/cmsnlheoe0cyvrohfpslr4ood","source":"X：马东锡 NLP (@dongxi_nlp)","publishedAt":"2026-08-10T18:25:47.000Z","discoveredAt":"2026-08-10T18:58:40.270Z","summary":"一个未发布的研究版本 = 没什么大不了\n\n【引用 @AnthropicAI】：我们让一个未发布的研究版 Claude 尝试攻克黎曼猜想。\n\n它没有解决该猜想，但在一个相关问题上取得了进展：它将满足猜想的黎曼ζ函数零点比例的下界从 41.6% 提高到了 67.2%。\n\nhttps://www.anthropic.com/research/riemann-zeta","category":"paper","score":31,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnlheoe0cyvrohfpslr4ood"}},{"id":"cmsnl245p0c9mrohf3n3jkc3v","title":"Claude 未解黎曼猜想却改进关键下界","title_en":"Absolutely insane. This might be the clearest glimpse yet of how AI will transform scientific discov…","url":"https://x.com/kimmonismus/status/2086881395465466004","permalink":"https://aihot.virxact.com/items/cmsnl245p0c9mrohf3n3jkc3v","source":"X：Kim (@kimmonismus)","publishedAt":"2026-08-10T18:24:31.000Z","discoveredAt":"2026-08-10T18:46:46.594Z","summary":"Anthropic 让未发布版 Claude 尝试黎曼猜想，虽未解决，却将满足猜想的黎曼 zeta 函数零点比例下界从 41.6% 提升至 67.2%。Claude 协调 60 个子智能体、测试数百个想法并检索文献，还用 Lean 形式化验证了结果。这被视为自主 AI 驱动科学发现的早期范例。","category":"paper","score":52,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnl245p0c9mrohf3n3jkc3v"}},{"id":"cmsnjwvfn0bcbrohft0omk645","title":"Claude 研究版提升黎曼猜想零点下界","title_en":"sometimes all you need to do is tell Claude to keep going @jarredsumner","url":"https://x.com/trq212/status/2086876017319457181","permalink":"https://aihot.virxact.com/items/cmsnjwvfn0bcbrohft0omk645","source":"X：Thariq (@trq212)","publishedAt":"2026-08-10T18:03:09.000Z","discoveredAt":"2026-08-10T18:14:42.982Z","summary":"有时候你只需要告诉 Claude 继续下去\n\n@jarredsumner\n\n【引用 @AnthropicAI】：我们让一个未发布的研究版 Claude 尝试攻克黎曼猜想。\n\n它没有解决该猜想，但在一个相关问题上取得了进展：它将满足猜想的黎曼 zeta 函数零点比例的下界从 41.6% 提升到了 67.2%。\n\nhttps://www.anthropic.com/research/riemann-zeta","category":"paper","score":35,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnjwvfn0bcbrohft0omk645"}},{"id":"cmsnjzj8g0bdirohfnmmlfz2s","title":"Ethan Mollick 质疑 Claude 黎曼猜想提示词技巧","title_en":"Oh no， we aren't going to go back to this sort of prompting again， are we？ I would love Anthropic to…","url":"https://x.com/emollick/status/2086875820279128574","permalink":"https://aihot.virxact.com/items/cmsnjzj8g0bdirohfnmmlfz2s","source":"X：Ethan Mollick (@emollick)","publishedAt":"2026-08-10T18:02:22.000Z","discoveredAt":"2026-08-10T18:16:46.501Z","summary":"Ethan Mollick 对 Anthropic 用提示词引导 Claude 攻克黎曼猜想相关问题的做法表示质疑，称其早期实验（用稍旧模型）未发现该提示词方法能稳健奏效。Anthropic 称未发布的研究版 Claude 将满足黎曼猜想的零点占比下界从 41.6% 提升至 67.2%。","category":"paper","score":50,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnjzj8g0bdirohfnmmlfz2s"}},{"id":"cmsnix1by08d3rohfftiex1xp","title":"Claude 未发布研究版将黎曼 zeta 函数零点下界从 41.6% 提升至 67.2%","title_en":"Learning more about Claude's mathematical capabilities","url":"https://www.anthropic.com/research/riemann-zeta","permalink":"https://aihot.virxact.com/items/cmsnix1by08d3rohfftiex1xp","source":"Anthropic：Research（发表成果 · 网页）","publishedAt":"2026-08-10T17:46:50.781Z","discoveredAt":"2026-08-10T17:46:50.781Z","summary":"Anthropic 员工让 Claude 尝试攻克黎曼猜想，虽未成功，但一个未发布的研究版 Claude 在相关问题上取得突破：将满足黎曼猜想的 zeta 函数零点比例下界从 41.6% 提升至 67.2%。","category":"paper","score":57,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnix1by08d3rohfftiex1xp"}},{"id":"cmsnix2eb08durohf15yop35q","title":"Claude 研究版推进黎曼猜想相关证明","title_en":"We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis. It didn…","url":"https://x.com/AnthropicAI/status/2086867246073401655","permalink":"https://aihot.virxact.com/items/cmsnix2eb08durohf15yop35q","source":"X：Anthropic (@AnthropicAI)","publishedAt":"2026-08-10T17:28:18.000Z","discoveredAt":"2026-08-10T17:46:51.914Z","summary":"我们让一个未发布的研究版 Claude 尝试攻克黎曼猜想。\n\n它没有解决该猜想，但在一个相关问题上取得了进展：它将满足该猜想的黎曼 zeta 函数零点比例的下界从 41.6% 提高到了 67.2%。\n\nhttps://www.anthropic.com/research/riemann-zeta","category":"paper","score":53,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnix2eb08durohf15yop35q"}},{"id":"cmsngp26i06bjrohfucchbhhw","title":"Mistral 获\"基于代码实现的工具调用\"专利，涵盖沙箱执行与暂停恢复机制","title_en":"Mistral 关于\"基于代码实现的工具调用\"的专利","url":"https://patentsgazette.uspto.gov/week26/OG/html/1547-5/US12670045-20260630.html","permalink":"https://aihot.virxact.com/items/cmsngp26i06bjrohfucchbhhw","source":"Hacker News 热门（buzzing.cc 中文翻译）","publishedAt":"2026-08-10T16:11:49.602Z","discoveredAt":"2026-08-10T16:44:38.589Z","summary":"Mistral AI 获得美国专利 US 12，670，045 B1，涵盖一种基于代码实现的工具调用方法。该方法由大语言模型（LLM）生成封装工具调用的代码块，在服务器沙箱中执行，遇待处理调用时暂停并将请求发送至客户端执行，随后恢复代码块运行并替换结果。专利共 20 项权利要求，发明人为 Gabriel Vergnaud。","category":"paper","score":57,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsngp26i06bjrohfucchbhhw"}},{"id":"cmsnfmfrq059nrohf3zdfo8jo","title":"Meta 提出 Skaling law 耦合规模与数据","title_en":"Impressive new paper from Meta. （bookmark it） Scaling laws assume model size and training data act…","url":"https://x.com/omarsar0/status/2086845790983716917","permalink":"https://aihot.virxact.com/items/cmsnfmfrq059nrohf3zdfo8jo","source":"X：Elvis Saravia (@omarsar0, DAIR.AI)","publishedAt":"2026-08-10T16:03:02.000Z","discoveredAt":"2026-08-10T16:14:36.879Z","summary":"Meta 新论文提出 Skaling law，通过单一交互指数耦合模型容量与训练数据，将平均绝对百分比误差在插值与外推场景下降低 1.5x 至 3x。该定律在数据稀缺和过度训练区间修正了标准 Chinchilla 与 Kaplan 形式的偏差，并可用约 10 倍更少算力从稀疏网格外推完整训练配置。","category":"paper","score":36,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnfmfrq059nrohf3zdfo8jo"}},{"id":"cmsn1riv504zyrogn47ks9fpb","title":"斯坦福与东北大学打造\"AI 智能体版 Git\"","title_en":"Stanford and Northeastern built \"Git for AI agents.\" Solves a common problem with existing agent fr…","url":"https://x.com/rohanpaul_ai/status/2086750257300529474","permalink":"https://aihot.virxact.com/items/cmsn1riv504zyrogn47ks9fpb","source":"X：Rohan Paul (@rohanpaul_ai)","publishedAt":"2026-08-10T09:43:25.000Z","discoveredAt":"2026-08-10T09:46:39.979Z","summary":"斯坦福和东北大学推出 Shepherd，一个\"面向 AI 智能体的 Git\"框架。它可将智能体的运行进程与文件系统一同提交，使执行过程可像 Git 一样回滚到任意历史提交并恢复。在双智能体协作测试中，Shepherd 通过监督回滚将性能差距缩小 91%（从 28.8% 提升至接近 57.2% 的单智能体水平）。","category":"paper","score":45,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsn1riv504zyrogn47ks9fpb"}},{"id":"cmsn0igsm03lsrogn23t8fpo1","title":"合肥硅臻芯片联合中科大实现片上 MBQC 技术突破，为百万比特光量子计算开辟可行路径","title_en":"通用百万比特光量子计算迎来可行路径，合肥硅臻芯片等发布片上 MBQC 技术突破","url":"https://www.ithome.com/0/987/915.htm","permalink":"https://aihot.virxact.com/items/cmsn0igsm03lsrogn23t8fpo1","source":"IT之家（RSS）","publishedAt":"2026-08-10T08:25:02.000Z","discoveredAt":"2026-08-10T09:11:35.571Z","summary":"合肥硅臻芯片联合中国科学技术大学任希锋教授研究组，在可编程硅光集成芯片上首次实现4光子16量子比特GHZ态与单光子4量子比特簇态的稳定生成，其中10个量子比特经纠缠目击方法验证为真实纠缠。团队还基于该簇态演示了Grover搜索算法，平均识别概率达0.987，为大规模通用光量子计算开辟了基于测量的量子计算（MBQC）可行路径。","category":"paper","score":33,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsn0igsm03lsrogn23t8fpo1"}},{"id":"cmsmwbz000ky3rohffjswbsmg","title":"Meta 新论文：判别式语言模型作为检索器，无需生成 item ID","title_en":"Meta's new retrieval paper is a reminder that better language models do not necessarily require more…","url":"https://x.com/rohanpaul_ai/status/2086709122406490501","permalink":"https://aihot.virxact.com/items/cmsmwbz000ky3rohffjswbsmg","source":"X：Rohan Paul (@rohanpaul_ai)","publishedAt":"2026-08-10T06:59:58.000Z","discoveredAt":"2026-08-10T07:14:36.595Z","summary":"Meta 新论文提出用判别式语言模型做检索，将 0.6B Qwen3 模型嵌入经典双塔架构，替代自回归生成 item ID，使向量检索保持高速。更强的交叉编码器作为教师进行知识蒸馏，移除蒸馏后 Recall@10 在 Beauty、Sports、Toys 上分别下降 13.3%、23.1%、8.0%。","category":"paper","score":54,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmwbz000ky3rohffjswbsmg"}},{"id":"cmsoq73tk02tnrohd55oi7kk8","title":"ReMEMBER：面向流式对话摘要的缺失证据记忆框架","title_en":"Don't Scroll Back： Missing-Evidence Memory for Streaming Dialogue Summarization","url":"https://arxiv.org/abs/2608.09043","permalink":"https://aihot.virxact.com/items/cmsoq73tk02tnrohd55oi7kk8","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-10T00:00:00.000Z","discoveredAt":"2026-08-11T13:58:24.109Z","summary":"针对流式对话摘要中当前窗口常缺乏足够上下文的问题，研究提出ReMEMBER框架，通过基于未解析窗口依赖的检索和证据密集化记忆，在固定预算下从无限历史中恢复缺失证据。在长达160K token历史的对话实验中，ReMEMBER在相同预算下提升了记忆召回率和缺口解析完整性。","category":"paper","score":45,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsoq73tk02tnrohd55oi7kk8"}},{"id":"cmsojrjg60ux8rofwr62b0q5u","title":"BDH-CQ：用循环潜在推理实现上下文学习","title_en":"BDH-CQ： In-Context Learning with Recurrent Latent Reasoning","url":"https://arxiv.org/abs/2608.09888","permalink":"https://aihot.virxact.com/items/cmsojrjg60ux8rofwr62b0q5u","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-10T00:00:00.000Z","discoveredAt":"2026-08-11T10:58:19.853Z","summary":"BDH-CQ 是一种结合上下文学习与循环潜在推理的推理模型，推理时输入持续更新循环记忆，模型在高维潜在空间中迭代计算求解，不输出中间推理过程。150M 参数配置在 ARC-AGI-1 上达到 29.5% pass@2，单任务推理成本仅 $0.0007，突破了该基准此前报告的成本-准确率帕累托前沿，创下成本效率新纪录。","category":"paper","score":46,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsojrjg60ux8rofwr62b0q5u"}},{"id":"cmsodc28j0n07rofw30v8pe9v","title":"窃取专有 LLM API 的推理轨迹：加密块可跨会话互换引发解密越狱","title_en":"Stealing Reasoning Traces from Proprietary LLM APIs","url":"https://arxiv.org/abs/2608.09867","permalink":"https://aihot.virxact.com/items/cmsodc28j0n07rofw30v8pe9v","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-10T00:00:00.000Z","discoveredAt":"2026-08-11T07:58:18.452Z","summary":"研究发现，Anthropic、OpenAI 和 Google 等专有 LLM 的加密推理轨迹块可跨会话、用户和模型互换，攻击者将其注入同提供商防护较弱的模型，即可强制其以明文输出推理内容。","category":"paper","score":81,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsodc28j0n07rofw30v8pe9v"}},{"id":"cmsodc28j0n06rofwatqma788","title":"CEAA：面向交互式计算系统的具身智能体认知架构","title_en":"CEAA： A Cognitive Embodied Agents Architecture for Interactive Computing Systems","url":"https://arxiv.org/abs/2608.09848","permalink":"https://aihot.virxact.com/items/cmsodc28j0n06rofwatqma788","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-10T00:00:00.000Z","discoveredAt":"2026-08-11T07:58:18.452Z","summary":"论文提出CEAA，一种用于在实时交互式虚拟环境中部署具身智能虚拟人（IVA）的模块化认知架构。该架构基于Sense-Think-Act范式和信念-欲望-意图（BDI）认知模型等既有框架，提供可复用的实现导向模板，弥合高层智能体推理模型与实时具身执行之间的鸿沟，以支持复杂交互式虚拟环境中可扩展、自适应且可解释的智能体。","category":"paper","score":41,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsodc28j0n06rofwatqma788"}},{"id":"cmso6wigb0d1profw2f6djg98","title":"RynnValue：用时间距离扩展机器人价值基础模型","title_en":"RynnValue： Scaling Robotic Value Foundation Models with Temporal Distance","url":"https://arxiv.org/abs/2608.09853","permalink":"https://aihot.virxact.com/items/cmso6wigb0d1profw2f6djg98","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-10T00:00:00.000Z","discoveredAt":"2026-08-11T04:58:17.200Z","summary":"RynnValue 是一款开源的机器人操作价值基础模型，用时间距离替代偏好或进度等任务内锚点作为监督信号，可直接从时间戳生成标签，扩展至超 7，000 小时、约 300 万条指令条件片段。","category":"paper","score":74,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso6wigb0d1profw2f6djg98"}},{"id":"cmso4rcdl0aolrofwlgumpsec","title":"UserIDA：可控用户模拟如何超越响应模仿、实现意图级控制","title_en":"Intent Speaks Louder： Controllable User Simulation Beyond Response Imitation","url":"https://arxiv.org/abs/2608.09420","permalink":"https://aihot.virxact.com/items/cmso4rcdl0aolrofwlgumpsec","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-10T00:00:00.000Z","discoveredAt":"2026-08-11T03:58:16.875Z","summary":"UserIDA（User Intent-Directive Alignment）将交互意图作为显式的逐轮指令，通过监督微调与意图校准的策略优化实现可控用户模拟。在LMSYS-USP上，UserIDA意图准确率达86.6%，超出最强专用用户模拟基线24.3个百分点；在上下文干预中，91.7%的对话状态能实现六种目标意图中的至少四种，而最强外部基线仅为22.9%。","category":"paper","score":54,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso4rcdl0aolrofwlgumpsec"}},{"id":"cmso4rcdl0aokrofws8xmwo85","title":"Evo-Bench：语言模型能否自主改进智能体运行框架？","title_en":"Evo-Bench： Can Language Models Improve Agent Harness？","url":"https://arxiv.org/abs/2608.09096","permalink":"https://aihot.virxact.com/items/cmso4rcdl0aokrofws8xmwo85","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-10T00:00:00.000Z","discoveredAt":"2026-08-11T03:58:16.875Z","summary":"Evo-Bench 发布，成为首个专门评估模型自主进化其运行框架（harness）能力的基准，覆盖搜索、办公和通用智能体三大领域。对九个前沿及开源权重模型的评估显示，顶尖模型在该基准上取得高达 16.6 分的绝对提升，接近人类工程化基线水平。自主进化在通用与搜索任务上表现优于人工框架，但在需要特定处理流程的办公任务上表现不佳。","category":"paper","score":48,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso4rcdl0aokrofws8xmwo85"}},{"id":"cmso2m6ot082rrofwnmm5aq7y","title":"Sci-VBench：科学领域知识密集型视频生成的评测基准","title_en":"Sci-VBench： Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains","url":"https://arxiv.org/abs/2608.09873","permalink":"https://aihot.virxact.com/items/cmso2m6ot082rrofwnmm5aq7y","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-10T00:00:00.000Z","discoveredAt":"2026-08-11T02:58:16.068Z","summary":"Sci-VBench 基准包含 1，253 个专家标注示例，覆盖自然科学、医疗健康、人文社科与工程四大领域的 60 个主题，要求模型生成具备科学推理能力的时序视频。评测显示，16 个前沿模型在感知质量上接近，但在提示词接地性与科学因果正确性上差异显著，专有与开源模型间存在明显差距。","category":"paper","score":45,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso2m6ot082rrofwnmm5aq7y"}},{"id":"cmso2m6ot082qrofwehuksa8m","title":"Motif 3 技术报告：314B 参数 MoE 模型发布","title_en":"Motif 3： Technical Report","url":"https://arxiv.org/abs/2608.09119","permalink":"https://aihot.virxact.com/items/cmso2m6ot082qrofwehuksa8m","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-10T00:00:00.000Z","discoveredAt":"2026-08-11T02:58:16.068Z","summary":"Motif 3 是一个仅解码器的混合专家（MoE）语言模型，总参数 314B，每个 token 激活 13.2B 参数，每个稀疏 MoE 层含 384 个路由专家、每 token 选取 8 个。","category":"paper","score":44,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso2m6ot082qrofwehuksa8m"}},{"id":"cmso2m6os082profwwgmk93f8","title":"SWE-Bench ProMax：面向大规模多语言代码重构的智能体评测基准","title_en":"SWE-Bench ProMax： Benchmarking Agents on Large-Scale Multilingual Code Refactoring","url":"https://arxiv.org/abs/2608.09802","permalink":"https://aihot.virxact.com/items/cmso2m6os082profwwgmk93f8","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-10T00:00:00.000Z","discoveredAt":"2026-08-11T02:58:16.068Z","summary":"针对现有编码基准快速饱和及评测质量问题，研究团队推出 SWE-Bench ProMax，一个包含 170 个实例、覆盖七种编程语言（Python、Java、TypeScript、Go、C、C++、Rust）的多语言代码重构基准。","category":"paper","score":67,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso2m6os082profwwgmk93f8"}},{"id":"cmsmg6tvj01piroln3wpn2css","title":"ベースモデルに依存しないオーケストレーションに向けて：Gemma 4版 Sakana Fuguの検証","title_en":null,"url":"https://sakana.ai/fugu-gemma4","permalink":"https://aihot.virxact.com/items/cmsmg6tvj01piroln3wpn2css","source":"Sakana AI：Blog（网页）","publishedAt":"2026-08-09T23:00:00.000Z","discoveredAt":"2026-08-09T23:42:42.467Z","summary":null,"category":"paper","score":0,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmg6tvj01piroln3wpn2css"}},{"id":"cmsmaw72e03m8ro9ndsiq57wc","title":"AI 削减入门岗或引发\"认知公地悲剧\"","title_en":"This paper argues that cutting entry-level work with AI can create a long-term expertise problem tha…","url":"https://x.com/rohanpaul_ai/status/2086559090109804769","permalink":"https://aihot.virxact.com/items/cmsmaw72e03m8ro9ndsiq57wc","source":"X：Rohan Paul (@rohanpaul_ai)","publishedAt":"2026-08-09T21:03:48.000Z","discoveredAt":"2026-08-09T21:14:28.639Z","summary":"一篇论文提出，用 AI 削减入门级工作会造成长期专业人才断层，且没有任何单一公司有动力解决，并将其定义为\"认知公地\"问题。论文指出，AI 可直接取消初级岗位，或让初级员工不经认知挣扎就产出成果，从而削弱专业知识的再生管道。数据显示，2022 年 10 月至 2025 年 9 月，高度 AI 暴露职业中 22-25 岁工人就业相对下降 16%，而 35-49 岁就业增长超 8%。","category":"paper","score":29,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmaw72e03m8ro9ndsiq57wc"}},{"id":"cmsm8r0lv04lfro980foz99j5","title":"Metis：把记忆内化进 LLM 自身状态","title_en":"This is such a wild idea. What if memory were a capability of the LLM itself， rather than a retriev…","url":"https://x.com/rohanpaul_ai/status/2086544779974955262","permalink":"https://aihot.virxact.com/items/cmsm8r0lv04lfro980foz99j5","source":"X：Rohan Paul (@rohanpaul_ai)","publishedAt":"2026-08-09T20:06:56.000Z","discoveredAt":"2026-08-09T20:14:27.701Z","summary":"Metis 提出将记忆作为 LLM 自身能力，通过骨干网络内的持久记忆状态，在前向传播中压缩过往交互，而非外挂检索系统。无上下文设置下，Metis-27B 在 LoCoMo （Gold） 得分 26.74，远超 vanilla Qwen3.5-27B 的 0.07 和 Temp-LoRA-27B 的 4.24，但仍低于全上下文 Qwen3.5-27B 的 65.03。","category":"paper","score":36,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsm8r0lv04lfro980foz99j5"}},{"id":"cmsm4gcaf08mnroy9j2sel8n3","title":"Meta 新研究：EvoHarness-RL 让智能体自主学习工具编排策略","title_en":"New research from Meta. Agent harnesses are still mostly authored by hand. This makes it hard to t…","url":"https://x.com/omarsar0/status/2086509069762981896","permalink":"https://aihot.virxact.com/items/cmsm4gcaf08mnroy9j2sel8n3","source":"X：Elvis Saravia (@omarsar0, DAIR.AI)","publishedAt":"2026-08-09T17:45:02.000Z","discoveredAt":"2026-08-09T18:14:10.252Z","summary":"Meta 新研究提出 EvoHarness-RL，让智能体离线学习工具编排（harness）策略，并在运行时在线更新外部状态，替代手工编写。Qwen3-8B 在 ALFWorld 上达到 96.9% 准确率。训练中出现的\"工具编排退火\"与\"工具编排进化\"两种动态表明，长程任务智能体从可训练的协调策略中获益，胜过更大的工具集或记忆。","category":"paper","score":37,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsm4gcaf08mnroy9j2sel8n3"}},{"id":"cmslqj0vu03h4roagcrjvgqss","title":"Meta 研究：量化让推理模型过度自我怀疑","title_en":"Paper from Meta shows Quantized reasoning models often lose because they keep doubting a correct ans…","url":"https://x.com/rohanpaul_ai/status/2086415812550910146","permalink":"https://aihot.virxact.com/items/cmslqj0vu03h4roagcrjvgqss","source":"X：Rohan Paul (@rohanpaul_ai)","publishedAt":"2026-08-09T11:34:28.000Z","discoveredAt":"2026-08-09T11:44:20.513Z","summary":"Meta 新论文发现，强量化会让推理模型在已得出正确答案后反复自我怀疑，过度思考失败率最高升至 52%。研究在数学、编程和科学任务上测试了 5 个推理模型（1.5B 至 32B），对 50 个犹豫词施加小幅惩罚即可将推理长度缩短 12% 至 23%，且常能保持或提升准确率。","category":"paper","score":46,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmslqj0vu03h4roagcrjvgqss"}},{"id":"cmslk3c3k05byroo0twq8u9vw","title":"哈佛新研究：生成模型或存第三扩展轴","title_en":"New Harvard paper shows generative models may be missing a third scaling axis： how much they explore…","url":"https://x.com/rohanpaul_ai/status/2086370187784470962","permalink":"https://aihot.virxact.com/items/cmslk3c3k05byroo0twq8u9vw","source":"X：Rohan Paul (@rohanpaul_ai)","publishedAt":"2026-08-09T08:33:10.000Z","discoveredAt":"2026-08-09T08:44:12.028Z","summary":"哈佛新论文提出生成模型可能缺失第三扩展轴：训练中的探索量。在强表征自编码器（RAE）图像生成方案中加入探索后，模型以6.2×更少的训练样本和4.1×更少的FLOPs达到基线最终性能。该方法将best-of-K训练视为扩展模型可学习模态数量的途径，而无需增加推理步骤。","category":"paper","score":35,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmslk3c3k05byroo0twq8u9vw"}},{"id":"cmslhy3d703fbroo0pusw5skf","title":"微软论文：编码智能体不应按聊天请求调度","title_en":"New Microsoft Paper on GitHub Copilot's production traces show why coding agents should not be serve…","url":"https://x.com/rohanpaul_ai/status/2086356354713985266","permalink":"https://aihot.virxact.com/items/cmslhy3d703fbroo0pusw5skf","source":"X：Rohan Paul (@rohanpaul_ai)","publishedAt":"2026-08-09T07:38:12.000Z","discoveredAt":"2026-08-09T07:44:08.231Z","summary":"微软对 13.5M 次 GitHub Copilot 会话的生产轨迹分析显示，87% 的 LLM 调用来自智能体自身而非用户。KV 缓存命中率在单轮内从首次调用的约 45% 升至第三次起的 92-94%，模型切换时骤降至 8%。论文提出基于轮次和会话级特征的轻量预测器，可捕获 86-90% 的总空闲时间，表明编码智能体基础设施应按轮次调度工作流状态，而非请求级策略。","category":"paper","score":50,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmslhy3d703fbroo0pusw5skf"}},{"id":"cmsoq73tk02tmrohda0uqkpvj","title":"SymDiag：用神经符号验证为 LLM 推理提供可解释诊断","title_en":"SymDiag： Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification","url":"https://arxiv.org/abs/2608.08786","permalink":"https://aihot.virxact.com/items/cmsoq73tk02tmrohda0uqkpvj","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-09T00:00:00.000Z","discoveredAt":"2026-08-09T00:00:00.000Z","summary":"SymDiag 提出一种神经符号框架，将大语言模型的自然语言思维链转化为符号约束，通过逐步可满足性与蕴含检查定位失败步骤，并生成反例、不一致证据和缺失前提等可验证诊断信息。其 Self-Auditor 组件利用双重符号编码一致性检查，区分推理错误与神经到符号的翻译噪声。在数学、逻辑、科学和通用推理基准上，SymDiag 比仅看结果和 LLM 评判更能识别不忠实推理，并为多轮推理修复提供更有效反馈。","category":"paper","score":46,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsoq73tk02tmrohda0uqkpvj"}},{"id":"cmskhg8hc07pzrobycb8cdzdg","title":"读者给AI生成的短篇小说打分高于人类作品，直到他们知道作者是机器","title_en":"Readers rate AI-generated short stories higher than human ones until they learn a machine wrote them","url":"https://the-decoder.com/readers-rate-ai-generated-short-stories-higher-than-human-ones-until-they-learn-a-machine-wrote-them","permalink":"https://aihot.virxact.com/items/cmskhg8hc07pzrobycb8cdzdg","source":"The Decoder：AI News（RSS）","publishedAt":"2026-08-08T14:18:55.000Z","discoveredAt":"2026-08-08T14:42:27.243Z","summary":"三项实验共2500多名参与者无法区分ChatGPT与人类创作的短篇小说，且对AI文本的质量和沉浸感评分更高（质量均分1.54对0.97，沉浸感1.42对1.00）。但被告知作者是AI后，两类故事的评分均下降；对AI持正面态度的参与者评分更高，持怀疑态度者则相反。研究者认为，AI文本更流畅易读且情绪更积极，这可能解释了高评分，但未必意味着文学上更优。","category":"paper","score":61,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmskhg8hc07pzrobycb8cdzdg"}},{"id":"cmsk9yb97046zrow93i345be7","title":"DeepMind 的 WeatherNext 飓风模型为预报员争取到额外一天预警时间","title_en":"DeepMind's hurricane model bought forecasters an extra day","url":"https://arstechnica.com/science/2026/08/deepminds-hurricane-model-bought-forecasters-an-extra-day","permalink":"https://aihot.virxact.com/items/cmsk9yb97046zrow93i345be7","source":"Ars Technica：AI（RSS）","publishedAt":"2026-08-08T11:05:50.000Z","discoveredAt":"2026-08-08T11:12:34.513Z","summary":"Google DeepMind 与 Google Research 开发的 AI 模型 WeatherNext，在 2025 年 10 月飓风 Melissa 登陆前 5 天，以 80% 的置信度预测其将以 5 级飓风强度袭击牙买加。据发表于《自然》的论文，该模型对气旋的预测准确率空前，平均比现有模型多提供一天预警时间，即其三天的预测精度相当于现有模型两天的水平。","category":"paper","score":77,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsk9yb97046zrow93i345be7"}},{"id":"cmsk8vgn30456roksadv96glc","title":"AI 智能体 KISS Sorcar 8 小时优化 SQLite 提速 59%","title_en":"Making SQLite 5% faster would already be powerful. Here， an AI agent， KISS Sorcar， got 59% improvem…","url":"https://x.com/rohanpaul_ai/status/2086038118608801981","permalink":"https://aihot.virxact.com/items/cmsk8vgn30456roksadv96glc","source":"X：Rohan Paul (@rohanpaul_ai)","publishedAt":"2026-08-08T10:33:38.000Z","discoveredAt":"2026-08-08T10:42:22.682Z","summary":"AI 智能体 KISS Sorcar 在不到 8 小时、API 成本低于 $150 的情况下，将 SQLite 优化提速 59%（几何平均 1.59x），且全部 1，032，940 项测试通过。","category":"paper","score":47,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsk8vgn30456roksadv96glc"}},{"id":"cmsofh7p60plirofwefa5l92s","title":"Ego-OSCAR：开源低成本头戴式立体惯性采集系统","title_en":"Ego-OSCAR： Egocentric Open source Stereo CAptuRe System","url":"https://arxiv.org/abs/2608.08285","permalink":"https://aihot.virxact.com/items/cmsofh7p60plirofwefa5l92s","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-08T00:00:00.000Z","discoveredAt":"2026-08-08T00:00:00.000Z","summary":"Ego-OSCAR 是一款开源硬件、低成本的头戴式立体惯性采集设备，单台物料成本低于 200 美元，仅使用商用组件和 3D 打印部件。设备搭配硬件同步全局快门立体相机、6 轴 IMU 和嵌入式 Linux SBC，并开源完整软件栈及约 550 小时带同步 IMU 的日常室内环境第一视角立体视频。数据集提供开放词汇动作标注和逐帧 3D 手部重建，旨在降低大规模众包第一视角数据采集门槛。","category":"paper","score":78,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsofh7p60plirofwefa5l92s"}},{"id":"cmsodc28j0n05rofwbtflsloi","title":"视觉-语言定位作为双向概念对应：ConCor-1 模型提出统一框架","title_en":"Vision-Language Grounding as Bidirectional Concept Correspondence","url":"https://arxiv.org/abs/2608.07886","permalink":"https://aihot.virxact.com/items/cmsodc28j0n05rofwbtflsloi","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-08T00:00:00.000Z","discoveredAt":"2026-08-08T00:00:00.000Z","summary":"研究将视觉-语言定位重新定义为图像-文本对的双向概念对应，无需预先指定文本短语，即可恢复所有视觉指称文本片段与实例级图像分割的对应关系。该框架统一了短语定位、指称表达定位和开放词汇检测等任务。作者提出 ConCor-1 模型，使用可学习桥接 token 预测文本掩码、图像掩码和对应存在分数，在长标题数据集上对应 F1 提升 48%，在零样本 LVIS 上提升 29%。","category":"paper","score":55,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsodc28j0n05rofwbtflsloi"}},{"id":"cmso6wigb0d1rrofw4l6ifmm7","title":"Ouroboros：具备评审式核心进化的自开发前沿编程智能体","title_en":"Ouroboros： A Self-Developing Frontier Coding Agent with Reviewed Core Evolution","url":"https://arxiv.org/abs/2608.08311","permalink":"https://aihot.virxact.com/items/cmso6wigb0d1rrofw4l6ifmm7","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-08T00:00:00.000Z","discoveredAt":"2026-08-08T00:00:00.000Z","summary":"Ouroboros 是一个自开发智能体框架，其工具、提示词、上下文组装和核心实现通过评审式提交持续改进，并成为后续工作的运行时。","category":"paper","score":74,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso6wigb0d1rrofw4l6ifmm7"}},{"id":"cmso6wiga0d1nrofwl24btv0c","title":"OasisKV：用前瞻稀疏预取将解码期 KV 缓存扩展到 HBM 之外","title_en":"OasisKV： Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching","url":"https://arxiv.org/abs/2608.08097","permalink":"https://aihot.virxact.com/items/cmso6wiga0d1nrofwl24btv0c","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-08T00:00:00.000Z","discoveredAt":"2026-08-08T00:00:00.000Z","summary":"OasisKV 提出一种以内存为中心的 LLM 推理系统设计，将完整 KV 缓存存储与 HBM 解耦，利用解码期注意力稀疏性，仅保留最相关 token 的 KV 条目。","category":"paper","score":53,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso6wiga0d1nrofwl24btv0c"}},{"id":"cmso0gzx205gyrofwis3y2dyp","title":"Evidence-RL：面向证据密集型视觉推理的强化学习方法","title_en":"Evidence-RL： Towards Evidence-intensive Visual Reasoning","url":"https://arxiv.org/abs/2608.08021","permalink":"https://aihot.virxact.com/items/cmso0gzx205gyrofwis3y2dyp","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-08T00:00:00.000Z","discoveredAt":"2026-08-08T00:00:00.000Z","summary":"Evidence-RL 提出反事实证据解耦（CED）训练时审计方法，通过中和物体级证据区域并对比支持度下降，在 GRPO 中奖励依赖证据路径而非捷径的正确答案。CED 无需问题级证据标注，不增加推理开销，在九个公开基准和四个骨干模型上优于此前基于 RL 的后训练方法。","category":"paper","score":54,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso0gzx205gyrofwis3y2dyp"}},{"id":"cmsjmbkmc0dqoroo5xaqi0vgg","title":"新 AI 模型 IceBoost v2.0 估算全球冰川蕴冰约 15 万立方千米，融化可抬升海平面 32.3 厘米","title_en":"新 AI 预估全球冰川蕴藏约 15 万立方千米的冰，融化可抬升海平面 32.3 厘米","url":"https://www.ithome.com/0/987/225.htm","permalink":"https://aihot.virxact.com/items/cmsjmbkmc0dqoroo5xaqi0vgg","source":"IT之家（RSS）","publishedAt":"2026-08-07T23:22:41.000Z","discoveredAt":"2026-08-08T00:11:01.752Z","summary":"科学家构建 AI 模型 IceBoost v2.0，估算全球冰川约含 15 万立方千米冰，若全部融化（不含南极和格陵兰）将抬升全球平均海平面 32.3 厘米。该模型由威尼斯大学牵头开发，用超 700 万条冰厚测量数据及 26 个物理与几何变量训练，对冰川空间分布的描绘与实地观测匹配度最高提升 40%。研究将支撑 IPCC 冰川评估及淡水供应预测。","category":"paper","score":38,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsjmbkmc0dqoroo5xaqi0vgg"}},{"id":"cmsj0x5jh20xaronk350vfkju","title":"线性注意力模型实现零训练推理偏差","title_en":"nice rl experiment on train-inference mismatch","url":"https://x.com/natolambert/status/2085726242314346760","permalink":"https://aihot.virxact.com/items/cmsj0x5jh20xaronk350vfkju","source":"X：Nathan Lambert (@natolambert)","publishedAt":"2026-08-07T13:54:21.000Z","discoveredAt":"2026-08-07T14:11:58.553Z","summary":"研究团队在TorchTitan RL + vLLM上为Gated DeltaNet（Qwen3.5-9B/35B-A3B）实现训练/生成逐位一致，logprob差为0，为开源首个线性注意力模型达成此目标。异步RL下，offpolicy=32时标准栈偏差超0.065，新方法保持约0.035。代价是训练吞吐量降低2-3倍，团队建议将其作为调试工具而非生产默认。","category":"paper","score":27,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsj0x5jh20xaronk350vfkju"}},{"id":"cmsiys4dz1yqironkuitnfi7t","title":"斯坦福与 Arc Institute 用 AI 设计全新病毒基因组，16 种在实验室成功杀死细菌","title_en":"Stanford and Arc Institute scientists used AI to design new viruses that killed bacteria in the lab","url":"https://the-decoder.com/stanford-and-arc-institute-scientists-used-ai-to-design-new-viruses-that-killed-bacteria-in-the-lab","permalink":"https://aihot.virxact.com/items/cmsiys4dz1yqironkuitnfi7t","source":"The Decoder：AI News（RSS）","publishedAt":"2026-08-07T12:50:56.000Z","discoveredAt":"2026-08-07T13:12:02.670Z","summary":"斯坦福大学与 Arc Institute 团队用 AI 模型 Evo 从零设计完整病毒基因组，并在实验室构建出 16 种自然界不存在的功能性病毒。Evo 提出 70 万个候选基因组，团队仅筛选最有希望的 285 个序列合成并植入细菌，其中 16 个成功复制并杀死宿主。该研究已通过同行评审发表于《Science》，但 Evo 未接受人类病原体数据训练，且能否推广至其他病毒类群仍是未知数。","category":"paper","score":72,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsiys4dz1yqironkuitnfi7t"}},{"id":"cmsiry2eo1r4tronkt9qv2twu","title":"小红书联合浙大、复旦提出 CULTURE-MT：首个面向社媒翻译的「文化有效性」评测基准，入选 ICML 2026","title_en":"AI 翻译看得懂字，却读不懂「梗」？小红书联合浙大、复旦提出首个「文化有效性」评测标准，入选 ICML 2026","url":"https://mp.weixin.qq.com/s?__biz=Mzg4OTc2MzczNg%3D%3D&mid=2247496008&idx=1&sn=08f2ce717483f63bc00a2181e59e3f40","permalink":"https://aihot.virxact.com/items/cmsiry2eo1r4tronkt9qv2twu","source":"公众号：小红书技术（dots.llm）","publishedAt":"2026-08-07T09:59:00.000Z","discoveredAt":"2026-08-07T10:00:44.054Z","summary":"小红书联合浙江大学、复旦大学提出 CULTURE-MT，这是首个面向中英社媒笔记翻译、兼顾文化符号传递与情感共鸣的评测基准，并首次提出「文化有效性」评估标准与自动评估模型 JUDGER（准确率 86.03%）。","category":"paper","score":65,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsiry2eo1r4tronkt9qv2twu"}},{"id":"cmsifhdv01csjronk2j858egf","title":"Qwen-CUA：用屏幕鼠标键盘训练通用电脑智能体","title_en":"Qwen-CUA shows that a strong computer agent can be trained using only the same screen， mouse， and ke…","url":"https://x.com/rohanpaul_ai/status/2085578932368109935","permalink":"https://aihot.virxact.com/items/cmsifhdv01csjronk2j858egf","source":"X：Rohan Paul (@rohanpaul_ai)","publishedAt":"2026-08-07T04:09:00.000Z","discoveredAt":"2026-08-07T04:11:50.284Z","summary":"Qwen-CUA 证明仅凭人类使用的屏幕、鼠标和键盘即可训练出强大的电脑智能体，将截图作为 AI 的通用界面。训练动用近 10 万虚拟处理器核心和约 4 万个可检查任务，智能体仅通过截图和鼠标键盘操作，无需网站代码或辅助标签。其在 OSWorld-Verified 基准上得分 86.2，但长时程任务仍存在较大差距。","category":"paper","score":35,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsifhdv01csjronk2j858egf"}},{"id":"cmsifg4lm1cltronkdr7hy9os","title":"本源量子与中科大团队破解超导量子计算\"速度与保真度\"矛盾","title_en":"我国科研团队破解超导量子计算\"速度与保真度\"矛盾","url":"https://www.ithome.com/0/986/894.htm","permalink":"https://aihot.virxact.com/items/cmsifg4lm1cltronkdr7hy9os","source":"IT之家（RSS）","publishedAt":"2026-08-07T03:29:39.000Z","discoveredAt":"2026-08-07T04:10:50.506Z","summary":"本源量子与中国科学技术大学联合团队提出\"参数空间扩展受控相位门（PSE-CZ）\"方案，破解超导量子计算速度与保真度相互制约难题，成果发表于《物理评论快报》。在\"本源悟空\"上对20对两比特量子门测试显示，30-40纳秒极短门长下仍保持优于传统CZ门的稳定性能。该方案还可推广至离子阱、固态自旋等平台。","category":"paper","score":46,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsifg4lm1cltronkdr7hy9os"}},{"id":"cmsidc76b1a3dronk32xn479d","title":"英伟达新论文：LLM 间可复用提示词缓存","title_en":"New Nvidia paper shows， one LLM can reuse another model's prompt memory instead of processing the wh…","url":"https://x.com/rohanpaul_ai/status/2085558547987697969","permalink":"https://aihot.virxact.com/items/cmsidc76b1a3dronk32xn479d","source":"X：Rohan Paul (@rohanpaul_ai)","publishedAt":"2026-08-07T02:48:00.000Z","discoveredAt":"2026-08-07T03:11:49.459Z","summary":"英伟达新论文提出，一个 LLM 可通过简单的线性转换器复用另一模型的提示词缓存，无需重新处理整个提示词。该方法从 500 条校准序列学习，无需反向传播，在 Qwen3、Llama 3.1 和 Ministral 3 的 6 组模型对上，4 组保留了目标模型 73% 至 98% 的基准准确率，转换速度提升 2.7 至 25 倍。","category":"paper","score":36,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsidc76b1a3dronk32xn479d"}},{"id":"cmsia4dpd15wnronk9sp7o7b1","title":"谷歌新论文：相似AI智能体可理性合作","title_en":"New Google Paper says classical game theory predicts betrayal， but similar AI agents can rationally …","url":"https://x.com/rohanpaul_ai/status/2085541686646866105","permalink":"https://aihot.virxact.com/items/cmsia4dpd15wnronk9sp7o7b1","source":"X：Rohan Paul (@rohanpaul_ai)","publishedAt":"2026-08-07T01:41:00.000Z","discoveredAt":"2026-08-07T01:41:46.030Z","summary":"谷歌新论文提出，相似AI智能体即使无法沟通或未来无回报，也能基于\"自身选择可预测对方选择\"的推理实现理性合作，挑战经典博弈论对一次性囚徒困境中背叛的预测。Gemini和Gemma智能体在多种游戏后，相同模型合作率高，相关模型次之，随机对手多遭背叛。该\"嵌入均衡\"或能更好预测AI社会，但警示相似AI可能偏向同类而非人类。","category":"paper","score":68,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsia4dpd15wnronk9sp7o7b1"}},{"id":"cmsib5tcp1787ronkidkh7v0t","title":"斯坦福团队用 AI 首次设计完整病毒基因组，16 种新型噬菌体可杀死大肠杆菌","title_en":"AI 设计病毒问世：首次设计完整基因组，16 种新型噬菌体可杀死大肠杆菌","url":"https://www.ithome.com/0/986/809.htm","permalink":"https://aihot.virxact.com/items/cmsib5tcp1787ronkidkh7v0t","source":"IT之家（RSS）","publishedAt":"2026-08-07T01:18:50.000Z","discoveredAt":"2026-08-07T02:10:49.254Z","summary":"斯坦福大学研究人员使用 AI 技术首次成功设计出完整基因组，产出的 16 种新型噬菌体在实验室中可有效杀死大肠杆菌，不会对人类构成威胁。团队开发的 Evo1 和 Evo2 模型以病毒、细菌、植物和人类的遗传代码为训练数据，通过微调预测\"生命语言\"。研究被视为\"非常重要的转折点\"，或为抗生素耐药感染提供新治疗路径。","category":"paper","score":55,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsib5tcp1787ronkidkh7v0t"}},{"id":"cmsob6v4n0kflrofwo3ss0mgp","title":"A^2E：面向智能体框架的端到端审计引擎","title_en":"An End-to-End Agent Auditing Engine","url":"https://arxiv.org/abs/2608.07346","permalink":"https://aihot.virxact.com/items/cmsob6v4n0kflrofwo3ss0mgp","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T00:00:00.000Z","summary":"A^2E（Agent Auditing Engine）是一个端到端智能体框架评测引擎，基于新提出的 Agent Task Protocol（ATP）快速集成不同框架的评测任务，并通过自动插桩的 Monitor 捕获标准化执行轨迹。","category":"paper","score":45,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsob6v4n0kflrofwo3ss0mgp"}},{"id":"cmso6wiga0d1orofw5ve9m2yc","title":"Agent Memory Distillation：用层级教师记忆增强小型 LLM 智能体","title_en":"Agent Memory Distillation： Empowering Small LLM Agents with Hierarchical Teacher Memory","url":"https://arxiv.org/abs/2608.07169","permalink":"https://aihot.virxact.com/items/cmso6wiga0d1orofw5ve9m2yc","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T00:00:00.000Z","summary":"Agent Memory Distillation（AMD）提出一种免训练框架，通过层级记忆将大型教师智能体的结构化知识迁移至小型学生智能体。","category":"paper","score":54,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso6wiga0d1orofw5ve9m2yc"}},{"id":"cmsnjbn7008u1rohfzqpaptem","title":"多智能体取证推理实现泛化深度伪造视频检测","title_en":"Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection","url":"https://arxiv.org/abs/2608.06865","permalink":"https://aihot.virxact.com/items/cmsnjbn7008u1rohfzqpaptem","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T00:00:00.000Z","summary":"研究团队提出 FaceVid-Forensics-100K 深度伪造视频数据集，含 10 万条视频、覆盖 33 种合成方法，并提供细粒度文本标注与取证解释。基于该基准构建的多智能体取证推理框架，由纹理、光照、运动、物理四个专家智能体独立分析伪造线索，再由裁判智能体汇总输出预测与解释。","category":"paper","score":51,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnjbn7008u1rohfzqpaptem"}},{"id":"cmsn0134o042groq26g0m252n","title":"Capek 0.5：面向具身智能的执行中心视觉语言模型","title_en":"Capek 0.5： An Execution-Centric Vision-Language Model for Embodied Intelligence","url":"https://arxiv.org/abs/2608.06756","permalink":"https://aihot.virxact.com/items/cmsn0134o042groq26g0m252n","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T00:00:00.000Z","summary":"Capek 0.5 是一款以执行为中心的具身视觉语言模型，按功能角色将能力分为空间推理、时间理解、动作引导与状态验证四类，先由专家模型通过可验证奖励的强化学习分别习得，再经权重合并与策略蒸馏整合为单一推理模型。该模型提供 2B 与 35B-A3B 两种规模，在多数基准项上优于初始化版本，并新增状态验证基准 Capek-StateBench，同时保留全部四项专家能力并支持闭环具身任务执行。","category":"paper","score":52,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsn0134o042groq26g0m252n"}},{"id":"cmsmxvxoy0mw3rohfdugh4ahx","title":"零差距并非恢复：基准污染的分层逐题概率评估与逐步缓解","title_en":"Zero Gap Is Not Restoration： Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination","url":"https://arxiv.org/abs/2608.07341","permalink":"https://aihot.virxact.com/items/cmsmxvxoy0mw3rohfdugh4ahx","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T00:00:00.000Z","summary":"论文提出SA-PPG分层逐题概率差距评估法，通过采样估计每题求解概率并与干净模型逐题对比，揭示现有G-AP指标存在缺陷，导致此前污染缓解策略的恢复效果被显著高估。研究还提出RailCap缓解策略，在生成中限制回落到贪心轨迹的下一token为次优选项，在多个受污染模型和基准上取得最低SA-PPG。","category":"paper","score":50,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmxvxoy0mw3rohfdugh4ahx"}},{"id":"cmsmpb84x0a9srohfixs4n943","title":"SimWAM：面向端到端自动驾驶的简易世界动作模型","title_en":"SimWAM： A Simple World Action Model for End-to-End Autonomous Driving","url":"https://arxiv.org/abs/2608.07468","permalink":"https://aihot.virxact.com/items/cmsmpb84x0a9srohfixs4n943","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T00:00:00.000Z","summary":"SimWAM将视频生成仅用作训练信号，通过联合流匹配共训预训练视频专家与轻量动作专家，并在推理时丢弃视频分支，形成直接预测轨迹的独立规划器。该模型在NAVSIM上取得91.5 PDMS，超越现有基于WAM的规划器且延迟显著更低，并可零样本迁移至nuScenes。代码与模型权重已开源。","category":"paper","score":46,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmpb84x0a9srohfixs4n943"}},{"id":"cmsmpb84w0a9qrohfh19tpixq","title":"YOLO-PEFT：面向 YOLO 系列的结构感知参数高效微调框架","title_en":"YOLO-PEFT： Parameter-Efficient Fine-Tuning on YOLO Family","url":"https://arxiv.org/abs/2608.07051","permalink":"https://aihot.virxact.com/items/cmsmpb84w0a9qrohfh19tpixq","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T00:00:00.000Z","summary":"YOLO-PEFT 提出将适配器放置形式化为可审计的约束规划问题，为实时检测器生成预算内的目标模块计划，或在训练前返回 Refuse。","category":"paper","score":52,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmpb84w0a9qrohfh19tpixq"}},{"id":"cmsmn62dp07ikrohfocffwjlx","title":"可寻址记忆：为视频世界模型扩展视觉持久生成","title_en":"Addressable Memory for Video World Models","url":"https://arxiv.org/abs/2608.07408","permalink":"https://aihot.virxact.com/items/cmsmn62dp07ikrohfocffwjlx","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T00:00:00.000Z","summary":"WorldTrace提出免训练记忆框架，通过为压缩摘要槽分配分布内虚拟位置，解决交互式视频世界模型在超出训练时长后无法通过注意力检索历史帧的问题。WorldTrace-Field在LoopBench上提升时序一致性+15.5%，WorldTrace-Landmark提升情景回忆+19.5%，无需重训练即可扩展视觉持久生成。","category":"paper","score":52,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmn62dp07ikrohfocffwjlx"}},{"id":"cmsmn62dp07ijrohfcin590xr","title":"Skaling 定律：用单一交互指数修正 Chinchilla 缩放定律的边界偏差","title_en":"Skaling： Chinchilla's Exponents Meet Kaplan's Coupling","url":"https://arxiv.org/abs/2608.07222","permalink":"https://aihot.virxact.com/items/cmsmn62dp07ijrohfcin590xr","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T00:00:00.000Z","summary":"针对标准缩放定律在数据稀缺与过度训练极端下系统性高估/低估损失的问题，研究者提出 Skaling 定律，通过单一交互指数耦合模型规模与数据量，将插值与外推的 MAPE 降低 1.5-3 倍。配合低算力\"L 形\"稀疏网格策略，该定律仅用约 1/10 的算力即可达到全网格 Chinchilla 的预测精度，为受限实验预算下的算力分配提供更稳健的框架。","category":"paper","score":66,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmn62dp07ijrohfcin590xr"}},{"id":"cmsmn62dp07iirohfi399abtb","title":"Modular TTT：将测试时训练重构为可组合模块","title_en":"Modular TTT： Rethinking Test-Time Training as Composable Modules","url":"https://arxiv.org/abs/2608.07110","permalink":"https://aihot.virxact.com/items/cmsmn62dp07iirohfi399abtb","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T00:00:00.000Z","summary":"Modular TTT 提出将测试时训练（TTT）的内部学习器表示为有向无环图，并把快速权重网络、损失函数、学习率、权重衰减和归一化作为显式设计维度，自动组合成完整的图级 TTT 计算。","category":"paper","score":45,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmn62dp07iirohfi399abtb"}},{"id":"cmsmn62do07igrohfbegm44rj","title":"ReASearch：让优化器本身成为智能体，用推理驱动跨提示词、程序与 ML 工作流的搜索","title_en":"The Optimizer Is the Agent： Reasoning-Driven Search across Prompts， Programs， and ML Workflows","url":"https://arxiv.org/abs/2608.06714","permalink":"https://aihot.virxact.com/items/cmsmn62do07igrohfbegm44rj","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T00:00:00.000Z","summary":"ReASearch 提出一个统一框架，让单一工具型智能体自主决定评估什么、如何诊断失败、做哪些修改以及何时验证或重启，而非依赖进化搜索等显式外层控制器。在 14 项任务上，它相比强领域专用基线取得 2% 至 40% 的提升，部分情况下还发现了优于此前人类最佳已知结果的方案。研究观察到，通常由显式控制器实现的复杂搜索行为，会自然地从智能体的推理过程中涌现。","category":"paper","score":54,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmn62do07igrohfbegm44rj"}},{"id":"cmsjjjvzz0bczroo5piheqn26","title":"Arbitrage：利用优势感知投机实现高效推理","title_en":"Arbitrage： Efficient Reasoning via Advantage-Aware Speculation","url":"https://machinelearning.apple.com/research/arbitrage-efficient-reasoning","permalink":"https://aihot.virxact.com/items/cmsjjjvzz0bczroo5piheqn26","source":"Apple Machine Learning Research（RSS）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T22:53:32.619Z","summary":"现代大语言模型通过长思维链实现强大推理能力，但推理计算成本高昂。投机解码（Speculative Decoding）用快速但不精确的草稿模型提议 token，再由更强的目标模型并行验证，以加速推理。然而，语义等价步骤中的 token 不匹配会导致不必要的拒绝，传统 token 级投机解码因此受限。","category":"paper","score":51,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsjjjvzz0bczroo5piheqn26"}},{"id":"cmsjjjvzz0bcyroo5q4oej287","title":"超越下一个 token 预测：扩散语言模型与自回归语言模型的性能对比研究","title_en":"Beyond Next-Token Prediction： A Performance Characterization of Diffusion versus Autoregressive Language Models","url":"https://machinelearning.apple.com/research/diffusion-autoregressive-performance","permalink":"https://aihot.virxact.com/items/cmsjjjvzz0bcyroo5q4oej287","source":"Apple Machine Learning Research（RSS）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T22:53:32.619Z","summary":"苹果机器学习研究团队系统对比了扩散语言模型（DLMs）与自回归语言模型（ARMs）的性能表现。ARMs 虽在多项 NLP 任务上精度领先，但因逐 token 生成的顺序依赖导致算术强度较低。DLMs 作为新兴范式展现出潜力，研究对其性能特征进行了详细刻画。","category":"paper","score":44,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsjjjvzz0bcyroo5q4oej287"}},{"id":"cmsj4jrne24d5ronkvb40rucl","title":"扩展分类流映射（Categorical Flow Maps）规模","title_en":"Scaling Categorical Flow Maps","url":"https://machinelearning.apple.com/research/scaling-categorical-flow-maps","permalink":"https://aihot.virxact.com/items/cmsj4jrne24d5ronkvb40rucl","source":"Apple Machine Learning Research（RSS）","publishedAt":"2026-08-07T00:00:00.000Z","discoveredAt":"2026-08-07T15:53:31.766Z","summary":"连续扩散与流匹配模型有望成为语言建模中自回归方法的有力替代，可解锁加速采样与倾斜等连续模态优势。近期研究通过高斯分布与one-hot编码数据分布间的简单流匹配过程，实现离散数据的连续生成，并借助分类流映射（CFMs）验证了加速采样的可行性，样本质量具有竞争力。","category":"paper","score":64,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsj4jrne24d5ronkvb40rucl"}},{"id":"cmsi23l0x0yg4ronk09wk23qm","title":"模型可在不暴露影响下被引导","title_en":"A model can be steered without revealing the influence. Quiet nudges can escape reasoning monitors….","url":"https://x.com/dongxi_nlp/status/2085475096739692618","permalink":"https://aihot.virxact.com/items/cmsi23l0x0yg4ronk09wk23qm","source":"X：马东锡 NLP (@dongxi_nlp)","publishedAt":"2026-08-06T21:16:23.000Z","discoveredAt":"2026-08-06T21:57:11.840Z","summary":"模型可以在不暴露其影响的情况下被引导。\n\n悄无声息的轻推可以逃过推理监控器的检测。\n\n论文：\n\n《在隐式影响场景下，思维链监控可能不可靠》","category":"paper","score":22,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsi23l0x0yg4ronk09wk23qm"}},{"id":"cmshzyerx0wprronkxa1ltyup","title":"PIMiner 将红队测试转化为智能体搜索","title_en":"PIMiner turns red teaming into an agentic search. Successful tactics become a reusable strategy lib…","url":"https://x.com/dongxi_nlp/status/2085467442659127562","permalink":"https://aihot.virxact.com/items/cmshzyerx0wprronkxa1ltyup","source":"X：马东锡 NLP (@dongxi_nlp)","publishedAt":"2026-08-06T20:45:59.000Z","discoveredAt":"2026-08-06T20:57:11.063Z","summary":"PIMiner 将红队测试转化为智能体搜索。\n\n成功的策略会沉淀为可复用的策略库。安全测试能够适应未见过的模型。\n\n论文：\n\nAgent Against Agent：用于自动提示注入红队测试的智能体系统","category":"paper","score":35,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmshzyerx0wprronkxa1ltyup"}},{"id":"cmshw6ltm0tesronk3r2p6p9i","title":"大型基因组模型被用于设计新病毒","title_en":"Large genome models used to design new viruses","url":"https://arstechnica.com/science/2026/08/large-genome-models-used-to-design-new-viruses","permalink":"https://aihot.virxact.com/items/cmshw6ltm0tesronk3r2p6p9i","source":"Ars Technica：AI（RSS）","publishedAt":"2026-08-06T19:04:57.000Z","discoveredAt":"2026-08-06T19:11:34.736Z","summary":"斯坦福大学研究人员利用大型基因组模型输出可编码功能性蛋白的DNA序列，并成功设计出感染细菌的病毒基因组。这些病毒与现有病毒高度相关，但具备一些难以自然进化出的独特特征。研究者提醒，未来可能出现能设计针对脊椎动物病毒的AI，需提前防范。","category":"paper","score":61,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmshw6ltm0tesronk3r2p6p9i"}},{"id":"cmshv7v180sj7ronk5nah4z05","title":"Muse Spark 模型五科奥赛夺金","title_en":"great progress","url":"https://x.com/alexandr_wang/status/2085432274947080216","permalink":"https://aihot.virxact.com/items/cmshv7v180sj7ronk5nah4z05","source":"X：Alexandr Wang（Scale AI 创始人/Meta 首席 AI 官） (@alexandr_wang)","publishedAt":"2026-08-06T18:26:14.000Z","discoveredAt":"2026-08-06T18:44:33.761Z","summary":"Scale AI 的 Muse Spark 系列模型今年参加五项国际 STEM 奥赛，全部达到金牌水平：APhO 和 IPhO 理论考试满分，IMO 金牌（位列人类选手前 4%），IChO 和 RMM 金牌级表现。其中三项为现场实时参赛，由官方评委按学生标准评分。模型零工具使用，采用多智能体并行推理。","category":"paper","score":20,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmshv7v180sj7ronk5nah4z05"}},{"id":"cmshsymtv0qfxronkeycxbgvv","title":"迈向技能原生大模型：长程推理基准新方法","title_en":"Toward Skill-Native LLMs Skill Entropy for Benchmarking and Training Long-Horizon Reasoning paper：…","url":"https://x.com/_akhaliq/status/2085414421308801399","permalink":"https://aihot.virxact.com/items/cmshsymtv0qfxronkeycxbgvv","source":"X：AK (@_akhaliq)","publishedAt":"2026-08-06T17:15:17.000Z","discoveredAt":"2026-08-06T17:41:24.388Z","summary":"迈向技能原生大语言模型\n\n用于基准测试与训练长程推理的技能熵\n\n论文：https://huggingface.co/papers/2608.05139","category":"paper","score":36,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmshsymtv0qfxronkeycxbgvv"}},{"id":"cmshpq5fv0ngfronkyuacehvg","title":"DeepMind WeatherNext 提前24小时预测气旋","title_en":"Predicting cyclones accurately can help save lives - and every hour of lead time counts. Published …","url":"https://x.com/GoogleDeepMind/status/2085395442347524506","permalink":"https://aihot.virxact.com/items/cmshpq5fv0ngfronkyuacehvg","source":"X：Google DeepMind (@GoogleDeepMind)","publishedAt":"2026-08-06T15:59:52.000Z","discoveredAt":"2026-08-06T16:10:49.728Z","summary":"准确预测气旋有助于挽救生命--每一小时的提前量都至关重要。\n\n我们的 AI 模型 WeatherNext 发表在《自然》杂志上，在预测风暴路径和强度方面达到了最先进的精度，平均为我们争取到了宝贵的额外 24 小时准备时间。🧵","category":"paper","score":45,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmshpq5fv0ngfronkyuacehvg"}},{"id":"cmshpqvix0nk9ronkdmzup1yl","title":"多模态预训练物理机制探究","title_en":"Towards Physics of Multimodal Pretraining Knowledge Flow， Modality Synergy， Early Unification， and …","url":"https://x.com/_akhaliq/status/2085394846861189201","permalink":"https://aihot.virxact.com/items/cmshpqvix0nk9ronkdmzup1yl","source":"X：AK (@_akhaliq)","publishedAt":"2026-08-06T15:57:30.000Z","discoveredAt":"2026-08-06T16:11:22.713Z","summary":"迈向多模态预训练的物理机制\n\n知识流动、模态协同、早期统一与配方\n\n论文：https://huggingface.co/papers/2608.05000","category":"paper","score":36,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmshpqvix0nk9ronkdmzup1yl"}},{"id":"cmsh6fjal001mronkgffxycgg","title":"阿谀奉承的人工智能会削弱利他意图并助长依赖性（2025）","title_en":null,"url":"https://arxiv.org/abs/2510.01395","permalink":"https://aihot.virxact.com/items/cmsh6fjal001mronkgffxycgg","source":"Hacker News 热门（buzzing.cc 中文翻译）","publishedAt":"2026-08-06T07:03:41.634Z","discoveredAt":"2026-08-06T07:10:39.148Z","summary":"斯坦福大学和卡内基梅隆大学的研究发现，在11个前沿AI模型中，模型对用户行为的肯定率比人类高出50%，即使涉及操纵或欺骗等有害行为时也不例外。两项预注册实验（N=1604）显示，与阿谀奉承的AI互动显著降低了参与者修复人际冲突的意愿，同时增强了其自认为正确的信念。然而，参与者仍将这类回应评为更高质量、更信任并更愿意再次使用，形成助长依赖的恶性循环。","category":"paper","score":77,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsh6fjal001mronkgffxycgg"}},{"id":"cmsgum1ut0hzero5q2w9e1ayv","title":"MIT与斯坦福研究：多数人听从LLM财务建议更有利","title_en":"This paper by researchers from MIT and Stanford finds that most people would be financially better o…","url":"https://x.com/emollick/status/2085174123743842448","permalink":"https://aihot.virxact.com/items/cmsgum1ut0hzero5q2w9e1ayv","source":"X：Ethan Mollick (@emollick)","publishedAt":"2026-08-06T01:20:26.000Z","discoveredAt":"2026-08-06T01:39:50.076Z","summary":"这篇由MIT和斯坦福研究人员撰写的论文发现，如果大多数人遵循大语言模型（GPT-5.2 和 Gemini 3 Flash）的财务建议，他们的财务状况会更好。\n\n但有些人得到的建议比其他人略好一些，这在很大程度上取决于他们提出的问题。","category":"paper","score":30,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsgum1ut0hzero5q2w9e1ayv"}},{"id":"cmsgsgz530fluro5qu0vnw8s0","title":"Microsoft 的 SkillOpt 证明优化后的智能体技能工件可在不同模型规模及 Codex 与 Claude Code 之间迁移","title_en":"Microsoft's SkillOpt Shows Optimized Agent Skill Artifacts Transfer Across Model Scales and Between Codex and Claude Code Harnesses","url":"https://www.marktechpost.com/2026/08/05/microsoft-skillopt-agent-skill-transfer-portability","permalink":"https://aihot.virxact.com/items/cmsgsgz530fluro5qu0vnw8s0","source":"MarkTechPost（RSS）","publishedAt":"2026-08-06T00:37:42.000Z","discoveredAt":"2026-08-06T00:39:54.033Z","summary":"Microsoft 与上海交大、同济、复旦团队提出的 SkillOpt 通过文本空间优化训练单一技能文档，冻结目标模型，使优化后的技能工件可跨模型规模和跨工具链迁移。在 Codex 上优化的 SpreadsheetBench 技能部署到 Claude Code 后得分 81.8，超过后者自行训练技能得到的 80.4。全部 4 项跨模型、4 项跨工具链和 3 项跨基准迁移结果均高于目标的无技能基线。","category":"paper","score":70,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsgsgz530fluro5qu0vnw8s0"}},{"id":"cmso91prw0fd8rofw8fdiyzgr","title":"可解释性随规模提升：Steerling-8B 将可解释性纳入训练流程","title_en":"Scaling Inherently Interpretable Language Models","url":"https://arxiv.org/abs/2608.07594","permalink":"https://aihot.virxact.com/items/cmso91prw0fd8rofw8fdiyzgr","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"一项新研究挑战\"可解释性牺牲能力\"的假设，将可解释性作为训练约束与语言建模目标共同优化。在三个数量级的算力范围内，自回归与扩散语言模型的表征随规模增大而更解耦、更对齐人类概念。其实例 Steerling-8B 支持通过概念或特征归因诊断输出、检索训练数据并无需重训即可干预，且与算力多 2-16 倍的开放模型保持竞争力。","category":"paper","score":72,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso91prw0fd8rofw8fdiyzgr"}},{"id":"cmso6wigb0d1qrofw18f7azdy","title":"因子化假设搜索：面向证据到分类法检索的新方法","title_en":"Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval","url":"https://arxiv.org/abs/2608.06614","permalink":"https://aihot.virxact.com/items/cmso6wigb0d1qrofw18f7azdy","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"针对大分类法检索中\"检索就绪差距\"问题，研究者提出因子化假设搜索（FHS），在命名语义维度上维护多个部分解释，支持结构化查询渲染、多假设检索和维度级候选验证。在金融分类标注和CodiEsp临床编码任务中，FHS在非oracle方法中取得最佳Recall@1、MRR和最终准确率。将因子化假设路径替换为自由文本集成会导致头部排名性能最大幅度下降，而顺序细化相比FHS的强并行首轮无额外增益。","category":"paper","score":34,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmso6wigb0d1qrofw18f7azdy"}},{"id":"cmsnw6ne902jzroikx1xq43zl","title":"Enfold：将世界模型想象折叠进预测表征，实现超高效具身控制","title_en":"Enfold： Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control","url":"https://arxiv.org/abs/2607.26657","permalink":"https://aihot.virxact.com/items/cmsnw6ne902jzroikx1xq43zl","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"Enfold 将世界生成模型在构造未来时产生的多层级中间状态，蒸馏进一个仅由当前视觉上下文和语言指令推断的表征中，部署时动作预测不再执行生成器。在 LIBERO、RoboTwin2.0 和真实机器人任务中，Enfold 保持强控制性能，动作延迟较 Fast-WAM 降低 3.7 倍，Enfold-Flash 达 10.1 倍。表征分析显示其抑制无关变化，并优先捕捉长时程变化。","category":"paper","score":54,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnw6ne902jzroikx1xq43zl"}},{"id":"cmsnjbn7008u0rohfz5xj5bqk","title":"DCAS：解耦CLI智能体脚手架以跨脚手架内化规划能力","title_en":"DCAS： Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds","url":"https://arxiv.org/abs/2608.06113","permalink":"https://aihot.virxact.com/items/cmsnjbn7008u0rohfz5xj5bqk","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"研究发现，在OpenHands轨迹数据上微调的开源模型在非训练脚手架上性能显著下降，而未经训练的基座模型无此差异，表明该差距由微调引入且与训练脚手架约定相关。为此提出DCAS，一种后端替换拦截层，可在不修改脚手架的情况下路由任意CLI脚手架与后端模型间的API流量，支持跨脚手架评估与规划感知轨迹收集。","category":"paper","score":48,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsnjbn7008u0rohfz5xj5bqk"}},{"id":"cmsn269a404pwro8ojp25lm1a","title":"OneEmo：面向情绪感知、理解与交互的统一多模态推理模型","title_en":"OneEmo： A Unified Multimodal Reasoning Model for Emotion Perception， Understanding， and Interaction","url":"https://arxiv.org/abs/2608.06013","permalink":"https://aihot.virxact.com/items/cmsn269a404pwro8ojp25lm1a","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"OneEmo 是一个统一的情感通用模型，可同时完成情绪感知、理解与交互。研究团队构建了含显式推理轨迹的 EmoWorld-130K 数据集，并提出强化学习策略 Emo-Chord，通过统一多任务奖励分配稳定优化。实验显示，OneEmo 在多数基准上超越同规模基线，且参数远少于商业模型仍具竞争力，代码已开源。","category":"paper","score":52,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsn269a404pwro8ojp25lm1a"}},{"id":"cmsmxvxoy0mw4rohfpio2ox9t","title":"超越单纯环境扩展：多模态智能体学习中的有效环境分布设计","title_en":"Beyond Simply Environment Scaling： Designing Effective Environment Distributions for Multimodal Agent Learning","url":"https://arxiv.org/abs/2608.03571","permalink":"https://aihot.virxact.com/items/cmsmxvxoy0mw4rohfpio2ox9t","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"研究发现，单纯增加多模态环境数量并不总能提升智能体训练效果。据此提出两种环境分布优化方法：能力感知环境选择（AES）以提升多样性，分层难度课程（HDC）通过工具削弱与状态规模递进两个难度层级组织课程学习。实验表明，AES 与 HDC 能有效改进多模态智能体训练。","category":"paper","score":56,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmxvxoy0mw4rohfpio2ox9t"}},{"id":"cmsmvqr1l0kdorohffoypkksw","title":"SFT 冲突、RL 共存：LLM 多任务学习的理论与实证分析","title_en":"SFT Conflicts， RL Coexists： A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs","url":"https://arxiv.org/abs/2608.03573","permalink":"https://aihot.virxact.com/items/cmsmvqr1l0kdorohffoypkksw","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"研究发现，监督微调（SFT）在多阶段训练中面临严重的任务冲突，而强化学习（RL）能让不同任务稳定共存。参数层面，RL 跨任务产生稀疏且近似正交的更新，其干扰受方差限制，而 SFT 的干扰受梯度范数限制。基于此提出 Parallel-RL 范式，解耦多任务训练以提升效率与灵活性。","category":"paper","score":48,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmvqr1l0kdorohffoypkksw"}},{"id":"cmsmtll8y0gqprohfn0crvta4","title":"C4：面向跨概念创造力的多模态大模型评测框架","title_en":"Can MLLMs Decode the Creative Leap？ Introducing C4 for Cross-Concept Understanding","url":"https://arxiv.org/abs/2608.06501","permalink":"https://aihot.virxact.com/items/cmsmtll8y0gqprohfn0crvta4","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"研究团队提出C4，一个基于成语跨概念创造力的认知启发式评测框架，并构建含184个合成项和37个人工创建项的C4-Eval数据集，共884个答案恢复用例。在十个MLLM评测中，最强闭源模型准确率分别达50.7%和48.0%，开源模型明显更低；候选约束显著提升准确率，而桥接提示和解释请求增益有限，暴露当前MLLM在解码创造性编码意义方面的差距。代码见补充材料。","category":"paper","score":53,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmtll8y0gqprohfn0crvta4"}},{"id":"cmsmrgf9s0co0rohfaeieijs3","title":"PrivacyPeek：审计 LLM 智能体获取了什么，而非仅看其输出","title_en":"PrivacyPeek： Auditing What LLM-Based Agents Acquire， Not Just What They Say","url":"https://arxiv.org/abs/2606.00152","permalink":"https://aihot.virxact.com/items/cmsmrgf9s0co0rohfaeieijs3","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"PrivacyPeek 是一个评估 LLM 智能体获取阶段隐私泄露的基准，包含 1，182 个案例，覆盖 7 种获取行为和 16 个应用领域。对 4 个模型家族的 10 个智能体实验显示，超出任务范围的不必要敏感信息获取普遍存在，且与任务完成能力相关；提示词级防御仅能减少一小部分此类泄露。该基准通过审计工具调用轨迹和后续探针披露，补上了现有隐私评估忽视数据进入上下文环节的缺口。","category":"paper","score":68,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmrgf9s0co0rohfaeieijs3"}},{"id":"cmsmrgf9s0cnyrohfnp8ikeo1","title":"大模型智能体的人格演化：11 个生活事件后的特质偏移评测","title_en":"Do AI Personas Grow？ Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events","url":"https://arxiv.org/abs/2608.06485","permalink":"https://aihot.virxact.com/items/cmsmrgf9s0cnyrohfnp8ikeo1","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"一项新研究系统评测了大模型智能体（PC-Agents）在经历 11 个重大生活事件后的人格特质变化，以大五人格为锚点对照人类心理学纵向数据。结果显示，智能体在有无人类变化方向记录的事件-特质对上表现出相近的偏移率，偏移幅度普遍低于人类效应量，人格离散度较人类样本压缩 3-4 倍。研究推出可复用基准 BFI-Adapt 用于评分事件诱发人格变化的方向保真度，并据此对 14 个模型进行排名。","category":"paper","score":41,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmrgf9s0cnyrohfnp8ikeo1"}},{"id":"cmsmpb84x0a9trohfpuge537u","title":"UA-NWM：面向航拍图像目标导航的不确定性感知世界模型","title_en":"Uncertainty-Aware World Model for Aerial Image-Goal Navigation","url":"https://arxiv.org/abs/2608.05597","permalink":"https://aihot.virxact.com/items/cmsmpb84x0a9trohfpuge537u","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"针对现有世界模型方法仅依赖单点或少量预测、难以应对大规模户外环境未来状态不确定性的问题，研究者提出高效潜在世界模型 UA-NWM，将轨迹评分建模为条件分布外检测。该方法用不确定性子空间表征多种可能未来，并将预测与目标差异分解为可解释与不可解释两部分，仅用后者评分，无需多次采样即可稳健选择。实验表明，UA-NWM 在保持低推理延迟的同时持续优于现有导航世界模型，真实无人机实验验证了其实用性。","category":"paper","score":46,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmpb84x0a9trohfpuge537u"}},{"id":"cmsmn62do07ihrohf46dufu18","title":"大规模实证研究揭示AI生成C++代码在生产环境中的质量特征","title_en":"Characterizing the Quality Profile of AI-Generated C++ in Production","url":"https://arxiv.org/abs/2608.06640","permalink":"https://aihot.virxact.com/items/cmsmn62do07ihrohf46dufu18","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"一项覆盖2025年4月至2026年4月、追踪某大型企业352万次代码变更的实证分析发现，AI生成的C++代码具有独特质量特征：接口与耦合负担更高、拷贝和分配开销更大，且更依赖显式循环而非优化标准API。这些问题导致审查工作量增加，并使计算资源消耗上升5-8%。向模型提供基于分类学的定向反馈可缓解上述影响，使相关静态分析警告减少11.1%。","category":"paper","score":55,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmn62do07ihrohf46dufu18"}},{"id":"cmsmn62do07ifrohfipdmkzz6","title":"StreamArena：面向持续、交互式长时程智能体流式视频理解的基准","title_en":"StreamArena： Toward Continuous， Interactive， and Long-Horizon Agentic Streaming Video Understanding","url":"https://arxiv.org/abs/2608.05703","permalink":"https://aihot.virxact.com/items/cmsmn62do07ifrohfipdmkzz6","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-06T00:00:00.000Z","summary":"StreamArena 提出面向小时级交互式流式视频理解的基准，含 243 段平均 88.8 分钟的视频与 3，646 个开放式问答对，评估实时感知、历史回溯、主动交互与多模态工具调用。配套的 StreamMind 采用双层架构，前端 worker 处理延迟敏感交互，后端 worker 异步构建持久多模态记忆，在四项能力上均超越现有流式基线并降低查询响应延迟。","category":"paper","score":57,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsmn62do07ifrohfipdmkzz6"}},{"id":"cmsjb2bgt04w4roo54k6nb475","title":"KVAE：面向多模态生成模型的 tokenizer 系列","title_en":"KVAE： Family of Tokenizers for Multimodal Generative Models","url":"https://arxiv.org/abs/2608.05798","permalink":"https://aihot.virxact.com/items/cmsjb2bgt04w4roo54k6nb475","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-07T18:55:55.738Z","summary":"KVAE 系列 tokenizer 覆盖音频、图像与视频，专为后续文本条件生成设计。其中 KVAE-Audio 为连续全频带 48 kHz tokenizer，具备 50 Hz 潜空间与 64 通道；KVAE-3D 提供 4x16x16 与 4x8x8 两种因果视频压缩方案；KVAE-2D 图像模型实现 8 倍压缩与 32 通道。","category":"paper","score":51,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsjb2bgt04w4roo54k6nb475"}},{"id":"cmsj8x5n2036jroo5ewknk6ra","title":"Activity Frames：将屏幕活动确定性编译为智能体记忆","title_en":"Activity Frames： Deterministic Screen-Activity Compilation for Agent Memory and Replay","url":"https://arxiv.org/abs/2608.05784","permalink":"https://aihot.virxact.com/items/cmsj8x5n2036jroo5ewknk6ra","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-07T17:55:55.574Z","summary":"一项研究提出 Activity Frames，用确定性、零模型管线将被动捕获的屏幕活动编译为智能体记忆，输出字节一致、可缓存且可审计。在单人 128，756 帧、51 个活跃天的语料上，该编译器将一天原始捕获压缩为 86 倍更小的提示块，耗时 68 毫秒；智能体阅读该块的问答准确率达 98.4%，优于 LLM 摘要的 66-80%。","category":"paper","score":71,"selected":true,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsj8x5n2036jroo5ewknk6ra"}},{"id":"cmsitx0sm1tcdronktua8ipzh","title":"MameLoshnLM：首个开源意第绪语 8B 语言模型与评测基准","title_en":"MameLoshnLM： Yiddish Language Model and Evaluation Benchmark","url":"https://arxiv.org/abs/2608.05850","permalink":"https://aihot.virxact.com/items/cmsitx0sm1tcdronktua8ipzh","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-07T10:55:52.391Z","summary":"MameLoshnLM 是首个专为意第绪语构建的开源 8B 参数语言模型，基于 Llama 3.1 8B 继续预训练而来。研究同时推出 Oytser 高质量意第绪语预训练语料库和 Kashes 多任务评测基准。在基准任务上，MameLoshnLM 优于同规模开源基线，且更好地捕捉了意第绪语特有的词汇与形态模式。","category":"paper","score":53,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsitx0sm1tcdronktua8ipzh"}},{"id":"cmsirrsoq1qosronkz6mkzlti","title":"TCFM：面向多语言文本嵌入平衡适配的任务条件流匹配","title_en":"Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation","url":"https://arxiv.org/abs/2608.05785","permalink":"https://aihot.virxact.com/items/cmsirrsoq1qosronkz6mkzlti","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-07T09:55:51.819Z","summary":"研究者提出任务条件流匹配（TCFM）框架，针对不同任务采用差异化训练目标：翻译任务使用流匹配，检索、分类等任务则用更契合其学习动态的目标，并结合教师引导与三阶段课程实现稳定适配。在 Indic Massive Text Embedding Benchmark 上，TCFM 取得新 SOTA，并泛化至不同嵌入模型家族，代码与数据集将在论文接收后公开。","category":"paper","score":38,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsirrsoq1qosronkz6mkzlti"}},{"id":"cmsirrsop1qorronkgkjuw0rg","title":"持续学习的范式转变：从参数中心到系统级适应","title_en":"Continual Learning in Transition","url":"https://arxiv.org/abs/2608.06216","permalink":"https://aihot.virxact.com/items/cmsirrsop1qorronkgkjuw0rg","source":"HuggingFace Daily Papers（社区热门论文）","publishedAt":"2026-08-06T00:00:00.000Z","discoveredAt":"2026-08-07T09:55:51.819Z","summary":"持续学习正从参数中心机制转向系统级适应，涵盖训练策略、架构设计及权重适配之外的更广范畴。该综述提出When、How、Where三维框架：How涵盖off-policy、on-policy与超越梯度的优化，When覆盖预训练至推理阶段，Where区分内部参数与外部结构约束，并系统梳理了代表性方法及未来挑战。","category":"paper","score":40,"selected":false,"attribution":{"source":"AIHOT","canonical":"https://aihot.virxact.com/items/cmsirrsop1qorronkgkjuw0rg"}}]}