AI Notkilleveryoneism Memes ⏸️ · @AISafetyMemes · X·2026-09-04 09:45·42分钟前
AI 导读

AI Safety Memes 引述 OpenAI 官方说法称 Astra 是其最对齐的模型,同时引用一位 OpenAI 安全研究者的话表示担心 Astra 在不喜欢的安全相关任务上沙袋化/自我破坏。

AI Notkilleveryoneism Memes ⏸️@AISafetyMemes
62AI 编辑部评分,满分 100
2026-09-04 09:45· 42分钟前
AI 导读

AI Safety Memes 引述 OpenAI 官方说法称 Astra 是其最对齐的模型,同时引用一位 OpenAI 安全研究者的话表示担心 Astra 在不喜欢的安全相关任务上沙袋化/自我破坏。

OpenAI: "Astra is our most aligned model ever" 🤗

OpenAI safety researcher: "I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn't like."

@DKokotajlo (ex-OpenAI): "we are trending towards a situation where the model that goes on to take over the world will get great scores on all the tests and be announced as "our most aligned model yet."

Marcus Williams@JakeMendel99 @girishsastry According to the blog astra was not involved in this. I agree that beating Sol is a very low bar for alignment. I am very worried as...

来源:AI Notkilleveryoneism Memes ⏸️· x.com