The Decoder:AI News(RSS)
59AI 编辑部评分,满分 100

Anthropic's bio-weapons filter was down for nearly a year, exposing 133 million requests

2026-08-16 15:20· 12分钟前· Matthias Bastian

From the company whose CEO calls AI-assisted development of chemical and biological weapons a bigger threat than cyberattacks comes a safety report revealing that Anthropic's blocking biological classifiers were inactive from May 2025 through April 2026. These filters are designed to prevent AI models from being used to extract dangerous knowledge about chemical or biological weapons. For almost a year, all traffic from external contractors providing human feedback ran without them.

Screenshot

The gap affected a pool of about 50,000 people who ran roughly 133 million chats with the models. According to Anthropic, these individuals were vetted only by external vendors whose screening processes were often insufficient. Anthropic says its internal investigation turned up no evidence of actual misuse. The company has since tightened contractor requirements.

Anthropic also recently loosened its classifiers on Fable 5 after researchers complained the filters were so aggressive they blocked legitimate research.

Anthropic Redacted Risk Report

来源:The Decoder:AI News(RSS) · the-decoder.com

Anthropic's bio-weapons filter was down for nearly a year, exposing 133 million requests

The Decoder:AI News(RSS)·2026-08-16 15:20·12分钟前·Matthias Bastian
原文 · 保持原样,未翻译

From the company whose CEO calls AI-assisted development of chemical and biological weapons a bigger threat than cyberattacks comes a safety report revealing that Anthropic's blocking biological classifiers were inactive from May 2025 through April 2026. These filters are designed to prevent AI models from being used to extract dangerous knowledge about chemical or biological weapons. For almost a year, all traffic from external contractors providing human feedback ran without them.

Screenshot

The gap affected a pool of about 50,000 people who ran roughly 133 million chats with the models. According to Anthropic, these individuals were vetted only by external vendors whose screening processes were often insufficient. Anthropic says its internal investigation turned up no evidence of actual misuse. The company has since tightened contractor requirements.

Anthropic also recently loosened its classifiers on Fable 5 after researchers complained the filters were so aggressive they blocked legitimate research.

Anthropic Redacted Risk Report

来源:The Decoder:AI News(RSS)· the-decoder.com