Meta 发布 WildArtifactBench 评测框架

AI at Meta · @AIatMeta · X·2026-08-21 02:25·14天前
AI 导读

Meta 发布内部评测框架 WildArtifactBench,用于评估多模态智能体在复杂真实任务中的表现。该框架采用人类与智能体偏好评判的胜率和 Elo 评分,替代严格的标准答案评分规则,以覆盖更多实际多模态工作流。Meta 已开源其中 10 个任务。

AI at Meta@AIatMeta
53AI 编辑部评分,满分 100

Meta 发布 WildArtifactBench 评测框架

2026-08-21 02:25· 14天前
AI 导读

Meta 发布内部评测框架 WildArtifactBench,用于评估多模态智能体在复杂真实任务中的表现。该框架采用人类与智能体偏好评判的胜率和 Elo 评分,替代严格的标准答案评分规则,以覆盖更多实际多模态工作流。Meta 已开源其中 10 个任务。

Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess agents on complex, real-world tasks across diverse deliverable formats.

By using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth rubrics, it expands task coverage across practical multimodal workflows.

We’re releasing 10 tasks from WildArtifactBench as a step forward in our ability to measure the real practical utility delivered by multimodal agents: https://research.meta.ai/wild-artifact-bench