Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We introduce FlavourBench, an automated benchmark in which a versioned culinary system supplies dense, executable ground truth. Each task presents eight ingredients and asks for a three-ingredient portfolio; before model execution, Epicure scores all 56 possible portfolios. We evaluate 27 frontier endpoints on an identical 534-task core spanning substitution, pairing, and constrained composition. Every ranked model has exactly 89 valid responses per panel and family (14,418 model-task cells total), eliminating differential missingness from the leaderboard. The FlavourBench Score is the equal-family mean of the frozen task scores. We use 50,000 anchor-cluster bootstrap replicates for simultaneous 95% score bands and 100,000 sign-flip draws for all 351 paired model contrasts, with Holm control. The two independently compiled panels correlate at r = 0.89 (rank rho = 0.80). Grok 4.6 has the largest point estimate at 65.1 (simultaneous 95% CI 61.0-69.2); 101 of 351 model pairs are resolved. The release includes the prompts, all portfolio score maps, raw responses, exact routes, content hashes, and an offline verifier that reconstructs every result.
FlavourBench:用可执行烹饪基准评测前沿语言模型
AI 导读
FlavourBench 推出自动化基准,以版本化烹饪系统提供可执行真值,在 534 项任务上评测 27 个前沿模型端点。Grok 4.6 得分最高,为 65.1(95% 置信区间 61.0-69.2),351 对模型对比中 101 对差异显著。基准发布含提示词、评分映射、原始响应及离线验证器,可复现全部结果。
HuggingFace Daily Papers(社区热门论文)
49
AI 编辑部评分,满分 100FlavourBench:用可执行烹饪基准评测前沿语言模型
FlavourBench 推出自动化基准,以版本化烹饪系统提供可执行真值,在 534 项任务上评测 27 个前沿模型端点。Grok 4.6 得分最高,为 65.1(95% 置信区间 61.0-69.2),351 对模型对比中 101 对差异显著。基准发布含提示词、评分映射、原始响应及离线验证器,可复现全部结果。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org