# NOLLI：用于诊断英语-韩语性能差距的难度校准谜题基准

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-08-05 08:00
- AIHOT 分数：46
- AIHOT 链接：https://aihot.virxact.com/items/cmsgxbipf02rhroxz4nvmaeuk
- 原文链接：https://arxiv.org/abs/2608.04397

## AI 摘要

NOLLI 是一个程序化生成的英语-韩语谜题基准，含 15 种谜题类型（25 个任务、7,500 个条目），每个实例可重新生成种子、有唯一解并确定性评分。在 12 个超过 3% 整体准确率下限的模型中，匹配的英韩准确率在 ±10 个百分点（TOST）内统计等效；韩语密码落后英语最多 68.7 个百分点，而韩文拼写块组成准确率可预测韩语密码准确率。

## 正文

We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically. Rather than equating harder with bigger, we calibrate difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. Its three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography. We evaluate 15 frontier, open-weight, and Korean-developed models; among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a +/- 10 pp margin (TOST), suggesting little cost from presentation language alone. Writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 pp, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy. These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.
