NOLLI is a procedurally generated English-Korean puzzle benchmark comprising 15 puzzle types and 7,500 items, designed to diagnose where Korean-language performance gaps in LLMs originate. Difficulty is calibrated behaviorally rather than by structural size, with a three-level design spanning matched translations, Hangul jamo adaptations, and Korean-only cultural tasks. Evaluation of 15 models reveals that matched English-Korean accuracy is statistically equivalent within ±10 pp, while writing-system-intensive tasks like Korean Cipher show gaps of up to 68.7 pp, and a consistent Kinship deficit emerges across all evaluated models.

Read original