Languages are not interchangeable labels. A single language contains differences in script, region, register, professional domain and cultural experience; translation scores, knowledge-test accuracy and real dialogue completion answer different questions. Evaluation must preserve these differences rather than erase them with one average.
Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, country, language, culture, benchmark or leaderboard. It is not certification, audit, procurement, investment, legal, policy or compliance advice. C-Eval and CMMLU are presented as Chinese research contributions to the method, not endorsements or permanent rankings.