AI Model Evaluation

MMLU

Large-scale multi-task language understanding benchmark

标签:

MMLU, short for Massive Multitask Language Understanding, is a benchmark for language comprehension capabilities of large-scale models. It is one of the most well-known semantic understanding benchmarks for large models, launched by researchers at UC Berkeley in September 2020. The test covers 57 tasks, including elementary mathematics, US history, computer science, and law. The tasks cover a wide range of knowledge, and the language is English, used to evaluate the basic knowledge coverage and comprehension ability of large models.

MMLU

相关导航