AI Model Evaluation

LLMEval3

Large Model Evaluation Benchmarks Developed by Fudan University NLP Lab

标签:

LLMEval is a large-scale model evaluation benchmark launched by the NLP Lab of Fudan University. The latest LLMEval-3 focuses on the evaluation of professional knowledge and ability, covering 13 disciplines and more than 50 sub-disciplines as defined by the Ministry of Education, including philosophy, economics, law, education, literature, history, science, engineering, agriculture, medicine, military science, management, and art, with a total of about 200,000 standard generative question-and-answer questions.

LLMEval3

相关导航