AI Model Evaluation
LLMEval3
Large Model Evaluation Benchmarks Developed by Fudan University NLP Lab
标签:AI Model EvaluationAI model evaluationLLMEval is a large-scale model evaluation benchmark launched by the NLP Lab of Fudan University. The latest LLMEval-3 focuses on the evaluation of professional knowledge and ability, covering 13 disciplines and more than 50 sub-disciplines as defined by the Ministry of Education, including philosophy, economics, law, education, literature, history, science, engineering, agriculture, medicine, military science, management, and art, with a total of about 200,000 standard generative question-and-answer questions.