AB
AiBoss
project

SuperGPQA - A knowledge reasoning benchmark suite jointly developed by Doubao's large model and M-A-P open source.

SuperGPQA is a knowledge reasoning benchmark test suite jointly launched by ByteDance's Doubao Big Model Team and M-A-P. It comprehensively covers 285 graduate-level disciplines and contains 26,529 professional questions. It addresses the limitations of traditional benchmark testing in terms of subject coverage...

What is SuperGPQA?

SuperGPQA is a comprehensive knowledge reasoning benchmark test suite launched by ByteDance's Doubao Big Model team in collaboration with M-A-P. It covers 285 graduate-level disciplines and includes 26,529 professional questions. Addressing the shortcomings of traditional benchmark tests such as incomplete subject coverage, questionable question quality, and limited evaluation dimensions, SuperGPQA is built collaboratively by experts and a large language model, ensuring high-quality and challenging questions. SuperGPQA includes both STEM and non-STEM subjects, with 42.33% of the questions requiring mathematical calculations or rigorous reasoning, effectively measuring the generalization ability and true reasoning level of the large language model.

Main functions of SuperGPQA

  • A comprehensive evaluation of the generalization ability of the Large Language Model (LLM)Covering 285 graduate-level disciplines (including long-tail disciplines), SuperGPQA can comprehensively measure LLM students' knowledge reserves and reasoning abilities in different fields.
  • Revealing the model's true reasoning ability42.33% of the questions require mathematical calculations or formal reasoning, ensuring that the test set effectively evaluates the model's performance on complex tasks, not just knowledge memorization.
  • Provide an interdisciplinary analysis frameworkSuperGPQA has broad subject coverage, encompassing both STEM (science, technology, engineering, and mathematics) and non-STEM (philosophy, literature, history, etc.) fields, providing a unified assessment tool for the performance of research models across different disciplines.
  • Filling the gap in long-tail subject evaluationTraditional evaluation sets do not cover long-tail disciplines (such as light industry, agriculture, service science, etc.). SuperGPQA makes up for this deficiency by covering all disciplines.
  • Provide a reference for model optimizationBased on the evaluation results on SuperGPQA, we identified the shortcomings of the model and optimized the model architecture and training methods.

SuperGPQA Technical Principles

  • Expert-LLM Collaborative Construction:
    • Source filteringExperts select and collect original questions from credible sources (such as textbooks and authoritative practice websites) to avoid the low-quality risks of crowdsourced labeling.
    • Transcription and NormalizationExperts standardized the language and format of the original questions to ensure that all questions used a consistent academic language and a standard multiple-choice question format.
    • Quality InspectionThe high quality and high discrimination of the questions are ensured through rule-based initial filtering, LLM-based quality checks (such as validity and domain relevance assessment), and expert review.
  • Multi-model collaborative validationDuring the quality inspection phase, multiple advanced LLMs (such as GPT-4, Gemini-flash, etc.) are used for multi-dimensional testing to reduce the risk of data leakage and improve the reliability and discrimination of the questions.
  • Interdisciplinary semantic structure designBased on visualization techniques such as t-SNE, we analyze the semantic structure of questions to ensure that the linguistic characteristics of different disciplines are preserved and semantic similarity is maintained in engineering and science questions.
  • High-difficulty task design42.33% of the questions require mathematical calculations or rigorous reasoning, ensuring that the test set effectively evaluates the model's performance on complex tasks, not just knowledge memorization.

SuperGPQA's project address

Application scenarios of SuperGPQA

  • Model performance evaluation: To comprehensively measure the knowledge and reasoning ability of a large language model across multiple disciplines.
  • Model optimization guidanceIt helps researchers identify model deficiencies and optimize training strategies.
  • Interdisciplinary analysisSupports comparative studies of model capabilities across different disciplines.
  • Educational ResearchUsed for developing intelligent educational tools and researching the application of AI in education.
  • Industry application testing: Provides testing tools for industry applications such as intelligent customer service and medical assistance.