AB
AiBoss
project

LalaEval - A joint venture between The Chinese University of Hong Kong and Lalamove, launching a model evaluation framework for specific sectors.

LalaEval is a human evaluation framework for domain-specific Large Language Models (LLMs), jointly developed by the Chinese University of Hong Kong and the Lalamove Data Science team. The framework utilizes a complete end-to-end protocol, covering domain specifications, standards, etc.

What is LalaEval?

LalaEval is a human evaluation framework for domain-specific Large Language Models (LLMs), jointly developed by the Chinese University of Hong Kong and the Lalamove Data Science team. The framework utilizes a complete end-to-end protocol, covering domain specification, standard establishment, benchmark dataset creation, evaluation rule construction, and the analysis and interpretation of evaluation results. Its core feature is the automatic correction of subjective human errors through controversy and score fluctuation analysis, generating high-quality question-answer pairs. LalaEval employs a single-blind testing principle to ensure the objectivity and fairness of the scoring. It has already been successfully applied in the logistics field.

LalaEval's main functions

  • Scope of the fieldDefine the scope and boundaries of a specific area, relating it to the organization's goals or business needs. In the logistics field, this means moving from the most basic sub-sectors (such as local freight) to broader sub-domains.
  • Capability indicator constructionDefine the capability dimensions for evaluating the performance, effectiveness, or applicability of LLMs, including general capabilities and domain capabilities. General capabilities include semantic understanding, contextual dialogue, and factual accuracy; domain capabilities involve understanding concepts and terminology, knowledge of industry policies, etc.
  • Evaluation set generationDevelop standardized tests and collect data from vetted information sources to evaluate under consistent conditions.
  • Evaluation criteria developmentDesign a detailed scoring scheme to provide a structured framework for human evaluators and ensure the scientific rigor and reliability of the assessment.
  • Statistical analysis of resultsThe system systematically examines and evaluates data during the evaluation process. Through analytical frameworks such as scoring controversy, question controversy, and scoring volatility, it automatically performs quality inspection of scoring results, secondary identification of low-quality QA, and quantitative attribution of the causes of scoring volatility.

LalaEval's technical principles

  • Single-blind testing principleDuring the evaluation process, the model's response was anonymized and presented to at least three human evaluators in a random order.
  • Analysis of Controversy and Rating FluctuationsLalaEval automatically detects and corrects subjective errors in human scoring by establishing three analytical frameworks: scoring controversy, question controversy, and scoring volatility.
  • Structured evaluation processLalaEval employs an end-to-end evaluation process, covering domain scope definition, capability indicator construction, evaluation set generation, evaluation standard formulation, and result statistical analysis.
  • Dynamic interactive deployment structureLalaEval's deployment structure emphasizes modularity and dynamic interaction, enabling flexible adjustments to the evaluation process based on different business scenarios and ensuring the framework's scalability across various domains.

LalaEval's project address

Application scenarios of LalaEval

  • Large-scale model evaluation in the logistics fieldLalaEval targets specific business scenarios such as intra-city freight. By defining the domain scope, constructing capability indicators, generating evaluation sets, and establishing assessment standards, LalaEval can scientifically evaluate the performance of large language models in the logistics industry, helping companies optimize their logistics business processes.
  • Invitation to review large modelsIn the driver invitation scenario, LalaEval evaluates the performance of large models in automated invitation tasks by simulating real-world dialogue scenarios.
  • Customization and optimization of large-scale enterprise modelsLalaEval provides enterprises with a standardized evaluation method that can dynamically generate evaluation sets based on their own business needs, reducing human subjectivity through automated analysis.
  • Scalability of cross-domain applicationsThe design follows the principles of modularity and dynamic interaction, and can be flexibly extended to other fields.