AB
AiBoss
project

AgentCLUE-ICabin - AI Agent Evaluation Benchmark for Automotive Smart Cockpits

AgentCLUE-ICabin is an AI agent benchmark focused on automotive smart cockpit scenarios, comprehensively evaluating the tool invocation capabilities of large language models within the smart cockpit. The benchmark is built upon 12 common vehicle usage scenarios, covering everything from daily...

What is AgentCLUE-ICabin?

AgentCLUE-ICabin is an AI agent benchmark focused on automotive smart cockpit scenarios, comprehensively evaluating the tool-calling capabilities of a large language model within a smart cockpit. The benchmark is built upon 12 common car usage scenarios, covering various travel needs from daily commutes to long-distance driving, fully aligning with the actual interaction scenarios of domestic users. The evaluation design includes 1 to 10 rounds of multi-turn interactive dialogue, with each round invoking at least one tool, comprehensively examining the model's interaction capabilities in complex environments.

AgentCLUE-ICabin employs an objective 0/1 evaluation mechanism, ensuring the fairness of evaluation results by comparing the consistency of called functions and the system state after execution. The toolset is divided into five categories: travel, vehicle control, entertainment, safety, and general, covering more than 70 functions ranging from navigation to seat adjustment. The evaluation process includes scenario collection, toolset construction, dialogue data generation, and answer verification, ensuring the scientific rigor and practicality of the evaluation.

Main functions of AgentCLUE-ICabin

  • Scene buildingBased on 12 common car usage scenarios, such as daily commuting, long-distance self-driving, and family travel, an evaluation set was built to cover diverse travel situations.
  • Multi-turn interactionDesign 1 to 10 rounds of multi-round interactive dialogue, with at least one tool invoked in each round, to simulate the continuous dialogue requirements in real cockpit use.
  • Tool callThe tools for the smart cockpit are divided into five categories: travel, vehicle control, entertainment, safety, and general use, covering more than 70 functions and comprehensively covering the core functions of the smart cockpit.
  • Evaluation mechanismThe 0/1 evaluation method is adopted, which judges the correctness by comparing the consistency of the called function and the system state after the function is executed, so as to ensure that the result is fair and objective.
  • Data generationThe system utilizes large models to generate multi-round interactive dialogue data, which is then manually verified and optimized to form accurate QA pairs for automotive intelligent cockpits, providing standard samples for evaluation.

The technical principle of AgentCLUE-ICabin

  • Scenario-driven multi-turn interaction design
    • Scene buildingBased on 12 common car usage scenarios (such as daily commuting, long-distance driving, and family travel), we have built an evaluation set that closely reflects actual user needs. These scenarios cover the diverse needs of users in different situations.
    • Multi-turn interactionDesign multi-turn interactive dialogues ranging from 1 to 10 rounds, with each round invoking at least one tool. This multi-turn interaction design simulates the continuous dialogue needs of users in actual use of a smart cockpit, examining the model's performance in complex interactions.
  • Tool ClassificationThe tools in the smart cockpit are categorized into five main types: travel, vehicle control, entertainment, safety, and general, encompassing over 70 specific functions. For example:
    • Travel service toolsFeatures include navigation, traffic updates, and gas station location search.
    • Intelligent vehicle control toolsAir conditioning control, window control, seat adjustment, etc.
    • Entertainment service toolsMusic playback, radio listening, movie watching, etc.
    • Security Service ToolsFeatures include tire pressure monitoring, sentry mode, and child lock control.
    • General toolsFeatures include: seat adjustment, steering wheel adjustment, and headlight adjustment.
  • Tool callThe model needs to call the corresponding tools according to user instructions and ensure the accuracy of the call and the correctness of the execution results.
  • An objective and fair evaluation mechanism
    • 0/1 evaluation methodThe evaluation method judges correctness by comparing the consistency between the functions called by the model and the reference answer, as well as the changes in the system state after the functions are executed. This method is more objective and fair, avoiding the bias of subjective scoring.
    • Multi-round feedback mechanismThe model has a maximum of 3 attempts in each round of dialogue. The system will provide error feedback based on the model's call results, and the model can adjust accordingly.
  • Dialogue data generation: Utilize large models to generate multi-round interactive dialogue data to simulate real user interaction scenarios with the smart cockpit.
  • Manual verification and optimizationThe generated dialogue data and answers will be manually verified and optimized to ensure the accuracy and usability of the data, forming precise QA pairs for the car's smart cockpit.
  • State trackingIn multi-round interactions, the system tracks and manages changes in the cockpit's state. The model needs to consider the impact of each operation on the system state to ensure the correctness of subsequent operations.
  • State comparisonDuring the evaluation process, the system compares the system state after the model operation with the expected state to ensure that the model operation not only calls the correct commands but also correctly changes the system state.

AgentCLUE-ICabin's core advantages

  • Scene comprehensivenessIt covers 12 typical car usage scenarios, such as daily commuting, long-distance self-driving, and family travel, fully meeting the actual needs of domestic users and ensuring that the evaluation results have high practicality and reference value.
  • Interaction complexityDesign 1 to 10 rounds of multi-turn interactive dialogue, with at least one tool invoked in each round, to simulate the continuous dialogue needs in real-world use, examine the model's performance in complex interactions, and improve the depth and breadth of the evaluation.
  • Evaluation objectivityThe evaluation mechanism adopts a 0/1 evaluation method, which judges the correctness by comparing the consistency of the called function and the system state after execution, so as to ensure that the evaluation results are objective and fair and avoid interference from subjective factors.
  • Tool richnessThe intelligent cockpit tools are subdivided into five categories: travel, vehicle control, entertainment, safety, and general, covering more than 70 specific functions, comprehensively covering the core functions of the intelligent cockpit, and providing the model with a wealth of calling options.
  • Data accuracyThe system utilizes large models to generate multi-turn interactive dialogue data, which is then manually verified and optimized to form accurate QA pairs. This ensures the high quality and accuracy of the evaluation data, providing a reliable basis for model training and evaluation.

Application scenarios of AgentCLUE-ICabin

  • daily commuteIt helps users check traffic conditions, play music, and receive news during their commute, improving the convenience and comfort of their journey.
  • Long-distance self-driving tourIt provides accurate navigation, seat massage, and gas station search functions for long-distance travel, ensuring a smooth journey and comfortable driving experience.
  • Family travelTo meet the needs of families traveling with children, features include child locks, rear-seat entertainment options, and information on family-friendly facilities along the route, ensuring children's safety and travel convenience.
  • Working in a carIt creates a mobile office space, supporting functions such as Bluetooth teleconferencing, voice notes, and in-vehicle WiFi to meet users' needs for working in their cars.
  • Daily shoppingServing daily shopping and strolling needs, it provides functions such as mall navigation, parking lot inquiry, and trunk opening, improving the convenience of shopping and travel.
  • Picking up and dropping off school childrenAddressing pain points in picking up and dropping off children at school, such as finding temporary parking spots, setting in-car temperature, and providing accurate navigation to school, thus optimizing the pick-up and drop-off process.