AB
AiBoss
News

MiniMax Open Source New Benchmark Set: Defining Production-Grade Standards for Coding Agents

MiniMax has open-sourced OctoCodingBench, a next-generation Coding Agent evaluation set, which for the first time shifts the evaluation focus from "correct results" to "compliance with process standards." The evaluation set uses two metrics—Check-level accuracy and Instance-level success rate—to systematically assess the AI programming assistant's ability to adhere to process constraints such as naming conventions, security rules, and team collaboration standards.