AB
AiBoss
News

Proximal releases FrontierSWE v2, expanding to 34 long-term coding tasks.

Proximal has released FrontierSWE v2, a long-term software engineering benchmark, expanding the number of tasks to 34, including 21 new challenges; each task allows for a maximum runtime budget of 20 hours, used to evaluate real-world engineering work over longer periods.

Proximal releases FrontierSWE v2, extending long-term software engineering benchmarks to 34 tasks.

The new version includes 21 new challenges and increases the runtime budget for a single task to a maximum of 20 hours, covering work that more closely resembles a real engineering environment, such as cross-file modification, testing, building, and long-chain debugging.

The official page also provides scores for different models on this benchmark. These results depend on the FrontierSWE v2 task set, execution environment, and scoring method, and are suitable for comparisons within the same caliber. They should not be directly equated with a comprehensive ranking of capabilities across all software development scenarios.

refer to:FrontierSWE official statement