News
Alitongyi launches a new reinforcement learning framework, EAPO
Alibaba's Tongyi Lab has launched a new reinforcement learning framework, EAPO (Evidence-Augmented Policy Optimization), which introduces an "evidence reward" mechanism. This mechanism shifts supervision from the answer itself to the evidence extraction process, addressing the illusion problem of "correct search but wrong answer" in long-text reasoning using large models. The framework, based on the Qwen3-30B model, has demonstrated superior performance on multiple authoritative long-text benchmark tests, outperforming large models such as GPT-OSS and Claude-Sonnet-4 with 120 parameters.