News
Xiaohongshu released and open-sourced its end-to-end document recognition model: FireRed-OCR
The Xiaohongshu team released and open-sourced the end-to-end document recognition model FireRed-OCR, based on the Qwen3-VL architecture. It pioneered a "three-stage progressive optimization" strategy and a "geometric + semantic" data factory to solve the "structural illusion" problem encountered by general VLM models when processing complex documents. The model achieved state-of-the-art (SOTA) performance in the authoritative benchmark OmniDocBench v1.5, with a comprehensive score of 92.9%, outperforming models such as Gemini-3.0 Pro.