News
DeepSeek Releases Multimodal Modeling Technology Report
DeepSeek released a multimodal large-scale model and published a technical report on GitHub, proposing a "visual primitive-based thinking" framework. This framework elevates spatial markers such as points and bounding boxes to "basic thinking units" for reasoning, enabling the model to possess accurate spatial reference and deduction capabilities, breaking through the bottleneck of traditional chain-like thinking in complex spatial reference tasks. The model has a compact architecture and high visual marker efficiency, and it rivals cutting-edge models such as GPT-5.4 and Claude-Sonnet-4.6 in counting and spatial reasoning benchmark tests.