News
Tencent Open Source General Multimodal Embedding Model: WeMM-Embedding
Tencent's WeChat Vision Team has open-sourced the general-purpose multimodal embedding model WeMM-Embedding, offering 2B, 4B, and 9B specifications. The model supports unified encoding of text, images, videos, visual documents, and interleaved multimodal inputs, and ranks first on the MMEB-v2 benchmark leaderboard. Based on the Qwen3.5 architecture, the model employs a two-stage training strategy and supports Matryoshka's flexible dimensional output.