AB
AiBoss
News

Tencent Open Source WeMM-Embedding Multimodal Embedding Model Family

Tencent's WeChat Vision Team released WeMM-Embedding, offering three versions: 2B, 4B, and 9B, which uniformly support text, images, videos, visual documents, and interleaved multimodal input.

Tencent's WeChat Vision Team released the WeMM-Embedding family of multimodal embedding models.Currently, three versions are available: 2B, 4B, and 9B, which can handle text, images, videos, visual documents, and interleaved multimodal inputs.

The model supports Matryoshka dimension clipping and provides examples of Transformers, Sentence Transformers, vLLM, and SGLang. The official documentation also states that audio input is not currently supported.

refer to:Tencent WeMM-Embedding repository