Tencent’s WeChat Vision team has open-sourced WeMM-Embedding, a family of multimodal embedding models that represents and matches text, images, videos and other content types. The team said the models are already deployed in WeChat Channels, Official Accounts, Moments and e-commerce services.
The release includes 2B, 4B and 9B versions. Tencent’s published evaluation results place the 9B model first among the listed models on both the MMEB-v2 and MMEB-v3 benchmarks.
Tencent has also released the model code, evaluation tools and weights through its public repository and model pages, allowing developers to use the models for multimodal search, retrieval and recommendation applications. [TechNode reporting]