Tencent Releases WeMM-Embedding for Multimodal Retrieval Tencent's WeChat Vision team released WeMM-Embedding, a family of multimodal embedding models in 2B, 4B, and 9B parameter sizes, supporting text, images, videos, visual documents, and interleaved inputs. The models support Matryoshka dimensions for adjustable vector widths, with reported average scores of 77.9, 79.2, and 80.6 on MMEB-v2 for the respective sizes. The repository includes inference examples, serving instructions, and evaluation code. Tencent’s WeChat Vision team has published WeMM-Embedding https://github.com/Tencent/WeMM-Embedding , a family comprising 2B, 4B and 9B embedding models. Each variant supports text, images, videos, visual documents and interleaved multimodal inputs; audio is not supported. For developers, the immediate consequence is a single repository containing model options, inference examples, serving instructions and evaluation code for retrieval across the supported inputs. The models obtain embeddings from the last-layer hidden state at a dedicated