WeChat's Open-Source Multimodal Embedding Model Tops MMEB-v2
Tencent's WeChat team open-sourced WeMM-Embedding, a multimodal embedding model whose 2B version tops the MMEB-v2 leaderboard, beating rivals four times its size.
11 verified stories covering multimodal AI, product updates and industry developments.
Tencent's WeChat team open-sourced WeMM-Embedding, a multimodal embedding model whose 2B version tops the MMEB-v2 leaderboard, beating rivals four times its size.
DeepSeek's new deepseek-v4-flash-vision-exp caps every image at 384 tokens with no extra vision fee, effectively putting a ceiling on GUI agent screenshot costs.
SenseTime open-sourced SenseNova-U1.5-8B-MoT with native 4K generation and precise editing, but its parameter count is reported three different ways.
Thinking Machines released Inkling-Small with open weights, 276B total parameters, 12B active parameters, multimodal input and benchmark scores close to the 975B Inkling.
MiniMax H3 combines 2K, 15-second video, native stereo audio, RMB 0.8-per-second pricing and a coming open-weight release.
ByteDance targets design workflows with Seedream 5.0 Pro is reframed for global technology readers, with the key numbers, product claims and open questions kept intact.
The multimodal version of Qwen3.7-Max can read images and video, but Alibaba is emphasizing autonomous iteration, tool use and self-testing.
Alibaba released Qwen3.7-Plus on June 2, highlighting multimodal input, tool use, code generation and autonomous testing as its next agent bet.
Alibaba's Qwen team has released Qwen3.5-Omni, its first true all-modal model that processes text, images, audio, and video within a single architecture, featuring real-time voice cloning and audio-video vibe coding.
Alibaba's Qwen3.5 open-weight model now covers 201 languages and dialects, introduces native multimodal processing, and signals a broader industry shift from chatbots to AI agents.
Moonshot AI's Kimi K2.5 introduces native multimodal understanding and an Agent Swarm feature that splits complex tasks into parallel sub-tasks handled by multiple agents, significantly boosting speed.