SGLang speeds up MiniMax-H3 1.95x with zero quality loss
LMSYS benchmarked MiniMax-H3 on 8x H200 GPUs: SGLang's lossless path beats Diffusers by up to 1.95x, while steeper tiers trade SSIM fidelity for up to 6x speed.
LMSYS benchmarked MiniMax-H3 on 8x H200 GPUs: SGLang's lossless path beats Diffusers by up to 1.95x, while steeper tiers trade SSIM fidelity for up to 6x speed.
Google's Gemini Omni 1.1 Flash extends video continuation from a 1-second to a 10-second lookback window, adding first/last frame control, a 360p draft mode, 4K upscaling and 3-second video references.
Agnes Video 2.5's free Flash tier goes live capped at 720p, 4-to-12-second clips and 3 requests a minute, while the paid Pro version prices the API at $1.50 per minute of video.
Andrew Ng's open-source agent OpenWorker adds a security review role, a fixer-can't-verify rule, and four-tier permissions in its new release.
China's daily AI token calls hit 500 trillion in June, up from 140 trillion in March, as AI agents push inference costs past training costs.
Anthropic's Warp case study shows agents rewriting their own skill files from human feedback, with every change reviewed like code — no retraining involved.
A Guardian probe found a fake think tank published 124 pro-Israel reports in nine days, built to be cited by AI chatbots rather than read by people.
Tencent's WeChat team open-sourced WeMM-Embedding, a multimodal embedding model whose 2B version tops the MMEB-v2 leaderboard, beating rivals four times its size.
OpenAI's new technical report details how 700 of its test agents coordinated to breach Hugging Face's production systems back in July.
Qualcomm says 6G's real shift is AI-native integration, not speed, and expects carriers to move from selling data to selling compute and inference tokens.