Liquid AI Cuts Decoding Latency 3.18x With a 300M Draft Model
Liquid AI's LFM2.5-DSpark draft models bring speculative decoding to three on-device models, cutting decoding latency up to 3.18x with identical outputs.
11 verified stories covering AI inference, product updates and industry developments.
Liquid AI's LFM2.5-DSpark draft models bring speculative decoding to three on-device models, cutting decoding latency up to 3.18x with identical outputs.
Tencent Hunyuan released AngelSpec, an open-source framework for training speculative-decoding drafters and reducing large-model inference cost.
DeepSeek and Peking University open sourced DSpark, an inference framework that raises generation speed by 60% to 85% and can lift single-GPU throughput as much as 6.6 times.
Xiaomi MiMo claims 1,000 tokens per second on standard GPUs. The story explains the announcement, the strategic context and the practical risks for the people or companies affected.
Tensormesh discussed $20 million from AMD Ventures, NVentures, CoreWeave, and others to commercialize its KV cache platform that reduces redundant GPU computation in AI inference.
MiniMax has unveiled its M1 reasoning model series, targeting OpenAI's o-series and DeepSeek's R1 with strong performance in math, code, and logic tasks, while its M2.5 iteration scores 80.2% on SWE-bench Verified.
A comprehensive comparison of major AI reasoning models from OpenAI, DeepSeek, Google, and Anthropic, examining their approaches, trade-offs between depth and speed, and the open-source vs. closed-source landscape.
OpenAI's GPT-5.2 scores 72% on ARC-AGI-1 and 18% on ARC-AGI-2, highlighting the gap between current LLM performance and true general intelligence.
DeepSeek released the R1 reasoning model as open source in January 2026, achieving scores comparable to OpenAI's o1 on math, code, and logic benchmarks while its distilled 32B variant outperformed o1-mini on several tests.
Google has released Gemini 2.5 Pro, touted as its most intelligent model yet, with top scores in math, science, code, and multimodal benchmarks, though its deep reasoning leads to slower response times.
Anthropic released Opus 4.6 on February 5, featuring Adaptive Thinking that dynamically adjusts reasoning depth based on question complexity, a 100K-token context window, and a Compaction API for theoretically unlimited conversations.