7 月 31 日,MiniMax 發布 H3。官方一邊給出 API 商業價格,一邊承諾開放模型權重,讓影片生成競爭從效果展示轉向可反覆試錯的成本。
H3 支援文字、圖片、影片、音訊一起輸入,輸出最高 2K 解析度、最長 15 秒,並帶原生立體聲。
開放承諾先到,權重還要等
“We plan to open up the model weights in the coming days, subject to applicable laws and regulations.”
Developers still need the license, commercial-use terms, hardware guidance and a technical report before treating H3 as a deployable open model.
0.8 元/秒,把成本擺上桌
Shanghai Securities News, The Beijing News and 21st Century Business Herald reported H3 at RMB 0.8 per second for 2K generation, about one-third of comparable flagship video models. A 15-second 2K clip is roughly RMB 12 before platform packaging.
技術路線押注統一上下文
MiniMax describes H3 through Contextual Omni Representation, H3-VAE, H3-Omni Transformer and In-Context Regeneration. The model handles video, audio, images and text as one creative context.
Most source material requires about 100,000 tokens of inference before being distilled to roughly 4,000 tokens on average. H3-VAE provides a 4x gain in effective sequence length, and workload separation lifted training throughput by nearly 30%.
The next checks are whether weights appear, whether the license works for companies, and whether the community can reproduce usable text, motion and audio control on non-official hardware.
參考來源:MiniMax 官方研究博客、上證報中國證券網、新京報、CocoLoop、21 財經;MiniMax 官網核驗 2K/15 秒/原生立體聲、開放權重承諾、0.8 元/秒價格口徑、768p 價格比較和訓練吞吐提升近 30%。