MiniMax chuẩn bị mở mô hình video H3

MiniMax công bố H3 ngày 31 tháng 7. Công ty vừa đưa ra giá API thương mại, vừa nói sẽ mở trọng số mô hình.

H3 nhận văn bản, hình ảnh, video và âm thanh trong cùng một ngữ cảnh. Mô hình tạo tối đa 2K, 15 giây, có stereo gốc.

Cam kết mở chưa phải tệp tải về

“We plan to open up the model weights in the coming days, subject to applicable laws and regulations.”

Developers still need the license, commercial-use terms, hardware guidance and a technical report before treating H3 as a deployable open model.

Giá làm câu chuyện cụ thể hơn

Shanghai Securities News, The Beijing News and 21st Century Business Herald reported H3 at RMB 0.8 per second for 2K generation, about one-third of comparable flagship video models. A 15-second 2K clip is roughly RMB 12 before platform packaging.

Đặt cược vào ngữ cảnh hợp nhất

MiniMax describes H3 through Contextual Omni Representation, H3-VAE, H3-Omni Transformer and In-Context Regeneration. The model handles video, audio, images and text as one creative context.

Most source material requires about 100,000 tokens of inference before being distilled to roughly 4,000 tokens on average. H3-VAE provides a 4x gain in effective sequence length, and workload separation lifted training throughput by nearly 30%.

The next checks are whether weights appear, whether the license works for companies, and whether the community can reproduce usable text, motion and audio control on non-official hardware.

Nguồn: blog nghiên cứu chính thức của MiniMax, Shanghai Securities News, The Beijing News, CocoLoop, 21st Century Business Herald; trang MiniMax dùng để kiểm chứng 2K/15 giây, stereo gốc, cam kết mở trọng số, giá 0,8 nhân dân tệ mỗi giây, so sánh 768p và mức tăng throughput huấn luyện gần 30%.