MiniMax, H3 영상 모델 개방 예고

MiniMax는 7월 31일 H3를 발표했다. 상용 API 가격과 모델 가중치 공개 계획이 동시에 나온 점이 핵심이다.

H3는 텍스트, 이미지, 영상, 오디오를 하나의 컨텍스트에 넣어 이해한다. 출력은 최대 2K, 최장 15초이며 네이티브 스테레오 사운드를 포함한다.

오픈 웨이트는 아직 약속 단계

“We plan to open up the model weights in the coming days, subject to applicable laws and regulations.”

Developers still need the license, commercial-use terms, hardware guidance and a technical report before treating H3 as a deployable open model.

가격이 제작 방식을 바꾼다

Shanghai Securities News, The Beijing News and 21st Century Business Herald reported H3 at RMB 0.8 per second for 2K generation, about one-third of comparable flagship video models. A 15-second 2K clip is roughly RMB 12 before platform packaging.

기술 베팅은 통합 컨텍스트

MiniMax describes H3 through Contextual Omni Representation, H3-VAE, H3-Omni Transformer and In-Context Regeneration. The model handles video, audio, images and text as one creative context.

Most source material requires about 100,000 tokens of inference before being distilled to roughly 4,000 tokens on average. H3-VAE provides a 4x gain in effective sequence length, and workload separation lifted training throughput by nearly 30%.

The next checks are whether weights appear, whether the license works for companies, and whether the community can reproduce usable text, motion and audio control on non-official hardware.

출처: MiniMax 공식 연구 블로그, 상하이증권보 중국증권망, 신경보, CocoLoop, 21세기경제보도; MiniMax 공식 사이트로 2K/15초 출력, 네이티브 스테레오, 오픈 웨이트 계획, 초당 0.8위안 가격, 768p 가격 비교, 학습 처리량 약 30% 향상을 확인.