MiniMax正式开源新一代通用视频模型MiniMax H3[1]。这是一个全模态生成系统,可生成最长15秒、最高2K分辨率的视频并带有原生立体声音频[1]。该模型在Artificial Analysis有声视频编辑榜单中以1130分的Elo成绩位列全球第一[1]。
MiniMax H3支持中文、英语、日语、韩语、法语、德语等11种语言[1],输出帧率为24FPS[1],支持21:9、16:9、4:3、1:1、3:4、9:16等多种画面比例[1]。该系统由H3-Context-IR、H3-Base、H3-Regenerate-2K三个模块组成[1]。
在生态适配方面,华为昇腾、摩尔线程、沐曦、海光信息、昆仑芯、天数智芯、壁仞科技、AMD、Intel等16家芯片及平台首日完成支持[1]。此外,Hugging Face、魔搭ModelScope、ComfyUI、RunningHub、fal、vLLM-Omni、SGLang等生态伙伴也加入适配阵营[1]。
MiniMax has officially released MiniMax H3, a new-generation multimodal video generation system that ranks first globally in audio-video editing benchmarks [1]. The model generates videos up to 15 seconds long at a maximum resolution of 2K with native stereo audio [1]. On the Artificial Analysis audio-video editing leaderboard, H3 achieved an Elo score of 1,130, securing the top position [1].
The system supports 11 languages including Chinese, English, Japanese, Korean, French, and German [1], with output frame rates of 24 FPS and compatibility across multiple aspect ratios—21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 [1]. H3's architecture comprises three core modules: H3-Context-IR, H3-Base, and H3-Regenerate-2K [1].
The model attracted rapid ecosystem adoption, with 16 technology partners completing integration support on the first day [1]. These partners include chip manufacturers and platforms such as Huawei Ascend, Moore Threads, Moxin, Haiguang Information, Kunlun Chip, Tianshu Zhixin, Bilixin Technology, AMD, Intel, Hugging Face, ModelScope, ComfyUI, RunningHub, fal, vLLM-Omni, and SGLang [1].