MiniMax H3是一个具有开源权重的多模态视频生成模型,今日正式发布并在ComfyUI中获得Day-0原生支持[1]。该模型支持文本、图像、视频和音频等多种输入方式,可生成具有真实立体声的视频输出,最高支持2K分辨率和15秒时长[1]。
通过优化技术的应用,MiniMax H3已实现在消费级GPU上的本地运行。模型内约40%的调制权重可被修剪并替换为查找表,内存占用从全精度的123.6 GB大幅降低至最小变体的42.5 GB,降幅达66%[1]。这项优化使得用户能够在RTX 3060等配置的设备上本地运行2K视频生成功能[1]。使用ComfyUI体验该模型需要更新至0.30.0或更新版本[1]。
MiniMax H3, a multimodal video generation model with open weights, has been released today with native support in ComfyUI [1]. The model accepts diverse inputs including text, images, videos, and audio, enabling multiple generation modes such as text-to-video, image-to-video, and reference-based video creation [1]. Output specifications include up to 2K resolution, a maximum duration of 15 seconds, and native stereo audio [1].
Through optimization techniques, approximately 40% of the model's modulation weights can be pruned and replaced with lookup tables, significantly reducing memory requirements [1]. This optimization brings down peak memory consumption from 123.6 GB in full precision to 42.5 GB in the minimal variant, representing a 66% reduction [1]. As a result, the model can now run locally on consumer-grade GPUs such as the RTX 3060 for 2K video generation [1]. Users need ComfyUI version 0.30.0 or later to access the model's native integration [1].