7月31日,国内多家大模型厂商同日推出新产品,在视频生成和代码开发领域密集推进。[1]MiniMax发布首款开源多模态生成模型H3,支持视频生成、编辑和多模态理解能力。[1]该模型单次最长生成15秒音画视频,生成2K视频的价格为每秒0.8元,约为业界同类旗舰产品价格的三分之一。[1]H3在Artificial Analysis带音频的视频编辑榜单中排名第一。[1]MiniMax计划近期开放H3的模型权重。[1]
字节跳动发布视频创作模型Seedance 2.5,将单次生成时长从上代的15秒扩展至30秒。[1]用户可输入最多30张图片、10段视频和10段音频作为参考素材。[1]DeepSeek宣布V4-Flash正式版API开启公测。[1]该版本与预览版采用相同模型架构,性能提升主要源于重新进行的后训练,升级了工具调用、代码开发和Agent执行能力。[1]DeepSeek-V4-Flash-0731在Terminal Bench 2.1上的得分达到82.7。[1]
On July 31, three major domestic AI model companies unveiled new offerings in a concentrated product launch day [1]. MiniMax introduced H3, its first open-source multimodal generation model, while ByteDance released Seedance 2.5, an upgraded video creation tool, and DeepSeek announced the official API public beta of V4-Flash with enhanced capabilities [1].
MiniMax's H3 model supports video generation, editing, and multimodal understanding, with the ability to generate up to 15 seconds of video with synchronized audio in a single pass [1]. The cost for generating 2K resolution video stands at 0.8 yuan per second, approximately one-third of comparable flagship models in the industry [1]. According to rankings on Artificial Analysis for video editing with audio, H3 secured the top position [1]. The company plans to open-source the model weights in the coming weeks, marking MiniMax's first open-source multimodal generation model [1].
ByteDance's Seedance 2.5 doubled the maximum single-generation video length from the previous version's 15 seconds to 30 seconds [1]. The model accepts up to 30 reference images, 10 video clips, and 10 audio segments as input materials [1]. DeepSeek's V4-Flash official version maintains the same model architecture as its preview iteration, with performance improvements driven by retraining [1]. The model achieved a score of 82.7 on Terminal Bench 2.1, demonstrating enhancements in tool calling, code development, and Agent execution capabilities [1].