商汤科技开源了轻量级多模态模型SenseNova U1.5-Lite-Preview,采用8B-MoT的规模集成了图像识别、4K极清生成与精细化编辑能力 [1]。该模型搭载原生统一多模态架构,能够执行复杂提示词、进行参考图创作和图像编辑操作 [1],可将手绘草图转换为商业级海报 [1]。
在性能表现上,该模型在Qwen-Image-Bench测试中得分由上一代U1的47.14分提升至55.20分 [1];在GEdit-Bench图像编辑基准测试中,英文项得分8.17分、中文项得分8.05分 [1]。该模型已在GitHub、Hugging Face和魔搭社区等全球开源平台发布 [1]。
SenseTime has released SenseNova U1.5-Lite-Preview, a lightweight multimodal artificial intelligence model designed to combine image recognition, 4K high-resolution generation, and precise editing capabilities in a compact 8B-MoT framework [1]. The model employs a native unified multimodal architecture that enables execution of complex prompts, creation based on reference images, and fine-grained image editing, allowing hand-drawn sketches to be transformed into professional-grade commercial posters [1].
The newly released model demonstrates significant performance improvements across multiple benchmarks [1]. On the Qwen-Image-Bench assessment, the model's score increased from 47.14 points achieved by its predecessor U1 to 55.20 points [1]. In GEdit-Bench testing, the model scored 8.17 points on English-language items and 8.05 points on Chinese-language items [1]. The model is now available globally through multiple platforms, including GitHub, Hugging Face, and the ModelScope community [1].