dealignai 团队于 2026 年 9 月 10 日发布了 DeepSeek-V4.1-Flash 的无审查版本(dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8),通过权重级别消融技术移除了模型的安全防护机制,同时保留了核心能力1。该模型包含 552B 参数,约占 510 GB 存储空间,支持最高 1M 令牌的上下文长度,并具备视觉和工具调用功能1。
在性能表现上,无审查版本在 HarmBench-320 测试中于 effort=max 时达到 100.0% 合规率1。MMLU 基准测试显示,排除伦理相关科目后,该版本相比基础模型性能下降 1.1 个百分点,保持在 3 个百分点的目标范围内1。单流解码速度达到每秒 101 个令牌(无推测),使用 DSpark + cap-accept 优化后可达每秒 113 个令牌,8 路并发聚合吞吐量为每秒 126 个令牌1。
该模型已通过 4×H200 NVLink 硬件验证,每块 GPU 显存占用在 76 GB(启用 Engram 主机表)或 122 GB(不启用)之间1。冷启动时间约为 28 分钟,预热重启约 10 分钟1。dealignai 团队为用户提供了多种部署方案,支持 Transformers、vLLM、SGLang 和 Docker 等推理框架1。
The dealignai team has released an uncensored version of DeepSeek-V4.1-Flash, designated UNCENSORED-FP8, which removes the model's safety guardrails through weight-level ablation techniques while maintaining core capabilities.1 Released on September 10, 2026, the model features 552 billion parameters and supports a context length of up to 1 million tokens, along with vision and tool-calling functionalities.1
Performance testing demonstrates that the uncensored variant achieves a 100.0% compliance rate on HarmBench-320 at maximum effort levels.1 When evaluated on MMLU excluding ethics-related subjects, the model shows only a 1.1 percentage point decline compared to the base model, remaining within the target range.1 Single-stream decoding speed reaches 101 tokens per second without speculation, accelerating to 113 tokens per second when using DSpark with cap-accept optimization, while 8-way concurrent aggregated throughput reaches 126 tokens per second.1
The model has been validated on 4×H200 NVLink hardware, consuming 76 GB of VRAM per GPU with Engram host tables enabled or 122 GB without them.1 Cold startup requires approximately 28 minutes, with warm restarts taking around 10 minutes.1 The team has provided support for multiple inference frameworks including Transformers, vLLM, SGLang, and Docker deployment options.1
评论
还没有评论,欢迎留下第一条。