Datamimic 是一个采用 MIT 许可证的开源合成数据生成和数据伪名化工具。1 该工具的社区版(CE)采用 Python 编写,支持 MCP 协议,兼容 PostgreSQL、MySQL、Oracle、MS SQL、SQLite、MongoDB 等多种数据库系统,同时支持 CSV、JSON、XML、XLSX、DbUnit、fixed-width 等多种数据格式。1
Datamimic 的核心特性是确定性数据生成:在相同的引擎版本、相同的数据模型和相同的种子条件下,每次运行在每台机器上都能生成字节完全相同的输出。1 平台的企业版(EE)提供了额外功能,包括 PII 扫描、多系统执行、行业专用消息模板(EDIFACT/SWIFT MT/HL7)、基于角色的访问控制、审计日志和任务调度等。1 该工具已在欧洲受管制的银行环境中部署使用,用于 Oracle、MongoDB 和 Kafka 管道中的确定性测试数据生成。1
Datamimic is an MIT-licensed open-source platform for generating synthetic data and anonymizing personally identifiable information, designed to integrate with AI agents across multiple database systems.1 The tool's Community Edition, written in Python and supporting the MCP protocol, provides deterministic data generation capabilities.1 Its defining feature ensures that identical engine versions, data models, and random seeds produce byte-for-byte identical outputs across repeated runs and different machines.1
The platform supports a wide range of database systems including PostgreSQL, MySQL, Oracle, MS SQL, SQLite, MongoDB, and file formats such as CSV, JSON, XML, and XLSX, along with specialized formats like DbUnit and fixed-width files.1 Beyond the open-source Community Edition, Datamimic offers an Enterprise Edition that adds advanced governance features including PII scanning, multi-system execution, industry-specific message templates for formats such as EDIFACT, SWIFT MT, and HL7, role-based access controls, audit logging, and task scheduling.1 The tool has already been deployed in regulated banking environments across Europe for generating deterministic test data in Oracle, MongoDB, and Kafka pipeline systems.1
评论
还没有评论,欢迎留下第一条。