Introducing the Synthetic Data SDK - Privacy Preserving Synthetic Data for AI/ML
Privacy regulations are tightening globally, making it increasingly challenging for organizations to access and share data while ensuring compliance.
AI-generated synthetic data is gaining traction as a privacy-safe solution for data access and sharing. This data is created from original datasets, maintaining privacy without compromising utility.
MOSTLY AI has recently released an efficient and flexible Synthetic Data SDK under a fully permissive Apache v2 license, empowering anyone to generate high-quality synthetic data with top-tier performance. Powered by the TabularARGN model architecture, the SDK achieves training times 10x to 100x faster than existing models, while acchieving a SOTA fidelity-privacy balance.
In this Session, we'll cover the fundamental concepts of synthetic data and demonstrate how easy it is to generate synthetic data directly from a Jupyter Notebook using the Synthetic Data SDK. Specifically, we will go through
- Installing the Synthetic Data SDK
- Loading original data into the SDK and locally creating a Generator
- Using a Generator to create different versions of synthetic data
- Uploading a Generator to the MOSTLY AI Platform and sharing it with the world
This will be a hands-on session - so come with your laptop and ideally a dataset that you'd like to synthesize!
This session took place in track Data Handling & Engineering and was classified suitable for intermediate python by the speaker.