In plain words: Computer-made data—fake images, scenes, and simulations—can train deep learning systems for self-driving, robots, biology, and language. The survey maps how it is made and how to fix the biggest catch: models trained on fake data stumble on real ones.
Abstract · Synthetic Data for Deep Learning
Synthetic data is an increasingly popular tool for training deep learning models, especially in computer vision but also in other areas. In this work, we attempt to provide a comprehensive survey of the various directions in the development and application of synthetic data. First, we discuss synthetic datasets for basic computer vision problems, both low-level (e.g., optical flow estimation) and high-level (e.g., semantic segmentation), synthetic environments and datasets for outdoor and urban scenes (autonomous driving), indoor scenes (indoor navigation), aerial navigation, simulation environments for robotics, applications of synthetic data outside computer vision (in neural programming, bioinformatics, NLP, and more); we also survey the work on improving synthetic data development and alternative ways to produce it such as GANs. Second, we discuss in detail the synthetic-to-real domain adaptation problem that inevitably arises in applications of synthetic data, including synthetic-to-real refinement with GAN-based models and domain adaptation at the feature/model level without explicit data transformations. Third, we turn to privacy-related applications of synthetic data and review the work on generating synthetic datasets with differential privacy guarantees. We conclude by highlighting the most promising directions for further work in synthetic data studies.
Sergey I. Nikolenko
arXiv:1909.11512 · cs.LG, cs.CR, cs.CV · submitted Sep 25, 2019
abstract · pdf · html · 156 pages, 24 figures, 719 references
It's pretty cool, because both the branches let you create custom generators. One we tinkered with was using T-Digests to profile large datasets to produce synthetic data, rather than just fake data. Basically we're working on something that lets you take in data and spit out something with identical statistical information inside. One thing you can use this for is stripping out confidential data (in theory at least). Another use case was to expand dataset sizes.