Glossary · AI engineering and governance
Synthetic data
Also known as: Artificial data, Simulated data
German: Synthetische Daten
In machine learning, synthetic data is artificially generated data, created by simulation, rendering, rule-based generation or generative models, that imitates properties of real data and is used to train, test or validate models when real data is scarce, expensive, confidential or lacks rare cases.
- Industrial AI
- AI
In one sentence
Synthetic data is artificially generated data that imitates real data and is used to train or test models when real data is scarce.
Example
To train a defect detection model, engineers render thousands of images of cracks on CAD models of the part under varied lighting and camera angles.
How it applies
- Engineering: Synthetic data helps cover rare faults, dangerous situations and new product variants before real data exists. Simulation models and digital twins are common sources.
- Validation: The gap between synthetic and real data can mislead a model. Validate models on real data before release, and do not rely on synthetic test data alone (Model validation).
- Documentation: Document how synthetic data was generated, which share of the training data it forms and which real-world conditions it may not represent.
Synthetic data vs. data augmentation
Data augmentation modifies real samples, for example by rotating or adding noise to images. Synthetic data is created without a direct real counterpart. Both extend the training set, and both can introduce patterns that do not occur in reality.