Glossary · IIoT, data and AI
Training data
Also known as: Training set, Training dataset
German: Trainingsdaten
In machine learning, training data is the data set from which a model learns its parameters, typically consisting of input examples and, for supervised learning, the associated labels or target values. It is kept separate from validation and test data, which are used to tune and evaluate the model.
- IIoT
- AI
In one sentence
Training data is the data set a machine learning model learns from, kept separate from the validation and test data used to evaluate it.
Example
A visual inspection model is trained on 12,000 labeled images of good and defective welds collected over three months on two lines.
How it applies
- Engineering: Training data should cover all relevant operating conditions: products, variants, shifts, seasons and fault types. Gaps show up later as poor performance in exactly those conditions.
- Quality: Label quality matters as much as quantity. Define labeling rules, check agreement between labelers and document known errors.
- Governance: For high-risk AI systems under the EU AI Act, data governance requirements apply to training, validation and testing data. Record the origin, selection and preparation of the data (Data lineage).
- Documentation: Describe the training data in the model documentation: sources, time range, size, class distribution and limitations, for example in a Model card.
Training data vs. test data
Training data is used to fit the model. Test data is held back and used only to estimate performance on unseen data. Mixing the two leads to overly optimistic results.