Glossary · System coordination, integration and orchestration
Data orchestration
Also known as: data pipeline orchestration, workflow orchestration (data)
German: Datenorchestrierung
In data engineering, data orchestration is the automated coordination of data movement and processing steps across systems, including scheduling, dependency handling, retries and monitoring, so that data arrives complete and in the right order where it is needed.
- System integration
In one sentence
Data orchestration coordinates data movement and processing across systems, with scheduling, dependencies, retries and monitoring.
Example
A nightly orchestration job extracts machine data from the historian, joins it with MES orders, computes OEE and loads the results into the reporting database.
How it applies
- Engineering: Orchestration tools describe pipelines as dependency graphs. Each step should be restartable and its inputs and outputs traceable, so a failed run can be resumed without duplicating data.
- Operation: Monitor completeness and timeliness, not only whether jobs ran. A pipeline that succeeds on empty input produces empty reports without error.
- Documentation: Document data sources, schedules, owners and downstream consumers of each pipeline. When data feeds AI models or regulatory reports, record data lineage so results can be traced back.
Data orchestration vs. data integration
Data integration makes data from different sources usable together, for example through mappings and a Canonical data model. Data orchestration schedules and supervises the steps that perform integration and processing.