Glossary
Data Pipeline
A data pipeline is the automated path data takes from a source system to wherever it is used, including the scheduling, transformation, error handling and recovery that keep it running unattended.
Glossary
A data pipeline is the automated path data takes from a source system to wherever it is used, including the scheduling, transformation, error handling and recovery that keep it running unattended.
A script moves data once. A pipeline handles a source being unavailable, a run overlapping the previous one, a partial failure halfway through, and the need to re-process yesterday after a bug is fixed.

When a transformation turns out to be wrong, the correction has to be applied to historical data as well as new, and a pipeline that can only move today's records makes that a manual project.
Pipelines break when source systems change, credentials rotate and schemas drift. Without a named owner and an alert that reaches them, the first sign of failure is usually a figure someone did not believe.
A script moves data once. A pipeline handles scheduling, overlapping runs, partial failures, alerting and re-processing history after a fix, which is most of the real work.
Changes in the source system: a renamed field, a changed format, a new null. They are silent by default, which is why validation on ingest matters more than clever transformation.
A data pipeline is the automated path data takes from a source system to wherever it is used, including the scheduling, transformation, error handling and recovery that keep it running unattended.