Fabric Data Factory Notes: designing pipelines for failure, not the happy path
I focused on making delivery decisions auditable and repeatable—documenting intent, success criteria, and rollback paths to reduce tribal knowledge.
32 articles
I focused on making delivery decisions auditable and repeatable—documenting intent, success criteria, and rollback paths to reduce tribal knowledge.
I worked on smoothing the handoff between data engineering and AI teams—standardizing feature contracts, embedding validation, and adding lightweight…
I turned implicit processes into explicit operating rules—defining owners, acceptance tests, and lightweight runbooks so teams can move confidently and…
I focused on making delivery decisions auditable and repeatable—documenting intent, success criteria, and rollback paths to reduce tribal knowledge.
I spent the day reducing cognitive overhead for engineers and analysts—introducing clearer table contracts, simpler failure modes, and concise runbooks that…
I spent the day reducing cognitive overhead for engineers and analysts—introducing clearer table contracts, simpler failure modes, and concise runbooks that…
I worked on smoothing the handoff between data engineering and AI teams—standardizing feature contracts, embedding validation, and adding lightweight…
Data Factory pipelines use a visual designer with activities that can be chained together.
The Copy Activity in Fabric Data Factory is the same copy engine as Azure Data Factory — the underlying data movement service optimised for throughput with…
Tomorrow we'll dive deeper into Copy Activity patterns. Data Factory in Fabric Data Pipeline Documentation Pipeline Expressions
Fabric Pipelines enable robust data orchestration with full control flow capabilities. Tomorrow, I will cover Dataflows Gen2.
Data Pipelines in Fabric are similar to Azure Data Factory pipelines but with deeper Fabric integration.
Lift and shift existing SSIS packages with minimal changes: Convert SSIS logic to native Azure Data Factory: SSIS Component Data Flow Equivalent OLE DB…
Data Flows automatically use the managed VNet when connected to private endpoints.
Think of it as the best of both worlds: visual design for maintainability, Spark for scale.
Data Flows Gen2 democratize data transformation, enabling both developers and data analysts to build scalable ETL solutions.
The Stored Procedure Activity bridges the gap between data movement and data transformation, enabling complex database logic to be orchestrated as part of…
The Web activity in ADF is the HTTP client that lets your pipeline call external REST APIs—triggering a Logic App, posting a Teams notification, calling a…
The Until activity in ADF is the polling loop that keeps running until a condition is true—useful for pipelines that trigger an external process and need to…
The ForEach Activity is the key to building scalable, efficient data pipelines that can process thousands of items in parallel while maintaining control and…
The Lookup Activity is the foundation for building intelligent, data-driven pipelines that adapt their behavior based on configuration and runtime data.
Dynamic content in ADF is the expression system that makes pipelines adaptive rather than static—generating file paths from date parameters, constructing…
Parameterized pipelines dramatically reduce maintenance overhead and enable rapid deployment of new data integration scenarios without code changes.
The Copy Activity in Azure Data Factory is the activity I configure more than any other—it's the data movement primitive that connects over 90 source and…
The Synapse Pipelines vs. ADF question comes up in almost every Synapse project I start. The honest answer depends on where the rest of your data platform…
The Self-Hosted Integration Runtime is the component that makes ADF viable for the majority of real enterprise projects—the ones where the source database…
ADF trigger design is where data pipeline reliability is won or lost. Schedule triggers are obvious—run at midnight, run every hour—but the interesting…
Mapping Data Flows are the feature that ended most of my "should I use PySpark or Spark SQL for this?" debates. Visual, yes—but reviewable in JSON and…
I have a love-hate relationship with visual ETL tools. Drag-and-drop is wonderful right up until you need to diff two pipelines in source control. ADF Data…
Best practice: use Managed Identity where supported. Linked Services are the foundation of Data Factory connectivity.
A Data Factory pipeline without a trigger is a stored procedure that nobody calls. Triggers are the boring half of orchestration that decides whether your…
A surprising number of "modernise my reporting" engagements I see start the same way: data scattered across an on-prem SQL Server, a SaaS CRM, three Excel…