Building Data Pipelines with Microsoft Fabric Data Factory
Data Factory pipelines use a visual designer with activities that can be chained together.
37 articles
Data Factory pipelines use a visual designer with activities that can be chained together.
AI-enhanced pipelines handle complex transformations and quality issues automatically.
AI-powered pipelines transform raw data into intelligent, enriched datasets. Design for both batch and streaming scenarios.
AI transforms data pipelines from rigid rule-based systems to adaptive, intelligent processes. Start with high-value, error-prone steps and expand from there.
CDC captures row level changes (inserts, updates, deletes) from source systems: The simplest option for supported sources: For more control over the CDC…
Automation is key to maintaining reliable data pipelines. This guide covers scheduling, orchestration, and automation patterns in Microsoft Fabric.
A Fabric pipeline isn't just a sequence of activities — it's an orchestration layer with conditional branching, parameter passing, loops, and error handling…
A month into production-style Fabric testing and Dataflow Gen2 has become my go-to recommendation for teams without Spark expertise who need repeatable data…
The Copy Activity in Fabric Data Factory is the same copy engine as Azure Data Factory — the underlying data movement service optimised for throughput with…
Tomorrow we'll dive deeper into Copy Activity patterns. Data Factory in Fabric Data Pipeline Documentation Pipeline Expressions
Dataflows Gen2 provide an accessible way to build data transformations. Tomorrow, I will cover Mirroring in Fabric.
Fabric Pipelines enable robust data orchestration with full control flow capabilities. Tomorrow, I will cover Dataflows Gen2.
Data Pipelines in Fabric are similar to Azure Data Factory pipelines but with deeper Fabric integration.
Lift and shift existing SSIS packages with minimal changes: Convert SSIS logic to native Azure Data Factory: SSIS Component Data Flow Equivalent OLE DB…
Traditional ETL: Delta Live Tables: DLT provides built in data quality enforcement: Handle CDC feeds with APPLY CHANGES: Delta Live Tables transforms data…
Think of it as the best of both worlds: visual design for maintainability, Spark for scale.
Data Flows Gen2 democratize data transformation, enabling both developers and data analysts to build scalable ETL solutions.
The COPY command replaced PolyBase as the recommended data loading mechanism for Synapse Dedicated SQL Pool because it's simpler to use and performs at the…
PolyBase is the T-SQL feature that makes Synapse Dedicated SQL Pool a data virtualisation engine, not just a data warehouse—it allows you to define external…
The Stored Procedure Activity bridges the gap between data movement and data transformation, enabling complex database logic to be orchestrated as part of…
The ForEach Activity is the key to building scalable, efficient data pipelines that can process thousands of items in parallel while maintaining control and…
The Lookup Activity is the foundation for building intelligent, data-driven pipelines that adapt their behavior based on configuration and runtime data.
Dynamic content in ADF is the expression system that makes pipelines adaptive rather than static—generating file paths from date parameters, constructing…
Parameterized pipelines dramatically reduce maintenance overhead and enable rapid deployment of new data integration scenarios without code changes.
The Copy Activity in Azure Data Factory is the activity I configure more than any other—it's the data movement primitive that connects over 90 source and…
The MERGE statement combines INSERT, UPDATE, and DELETE operations in a single atomic transaction: Handle complex business logic with conditional updates:…
The Synapse Pipelines vs. ADF question comes up in almost every Synapse project I start. The honest answer depends on where the rest of your data platform…
ADF trigger design is where data pipeline reliability is won or lost. Schedule triggers are obvious—run at midnight, run every hour—but the interesting…
Indexers handle the ETL process for search: Requires a column that increases monotonically: Enable change tracking on your database: Source JSON: Source…
Data flows in Synapse run on Spark clusters and offer: Visual drag and drop interface 80+ built in transformations Schema drift handling Data preview and…
Mapping Data Flows are the feature that ended most of my "should I use PySpark or Spark SQL for this?" debates. Visual, yes—but reviewable in JSON and…
Dataflows are collections of entities (tables) that are created and managed in workspaces in the Power BI service. They allow you to: Connect to various…
Delta Lake has gone from "interesting open source project" to the default storage layer for data lakes inside about eighteen months. ACID, schema…
I have a love-hate relationship with visual ETL tools. Drag-and-drop is wonderful right up until you need to diff two pipelines in source control. ADF Data…
A Data Factory pipeline without a trigger is a stored procedure that nobody calls. Triggers are the boring half of orchestration that decides whether your…
Dataflows bridge the gap between self-service BI and governed enterprise data.
A surprising number of "modernise my reporting" engagements I see start the same way: data scattered across an on-prem SQL Server, a SaaS CRM, three Excel…