Designing Better Lakehouse Flows in Fabric: turning messy raw zones into reliable products
I turned implicit processes into explicit operating rules—defining owners, acceptance tests, and lightweight runbooks so teams can move confidently and…
81 articles
I turned implicit processes into explicit operating rules—defining owners, acceptance tests, and lightweight runbooks so teams can move confidently and…
I tightened system boundaries so quality checks trigger earlier, catching regressions before downstream systems consume bad data.
I tightened system boundaries so quality checks trigger earlier, catching regressions before downstream systems consume bad data.
I worked on smoothing the handoff between data engineering and AI teams—standardizing feature contracts, embedding validation, and adding lightweight…
I spent the day reducing cognitive overhead for engineers and analysts—introducing clearer table contracts, simpler failure modes, and concise runbooks that…
I worked on smoothing the handoff between data engineering and AI teams—standardizing feature contracts, embedding validation, and adding lightweight…
I turned implicit processes into explicit operating rules—defining owners, acceptance tests, and lightweight runbooks so teams can move confidently and…
I tightened system boundaries so quality checks trigger earlier, catching regressions before downstream systems consume bad data.
I worked on smoothing the handoff between data engineering and AI teams—standardizing feature contracts, embedding validation, and adding lightweight…
I tightened system boundaries so quality checks trigger earlier, catching regressions before downstream systems consume bad data.
I focused on making delivery decisions auditable and repeatable—documenting intent, success criteria, and rollback paths to reduce tribal knowledge.
I turned implicit processes into explicit operating rules—defining owners, acceptance tests, and lightweight runbooks so teams can move confidently and…
I spent the day reducing cognitive overhead for engineers and analysts—introducing clearer table contracts, simpler failure modes, and concise runbooks that…
I spent the day reducing cognitive overhead for engineers and analysts—introducing clearer table contracts, simpler failure modes, and concise runbooks that…
I spent the day reducing cognitive overhead for engineers and analysts—introducing clearer table contracts, simpler failure modes, and concise runbooks that…
I spent the day reducing cognitive overhead for engineers and analysts—introducing clearer table contracts, simpler failure modes, and concise runbooks that…
I worked on smoothing the handoff between data engineering and AI teams—standardizing feature contracts, embedding validation, and adding lightweight…
I spent the day reducing cognitive overhead for engineers and analysts—introducing clearer table contracts, simpler failure modes, and concise runbooks that…
I spent the day reducing cognitive overhead for engineers and analysts—introducing clearer table contracts, simpler failure modes, and concise runbooks that…
Here's what it is, how it works, and a pattern that's running in production. Traditional analytics is batch-oriented. Data lands in storage, gets processed…
That's a missed opportunity. Kusto Query Language. Developed by Microsoft. Used across Azure Monitor, Azure Data Explorer, Microsoft Sentinel, and now…
When I tell clients "everything in Fabric uses Delta Lake format," the room divides. Data engineers nod. Everyone else says "what?"
It worked, but managing five services with five billing models and five sets of credentials was painful.
I've built data pipelines for 8 years. Every production deployment humbled me. Here's what the tutorials don't teach.
We migrated from Synapse to Fabric. It took twice as long as expected. Here's what I wish someone had told me.
Clients ask me this weekly: "Should I use Databricks or Fabric?" My answer: It depends. Let me explain.
OneLake is brilliant. One storage layer for everything. No more copying data between Synapse, Data Lake, and Power BI. It just works.
After a year of implementing data solutions on Azure, I've compiled the most important lessons that can save you time, money, and headaches. These come from…
Understanding data lineage is essential for production AI systems. Here's what you need to know.
Understanding data governance is essential for production AI systems. Here's what you need to know.
Understanding data lakehouse ai is essential for production AI systems. Here's what you need to know.
AI-powered table extraction handles complex layouts and normalizes data automatically.
A semantic layer enables AI to understand business terminology and generate accurate queries.
The future is bright for data and AI professionals who stay current and adapt to these changes.
Analytics modernization is a journey, not a destination. Start with clear goals, measure progress, and continuously evolve your analytics capabilities.
This unification simplifies the developer experience and provides a clear path from experimentation to production.
Feature engineering transforms raw data into meaningful inputs for machine learning models. Databricks provides powerful tools for building and managing…
Databricks AI/BI combines the power of the lakehouse with AI-driven analytics. This guide explores how to leverage AI/BI for intelligent data analysis.
Auto-generated insights use AI to automatically discover patterns, anomalies, and trends in your data. This guide covers building insight generation systems.
Natural language interfaces for data go beyond simple Q&A to enable complex analytical conversations. This guide explores advanced natural language query…
Microsoft Fabric's REST APIs enable automation and integration with your existing tools. This guide covers the key APIs and common use cases.
Microsoft Fabric's Git integration brings modern version control practices to analytics. This guide covers setting up and using Git with Fabric workspaces.
AI can transform raw data into polished, narrative reports. This guide covers building automated report generation systems that combine data analysis with…
Natural language to SQL (NL2SQL) transforms how users interact with databases. This guide covers implementation strategies, from simple approaches to…
Microsoft Fabric Copilot has received significant upgrades, making AI-assisted data analysis more powerful and contextually aware. Here's what's new and how…
Microsoft Fabric continues to evolve as a unified analytics platform. April 2024 brings significant updates across data engineering, science, and business…
2023 closed with more platform and model announcements than any year prior. My central thesis is simple: treat AI as a product — instrument, govern, and…
From enterprise pilots to platform bets, 2023 was the year AI stopped being optional. These are the key takeaways I share with clients when they ask what's…
Predictions are inherently risky, but after a year working closely with foundation models and platform teams, a few trends feel probable: larger context…
LLM-powered SQL generation democratizes data access. With proper validation and safety measures, it enables anyone to query databases using natural language.
Traditional databases are optimized for exact matches and range queries. Vector search requires finding approximate nearest neighbors in high-dimensional…
Closing out 2022 from Australia on New Year's Eve, the technical predictions for 2023 feel more consequential than they ever have—because November 30…
Enterprise features like private endpoints, managed identity, and content filtering make these models production-ready.
Knowledge Store enables rich analytics and knowledge graph construction from AI-enriched content.
Temporal joins and windows enable sophisticated event correlation, essential for IoT analytics, fraud detection, and user behavior analysis.
Static : Loaded once at job start Periodic : Refreshes at intervals Dynamic : SQL queries for latest data Reference data joins combine streaming data with…
The new visual editor allows building streaming pipelines without SQL: 1. Drag and drop input/output connections 2. Visual transformations and aggregations…
Elastic tables are ideal for: IoT sensor data Application telemetry Audit logs High frequency transactions Time series data Define an elastic table schema:…
Virtual tables provide: Real time access to external data Native Dataverse experience Security role support Power Platform integration Implement a custom…
March 2022 was a productive month for the Azure data and AI ecosystem: the Azure OpenAI Service access expansion brought GPT-3 and Codex to more enterprise…
Power BI in 2022 offers unprecedented flexibility for enterprise analytics with features that scale from simple reports to complex enterprise deployments.
Arc-enabled data services bring cloud-native database capabilities to any environment. Whether you need SQL Server or PostgreSQL, you can run managed…
Azure Synapse Link for Dataverse eliminates the traditional barriers between operational and analytical systems. Business users get insights from live data…
Premium capacities offer: Dedicated compute : No sharing with other tenants Larger datasets : Up to 400GB per dataset Paginated reports : Enterprise…
Dataverse for Teams democratizes data platform capabilities. Teams users can now build real applications with proper data management without needing IT…
Microsoft Ignite Fall 2021 brought a wave of announcements across Azure. While hybrid work and Microsoft Teams dominated the headlines, the data and AI…
Azure offers several Data Box options: Data Box Disk : Up to 40 TB per order (8 TB per disk) Data Box : 100 TB storage capacity Data Box Heavy : Up to 1 PB…
Set up the data sharing infrastructure: Configure datasets to be shared: Invite consumers to access shared data: Set up automated data synchronization: On…
Create and configure an Azure Purview account: Register various data sources for lineage tracking: Set up scans to discover and catalog data: Azure Data…
Premium Gen2 introduces several architectural improvements: Autoscale compute resources based on demand No need for capacity management or monitoring v…
Databricks SQL Analytics provides: Native SQL interface to query Delta Lake tables SQL Endpoints with auto scaling compute Built in visualization and…
The MERGE statement combines INSERT, UPDATE, and DELETE operations in a single atomic transaction: Handle complex business logic with conditional updates:…
First, let us create a new Spark pool with version 3.0 in Azure Synapse: One of the most significant improvements in Spark 3.0 is Adaptive Query Execution.…
For development or small workloads: For production workloads, use AKS with persistent volumes:
TimescaleDB is the time-series database I recommend to teams that are already comfortable with PostgreSQL and don't want to learn a new query language or…
The Cosmos DB integrated cache is the feature I've been waiting for since the first time I watched a read-heavy application burn through its RU budget…
Graph Data Connect solves this by delivering bulk extracts of Microsoft 365 data directly to Azure Data Factory.
Dataflows are collections of entities (tables) that are created and managed in workspaces in the Power BI service. They allow you to: Connect to various…
Flexible Server is the deployment option I now default to when a client asks "which PostgreSQL on Azure should I use?" The original Single Server mode…
"Why is the dataset refresh taking three hours?" is the question that eventually leads every Power BI shop to incremental refresh. The first time I switched…
Right-click table → Incremental refresh Incremental refresh transforms multi-hour refreshes into minutes.