Databricks vs Fabric: Which Should You Choose?
Clients ask me this weekly: "Should I use Databricks or Fabric?" My answer: It depends. Let me explain.
63 articles
Clients ask me this weekly: "Should I use Databricks or Fabric?" My answer: It depends. Let me explain.
Databricks unifies data, ML, and AI on a single lakehouse platform.
1. Use shared storage : ADLS as the common data layer 2. Standardize on Delta Lake : Compatible format for both 3. Optimize for readers : Both platforms…
Databricks Foundation Model APIs provide enterprise-ready access to state-of-the-art LLMs. This guide covers using these APIs for building AI applications.
Serverless model serving eliminates infrastructure management while providing cost-effective, scalable ML inference. This guide covers implementing…
Databricks Model Serving continues to evolve with new features for deploying and scaling ML models. This guide covers the latest updates and best practices.
Online feature serving enables real-time ML inference by providing low-latency access to features. This guide covers setting up and using Databricks online…
Feature engineering transforms raw data into meaningful inputs for machine learning models. Databricks provides powerful tools for building and managing…
Unity Catalog extends governance to machine learning assets. Manage models, features, and experiments with the same rigor as your data.
Taking Vector Search to production requires careful consideration of performance, reliability, and maintenance. This guide covers production-ready patterns.
Databricks Vector Search enables semantic similarity search over your lakehouse data. Build RAG applications, recommendation systems, and intelligent search…
The aiquery() function enables custom LLM interactions directly in SQL. Unlike specialized functions, it allows you to craft any prompt and get intelligent…
Databricks SQL AI functions bring large language model capabilities directly into SQL queries. Process text, generate insights, and enrich data without…
Databricks dashboards (including Lakeview) provide powerful visualization capabilities integrated with the lakehouse. This guide covers building effective…
Genie Spaces enable non-technical users to explore data using natural language. This guide covers setting up and optimizing Genie Spaces for your organization.
Databricks AI/BI combines the power of the lakehouse with AI-driven analytics. This guide explores how to leverage AI/BI for intelligent data analysis.
Azure Databricks continues to evolve with powerful AI/BI features. April 2024 brings significant updates that blur the line between data engineering and…
Error: {type(error).name}: {str(error)} response = openai.ChatCompletion.create( engine="gpt-4", messages=[{"role": "user", "content": prompt}] ) return…
Azure Databricks with Azure OpenAI creates intelligent data platforms where natural language becomes the interface for data engineering tasks.
Unity Catalog provides a unified governance layer across all your Databricks workspaces. Metastore: The top-level container for all metadata. You create one…
For large datasets, use streaming to incrementally update: Materialized views in Delta Live Tables provide: Precomputed results for fast query performance…
DLT streaming tables use append mode by default. For aggregations: DLT handles checkpointing automatically: Streaming tables in Delta Live Tables provide:…
DLT's applychanges function processes these into clean, current-state tables.
The expression must evaluate to true for valid records.
Traditional ETL: Delta Live Tables: DLT provides built in data quality enforcement: Handle CDC feeds with APPLY CHANGES: Delta Live Tables transforms data…
Key features: Jobs : Multi task workflows with dependencies Triggers : Schedule, file arrival, or API based Compute : Job clusters or serverless Monitoring…
Databricks AutoML automatically: Prepares and preprocesses data Engineers features Selects algorithms Tunes hyperparameters Evaluates and compares models…
The registry organizes models with: Registered Models : Named model artifacts Model Versions : Specific iterations of a model Stages : Lifecycle states…
Key features: Serverless : No infrastructure management Auto scaling : Handles variable traffic automatically Low latency : Sub second response times…
The Feature Store solves these problems.
The ML platform includes: Feature Store : Centralized feature management AutoML : Automated model training and selection MLflow : Experiment tracking and…
Photon is Databricks' native vectorised query execution engine written in C++—a replacement for the JVM-based Apache Spark execution engine for SQL and…
Databricks SQL offers: SQL Endpoints : Serverless compute for SQL queries Query Editor : Web based SQL IDE Dashboards : Built in visualization and reporting…
The protocol is open source, meaning recipients don't need Databricks to access shared data.
A complete governance framework addresses: Access Control : Who can access what data Data Quality : Ensuring data accuracy and completeness Data Lineage :…
Unity Catalog solves these problems with a unified approach.
The lakehouse architecture proved its value in 2021. Organizations are consolidating their data warehouses and data lakes into unified lakehouses, reducing…
Databricks Repos in production use means the code running in your prod workspace is explicitly linked to a specific Git commit—not "whatever notebooks…
Databricks Git integration (Repos) connects the workspace directly to GitHub, GitLab, Azure DevOps, or Bitbucket, replacing the manual pattern of…
Databricks notebook workflows—chaining notebooks together using dbutils.notebook.run()—are the stepping stone between "notebooks as scripts" and proper…
The Databricks REST API exposes every workspace capability that the UI and CLI provide—and then some—making it the integration point for external…
The Databricks CLI is the command-line interface for Databricks workspace operations—running jobs, managing clusters, deploying libraries, uploading…
The job cluster versus all-purpose cluster decision in Databricks is primarily a cost decision: all-purpose clusters stay running between tasks (you pay for…
Databricks cluster policies are the governance layer that prevents the classic enterprise data platform problem: data scientists spinning up 64-node GPU…
Azure Databricks cluster configuration is where cost and performance trade-offs become very concrete: the wrong cluster type for a workload is either money…
Databricks SQL Analytics provides: Native SQL interface to query Delta Lake tables SQL Endpoints with auto scaling compute Built in visualization and…
The MERGE statement combines INSERT, UPDATE, and DELETE operations in a single atomic transaction: Handle complex business logic with conditional updates:…
MLflow became the experiment tracking tool I recommend to every ML team regardless of their cloud platform choice. The core value proposition is simple…
Notebooks support multiple languages in a single document: Python PySpark and pandas Scala Native Spark SQL Spark SQL R SparkR and local R Markdown…
The appeal of Spark Structured Streaming is that you write it almost identically to a batch Spark job. Same DataFrame API, same transformations, same Spark…
Delta Lake time travel is the feature that makes "oops" survivable. An engineer runs a DELETE with a typo in the WHERE clause. A batch job writes corrupt…
Databricks workspace governance was the problem nobody thought about until there were thirty workspaces, fifteen clusters running overnight, and six teams…
MLflow is the closest thing to a standard the ML tooling space has right now—experiment tracking, run metadata, model registry, and a deployment abstraction…
Delta Lake has gone from "interesting open source project" to the default storage layer for data lakes inside about eighteen months. ACID, schema…
SQL Analytics is Databricks' bid for the BI persona who never wanted to learn PySpark. A T-SQL endpoint over Delta Lake, a query editor that looks like…
Artificial Intelligence (AI) and Machine Learning (ML) are trending topics right now. In 2021, there are countless of ways to have a form of "AI" in your…
Spinning up a Databricks workspace is the easy part. Making it production-ready is where I see the most avoidable rework. Cluster policies, AAD passthrough…
Data lakes have a long-standing reputation problem: cheap to fill, painful to trust. Schema drift, half-written files from failed jobs, "is this row a…
Delta Lake: the foundation of the modern lakehouse.
Notebooks are great until you need them to run on Tuesday at 2am, retry on failure, and chain into the next step. That's where Databricks Jobs come in.…
Incrementally process files as they arrive. Structured Streaming makes real-time processing accessible with familiar DataFrame semantics.
MLflow makes ML experiments reproducible and models traceable.
Anyone who's run a "data lake" for any length of time has hit the same wall: parquet files everywhere, no transactional guarantees, partial-write disasters…