Lesson 8 – Lakehouse architecture

Microsoft Fabric is a unified Data and AI analytics platform that helps organisations ingest, store, process, analyse, and visualise data. One standout feature is the “lakehouse,” a smart way of managing data that combines the best of both data lakes and data warehouses. With a lakehouse, you can keep and analyze various types of data all in one place, using affordable cloud storage and open formats. It comes with handy features like transactions (keeping track of changes), schema enforcement (ensuring data consistency), and support for business intelligence (helping with data analysis for decision-making).

🚀 Learn It in 5 Minutes

Skip the long read.
See the key concepts in one easy-to-follow infographic.

View Infographic →

Microsoft Fabric Lakehouse is built on OneLake, which serves as a unified storage layer. This enables different Fabric experiences such as Data Engineering, Data Science, Data Warehouse, Real-Time Intelligence and Power BI to work with the same data without creating multiple copies.

Source: Microsoft learn

A Microsoft Fabric Lakehouse can be understood through the following key architectural areas:

Ingestion Layer

  • Gathers data from diverse sources.
  • Transforms data into a format suitable for storage and analysis in a lakehouse.
  • Accommodates both batch and streaming data.
  • Handles structured and unstructured data.
  • Sources data from a variety of applications, devices, and sensors.
  • Fabric provides options such as Data Factory Pipelines, Dataflow Gen2, Notebooks, Eventstream, OneLake Shortcuts to bring or reference data in OneLake.

Storage Layer

  • Stores Lakehouse data in OneLake using open formats such as Parquet and Delta Lake.
  • OneLake provides the unified storage foundation for Microsoft Fabric and is built on Azure Data Lake Storage technology.
  • Enables accessibility to the data through a range of tools and languages, including SQL, Python, Scala, R, and more.
  • In Microsoft Fabric, Lakehouse data is stored as Delta Lake tables in OneLake, providing ACID transactions, schema evolution, versioning, and improved reliability.
  • A Fabric Lakehouse provides both a Tables area for managed Delta tables and a Files area for files and folders.

Metadata Layer

  • Tracks and manages various aspects of data in the Lakehouse, including schema, version, lineage, and data quality.
  • Facilitates features such as ACID-compliant transactions, schema enforcement and evolution, data validation, and time travel.
  • An illustration of a metadata layer is the open-source Delta Lake project.
  • Microsoft Fabric also integrates governance, lineage, and discovery capabilities across the platform.

Processing and Query Layer

  • Provides different ways to process and query data stored in the Lakehouse.
  • Apache Spark and Notebooks can be used for large-scale data engineering, transformation, and data science workloads.
  • Data Factory pipelines and Dataflow Gen2 can be used or orchestration and low-code data transformations.
  • Every Fabric Lakehouse includes a SQL analytics endpoint that provides a read oriented T-SQL experience over supported Delta Tables.
  • These different engines can work with the same underlying data stored in OneLake without requiring separate copies for every workload.

Consumption Layer

  • Enables users to consume and analyse data within the lakehouse.
  • Supports various applications and tools, including business intelligence, data visualization, data science, machine learning, and reporting.
  • Incorporates security and governance features such as authentication, authorization, encryption, and auditing.
  • Every Fabric Lakehouse includes a SQL Analytics Endpoint, allowing users to query Lakehouse tables using T-SQL. Power BI semantic models can use Direct Lake to analyse Delta tables stored in OneLake with less reliance on traditional Import-mode data refreshes.
  • Semantic models can be created explicitly over the Lakehouse tables required for reporting.

The lakehouse architecture also allows data to be consolidated and unified for different use cases, such as engineering, data science, machine learning, and business intelligence, in a single system.

Medallion Architecture

In the lakehouse architecture, data undergoes a continuous process of incremental improvement, enrichment, and refinement as it travels through various stages of staging and transformation. This particular approach is commonly denoted as a medallion architecture.

Source: Microsoft learn

The medallion architecture is like a step-by-step process for handling data in a lakehouse. It is a widely adopted architecture pattern supported in Microsoft Fabric where data progressively improves in quality as it moves through different layers. This method makes sure data is trustworthy and goes through checks and improvements before being stored in a way that makes it easy to analyze.

The medallion architecture typically involves three main layers, often referred to as bronze, silver, and gold.

Here’s a breakdown of each layer:

Bronze Layer (Raw Zone)

Description: This is where raw, untouched/unvalidated data is initially stored.

Characteristics:

  • Data in its original form, straight from various sources without much processing. Data here is unvalidated.
  • It preserves the raw state of the original data source.
  • The volume of data in this layer increases gradually over time.
  • Data is a combination of both streaming (continuous flow) and batch (grouped) transactions.

Purpose: Acts as the starting point for data ingestion, holding the unaltered source data.

Silver Layer (Enriched Zone)

Description: In this layer, the data undergoes validation and basic processing.

Characteristics:

  • Validation checks for accuracy, consistency, and completeness are applied. Metadata may be added.
  • Data in this layer can be trusted for downstream analytics.
  • Data may also be cleaned, deduplicated, standardised, and combined with related datasets.

Purpose: Ensures that the data is reliable and conforms to specific quality standards.

Gold Layer (Curated Zone)

Description: The curated, business-ready layer where data is enhanced with business logic, transformations, aggregations and analytical structures.

Characteristics: Optimized for efficient analytics, may include aggregated or calculated data.

Purpose: Represents a high-quality version of the data ready for consumption, often serving as a trusted source for decision-making.

Best Practice: In Microsoft Fabric, Gold layer data is commonly consumed through Power BI Direct Lake or a Fabric Warehouse, depending on reporting and business requirements.

Depending on governance and workload requirements, Bronze, Silver, and Gold can be implemented within a Lakehouse, across separate Lakehouses or workspaces, or by combining Lakehouse and Warehouse capabilities.

In summary, the three layers of the medallion architecture (bronze, silver, and gold) represent different stages in the data processing pipeline, with each stage adding value and improving the quality of the data for various purposes, from raw storage to cleaned, curated, analytics-ready information.