All projects

Microsoft Fabric

Fabric Data Platform Template

A reusable end-to-end Microsoft Fabric reference implementation: ingestion, Bronze, Silver and Gold layers, serving, Git, CI/CD and observability, built with practical engineering patterns.

Status
Planned
Level
Intermediate

Planned project: this page describes the intended design. No code, repository or demo has been published yet.

Project overview

Who this is for

  • Students who want to build a complete Fabric platform step by step, not just isolated demos
  • Data engineers designing or reviewing a Fabric platform for their organisation
  • Architects who need a reference for environments, deployment and operational boundaries

What you will learn

  • How workspaces, Lakehouses, Warehouses, notebooks and pipelines fit together in one platform
  • How to implement Bronze, Silver and Gold layers with clear responsibilities and quality gates
  • How to load data incrementally and keep pipelines idempotent
  • How to choose partitioning and file-size settings for Delta tables
  • How to log runs and debug failures across layers
  • How to work with Git integration and promote changes from dev to test to production

Tech stack

  • Microsoft Fabric
  • Fabric Lakehouse
  • Fabric Warehouse
  • Fabric Notebooks
  • Fabric Data Pipelines
  • SQL analytics endpoint
  • Delta Lake
  • Parquet
  • PySpark
  • Spark SQL
  • Power BI
  • Git
  • GitHub
  • Bicep
On this page
  1. The problem
  2. Architecture
  3. How it works
  4. Implementation walkthrough
  5. Student path
  6. Professional considerations
  7. Common mistakes
  8. Next improvements

The problem

Most Microsoft Fabric examples show one item at a time: a notebook that reads a CSV, a pipeline that copies a table, a report on top of a Lakehouse. Real platforms are harder because the pieces have to work together. Layers need owners and contracts. Loads must be incremental and safe to rerun. Changes must move between environments without manual edits. And someone has to be able to tell, at 7 a.m., which step failed and why.

This project is a single, coherent reference implementation of those concerns. Students can follow it from an empty workspace to a working platform, and professionals can reuse it as a set of patterns for their own environments.

Architecture

  1. Sources

    • Sample operational database
    • CSV and Parquet files
    • REST API
  2. Ingestion

    • Data pipelinesCopy and orchestration
    • NotebooksAPI and file ingestion
  3. Bronze

    • Bronze LakehouseRaw Delta tables + landed files, with batch metadata
  4. Silver

    • Silver LakehouseValidated, deduplicated, conformed entities
  5. Gold

    • Gold WarehouseDimensional model and serving tables
  6. Serving

    • SQL analytics endpoint
    • Semantic model
    • Reports
    • API

Cross-cutting concerns

  • Git integration
  • CI/CD
  • Run logging
  • Monitoring
  • Configuration
  • Workspace security
Target architecture of the template. Every tier has one responsibility; logging, configuration, security and deployment apply to all of them.

The layers follow the pattern described in Medallion Architecture Explained. Each layer is a separate item with its own permissions, so the boundary between raw and curated data is enforced by the platform, not only by naming. Gold is a Warehouse in the template because its consumers expect T-SQL and multi-table transactions. A Gold Lakehouse is documented as an alternative.

How it works

  1. TriggerSchedule or manual run
  2. Ingest batchNew data since the watermark
  3. Bronze appendBatch id, load time, source
  4. Silver MERGEBy business key, with quarantine
  5. Gold refreshOnly affected facts and aggregates
  6. Quality gateRow counts, keys, reconciliation
One scheduled run. Each step records its outcome in the run log, so a failure points to a single step and batch.

A run is driven by one orchestration pipeline with a run identifier. Ingestion lands new source data in Bronze with the run and batch identifiers attached. Silver notebooks read only the new Bronze batches, validate them, quarantine failing records, and merge the rest by business key. Gold refreshes only the facts and aggregates affected by the changed keys. A final quality gate compares row counts and key totals between layers before the run is marked successful.

Configuration such as connection names, workspace ids and source lists lives in parameters and a configuration table, not in notebook code. This is what lets the same code run in dev, test and production.

Implementation walkthrough

The template will be built and published in milestones. Each one ends with something that runs:

  1. Workspaces and Git: dev, test and production workspaces, naming conventions, Git integration on dev.
  2. Bronze ingestion: a pipeline and a notebook landing three sample sources with ingestion metadata.
  3. Silver transformations: validation rules, deduplication, quarantine table and MERGE by business key.
  4. Gold model: a small dimensional model in the Warehouse and a semantic model on top.
  5. Incremental runs: watermarks, idempotent reruns and late-arriving data.
  6. Observability: a run log table, failure alerts and a simple operational report.
  7. Deployment: promotion from dev to test to production, with environment-specific configuration.
  8. Performance: partitioning and file-size choices, compaction and table maintenance at larger volumes.
  1. Feature branchChange in a dev workspace
  2. Pull requestReview and checks
  3. DevGit-connected workspace
  4. TestDeployed, with test data
  5. ProductionDeployed, approved release
Planned deployment flow. Code reaches production only through Git and the deployment process, never by editing production items.

Student path

For students

Prerequisites

  • Access to a Microsoft Fabric capacity or trial, and permission to create workspaces
  • Basic SQL; some Python helps but is not required at the start
  • A GitHub account for the Git milestones

Concepts to understand first

Read Medallion Architecture Explained before milestone 2. The Fabric learning path introduces workspaces, Lakehouses, notebooks and Delta tables in the order this project uses them.

Guided steps

Follow the milestones in order. Each milestone will have its own instructions, a checklist of what should exist when you finish, and a short “how do I know it worked” section.

Exercises

  • Add a fourth source to Bronze without changing the Silver code.
  • Write a validation rule that quarantines orders with a missing customer, and check the quarantine table.
  • Rerun the same batch twice and prove that Silver and Gold do not change.

Try this next

Replace the Gold Warehouse with a Gold Lakehouse and compare the developer experience, then read the planned Lakehouse vs Warehouse material.

Expected outcomes

You will be able to explain what each Fabric item in the platform does, build and rerun a layered pipeline, and use Git to move a change between environments.

Professional considerations

For professionals

Architecture decisions

One Lakehouse per layer keeps permissions and ownership explicit; a single Lakehouse with schemas is simpler for small teams. The template documents both and uses the first. Gold in a Warehouse favours T-SQL consumers; Gold in a Lakehouse favours Spark-heavy teams.

Environments and deployment

Dev, test and production are separate workspaces, ideally on separate capacities for production isolation. Only dev is edited directly. Environment differences live in configuration, so deployment never requires changing code.

Scale

Bronze tables are partitioned by load date only when volumes justify it; Silver and Gold start unpartitioned. Incremental processing, compaction and table maintenance are part of the design rather than added after the first slow run.

Security

Workspace roles separate engineers from consumers. Consumers read Gold through the SQL analytics endpoint or the semantic model and never get access to Bronze. Secrets are not stored in notebooks.

Observability and failure modes

Every run writes a row per step to a run log: run id, batch id, rows read and written, duration and outcome. This answers which batch failed, in which layer, and what changed. Failure modes covered include source unavailability, schema drift, duplicate deliveries, partial runs and reruns.

Cost and maintainability

Incremental processing keeps compute proportional to change rather than to data size. Shared notebook utilities for logging, configuration and MERGE keep layer code short and consistent.

Alternatives considered

  • A single all-in-one notebook per source: simple at first, but hard to test and impossible to rerun per layer.
  • Full reloads everywhere: correct and easy, but cost and duration grow with data size.
  • Fewer layers: valid for simple workloads; the template uses three because it is a reference for multi-source platforms.

Common mistakes

  • Building all three layers before the first end-to-end run works with one source.
  • Hard-coding workspace ids or connection names, which breaks every deployment.
  • Merging into Silver without deduplicating the incoming batch first.
  • Letting reports read Silver “temporarily”, which makes it a contract nobody planned.
  • Logging only failures. Successful runs need row counts too, or there is nothing to compare against.

Next improvements

  • Publish the repository and milestone 1.
  • Add a sample dataset large enough to show partitioning and compaction effects.
  • Add a workshop version of the student path with timings and checkpoints.

Tags

  • Microsoft Fabric
  • Medallion Architecture
  • Delta Lake
  • Data Architecture
  • Fabric Lakehouse
  • Fabric Warehouse
  • Incremental Loads
  • CI/CD
  • Logging
  • Partitioning