Microsoft Fabric
How to Make Microsoft Fabric Notebooks Faster
A practical workflow for reducing Fabric notebook startup, read, shuffle, transformation and Delta write time without guessing.
On this page
A slow notebook is not one problem. Time can be spent acquiring a Spark session, installing libraries, scanning files, shuffling rows, waiting on skewed tasks, recomputing the same lineage, or writing many small Delta files. The useful question is not “Which Spark setting makes this faster?” but “Where did this run spend its time?”
Problem
A production pipeline calls a notebook that sometimes completes quickly and sometimes takes much longer. The notebook reads a broad history, performs joins and aggregations, and overwrites a Delta target. Before changing capacity or adding cache calls, separate orchestration delay, session startup, Spark execution, and output writing.
Record at least the pipeline run, notebook run, input window, input rows or files, output rows, and duration of major stages. Compare a slow run with a normal run that processed similar data. Different volumes are different experiments.
Why it happens
Spark is lazy: DataFrame transformations build a plan, while actions such as count, display, collection, and writes execute it. Several debugging actions can therefore run the same expensive lineage repeatedly. Wide operations—including joins, groupings, distinct operations, and repartitioning—can move data across executors through a shuffle. A small number of very large partitions can leave most tasks finished while a few stragglers dominate elapsed time.
Input design matters too. Reading an entire table and filtering afterward may prevent useful partition pruning or data skipping. Selecting every column increases I/O and serialization. Thousands of small files add listing and scheduling overhead. On the output side, forced repartitioning and repeated tiny appends can exchange one problem for another.
Fabric adds a session layer. Compatible workloads may reuse high-concurrency sessions; other configurations require a new session. Starter pools can reduce acquisition time when their constraints fit, but they are a best-effort fast path rather than a guarantee. Custom configuration, libraries, environments, and networking can change startup behavior. Microsoft’s Spark monitoring guidance and session detail views help separate those causes.
Practical recommendations
Begin with a timing breakdown. If most time occurs before the first Spark job, investigate session acquisition and libraries. If a stage dominates, inspect its jobs, tasks, input, shuffle, spill, skew, and output. If the write dominates, inspect partition counts, file sizes, write mode, and downstream table design.
Read only the required columns and time range. Put selective filters next to the read and use predicates aligned with the table’s physical design. Confirm in the executed plan that pruning or skipping occurred; source syntax alone does not prove it.
Reduce data before wide operations. Aggregate at the required grain before joining when that preserves correctness. Broadcast only genuinely small reference data after measuring it. Avoid collect() and toPandas() for large results because they move data to the driver. Use limited samples for inspection.
Cache only a DataFrame that is expensive to recompute and reused by multiple actions. Materialize it deliberately, verify storage pressure, and unpersist it when finished. Caching a one-use DataFrame adds work and memory pressure without saving recomputation.
For Delta output, choose append, overwrite, or merge from the data contract. Avoid repartitioning merely to produce a preferred file count: repartition causes a shuffle, while coalesce reduces partitions without a full redistribution but can create imbalance. Evaluate Optimize Write, compaction, V-Order, and Z-Order according to write frequency and consumers. Microsoft notes that V-Order can help cross-engine reads while adding Spark write cost, so it is not a universal notebook speed switch.
Example pattern
The example keeps projection and filtering close to the source, performs one reusable aggregation, and validates output without collecting full data.
from pyspark.sql import functions as F
from pyspark.storagelevel import StorageLevel
run_date = "2026-10-02"
orders = (
spark.read.table("SalesLakehouse.silver.orders")
.select("order_id", "customer_id", "order_date", "status", "amount")
.where(F.col("order_date") == F.lit(run_date))
.where(F.col("status") == F.lit("Completed"))
)
daily_customer = (
orders.groupBy("customer_id")
.agg(
F.countDistinct("order_id").alias("order_count"),
F.sum("amount").alias("revenue"),
)
.persist(StorageLevel.MEMORY_AND_DISK)
)
output_rows = daily_customer.count()
(
daily_customer.write
.format("delta")
.mode("append")
.saveAsTable("SalesLakehouse.gold.daily_customer_sales")
)
print({"run_date": run_date, "output_rows": output_rows})
daily_customer.unpersist()
This is a pattern, not a prescription. If daily_customer is used only by the write, remove persistence and the separate count or replace validation with a metric produced elsewhere. If reruns are possible, a blind append may duplicate the date; the idempotent write design belongs to the pipeline contract.
Use explain("formatted") before and after an important change. Look for filters near the scan, unexpected exchanges, join strategies, and repeated branches. Then confirm with executed Spark jobs because estimates and plans do not show every runtime condition.
Common mistakes
- Changing executor or shuffle settings before identifying the slow stage.
- Calling
count, display, and write on the same uncached expensive lineage. - Reading a complete history for an incremental run.
- Selecting every column “just in case.”
- Broadcasting a lookup based on row count without considering bytes.
- Caching everything and creating eviction or spill pressure.
- Repartitioning every output to one file, which destroys parallelism.
- Treating session reuse as isolation-free; shared compute still needs compatible configuration and workload planning.
- Comparing runs with different input volumes and calling the difference a regression.
What to check next
If startup dominates, inspect session details, environment and library application, starter-pool eligibility, and whether compatible pipeline notebooks should use high concurrency with a session tag. If a Spark stage dominates, examine skew, shuffle read/write, spill, task duration, and the physical plan. If reads dominate, inspect file count, partition pruning, data skipping, and incremental boundaries. If writes dominate, inspect distribution, small files, table properties, and competing consumers.
Change one understood variable, rerun comparable input, and keep correctness controls alongside performance measures. A faster notebook that loses unmatched rows, duplicates a partition, or changes business grain is a failed optimization.