Community Pulse

What people are discussing across the data engineering community.

  • Source Microsoft Fabric Community

    AI-assisted development in Fabric

    This is a vendor announcement, not a community thread: Microsoft’s FabCon Europe 2026 post on the Fabric Community blog. It describes agentic Copilot in notebooks, which turns a developer’s intent into governed, Fabric-native operations, and a data engineering agent in preview, built on the Osmos acquisition. Instead of step-by-step chat, developers describe an outcome and grant access to Fabric assets; the agent plans, executes and validates work over hours or days, creating notebooks, transforming data and updating schemas on Fabric Spark. Microsoft positions it for long-running efforts such as migrations, schema harmonization and data quality remediation. These are preview capabilities as described by Microsoft; production experience is still limited.

    My take

    AI can accelerate boilerplate, exploration and repetitive implementation, but production data engineering still needs human ownership of architecture, security, schemas, performance and validation. Generated notebooks should be reviewed like any other code.

    • AI
    • LLM
    • Microsoft Fabric
    • Fabric Notebook
  • Source Reddit

    CI/CD and Git deployment in Microsoft Fabric

    A team new to Git shared its design: workspaces per layer and environment, Dev workspaces connected to a dev branch, Test and Prod updated only by GitHub Actions using fabric-cicd, variable libraries for configuration, and SQL database projects for the Warehouse (which fabric-cicd did not deploy), with a manual review of the Warehouse diff before it runs. Report developers stay on deployment pipelines in separate workspaces. Replies challenged keeping Bronze in Prod only, since structural differences between environments tend to cause trouble later; suggested separating storage items from engineering items so feature branches can share Dev data; recommended dbt; and argued that branch policies and Git discipline matter more than the architecture. Since then, Microsoft’s August 2026 update has previewed Warehouse CI/CD with DacFx and an API for branched-workspace relationships.

    My take

    Git should be the source of truth for code and deployable artifacts, while environment-specific configuration should remain externalized. Avoid treating a Fabric workspace as the source of truth. CI/CD should promote versioned artifacts into environments rather than depend on manual workspace edits.

    • CI/CD
    • Git
    • Microsoft Fabric
    • DevOps
  • Source Reddit

    Delta Lake interoperability, schema and write behavior

    The question: when promoting Delta schema changes from Dev to Prod, rely on mergeSchema or overwriteSchema in the writer, or alter tables explicitly? The poster found switching mergeSchema on for one release and off again clunky. Answers differed: treat a table like a public API and never drop or retype columns (add a new column instead); keep mergeSchema on but control the columns explicitly before every write, except for Gold tables behind Direct Lake models; or route all writes through shared functions that enforce a YAML table definition. Related threads add the multi-engine side: delta-rs and Polars are fast and light for small workloads, but lag on Delta features (MERGE needed package upgrades in the Fabric runtime, Polars cannot write deletion vectors or V-Order), and one reply stated there are no plans to support Fabric-specific features in single-node libraries.

    My take

    When multiple engines can write the same Delta table, the operational contract matters more than the library choice. Be explicit about schema ownership, supported protocol/table features, evolution rules and writer compatibility. Automatic schema merging is convenient, but it should not replace deliberate schema management in production.

    • Delta Lake
    • Schema Evolution
    • Fabric Lakehouse
    • Python
  • Source Reddit

    Fabric architecture for a small team

    A poster shared a design for a small team: three repositories split by ownership (ingestion, analytics, reports), only Dev and Prod with no staging, Dev synced from main through Git integration, Prod promoted with deployment pipelines, a workspace pair per layer, metadata-driven Bronze, and shortcuts into a Gold workspace where Warehouse versus Lakehouse was still open. Feedback was mixed. Some warned about deployment pipelines and Warehouse deployment and suggested fabric-cicd, dbt or SQL database projects; the poster said deployment pipelines had worked so far. Others advised just starting and refactoring instead of designing everything up front, warned that capacity limits can push small teams toward larger SKUs, and offered different Gold choices depending on SQL versus Python skills and reporting performance.

    My take

    Small teams benefit from fewer moving parts. The goal should not be to copy an enterprise architecture with dozens of workspaces and environments. Separate things where ownership, risk or deployment boundaries justify it; otherwise optimize for simplicity and maintainability.

    • Data Architecture
    • Microsoft Fabric
    • Fabric Workspace
    • Medallion Architecture
  • Source Reddit

    Moving notebooks between Dev, Test and Production

    An engineer coming from Databricks Asset Bundles asked how Fabric handles environments. Replies agree on one workspace per environment but split on promotion. Several prefer the fabric-cicd Python library run from Azure DevOps or GitHub Actions, with Git integration only on per-developer feature workspaces, and describe deployment pipelines as tedious (rules per item, rebuilding the pipeline to change stages, unclear rollback); others say deployment pipelines work well for them, and one reply notes there is no polished equivalent yet. Environment-specific values belong in variable libraries, which one person found laborious to retrofit. A related thread on notebooks gives the practical version: replace the default Lakehouse and hard-coded ABFS paths with variables or paths resolved at run time, and expect a few manual fixes the first time new items reach Test or Prod.

    My take

    Treat environment binding as configuration, not notebook logic. Notebooks should not require manual edits every time they move between Dev, Test and Production. Keep workspace-specific values externalized through variables, configuration or deployment tooling, and make environment promotion repeatable.

    • Fabric Notebook
    • Fabric Workspace
    • CI/CD
    • Git

Question of the Week

Ways to participate

Community is a conversation. These are the ways to take part as they become available.

  • Suggest a topic

    Point me to a discussion or a production problem worth a closer look.

  • Ask a technical question

    Questions from real projects often become articles or Community topics.

  • Contribute to a project

    Help build the open projects: code, docs, samples and tests.

  • Help with translations

    Improve the French and Spanish versions of articles and pages.

  • Report an issue

    Spotted an error, a broken link or an outdated detail? Let me know.

  • Share what you built

    Built something with an article or project from this site? Share the result.

Links will be added here once each channel is set up.

Reading on the topics I am tracking in the community.

Articles

Projects

Talks

Have a technical topic worth discussing?

If a discussion here matches a problem you are working on, or you think my take is missing something, I would like to hear about it.

Suggest a topic LinkedIn (opens in a new tab)