{"slug": "managing-dependency-hell-setting-up-custom-python-environments-in-fabric", "title": "Managing Dependency Hell: Setting Up Custom Python Environments in Fabric", "summary": "Microsoft Fabric supports two scopes for custom Python environments — Workspace-level and Item-level — to decouple project-specific dependencies from the default Spark runtime, according to a technical guide on managing dependency hell in Fabric. The guide states that installing unmanaged pip dependencies inline at the start of every notebook run wastes Compute Units on an F64 SKU capacity, lengthening startup times and inflating compute costs. Workspace-level environments set a baseline for every Spark artifact in a workspace, while Item-level environments serve shared development or sandbox workspaces.", "body_md": "In the rapidly evolving landscape of modern data engineering, we have developed an almost reflexive habit of reaching for Apache Spark the moment we need to ingest, move, or process data. **Microsoft Fabric** provides a robust, pre-configured default Spark runtime that works beautifully out of the box. For many standard data transformation tasks, this default environment, packed with popular libraries like Pandas, NumPy, and basic machine learning tools, is more than sufficient to get our initial workloads off the ground.\n\nHowever, as we mature our enterprise data architecture and transition from exploratory data analysis to production-grade machine learning, we quickly discover that a “*one-size-fits-all*” environment is inherently limiting. Our advanced data science teams often require highly specific, sometimes cutting-edge versions of libraries like prophet for time-series forecasting or scikit-learn for predictive modeling. Furthermore, we frequently rely on proprietary, custom-built internal packages that encapsulate our company’s unique business logic.\n\nWhen we rely solely on the default Fabric runtime, we expose ourselves to “*dependency hell*”. A platform update might silently upgrade a core library, instantly breaking our strictly validated predictive models. Conversely, our engineers might find themselves unable to leverage a new feature in a recent library release because the default environment is locked to an older version. We must also remain acutely aware of our compute consumption and the financial mechanics underlying our workspaces; when we purchase a Microsoft Fabric capacity (for instance, an F64 SKU), we are provisioning a pool of Compute Units. Forcing our compute engines to manually install unmanaged pip dependencies at the start of every single notebook run via inline commands is an inefficient use of these precious Compute Units. It elongates startup times and inflates our compute costs.\n\nTo achieve true reproducibility, stability, and operational efficiency, we must decouple our project-specific dependencies from the default platform runtime. We need a solution that respects our architectural elegance and our financial constraints. By shifting our strategy toward managed, custom Python environments, we unlock engineering patterns that are fundamentally leaner, significantly faster, and far more cost-effective for our enterprise workloads.\n\nBefore we begin authoring configuration files, we must strategically determine exactly *where* our custom environment should reside within our Microsoft Fabric architecture. Fabric offers us two distinct scopes for environment application: Workspace-level environments and Item-level environments. Understanding the distinction is critical for designing a multi-team operating model that enforces separation of duties across development, test, and production workspaces.\n\n**Workspace-Level Environments**\n\nWhen we configure an environment at the **Workspace level**, we are establishing a new baseline for every single Spark-based artifact contained within that workspace. Every new PySpark notebook, every Spark Job Definition, and every data pipeline that invokes Spark compute will default to this custom environment.\n\nWe utilize Workspace-level environments when we have dedicated a specific Fabric workspace to a unified, cohesive project. For example, if we maintain a dedicated “Financial Forecasting Production” workspace, it is highly logical to attach our financial_forecasting_env directly to the workspace settings. This ensures that any engineer collaborating in this space is automatically utilizing the exact same library versions, virtually eliminating the infamous \"it works on my machine\" paradigm. This approach aligns perfectly with our goal of moving from ad-hoc setups to a governed, scalable platform.\n\n**Item-Level Environments** Conversely, “one workspace for everything” often fails at scale. In reality, we frequently utilize shared development or sandbox workspaces where multiple data scientists are working on disparate projects simultaneously. One engineer might be developing a natural language processing model requiring heavy PyTorch dependencies, while another is building a lightweight data pipeline using standard PySpark.\n\nIf we apply a massive, heavy environment at the workspace level in this scenario, we force the data engineer to suffer through longer compute initialization times for libraries they will never use. This is where Item-level environments become invaluable. Microsoft Fabric allows us to attach a custom environment specifically to an individual Notebook or a single Spark Job Definition. By overriding the workspace default, we ensure that our compute is perfectly tailored to the specific task at hand. This granular control allows us to maintain strict isolation between competing project dependencies, ensuring that a library upgrade in one notebook does not inadvertently sabotage a model in another.\n\nTo define our custom Python environment, we rely on the industry-standard environment.yml configuration file. This file serves as the declarative blueprint for our compute instances. By uploading this via the Fabric UI, we can seamlessly standardize the compute environment across all notebooks and job definitions. Let us examine the foundational structure of our environment file.\n\n```\n# Environment configuration file (environment.yml) for Fabric Workspace# Upload this via the Fabric UI to standardize the compute environment across all notebooksname: financial_forecasting_envchannels:  - conda-forge  - defaultsdependencies:  - python=3.10  - pandas=2.1.0  - numpy=1.25.2  - prophet=1.1.4  - pip:    - azure-storage-blob==12.18.0    - dbldatagen==0.3.1    # Custom internal package    - https://internalrepo.com/packages/mycompany_utils-1.0.2-py3-none-any.whl\n```\n\nWhen we construct this file, we are orchestrating a delicate balance between two primary package management systems: Conda and Pip.\n\n**The Role of Channels and Conda Dependencies**\n\nIn the channels section, we explicitly declare where the Conda package manager should look for our requested libraries. We prioritize conda-forge, a community-driven repository that typically hosts the most up-to-date and extensively tested binaries. We follow this with defaults as a fallback mechanism.\n\nUnder the dependencies block, we list our core libraries. It is an absolute best practice in enterprise data engineering to explicitly pin our version numbers (e.g., pandas=2.1.0). If we omit the version number, the environment resolver will simply fetch the latest available version at the time of environment creation. While this seems convenient, it introduces a severe risk of non-determinism. A pipeline that runs perfectly today might fail next month if the environment is rebuilt and inadvertently pulls a major version update with deprecated functions. By pinning our versions, we guarantee that our environment remains mathematically identical every time it is provisioned.\n\n**Integrating Pip Dependencies**\n\nWhile Conda is exceptional at resolving complex dependency trees involving C/C++ libraries (which is vital for libraries like NumPy and Pandas), not every package exists in the Conda repositories. This is why we include a nested - pip: section.\n\nWe use Pip for pure Python packages or specific proprietary connectors, such as azure-storage-blob==12.18.0. It is crucial that we allow Conda to install the heavy, system-level dependencies first, and then allow Pip to install the remaining packages on top. This order of operations prevents dependency conflicts where Pip might blindly overwrite a lower-level library that Conda previously optimized for our Spark cluster. By thoughtfully blending Conda and Pip, we achieve a highly resilient, comprehensive environment that caters perfectly to our advanced analytical requirements.\n\nAs we mature our enterprise data architecture, we inevitably develop proprietary code. Whether it is a highly specific data cleansing utility, a custom statistical scoring function, or a securely wrapped authentication module for our internal APIs, we must find a way to distribute this logic across our distributed Spark clusters without pasting the same code into dozens of disparate notebooks.\n\nWe solve this by packaging our internal Python code into distributable wheel (.whl) or source archive (.tar.gz) files. However, integrating these highly secure, private libraries into a cloud-based compute environment requires strategic handling.\n\n**Direct UI Uploads and Workspace Artifacts**\n\nMicrosoft Fabric gracefully simplifies this process through its environment management UI. Instead of hosting our proprietary packages on a public repository, we can upload our .whl files directly into the Fabric Environment artifact. Once uploaded, we simply reference the file within our environment configuration, or allow the Fabric platform to automatically install the uploaded artifact during the cluster initialization phase.\n\n**Enterprise CI/CD Integration** For a locked-down enterprise scenario, manually clicking and dragging files into a web browser is insufficient. The challenge of maintaining deployment pipelines behind strict corporate firewalls is a well-documented hurdle. To achieve true zero-manual-deployment automation safely, we integrate our environment management into our overarching continuous integration and continuous deployment (CI/CD) pipelines.\n\nWe can store our custom .whl files in a secure artifact repository (such as Azure Artifacts) and utilize Fabric's REST APIs or Git integration to dynamically update our workspace environments. As demonstrated in our code snippet, we can even direct Pip to fetch a wheel file from a secured internal HTTPS endpoint ([https://internalrepo.com/packages/](https://internalrepo.com/packages/)...). By automating the foundational structure of our orchestration pipelines and environments, we free our data engineers to focus on higher-level architectural challenges rather than manual configuration. This ensures that our intellectual property remains securely within our corporate boundaries while remaining effortlessly accessible to our approved compute engines.\n\nWhen we first deploy our workloads in Microsoft Fabric, everything feels seamless, but without a strategic financial operations (FinOps) framework, this frictionless platform can lead to spiraling compute costs. Every action we take within the platform consumes a fraction of our provisioned Compute Units. Therefore, the manner in which we manage our custom Python environments has a direct and measurable impact on our monthly expenditure.\n\nWhen we initiate a PySpark notebook for a data load or machine learning task, we are subjected to session initialization delays, frequently referred to as cold starts. There is an inherent latency in allocating containers, loading extensive libraries, and establishing the distributed environment. If we rely on inline %pip install commands at the top of our notebooks, we force the Spark driver and every single executor node to independently reach out to the internet, download packages, and build dependencies *every single time* the notebook is executed. If a pipeline runs every fifteen minutes, this accumulated delay becomes a significant operational bottleneck and a massive waste of Compute Units.\n\n**The Power of Pre-Compiled Environments**\n\nBy transitioning to a formal Fabric Environment artifact (using our environment.yml), we fundamentally alter this paradigm. When we publish a custom environment in Fabric, the underlying engine pre-compiles and caches the environment image. It resolves the dependency tree exactly once during the publication phase.\n\nWhen our scheduled pipeline triggers a Spark session, the compute nodes simply mount the pre-cached environment image. The cold start time is drastically reduced because the containers do not need to compile or download anything at runtime; the libraries are simply present and ready to execute. This optimization is akin to Microsoft’s proprietary V-Order compression, which reorganizes data at write time to dramatically accelerate downstream read speeds. Just as V-Order optimizes our storage layer for immediate analytical consumption, pre-compiling our custom Python environments optimizes our compute layer for immediate execution.\n\nTo further minimize overhead, we continually monitor our environments to ensure we are not bloating our compute with unused dependencies. We must resist the urge to create a single, monolithic environment containing every library our organization has ever used. By creating lean, purpose-built environments tailored to specific domains (e.g., one environment specifically for time-series forecasting, another specifically for natural language processing), we keep our cache sizes small, our startup times blindingly fast, and our Compute Unit consumption strictly controlled.\n\nWe no longer need to choose between the flexibility of open-source Python development and the rigid governance of a managed enterprise platform. In Microsoft Fabric, we possess the precise tools required to harmonize our advanced data science requirements with our operational strictures. By thoughtfully designing our environment.yml configurations, isolating our dependencies across appropriate workspace or item levels, and securely integrating our proprietary internal packages, we eradicate the chaos of dependency hell.\n\nFurthermore, by embracing pre-compiled environment artifacts, we eliminate the costly inefficiencies of runtime library installations, thereby maximizing the value of our provisioned compute capacity. As we continue to refine our analytics architecture, establishing these robust environment management practices ensures that our data engineering pipelines and machine learning models remain highly performant, exceptionally secure, and relentlessly reliable day after day.\n\n*Hey, I am* *Sandip Palit**, from Kolkata, India. I love to explore what’s new in the Data Science space and share it with the community. I am a* *Fabric Super User**, and in this* *Microsoft Fabric* *Playlist, I will share my learnings on Microsoft Fabric and the tips and tricks of using it effectively..*\n\n*Thank You for reading this article. Please feel free to share your thoughts in the comments section, and give this article a 🌟.*\n\n[Managing Dependency Hell: Setting Up Custom Python Environments in Fabric](https://pub.towardsai.net/managing-dependency-hell-setting-up-custom-python-environments-in-fabric-18af52efd4fd) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.", "url": "https://wpnews.pro/news/managing-dependency-hell-setting-up-custom-python-environments-in-fabric", "canonical_source": "https://pub.towardsai.net/managing-dependency-hell-setting-up-custom-python-environments-in-fabric-18af52efd4fd?source=rss----98111c9905da---4", "published_at": "2026-09-29 04:21:07+00:00", "updated_at": "2026-09-29 04:46:49.418972+00:00", "lang": "en", "topics": ["mlops", "developer-tools", "ai-infrastructure"], "entities": ["Microsoft Fabric", "Apache Spark", "Pandas", "NumPy", "prophet", "scikit-learn", "PySpark", "F64 SKU"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/managing-dependency-hell-setting-up-custom-python-environments-in-fabric", "markdown": "https://wpnews.pro/news/managing-dependency-hell-setting-up-custom-python-environments-in-fabric.md", "text": "https://wpnews.pro/news/managing-dependency-hell-setting-up-custom-python-environments-in-fabric.txt", "jsonld": "https://wpnews.pro/news/managing-dependency-hell-setting-up-custom-python-environments-in-fabric.jsonld"}}