{"slug": "moving-your-local-airflow-to-gcp-for-under-150-a-month", "title": "Moving Your Local Airflow to GCP for under $150 a month", "summary": "A technical team migrated its self-managed Apache Airflow orchestration from on-prem Linux boxes to Google Cloud Platform, building a self-managed Airflow VM plus a config service for under $150 a month after finding Google's managed Cloud Composer would cost roughly $350 a month. The design keeps Airflow from calling the config service directly, sharing only a versioned GCS bucket that doubles as change history, and secures the config app behind an HTTPS load balancer with Identity-Aware Proxy.", "body_md": "My team runs ML training and prediction pipelines on Vertex AI. For a long\n\ntime, the thing telling those pipelines when to run was Apache Airflow,\n\ninstalled by hand on a few on-prem Linux boxes. From there it orchestrated\n\nour hybrid cloud setup from one place.\n\nNext to Airflow sat our own configuration management system. It lets people\n\nchange a pipeline's settings on demand, and every change is tracked and\n\nlogged over time, so if a model behaves differently on Tuesday, we can see\n\nwhat changed on Monday.\n\nBoth of these lived on Linux boxes we maintained ourselves, and those boxes\n\nhad to go: our organization is phasing out its on-prem setup and moving fully\n\ncloud native. We are in the middle of that move right now. The on-prem\n\npipelines are moving out to different clouds, and our own pipelines and\n\ncompute are moving into GCP. With the work landing in GCP, putting the\n\norchestrator there too was the natural answer.\n\nThe goal was to make that move without ending up any less safe or less\n\nreliable than what we had.\n\nIn this post I want to share how we did it, chapter by chapter. The code is\n\nin the `self-managed-airflow-on-gcp`\n\nrepo, and everything below is a trimmed-down, working version of it.\n\nIf you ask \"how do I run Airflow on GCP?\", the answer is Cloud Composer.\n\nIt is managed Airflow: Google runs the scheduler, the database and the\n\nworkers, and you drop DAGs into a bucket.\n\nWe built it first. We stood up a Composer 3 environment with Terraform,\n\npointed our DAGs at it, and it worked.\n\nThe cost was the problem. With the pricing calculator, a small Composer\n\nenvironment came to about **$350 a month**. Composer 3 bills for compute\n\nunits every hour the environment is alive, whether a DAG runs or not. Our\n\npipelines do not need a cluster awake around the clock to schedule a handful\n\nof jobs, because the heavy work happens in Vertex AI anyway.\n\nMy search was not done there.\n\nOur Airflow does not do the heavy work. It decides when work happens and\n\nhands it to Vertex AI or BigQuery. A scheduler like that fits on one small\n\nVM.\n\nSo we built the same thing a second time:\n\nThe estimate for this came in **under $150 a month**. About $20 of that is\n\nthe HTTPS Load Balancer in front of our config app (GCP charges a flat\n\n$0.025/hour for the first five forwarding rules, roughly $18/month, plus\n\ndata). The rest is the VM, its disk, and Cloud NAT.\n\nThat number is not the whole cost. With Composer, Google upgrades Airflow and\n\npatches the machine. On a VM, we do. We had both versions running side by\n\nside, and the team picked the self-managed one: we are a technical team that\n\nalready ran Airflow ourselves, so the extra work was work we knew.\n\nCheaper only counts if it is also safe and reliable, so the rest of this post\n\nis about how we got there.\n\nThere are two components we deploy:\n\n`config_hq`` airflow-vm`\nThe part I like most in this design is that Airflow never calls `config_hq`.\n\nThe two only share a bucket.\n\nOur org policy disables Cloud Run's default `run.app` URL, so the service is\n\ndeployed with `default_uri_disabled = true` and there is no hostname for a\n\nDAG to call anyway. Instead, `config_hq` writes every save as a new object\n\nnamed `config/<timestamp>-<id>.txt` in a versioned bucket. That bucket is\n\nalso our change history. The VM gets read-only access to it, and the DAG\n\nreads it directly:\n\n``` python\ndef fetch_configs(**context):\n    hook = GCSHook()\n    configs = {\n        blob_name: hook.download(bucket_name=CONFIG_HQ_BUCKET, object_name=blob_name).decode(\"utf-8\")\n        for blob_name in hook.list(bucket_name=CONFIG_HQ_BUCKET)\n    }\n    context[\"ti\"].xcom_push(key=\"configs\", value=configs)\n```\n\nIf `config_hq` is down, Airflow still reads the last config that was saved.\n\nBoth components are closed by default.\n\n**`config_hq`** is only reachable through an External HTTPS Load Balancer\n\nwith Identity-Aware Proxy (IAP) on the backend:\n\n`INGRESS_TRAFFIC_INTERNAL_LOAD_BALANCER`, so\nthe Load Balancer is the only way in.`iap_authorized_members`.\n**The VM** has no external IP at all (org policy\n\n`constraints/compute.vmExternalIpAccess`), so there is no public path to it.\n\nGetting in means passing two separate gates:\n\n`35.235.240.0/20`.` constraints/compute.requireShieldedVm`.\nFor outbound traffic, Private Google Access covers `*.googleapis.com`, and\n\nCloud NAT handles everything else the VM needs for installs and `git pull`\n\n(GitHub, PyPI, and `dev.azure.com` for the git remote).\n\nThe repo has one sample pipeline. It shows the same training and prediction\n\nflow we use for our own models.\n\nThe model code lives in `app/` as a Python package, `ml_experiment`. It\n\ntrains a logistic regression on BigQuery's public penguins dataset. We build\n\nit into a wheel, upload it to the `ml-artifacts` bucket, and a DAG submits\n\nit to Vertex AI as a custom training job inside Google's prebuilt\n\n`sklearn-cpu.1-0` container. Here is a full run, from saving config to the\n\njob finishing:\n\nThe Vertex job runs as the VM's own service account, which holds\n\n`roles/aiplatform.user`, `roles/bigquery.jobUser`, and write access to its\n\nstaging bucket. Airflow gets its GCP credentials from the VM's metadata\n\nserver, so there are no key files on the machine.\n\nThe platform is Terraform, one folder per component:\n\n```\nterraform/\n├── config_hq/     # Cloud Run, Load Balancer, IAP, config bucket\n└── airflow_vm/    # VM, firewall, Cloud NAT, ml-artifacts + vertex-staging buckets\n```\n\nEach one uses Terraform workspaces, so dev, test and prod get their own\n\ncopies inside the same project, with the env name as a suffix\n\n(`airflow-vm-prod`, `example-project-config-hq-prod`). A precondition\n\nrefuses to apply on the unnamed `default` workspace, so nobody creates\n\nresources without an environment by accident.\n\nTerraform only builds a bare VM. Airflow goes on with Fabric, from\n\n`deployer/fabfile.py`, over the IAP tunnel. These are the tasks we use:\n\n| Task | What it does | \n|---|---|\n| `deploy_from_scratch` | Clones the repo, sets up Postgres and the conda env, installs Airflow, starts it, and creates the `config_hq_bucket` variable and`google_cloud_default` connection | \n| `light_deploy` | `git pull` plus a rebuild of the app package. This is our normal deploy. | \n| `restart_airflow` | Stops and starts the scheduler, DAG processor and API server | \n| `complete_teardown` | Stops everything and removes the database, env and project files | \n\nAirflow's `AIRFLOW_HOME` points at the cloned repo, so a DAG change is a\n\ncommit followed by `light_deploy`. `config_hq` has no deploy script at all.\n\n`terraform apply` builds the container with Cloud Build, pushes it, and\n\nrolls out a new Cloud Run revision.\n\n**Access** lives in Terraform variables, in three tiers:\n\n| Variable | Grants | \n|---|---|\n| `iap_tunnel_members` | open the tunnel | \n| `oslogin_members` | log in as a normal user | \n| `oslogin_admin_members` | log in with sudo (empty by default) | \n\nWe put Google Groups in there, not individual people, so adding someone to\n\nthe team is a group change and not a Terraform change. If you ever grant\n\naccess by hand in an emergency, add it to `terraform.tfvars` afterwards, or\n\nit drifts.\n\n**If the VM dies**, almost everything that matters lives somewhere else. The\n\nDAGs and the ML code are in git. Configs are in the versioned `config_hq`\n\nbucket. Model wheels and Vertex AI outputs are in their own buckets. Getting\n\nback is `terraform apply` on `airflow_vm` and then `deploy_from_scratch`,\n\nwhich is safe to re-run if it fails partway. What does not come back is\n\nAirflow's own Postgres database, with its run history. It lives on the VM's\n\ndisk, and nothing in the repo backs it up yet.\n\n**Tearing down** goes in reverse: `terraform destroy` on `airflow_vm` first,\n\nbecause it depends on `config_hq`'s bucket, then `config_hq`. The IAP brand,\n\nthe OAuth client and the Terraform state objects stay behind and have to be\n\nremoved by hand.\n\nIt is running now and costs less than half of what Composer would.\n\nThere are still chores. Airflow upgrades and OS patching are on us. The\n\nPostgres backups mentioned above are still missing. And `config_hq` uses a\n\nself-signed certificate until it gets a real domain, so the browser shows a\n\nwarning every time someone opens it.\n\nI hope you have enjoyed this one. Feel free to share your comments,\n\nespecially if you made the opposite call and stayed on Composer.\n\nRegards,\n\nErfan", "url": "https://wpnews.pro/news/moving-your-local-airflow-to-gcp-for-under-150-a-month", "canonical_source": "https://dev.to/erfankashani/moving-your-local-airflow-to-gcp-for-under-150-a-month-2i1j", "published_at": "2026-09-28 02:17:33+00:00", "updated_at": "2026-09-28 02:48:37.742951+00:00", "lang": "en", "topics": ["mlops", "ai-infrastructure", "developer-tools"], "entities": ["Google Cloud", "Apache Airflow", "Cloud Composer", "Vertex AI", "BigQuery", "Terraform", "Identity-Aware Proxy", "Google Cloud Storage"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/moving-your-local-airflow-to-gcp-for-under-150-a-month", "markdown": "https://wpnews.pro/news/moving-your-local-airflow-to-gcp-for-under-150-a-month.md", "text": "https://wpnews.pro/news/moving-your-local-airflow-to-gcp-for-under-150-a-month.txt", "jsonld": "https://wpnews.pro/news/moving-your-local-airflow-to-gcp-for-under-150-a-month.jsonld"}}