# Moving Your Local Airflow to GCP for under $150 a month

> Source: <https://dev.to/erfankashani/moving-your-local-airflow-to-gcp-for-under-150-a-month-2i1j>
> Published: 2026-09-28 02:17:33+00:00

My team runs ML training and prediction pipelines on Vertex AI. For a long

time, the thing telling those pipelines when to run was Apache Airflow,

installed by hand on a few on-prem Linux boxes. From there it orchestrated

our hybrid cloud setup from one place.

Next to Airflow sat our own configuration management system. It lets people

change a pipeline's settings on demand, and every change is tracked and

logged over time, so if a model behaves differently on Tuesday, we can see

what changed on Monday.

Both of these lived on Linux boxes we maintained ourselves, and those boxes

had to go: our organization is phasing out its on-prem setup and moving fully

cloud native. We are in the middle of that move right now. The on-prem

pipelines are moving out to different clouds, and our own pipelines and

compute are moving into GCP. With the work landing in GCP, putting the

orchestrator there too was the natural answer.

The goal was to make that move without ending up any less safe or less

reliable than what we had.

In this post I want to share how we did it, chapter by chapter. The code is

in the `self-managed-airflow-on-gcp`

repo, and everything below is a trimmed-down, working version of it.

If you ask "how do I run Airflow on GCP?", the answer is Cloud Composer.

It is managed Airflow: Google runs the scheduler, the database and the

workers, and you drop DAGs into a bucket.

We built it first. We stood up a Composer 3 environment with Terraform,

pointed our DAGs at it, and it worked.

The cost was the problem. With the pricing calculator, a small Composer

environment came to about **$350 a month**. Composer 3 bills for compute

units every hour the environment is alive, whether a DAG runs or not. Our

pipelines do not need a cluster awake around the clock to schedule a handful

of jobs, because the heavy work happens in Vertex AI anyway.

My search was not done there.

Our Airflow does not do the heavy work. It decides when work happens and

hands it to Vertex AI or BigQuery. A scheduler like that fits on one small

VM.

So we built the same thing a second time:

The estimate for this came in **under $150 a month**. About $20 of that is

the HTTPS Load Balancer in front of our config app (GCP charges a flat

$0.025/hour for the first five forwarding rules, roughly $18/month, plus

data). The rest is the VM, its disk, and Cloud NAT.

That number is not the whole cost. With Composer, Google upgrades Airflow and

patches the machine. On a VM, we do. We had both versions running side by

side, and the team picked the self-managed one: we are a technical team that

already ran Airflow ourselves, so the extra work was work we knew.

Cheaper only counts if it is also safe and reliable, so the rest of this post

is about how we got there.

There are two components we deploy:

`config_hq`` airflow-vm`
The part I like most in this design is that Airflow never calls `config_hq`.

The two only share a bucket.

Our org policy disables Cloud Run's default `run.app` URL, so the service is

deployed with `default_uri_disabled = true` and there is no hostname for a

DAG to call anyway. Instead, `config_hq` writes every save as a new object

named `config/<timestamp>-<id>.txt` in a versioned bucket. That bucket is

also our change history. The VM gets read-only access to it, and the DAG

reads it directly:

``` python
def fetch_configs(**context):
    hook = GCSHook()
    configs = {
        blob_name: hook.download(bucket_name=CONFIG_HQ_BUCKET, object_name=blob_name).decode("utf-8")
        for blob_name in hook.list(bucket_name=CONFIG_HQ_BUCKET)
    }
    context["ti"].xcom_push(key="configs", value=configs)
```

If `config_hq` is down, Airflow still reads the last config that was saved.

Both components are closed by default.

**`config_hq`** is only reachable through an External HTTPS Load Balancer

with Identity-Aware Proxy (IAP) on the backend:

`INGRESS_TRAFFIC_INTERNAL_LOAD_BALANCER`, so
the Load Balancer is the only way in.`iap_authorized_members`.
**The VM** has no external IP at all (org policy

`constraints/compute.vmExternalIpAccess`), so there is no public path to it.

Getting in means passing two separate gates:

`35.235.240.0/20`.` constraints/compute.requireShieldedVm`.
For outbound traffic, Private Google Access covers `*.googleapis.com`, and

Cloud NAT handles everything else the VM needs for installs and `git pull`

(GitHub, PyPI, and `dev.azure.com` for the git remote).

The repo has one sample pipeline. It shows the same training and prediction

flow we use for our own models.

The model code lives in `app/` as a Python package, `ml_experiment`. It

trains a logistic regression on BigQuery's public penguins dataset. We build

it into a wheel, upload it to the `ml-artifacts` bucket, and a DAG submits

it to Vertex AI as a custom training job inside Google's prebuilt

`sklearn-cpu.1-0` container. Here is a full run, from saving config to the

job finishing:

The Vertex job runs as the VM's own service account, which holds

`roles/aiplatform.user`, `roles/bigquery.jobUser`, and write access to its

staging bucket. Airflow gets its GCP credentials from the VM's metadata

server, so there are no key files on the machine.

The platform is Terraform, one folder per component:

```
terraform/
├── config_hq/     # Cloud Run, Load Balancer, IAP, config bucket
└── airflow_vm/    # VM, firewall, Cloud NAT, ml-artifacts + vertex-staging buckets
```

Each one uses Terraform workspaces, so dev, test and prod get their own

copies inside the same project, with the env name as a suffix

(`airflow-vm-prod`, `example-project-config-hq-prod`). A precondition

refuses to apply on the unnamed `default` workspace, so nobody creates

resources without an environment by accident.

Terraform only builds a bare VM. Airflow goes on with Fabric, from

`deployer/fabfile.py`, over the IAP tunnel. These are the tasks we use:

| Task | What it does | 
|---|---|
| `deploy_from_scratch` | Clones the repo, sets up Postgres and the conda env, installs Airflow, starts it, and creates the `config_hq_bucket` variable and`google_cloud_default` connection | 
| `light_deploy` | `git pull` plus a rebuild of the app package. This is our normal deploy. | 
| `restart_airflow` | Stops and starts the scheduler, DAG processor and API server | 
| `complete_teardown` | Stops everything and removes the database, env and project files | 

Airflow's `AIRFLOW_HOME` points at the cloned repo, so a DAG change is a

commit followed by `light_deploy`. `config_hq` has no deploy script at all.

`terraform apply` builds the container with Cloud Build, pushes it, and

rolls out a new Cloud Run revision.

**Access** lives in Terraform variables, in three tiers:

| Variable | Grants | 
|---|---|
| `iap_tunnel_members` | open the tunnel | 
| `oslogin_members` | log in as a normal user | 
| `oslogin_admin_members` | log in with sudo (empty by default) | 

We put Google Groups in there, not individual people, so adding someone to

the team is a group change and not a Terraform change. If you ever grant

access by hand in an emergency, add it to `terraform.tfvars` afterwards, or

it drifts.

**If the VM dies**, almost everything that matters lives somewhere else. The

DAGs and the ML code are in git. Configs are in the versioned `config_hq`

bucket. Model wheels and Vertex AI outputs are in their own buckets. Getting

back is `terraform apply` on `airflow_vm` and then `deploy_from_scratch`,

which is safe to re-run if it fails partway. What does not come back is

Airflow's own Postgres database, with its run history. It lives on the VM's

disk, and nothing in the repo backs it up yet.

**Tearing down** goes in reverse: `terraform destroy` on `airflow_vm` first,

because it depends on `config_hq`'s bucket, then `config_hq`. The IAP brand,

the OAuth client and the Terraform state objects stay behind and have to be

removed by hand.

It is running now and costs less than half of what Composer would.

There are still chores. Airflow upgrades and OS patching are on us. The

Postgres backups mentioned above are still missing. And `config_hq` uses a

self-signed certificate until it gets a real domain, so the browser shows a

warning every time someone opens it.

I hope you have enjoyed this one. Feel free to share your comments,

especially if you made the opposite call and stayed on Composer.

Regards,

Erfan
