My team runs ML training and prediction pipelines on Vertex AI. For a long
time, the thing telling those pipelines when to run was Apache Airflow,
installed by hand on a few on-prem Linux boxes. From there it orchestrated
our hybrid cloud setup from one place.
Next to Airflow sat our own configuration management system. It lets people
change a pipeline's settings on demand, and every change is tracked and
logged over time, so if a model behaves differently on Tuesday, we can see
what changed on Monday.
Both of these lived on Linux boxes we maintained ourselves, and those boxes
had to go: our organization is phasing out its on-prem setup and moving fully
cloud native. We are in the middle of that move right now. The on-prem
pipelines are moving out to different clouds, and our own pipelines and
compute are moving into GCP. With the work landing in GCP, putting the
orchestrator there too was the natural answer.
The goal was to make that move without ending up any less safe or less
reliable than what we had.
In this post I want to share how we did it, chapter by chapter. The code is
in the self-managed-airflow-on-gcp
repo, and everything below is a trimmed-down, working version of it.
If you ask "how do I run Airflow on GCP?", the answer is Cloud Composer.
It is managed Airflow: Google runs the scheduler, the database and the
workers, and you drop DAGs into a bucket.
We built it first. We stood up a Composer 3 environment with Terraform,
pointed our DAGs at it, and it worked.
The cost was the problem. With the pricing calculator, a small Composer
environment came to about $350 a month. Composer 3 bills for compute
units every hour the environment is alive, whether a DAG runs or not. Our
pipelines do not need a cluster awake around the clock to schedule a handful
of jobs, because the heavy work happens in Vertex AI anyway.
My search was not done there.
Our Airflow does not do the heavy work. It decides when work happens and
hands it to Vertex AI or BigQuery. A scheduler like that fits on one small
VM.
So we built the same thing a second time:
The estimate for this came in under $150 a month. About $20 of that is
the HTTPS Load Balancer in front of our config app (GCP charges a flat
$0.025/hour for the first five forwarding rules, roughly $18/month, plus
data). The rest is the VM, its disk, and Cloud NAT.
That number is not the whole cost. With Composer, Google upgrades Airflow and
patches the machine. On a VM, we do. We had both versions running side by
side, and the team picked the self-managed one: we are a technical team that
already ran Airflow ourselves, so the extra work was work we knew.
Cheaper only counts if it is also safe and reliable, so the rest of this post
is about how we got there.
There are two components we deploy:
config_hq`` airflow-vm
The part I like most in this design is that Airflow never calls config_hq.
The two only share a bucket.
Our org policy disables Cloud Run's default run.app URL, so the service is
deployed with default_uri_disabled = true and there is no hostname for a
DAG to call anyway. Instead, config_hq writes every save as a new object
named config/<timestamp>-<id>.txt in a versioned bucket. That bucket is
also our change history. The VM gets read-only access to it, and the DAG
reads it directly:
def fetch_configs(**context):
hook = GCSHook()
configs = {
blob_name: hook.download(bucket_name=CONFIG_HQ_BUCKET, object_name=blob_name).decode("utf-8")
for blob_name in hook.list(bucket_name=CONFIG_HQ_BUCKET)
}
context["ti"].xcom_push(key="configs", value=configs)
If config_hq is down, Airflow still reads the last config that was saved.
Both components are closed by default.
config_hq is only reachable through an External HTTPS Load Balancer
with Identity-Aware Proxy (IAP) on the backend:
INGRESS_TRAFFIC_INTERNAL_LOAD_BALANCER, so
the Load Balancer is the only way in.iap_authorized_members.
The VM has no external IP at all (org policy
constraints/compute.vmExternalIpAccess), so there is no public path to it.
Getting in means passing two separate gates:
35.235.240.0/20. constraints/compute.requireShieldedVm.
For outbound traffic, Private Google Access covers *.googleapis.com, and
Cloud NAT handles everything else the VM needs for installs and git pull
(GitHub, PyPI, and dev.azure.com for the git remote).
The repo has one sample pipeline. It shows the same training and prediction
flow we use for our own models.
The model code lives in app/ as a Python package, ml_experiment. It
trains a logistic regression on BigQuery's public penguins dataset. We build
it into a wheel, upload it to the ml-artifacts bucket, and a DAG submits
it to Vertex AI as a custom training job inside Google's prebuilt
sklearn-cpu.1-0 container. Here is a full run, from saving config to the
job finishing:
The Vertex job runs as the VM's own service account, which holds
roles/aiplatform.user, roles/bigquery.jobUser, and write access to its
staging bucket. Airflow gets its GCP credentials from the VM's metadata
server, so there are no key files on the machine.
The platform is Terraform, one folder per component:
terraform/
├── config_hq/ # Cloud Run, Load Balancer, IAP, config bucket
└── airflow_vm/ # VM, firewall, Cloud NAT, ml-artifacts + vertex-staging buckets
Each one uses Terraform workspaces, so dev, test and prod get their own
copies inside the same project, with the env name as a suffix
(airflow-vm-prod, example-project-config-hq-prod). A precondition
refuses to apply on the unnamed default workspace, so nobody creates
resources without an environment by accident.
Terraform only builds a bare VM. Airflow goes on with Fabric, from
deployer/fabfile.py, over the IAP tunnel. These are the tasks we use:
| Task | What it does |
|---|---|
deploy_from_scratch |
Clones the repo, sets up Postgres and the conda env, installs Airflow, starts it, and creates the config_hq_bucket variable andgoogle_cloud_default connection |
light_deploy |
git pull plus a rebuild of the app package. This is our normal deploy. |
restart_airflow |
Stops and starts the scheduler, DAG processor and API server |
complete_teardown |
Stops everything and removes the database, env and project files |
Airflow's AIRFLOW_HOME points at the cloned repo, so a DAG change is a
commit followed by light_deploy. config_hq has no deploy script at all.
terraform apply builds the container with Cloud Build, pushes it, and
rolls out a new Cloud Run revision.
Access lives in Terraform variables, in three tiers:
| Variable | Grants |
|---|---|
iap_tunnel_members |
open the tunnel |
oslogin_members |
log in as a normal user |
oslogin_admin_members |
log in with sudo (empty by default) |
We put Google Groups in there, not individual people, so adding someone to
the team is a group change and not a Terraform change. If you ever grant
access by hand in an emergency, add it to terraform.tfvars afterwards, or
it drifts.
If the VM dies, almost everything that matters lives somewhere else. The
DAGs and the ML code are in git. Configs are in the versioned config_hq
bucket. Model wheels and Vertex AI outputs are in their own buckets. Getting
back is terraform apply on airflow_vm and then deploy_from_scratch,
which is safe to re-run if it fails partway. What does not come back is
Airflow's own Postgres database, with its run history. It lives on the VM's
disk, and nothing in the repo backs it up yet.
Tearing down goes in reverse: terraform destroy on airflow_vm first,
because it depends on config_hq's bucket, then config_hq. The IAP brand,
the OAuth client and the Terraform state objects stay behind and have to be
removed by hand.
It is running now and costs less than half of what Composer would.
There are still chores. Airflow upgrades and OS patching are on us. The
Postgres backups mentioned above are still missing. And config_hq uses a
self-signed certificate until it gets a real domain, so the browser shows a
warning every time someone opens it.
I hope you have enjoyed this one. Feel free to share your comments,
especially if you made the opposite call and stayed on Composer.
Regards,
Erfan