# How Honeycomb Private Cloud Drinks From the Fire Hose

> Source: <https://www.honeycomb.io/blog/how-honeycomb-private-cloud-drinks-from-the-fire-hose>
> Published: 2026-10-05 13:00:00+00:00

# How Honeycomb Private Cloud Drinks From the Fire Hose

This post walks through how the Honeycomb Private Cloud team built an automated drift report, refactored their own config to make diffing possible, and layered on an AI triage skill to keep pace with changes across the org—plus the guiding principles that made the automation trustworthy.

By: [Fred Hebert](https://www.honeycomb.io/author/fred-hebert)

#### Embracing the Code Review Bottleneck

Faced with an endless stream of AI-generated code reviews, our team made the counterintuitive choice to lean into the bottleneck rather than reduce it. Surprisingly, velocity held up, knowledge sharing improved, and we developed a collective system ownership that stuck.

[Read Now](https://www.honeycomb.io/blog/embracing-code-review-bottleneck)

Back in July, I wrote about how the Tenant team (the team behind [Honeycomb Private Cloud](https://www.honeycomb.io/platform/private-cloud) (HPC)) has [embraced the code review bottleneck to focus more of its work](https://www.honeycomb.io/blog/embracing-code-review-bottleneck). One of the other challenges we have is that we're downstream of almost all the other teams at Honeycomb, meaning that we have to package up everyone's code and services, and how it gets provisioned! This is something impossible to handle through code review since there are so many engineers on other teams, and so few of us. The strategy needed to be different, and I wanted to show you what we did to manage the process.

## Carrying shovels behind the parade

At a past job, I used to hang out with someone from the data team who would complain about how challenging it was to be an organizational afterthought (I'm paraphrasing). They had to take data produced by all the other teams—teams who didn't care about the data team's projects—and hammer it into a format that would work for all the reporting and mining that was required.

The HPC team is in a similar situation, in that most of our organization works on the SaaS offering, delivers dozens of changes every hour, and has the ability to “buy versus build” that cannot carry to running Honeycomb in a customer organization. We had already set ourselves up to consume all of the build artifacts from different teams, but their configuration could be a problem.

There are elements such as being able to change policies in IAM or pick names for AWS resources that will become too long when they expand in a different environment (for example, `production-eu1-name-of-my-resource-with-generated-policy` could very much break once it goes into `hnypc-customername-region-name-of-my-resource-with-generated-policy`). Circular dependencies can exist. Breaking or even just clickopsed operations can be a one-off thing that won't work as nicely when dozens of customer installs require a similar one. A hardcoded bucket name is globally unique and needs to be made configurable at other levels. These have been handled mostly by moving checks that detects these breaking patterns into CI at the production level or into the processes required (such as vendor reviews) so that our concerns are pushed into the feedback loop of engineers before code makes it to production.

But there was a more challenging pattern that we couldn't easily push back to engineering teams: new services with new features that require new configuration values need to be integrated to prevent our installs from breaking because we didn't port and repackage things properly for HPC.

# Read our O’Reilly book, Observability Engineering

Get your free copy and learn the foundations of observability,

right from the experts.

## Manual workflows

Our first approach for many of these elements was to do it the hard way. Lurk across channels, get flagged onto more reviews, keep track of what is going on, and see what breaks as you go. This does not scale, but is necessary to get a good grasp of the rate of change and what teams care about, compared to what *we* care about.

In doing so, patterns emerged:

- Certain change types are more relevant to you than others, particularly those that will be breaking changes due to changing interfaces or the meaning of configuration values.
- Some change types are not risky but create a subtle accumulation of drift; think for example of adjusting autoscaling to workloads changed by features but in a way you do not get to see by just looking at how hot CPU is.
- Some will slowly bubble your way, like a new service being deployed to pre-production environments before it makes it to production.

By tracking these, we were able to identify a few high-leverage points and files to monitor:

- Configuration files for services when new keys are added (while not bothering with value changes).
- New services being defined.
- Changes to the input (variables) or output definitions of terraform modules.
- Key resource version upgrades such as database or Kubernetes cluster versions in pre-production environments.

These were somewhat easier to track. When they change, we could look at the intent behind the change and see whether we needed to act on it, if we were missing steps, or if it was already covered. But constant vigilance is also troublesome. What would be a good way to automate this?

## Simple automation

The first decision to make about this was to figure out what was the proper level of automation. In an organization with a heavy drive to use AI for everything, two options were available:

1. Ask an AI to look over the code and find what change, possibly asking it to also fix things.
2. Use AI to create a straightforward script that generates a condensed diff to tackle the vigilance aspect.

There are many steps in keeping up with changes, looking at their scope, and acting on them, and so there are ways for the process to go subtly wrong. We chose to start with the second half and generated something we called the *drift report*.

The scripts were relatively simple: give a time range, grab diffs from multiple repositories, filter out any changeset that does not pertain to the high leverage points, categorized, with an option to show the actual diff.

We could then run that report at regular intervals and use it to schedule catch-up work. If we skipped the diff for a couple of weeks, we could run it for that time period, spin off projects for some of the work, or see that nothing needed to be done.

## Slightly better automation

It didn't necessarily take long to notice that what now became a bit more challenging was the need to compare what had changed to what we had in place in our Private Cloud software. Some comparisons were easy (where declarations were in dedicated files with a similar format as what our SaaS team used) and some were more challenging (when declarations were in Go code on our end but declarative for the Infra team).

If this was tricky for us, it would be tricky for automation as well. So, the next step to improve our automation was to re-structure and refactor our own code be more explicit, more amenable to diffing and comparisons. This meant re-normalizing all of our configuration overrides away from code and into files, even if the files had to be empty be to show “nothing overridden here.”

This forced our project structures to be a semi-explicit contract. As long as you provided a roughly coherent declarative mechanism, the drift reporting could then show not only what changed, but also incorporate the difference in values or whether keys existed or not *on our end*.

The diff was still fairly compact, but could now anchor *what changed there* with *what exists here* so you could declare at report generation time that something would risk clashing, breaking, or was already covered. This information also made it easier to generate much richer diffs:

Every time we ported some of these changes early, we reused the plan-based approach described in [embracing the code review bottleneck](https://www.honeycomb.io/blog/embracing-code-review-bottleneck).

## Increased automation

The plans for migrations were saved to our repositories, and we started being able to provide the changelog report entry, and became able to say, “This is a migration with the same flavor as this other one, let's do it again” and delegate that part to AI.

As we started trusting it more and provided further improvements (such as increasing PR or commit numbers when relevant), we managed to catch up and eventually turned the report into a daily occurrence. Triaging was a bit more work, but one of my teammates went in and gave very high-level instructions by writing a ‘drift-triage’ skill that mostly contains instructions similar to what we'd tell each other, with more gotchas written down about the project structure. Here are bits of the skill:

Work a drift report item by item. For **each** item, decide one of:**1. Tenant must act** → file a Linear ticket, then comment the doc line with the ticket number.**2. Not relevant to tenant** → comment the doc line explaining *why* it's a no-op.**3. Already handled** → comment the doc line noting where/when it landed.

Discuss each item with the user before filing or commenting. Go one at a time.

Don't stop at “does the tenant consume this?” — that's necessary but not sufficient. Ask, in order:**1. Is it breaking?** New *optional* var / additive key with a default = non-breaking; the tenant keeps applying. Removed/renamed/required = breaking, needs action.**2. Does the tenant already consume it?** Read the tenant's actual module call / chart override. Tenants are often already ahead of a change (a var they already pass, a DB already mirrored).**3. What was the tuning intent, and does it apply to the tenant?** This is the step people skip. A SaaS change usually encodes an intent (perf, cost, security, a bug fix). Even if the tenant doesn't consume the literal knob, decide whether the *intent* should reach the tenant's own values.

**Use the report's provenance and template signals.** Findings carry an `upstream commit: <sha> <subject>` line whenever the changelog supplied one, and the subject usually names the feature outright — that beats re-deriving intent from a values diff. […]**Nested keys are reported as dotted paths** (`partitions.newColumn`). The tenant overriding sibling keys under the same parent does **not** mean it overrides the new one — deep-merge fills the gap from chart defaults.

We ran frequent drift reports as pull requests in our repo, where we could create tickets and merge or close the PR once the follow-up work was handled and scheduled.

This skill eventually got good enough to be trusted, and at that point, my coworker added it back to the pull request. It then started suggesting which tickets would need to be opened, and with what content.

A final step was to support auto-creating the tickets based on the drift report, and then we could pick how to dispatch each of the tickets.

## Guiding principles

We're now at a step where the report runs daily and most of the tasks that follow known patterns are auto-tracked and kept to date mostly on their own. Some unusual things still happen and can take days if not sometimes weeks of work—we don't track everything everywhere, after all.

But some key patterns that proved really useful here are classics that applied even before AI:

- Provide as much determinism as you can at every step of the way so the potential for AI drift is minimized. You can abandon or improve portions of your workflow in isolation.
- Gradually grow your automation based on building an understanding of the problem so that you capture early where inefficiencies are.

Some design principles that apply more specifically to AI and automation design are:

- It is more effective to use LLMs to support re-structuring the problem space to make analysis easier than it is to use LLMs of increasing power and autonomy to automate cumbersome work.
- Favor *coordination* over*autonomy* . Focusing workflows around our needs for coordination and their related mechanisms provides a good structuring mechanism for automation as well.

In line with *leaning into a bottleneck* (since it's possibly a bottleneck *because* it matters), a key part of this approach is that instead of transferring problems to AI because they are challenging or cumbersome, we instead try to make the tasks within the bottleneck *easier* through better or different representations. Rather than making sure the code being up to date happens on its own, we try to narrow down which signals are key for us and make surfacing them easier. If we find that we struggle to keep up because synchronization is costly to repair late in a workflow, can we front-load synchronization to only keep quick checkpoints and overall be more aligned the whole time?

Basically, if we're stuck solving puzzles, we can either try to make a puzzle-solving machine to make puzzle-solving faster, or we can re-represent the components of the puzzle such that they're less demanding, less of a puzzle overall. This requires a willingness to experiment, take a step back, and change how you work, but it quickly becomes a natural habit.

This helped us keep up with changes, but also get immediate gains to our workflows that gradually compounded. Every deterministic step is easier to inspect and modify. The AI workflows are getting more convenient, but the building blocks beneath them remain independently useful to us. And if we do not like what the AI does, we can more reliably steer it by changing the underlying deterministic reports than by adding increasing sets of instructions to a prompt.

This means that issues with our workflows degrade far more gracefully, and components of the workflows themselves can be repurposed and adjusted far more granularly as well.
