{"slug": "how-we-systematically-improved-our-reliability", "title": "How we systematically improved our reliability", "summary": "Neon adopted Databricks' postmortem standard and rolled out reliability improvements across its Lakebase Postgres control plane after two incidents in one week earlier this year, both triggered by an external dependency failure that caused a self-reinforcing loop and prevented the control plane from recovering on its own. Lakebase Postgres runs across Neon and Databricks surfaces, AWS, Azure and GCP, in more than 30 regions with up to 15 cells per region, driving the number of environments above 70. Neon said its stability is already better because of the practices, which structure postmortems around impact — affected customers with denominator and percentage, affected endpoints and regions, error rate and latency against a baseline, duration, and measurement confidence.", "body_md": "#### Neon is a complete set of cloud backend primitives\n\n[We just announced](https://neon.com/blog/neon-backend-is-ga) that Object Storage, Managed Better Auth, Functions, and AI Gateway are now generally available. But even as our toolset grows, Lakebase Postgres remains the core of the backend, and the rest of the primitives depend on it. We'll continue to double down on database work, not only on new features but in reliability improvements as well.\n\nLet's start with a simple problem: one database, one process, one host. That is a reliability problem we understand. We know how to connect, inspect, restart, and read logs.\n\nNow, copy and paste that millions of times. Processes crash. Machines go down. Networks flap. At the million-database scale, these failures are routine. And Lakebase deals with these failures every day.\n\nLakebase Postgres runs a very large fleet across two surfaces (Neon and Databricks) and three clouds: AWS, Azure, and GCP. We are in more than 30 regions. We also use cells - isolated deployments that let us scale horizontally and reduce blast radius inside a region. A region can have up to 15 cells. That drives the number of environments above 70.\n\nWhat started as a database problem is now a global operations problem.\n\nIn Lakebase Postgres, a set of services manages the fleet. We call them the control plane. They provision and configure databases, monitor health, recover databases when needed, and manage high availability, branch creation, scale to zero, and other features. Control-plane failures can prevent database creation, waking, and recovery. When those paths fail, affected customers may be unable to connect. The control plane itself has to stay resilient to failure at fleet scale.\n\nEarlier this year, we had two incidents in one week. Both started with an external dependency failing; that failure triggered a self-reinforcing loop that took down the control plane and kept it from recovering on its own.\n\nAfter those incidents, we doubled down on reliability. This post summarizes the reliability work we've been rolling out since then to keep uptime high as our fleet, product, and customer base grow.\n\nThis work never really ends, and new improvements will keep shipping. But we're pleased with the results so far: our stability is already better because of these practices.\n\n## The first step: writing effective postmortems\n\nA good postmortem reconstructs what happened, finds the root causes, and turns them into fixes and lessons. Databricks has a great standard practice for postmortems that Neon adopted as one of the transitions of moving from startup-scale practices to enterprise-grade reliability.\n\nOur postmortems are now structured around these two general principles:\n\n- An effective postmortem exists to reduce future customer harm. It should work as a risk-reduction tool vs just a record of events.\n- A reviewer who was not involved in the response should be able to understand what customers experienced, trace that outcome through the system, check the evidence, and decide whether the proposed actions are proportionate and complete.\n\n### We start with the impact\n\nThe first thing we write is the impact of the incident. This section is the foundation of the whole postmortem, and it analyzes\n\n- **The who -** affected customers, accounts, workspaces, or projects, with the denominator and percentage\n- **The what -** affected endpoints, APIs, regions, cells, and control-plane and data-plane operations, and what was not affected\n- **How bad -** error rate, latency against a baseline, failed operations\n- **How long -** start, end, duration, and whether impact was continuous or intermittent\n- **How sure -** measurement source, method, known blind spots, and whether each figure is exact or estimated\n\nGetting the numbers right matters for three reasons:\n\n- **Proportionality.** If you record \"customers experienced latency\", this may mean very different things. A 10% latency increase on one non-critical method and a 1,000x increase across all control-plane and data-plane operations can both be described as \"latency\", but both incidents need very different remediation.\n- **Unknown impact is still information.** Do not manufacture precision. If the system can't produce a count, say so, explain why, bound it where the evidence allows, and treat the gap as an observability action item.\n- **Internal customers count too.** On-call engineers, support specialists, and product managers are affected by incidents. Measure that in time wasted and surprise interruptions, especially nightly pages.\n\n#### Using AI effectively\n\nAI agents are useful for digging through logs to work out how many customers were affected. But an agent can get numbers wrong, so we never paste its figures straight into a postmortem. Instead, we ask it to build a notebook that shows each number next to the exact query that produced it. That way, an engineer can rerun every query and check the result.\n\nEven better is not needing AI for this at all. If a service has dashboards that track its reliability targets (SLOs) and the metrics behind them (SLIs), anyone can open them and see exactly who was affected, and the answer is the same every time.\n\n### We measure how long it took us to detect and mitigate\n\nNext, we look at how we responded. We record three timelines:\n\n- **Time to detection:** the time from when customers were first affected to when a person first reacted\n- **Time to mitigation** : the time from detection to the moment customer harm actually stopped. We break this into smaller steps (reaching the right engineer, figuring out the problem, deciding what to do, doing it, and waiting for it to take effect) so we can see where the time went.\n- **Time of recovery:** we confirm recovery from the customer's side\n\n**For each of these time spans, we ask ourselves: what would it take to cut this time in half?** The answer should be specific - e.g. we may need a new signal, alert, runbook, permission, or piece of automation.\n\n### We find the root cause\n\nWe use the [5 Whys](https://asq.org/quality-resources/five-whys) framework, starting from the impact: \"Why were [number] customers unable to use [feature] for [duration]?\" Each answer becomes the next question, until we reach something specific in the system that we can fix. Every answer has to point to evidence and name the safeguard that was missing or didn't work. Labels like \"human error\" or \"bad deployment\" aren't answers, because they don't explain why the system let the problem reach customers.\n\nIncidents rarely have one cause. In complex systems, several things usually go wrong at once. So the Whys often branch into a tree:\n\n- **why it failed**\n- **why it affected so many customers**\n- **why it lasted so long**\n- **and why it took so long to detect and stop**\n\nWhen the evidence can't settle which explanation is right, we label the options as hypotheses.\n\n\"We don't know\" is only acceptable in two cases:\n\n1. The evidence doesn't exist - for example because the logs weren't kept. We say what's missing and add an action item so we have the data next time.\n2. A full answer would take far too much work. We only stop if we can show our fixes cover the worst realistic outcome.\n\n#### Dependency failures are also our problem\n\nWhen something we depend on fails, \"the vendor had an outage\" is not a cause we can act on. To our customers, that dependency is part of our product. So we look for the missing safeguard on our side, such as a timeout, a fallback, a cache, or isolation from the failing service.\n\n### We write action items that can actually be closed\n\nEvery problem we find ends in an action item, a clear reason no action is needed, or an accepted risk with an owner. Good action items are:\n\n- **Specific and testable:** not \"improve monitoring,\" but which alert to add, when it fires, who it pages, and how we'll prove it works\n- **Small enough to finish:** days for urgent fixes, weeks for most items, a few months at most\n- **Worth doing:** we prioritize by customer impact and by how likely the problem is to happen again, and skip items that cost a lot for little benefit\n- **Linked from the text:** each problem points to its action item where it's mentioned, like \"[AI10: check that all alerts link to a runbook].\" If we can't link one, we've found a gap\n\n### Last, we write the executive summary\n\nThe executive summary appears first in the document, but we write it last, once the rest of the analysis backs up its numbers. It should make sense on its own: what happened, how many customers were affected and for how long, why it happened, how we stopped it, and what risk remains.\n\nOnce the draft is ready, it gets two reviews. The first is a team review, ideally including someone who wasn't involved in the incident. The second is an independent review by a senior reviewer, who checks the whole document.\n\n## After the postmortems, come the fixes\n\n**The postmortems for the two incidents at the start of this post led to more than 30 follow-up fixes over three months.** They fell into three areas: how the control plane uses its own database, bugs and defaults in control-plane code, and weaknesses in the system design that only showed up under incident load.\n\n### Database behavior in the control plane\n\nThe control plane keeps its own state in a Postgres database. When that database has problems, provisioning, health checks, and recovery slow down or fail across the customer fleet.\n\nWe fixed several query-plan regressions and added safeguards so they don't come back. We also started blocking long-running transactions, so a stuck query can't hold up the control plane. Finally, we tuned a few Postgres parameters, including turning off parallel query execution, which had caused some of the plan and runtime problems.\n\n### Control-plane bugs and defaults\n\nNext, we went through control-plane code and runtime settings. We found problems in:\n\n- An open-source dependency we use\n- Queue processing\n- TCP timeout settings\n- Connection-pool defaults\n\nMost of these weren't in code we wrote. They were defaults and behavior we inherited. A widely used library can still ship defaults that don't suit a control plane managing millions of databases, so we review and set those values for our workload.\n\n### System design\n\nThe last group of fixes was architectural, and these problems showed up most clearly during the incidents themselves:\n\n- Our backpressure mechanisms, which slow down incoming work when the system is overloaded, didn't work as they should\n- Under load, we had starvation and fairness issues: some work kept making progress while other work waited\n\n## We also built a process to find failures earlier\n\nThose 30+ fixes addressed what went wrong in the two incidents. But fixes that come out of a postmortem arrive late - by the time we write them, customers have already been affected.\n\n**So we stepped back and asked where reliability failures come from in the first place, so we could catch them before the next incident.**\n\nWe discover that they come from three places:\n\n- **The environment:** infrastructure, networks, dependencies, hardware. Anything we depend on will eventually fail. Both of our incidents started here, with an external dependency failing.\n- **Changes we make:** code, configuration, migrations, operations.\n- **Client workload:** traffic, usage patterns, invalid input, retries. Even predictable patterns, such as hourly cron spikes, can become dangerous at scale.\n\nKnowing where failures come from, we built a system around two ideas:\n\n1. Every change gets a risk score\n2. The higher the score, the more reliability checks the change has to pass as it moves through design, build, release, and operation\n\n### Risk decides how much rigor a change gets\n\nWe score every change like this:\n\nRisk = probability × impact\n\nProbability reflects how big and how new the change is, and how many dependencies it touches. Impact reflects which customer workflows it could break, how many customers it could reach, and how long recovery would take.\n\n- Low risk: a change to an internal dashboard\n- High risk: a control-plane migration that affects database creation\n\nThe score decides how many reliability checks a change has to pass. High-risk changes get more scrutiny, and low-risk changes stay fast.\n\n### Reliability checks at every stage\n\nEvery change already goes through the same four stages: design, build, release, and operate. We added reliability checks to each one, scaled to the change's risk score, so that all three sources of failure get looked at before a change reaches customers.\n\n### Design\n\nBefore we write much code, we ask how the change could fail:\n\n- An AI skill we built reviews the design from a reliability point of view\n- We give the change its risk score\n- For high-risk changes, we run a premortem: we imagine the feature has already caused an incident, then work out how that could have happened and how to prevent it\n\n### Build\n\nWhile we write the code, we make sure the likely failures are tested and that we can undo the change:\n\n- We review the code for reliability\n- We run load and failure tests\n- We check that rollback works\n\n### Release\n\nAs we roll out, we watch the fleet and stop if the change misbehaves:\n\n- We ramp up in stages\n- The feature owner watches an SLO during the rollout\n- High-risk changes go through an operational readiness review\n\n### Operate\n\nOnce the change is live, the team that gets paged has to be able to run it:\n\n- The feature owner hands over what the on-call team needs to know\n- We run on-call drills\n- The team reviews every dip in reliability, even the smallest\n\n## What's next\n\nThe process we've described is already preventing incidents, and we're extending it with more analysis and more automation, including\n\n- **Failure mode and effects analysis (FMEA):** we pick a component, imagine it failing, trace the impact on customers, and fix the weak spots before it happens.\n- **Fault tree analysis (FTA):** we start from a failure customers would see and work backward to every component that could cause it.\n- **Automated stress and failure testing:** we make load, limit, and recovery tests run on their own, so they don't depend on someone remembering to run them.\n\nYou trust us with your data and your uptime. We take that seriously. Reliability work doesn't have a finish line: we'll keep improving and sharing what we learn as we go.", "url": "https://wpnews.pro/news/how-we-systematically-improved-our-reliability", "canonical_source": "https://neon.com/blog/how-we-systematically-improved-our-reliability", "published_at": "2026-09-30 12:00:00+00:00", "updated_at": "2026-09-30 17:19:07.708902+00:00", "lang": "en", "topics": ["ai-infrastructure", "mlops"], "entities": ["Neon", "Databricks", "Lakebase Postgres", "AWS", "Azure", "GCP"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-we-systematically-improved-our-reliability", "markdown": "https://wpnews.pro/news/how-we-systematically-improved-our-reliability.md", "text": "https://wpnews.pro/news/how-we-systematically-improved-our-reliability.txt", "jsonld": "https://wpnews.pro/news/how-we-systematically-improved-our-reliability.jsonld"}}