# GitHub’s Commit Volume Doubled to 2.9B/Month — Then It Broke

> Source: <https://byteiota.com/githubs-commit-volume-doubled-to-2-9b-month-then-it-broke/>
> Published: 2026-10-05 17:17:52+00:00

On August 17, GitHub went dark for nearly eight hours. Not because of a breach — but because AI coding agents had become so prolific that GitHub’s own infrastructure couldn’t keep up. Monthly commits had doubled from 1.4 billion to 2.9 billion in just four months. A single misconfigured Istio sidecar policy cascaded into authentication failures that locked enterprise teams out of their code during business hours. The platform that 100 million developers depend on broke under the weight of the very workflow it helped build.

## The Numbers Are Genuinely Staggering

GitHub now processes **2.9 billion commits per month** — equivalent to 275 million per week. To put that in context: GitHub handled roughly 1 billion total commits *in all of 2025*. Now it processes nearly three times that amount every single month. The company is on pace for 14 billion commits in 2026.

The driver is obvious. Claude Code alone generates 4.5% of all public GitHub commits — about 2.6 million per week, up 25x from just a year ago. AI agents opened 17 million pull requests in March 2026, a 4x increase from six months prior. Coding agents overall grew 1,400% this year. [The New Stack reported](https://thenewstack.io/github-2-9b-monthly-commits/) that GitHub added 3 million CPU cores and 120 petabytes of storage in response — and still wasn’t fast enough.

## What Actually Happened on August 17

The [post-mortem from GitHub](https://github.blog/news-insights/company-news/the-august-17-outage-and-the-work-ahead/) is worth reading in full, but the short version: an Istio sidecar pod hit its concurrency limit. The policy watching it was configured to monitor the host service — not the sidecar. One failure cascaded to four HAProxy nodes that exhausted their flow limits, degrading the entire gateway authentication path. Optimistic retry logic worsened the damage by flooding internal load balancers.

For nearly eight hours, error rates hit 20% for web and API requests. Archive and raw-content downloads reached 50% error rates. Enterprise users with SAML or OIDC authentication couldn’t log in at all. GitHub’s own Copilot product made recovery slower — its client-side retry loop amplified traffic exactly when the system was most fragile. The tool GitHub sells as a productivity multiplier was part of the problem.

## The Emergency Scale-Up

The response reveals how badly GitHub misjudged the growth curve. In October 2025, the company planned for 10x capacity growth. By February 2026, internal projections showed 30x was needed. GitHub has since migrated Azure from handling 12% of platform traffic in May to **58% today**. Overflow is routing through AWS — Microsoft is sending its own platform’s traffic through a direct competitor’s cloud infrastructure.

GitHub COO Kyle Daigle acknowledged the scale of the challenge in a [Latent Space interview](https://www.latent.space/p/github), saying infrastructure rewrites should produce “step-change improvements within three months.” Reasonable — but the August outage consumed an entire year’s worth of enterprise SLA budget in a single day for some customers.

## The Part Nobody Is Talking About

The infrastructure story is the obvious one. The verification gap is harder to see and more dangerous.

Commits doubled. Commit verification capacity didn’t. AI agents rarely sign their commits. When 80% of pull requests are opened by agents — and other agents are reviewing them — traditional trust signals collapse. Reputation, commit history, human review: these heuristics assume human-paced development. They break at agent velocity.

“The industry’s hardest unsolved problem is trust — when agents write 80% of pull requests and other agents review them, traditional social signals like reputation and commit history break down.”

Kyle Daigle, COO, GitHub

Plugin4Shell, disclosed in September 2026, demonstrated the attack surface: a supply chain vulnerability that lets attackers swap plugin code even when coding agents have locked plugins to specific reviewed commit hashes. Volume makes anomaly detection harder too — everything looks anomalous now. [IncidentHub’s reliability analysis](https://blog.incidenthub.cloud/github-reliability-outage-history-2025-2026) tracked 257 incidents between May 2025 and April 2026 — roughly one significant disruption per week.

## What Developers Should Do Now

Stop treating GitHub as a guaranteed utility. It isn’t one anymore. Here’s what that means in practice:

- **GitHub Actions redundancy:** Actions had 57 outages from May 2025 through April 2026. If your deployment pipeline lives entirely in Actions, you have a single point of failure. Build fallback triggers or evaluate[Semaphore](https://semaphoreci.com) , which benchmarks at roughly twice the throughput for agent-native workloads.
- **Status monitoring:** Automate GitHub status checks in your pipeline. If GitHub’s status API returns degraded, stop deploying — don’t flood an already-struggling platform with retries.
- **Commit signing for agents:** Require GPG or SSH commit signing from your CI/CD agents. The Plugin4Shell attack class exploits unsigned agent commits. This is a cheap safeguard.
- **Review your SLAs:** If your team builds customer-facing services on GitHub Actions, verify your contracts account for platform outages outside your control. The August 17 incident broke annual uptime budgets in a single day.

## The Structural Problem

GitHub built its infrastructure for human-paced commits — individual developers pushing work over hours and days. It now runs under machine-paced commits from thousands of simultaneous agents. That mismatch isn’t fixed by adding CPU cores. It requires rethinking how authentication, rate limiting, commit verification, and trust signals work when the “developer” is a fleet of agents rather than a person.

GitHub isn’t going anywhere. Its network effects — stars, forks, dependency graphs, developer identities — are real and durable. But the August 17 outage wasn’t just a bad day. It was a signal that the infrastructure model underpinning all of software development is under genuine structural stress. Plan accordingly.
