GitHub has attributed this week’s nearly eight-hour outage to a chain of infrastructure failures that began with saturated load balancers and was amplified by retry behavior in Visual Studio Code, according to The Register’s report on GitHub’s account of the incident. The incident began at 13:28 UTC on August 17 and was not fully resolved until 21:15 UTC, a duration of 7 hours and 47 minutes, The Register reports. GitHub said the disruption produced elevated errors across Issues, Pull Requests, application programming interfaces, Actions, and Copilot — a broad enough footprint to affect both code collaboration and automated development workflows. The immediate technical failure, according to the report, was network saturation on load balancers in GitHub’s Central US facility. GitHub traced that saturation to an Istio sidecar reaching its concurrency limit. In a service-mesh architecture, sidecars often sit alongside application services to manage traffic and policy; if that layer hits a limit the rest of the system does not correctly observe, capacity planning can fail in non-obvious ways. The autoscaling layer did not catch the right constraint. The Register reports that GitHub’s policy monitored the host service, but not the Istio sidecar’s concurrency limit. That blind spot meant the system did not add capacity as the sidecar limit was reached, allowing a cascading failure to develop. Retry logic then made the problem worse. GitHub said, according to The Register, that optimistic retry behavior overloaded internal load balancers. A delayed response from one internal endpoint also triggered what GitHub called a latent retry bug in Visual Studio Code, amplifying traffic to the Copilot Token Service by roughly 10x and slowing recovery for that service. Recovery was staggered. The Register reports that most services recovered by 16:36 UTC, and Actions recovered by 18:03 UTC. The Copilot Token Service, which was affected by the amplified token traffic, did not recover until 21:02 UTC, shortly before the incident was fully resolved at 21:15 UTC. GitHub’s engineers mitigated the incident by temporarily reducing gateway retries through a code change and configuring load balancers to reject inbound Copilot Token Service requests with HTTP 403 responses, according to the report. GitHub also said scraping attacks on codeload endpoints complicated recovery. For follow-up, GitHub said it will correct the autoscaling policies, review retry limits, audit Istio concurrency settings, and address the Visual Studio Code behavior that amplified Copilot token traffic. Those are targeted fixes to the failure chain GitHub described: missed saturation signals, retry amplification, and service-specific recovery delays. Who benefits: Infrastructure teams that audit service-mesh limits, retry budgets, and autoscaling triggers can use this incident as a concrete failure pattern to test against. Vendors building developer infrastructure alternatives may also gain attention when a major platform outage disrupts normal workflows. Who's exposed: Teams heavily dependent on GitHub-hosted workflows, GitHub Actions, APIs, or Copilot were exposed to a single platform-level incident. Systems with aggressive client retries or incomplete monitoring of sidecar limits face similar amplification risks.