# GitHub Actions large hosted runner availability degraded by VM provisioning failures

DevFeed: [GitHub Actions large hosted runner availability degraded by VM provisioning failures](<https://devfeed.tech/articles/disruption-with-some-github-services-76162.md>)

Original publisher: [Read original article](<https://www.githubstatus.com/incidents/v4d2jbm842p4>)

Published: 2024-12-03T04:11:00Z

Content type: news

Language: en

Sources: [GitHub Status - Incident History](<https://devfeed.tech/sources/github-status-incident-history.md>)

Topics: [self-hosted runner](<https://devfeed.tech/topics/self-hosted-runner.md>), [GitHub](<https://devfeed.tech/topics/github.md>), [Monitoring & Alerting](<https://devfeed.tech/topics/monitoring-alerting.md>), [Availability](<https://devfeed.tech/topics/availability.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [github](<https://devfeed.tech/tags/github.md>), [impact](<https://devfeed.tech/tags/impact.md>), [jobs](<https://devfeed.tech/tags/jobs.md>), [performance](<https://devfeed.tech/tags/performance.md>), [resiliency](<https://devfeed.tech/tags/resiliency.md>), [runner](<https://devfeed.tech/tags/runner.md>)

## AI overview

On Dec 3, failures in background VM provisioning jobs degraded GitHub Actions large hosted runners for an hour. About 13.5% of workflows requiring large runners were affected on average, peaking at 46%; standard and Mac runners were unaffected. The incident recurred after earlier changes failed to address broader agent job health problems. GitHub began regularly recycling agents and improving detection while working on job health and agent resiliency.

## Source excerpt

Between Dec 3 03:35 UTC and 04:35 UTC, availability of large hosted runners for Actions was degraded due to failures in background VM provisioning jobs. This was a shorter recurrence of the issue that occurred the previous day. Users would see workflows queued waiting for a large runner. On average, 13.5% of all workflows requiring large runners over the incident time were affected, peaking at 46% of requests. Standard and Mac runners were not affected. Following the Dec 1 incident, we had disabled non-critical paths in the provisioning job and believed that would eliminate any impact while we understood and addressed the timeouts. Unfortunately, the timeouts were a symptom of broader job health issues, so those changes did not prevent this second occurrence the following day. We now understand that other jobs on these agents had issues that resulted in them hanging and consuming available job agent capacity. The reduced capacity led to saturation of the remaining agents and significant performance degradation in the running jobs. In addition to the immediate improvements shared in the previous incident summary, we immediately initiated regular recycles of all agents in this area while we continue to address the issues in both the jobs themselves and the resiliency of the agents. We also continue to improve our detection to ensure we are automatically detecting these delays.