New Multi-region uptime checks and custom-domain status pages

Buildkite Outage History

Daily status observations, past incidents, and reported issue history for Buildkite.

Checking current status...
79.1% of 91 observed days had no reported issue

90-Day Trend

May 28Aug 25

Monthly Status Summary

Month Issue-free days Days Tracked Days with Issues
August 2026 76% 25 6
July 2026 77.4% 31 7
June 2026 83.3% 30 5
May 2026 80% 5 1

This percentage summarizes normalized provider-status observations by calendar day. It is not duration-based, component-weighted, or contractual uptime. See the methodology and limitations.

Daily Status (Last 91 Days)

May 27 Today
Operational Degraded Partial Outage Major Outage Maintenance No Data

Incident History

August 2026
Buildkite service disruption
minor 73h 22m

Started:

Hosted Agents
monitoring
We have confirmed that all background workers have caught up as of 1:13am UTC. Hosted Agents dispatch has returned to normal. We'll continue to to monitor the service and are checking remaining retries.
investigating
We’re investigating elevated latency in background worker processing between approximately 12:39am and 1:02am UTC. Processing has since caught up and was near real time as of 1:13am UTC. We’re continuing to monitor the service and are checking remaining retries. Some builds or hosted-agent-related operations may have experienced delays during this period. We’ll provide another update as soon as we have confirmed the service is stable and have more information about the cause. We apolo...
investigating
We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.
Slow web UI
minor 154h 27m

Started:

monitoring
We have identified and fixed the issue. We are monitoring and seeing signs of improvement.
identified
We've identified the cause of the issue and are actively mitigating.
investigating
We've spotted that something has gone wrong and there's elevated latency. We're currently investigating the issue, and will provide an update soon.
Buildkite service disruption
minor 191h 25m

Started:

monitoring
We have reverted the change which introduced the elevated error rate. We believe the issue to be resolved but will continue to monitor.
identified
For a subset of customers, <1% of their Agent API traffic is experiencing issues. We are rolling back the change that caused this issue while we investigate it further.
identified
We're seeing increased latency and error rates for a subset of our customers. We're currently addressing the cause and will provide status updates as they become available.
investigating
We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.
Slow web UI
minor 219h 8m

Started:

monitoring
The integration between Pipelines and Test Engine has been re-enabled, restoring the Tests tab on build pages. We’re monitoring the recovery to confirm the integration remains stable. Test results continue to be available directly through Test Engine.
monitoring
The integration between Pipelines and Test Engine remains temporarily disabled,  and we're actively working to bring this back online. This affects the Tests tab on build pages. We’ll provide another update as recovery progresses. Test results are available directly through Test Engine
monitoring
We have identified and fixed the issue. We are monitoring and seeing signs of improvement. We've temporarily disabled the integration between Pipelines and Test Engine. This affects the Tests tab on build pages. We’ll provide another update as recovery progresses. Test results are available directly through Test Engine
identified
We've identified the cause of the issue and are actively mitigating.
investigating
Our Web UI is slow at the moment, we are investigating the issue.
investigating
We've spotted that something has gone wrong and there's elevated latency. We're currently investigating the issue, and will provide an update soon.
Latency on Pipelines
major 404h 58m

Started:

monitoring
The fix has been fully deployed and we are seeing signs of recovery. All events impacted by this issue will be automatically retried. We are continuing to monitor the fix for stability.
identified
We've identified that the impact is more widespread that we initially understood - scheduled builds as well as job dispatch. The fix that is rolling out now is expected to address all impact.
identified
We're experiencing latency with processing job dispatch, pipeline uploads and incoming webhooks. We have identified the cause and are rolling forward with a fix now.
investigating
We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.
July 2026
Delayed notifications
minor 533h 31m

Started:

identified
We are investigating delays to build and job notifications for a subset of customers.
June 2026
Increased latency and error rates
major 1506h 4m

Started:

monitoring
A fix has been deployed, services are recovering.
identified
We've identified the issue, we're continuing working on rolling out a fix. Impact is restricted to the builds create API for a subset of tenants.
identified
We're seeing increased latency and error rates for a subset of our customers. We're currently investigating and will provide status updates as they become available.
Increased latency on REST and GraphQL APIs
major 1658h 51m

Started:

monitoring
We've isolated the issue to elevated load on our REST API service and have mitigated the issue. The agent and stacks API isn’t affected.
monitoring
We've isolated the issue to elevated load on our REST API service and are working to mitigate. The agent and stacks API isn’t affected.
investigating
We're observing increased latency for all our customers. We're currently investigating and will provide status updates as they become available.
Increased latency and error rates
minor 1681h 36m

Started:

investigating
We're observing increased latency and error rates for a subset of our customers on the Agent API. We're currently investigating and will provide status updates as they become available.
May 2026
Delayed notifications
major 1997h 49m

Started:

identified
We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.
investigating
We are investigating delays to build and job notifications for a subset of customers.
Increased latency and error rates
2056h 12m

Started:

monitoring
We've identified the problem and have completed the remediation steps, we are now monitoring as service resumes.
identified
We're observing increased latency and error rates for a subset of our customers. We're currently remediating and will provide status updates as they become available.
Delayed notifications
major 2193h 29m

Started:

monitoring
We are seeing recovery across affected customers and continue to monitor
identified
We have identified the issue and applied mitigations and are monitoring recovery We have determined that only a subset of customers are affected by the notification latency.
investigating
We are investigating delays to notifications across all customers
Delayed Test Engine ingestion processing
minor 2323h 18m

Started:

monitoring
Ingestion of Test Engine execution data from an internal queue to a data store stalled, has been resumed, and is working through the backlog. Visibility of test executions from the past hour hours will be delayed for approximately a further one hour. This has been a recurring issue; an architectural change is coming soon to eliminate this failure mode.
Error rates increasing
2362h 54m

Started:

investigating
We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.
Delayed Test Engine ingestion processing
minor 2389h 9m

Started:

monitoring
We are currently experiencing delayed processing of Test Engine data. We have identified and applied a fix for the issue but are expecting to continue to experience delays while we clear the ingestion backlog
Delayed Test Engine ingestion processing
minor 2476h 59m

Started:

investigating
We are currently experiencing delayed processing of Test Engine data. We have identified and applied a fix for the issue but are expecting to continue to experience delays while we clear the ingestion backlog. At the current processing rate we expect the backlog to be cleared by approximately Sat 09 May 2026 00:00 UTC
Delayed Test Engine ingestion processing
minor 2485h 48m

Started:

monitoring
We are currently experiencing delayed processing of Test Engine data. We have identified and applied a fix for the issue but are expecting to continue to experience continued processing delays while we clear the ingestion backlog. At the current processing rate we expect the backlog to be cleared by approximately Fri 08 May 2026 13:30 UTC.
AWS us-east-1 single availability zone outage
minor 2496h 57m

Started:

monitoring
Despite the ongoing AWS incident, our own services are now stable. We are continuing to monitor our services closely, and are ready for further action should the need arise. We are also watching AWS services closely as they recover.
investigating
We are continuing to move infrastructure resources out of the affected AWS Availability Zone. Brief latency and error blips may continue while these manual failovers occur. (Apologies if you receive duplicated notifications for this update.)
investigating
We are continuing to move infrastructure resources out of the affected AWS Availability Zone. Brief latency and error blips may continue while these manual failovers occur.
investigating
We are continuing to move infrastructure resources out of the affected AWS Availability Zone. Brief latency and error blips will unfortunately continue while these manual failovers occur.
investigating
We are actively moving resources out of us-east-1c. Similar brief latency and error blips will be visible to customers while these manual failovers occur.
investigating
We have provisioned additional capacity in unaffected availability zones so that they are able to support the additional load. Automatic failovers continue to occur where necessary. Some latency and transient errors will be visible to customers.
investigating
We are continuing to actively monitor the impacts of this availability zone outage for Buildkite customers. Some transient errors are visible due to availability zone failover events.
investigating
A small subset of our customers are experiencing delayed notifications. We are actively provisioning additional capacity for these customers. Availability zone automatic failovers are occurring in response to the outage, and this is causing some brief error blips for some customers.
investigating
We're aware that AWS is reporting availability zone failures in us-east-1. We are monitoring the situation but so far there is no customer impact.
Delays in job dispatch, webhook processing, and outbound webhooks
major 2499h 24m

Started:

monitoring
We have now monitoring the incident. We are seeing most customers have recovered, and some showing signs of recovery.
identified
We've identified the issue and are working on applying mitigations. At this time we can confirm inbound and outbound webhooks, and notifications are delayed.
investigating
We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.
Jobs not starting on hosted agents and agent-stack-k8s
major 2513h 21m

Started:

identified
We're currently seeing recovery at 50% rate. We'll provide next update soon.
investigating
We've identified issue with job acquiring endpoint. We're rolling back now. We'll provide next update in ~20 minutes.
investigating
We've spotted that something has gone wrong. We're currently investigating the issue with new builds not starting.
Test Engine: Delayed processing of test result ingestion
minor 2542h 11m

Started:

monitoring
We've identified the issue and the system is currently processing the backlog of test executions
investigating
A process writing test results to our Test Engine data store stalled, we've restarted the process and are seeing it catching up. We expect to be fully caught up on the backlog within the next couple of hours.
Delayed notifications
minor 2577h 7m

Started:

monitoring
Applied remediations have resolved the previous notification delays affecting a subset of our customers. We're continuing to monitor the affected services for stability.
identified
We've identified the source of the notification delays affecting a subset of our customers. Our engineers are applying remediations to reduce these delays.
investigating
We are investigating delays with build and job notifications for a subset of customers.
Increased latency and error rates
2588h 6m

Started:

investigating
We're observing increased latency and error rates in the Agent API for a subset of our customers. We're currently investigating and will provide status updates as they become available.
April 2026
Increased latency and error rates
minor 2696h 25m

Started:

monitoring
We have identified and fixed the issue with the underlying database for a subset of customers. We are now monitoring the issue.
investigating
We're observing increased latency and error rates for a subset of our customers. We're currently investigating and will provide status updates as they become available.
Increased dispatch latency and error rates
minor 2720h 8m

Started:

monitoring
We have mitigated the issue causing increased Hosted Agents dispatch latency and intermittent timeout errors for a subset of customers. We identified abnormal workload activity that was placing elevated load on a supporting service, and have now blocked that activity and applied additional protections. Service metrics have returned to normal, and we are continuing to monitor closely.
identified
The issue has been identified and a fix is being implemented.
investigating
We're observing increased error rates and dispatch latency for a subset of our customers. We're currently investigating and will provide status updates as they become available.
Auth failures with remote MCP server
minor 2860h 50m

Started:

monitoring
We have rolled back a change on the remote MCP server that was contributing to authentication failures.
investigating
We are continuing to investigate errors when authenticating to the remote MCP server.
investigating
We are currently investigating reports of authentication failures with the remote MCP server.
Delayed processing of test execution
minor 2879h 37m

Started:

monitoring
We noticed a lag in data processing, but our systems are operational and currently working through the backlog. We expect to be fully caught up within the next couple of hours.
Degraded performance and increased error rates
major 3195h 43m

Started:

monitoring
We have identified and fixed the issue. We are monitoring and seeing signs of improvement.
investigating
We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon.
March 2026
Hosted Agents jobs immediately cancelled
minor 3402h 17m

Started:

identified
We have identified the issue and are rolling out a fix.
investigating
We have received reports from customers that they are unable to start builds on Hosted Agents. Their builds are immediately cancelled. We are investigating.
504 errors viewing builds
minor 3499h 7m

Started:

monitoring
The deploy to revert this change is complete and builds are loading normally. We will continue to monitor for any other issues.
identified
We've identified a change which we think is the cause of this issue, and we're in the process of reverting it.
investigating
We're seeing an increase in 504 errors when viewing pipeline builds. We're investigating this now.
Increased Delays with Hosted Agents
minor 3539h 43m

Started:

monitoring
The networking issue has been resolved, dispatch of Hosted Agents has returned to normal levels and no further issues with Git cloning. We are monitoring the situation.
identified
The issue has been identified to be related to Networking and affecting Git Mirror cloning.
investigating
We are currently investigating this issue.