New Multi-region uptime checks and custom-domain status pages
Services Pricing Dashboard

Fly.io Outage History

Daily status observations, past incidents, and reported issue history for Fly.io.

Checking current status...
50.5% of 91 observed days had no reported issue

90-Day Trend

May 26Aug 23

Monthly Status Summary

Month Issue-free days Days Tracked Days with Issues
August 2026 65.2% 23 8
July 2026 58.1% 31 13
June 2026 40% 30 18
May 2026 14.3% 7 6

This percentage summarizes normalized provider-status observations by calendar day. It is not duration-based, component-weighted, or contractual uptime. See the methodology and limitations.

Daily Status (Last 91 Days)

May 25 Today
Operational Degraded Partial Outage Major Outage Maintenance No Data

Incident History

August 2026
Network Issues in LAX Region
major 46m

Started:

Customer Applications
monitoring
Upstream networking issues have resolved.
investigating
We are investigating network issues in the Los Angeles region. Apps may experience higher latency or be unreachable at this time.
Oauth/Macaroon Errors from flyctl
25m

Started:

monitoring
A fix has been deployed and this error should no longer be occurring. We're monitoring to ensure full recovery.
identified
We have identified an issue causing authentication errors for some operations from `flyctl`. These operations are failing with an error like: `This endpoint no longer accepts legacy OAuth tokens (starting with `fo1_`). Please use a macaroon token (starting with `fm2_`) instead. We have identified the issue and are rolling out a fix
MPG (v1) partially down in ORD
major 34m

Started:

Management Plane - ORD
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are continuing to investigate this issue.
investigating
We had an issue with the ord-0 Fly Kubernetes cluster, and many MPG clusters are failing to restart. Our MPG team is actively working on it.
No capacity in ARN
60h 49m

Started:

Deployments
investigating
New machines may fail to create in ARN because we lack capacity.
Secrets service outage
major 126h 38m

Started:

monitoring
We have failed over the secrets database to a replica, and the Machines API appears healthy now. We are monitoring for any further issues.
identified
We are working to recover our secrets service after a failed deployment. Apps continue to run, but it is not possible to create new apps or update secrets at this time.
IPv6 Networking Issues
minor 152h 24m

Started:

monitoring
We are continuing to monitor for any further issues.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating degraded ipv6 networking on a subset of hosts
Increased app-not-found errors
major 263h 19m

Started:

monitoring
A fix has been implemented and we are monitoring the results
identified
We’ve deployed an additional mitigation to further reduce Corrosion retry pressure and are seeing improvement; we’re continuing to monitor while remaining affected nodes catch up.
identified
We’ve applied a mitigation to reduce the impact from Corrosion batch insertion retries and are continuing to monitor while affected nodes catch up.
identified
We have identified the issue as failed insertions in a subset of Corrosion batches. These failures trigger retries, which can cause timeouts for other batches.
investigating
We are currently investigating app-not-found errors returned by Machines API calls made shortly after creating new applications.
MPG IAD data plane degraded for new clusters
minor 341h 43m

Started:

identified
High CPU pressure on a shared etcd instance is causing lags on MPG creation in the IAD region
MPG creation is failing in GRU due to lack of capacity
major 367h 5m

Started:

monitoring
We tweaked hosts to allow for more machine allocation. We'll be monitoring the region over the next hours.
identified
New MPG clusters may fail to create in GRU because we lack capacity.
Certificate issuance delays
minor 371h 15m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
We have the size of TLS certificate issuance backlog under control, but are still seeing some remaining issues and are currently working to clean up the edge cases.
identified
We believe we have identified the issue and are releasing a fix.
investigating
We're currently investigating issues related to certificate issuance. Certificates may be delayed for new custom domains.
Managed Postgres v2 control plane issues in iad
minor 441h 7m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are investigating an issue with the control plane for Managed Postgres v2 in the IAD region. Creating new v2 clusters in the IAD region may fail at this time. Existing clusters continue to run, but may experience lagging backups.
July 2026
Capacity issues in CDG
465h 24m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
Creating machines in CDG region may fail at this time with an "no capacity available in cdg" message. Existing apps continue to run.
Outbound email issues
minor 468h 34m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
Our dashboard is failing to send outbound email. Emails such as new account verification or password reset may fail to send at this time.
Increased API latency
minor 517h 27m

Started:

investigating
We are investigating some database issues causing high latency on some API endpoints and dashboard operations. You may experience intermittent "503 service unavailable" errors at this time. Currently deployed apps continue to run.
High number of 5XX on the Machines API and dashboard
critical 738h 58m

Started:

monitoring
We are still working on fixing degraded Managed Postgres clusters.
monitoring
Some Managed Postgres v1 clusters are degraded. We are working on fixing them. Managed Postgres v2 is unaffected.
monitoring
A fix has been implemented and we are monitoring the results.
identified
We've identified an internal service providing authentication to our Machines API has failed, our team is currently looking at our options for restoring this service. Existing Machines/Apps will continue to run as normal. Thank you for your patience.
identified
We've identified an internal service providing authentication to our Machines API has failed, our team is currently looking at our options for restoring this service. Existing Machines/Apps will continue to run as normal. Thank you for your patience.
investigating
We are continuing to investigate this issue.
investigating
Existing machines are unaffected. We are investigating the issue.
Egress IPv6 issues in BOM
minor 766h 20m

Started:

identified
We have identified an upstream issue that is preventing egress IPv6 addresses in BOM from reaching parts of the internet, and we're currently working with an upstream provider to resolve this issue. Normal IPv6 addresses remain unaffected.
App creation failing
minor 830h 14m

Started:

identified
An issue with our Machines API is causing app creations to fail in some cases. We are working on a fix.
Edge proxy issues
minor 830h 35m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are investigating increased connection latency and "connection reset" errors from our edge proxy. Apps continue to run, but requests may experience increased connection latency or fail at this time.
Partial outage in SJC
major 875h 17m

Started:

monitoring
A fix has been implemented and we are monitoring the results. Apps should be reachable at this point in time.
identified
A subset of hosts in SJC are currently offline. Some apps may be unreachable at this time.
Some DFW hosts offline
minor 891h 43m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
A subset of hosts in DFW are currently offline, and we're investigating the issue.
Registry performance issues
minor 994h 35m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
The fly.io registry is currently experiencing capacity constraints that reduced performance and may lead to temporary high latency or failed pushes. We are currently working to add capacity and restore service.
Delays starting Depot Builders in IAD
minor 994h 48m

Started:

monitoring
A fix has been implemented and we are seeing improvements in builder performance, latency, and error rates. We are continuing to monitor for a full recovery. Customers still seeing issues can trigger a fly-hosted builder based deploy with `fly deploy --depot=false`.
identified
We have identified an issue causing delays or failures when starting depoting builders located in the IAD region. Customers with builders in IAD may see delays or timeouts starting builds during `fly deploy`. We are working on a fix. In the meantime customers can trigger a fly-hosted builder based deploy with `fly deploy --depot=false`. You can also change your builder region away from IAD via the `settings` tab of your Fly dashboard. We recommend DFW or ORD as alternate builder regions at ...
Partial Outage in ORD
major 1137h 9m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
We've identified the issue as a networking hardware failure impacting a subset of hosts at one of our Upstream providers in ORD. We are working with our provider to restore connectivity.
investigating
We are investigating an issue with one of our upstream providers in ORD. Machines across a subset of hosts may be unreachable or not running correctly. Deploys with machines on these hosts may fail at this time. Some Managed Postgres clusters in ORD region may be unavailable or see connectivity issues at this time.
Partial outage in ORD
major 1153h 58m

Started:

monitoring
Customer workloads are now starting and we're monitoring the affected hosts. Affected Managed Postgres instances will be investigated.
identified
Power restoration is ongoing and we're making sure the hosts are healthy before starting customer workloads to avoid issues. Customer impact remains and updates to come.
identified
Power restoration work in a subset of ORD is still in progress and impact remains ongoing for a subset of hosts and some Managed Postgres clusters.
identified
Our provider has advised us their facilities team is working on restoring the power, we'll provide another update as soon as we learn more.
identified
We've identified and reported power issues with one of our upstream providers in ORD. We're waiting for an update from our upstream for a resolution. Some Managed Postgres clusters in ORD will be unavailable due to placement.
investigating
We are investigating an issue with one of our upstream providers in ORD. Machines across a subset of hosts may be unreachable or not running correctly. Deploys with machines on these hosts may fail at this time. Some Managed Postgres clusters in ORD region may be unavailable at this time.
Errors issuing new SSL certificates
critical 1156h 30m

Started:

monitoring
A fix has been implemented upstream and certificates are being issued successfully. We will continue to monitor.
identified
The issue has been identified and we are awaiting a fix.
investigating
We are currently investigating errors when issuing new SSL certificates for hostnames.
Static Egress IPv6 issues in NRT
minor 1189h 2m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are investigating issues with static egress IPv6 addresses in NRT region. Apps using static egress IPs may experience connectivity failures to some destinations.
Elevated API Errors
major 1195h 54m

Started:

monitoring
Background jobs have caught up and the API is fully operational. We are continuing to monitor service health.
identified
A fix has been put in place, and we are no longer seeing elevated API errors. Some dashboard actions will be delayed while background processing catches up.
identified
The cause of the errors has been identified and we are working on a fix.
investigating
We are investigating elevated errors with our GraphQL API and background job processing
June 2026
Egress IP issues in SIN and NRT
minor 1212h 23m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
We are aware of egress IP issues in SIN and NRT and are working on a fix. Some machines in SIN and NRT using egress IPs may temporarily lose connectivity or otherwise see degraded performance.
Delayed Metrics
major 1249h 6m

Started:

identified
Almost all metrics have caught up aside from a small handful in `sin` and `syd`. We're continuing to monitor and expect these to complete in the next few hours.
identified
Backlogged metrics are still being processed. We're bringing extra processing capacity online to speed up the process.
identified
Backlogged metrics are still being processed and ingestion delays persist for some customers
identified
Backlogged metrics are still being processed and ingestion delays persist for some customers, but we’re continuing to see gradual improvement as the backlog continues to drain.
identified
Backlogged metrics data is still being processed, continue to monitor the situation while the system catches up.
identified
Metric ingestion is still delayed for some customers. We’re seeing gradual improvement and continue to monitor the situation while the system catches up.
identified
We are still in the process of increasing the metrics cluster throughput to catch up with metric backlog.
identified
We are continuing to see delayed metrics due to resource contention on a subset of metrics ingestion hosts, and we are working to rebalance ingestion traffic and reduce the backlog.
identified
We are continuing to see delayed metric exports from a number of hosts. Users will see delayed or missing metrics for machines on impacted hosts at this time.
identified
We've identified an issue on multiple hosts causing delayed metric ingestion into our hosted fly-metrics.net dashboards. We are working on a fix.
investigating
We are investigating issues with customer facing metrics in the fly-metrics.net dashboard. Users may see delayed or missing metrics at this time.
Metrics currently experiencing issues
major 1265h 58m

Started:

investigating
We are currently investigating an issue with our metrics cluster.
IPv6 Connectivity Issues in EWR
major 1299h 52m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
One of our upstream providers is experiencing IPv6 network connectivity problems in EWR. Apps with machines on affected hosts may have impacted connectivity to certain IPv6 destinations while they investigate and resolve this issue.
Deploys defaulting to Fly-hosted Builders
minor 1332h 27m

Started:

monitoring
We are seeing improvements in Depot builder provision times and are switching the default deploy strategy back to them. We will continue to monitor builder performance closely. Users with a preference can trigger a Fly builder based deploy with `fly deploy --depot=false` or use `fly deploy --depot=true` to force a depot-based deployment.
investigating
We are investigating delays provisioning Depot backed builders for deploys. We have switched the default `fly deploy` strategy to use fly hosted builders at this time. Users can still trigger a depot based deploy with `fly deploy --depot=true`
Elevated control plane latency
minor 1333h 58m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
investigating
We're addressing elevated control plane latency and saturation affecting the BOM and NRT regions. Apps with machines in this region might experience longer response times and possible timeouts (502 errors).
Degraded networking in North America
minor 1368h 27m

Started:

identified
Some 6PN Private Networking traffic remains impacted into and out of our LAX region, pending upstream resolution.
identified
Most networking is largely healthy between primary North American regions. Some Machines may see ongoing packet loss and higher latency communicating with other Machines on certain routes. We're continuing to monitor the backbone health upstream.
investigating
We are currently investigating degraded network performance between sites in NA due to an upstream incident
Network issues in SIN, NRT
major 1397h 53m

Started:

identified
Our upstream provider is continuing to experience network issues in SIN and NRT regions. Apps running in those regions may be unreachable or experience high packet loss at this time.
SIN, NRT network issues
major 1400h 59m

Started:

identified
We are continuing to experience network issues with an upstream provider in SIN and NRT regions.
monitoring
A fix has been implemented and we are monitoring the results.
identified
Our upstream provider has identified the issue and is working on a fix. Apps running in NRT (Tokyo) region may also have issues reaching certain destinations at this time.
investigating
We are investigating an upstream network issue in the SIN (Singapore) region. Apps may be unreachable or have higher packet loss.
Log search unavailable
minor 1472h 3m

Started:

monitoring
We've applied a fix for this issue. Historical logs are currently backfilling. We will post an update once logs have finished backfilling and current logs are being ingested normally.
investigating
Log search is available; however, new app logs since ~1 hour ago are missing and new logs are not being ingested. We are continuing to investigate.
investigating
We are investigating an issue causing application log search to be unavailable. This is affecting the Fly Metrics log search panels, and historical application logs initially returned from the `fly logs` command. Streaming logs using `fly logs`, the Live Logs page in the dashboard, and Fly Log Shipper services continue to work as expected.
Network Issues in SIN
major 1534h 48m

Started:

monitoring
Network connectivity in SIN has been fully restored. We're continuing to monitor.
identified
We are seeing recovery of network connectivity between SIN and most destinations. We're continuing to work with our upstream provider to resolve the remaining issues.
identified
Some machines in SIN are unreachable. A few Managed Postgres clusters may fail to fail-over or update. We are in the process of fixing this with our upstream provider.
investigating
We are currently investigating network connectivity issues in the SIN region. Hosted apps may be unavailable.
Macaroon Auth + Machines API Issues
critical 1570h 39m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
We have deployed another change and are seeing wider improvements in platform stability across all regions. Performance is trending to normal, though users may still see some degradation at this time. We are continuing to closely monitor to ensure full, stable recovery. We will provide another update in 15m.
identified
We are seeing elevated cluster errors with Managed Postgres clusters as the MPG control plane recovers from the API outage. MPG Users may see elevated rates of failing or slow connections, as well as increased primary/replica failovers. The managed postgres team is addressing any degraded clusters. We will provide a further update within 15m.
identified
We continue to seeing degraded performance and increased errors with the Machines API and other platform features at this time. We are continuing to work on fully restoring service.
identified
An initial fix has been deployed and we are starting to see platform features recover. Users may still see degraded performance and intermittent failures at this time. We are continuing to address the issue to ensure a full stable recovery.
identified
We have identified the cause of the issue and are working on deploying a fix. Impacted features remain unavailable or degraded at this time. Already running customer applications/machines remain available. MPG clusters remain generally reachable and healthy, however new clusters cannot be provisioned and failovers may not complete. We will provide another update within 15 minutes.
identified
We are continuing to address this issue. Platform authentication with macaroon based tokens is currently failing. Platform features that authenticate with macaroons including Machines API operations, Dashboard logins, some flyctl commands, fly-metrics.net Grafana, and deployments are failing at this time. Existing, running customer applications and machines remain reachable and running. We will provide another update within 15 minutes
investigating
We are investigating issues with Macaroon based authentication. This is impacting parts of the Machines API, Fly.io Dashboard, some flyctl operations and other platform features that rely on this.
MPG cluster provisioning is broken
1579h 48m

Started:

investigating
New MPG cluster provisioning is broken. Existing MPG clusters are not affected. Newly created organizations may see errors while SSHing into their machines. We are investigating the issue.
Elevated Sprites error rates in SIN
major 1655h 33m

Started:

monitoring
A fix has been implemented and we are seeing error rates for sprites in SIN normalize. We are continuing to monitor to ensure full recovery.
investigating
We are investigating elevated 500 / internal server error rates with Sprites in the SIN region. Users may see increased errors when accessing sprites located in this region, or for requests to the Sprites API originating from SIN
Increased network latency in North America
minor 1680h 53m

Started:

monitoring
Our upstream paths have been fixed. We are monitoring the results.
identified
We are working with our upstream network provider to address periodic loss of connectivity over transit in ord
identified
We are seeing a recurrence of elevated latency across some hosts in ORD impacting the MPG control plane and a subset of clusters there. We are working to address this.
monitoring
The network between our regions has been performing well for the majority of traffic. We're still continuing to monitor a few impacted routes in North America that may be seeing elevated latency and packet loss.
monitoring
Impacted NA backbones have been sidestepped where possible, and we're continuing to monitor network health.
monitoring
Managed Postgres in ORD has returned to normal operation. We continue to see slightly elevated latencies and loss over transits in North America. We are working with our upstream network providers to improve performance.
investigating
We are investigating elevated network instability across some hosts in ORD. Apps and managed postgres clusters on impacted hosts may see elevated latency or networking errors at this time.
Ingress Traffic issues in GRU
major 1686h 30m

Started:

identified
Some of our edge nodes in GRU has suffered an error that crashed some of the critical services. We're currently working to bring them back online. Some traffic entering through GRU (i.e. users connecting from around GRU) may be temporarily affected: connections may see increased latency or be occasionally dropped.
Emergency maintenance of Petsem causing some control plane errors
major 1711h 49m

Started:

monitoring
The maintenance has been completed and control plane functions should recover to normal. We're monitoring for any further complications.
identified
We're performing an emergency maintenance on Petsem, our secrets management service. Some control plane write operations may temporarily fail, for example, creating new apps or secrets. Existing apps and machines should keep functioning without issues.
Managed Postgres Control Plane Issues in IAD
major 1715h 11m

Started:

monitoring
An initial fix has been implemented and connectivity to all impacted clusters has been restored. We are continuing to monitor to ensure stable recovery.
identified
We are continuing to address this issue. Some clusters in IAD are unavailable at this time, some users may have seen unexpected cluter restarts. We are working on restoring normal performance for all clusters in IAD
investigating
We are investigating MPG control plane instability in a subset of the IAD region. A small number of clusters in the region may have seen unexpected failovers or connection issues over the past 30m.
egress ips are broken in ORD
minor 1720h 53m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
Egress ips are broken in most of ORD, we are currently investigating this issue
Capacity issues in ARN region
minor 1743h 37m

Started:

investigating
The ARN region is low on available host capacity. Creating new machines, or starting currently stopped/suspended machines, may fail at this time. We are working on provisioning new host capacity in the region. Please consider using nearby regions if possible.
Consul cluster degradation
1836h 25m

Started:

monitoring
We have restored the degraded Consul cluster and are monitoring for stability. All affected functionality should now be working correctly: Unmanaged Postgres and LiteFS with dynamic leases.
identified
One of our Consul clusters is in degraded state due to a failed node. This can cause issues with LiteFS primary node selection, Unmanaged Postgres (14.x and older *only*), and creation of new Unmanaged Postgres clusters. Impact is limited to these legacy products and does not affect deployments, running Fly applications in general, or Managed Postgres clusters.
Issues with flyctl ssh console and Machines OIDC
minor 1903h 36m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We're currently investigating an issue affecting flyctl ssh console functionality and machines' OIDC tokens.
May 2026
IPv6 outage for some machines in ORD
major 1939h 34m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We're working with our upstream providers to investigate an IPv6 networking failure in ORD.
Private networking issues in SYD
minor 1956h 40m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
Due to an upstream provider issue, Private Networking (6PN) is currently degraded in SYD region. Communication between Machines in SYD region and Machines in other regions may fail at this time. Newly created Machines in SYD may fail to sync to other regions (may not show up in Machines API List endpoint, or state may be incorrect). Additionally, TLS certificate resolution and Machines API authentication may currently be degraded in the SYD region. We are working with our upstream providers...
Elevated deployment errors
minor 1967h 26m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
We identified the issue and are working on a fix.
investigating
We're investigating an increase in deployment errors affecting some users. At this time, creating or updating Machines may erroneously fail with the message: "We require your billing information, please add it at https://fly.io/dashboard/<org>/billing".
Networking issues in ORD
minor 1974h 59m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
We are aware of increased latency and connection drops for clients located near Chicago (ORD) and are currently working on a fix.
Elevated latency for anycast ingress in ORD
minor 1987h 9m

Started:

identified
We are seeing elevated latency for requests hitting the ORD edges.
Networking issues in ORD
minor 1987h 9m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently investigating increased latency and dropped connections in ORD (Chicago).
Networking issues in ORD
minor 1993h 26m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently investigating increased latency and dropped connections in ORD (Chicago).
Increased latency in SJC
minor 1997h 0m

Started:

monitoring
We are continuing to monitor for any further issues.
monitoring
A fix has been implemented and we are monitoring the results.
identified
We're still seeing connection issues originating from SJC/LAX/West Coast US and are still investigating.
monitoring
We have implemented a mitigation for the issue and are monitoring for stability. This issue happened on the edge hosts in SJC, which means that any traffic near/around SJC would have been affected during the incident. We will provide a more detailed write-up on our Infra Log (https://fly.io/infra-log/) later once we have a better picture of the entire incident.
investigating
We are currently investigating increased latency in SJC.
Private networking issues in SYD
major 2030h 14m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
Due to an upstream provider issue, Private Networking (6PN) is currently degraded in SYD region. Communication between Machines in SYD region and Machines in other regions may fail at this time. Newly created Machines in SYD may fail to sync to other regions (may not show up in Machines API List endpoint, or state may be incorrect). Additionally, TLS certificate resolution and Machines API authentication may currently be degraded in the SYD region. We are working with our upstream providers...
Elevated GraphQL API Latency
minor 2038h 18m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are investigating elevated API Latency. Users may see delays or errors creating apps, as well as on some dashboard pages.
Capacity issues in EWR region
2048h 35m

Started:

identified
The EWR region is low on available host capacity. Creating new machines, or starting currently stopped/suspended machines, may fail at this time. We are working on provisioning new host capacity in the region. Please consider using nearby regions such as ord if possible.
Networking performance degraded in BOM and SJC
minor 2137h 8m

Started:

monitoring
Network performance has been restored and we're continuing to monitor.
investigating
We're currently looking into this issue.
IPv6 outage for some machines in SIN
minor 2148h 37m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating this issue.
Network issues in SIN region
major 2167h 38m

Started:

monitoring
Our upstream provider has implemented a fix and all apps and Managed Postgres clusters are now reachable. For the time being, ingress traffic is being re-routed to other regions, so users around the Singapore area may experience higher latency.
investigating
We are investigating network issues in the Singapore region. Apps may experience higher latency or be unreachable at this time. Some Managed Postgres clusters may be unreachable.
IPv6 outage for some machines in ORD
2174h 7m

Started:

investigating
We're working with our upstream providers to investigate an IPv6 networking failure in ORD. Impacted apps may wish to temporarily provision additional capacity in nearby regions.
Networking issues in SIN
minor 2202h 3m

Started:

identified
All MPGs in SIN are reachable again. We are seeing high latency to the affected provider. This may affect a subset of machines hosted in SIN.
identified
One of our upstreams is experiencing high packet loss and latency. We are actively working with them.
investigating
Machines may see high packet loss. Some MPGs are unable to connect to their config store and may be unreachable right now.
Networking issues with egress IP addresses in SYD
minor 2205h 3m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
We've identified the issue and are working on a fix.
investigating
We are currently investigating an issue affecting networking for new machines whose apps have assigned egress IP addresses in our SYD region
Issues with the Fly.io dashboard
major 2210h 50m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
We are continuing to work on a fix for this issue.
identified
The issue has been identified and a fix is being implemented.
investigating
We're currently investigating an issue where the Fly.io dashboard is failing to load in some cases.
Logs issues in IAD
major 2214h 53m

Started:

investigating
We are investigating an issue with logs and metrics in IAD region. New logs and metrics from machines in IAD region may be missing, but past logs/metrics are still accessible. Apps continue to run.
Proxy issues in SIN region
major 2222h 53m

Started:

monitoring
A fix has been implemented and we are seeing proxy performance in SIN return to normal. All Managed Postgres clusters in the region are reachable. We are continuing to monitor to ensure stable recovery.
investigating
We are investigating issues with fly-proxy on a subset of hosts in the Singapore region. Apps are still running, but requests to/from some apps may fail, and some Managed Postgres clusters may be inaccessible.
Some Managed Postgres clusters in FRA are unreachable
major 2293h 24m

Started:

monitoring
All affected clusters have recovered
identified
Some of unreachable clusters are showing recovery. We are still fixing the root cause.
identified
The issue has been identified and a fix is being implemented.
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating this issue.
fly ssh console returns error 500
2313h 40m

Started:

monitoring
We've deployed a fix and are monitoring as error rates normalize. `fly ssh console` should be working now.
identified
A problem with our vault used to issue temporary certificates for SSH sessions is causing calls to `fly ssh console` and `fly console` to return error 500. Our team has identified the cause and is deploying a fix.
log search unavailable
major 2406h 6m

Started:

monitoring
A fix has been implemented and we are monitoring the results. Logs should again be available through Log search in Grafana.
investigating
Log search in Grafana is currently unavailable. You may see `failed to make http request: 503` errors when accessing logs from fly-metrics.net at this time. App logs are still available using the `fly logs` command and in the Fly.io dashboard.
Upstash Redis Outage
major 2411h 32m

Started:

monitoring
A fix has been implemented and we are seeing Upstash Redis connectivity return to normal across all regions. We continuing to monitor to ensure stable recovery.
investigating
We are continuing to work with Upstash on this issue. We have received reports of partial recovery for some users, however we are still seeing higher levels of degraded or failing connections connecting to Upstash Redis databases at this time.
investigating
We are continuing to work with Upstash on this issue.
investigating
We are working with Upstash to investigate issues with their Fly hosted Redis service. Users may see degraded or failing connections connecting to their Upstash Redis databases at this time.
Certificate Issuance failing due to LetsEncrypt Outage
critical 2479h 24m

Started:

monitoring
We are seeing recovery and certificates are now issuing normally. We are continuing to monitor to ensure full recovery.
identified
Due to a service outage at LetsEncrypt, creating new certificates with `fly certs add` is failing. Existing certificates and `*.fly.dev` preview certificates are not impacted. For additional details please see LetsEncrypt's statuspage https://letsencrypt.status.io/pages/incident/55957a99e800baa4470002da/69fe2d6698ca07050eb4b1b3
Connectivity issues in SJC
major 2504h 19m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
Some hosts in SJC are currently experiencing an upstream network issue. Apps running on these hosts may be temporarily unavailable.
Intermittent machines issues in BOM
minor 2511h 52m

Started:

identified
Creation and updating of machines in BOM are affected. Some metrics and logs for resources in BOM may be delayed.
Elevated error rates on List Machines endpoint
minor 2546h 2m

Started:

investigating
We are currently investigating this issue.
Errors Setting and Updating Secrets on Apps
major 2556h 50m

Started:

investigating
Creation of new apps or changing secrets on existing apps fails
investigating
Creation of new apps or changing secrets on existing apps fails
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating this issue.
Log search unavailable
minor 2575h 12m

Started:

monitoring
We have a mitigation in place and are monitoring results.
investigating
Log search in Grafana is currently unavailable. You may see `failed to make http request: 502` errors when accessing logs from fly-metrics.net at this time. App logs continue to be available using the `fly logs` command and in the Fly.io dashboard.
April 2026
flyctl deploy creating new app instances
minor 2714h 18m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
investigating
We're investigating an issue where fly deploy is creating new Fly machine instances rather than updating existing ones, leading to apps with a mixed state. We're currently investigating the issue. As a workaround, please try removing the "processes = [ "app" ]" line from your fly.toml configuration file and redeploying. Another workaround is to downgrade flyctl to 0.4.40 - this should resolve the issue in the meantime.
Slow machines operations in IAD region
minor 2811h 23m

Started:

monitoring
Network packet loss has returned to normal levels. We are monitoring the Machines API for stability.
investigating
We are continuing to investigate this issue.
investigating
We are deploying a partial mitigation while we continue investigating.
investigating
We are currently investigating the issue. Only a portion of machines within the region are impacted.
Errors when adding or editing Github integrations for deployments
major 2843h 3m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
We are continuing to work on a fix for this issue.
identified
The issue has been identified and a fix is being implemented.
investigating
We're investigating reports of "500" errors when trying to add a new Github integration or edit an existing Github integration in Fly.io/dashboard. This only affects "Launch an app from Github" or trying to change settings for an app set up this way. Existing integrations continue to work normally. It does not affect deploys done with `flyctl` or existing, running apps.
Errors (5xx, timeouts) in Fly.io dashboard
major 2846h 51m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
investigating
We are investigating issues with web dashboard.
Increased latency in SIN
minor 2915h 40m

Started:

identified
We are currently working on resolving increased latencies in our Singapore region.
TLS certificate issues
major 2989h 3m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are investigating an issue with the Vault server that stores TLS certificates. Provisioning new TLS certificates may fail, and connecting to domains whose existing certificate has not yet been cached may fail.
Network issues in SYD
3039h 1m

Started:

monitoring
We've identified the issue and applied a fix. All services should be working as normal.
investigating
We're currently investigating some networking issues in SYD. This is affecting a number of our central services.
Heightened latency in ORD
3103h 18m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently investigating heightened network latency in ORD.
Managed Postgres control plane instability in NRT (Tokyo)
minor 3151h 26m

Started:

monitoring
A fix has been implemented and we are seeing MPG performance in NRT normalize. We are continuing to monitor to ensure a stable recovery
identified
The issue has been identified and a fix is being implemented. Users with clusters in NRT may continue to see instability at this time
investigating
We are investigating instability in the MPG control plane in the NRT (Toyko, Japan) region causing unexpected cluster failovers. Clusters return to health shortly after, but some users with clusters in NRT may see dropped connections or degraded performance at this time.
Unavailable hosts in ORD region
major 3174h 40m

Started:

investigating
Some hosts in our Chicago (ORD) region are currently inaccessible. We are working with our provider to resolve this issue. To see if you are affected, please visit the personalized status page: https://fly.io/status A small amount of Managed Postgres clusters may also be inaccessible at this time.
Managed Postgres Control Plane Issues in SYD
major 3190h 19m

Started:

identified
We are seeing an improvement in control plane performance in the SYD region. Some clusters in the region currently are showing degraded standby nodes and we are working to bring those back to full health.
investigating
We are investigating elevated control plane issues for Managed Postgres clusters in SYD. The majority of clusters appear to be running fine, but new creates, backup restores, and upgrades may show errors or take longer than usual to complete. Some clusters will have seen a failover event from primary to standby.
Metrics currently experiencing issues
major 3209h 34m

Started:

monitoring
We are continuing to monitor for any further issues.
monitoring
We have implemented a fix. We're monitoring the cluster for further issues.
investigating
We are currently investigating an issue with our metrics cluster.
GraphQL API / Dashboard Issues
critical 3227h 0m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
We have restored GraphQL and dashboard availability, but some actions (e.g. app state updates) may still be delayed.
investigating
We are investigating issues with our GraphQL API and web dashboard
March 2026
Low Capacity in SIN and AMS regions
3443h 9m

Started:

monitoring
We've freed up additional room in the SIN and AMS regions and are monitoring capacity.
monitoring
We've freed up additional room in the SIN and AMS regions and are monitoring capacity.
identified
We are currently investigating capacity issues in SIN and AMS regions that are affecting: - Machine Create and Start events - Deployments, due to affected, degraded Remote Builders - Sprite startup from cold state
identified
This may also affect: - Remote builders in AMS and SIN regions, which could currently be experiencing degraded performance or failures. - Sprites starting from a cold state, which may experience failures in starting
identified
We are currently investigating elevated errors when creating and starting machines in the SIN and AMS regions. Choosing other regions to create or deploy may help in the meantime
Low capacity in IAD
minor 3488h 0m

Started:

monitoring
With the additional capacity we've brought online, machine start failure rates in IAD have now recovered. We'll continue to monitor IAD capacity.
identified
We've brought some additional capacity online in IAD and are seeing improvements, and we're continuing to work on adding more and freeing up additional room.
investigating
We're continuing to evaluate our options for increasing short-term capacity in the IAD region.
investigating
We're currently investigating capacity issues in IAD that is preventing machine starts (machine creates are currently unaffected). This may result in deploys failing to complete (even for apps outside of the IAD region). As a workaround, using legacy Fly builders explicitly located in another region (i.e., `FLY_REMOTE_BUILDER_REGION=lhr fly deploy --depot=false --recreate-builder`) may help in the meantime.
Machine Creates Failing in ORD Region
major 3514h 47m

Started:

monitoring
We've implemented a fix and have seen error rates for machine creates in ORD drop off. We're continuing to monitor the results.
identified
We've identified the cause of this increased failure rate and a fix is in progress. We are seeing most creates in ORD succeed at this time, though failure rate is still above baseline.
investigating
We are continuing to investigate this issue. We are seeing 408 errors decreasing in ORD, though still above baseline.
investigating
We are currently investigating elevated errors creating machines in the ORD (Chicago, Illinois) region. Users may see `failed to launch VM: request returned non-2xx status: 408` errors when creating, updating, or scaling machines in ORD. Existing, already running machines in the ORD region continue to run as normal.
Network issues in FRA region
critical 3517h 32m

Started:

identified
Some Managed Postgres clusters in FRA region are still unreachable, we are investigating this issue.
monitoring
Apps and Managed Postgres clusters in FRA region should be back online at this time. We are monitoring for any further issues.
investigating
We are investigating network issues in FRA region. Apps and/or Managed Postgres clusters in the region may be inaccessible at this time.
Backend errors when trying to use Grafana to view logs
3586h 51m

Started:

monitoring
We've deployed a fix and are monitoring the results. Logs are now be visible on Grafana.
identified
Using the Logs panel in Grafana at https://fly-metrics.net/ will show a 502 error from the backend and won't show any logs. You can use `fly logs` or the live log viewer directly on https://fly.io/dashboard to view streaming logs for the time being.
investigating
Using the Logs panel in Grafana at https://fly-metrics.net/ will show a 502 error from the backend and won't show any logs. You can use `fly logs` or the live log viewer directly on https://fly.io/dashboard to view streaming logs for the time being.
Machines failing to start in DFW
minor 3666h 43m

Started:

monitoring
Machine start success rates in DFW have improved but we are continuing to monitor and make further adjustments. We will provide updates as the situation progresses.
monitoring
In addition to freeing up existing capacity, the team has provisioned new capacity in DFW and we are monitoring the results.
monitoring
We freed up some capacity on our workers to allow for successful Machine starts.
investigating
The Machines start failure rate is elevated in DFW.
Metrics currently experiencing issues
critical 3691h 40m

Started:

monitoring
We have implemented a fix. There has been approximately 1h of lost metrics from 06:07UTC. We're monitoring the cluster for further issues
investigating
We are currently investigating an issue with our metrics cluster.
IPv6 networking issues in SJC region
major 3705h 56m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are investigating intermittent network issues in SJC region impacting outbound public IPv6 access from Machines. Connecting to IPv6 internet resources from apps hosted in SJC region may be slow or fail at this time. IPv4 access, as well as 6PN private networking, are unaffected.
Fly ssh console command failing
minor 3707h 56m

Started:

identified
We have identified an issue causing new `fly ssh console` connections to fail with 500 errors. A fix is in progress.
Connection Issues in SJC
minor 3708h 2m

Started:

monitoring
Between 13:55 and 14:03 UTC machines and MPG clusters hosted in the SJC region saw elevated connection errors. Users may have seen errors connecting to or from most machines in the region, as well as with deployments or updates to machines in the region. Networking has returned to normal in the region, and we are continuing to monitor closely to ensure stable recovery.
Machines failing to start in DFW
major 3712h 11m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
The team is currently rolling out additional capacity in DFW which should help ease Machine start failures across the region.
investigating
We are investigating reports of machines failing to start in the DFW (Dallas) region with "insufficient memory" errors. This may cause deployment failures for applications running in DFW. Our team is actively working to restore full capacity in the region. If you are affected, deploying to an alternate region may serve as a temporary workaround. We will provide updates as the situation progresses.
Elevated 502 errors when starting Sprites in LAX and ORD
minor 3748h 9m

Started:

investigating
We're currently investigating an elevated number of 502 errors when attempting to start Sprites in LAX and ORD.
Sprite Operations: 401 errors for certain organizations
3804h 36m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
Organizations with numerical prefixes might experience failing sprite operations ( like creating a sprite, listing sprites, etc... ) due to 401 errors
monitoring
Root cause has been identified and a fix has been applied
investigating
Organizations with numerical prefixes might experience failing sprite operations ( like creating a sprite, listing sprites, etc... ) due to 401 errors
monitoring
Organizations with numerical prefixes might experience failing sprite operations ( like creating a sprite, listing sprites, etc... ) due to 401 errors
Sprites Operations: 401 errors for certain organizations
3805h 38m

Started:

monitoring
Organizations with names prefixed with numerical digits may experience 401 errors. Affected operations include actions such as Sprite creation, listing, etc... A fix has been implemented and we are monitoring the results!