New Multi-region uptime checks and custom-domain status pages

Astronomer Outage History

Daily status observations, past incidents, and reported issue history for Astronomer.

Checking current status...
84.6% of 91 observed days had no reported issue

90-Day Trend

May 28Aug 25

Monthly Status Summary

Month Issue-free days Days Tracked Days with Issues
August 2026 64% 25 9
July 2026 90.3% 31 3
June 2026 93.3% 30 2
May 2026 100% 5 0

This percentage summarizes normalized provider-status observations by calendar day. It is not duration-based, component-weighted, or contractual uptime. See the methodology and limitations.

Daily Status (Last 91 Days)

May 27 Today
Operational Degraded Partial Outage Major Outage Maintenance No Data

Incident History

August 2026
Incorrect scheduler heartbeat metrics being reported
major

Started:

investigating
Our metrics pipeline is falsely displaying spikes in scheduler heartbeat metrics, despite underlying component health.
Intermittent task instances failures.
major 45h 52m

Started:

monitoring
Fix is being rolled out across all astro deployments.
investigating
We have identified an issue affecting an internal rollout related to the `pre_execute` Airflow hook in Astro Hosted. We have identified a fix and are rolling it out across the fleet. We are monitoring the rollout and will provide further updates as it progresses.
KPO podspec changes for GPU-enabled nodes may cause KPO pod startup failures
minor 370h 58m

Started:

monitoring
We have rolled back the changes that triggered this issue and are monitoring the issue for confirmation that this issue is resolved.
identified
We've identified a change that enabled GPU support for KubernetesPodOperator pods had an unintended side effect when reading pod specs without proper limits set in the expected format. We're working on a fix, if you notice any unexpected failures for KPO tasks in your Astro deployments, please open a support ticket.
Intermittent Delays Managing Airflow Deployments in Azure Central US
minor 373h 11m

Started:

monitoring
Affected clusters have recovered and are operating normally. We are continuing to monitor the environment to ensure service stability.
investigating
We are investigating intermittent connectivity issues affecting clusters in Azure Central US when calling Astronomer APIs. This may prevent or delay attempts to create, update, or delete Airflow Deployments. Our engineering team is actively investigating the issue and working to restore normal service.
July 2026
Platform Log ingestion issue
minor 490h 6m

Started:

monitoring
We are continuing to see progress as the log injestions works through the backlog. Available logs are now about ~20mins old, and we expect to see that latency return to normal baseline in the next 2 hours.
monitoring
We continue to see log ingestion improving on the platform, and we're continuing to monitor the situation.
monitoring
We are beginning to see ingestion catch up with backlog log load. Customers may begin to see some logs becoming available in the Cloud UI
identified
We are continuing to investigate and monitor this issue.
identified
We have identified the source of increased log generation volume, which is causing issues on the ingestion side. We're implementing mitigations now and are monitoring the results.
investigating
We are continuing to investigate this issue, and will provide another update in 30mins or as new information is available.
investigating
We are currently investigating an issue with platform log ingestion for Astro. Logs may not be available in the Cloud UI.
Node Scaling Issue in Azure EastUS2 Region
minor 1035h 50m

Started:

investigating
We have detected node scaling issues in the Azure EastUS2 region and are currently investigating.
Node Scaling issues in AWS us-east-1 region
major 1070h 33m

Started:

monitoring
We are no longer seeing any impacting clusters in the AWS us-east-1 region. We're continuing to monitor the situation.
monitoring
AWS has now acknowledged a us-east-1 EC2 Launch Template API incident impacting EKS node provisioning/scaling (incl. Karpenter). We are continuing to monitor the situation.
investigating
AWS has published a status update for this issue. We are actively monitoring their progress and will continue to share updates. <https://health.aws.amazon.com/health/status?path=service-history>
investigating
We are investigating node scaling issues in the AWS us-east-1 region. Customers may experience issues with worker scale-up and new task startup.
June 2026
Azure clusters may experience issues with worker scale up and task pod scheduling
minor 1562h 42m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue, according to Azure status page ---- Investigating a spike in 401 authentication errors We are actively investigating customer alerts of 401 authentication errors when pulling images from mcr.microsoft.com. We've identified a potential underlying factor related to a recent update made to a backed service and we are actively working on determining mitigation workstreams. More information will be provided shortly. This message was last updated at 23:04 UTC on 15 June 2026
investigating
We are currently investigating this issue.
astro deploys not reflecting on certain Astro clusters
major 1569h 45m

Started:

monitoring
We are monitoring the affect of the applied fix and are observing error rates coming down/ceasing for affected clusters.
identified
We've identified an issue with `astro deploy`s not updating manifests in Astro Hosted clusters, and we're currently applying a workaround and are continuing to monitor the situation.
May 2026
Increased Task Failures in US-East-1 (AWS) - AWS Incident
minor 2490h 25m

Started:

monitoring
We are continuing to monitor the situation as AWS works toward full recovery in the affected Availability Zone. At this time, we are seeing signs of stabilization, and most transient failures should continue to self-resolve. We will provide further updates once AWS confirms that the issue has been fully resolved.
investigating
We are currently observing elevated task failures and latency for some deployments running in the AWS us-east-1 region. This is related to an ongoing AWS incident affecting a single Availability Zone (use1-az4), where EC2 and EBS resources have experienced impairments. What to expect: You may see intermittent task failures or retries in your deployments. In most cases, these failures are transient and should self-resolve automatically as AWS continues recovery. What you should do: No immedi...
Delayed delivery for some time-based alerts
minor 2511h 37m

Started:

monitoring
A fix has been implemented, and we are monitoring the results.
identified
A fix has been prepared and is moving through deployment. We will provide another update once the rollout is complete and validation is underway.
identified
We have identified the cause of the delayed delivery for some time-based alerts and are implementing a fix. We will provide another update once the fix has been deployed and validated.
investigating
We are investigating delayed delivery for some time-based alerts. The delay is primarily visible for DAG Timeliness alerts, and may also affect Observe-related SLA, Proactive SLA, and Data Quality monitor alerts. DAG Duration and Task Duration alerts do not appear to be affected at this time.
Deployments on Astronomer Runtime 3.2-3 are returning 403 errors when accessing dags page.
major 2526h 13m

Started:

investigating
Attempting to access the dags page in the Airflow UI results in a 403 Forbidden error. This should not be affecting task execution.
Dashboard cost breakdown data delays
minor 2583h 13m

Started:

investigating
We are continuing to investigate this issue.
investigating
We’re investigating an issue affecting Dashboard cost breakdown data. For affected customers, cost breakdown information may appear stale and may not have updated since May 1, 2026. Our team is actively investigating the cause and working to restore current data. We’ll share another update as we have more information.
April 2026
Airflow UI showing 403s for some customers
major 2740h 4m

Started:

investigating
We are continuing to investigate this issue.
investigating
It seems to be only in AWS clusters for now. We have a workaround we can apply while we investigate.
Azure East US multiservice outage impacting Astro deployments in the region
minor 2816h 40m

Started:

monitoring
Azure has fixed the issue, and services are back to normal in East US. We are monitoring to make sure everything stays stable.
identified
Azure East US has reported multi-service impact that is affecting Astro deployments in the region. For more information on Azure outage, please visit: https://azure.status.microsoft/en-us/status
Terminating Workers Accepting Tasks Causing Failures in Astro Executor
minor 2994h 4m

Started:

monitoring
The fix has been implemented, and we are now monitoring the deployments.
identified
We have identified the issue and we are rolling out the fix for the affected deployments.
investigating
We have identified an issue with Astro Executor deployments where terminating workers can continue to accept new tasks, which can, in some cases, can cause these tasks to fail.
Runtime 3.2-1 Yanked - Incompatible with Env Manager
major 3008h 4m

Started:

identified
We have identified that Runtime 3.2-1 is incompatible with the Astro Environment Manager. Any Connections or Variables stored at the Workspace level will not be available on deployments running 3.2-1. For this reason, we have disallowed the use of 3.2-1 for any deployments which are not already on that version. We are working to release a 3.2-2 version that is properly compatible as quickly as possible.
Degraded service for some deployments in Azure West Europe
3076h 25m

Started:

identified
We are currently investigating an issue affecting deployments in our Azure West Europe region. Some customer deployments are experiencing degraded performance due to a compute resource constraint in our shared infrastructure. Our engineering team has identified that the region has reached its vCPU quota limit for a specific compute type, which is preventing new resources from being provisioned. We have opened a high-priority support request with Azure to increase this quota and are activel...
Astro Alerts Degraded Performance
minor 3099h 12m

Started:

monitoring
We have applied a fix for the Astro Alerts degraded performance and currently monitoring it.
identified
We have identified the issue with Astro Alert's degraded performance and currently working on a fix.
investigating
We are currently investigating an issue causing false positive alerts via Astro Alerts. Our team is actively investigating the issue.
Deployment changes fail due to an unrelated "hibernation" error
major 3367h 37m

Started:

identified
We are continuing to work on a fix for this issue.
identified
The issue has been identified and a fix is being implemented.