New Multi-region uptime checks and custom-domain status pages

Grafana Cloud Outage History

Daily status observations, past incidents, and reported issue history for Grafana Cloud.

Checking current status...
18.7% of 91 observed days had no reported issue

90-Day Trend

May 28Aug 25

Monthly Status Summary

Month Issue-free days Days Tracked Days with Issues
August 2026 40% 25 15
July 2026 19.4% 31 25
June 2026 3.3% 30 29
May 2026 0% 5 5

This percentage summarizes normalized provider-status observations by calendar day. It is not duration-based, component-weighted, or contractual uptime. See the methodology and limitations.

Daily Status (Last 91 Days)

May 27 Today
Operational Degraded Partial Outage Major Outage Maintenance No Data

Incident History

August 2026
K6 Test Outage
critical 154h 44m

Started:

monitoring
Telemetry push APIs went down at 15:06 UTC, causing tests to lose telemetry data, like metrics, logs, traces, and browser screenshots. This issue has now been mitigated, and we are monitoring to determine root cause as well as full resolution.
Some Grafana Instances Unavailable
critical 366h 58m

Started:

monitoring
We have rolled out a change to the affected regions to restore services.
investigating
We have identified the cause: our cloud provider has exhausted compute capacity in the affected region. We are working directly with them and are moving affected workloads to alternative capacity to restore service.
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating an issue that causing some Grafana deployments in prod-us-central-0 and prod-us-central-3 to become unavailable.
Logs latency increase within prod-eu-west-3
minor 375h 35m

Started:

investigating
There's still a minor latency increase in this cell however we have identified that it is improving although not back to normal as of yet. We will continue to look into this.
investigating
We have identified minor latency increases in write endpoints within the prod-eu-west-3. Our team is currently monitoring and investigating this.
K6 - Cloud test-run issues
minor 450h 43m

Started:

monitoring
A critical piece of infrastructure behaved poorly after a reboot. It has now been properly recovered
investigating
We are currently investigating an issue which is resulting in some test runs to abort and metrics data to be incomplete for a subset of tests
July 2026
Degraded Performance: Stack Provisioning Failures within certain reigons (PDC Setup)
minor 471h 17m

Started:

investigating
We are currently investigating an issue impacting stack provisioning. Attempting to set up PDC on newly created stacks will currently fail across several regions. We are currently looking into the cause of this.
PDC Authentication Issues
major 491h 23m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are continuing to investigate this issue.
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating an issue impacting PDC authentication. We will provide additional updates as they become available.
Issues with Billing/Usage Dashboard Metrics and Panels.
major 492h 53m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
We have identified the issue, and are working on deploying a fix.
investigating
We are investigating an outage for the Billing / Usage dashboard metrics and panels. This appears to be partially affecting organizations using Grafana Cloud. We are working on identifying and resolving the issue
IRM Performance Degradation in EU Region
minor 508h 22m

Started:

identified
We have identified the underlying issue affecting Grafana IRM in our EU region and are actively working to restore normal service. Users may continue to experience a degraded or unresponsive IRM UI, reduced availability of the public API, and delays in notifications while mitigation efforts are underway. Our engineering team is working to restore full functionality as quickly as possible. We will provide another update as more information becomes available.
investigating
We are currently investigating an issue affecting Grafana IRM in our EU region. Users may experience a degraded or unresponsive IRM UI, reduced availability of the public API, and delays in notifications. Our engineering team is actively investigating the issue and working to restore normal service. We will provide another update as more information becomes available.
Partial OTLP Write Outage in prod-us-east-3
major 534h 29m

Started:

investigating
We are investigating an issue affecting OTLP ingestion in the prod-us-east-3 region. Customers may experience intermittent failures when writing telemetry to the OTLP endpoint. Based on current information, logs are confirmed to be affected, and metrics and traces may also be impacted. Our investigation indicates intermittent failures began around 08:00 UTC, with the primary period of impact occurring between 15:30 UTC and 17:45 UTC. We are continuing to investigate the scope and root cause ...
Write Outage
critical 633h 7m

Started:

monitoring
Healthy as of 17:00 UTC. We are continuing to monitor.
investigating
We are currently investigating a Write outage ongoing, starting at 16:30 UTC for the impacted region.
Errors creating new Slack integration for Grafana IRM
major 640h 51m

Started:

identified
We have verified the fix, and we are starting to roll it out
identified
We have identified the root cause of the problem, and we are working on a fix
investigating
We are investigating a possible error creating Slack integrations. Will update the status soon.
Intermittent write outage from 21:20-21:26, 21:54-21:56, and 22:22-22:23
minor 651h 8m

Started:

monitoring
status is health in prod-us-east-3.loki-prod-042 and we're continuing to monitor
Metrics Write Path Errors
major 656h 55m

Started:

monitoring
Error rates have dropped, and we are monitoring this issue for any recurrence.
investigating
We are currently investigating an elevated rate of error and latency in the impacted write path.
Stacks using SCIM user provisioning are currently unable to log into Grafana
major 663h 56m

Started:

identified
We have applied a fix and are currently awaiting feedback from affected customers.
identified
The issue has been identified and we are currently working on the mitigations now. The scope is much more limited than initially thought only specific SCIM configurations are impacted.
investigating
We are currently facing an issue where Stacks using SCIM user provisioning are currently unable to log into Grafana. We are currently investigating this issue and working on a fix.
K6 - Cloud output test-runs are failing to fetch the script logs
minor 687h 41m

Started:

monitoring
We have identified the cause and applied a fix which has has resulted in significant recovery. We will continue to monitor this before resolving.
investigating
We are currently investigating an issue cloud output test-runs which is resulting in failure to fetch the script logs. We are working on identifying the cause and working on a fix.
Cloud Log Exporter Unavailable
major 703h 25m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
Cloud Log Exporter is temporarily not available in the marked regions. We are currently investigating this issue and will update as soon as we have more info to share.
PDC Issues
major 732h 20m

Started:

monitoring
A fix has been implemented, and we are observing recovery. We will continue to monitor the results.
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating an issue that is causing issues with PDC in the prod-eu-west-2 region. We will provide another update in 1-2 hours.
Grafana Cloud login issues for users with role of None
minor 799h 46m

Started:

identified
We are in the process of implementing a fix and are monitoring the rollout. Thank you for your patience.
identified
We’ve identified the cause of the issue impacting users with the role of None from logging in. Our team is currently implementing a fix.
investigating
We’re currently investigating an issue with user logins with the None role in Grafana Cloud. Our team is actively working to identify the cause. Thank you for your patience.
Adaptive Metrics aggregation delay in eu-west-0 region.
minor 829h 13m

Started:

monitoring
Services are fully recovered now. We're monitoring to be sure the issue won't re-occur.
identified
We're currently facing an issue with Adaptive Metrics aggregation delay in the eu-west-0 region (GCP Belgium). The issue started at around 11:50 UTC, but we were able to identify the issue and the appropriate fix is already deployed. We're seeing services recovering. More updated to come soon.
Partial Outage in prod-eu-west-2
major 843h 44m

Started:

monitoring
We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
investigating
We are continuing to investigate this issue.
investigating
We are investigating a broader issue affecting multiple Grafana Cloud products in prod-eu-west-2. This appears to be caused by a third-party provider issue rather than a Loki-specific problem. Impact is currently inconsistent: some components are affected while others continue to function normally. We have also seen some impact to Tempo write paths. We are continuing to investigate and will share another update as soon as we have more information.
investigating
We are investigating an issue affecting Loki queries in prod-eu-west-2. We first observed this behavior at approximately 21:57 UTC. Affected users may see elevated query errors, timeouts, or intermittent failures when running Loki queries in this region. The issue appears to be improving, but it is not fully resolved yet. We are continuing to investigate and will share another update as soon as we have more information.
Delayed Aggregated Metrics (prod-us-central-0)
minor 846h 50m

Started:

monitoring
We have applied the mitigation and are monitoring as the backlog catches up. We will post another update once that is complete. Thank you for your patience.
identified
We are currently investigating a delay in aggregated metric results for tenants using Adaptive Metrics in the prod-us-central-0 region. Our engineering team has identified the issue and is actively working on a mitigation. We will provide further updates as the investigation progresses.
Some Reports of Grafana Not Loading.
major 848h 1m

Started:

identified
The fix is in the process of being rolled out.
identified
We believe we have found the cause and are working on remediation. It is also worth mentioning that only stacks on the "slow" release channel are impacted.
investigating
We are currently investigating an issue impacting a small subset of stacks in the impacted regions. We will provide more details as they become available.
Mimir Write Performance Degradation
minor 869h 41m

Started:

monitoring
From 19:24 until 19:40 UTC, backend performance degradation impacted ingestion.  We are currently monitoring.
Mimir Partial Write Outage
major 872h 59m

Started:

monitoring
The outage is recovered as of 16:56 UTC, and we're continuing to monitor.
identified
We're investigating backend degradation which has resulted in a partial write outage beginning around 16:40 UTC.  This degradation has also impacted reads and rule evaluation.  Issue has been identified and we are working on mitigation.
Issue with Dashboard Views Being Registered
minor 881h 58m

Started:

identified
We are currently working on a fix related to this incident. There are no new updates to share at this time.
identified
We continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.
identified
We continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.
identified
We continue to work on resolving this issue. We'll provide another update within the next 24 hours, or sooner if additional information becomes available.
identified
We continue to work on resolving this issue. We'll provide another update within the next 24 hours, or sooner if additional information becomes available.
identified
We've identified the cause of the issue impacting dashboard view and error counts in the Dashboards and Folder list. Our team is currently working on implementing a fix and validating the solution. We will provide another update within the next 24 hours, or sooner if we have additional information to share.
investigating
This is related to https://status.grafana.com/incidents/rhrk2ck6ly0y which was resolved by mistake. For some dashboards, the Views/Error counts in the Dashboards/Folder list renders as `-` and never updates, even after dashboards are viewed repeatedly. This is ultimately causing inaccurate or missing view counts. We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide ...
IRM Mobile App forcing some users to logout
minor 965h 41m

Started:

monitoring
A new iOS mobile app version (v2.39.5) has been released, which includes a fix for this issue. We are monitoring the rollout and its impact to ensure the issue has been fully resolved. If you continue to experience any problems after updating to the latest version, please let us know.
identified
A fix has been submitted for review and will be released as Grafana Mobile v2.39.5 once it is approved. If your app is currently on v2.39.3, we recommend not upgrading to v2.39.4 and instead waiting for v2.39.5 to become available. Users who have already been signed out by this issue can log back into the app and continue using it. We will provide another update once v2.39.5 is available.
investigating
Our investigation has narrowed the impact to version 2.39.4 of the Grafana mobile app on iOS. Users should avoid upgrading to this version until an updated release is available. As a precaution, we also recommend ensuring you have alternative notification methods configured (such as SMS, email, or phone calls) if you rely on mobile push notifications for alerting. We are actively working on a resolution and will provide another update as soon as more information is available.
investigating
We are aware of an issue in the latest mobile app release that is causing some users to be signed out and asked to log in again. We have reproduced the behavior and are investigating and working on a fix. We will share another update as soon as we have more information.
Issue with Dashboard Views Being Registered
minor 970h 11m

Started:

investigating
We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soon as more information becomes available.
investigating
For some dashboards, the Views/Error counts in the Dashboards/Folder list renders as `-` and never updates, even after dashboards are viewed repeatedly This is ultimately causing inaccurate or missing view counts.
PDC Degraded performance
critical 998h 32m

Started:

monitoring
We are now seeing recovery of PDC traffic to the affected customers. Our team will continue to monitor across shifts.
identified
We are continuing to see disruptions across multiple deployments, ranging from degraded performance to full outages for some customers. Additional resources have been engaged to mitigate this issue. We will post updates as they become available.
monitoring
We are continuing to monitor for any further issues.
monitoring
We are seeing recovery in some deployments, while others not just yet. We are continuing to monitor this incident.
monitoring
A fix has been applied and we are seeing recovery. We will continue to monitor this.
investigating
We are still currently investigating the issue.
investigating
We are currently facing performance degradation on PDC service hosted on Multiple clusters. Our Engineering Team is currently working on fixing the issue, we do apologize for any inconvenience.
Delayed ingestion and recording rule evaluation failures for Mimir in prod-ap-south-1
minor 999h 41m

Started:

monitoring
A fix has been implemented. We are currently monitoring the results.
investigating
We are observing delayed ingestion and recording rule evaluation failures for Mimir in prod-ap-south-1. As of yet we have not noticed any customer impact however we are currently observing the cell.
Grafana rulers crash-looping on prometheus
minor 1021h 52m

Started:

monitoring
We are in the process of rolling out the fix. Our engineers are monitoring the progress.
investigating
We are investigating issues with the grafana-ruler service on prod-us-east-2 and prod-us-west-0 which are causing periodic crash conditions. A code fix is currently being deployed to mitigate this
Cannot access dashboard set up as a home page
minor 1023h 41m

Started:

investigating
We are currently investigating an issue affecting some customers who are unable to load or access their dashboards when configured as their home page. We will update once we have more information on this.
Grafana Cloud IRM alert groups failing to produce alerts in us-east-3 region.
major 1094h 31m

Started:

monitoring
The work on bringing services back to healthy state is completed and alerts should be created correctly, without a delay now. We're keeping the incident open and monitoring for any potential hiccups that might occur.
identified
Some IRM alert groups in us-east-3 region may not be able to produce alerts and could be sluggish/misbehaving in general. The issue started around 01:00 UTC on July 3rd. The issue has been identified and the cause of the issue has been fixed. Our team is actively working on getting the service back to healthy state.
Some Queries Failing
minor 1159h 58m

Started:

identified
We are continuing to work on rolling back a PR responsible for this behavior. Once we have more information, we will share it here. Thank you for your patience.
identified
Queries with drop __error__, including log volume histogram queries (the queries that generate the histogram visualization in Grafana), are failing due to a bug with series limit checks. The Root cause has been identified, and we are working on a fix.
Loki and Frontend Observability - Major Outage in prod-us-central-0 region
critical 1173h 13m

Started:

investigating
The AWS Logs integration in the same region is affected as well. We will provide further updates as our investigation progresses.
investigating
We are currently investigating a major outage in Loki writes and Frontend Observability in the prod-us-central-0 region. Our Engineering team is investigating this and we will provide further updates as our investigation progresses.
Elevated Loki Query Bytes Reporting
minor 1182h 26m

Started:

investigating
We are investigating an issue where some customers may see higher Loki query byte usage reported than was actually consumed. This affects usage reporting only; there is no impact to query execution or service availability. The issue began at approximately 13:20 UTC and is ongoing. We expect the issue to be resolved soon and will provide another update as more information becomes available.
June 2026
Confluent API Outage
major 1237h 31m

Started:

investigating
We are investigating an issue affecting Confluent metrics ingestion across all regions. Due to an elevated error rate on the Confluent side, some metrics may not be ingested, resulting in potential data loss. We are actively investigating the issue and will provide updates as more information becomes available.
Mimir read errors and high latency in prod-eu-west-0
minor 1238h 31m

Started:

identified
We've identified a possible cause, and a mitigation is in place to prevent further occurrences.
investigating
The errors and latency have now recovered, we continue investigating the root cause.
investigating
The errors are recovering, and we are still looking into the root cause of this.
investigating
We are currently investigating an issue with Mimir in prod-eu-west-0 we are seeing read errors and high latency. This incident is currently ongoing. The errors are recovering but we are currently looking into the route cause of this.
Rule evaluation error on cluster prod-gb-south-0
minor 1239h 35m

Started:

monitoring
A fix has been applied and we are currently monitoring results.
investigating
We are currently investigating Rule Evaluation errors on the cluster prod-gb-south-0 which is leading to error codes showing within the stacks. We are looking into the issue and will update accordingly.
K6 - Test run metrics processing is delayed
minor 1309h 0m

Started:

investigating
We've improved the metric ingestion delay time and are working on additional fixes to bring it down to expected range. Customers can currently expect a delay of 2 to 5 minutes before their test run metrics show up (after starting a test).
investigating
Update: Changed incident title to "Test run metrics processing is delayed" We have found the issue and are working on deploying the fix.
investigating
We are experiencing intermittent delays with secondary metrics processing for k6 Cloud test runs due to heavy load. We don't expect any data loss or impact on user runs, but results may take longer time to appear in UI.
investigating
A small update: The issue is isolated to new test runs, and users can go see the metrics of all the previous test runs
investigating
We’re currently investigating an issue causing metrics not to appear during test runs. . Our team is actively working to identify the cause. Thank you for your patience.
Rule Evaluation Outage in prod-us-central-0
major 1473h 30m

Started:

monitoring
We’re continuing to track progress post-mitigation. While we don’t have new information to share yet, our team remains actively engaged.
monitoring
We had an outage affecting rule evaluations between 15:16-15:59 UTC in the prod-us-central-0 region. Our team quickly identified the issue and has since mitigated. The engineering team is monitoring.
Issues with actions in the Grafana IRM mobile app
major 1497h 7m

Started:

monitoring
We've verified a fix in our staging environment to restore functionality to the mobile app. The fix is currently being deployed to production. Thanks for your patience as we continue to roll this out and monitor the resolution.
identified
We're noticing an uptick in users being unable to respond to actions on the mobile app (acknowledging and silencing alerts, for example). Users working in the web UI should not be affected. Ingestion and notification delivery are working as expected. We have a fix in place and are in the process of deploying.
Potential Issues Loading Grafana for Users in India
major 1498h 51m

Started:

monitoring
Error rates have remained near zero, and we continue to monitor.
monitoring
We are continuing to monitor for further issues.
monitoring
We have deployed additional mitigations that should help with remaining errors. We are continuing to monitor error rates.
monitoring
We’ve verified and begun to implement a fix that will improve loading errors. We are continuing to roll this out to all regions and monitor for efficacy.
monitoring
We're actively monitoring this issue and working with our 3rd party provider. The next update will be sent on Monday unless there's new information to share.
monitoring
Due to the linked GCP outage below, users located in India may have trouble loading parts of Grafana. https://status.cloud.google.com/incidents/5fGQt4VbkDnr3Yp8PXPr We are continuing to work with our CSP on this investigation. Impacted users may receive intermittent error messages such as "Error Loading" or "Failed to load Assets". To be clear, it does not matter the region the stack is located, but the geography where the user is physically in.
monitoring
Due to the linked GCP outage below, users located in India may have trouble loading parts of Grafana. https://status.cloud.google.com/incidents/5fGQt4VbkDnr3Yp8PXPr Impacted users may receive intermittent error messages such as "Error Loading" or "Failed to load Assets". To be clear, it does not matter the region the stack is located, but the geography where the user is physically in. We continue to work with our CSP on this investigation.
investigating
Due to the linked GCP outage below, users located in India may have trouble loading parts of Grafana. https://status.cloud.google.com/incidents/5fGQt4VbkDnr3Yp8PXPr Impacted users may receive error messages such as "Error Loading" or "Failed to load Assets". To be clear, it does not matter the region the stack is located, but the geography where the user is physically in. We are currently investigating this issue from our end, and will provide updates as they are available.
Degraded k6 cloud UI performance
critical 1502h 43m

Started:

monitoring
We are continuing to monitor for any further issues.
monitoring
The root cause of the issue has been identified and a fix has been successfully deployed. We are observing widespread improvements across all systems. Our team is currently monitoring the environment to ensure performance remains stable.
investigating
We are continuing to investigate this issue.
investigating
We’re currently investigating an issue resulting in degraded k6 cloud UI performance and API response time. Our team is actively working to rectify this issue.
Frontend Observability - Suspected commit feature not working as expected
minor 1505h 20m

Started:

investigating
We’re currently investigating an issue affecting Frontend Observability product. The "Suspected commit" feature is not currently working as expected. Ingestion and querying is unaffected by this. Our team has identified the cause and is actively working on a fix. Thank you for your patience.
Loki data source-managed alert rules not visible in the Grafana Cloud Alerting UI
major 1517h 51m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
We are continuing to deploy the fix and monitor recovery efforts. As part of the rollout, we identified an issue that required adjustments to our deployment plan, which has extended the timeline for mitigation. Work remains actively underway, and we will share additional updates as progress continues.
identified
Deployment of the fix is still in progress. We are continuing to monitor the rollout and validate recovery across affected systems. We will share further updates as they become available.
identified
Our Engineering Team has implemented a fix which is now being rolled out. We will continue to monitor the situation and update as soon as we have more information.
identified
We have identified an issue where alert rules and alerts managed directly in a Loki data source (data source-managed alerting) are not displayed in the Grafana Cloud Alerting UI. Rules created via Prometheus/Mimir data sources and Grafana-managed alert rules are not affected. Impact is limited to visibility and management in the UI. Affected alert rules continue to evaluate and send notifications normally — there is no impact to alert delivery. Workaround: Loki alert rules can still be vi...
Grafana Dashboards page not displaying when set to ‘View by Folders’
major 1695h 19m

Started:

investigating
We’re currently investigating an issue affecting The Grafana Dashboards page. When set to view by folders, is currently experiencing an issue where no dashboards are shown. Our team is working on fixing the problem. In the meantime, switching to ‘View as list’ allows access to dashboards as usual”.
Investigating Issues with Data Source-Managed Alerting
major 1706h 58m

Started:

monitoring
Our team has implemented a fix and we are currently monitoring the results of this.
investigating
We are currently investigating an issue affecting data source-managed alerting management functionality in Grafana Cloud. Customers may experience problems viewing, creating, updating, or managing alerts through Grafana when using data source-managed alerting. This issue is limited to alert management functionality within Grafana. Alert evaluation and backend alerting services continue to operate normally. Direct alerting APIs for Mimir and Loki remain fully operational and are unaffected. ...
IRM Degraded Performance
minor 1743h 23m

Started:

monitoring
We've released a fix to the IRM app that should restore service for affected customers with issues related to labels. Thanks for your patience while investigating. We're continuing to monitor as we confirm the resolution in place.
identified
We are continuing to work on a fix for this. To further clarify, this issue is not about accessing IRM or alert ingestion/notification/delivery, but rather with handling labels.
identified
The degraded performance is about labels, and we have seen this degradation in more regions.
identified
We are continuing to work on a fix for this issue.
identified
The issue has been identified and a fix is being implemented.
investigating
We are continuing to investigate this issue.
investigating
We are experiencing access issues in IRM as there are elevated 500 API responses in prod-us-central-0.
Brief Rule Evaluation Failures in prod-eu-west-3
major 1776h 29m

Started:

monitoring
The incident has been mitigated, and services are operating normally. We continue to monitor the service to ensure full stability.
monitoring
The incident has been mitigated, and services are operating normally. We are currently monitor the service to ensure full stability.
investigating
We’re making ongoing progress on the investigation alongside our upstream provider.
investigating
We are continuing to investigate this issue.
investigating
Intermittent spikes in rule evaluations continuing.
investigating
From 00:20:00 to 00:27:00 and again 00:32:00 to 00:38:00 there were brief spikes in rule evaluation failures. Engineers are investigating.
Permissions Issues with IRM
critical 1809h 57m

Started:

monitoring
Continuing to monitor progress. Most customers affected should have all services restored, with a few remaining customers receiving updates as the rollout finishes out. Thanks again for your patience.
monitoring
A fix has been released to prod and rolling out across the fleet for IRM, restoring access to affected customers. Thanks for your patience through this work. We're continuing to monitor to confirm we've returned to a steady state.
monitoring
We've identified an earlier regression in one of our recent code changes that was affecting resolution of our previous fix. We're deploying this change now and applying a hot fix in the interim to restore access quickly
monitoring
A fix is being deployed now, and we are monitoring the progress.
identified
We've identified an issue with RBAC, and are working on a fix to restore permission services for those affected.
investigating
Our engineering team is still investigating this issue. We do not have any new information to share at this time, but will continue to provide timely updates.
investigating
We are currently investigating an issue impacting permissions for IRM. As a result, users are not currently getting paged. We will provide updates as they become available.
Silences not Working as Expected
major 1812h 6m

Started:

identified
We have identified an issue causing Silences to not work as expected in the Cloud (Mimir) Alertmanager. Grafana Alertmanager is working ok, this is only affecting Data source-managed alerts.
Grafana Assistant Skills Page Blank
major 1833h 25m

Started:

identified
The issue has been identified, and we are working on a fix.
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating an issue affecting the Skills page of Grafana Assistant. Impacted deployments will encounter a blank screen when attempting to access this page. At this time, we have observed partial impact in the us-east-0 and us-central-0 regions, and will provide an update here if the scope of impact expands.
K6 Test Runs Degraded Performance
minor 1853h 29m

Started:

monitoring
We have applied a fix, and are monitoring the results.
investigating
We are currently investigating an issue causing k6 test runs to take longer than expected to complete, or to time out within Grafana Cloud.
Synthetic Scripted/Browser checks failure
major 1862h 31m

Started:

identified
We are in the process of deploying a fix for this issue.
investigating
We’re currently investigating an issue affecting Synthetic Monitoring where updates for Scripted/Browser checks might fail. Our team is actively working to identify the cause. Thank you for your patience.
tempo prod-25 write-path-down
minor 1874h 54m

Started:

identified
Between 21:20 and 22:40 UTC, writes to tempo-prod-25 failed due to an outage. tempo-prod-24 was also affected during an overlapping window from 22:32 to 22:40 UTC."
Alert manager unavailable in prod-us-central-0
minor 1902h 15m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
Starting at 18:30 UTC, we noticed alert manager unavailability limited to prod-us-central-0 which affects grafana-managed and datasource-managed alerting, causing disruption to updating alertmanager config and limited disruption to alert sending. We have identified the cause and are in the process of remediation.
May 2026
Grafana Loki Log Query Issues
major 1985h 5m

Started:

monitoring
We have identified the cause of this incident and a fix has been applied. Normal functions are returning. We are currently monitoring the recovery process.
investigating
We’re currently investigating an issue affecting Loki queries in Grafana. We have had reports from customers showing the logs are not loading or showing missing logs. Our team is actively working to identify the cause. Thank you for your patience.
Prometheus Datasource Errors/Outage in prod-us-east-0
major 2021h 46m

Started:

investigating
We are seeing recovery across affected Prometheus datasources, and error rates have significantly improved. The service is recovering without any required customer action, and our team continues to monitor stability while we investigate the underlying cause. We’ll provide another update as we learn more.
investigating
We continue to investigate an issue affecting Prometheus datasources causing intermittent timeouts and unexpected errors, primarily impacting alert rule evaluations. Our team is actively working to identify the cause. Thank you for your patience.
investigating
We’re currently investigating an issue affecting Prometheus datasources causing 500 internal or Unexpected errors. Our team is actively working to identify the cause. Thank you for your patience.
Grafana K6 metrics processing and test runs degradation
minor 2249h 45m

Started:

monitoring
We've stabilized the system and test runs no longer result in timeout. There is a small delay (a few minutes) in processing metrics at the end of the test run, but most users shouldn't be too negatively impacted by that. We expected the delay/lag to also resolve within the next 30-60 minutes.
investigating
We have identified that test runs are getting timed out as a result of the issue This issue first occurred on May 05/15/2026 at 8:00PM UTC.
investigating
We’re currently investigating an issue that is resulting in degraded performance in metrics processing and test run metrics may take longer than usual to show up. Our team is actively working to identify the cause. Thank you for your patience.
Intermittent Errors and High latency Writing to Cloud Metrics, Cloud Logs and Cloud Traces
minor 2369h 18m

Started:

monitoring
We continue to see signs of recovery and improved stability across impacted services. Our teams continue to closely monitor the situation while working with the cloud provider.
monitoring
We continue to see signs of recovery and improved stability across impacted services. Our teams continue to closely monitor the situation while working with the cloud provider.
monitoring
We are seeing signs of recovery and improved stability across impacted services over the past hour. Our teams continue to closely monitor the situation while working with the cloud provider.
investigating
We have identified expanded impact affecting Grafana Cloud Logs and Grafana Cloud Traces in addition to Cloud Metrics, causing intermittent errors and increased latency when writing data. Our teams continue working on a fix and investigating the issue with the cloud provider’s support team.
investigating
We’re continuing to investigate the issue causing intermittent errors and high latency when writing to Cloud Metrics. We are in contact with the cloud provider’s support team, and they are investigating the issue alongside us.
investigating
We’re currently investigating an issue causing intermittent errors and high latency when writing to Cloud Metrics. Our team is actively working to identify the cause. Thank you for your patience.
"Failed to Load Dashboard" Errors
major 2404h 30m

Started:

identified
The fix is currently being rolled out to all impacted environments.
identified
Our teams continue working on a fix for this issue. We do not have additional information to share at this time, but we will continue to provide updates as progress is made.
identified
We are continuing to work on a fix for this issue. While we do not have additional updates to share at this time, our teams remain actively engaged and we will provide further updates as soon as they become available.
identified
Customers on Grafana Cloud may see an error on dashboard panels with "Failed to load dashboard ... json unmarshal number ...". We have identified the issue and are working to deploy out the fix.
SSL/TLS Connectivity Issues
major 2405h 20m

Started:

investigating
We are currently investigating reports of service disruption affecting a subset of customers. Customers may experience intermittent connectivity issues, degraded performance, or SSL/TLS certificate validation errors when accessing affected services. Our engineering teams are actively working to identify the scope of impact and restore full functionality as quickly as possible. We will continue to provide updates as more information becomes available.
Cloud Metrics -High Write Latency and Errors in prod-us-central-7
minor 2476h 52m

Started:

monitoring
From approximately 20:40-21:00 UTc, we experienced an issue affecting Grafana Cloud Metrics in prod-us-central-7. Affected users may have experienced high latency and/or errors during ingestion and rule evaluation. Our team has identified the cause and mitigated. We are currently monitoring for long-term stability.
Metrics read errors in prod-ap-south-1 region
critical 2514h 51m

Started:

monitoring
Engineering has released a fix and as of 07:50 UTC, customers should no longer experience errors when querying metrics. We will continue to monitor for recurrence and provide updates accordingly.
investigating
From approximately 06:24 UTC, we were alerted to an issue with read errors in mimir-prod-43. Users with instances hosted in the prod-ap-south-1 region experiencing this issue may encounter an error message when querying metrics. Engineering is actively engaged and assessing the issue. We will provide updates accordingly.
Datasource Query Performance Issues
minor 2526h 1m

Started:

investigating
We’re currently investigating an issue affecting Datasource query performance in prod-us-east-4. Our team is actively working to identify the cause. Thank you for your patience.
Elevated Error Rate of Browser Checks in PoP Oregon
minor 2553h 58m

Started:

monitoring
We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
identified
We’ve identified the cause of the issue impacting browser checks. Our team is currently implementing a fix.
investigating
We’re currently investigating an issue affecting browser checks in the PoP Oregon region. Our team is actively working to identify the cause. Thank you for your patience.
k6 Partial Outage
major 2571h 11m

Started:

monitoring
We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
investigating
After further investigation, this issue may also be affecting Synthetic Monitoring. We continue to identify the cause and will update as soon as we have more information.
investigating
We’re currently investigating an issue affecting k6. Our team is actively working to identify the cause. Thank you for your patience.
Ingestion Errors for AWS Cloud Provider Observability Metric Streams in prod-us-central-7
major 2656h 54m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are continuing to investigate this issue.
investigating
We are investigating an issue with ingesting Metrics for AWS Cloud Provider Observability with Metric Streams. Users experiencing this issue may encounter ingestion errors in the "prod-us-central-7" region only starting from ~06:30UTC. Engineering is actively engaged and assessing the issue. We will provide updates accordingly.
April 2026
Investigating Issues Saving SQL Datasource Credentials
minor 2719h 23m

Started:

monitoring
We’ve identified the cause of the issue impacting SQL datasources. Our team is currently implementing a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
investigating
We are currently investigating reports of issues affecting SQL-based data sources where users are unable to save credentials. This appears to impact a subset of customers and may be occurring across multiple regions. We are actively working to determine the scope and root cause. We will provide updates as more information becomes available.
Gateway Slowness Detected in Prod (US-East-1)
minor 2728h 49m

Started:

investigating
Successful requests have dropped, users may not be able to access their instances.. The issue is under investigation.
InfluxDB Datasource - Intermittent Failures
major 2745h 1m

Started:

monitoring
We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
identified
We’ve identified the cause of the issue impacting the InfluxDB datasource. Our team is currently implementing a fix.
investigating
We’re currently investigating an issue affecting the InfluxDB plugin. Some users may see intermittent failures. Our team is actively working to identify the cause. Thank you for your patience.
Cloudwatch Datasource Outage
major 2843h 42m

Started:

monitoring
We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
investigating
We’re currently investigating an issue affecting Cloudwatch datasources. Our team is actively working to identify the cause. Thank you for your patience.
Restrictions on Alerts & Reports for Grafana Cloud Free/Trial Users
minor 2908h 56m

Started:

monitoring
Grafana Labs is implementing measures to safeguard the Grafana Cloud platform against ongoing unauthorized use while preserving the capabilities relied upon by our community. Effective immediately, we have made the following modifications to the platform: Alerting Email alerting has been disabled for new Grafana Cloud Free and Trial accounts; however, all other integrations such as webhooks remain functional. Additionally, Cloud Alertmanager is now disabled for Grafana instances in these acc...
monitoring
We are continuing to monitor for any further issues.
monitoring
Grafana Labs is taking steps to safeguard our Grafana Cloud platform against unauthorized use while maintaining the Grafana Cloud Free and Trial tiers of service our users and the community have come to rely on. As of Monday April 20, alerting and reporting capabilities have been disabled in new Grafana Cloud Free and trial stacks. We are working towards deploying improvements and restoring those functionalities in a way that keeps our platform secure and open for all of our users.
Elevated 429 Errors Impacting Metrics Querying Across Multiple Regions
critical 2916h 0m

Started:

investigating
The issue is now confirmed to be widespread, affecting Prometheus across all regions. Customers may continue to experience elevated 429 (rate limit) errors, particularly when querying metrics, with failures or inconsistent responses possible. Our engineering team remains fully engaged and is actively working on mitigation and resolution efforts with the highest priority.
investigating
We are currently experiencing a major incident causing elevated 429 (rate limit) errors across multiple regions, primarily impacting metrics querying. This is a high-priority issue, and our engineering team is actively engaged and working urgently to identify the root cause and restore full service as quickly as possible. Customers may experience widespread failures or delays when querying metrics during this time. We understand the significant impact this may have and will continue to prov...
Query Caching - Degraded Performance
minor 2980h 45m

Started:

monitoring
Currently prod-us-east-0 and prod-eu-west-3 have recovered, and we are continuing to monitor prod-us-central-0 which is in the process of recovery.
investigating
As of 20:52 UTC, we are currently investigating degraded Query Caching performance in multiple regions. For datasources where query caching is configured, some queries may take longer than usual. Our team is actively working to identify the cause. Thank you for your patience.
Issues on Stack creation
minor 3013h 17m

Started:

monitoring
The issue is fixed and we are currently monitoring the service.
identified
Since today 16th at ~12:11UTC we are seeing issues on stack creation across all our regions. Customers will experience error message when attempting to create a stack. Our engineering team has identified the source of the issue as external to Grafana (provider), and they are tracking its recovery.
Degraded Ticket Visibility in Support System
minor 3034h 1m

Started:

monitoring
We are currently experiencing an issue with our ticketing system provider that is affecting how tickets appear within our internal support views. We are continuing to receive all new tickets successfully, and no requests are being lost at this time. Our team is actively monitoring the situation and working to ensure all incoming requests are reviewed, including those that may not be immediately visible in standard views. We will provide further updates as we receive more information from o...
K6 Sporadic DNS Issues
minor 3064h 46m

Started:

monitoring
Our engineering team has deployed a fix and we are currently monitoring the behaviour of the system until full resolution.
monitoring
We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
identified
We are having sporadic DNS issues that occasionally affect the start of cloud test runs, causing them to abort. We are currently working to resolve. The issue has been occurring since April 9.
Grafana Cloud Logs - Write degradation in us-east-3
major 3146h 15m

Started:

investigating
We are seeing issues on the write path for Loki in cluster in us-east-3, and we are actively investigating this issue.
Tempo Write Outage
major 3150h 26m

Started:

monitoring
We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time. We’ll update again within an hour.
investigating
We are currently investigating a write outage affecting prod-us-east-3. The issue began at 18:50 UTC. Users may experience errors, timeouts, or unavailability while we work to identify the cause and restore service.
K6 Browser Testing/Timeline Not Available
minor 3176h 35m

Started:

investigating
We’re currently investigating an issue affecting browser testing. Users running browser tests will not be able to see the browser timeline. Our team is actively working to identify the cause and will share an update within two hours. Thank you for your patience.
Unable to Edit Notification Policies
minor 3226h 51m

Started:

identified
We’ve identified the cause of the issue impacting notification policies. Our team is currently implementing a fix. We’ll provide another update in 2 hours or sooner if the situation changes.
identified
We’ve identified the cause of the issue impacting notification policies. Our team is currently implementing a fix. We’ll provide another update in 2 hours or sooner if the situation changes.
investigating
We’re currently investigating an issue affecting notification policies. Our team is actively working to identify the cause and will share an update within 2 hours. Thank you for your patience.
Notification Policies and Contact Points Missing in UI on the Slow Release Channel
minor 3251h 20m

Started:

monitoring
We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time. We’ll update again within 2 hours.
identified
We’ve identified the cause of the issue impacting the Notification Policy and Contact Point UI. Our team is currently implementing a fix. We’ll provide another update when the fix is deployed and we monitor the expected improvement.
investigating
We’re continuing to investigate the issue with the alerting UI. While we don’t have new information to share yet, our team is working to identify the root cause. Next update in 2 hours.
investigating
We’re currently investigating an issue affecting notification policies and contact points for instances on the slow release channel. Alerting API calls for contact points and notification policies return data as expected, so this appears to be limited to the UI. Our team is actively working to identify the cause and will share an update within 1-2 hours. Thank you for your patience.
Partial K6 Test Run Outage
major 3322h 39m

Started:

investigating
We're experiencing an outage affecting test runs that use k6 extensions. The issue prevents users from executing these types of test runs both locally and in Grafana Cloud. Test runs that do not use extensions are not affected by this incident.
AWS integration Degraded Performance
minor 3365h 51m

Started:

investigating
We are investigating a noticeable drop in active series for the AWS integration that began around 18:15 UTC. This issue may cause scrapes to hit rate limits, which can result in individual data points not being collected for the serverless integration. The impact is intermittent and may affect any customer using the AWS integration, regardless of region. We are currently working to identify the cause and will provide an update as soon as we have more information.
Query degradation and possible rule evaluation failure on prod-eu-west-0.cortex-prod-01
minor 3376h 12m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are continuing to investigate this issue.
investigating
We are currently observing delays in ingesting data, possibly causing partial query results and failed rule evaluations for prod-eu-west-0.cortex-prod-01 metrics cell.
March 2026
Some of the CloudWatch queries are failing
major 3400h 20m

Started:

monitoring
We are continuing to monitor for any further issues.
monitoring
Some of the CloudWatch queries were failing. Started at 08:37 UTC Monitoring from 09:21 UTC
Some Grafana Instances Unavailable
major 3492h 33m

Started:

monitoring
We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time. We’ll update again in 1 hour.
identified
We’ve identified the cause of the issue impacting the instances. Our team is currently implementing a fix. We’ll provide another update in 1–2 hours, or sooner, if the situation changes.
investigating
We’re continuing to investigate the issue with Grafana instances. While we don’t have new information to share yet, our team is working to identify the root cause. Next update in 1-2 hours.
investigating
We’re continuing to investigate the issue with Grafana instances. While we don’t have new information to share yet, our team is working to identify the root cause. Next update in 1-2 hours.
investigating
We’re currently investigating an issue which is affecting primarily users on the Free tier. Impacted users will be met with a "your Grafana instance is loading" message indefinitely. Our team is actively working to identify the cause and will share an update within 1-2 hours. Thank you for your patience.
Prometheus writes in prod-eu-west-3 are degraded
critical 3539h 57m

Started:

monitoring
We are continuing to monitor for any further issues.
monitoring
We have deployed mitigation and seen improvement in write failures over the past week. We are still seeing intermittent spikes in latency and continue to monitor.
monitoring
We are still seeing intermittent issues and continue to seek a resolution
monitoring
We are continuing to monitor for any further issues.
monitoring
We are continuing to monitor this through the weekend.
monitoring
We are continuing to monitor the previously impacted environments.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are continuing to investigate this issue.
investigating
The metric writes issue reported in https://status.grafana.com/incidents/gfshj17lxj5z is still ongoing. Our Engineering team is actively investigating this and we will provide further updates as our investigation progresses.
Prometheus writes, Logs, and Synthetic Monitoring in prod-eu-west-3 are degraded
minor 3569h 1m

Started:

investigating
This is also now impacting Logs and Synthetic Monitoring in prod-eu-west-3. For Synthetic Monitoring, users might observe errors pushing check execution metrics, and this can eventually lead to missing data. In addition, users might observe errors evaluating Synthetic Monitoring provisioned alert rule evaluations, and this can lead to missed alerts. For Logs, there is no immediate impact on alerts, however, remote writes to Mimir is delayed which means users may see gaps in their recordin...
investigating
We are moving this back to 'Investigating' as we are now observing a substantial drop in successful ingestion and increase in write path errors, and elevated rule evaluation latency and error. Reads are mostly fine. Our Engineering team is actively investigating this and we will provide further updates as our investigation progresses.
monitoring
We have not observed any recent errors, but we will continue to monitor while we work with our CSP.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently experiencing degraded writes for mimir-prod-22 in prod-eu-west-3 since 08:45Z.
Grafana Assistant Unavailable in prod-us-east-0
major 3585h 5m

Started:

identified
The issue has been identified, and we are implementing a fix.
investigating
The impact extends beyond the TOS check. Assistant is completely unavailable in the impacted region.
investigating
We are continuing to investigate this issue.
investigating
We are aware of an issue currently impacting Grafana Assistant. Impacted users are met with a request to accept the TOS, however the plugin is failing upon accepting. Our engineering are currently investigating this issue.
Authentication API Database Down in prod-eu-west-2 and prod-eu-west-4
major 3659h 8m

Started:

investigating
We have observed impact in prod-eu-west-4 as well.
investigating
We are currently investigating an issue impacting the main database for Authentication API's in the prod-eu-west-2 region. Writes are currently failing, but reads are operational.
Various Datasource Issues
major 3681h 22m

Started:

monitoring
We are continuing to monitor for any further issues.
monitoring
We have observed recovery for the Cloudwatch Datasource. We are now seeing failures for the following Datasources: Aurora Opensearch X-Ray Timestream Redshift Sitewise A fix for the above is being rolled out now, and we will monitor progress. We will also change the name of this incident from "Cloudwatch Datasource Issues" to "Various Datasource Issues" to more accurately reflect impact.
monitoring
We have identified the issue, and are rolling out the fix. We are already seeing improvements and will continue to monitor progress.
investigating
We are currently investigating an issue impacting the CloudWatch Datasource causing failures.
Degraded performance of Grafana Cloud k6 test runs
major 3686h 51m

Started:

investigating
Some customers are seeing degraded performance and errors from certain v6 API endpoints. We are investigating the issue.
Grafana Cloud Logs - Write degradation in Azure Netherlands (eu-west-3)
minor 3831h 41m

Started:

investigating
We are continuing to investigate this issue with our CSP, and will provide updates as they become available.
investigating
We are seeing issues on the write path for Loki in cluster Azure Netherlands (eu-west-3). Impact will reflect in degradation of logs ingestion on that cluster. Our engineering team is already working on restoring the service.
Increased number of Aborted-by-Systems with a k6 binary building errors
major 3834h 28m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
investigating
We are seeing an increased number of Aborted-by-Systems with a k6 binary building error. We are investigating the issue. The first occurrence of this happened back on March 9, has now been identified as a blocking issue for some customers.
Rule Evaluation Outage in prod-us-west-0
major 3872h 58m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently investigating an issue impacting rule evaluation for a subset of customers in the prod-us-west-0 region. We will provide updates as they become available.
Grafana Cloud Logs - Write degradation in Azure Netherlands (eu-west-3)
minor 3881h 37m

Started:

investigating
We are also reporting impact to Faro performance in the same region. We are continuing to investigate this issue.
investigating
We are seeing issues on the write path for Loki in cluster Azure Netherlands (eu-west-3). Impact will reflect in degradation of logs ingestion on that cluster. Our engineering team is already working on restoring the service.
Complete outage in prod-me-central-1
critical 4099h 25m

Started:

investigating
The TLS certificates serving prod-me-central-1 endpoints expire on May 30, 2026. Replacement certificates have been imported, but the ongoing AWS regional incident is preventing them from propagating to all load balancer nodes, so customers may see certificate errors after that date until AWS restores normal operation. We do not have any additional updates to share at this time. Our team is actively monitoring the situation and will provide further information as it becomes available. In th...
investigating
AWS UAE - prod-me-central-1: Public Probe checks might suffer degraded experience. We recommend migrating checks from the UAE probe to the next nearest probe suitable for your use case.
investigating
We do not have any additional updates to share at this time. Our team is actively monitoring the situation and will provide further information as it becomes available. In the meantime, please continue to refer to the AWS Status Page for the most detailed and up-to-date information.
investigating
We are continuing to investigate this issue.
investigating
We have not received any further updates from AWS at this time. However, we are actively monitoring the outage and will provide additional information as it becomes available. Also, please continue to refer to the AWS status page for more detailed updates. https://health.aws.amazon.com/health/status All the guidance previously included about stack migration is still relevant. Please reach out to our Support team if you have any questions.
investigating
We are actively monitoring the situation, but at this time there are no new updates to share. The next update will be provided once we have more information to share. Please reach out to our Support team if you have any questions.
investigating
We are continuing to investigate this issue.
investigating
Please continue to refer to the AWS status page for more detailed updates specific to AWS. https://health.aws.amazon.com/health/status AWS are recommending that affected customers move workloads to alternate regions, and we are recommending the same. Customers who are impacted and who cannot wait for a restoration of service are asked to: 1. Create a Grafana Cloud stack in an alternate region 2. Update clients to send telemetry to the new region, if using Grafana Alloy then you can use Fle...
investigating
AWS are recommending that affected customers move workloads to alternate regions https://health.aws.amazon.com/health/status and we are recommending the same. Customers who are impacted and who cannot wait for a restoration of service are asked to: 1. Create a Grafana Cloud stack in an alternate region 2. Update clients to send telemetry to the new region, if using Grafana Alloy then you can use Fleet Management https://grafana.com/docs/grafana-cloud/send-data/fleet-management/introduction/...
investigating
Customers are recommended to configure a new blank stack in an alternative Grafana Cloud region and to reconfigure their clients (such as Grafana Alloy) to send telemetry to that region, Fleet Management can be used for this purpose https://grafana.com/docs/grafana-cloud/send-data/fleet-management/introduction/
investigating
We are updating this incident to reflect a complete outage in prod-me-central-1, due to an on-going AWS UAE data center issue. We will provide further updates accordingly.
investigating
We are observing write and read outage errors across all databases (metrics, logs, traces) in prod-me-central-1, due to an on-going AWS UAE data center issue. We will provide further updates accordingly.
investigating
We are observing write and read outage errors across all databases (metrics, logs, traces) in prod-me-central-1, due to an on-going AWS UAE data center issue. We will provide further updates accordingly.
investigating
We are seeing elevated write and read path errors in prod-me-central-1, due to an on-going AWS UAE data center issue. We will provide further updates accordingly.
February 2026
Grafana Cloud Metrics - Intermittent Write Latency in prod-us-central, prod-us-central-5, and prod-eu-west-0
minor 4206h 15m

Started:

monitoring
We are rolling out a mitigation across the environments in these regions, and preemptively where possible to ensure it doesn’t spread elsewhere.
monitoring
We have seen an increase in latency in our cloud providers services, and are rolling out a change to mitigate the issue. We are monitoring.
monitoring
We are continuing to investigate this issue alongside the CSP, and have taken steps to escalate through the appropriate channels. The mitigation in place continues to work as expected, and any notable updates will continue to be shared here for tracking.
monitoring
We are continuing to investigate this issue alongside the CSP. Any notable updates will continue to be shared here for tracking.
monitoring
We've implemented mitigation in place and are continuing to monitoring and investigating this issue.
investigating
We have begun rolling out mitigation steps to reduce write latency in the prod-us-central-0 and prod-us-central-5 regions. While these measures are expected to improve performance, we are continuing to investigate the underlying root cause of the issue. We will provide additional updates as more information becomes available.
investigating
Since February 19, we have been investigating an intermittent issue causing increased write latency in the prod-us-central-0 and prod-us-central-5 regions. The issue does not affect all traffic but may result in delayed write operations for some customers. Our engineering team is actively working to identify the root cause and stabilize performance. We will share additional updates as progress is made.