New Multi-region uptime checks and custom-domain status pages

DataRobot Outage History

Daily status observations, past incidents, and reported issue history for DataRobot.

Checking current status...
85.7% of 91 observed days had no reported issue

90-Day Trend

May 28Aug 25

Monthly Status Summary

Month Issue-free days Days Tracked Days with Issues
August 2026 96% 25 1
July 2026 93.5% 31 2
June 2026 80% 30 6
May 2026 20% 5 4

This percentage summarizes normalized provider-status observations by calendar day. It is not duration-based, component-weighted, or contractual uptime. See the methodology and limitations.

Daily Status (Last 91 Days)

May 27 Today
Operational Degraded Partial Outage Major Outage Maintenance No Data

Incident History

August 2026
Codespace Session Validation Failure in Generative AI Playground
minor 37h 49m

Started:

Website Website Website API API API Predictions Predictions Prediction AutoML AutoML AutoML AI Catalog and Data Ingest AI Catalog and Data Ingest AI Catalog and Data Ingest AI Apps AI Apps AI Apps MLOps MLOps MLOps Pipeline Generative AI LLM Playground Notebooks Generative AI VDB Builder Generative AI LLM Playground Notebooks Generative AI VDB Builder Generative AI LLM Playground Generative AI VDB Builder
identified
Engineering has implemented a fix and is currently working on deploying it across the Production clusters.
identified
Codespace session in the Generative AI Playground fails due to a validation error. Engineering is working on creating a fix.
July 2026
Issues with Vector Database Creation
minor 632h 36m

Started:

monitoring
Engineering team has applied the fix and Vector Database creation is operational now. The team is monitoring the situation.
investigating
The DataRobot Engineering team observed issues while creating a DataRobot hosted Vector Database. Currently, this is affecting all the DataRobot MTS environments and engineering is investigating the root cause.
Issues with prediction requests against custom model deployments
minor 1165h 5m

Started:

identified
The team root caused the issue causing HTTP 403 responses for Custom Model deployments in EU region.  Mitigations were applied and the team is verifying the fix stabilized the service.
investigating
We are investigating issues in MTSaaS within EU region impacting prediction requests against Custom Models deployments.
May 2026
Batch Jobs queued up in US MTS
minor 1985h 32m

Started:

monitoring
We are continuing to monitor for any further issues.
monitoring
The remediation script has been deployed and Engineering is actively monitoring the situation. Batch jobs may take longer than usual to show as completed until a permanent fix is rolled out in the next production deployment.
identified
The affected jobs have been resolved and the issue is mitigated. Engineering is actively working on a permanent fix and testing is currently underway
identified
Batch jobs that are currently in the queue are completed successfully; however, their completion status is not being updated correctly. As a temporary workaround, we are manually marking these jobs as completed while we work on implementing a solution.
Delay in Feature Drift Statistics Processing
minor 2009h 38m

Started:

identified
Feature drift statistics are currently experiencing a processing delay of approximately 1 hour. No data has been lost, and all metrics will reflect accurate values once processing catches up. Our team is actively working to resolve this.
Customers Experiencing Errors with New Custom Model Creation.
minor 2061h 0m

Started:

monitoring
We are continuing to monitor for any further issues.
monitoring
Engineering has applied a fix in the US SAAS environment which resolved the issue. At the time of issue, some users might have experienced issues with Custom Apps, Data upload and custom model creation. The issue is contained.
investigating
We are experiencing a service interruption with Custom Models functionality in US SAAS environment. Predictions to existing deployments are working fine, but users cannot create new custom models. Engineering is investigating the issue and will provide updates as we make further progress.
Widespread intermittent service issues for new workloads in US Production
minor 2497h 53m

Started:

monitoring
Engineering resolved the underlying issue with workload scheduling and is monitoring the cluster.
identified
We are continuing to experience issues launching new workloads for Custom Models and Custom Applications in US Production. This is connected to an ongoing AWS outage. Our team is exploring multiple mitigation options.
identified
We are continuing to experience issues launching new workloads for Custom Models and Custom Applications in US Production. This is connected to an ongoing AWS outage. Our team is exploring multiple mitigation options.
monitoring
We are currently experiencing intermittent service issues in US Production, which are primarily affecting the launch of new workloads for Notebooks, Custom models, and Custom Applications. This issue does not impact existing workloads. This disruption is strongly correlated with an ongoing AWS Availability Zone outage (https://health.aws.amazon.com/health/status), causing resource allocation failures. The team is actively monitoring the situation and tracking updates from AWS.
April 2026
Delay in processing actual messages
minor 3159h 44m

Started:

monitoring
Processing actual messages on JP MTS is delayed due to autoscaling malfunction. Engineering scaled up the deployment to alleviate the issue. Root cause mitigation in progress
Elevated Errors on Managed AI Cloud
major 3172h 13m

Started:

monitoring
Engineering has applied changes to mitigate the elevated error rates. Services are now operating normally. We are continuing to monitor the system while investigating the cause of the issue.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We're experiencing an elevated level of errors and are currently looking into the issue.
March 2026
Degraded Performance on DataRobot MTS due to Quay outage
minor 3413h 26m

Started:

identified
We are continuing to work on a fix for this issue.
identified
Our engineering team has found the the Quay outage currently happening is causing degraded performance across the DataRobot platform. Engineering is currently monitoring the situation.
Performance Degradation on Managed AI Cloud
minor 3824h 19m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are experiencing performance degradation on Managed AI Cloud.
Intermittent UI disruptions on Managed AI Cloud
minor 3870h 4m

Started:

monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently investigating this issue.