Monitor customer paths, not provider logos
Using more than one cloud does not automatically make an application resilient. The useful unit of monitoring is the customer path: sign-in, checkout, file upload, message delivery, or another operation the business depends on. Each path can cross cloud infrastructure, DNS, identity, payments, and other SaaS services.
List those dependencies and record what a failure looks like from the customer’s point of view. The result is more actionable than a dashboard containing every product your company has purchased.
Combine three kinds of evidence
A practical monitoring plan uses:
None is sufficient by itself. Provider status can explain a broad event but may not match your region or account. An endpoint check may be green while a specific workflow is broken. Internal telemetry may show errors without identifying an upstream cause.
Assign ownership and response
For each important path, name an owning team and the alert destination. Describe the first useful response: verify a transaction, pause a queue, disable a feature, or open the relevant vendor support case. If nobody is expected to act on an alert, it probably belongs in a dashboard or report instead of an urgent channel.
Use severity based on customer impact rather than the vendor’s label alone. A major provider incident may not affect your system, while a narrow regional issue can stop a critical path.
Be honest about failover
Multi-region and multi-cloud failover introduce cost and operational risk. Monitor the readiness of the failover path itself: data replication, identity, network routes, capacity, and required secrets. Exercise the procedure before treating it as a control.
For some workflows, queueing, a read-only mode, or a clear customer message is safer than moving traffic during an incident. Record that decision in the runbook.
Review the map when the system changes
Add dependency review to architectural changes and incident follow-up. Remove services that are no longer used, update owners, and test notification routes. A smaller current inventory is more valuable than a detailed diagram nobody trusts.
ServiceAlert can group vendor status and your own checks around the service they support, giving responders a shared operational view without replacing application telemetry or tested recovery plans.