Two Providers Down at Once Is Usually One Provider and a Shared Dependency
    Cloud Governance & Security

    Two Providers Down at Once Is Usually One Provider and a Shared Dependency

    Two providers rarely fail independently — usually one fails and a second inherits it through a dependency nobody mapped. Here's the runbook check that would have caught it.

    Nathan Barrett
    Nathan Barrett

    Chief Product Officer

    August 15, 2026
    8 min read
    Share:

    Most incident plans assume one thing breaks at a time. The last fourteen months of public postmortems disproved that assumption, repeatedly, and not in the way most runbooks expect.

    One correction first, since it matters for anyone scheduling a post-incident review this week. As of 14 August 2026, Cloudflare's status page showed all systems operational. Recent entries were limited to minor, localised events:

    • A Durable Objects and Cloudflare Workflows availability drop, resolved 4:05 PM on 14 August.
    • Increased network congestion in the Eastern US.
    • Increased HTTP 5xx errors in Kuwait, Bangkok, Jakarta and Dammam.
    • Network performance issues in Querétaro, Mexico.

    Third-party monitoring on the same day recorded Google Cloud as operational, with one outage in the prior 30 days: the 15 July 2026 europe-west4-a cooling failure. No documented simultaneous Google Cloud and Cloudflare failure happened yesterday. The dual-provider event people are thinking of is 12 June 2025.

    That older incident is still the more useful teacher. So this piece works from it, and from what followed.

    What actually happened in June 2025

    Google Cloud attributed its global disruption to an invalid automated quota update. That update was pushed to its API management system and caused external API requests to fail. Its remediation commitment: stop metadata propagating globally without proper protection, testing and monitoring.

    Cloudflare went down in the same window for 2 hours and 28 minutes. The outage hit all customers using Workers KV (Cloudflare's key-value storage service), WARP, Access, Gateway, Images, Stream, Workers AI, Turnstile and Challenges, AutoRAG, Zaraz and parts of the dashboard. Workers KV saw 90.22% of requests fail. Cloudflare Access failed 100% of identity-based logins for self-hosted, SaaS and infrastructure application types, because Access is built to fail closed when it can't fetch policy configuration or user identity.

    The cause is worth reading twice. Cloudflare said the storage infrastructure behind Workers KV is partly "backed by a third-party cloud provider, which experienced an outage today." It then added that "while the proximate cause (or trigger) for this outage was a third-party vendor failure, we are ultimately responsible for our chosen dependencies and how we choose to architect around them."

    So this wasn't two providers failing independently. One provider failed, and a second inherited that failure through a dependency almost none of its customers had mapped. Capgemini's Pradeep Sanyal called the single point of failure in Workers KV storage "especially instructive." That's exactly the blind spot dependency mapping exists to close.

    Why regional failover would not have saved you

    The failure modes in the recent record propagate globally, not regionally.

    Cloudflare described its 18 November 2025 outage as its worst since 2019. It began at 11:20 UTC after a database permissions change caused a Bot Management feature file to double in size and exceed a software size limit. Core traffic was largely flowing normally by 14:30 UTC; everything was back to normal at 17:06 UTC. Cloudflare initially, and wrongly, suspected a hyper-scale DDoS attack. Downdetector itself was down during that window, along with platforms including X and ChatGPT.

    Configuration and metadata that propagate faster than any blast-radius control (a limit on how far a failure can spread before it's contained) aren't a capacity problem. Rehearsing a region evacuation does nothing for them. What helps is a pre-specified degradation path, and dependency mapping is what makes writing that path possible.

    Degradation is a tradeoff, not a free win

    Cloudflare activated kill switches (mechanisms that quickly disable a feature) during the June 2025 incident so end users wouldn't be blocked. Its own postmortem records the cost: while active, the Turnstile siteverify API could redeem valid tokens more than once, potentially letting a bad actor reuse a previously valid token.

    That's the honest shape of failing open. Usually the right call. It carries a security cost, and you should decide to accept that cost in advance, in writing, with a named owner, rather than mid-incident on a bridge call. Dependency mapping is what surfaces these decisions early enough to make them calmly.

    The dependency map most runbooks are missing

    Grant Thornton's November 2025 commentary argues that standard vendor assessments and SLAs rarely show the full picture, and recommends dependency mapping to find single points of failure across the digital ecosystem, including the sub-vendors of major cloud and CDN providers. Dependency mapping is the artifact to produce this week. Treat it as a living document, not a one-time audit.

    For teams running agentic and customer-facing paths on Google Cloud, the map has to cover at least these layers:

    • The edge provider fronting your application, and what it does when its control plane is unavailable.
    • Your identity and authorisation path, and whether it fails open or closed. Access failed closed by design in June 2025. Defensible, and it still meant no logins.
    • Model endpoints, which most legacy DR runbooks never listed as a failure domain. Google Cloud's service health history records a 27 Feb 2026 incident where Vertex AI Gemini API customers hit increased error rates on the global endpoint, lasting 1 hour 58 minutes.
    • Data and storage paths, including any managed service whose substrate belongs to someone else.
    • Your detection path: status pages, monitoring, and the person who tells the customer.

    Health-check scope belongs on that map too. Origin-only health checks stay green while a global control-plane failure takes user traffic down. A dashboard full of healthy origins tells you nothing about whether customers can log in or get a model response.

    On detection, Google Cloud promotes Personalized Service Health for seeing incidents affecting your own projects, with custom alerts, API data and logs. It's a more targeted signal than the public status page, which is the wrong instrument for answering "is this us."

    Endpoint choice is now a resilience decision

    If you run Claude on Vertex AI, this is a concrete post-incident action. Google Cloud announced U.S. and EU multi-region endpoints for Claude on Vertex AI on 15 April 2026 in public preview; an editor's note on that post states general availability was announced on 15 May 2026. Endpoint selection is itself a finding dependency mapping should surface before an incident forces the choice.

    Google's own comparison of the three options:

    • Global endpoints: maximum availability with global failover, and no data residency guarantees.
    • Multi-region endpoints: high availability with multi-region failover, restricted to a geography.
    • Regional endpoints: dependent on single-region capacity, with multi-zone failover.

    Google recommends multi-region as the default for production workloads with U.S. or EU residency requirements. Confirm current endpoint availability for the specific Claude models you run before rewriting routing. Google documents that global endpoints for partner models can carry a separate set of quotas and don't support data residency requirements.

    What to change in the runbook this week

    Google Cloud's disaster recovery planning guidance structures the work around designing to recovery goals, designing for end-to-end recovery, making tasks specific, maintaining more than one data recovery path, and testing regularly. Translated into this week's actions, starting with dependency mapping:

    1. Map the real dependency graph for your two or three most customer-visible paths, including sub-vendors.
    2. Re-scope your health checks. After an incident, our CloudOps team re-scopes health checks so they hit the model endpoint and the identity path, not just the load balancer origin.
    3. Write the degradation behaviour for every AI feature before you need it: cached responses, a fallback model, or disable-and-explain. Record the security consequence of each.
    4. Set your detection target and measure it. Track time-to-detection and who told the customer first.
    5. Rehearse the fallback and keep the report. We rehearse a regional evacuation on a fixed cadence with customers and keep the report as the compliance artifact. A Gresham analysis of DORA in 2026 argues regulators know full dual-cloud redundancy is unrealistic, and that a firm with a long outage but documented testing of a fallback plan faces significantly less regulatory heat than one with no contingency at all.
    6. Cost the diversification before you commit to it. Sanyal's caution stands: diversification brings its own costs and complexities. It's a board conversation, not only an engineering one.

    One note on figures: sources disagree on the duration of the 15 July 2026 europe-west4-a event. Google's summary row shows 12 hours 28 minutes for Bare Metal Solution; Computing reported 14 hours 55 minutes from Google's incident report. If a duration goes into your review, attribute it.

    If your dependency map doesn't exist yet, that's the deliverable to start with. Our CloudOps team runs this dependency map as part of the Workshop & Assessment engagement, alongside a prioritised remediation plan, before anyone commits to re-architecting.

    Ready to transform your enterprise with AI?

    Book a free AI Enablement Session with our team to discuss how agentic workflows can accelerate your business.

    Book Your Session