On Wednesday the eighth of March, Datadog lost its web application, its data ingestion and its alert notifications across multiple regions at once, and roughly thirty hours passed before live data was back everywhere. This is about what a customer could and could not know while that ran, and what the fallback would have cost. What I take from it is that the instrument must not share a failure domain with the patient, and that nobody prices a local fallback until the morning they need one.
12 March 2023·7 min read·operationsdependencies
The first entry on the status page went up at 01:31 Eastern and said that Datadog was investigating loading issues on its web application. Two hours later the same page reported delayed data ingestion across all data types, and the phrase that matters more than any of it: monitor notifications are delayed. That is the polite form of a sentence nobody wants to write, which is that the alerts were not going out.
An instrument can fail two ways, and only one of them is the loud one. It can read wrong, which you notice, because the number is absurd and somebody says so. Or it can stop reading, which looks exactly like a quiet morning. There is no pager, nothing red, nothing to click. Silence is also what a healthy system emits, and no amount of screen real estate distinguishes the two.
So the interesting question about that Wednesday is not how long the outage ran. It is what a customer could and could not know while it ran. If your own service had picked that morning to fall over, the first report would have come from a person, by whatever channel people use when they cannot reach you, and it would have arrived some unknown number of minutes after the fact.
Here is the whole argument in one sentence, and it is not really about observability vendors. The tool you use to see must not fail for the same reasons as the thing you are looking at, because the one property you were paying for is the property that goes missing exactly when you need it.
Here is the timeline, as posted. First entry at 01:31 Eastern on the eighth. The issue reported as identified at 07:08. All services able to receive and query live data at 08:00 Eastern on the ninth, so roughly thirty hours from first symptom to working instruments. Backfilling historical data ran on into the tenth with updates every two to three hours, which means the history took longer to come back than the present did.
There is only one stated cause so far. The summary posted at 15:39 Eastern on the eighth says a system update on a number of hosts controlling the compute clusters caused a subset of those hosts to lose network connectivity, and that clusters then entered unhealthy states and took internal services and datastores with them. That is the sum of it. No component is named, and four days later a customer deciding whether to change anything on their own side has one sentence to reason from.
Something is owed and has not landed. The same summary promises a more detailed analysis post-recovery. I have no complaint about that pace; a write-up worth reading takes weeks, and the alternative is a guess published early and corrected later. It does mean I am arguing from a paragraph. I have tried to make an argument that would not change when the detail arrives, and I will say plainly that this is a thing to be a little suspicious of.
The regions went together. The summary describes issues across multiple products and regions, and points customers at the separate status page for each region. Those regions run separate infrastructure, and being separate is the entire reason they are sold as separate. Whatever the cause turns out to be, it crossed the boundary that exists to stop things crossing, which is the part I would want explained first.
Consider what an alerting pipeline actually promises. It does not promise to tell you when things are bad. It promises to tell you when things are bad and the pipeline is working, and the second half is unstated because it is normally so reliable that nobody prices it. Delayed monitor notifications collapse the two cases into one output, and the output is nothing at all.
The one structure that catches this is a heartbeat, and it has to run somewhere else to work. You emit a signal on a schedule and something outside your estate complains when the signal stops. That check is cheap, it is boring, and it is the only alert in the whole system whose job is to notice that the other alerts have gone quiet. Very few teams have one.
There is a small piece of evidence sitting in plain sight here. The status page everyone was refreshing that morning is not run by Datadog; it is Atlassian's product, on somebody else's infrastructure, with its own domain and its own notification path. That separation is deliberate, and it is the whole idea of this essay rendered in one design decision that the industry made years ago and then applied to exactly one page.
The second-order cost showed up after the lights came back. Customers were told to expect gaps in historical data for parts of the previous day, and gaps in history are not a cosmetic problem. Your capacity math for that week has a hole in it, your error budget for the month is unmeasurable, and any incident of your own that happened inside that window will be investigated without the data you would normally investigate it with.
The instrument runs where the patient runs. Most monitoring, hosted or otherwise, sits in the same handful of clouds as the thing it watches, reaches you over the same networks, and authenticates through the same identity provider. Each of those is a wire connecting the observer to the observed, and the picture only stays trustworthy while every wire holds.
And it boots the same operating system, from the same archive. Packages arrive from a distribution on the distribution's schedule. That is a change to your production estate, initiated by people who have never heard of you, applied to your monitoring and your application by the same tooling on the same day. Keeping hosts current is correct, and the correctness of it is not what determines the blast radius.
A shared change window is a synchronization primitive. If every region applies changes inside the same hour, you have built a clock whose only function is to make failures simultaneous. Independence at the network layer means very little when the schedule above it is common. Staggering costs nothing except patience, and it is the cheapest correlation you will ever break.
Independence is asserted, rarely measured. Nobody produces the list, so here it is: same image, same package source, same config service, same certificate authority, same time source, same deployment pipeline, same change window. Write it down for your monitoring against your production, and the honest answer is usually that the two are the same system wearing different names.
I would buy hosted observability again tomorrow, and I would not treat the choice as close. Self-hosting means you now operate a storage system with awkward retention characteristics, you carry a pager for the thing that carries your pager, and the day your cluster falls over is the day you have no telemetry about your telemetry. That is a worse position, and it is worse on most days rather than one.
But the axis in that argument is wrong. It is not hosted against self-hosted. It is whether any signal path exists that does not pass through the vendor, and the useful version of that is much smaller than a second observability stack. Two days of retention, a handful of vital signs, one notification route that shares nothing with the main one. A small instance and an afternoon.
It does not get built because it has no owner and no visible return. On a budget line it reads as a duplicate monitoring system, which is the easiest thing in the world to decline, and its value cannot be stated in advance because it is denominated in an outage that has not happened. The correct name for it is not a monitoring system. It is a flashlight, and nobody justifies a flashlight by its feature list.
The honest limit is that a fallback nobody looks at is not a fallback. If no one has opened it in six months, the credentials have rotted, the dashboard is empty, and the morning you reach for it is the morning you find out. So the test is not whether the thing exists. It is whether anyone touched it recently, on purpose, when nothing was wrong.
Alert on the absence of data. A heartbeat out to a service that is not your monitoring vendor, checked by something that is not your monitoring vendor. It is the only alert that fires when the alerting is broken, and it takes an afternoon. Everything else in this list is optional; this one I would not run without.
Put the last-resort path on different everything. Different provider, different account, different credentials, different notification channel. The value comes entirely from what it does not share, so every dependency it reuses for convenience is a dependency that takes it down with the main system. Convenience is the whole failure mode here.
Decide in advance what you will not do blind. The deploy you would hold, the failover you would not trigger, the scaling change you would not make with no visibility. Writing that list on a calm afternoon is most of the value, because the alternative is arguing about it at 03:00 with people who cannot see and know they cannot see.
Keep raw telemetry at the edge longer than feels necessary. If the agent buffers, give it room; if the host keeps logs, keep them a while. When the pipe is down for a day you want the day back, and the difference between recovering it and losing it is usually a retention setting somebody trimmed to save a little disk.
I want to keep this proportionate. This was a bad Wednesday, not a scandal. A system update did something the change note did not describe, which has happened to every operator who has ever run a fleet, and the only thing separating the ordinary version from this one is how many machines were on the far side of it. Datadog posted more, and sooner, than most vendors would have, and the promised analysis is not late yet.
What stays with me is smaller. A dashboard is not a window. It is a photograph of a window, developed somewhere else and mailed to you, and on the eighth the mail did not come. The response is not a better photograph. It is one small pane of real glass, close enough to touch, showing almost nothing except whether the lights are still on.