On Wednesday the eighth of March, Datadog lost its web application, its data ingestion, and its alert notifications across multiple regions at once, and roughly thirty hours passed before live data was back everywhere. A customer could know some things while that ran and not others. The fallback had a price. What I take from it is that the instrument must not share a failure domain with the patient, and that nobody prices a local fallback until the morning they need one.
12 March 2023·7 min read·operationsdependencies
The first entry on the status page went up at 01:31 Eastern and said that Datadog was investigating loading issues on its web application. Two hours later the same page reported delayed data ingestion across all data types, and the phrase that matters more than any of it: monitor notifications are delayed. Delayed notifications are the polite form of a sentence nobody wants to write. The alerts were not going out.
An instrument can fail two ways, and only one of them is the loud one. It can read wrong, which you notice, because the number is absurd and somebody says so. Or it can stop reading, which looks exactly like a quiet morning. No pager, nothing red, nothing to click. Silence is also what a healthy system emits, and no amount of screen real estate distinguishes the two.
So the duration is the least interesting number from that Wednesday. The question is what a customer could and could not know while it ran. If your own service had picked that morning to fall over, the first report would have come from a person, by whatever channel people use when they cannot reach you, and it would have arrived some unknown number of minutes after the fact.
The argument is one sentence long, and it is not really about observability vendors. The tool you use to see must not fail for the same reasons as the thing you are looking at, because the one property you were paying for is the property that goes missing exactly when you need it.
Each of these stamps comes from the status page. First entry at 01:31 Eastern on the eighth. The issue reported as identified at 07:08. All services able to receive and query live data at 08:00 Eastern on the ninth, so roughly thirty hours from first symptom to working instruments. Backfilling historical data ran on into the tenth with updates every two to three hours, which means the history took longer to come back than the present did.
Only one cause has been stated so far. The summary posted at 15:39 Eastern on the eighth says a system update on a number of hosts controlling the compute clusters caused a subset of those hosts to lose network connectivity, and that clusters then entered unhealthy states and took internal services and datastores with them. That is the sum of it, and no component is named, and four days later a customer deciding whether to change anything on their own side has one sentence to reason from.
Something is owed and has not landed. The same summary promises a more detailed analysis post-recovery. I have no complaint about that pace; a write-up worth reading takes weeks, and the alternative is a guess published early and corrected later. It does mean I am arguing from a paragraph. I have tried to make an argument that would not change when the detail arrives, and I will say plainly that this is a thing to be a little suspicious of.
The regions went together. The summary describes issues across multiple products and regions, and points customers at the separate status page for each region. Those regions run separate infrastructure. Being separate is the entire reason they are sold as separate. Whatever the cause turns out to be, it crossed the boundary that exists to stop things crossing, which is the part I would want explained first.
Consider what an alerting pipeline actually promises. It does not promise to tell you when things are bad. It promises to tell you when things are bad and the pipeline is working, and the second half is unstated because it is normally so reliable that nobody prices it. Delayed monitor notifications collapse the two cases into one output, and the output is nothing at all.
The one structure that catches this is a heartbeat, and it has to run somewhere else to work. You emit a signal on a schedule and something outside your estate complains when the signal stops. That check is cheap, it is boring, and it is the only alert in the whole system whose job is to notice that the other alerts have gone quiet. Very few teams have one.
A small piece of evidence sits in plain sight here. The status page everyone was refreshing that morning is Atlassian's product, on somebody else's infrastructure, with its own domain and its own notification path. That separation is deliberate, and it is the whole idea of this essay rendered in one design decision that the industry made years ago and then applied to exactly one page.
The second-order cost showed up after the lights came back. Customers were told to expect gaps in historical data for parts of the previous day, and gaps in history are not a cosmetic problem. Your capacity math for that week has a hole in it, your error budget for the month is unmeasurable, and any incident of your own that happened inside that window will be investigated without the data you would normally investigate it with.
The instrument runs where the patient runs. Most monitoring, hosted or otherwise, sits in the same handful of clouds as the thing it watches, reaches you over the same networks, and authenticates through the same identity provider. Each of those is a wire connecting the observer to the observed, and the picture only stays trustworthy while every wire holds.
And it boots the same operating system, from the same archive. Packages arrive from a distribution on the distribution's schedule. A package update is a change to your production estate, initiated by people who have never heard of you, applied to your monitoring and your application by the same tooling on the same day. Keeping hosts current is correct, and the correctness of it is not what determines the blast radius.
A shared change window is a synchronization primitive. If every region applies changes inside the same hour, you have built a clock. Its only function is to make failures simultaneous. Independence at the network layer means very little when the schedule above it is common. Staggering costs nothing except patience, and it is the cheapest correlation you will ever break.
Independence is asserted, rarely measured. The list rarely gets produced, so here it is: same image, same package source, same config service, same certificate authority, same time source, same deployment pipeline, same change window. Write it down for your monitoring against your production, and the answer is usually that the two are the same system wearing different names.
I would buy hosted observability again tomorrow, and I would not treat the choice as close. Self-hosting means you now operate a storage system with awkward retention characteristics, you carry a pager for the thing that carries your pager, and the day your cluster falls over is the day you have no telemetry about your telemetry. Carrying it yourself is a worse position, and it is worse on most days rather than one.
But the axis in that argument is wrong. Hosted against self-hosted decides nothing. What matters is whether any signal path exists that does not pass through the vendor, and that path is much smaller than a second observability stack. Two days of retention, a handful of vital signs, one notification route that shares nothing with the main one. A small instance and an afternoon.
It does not get built because it has no owner and no visible return. On a budget line it reads as a duplicate monitoring system, which is the easiest thing in the world to decline, and its value cannot be stated in advance because it is denominated in an outage that has not happened. The correct name for it is a flashlight, and nobody justifies a flashlight by its feature list.
The limit is that a fallback nobody looks at is not a fallback. If no one has opened it in six months, the credentials have rotted, the dashboard is empty, and the morning you reach for it is the morning you find out. So the test is whether anyone touched it recently, on purpose, when nothing was wrong. Existing does not count for much.
Alert on the absence of data. A heartbeat out to a service that is not your monitoring vendor, checked by something that is not your monitoring vendor. That heartbeat is the only alert that fires when the alerting is broken. It takes an afternoon. Everything else in this list is optional; this one I would not run without.
Put the last-resort path on different everything. Different provider, different account, different credentials, different notification channel. The value comes entirely from what it does not share, so every dependency it reuses for convenience is a dependency that takes it down with the main system. Convenience is the whole failure mode here.
Decide in advance what you will not do blind. The deploy you would hold, the failover you would not trigger, the scaling change you would not make with no visibility. Writing that list on a calm afternoon is most of the value, because the alternative is arguing about it at 03:00 with people who cannot see and know they cannot see.
Keep raw telemetry at the edge longer than feels necessary. If the agent buffers, give it room; if the host keeps logs, keep them a while. When the pipe is down for a day you want the day back, and the difference between recovering it and losing it is usually a retention setting somebody trimmed to save a little disk.
I want to keep this proportionate. This was a bad Wednesday, and scandal is the wrong size of word for it. A system update did something the change note did not describe, which has happened to every operator who has ever run a fleet, and the only thing separating the ordinary version from this one is how many machines were on the far side of it. Datadog posted more, and sooner, than most vendors would have, and the promised analysis is not late yet.
The part that stays is smaller. A dashboard is a photograph of a window, developed somewhere else and mailed to you, and on the eighth the mail did not come. The response is one small pane of real glass, close enough to touch, showing almost nothing except whether the lights are still on.