Intel's 13th and 14th generation desktop parts pass acceptance testing and then degrade in service, and the remedy on offer this month is two extra years of warranty rather than a recall. This is about what that does to fleet planning: every spares pool and redundant pair assumes units fail independently, and a defect in the design means they do not. They are reading the same clock, started on the same day.
9 August 2024·7 min read·homelabdependencies
A machine starts crashing, so you look at what changed. That is the training and it is good training: a driver, a kernel, a firmware level, the thing that went out on Tuesday. Nothing changed. You pull the memory and swap it, run the burn-in overnight, and it passes. You run it again against the real workload and it fails in forty minutes, with a decompression error naming a file that is not corrupt.
That has been a reproducible experience on Intel's 13th and 14th generation desktop processors since the winter. It surfaced in games built on Unreal Engine and in video encoding, which is why it spent months filed under graphics drivers, since a shader compilation error is a message the GPU emits. In April, Nvidia put a note in a driver release telling anyone seeing it to contact Intel, which is an unusual thing for one vendor to say about another.
The distinguishing property took longer to accept than it should have. These processors are not defective when they arrive. They become defective. Intel's statement on 22 July named exposure to elevated operating voltage, requested by the processor through its own microcode, as the root cause. Asked directly by The Verge on 26 July, Intel did not deny that the damage already done is permanent.
So here is the sentence the rest of this is about. Acceptance testing tells you that a part works, and it can tell you nothing whatsoever about a part that works now and will not in six months. Every process I have seen for putting hardware into service answers the first question and is quietly assumed to have answered the second.
It looked like a software problem for months. Reports collected through late February of K-series Core i7 and i9 parts crashing in DirectX 12 titles and in HandBrake. The early theories were motherboard power settings and graphics drivers, both reasonable, because the failures were intermittent and every machine was configured differently. Nothing in the reporting looked like silicon, and a plausible reading was that some enthusiasts had overclocked their machines and were surprised.
April: the vendors stop pointing at each other. Nvidia's driver note sends affected users to Intel, and by 18 April Intel has a root-cause investigation running. Motherboard vendors ship an Intel Baseline Profile to pull power limits back toward specification. It helps some people and not others, which in retrospect is the tell: a workaround that partly works is usually addressing a contributing factor rather than a cause.
June and July: the scope keeps widening. Microcode 0x125 lands in June against a bug in the Enhanced Thermal Velocity Boost algorithm, and Intel says plainly that it is a contributing factor rather than the root cause, which is an unusual thing to say while shipping a fix. By 11 July the same failures are being reported on W680 boards in game server fleets.
Late July: the boundary is drawn far wider than expected. On 22 July Intel names elevated voltage and targets a production microcode patch for the middle of August. Four days later it tells The Verge that any 13th or 14th generation desktop part with a base power of 65 W or higher could be exposed, not only the top i9 parts everyone had been discussing, and confirms a separate via oxidation manufacturing issue corrected at an unspecified point in 2023.
August: the remedy is commercial rather than technical. Intel announces a two-year warranty extension on boxed parts on 1 August, extends it to tray and OEM processors on the 5th, and publishes the list of covered models on the 6th. There is no recall, no pause in sales, and no serial number range a buyer can check. The answer to an unrepairable hardware defect is a longer period in which you may ask for a replacement.
A digital circuit has a minimum voltage below which it stops producing correct results at a given frequency, and that threshold is not a constant. Sustained high voltage and temperature slowly change the transistors, the minimum the part needs creeps upward, and eventually it crosses the voltage the part is being given. Nothing melts and nothing shorts. The chip now needs more than it is getting, in exactly the conditions where it needs the most.
Which is why the symptom is so bad at identifying itself. A failed capacitor gives you a machine that does not post. A part sitting just under its required voltage gives you one compilation that fails, a checksum that does not match, a process exiting with a code nobody has seen. Those get attributed to software for months, correctly by the standards of ordinary debugging, because a wrong answer from a CPU is the hypothesis you are trained to reach for last.
And it is cumulative, which is the part that breaks testing outright. Every hour at elevated voltage moves the threshold further, so the part is worse this month than last and the acceptance test you ran on day one measured a different component from the one you now have. A soak test proves the part is above its threshold today. It says nothing about the rate at which the threshold is moving, and the rate is the entire defect.
The spares model assumes failures are independent. Every spares calculation I have seen, formal or in somebody's head, treats each unit as rolling its own dice. That is why a small pool covers a large population: the chance of many failing at once is a product of small numbers, and those are tiny. A systematic defect deletes the multiplication. The units are not rolling dice, they are reading the same clock, and the pool is sized for the wrong distribution.
Redundancy protects against the failure it was designed for. A second machine covers you when the first fails for reasons of its own. It covers nothing when both hold parts from the same production run, bought on one purchase order, running the same workload at the same duty cycle since the same week. That pair has been aging in lockstep by design, because buying identically is the procurement discipline everyone rewards, and identical purchasing is what turns independent risk into correlated risk.
The bathtub curve puts the risk in the wrong place. The received model says failures cluster at the start, from manufacturing defects, and again at the end, from wear. Burn-in exists to catch the first group and it works. This defect lands in the flat middle of the curve, where the model says the hazard rate is low and constant and where nobody schedules anything. Puget reports seeing these failures only after six months, which is precisely when a fleet stops paying attention.
A warranty is a clock, and it started before you noticed. The extension is two years on top of the standard three, which is real money and worth having. It is also measured from purchase, while the aging is measured in hours under load, and those two clocks run at different speeds for a machine that is busy. A part bought early in the cycle and worked hard has spent much of its degradation budget inside a warranty that spends itself in calendar time.
The alarming figures are secondhand and belong to specific populations. Puget Systems, writing on 2 August, notes that some game development studios and cloud gaming providers have described failure rates upwards of fifty percent. Those are workloads that hold a processor at high sustained load essentially forever, which is the exact condition the defect responds to. It is a real number for that population and it is not a number about the world.
Puget's own data is the more useful document, because it is a builder publishing its returns rather than a user reporting a bad month. It reports roughly five to seven field failures a month, elevated against its own history and, in its words, not a show-stopper. The comparison it draws is the one I did not expect: 14th generation failures running below what it recorded for 11th generation parts in 2021, which nobody wrote about at the time, and below Ryzen 5000 as it measured them.
That complicates what I have been arguing, and I would rather say so than route around it. Puget explains why its experience is muted: it distrusts motherboard defaults on principle and applies its own conservative power settings, at a cost it puts at one to two percent of performance. So the exposure was not fixed at the factory. It was largely set by a configuration decision made years earlier for unrelated reasons, and the organizations that had already given up that last two percent were buying a longer life without knowing it.
Read the warranty as a term, not a reassurance. The extension is the whole remedy on offer, so its details are the product: what it covers, how a claim is proven, how long a replacement takes, and whether the replacement is the same part with the same defect. Intel's answer on what proof an RMA requires was, as of late July, not yet given. A warranty with an undefined claims process is a promise with unbounded latency, and latency is what a fleet feels.
Run the parts inside the envelope, on purpose. The decision that most changed anyone's exposure was refusing motherboard defaults and holding to specified power and voltage limits. That is not exotic, it is boring configuration discipline, and it cost a couple of percent in benchmarks nobody would have noticed. It is worth asking which of your defaults were chosen by a vendor optimizing for a review score rather than by you optimizing for the machine still working next year.
Stagger what you buy, or know that you did not. Mixed sourcing is genuinely expensive: more part numbers, more images, more qualification, more spares. It is also the only thing that turns one bad production run into a partial outage instead of a total one. I would not run a mixed fleet everywhere, and I would not run a single-source one in the tier where a correlated failure is intolerable. Deciding which tier is which is the actual work.
Know your own exposure before you need to. Intel drew the line at 65 W base power and published a model list on 6 August. Answering "how many of ours are on that list" ought to be a query, and in most places it is an afternoon of walking around, because inventory systems record what was purchased rather than what is installed. The gap between those two is where every hardware advisory turns into a project.
I want to be careful about the size of the claim. This is one defect in one product line, the measured rates from the one builder publishing them are elevated rather than catastrophic, and an earlier generation apparently did worse without anyone noticing. Intel took months to find it because it is genuinely hard to find, and I do not think anyone in that seat moves much faster. The parts were sold to specification and most of them work.
What stays is the assumption underneath, which almost nobody writes down and everybody depends on: hardware fails one unit at a time, for reasons belonging to that unit, which is why a small spares pool and a redundant pair are enough. When the cause lives in the design rather than the individual, every part in the room is running the same countdown, started on the same day, and the second machine in the rack is not a spare. It is a copy of the first, six months older than it looks.