Four 3090s and a bill for about $6,900, against an hourly rate and somebody else's cooling problem. This is the arithmetic I actually did, including the costs that don't appear on either quote.
14 April 2026·13 min read
Every conversation about AI infrastructure eventually reduces to the same question, whether it's a homelab or a Series B company: do we own the hardware or do we rent it. The answers people give are usually about identity rather than arithmetic. Cloud people say rent, infrastructure people say own, and both cite TCO without showing the working.
I've now made this decision with my own money, which sharpens it. What follows is the model I used, the numbers I got, and the places where I think the conventional answer is wrong in both directions.
The short version, so nobody has to scroll: for inference the ownership case is much stronger than people assume, and for training it is much weaker. Most of the argument on the internet is bad because it doesn't separate those two workloads.
Start with renting, because it's the option that looks obviously sensible and mostly is.
The strongest argument for renting, and it's very strong. A training run wants eight H100s for eleven hours and then wants nothing for three weeks. Owning that capacity means owning a very expensive space heater for the other 97% of the month. Renting turns a capital problem into a usage problem, and that's the right shape for bursty work.
Power delivery, cooling, driver compatibility, dead hardware at 2am, and the fact that the interesting failures in GPU infrastructure are thermal and electrical rather than computational. Renting is partly a purchase of not having to care about any of that.
H100s and their successors are not a thing an individual acquires, and interconnect is the real differentiator at scale. If the workload really needs NVLink-class bandwidth between eight cards, that is a rental-only capability regardless of budget.
This is the cost people underestimate, and it's not financial. When the meter is running you ration. You think twice before pointing an agent at a big refactor, you don't leave an experiment running overnight to see what happens, and you batch things that should be interactive. That hesitation has a real cost in what you learn, and it never shows up on the invoice.
Now the other side, and specifically the costs that don't appear on the parts list.
My build came to about $6,900 for four 3090s, an EPYC platform, 512 GB of RAM and the rest of it. That number is knowable in advance and it is the only one most people budget for.
Four cards at full board power is roughly 1.7 kW of GPU before the rest of the machine draws anything. I run them power-limited to 280 W for inference, where the throughput loss is barely measurable, precisely because the difference between capped and uncapped is a monthly bill rather than a benchmark.
The part nobody warns you about. The room hit 105 °F on a long run, which turned into a portable air conditioner, a thermostat and a Home Assistant automation driven by GPU load. Every watt that goes into the cards comes back out as heat that something has to remove, and that something also draws power. Owning compute means owning a thermal problem, and I budgeted zero for it.
Two dedicated 20 A circuits, because the rig doesn't plug into a wall outlet in any responsible sense. A one-time cost, but a real one, and one that doesn't appear in a single build-vs-buy spreadsheet I've ever seen.
Consumer GPUs are a depreciating asset in a market where the depreciation schedule is set by whatever NVIDIA announces next. The 3090 held value unusually well because of its VRAM, which was luck rather than judgment. Model this as "worth meaningfully less in two years" and the ownership case gets harder, honestly.
Building it, tuning it, fixing passthrough, chasing driver issues, and the evening spent working out why a card wouldn't initialise. I enjoy this, which is why I discount it, and discounting it is exactly the error I'd flag in someone else's business case.
Everything above treats this as an arithmetic problem, and for a homelab it mostly is. In a business it isn't, because there's a term most build-vs-buy models leave out entirely: what the data is worth to someone who shouldn't have it, and what it costs you when they do.
Renting compute means your data goes somewhere else to be processed. That is not automatically bad, and the major providers are almost certainly better at security than you are. But it does two things to your risk profile that a spreadsheet of hourly rates will not show you. It increases the number of parties who could be breached on your behalf, and it moves part of your control from something you operate to something you contract for.
So the model needs a second calculation alongside the cost one. Not "what does inference cost per hour", but "what is the expected cost of this data being exposed, and how does each option change the probability".
Every dataset gets a sensitivity tier before anything else happens. Public and synthetic data can go anywhere and the decision is purely economic. Internal and commercially sensitive data, meaning source code, roadmaps, pricing models and customer lists, sits in the middle, where contracts and controls do the work. Regulated data is its own category: protected health information, payment data, anything with a statutory notification duty attached. The tiers matter because the right answer is different for each, and a blanket "we use the cloud" or "we keep it in-house" policy is wrong for most of the estate.
This is where people wave their hands, and there is no need to. IBM's 2026 Cost of a Data Breach study puts the global average at $4.99 million, up 12% and the highest in the study's 21-year history, with the US average at $11.5 million. Healthcare has been the costliest sector for thirteen consecutive years, averaging $6.64 million per incident, and the verified baseline works out around $398 per record. Multiply that per-record figure by the size of your dataset before you decide that an inference endpoint is a procurement decision.
The usable form is probability times impact. If a breach costs $6.6 million and an architecture moves your annual likelihood by even a fraction of a percent, the expected cost of that architectural choice runs to tens of thousands of dollars a year. Set that against the difference between renting and owning and the comparison changes shape entirely. It also cuts both ways honestly: if owning means the data sits on a box you patch irregularly, in a room with physical access you haven't thought about, the expected cost may well be higher in-house.
The tier where the calculation stops being economic. When I was building healthcare software, protected health information leaving our controlled environment was not a cost to be weighed. It was a line, and the reason was contractual, regulatory and ethical before it was financial. If the answer is "this data cannot go to a third party", you have not made a build-vs-buy decision at all. You've had one made for you, and the only remaining question is how to own it well.
A BAA, a data processing agreement, a no-training-on-your-data clause, SOC 2 reports and a defined subprocessor list are all real risk reduction and should be treated as such. They are also a promise about behaviour rather than a technical constraint on it, they extend to subprocessors you did not choose, and their remedy after a breach is commercial rather than preventive. Owning the hardware converts a contractual control into a physical one, and that difference is worth real money in specific circumstances even though it never appears in a rate card.
Scaled down, this is exactly why my rig exists. Code and data never leave the network, nothing is logged by a third party, and nothing trains on my prompts. The stakes are obviously smaller than a hospital system's. The structure of the reasoning is identical, and I'd rather practise it on something where the worst case is my own inconvenience.
The point isn't that renting is unsafe. It's that "cheaper per hour" and "cheaper in expectation" are different quantities, and only one of them appears on the invoice you're comparing. A build-vs-buy analysis that prices compute and ignores data classification is a shopping list wearing a business case's clothes.
Any honest build-vs-buy in 2026 has to deal with the fact that the rental side is not standing still, and the clearest evidence of that is AWS's own silicon. This section is longer than the rest because it is the single strongest argument against owning, and I would rather make that argument properly than wave at it.
AWS made Trn3 UltraServers generally available at re:Invent in December 2025. The chip is on a 3 nm process with NeuronCore-v4 cores, and per chip it carries 2.52 PFLOPs of FP8 compute, 144 GB of HBM3e and 4.9 TB/s of memory bandwidth. Chip-to-chip is NeuronLink-v4 at up to 2.5 TB/s aggregate per device. Sit with the memory number for a second: a single Trainium3 chip has 144 GB of HBM. My entire four-card rig has 96 GB of GDDR6X, split into four islands that can't pool.
A fully configured Trn3 UltraServer connects 144 chips into one integrated memory domain, up from 64 in the previous generation, aggregating 362 FP8 PFLOPs, 20.7 TB of HBM3e and 706 TB/s of memory bandwidth. One memory domain is the phrase to notice. The central constraint on my hardware is that 96 GB is four separate 24 GB islands and anything that doesn't fit a shard pays the interconnect tax. AWS is selling the exact thing consumer hardware structurally cannot provide, and above that sits EC2 UltraClusters 3.0, scaling to a million chips.
AWS puts Trn3 UltraServers at up to 4.4x the compute of Trainium2 UltraServers, 3x higher throughput per chip and 4x lower response times on OpenAI's GPT-OSS workloads, with roughly 40% better energy efficiency. These are vendor benchmarks on vendor-chosen workloads and should be read as such. The direction is not in doubt even if the multiples are generous.
Pricing is where this gets murky, and I would rather show the spread than fake precision. Published figures for the previous generation trn2.48xlarge range from about $8.60 per hour on demand down to roughly $4.80 depending on source, region and commitment. Comparable H100 capacity on p5.48xlarge is quoted anywhere from around $55 to nearly $100 per hour for eight GPUs. AWS's own claim is 30 to 40% better price-performance than P5, with a more aggressive pitch of H100-equivalent throughput at a quarter of the cost on specific workloads, and third-party analyses land at 50 to 70% lower cost per billion tokens against H100 clusters. I could not find published on-demand pricing for Trn3 in what I read, which is itself worth knowing before you plan around it.
AWS teased Trainium4 alongside the Trainium3 launch, and the preview figures are the reason this section exists. At least 6x the FP4 performance of Trainium3, 3x the FP8 performance, four times the HBM bandwidth and twice the HBM capacity. More interesting strategically, it integrates with NVIDIA NVLink 6 and NVLink Fusion and fits NVIDIA MGX rack architectures, with NVLink Fusion offering all-to-all connectivity across up to 72 accelerators at 3.6 TB/s per accelerator. No launch date and no pricing have been disclosed, so anyone quoting Trainium4 economics is guessing.
The counterweight, and it is real. Trainium runs through the Neuron SDK. PyTorch and TensorFlow are supported, but getting good performance means adapting code to that stack, and that adaptation is an asset that does not travel. Choosing Trainium for the price-performance means accepting a migration cost if you later leave, and that cost belongs in the same spreadsheet as the hourly rate. Owning commodity NVIDIA hardware is, among other things, a bet on portability.
Here is what all of that does to my argument, and it is not comfortable. Every hour my rig sits in that room, it is exactly as capable as the day I built it. The rental side is on a curve: Trainium2 to Trainium3 in a generation, Trainium4 promising six times the FP4 throughput of a chip that only shipped in December. Owning fixes your capability at the purchase date. Renting rides the improvement, and you get it without buying anything.
That is the honest case against the thing I did, and for a training workload I think it wins outright. A single Trainium3 chip has 50% more memory than my entire build, in one address space, with 4.9 TB/s of bandwidth behind it. I cannot get there. Nobody gets there with consumer cards, at any quantity, because the limit is interconnect and memory architecture rather than money.
What it does not change is the inference case, and that is the whole reason the decision still holds. My duty cycle is steady, my data does not leave the network, and $0 per token does not get cheaper when someone else's silicon does. The Trainium curve makes renting better and better for the workload I was never going to run at home, which is a good outcome. It just isn't an argument about the workload I do run.
With those in hand, the comparison stops being one number against another.
The naive version is straightforward. Take the hourly rate for comparable rented capacity, divide the build cost by it, and get a break-even in hours. Do that for inference-class workloads and the payback lands in months rather than years, and that's the number people quote when they argue for owning.
The honest version adds power at your actual rate, cooling at roughly the same again, amortised electrical work, depreciation over a realistic horizon, and some value for your own time. Doing that pushed my break-even out meaningfully, and it still cleared, because the thing I do most is inference and I do a lot of it.
But the sensitivity is what matters, not the result. The ownership case is dominated by utilisation. At high, steady utilisation owning wins comfortably. At bursty utilisation renting wins comfortably. There is no clever framing that escapes that, and any TCO argument that doesn't lead with duty cycle is selling something.
So the useful output is a decision rule rather than a number.
Constant, predictable, unglamorous token generation is the ownership sweet spot. A fixed cost you've already paid, no meter, no rationing, and the freedom to leave things running. This is most of what an individual or a small team actually does with models day to day.
The argument that has nothing to do with money and often decides it anyway. Regulated data, client code, anything where "it went to a third party's inference endpoint" is a conversation you don't want to have. In healthcare work this alone would have settled it.
Training is bursty, interconnect-bound, and the exact profile that punishes owning. Consumer cards make it worse: no NVLink means gradient traffic crawls over PCIe, so multi-GPU training scales poorly no matter how many cards you buy. For a serious training run, rent the H100s by the hour and be glad you did.
If you don't yet know the shape of the workload, buying hardware is buying a guess. Rent until the duty cycle is boring, then buy against a known number. Every ownership case I've seen go wrong was built on a projected workload rather than an observed one.
The least quantifiable and possibly the most valuable. Owning the hardware taught me things about quantization, VRAM, interconnect and thermals that I would never have learned renting, because renting hides exactly those constraints. That knowledge transfers directly to making better rental decisions.
The conclusion I'd defend is narrower than the one people usually want. Owning GPUs bought me unlimited private inference at a fixed cost, and it did not buy me a training cluster. Being clear-eyed about which of those you're purchasing is most of the decision.
And the thing I got wrong is worth stating plainly, because it's the same mistake twice: I modelled the hardware and ignored the environment. Power, cooling, circuits. Every one of those was a surprise, and none of them should have been, because they're the first questions I'd ask about a rack at work.