Own the GPUs, or Rent Them

Own the GPUs, or Rent Them

Four 3090s and a bill for about $7,900, against an hourly rate and somebody else's cooling problem. What follows is the arithmetic I actually did, including the costs that don't appear on either quote.

The Question

Every conversation about AI infrastructure eventually reduces to the same question, whether it's a homelab or a Series B company: do we own the hardware or do we rent it. The answers people give are usually about identity instead of arithmetic. Cloud people say rent, infrastructure people say own, and both cite TCO without showing the working.

I've now made this decision with my own money, which sharpens it. The model I used and the numbers I got are below, along with the places where I think the conventional answer is wrong in both directions.

The short version, so nobody has to scroll: for inference the ownership case is much stronger than people assume, and for training it is much weaker. Most of the argument on the internet is bad because it doesn't separate those two workloads.

The Case For Renting

The strongest argument for renting is that you pay for the peak instead of the average. A training run wants eight H100s for eleven hours and then wants nothing at all for three weeks. Owning that capacity means owning a very expensive space heater for the other ninety-seven percent of the month. Renting turns a capital problem into a usage problem. That is the right shape for bursty work.

It also buys you out of the hard parts. Power delivery, cooling, driver compatibility, dead hardware at two in the morning, and the fact that failures in GPU infrastructure are thermal and electrical, hardly ever computational. None of that is your problem when you rent, and all of it is when you do not.

Then there is hardware you cannot buy at any price. H100s and their successors are not something an individual acquires. Interconnect is the differentiator at scale. If a workload genuinely needs NVLink-class bandwidth between eight cards, renting is the only way to get it.

The cost I underestimated is not financial. When the meter is running you ration. You think twice before pointing an agent at a big refactor, you do not leave an experiment going overnight just to see what happens, and you batch things that ought to be interactive. That hesitation shapes what you learn, and it never appears on the invoice.

What Owning Actually Costs

The hardware is the smallest surprise in the whole exercise. My build came to about $7,900 for four 3090s, an EPYC platform, 512 GB of RAM, and the rest of it. That figure is knowable in advance, and it is the only one most people budget for.

Power is a real line item. Four cards at full board power is roughly 1.7 kW of GPU before the rest of the machine draws anything, which is why I run them capped at 280 W for inference. The throughput loss at that cap is barely measurable, and the difference between capped and uncapped shows up on the monthly bill instead of in a benchmark.

Cooling turned out to be the bigger line item. I budgeted nothing at all for it. The room hit 105 °F on a long run, which became a portable air conditioner, a thermostat, and a Home Assistant automation driven by GPU load. Every watt that goes into the cards comes back out as heat that something has to remove, and that something draws power of its own.

The electrical work happened once and was straightforward enough: two dedicated 20 A circuits, because the rig does not plug into a wall outlet in any responsible sense. One-time, real, and absent from every build-versus-buy spreadsheet I have ever read.

Depreciation is brutal, and almost nobody models it. Consumer GPUs depreciate on a schedule set by whatever NVIDIA announces next. The 3090 held its value unusually well because of its VRAM, which was luck, and I will not claim judgment for it, and if you model the cards as worth meaningfully less in two years the ownership case gets harder.

My own time is the largest hidden cost. It is the one I am least honest about. Building it, tuning it, fixing passthrough, chasing driver issues, and one entire evening spent working out why a card would not initialize. I enjoy all of that, which is why I discount it, and discounting it is exactly the error I would flag in somebody else's business case.

Data Sensitivity Is Part Of The Model

Everything above treats this as an arithmetic problem, and for a homelab it mostly is. In a business it isn't, because there's a term most build-vs-buy models leave out entirely: what the data is worth to someone who shouldn't have it, and what it costs you when they do.

Renting compute means your data goes somewhere else to be processed. Sending it out is not automatically bad, and the major providers are almost certainly better at security than you are. But it does two things to your risk profile that a spreadsheet of hourly rates will not show you. It increases the number of parties who could be breached on your behalf, and it moves part of your control from something you operate to something you contract for.

So the model needs a second calculation alongside the cost one. Not "what does inference cost per hour", but "what is the expected cost of this data being exposed, and how does each option change the probability".

Classify before you calculate, and do it before anything else happens. Public and synthetic data can go anywhere, and the decision there is purely economic. Internal and commercially sensitive data, meaning source code, roadmaps, pricing models, and customer lists, sits in the middle, where contracts and controls do the work. Regulated data is its own category: protected health information, payment data, anything carrying a statutory notification duty. The tiers matter because the right answer is different for each, and a blanket policy in either direction is wrong for most of an estate.

Then put a number on the downside. That is where people wave their hands, and they have no need to. IBM's 2026 Cost of a Data Breach study puts the global average at $4.99 million, up 12% and the highest in the study's 21-year history, with the US average at $11.5 million. Healthcare has been the costliest sector for thirteen consecutive years at $6.64 million per incident, and the verified baseline works out at around $398 per record. Multiply that per-record figure by the size of your dataset before deciding that an inference endpoint is a procurement question.

The usable form multiplies probability by impact instead of leaning on the worst case. If a breach costs $6.6 million and an architecture moves your annual likelihood by even a fraction of a percent, the expected cost of that choice runs to tens of thousands of dollars a year. Set that against the gap between renting and owning and the comparison changes shape entirely. It cuts both ways, too: if owning means the data sits on a box you patch irregularly, in a room whose physical access you have never thought about, the expected cost is probably higher at home.

Some data has no price. That is the tier where the calculation stops being economic. When I was building healthcare software, protected health information leaving our controlled environment was a line, and the reason was contractual, regulatory, and ethical before it was financial. If the answer is that this data cannot go to a third party, you have not made a build-versus-buy decision at all. You have had one made for you, and the only question left is how to own it well.

Contracts are controls, and they are not the same kind of control. A BAA, a data processing agreement, a no-training-on-your-data clause, SOC 2 reports, and a defined subprocessor list are all real risk reduction and should be counted as such. They are also a promise about behavior, with no technical constraint underneath it, they extend to subprocessors you did not choose, and their remedy after a breach compensates you without preventing one. Owning the hardware converts a contractual control into a physical one, which is worth real money in specific circumstances and appears on no rate card.

Scaled down, this is exactly why my rig exists. Code and data never leave the network, nothing is logged by a third party, and nothing trains on my prompts. The stakes are obviously smaller than a hospital system's, the structure of the reasoning is identical, and I would rather practice it somewhere the worst case is my own inconvenience.

The point is that "cheaper per hour" and "cheaper in expectation" are different quantities, and only one of them appears on the invoice you're comparing. A build-vs-buy analysis that prices compute and ignores data classification is a shopping list wearing a business case's clothes.

What Trainium Changes

Start with what Trainium3 actually is. AWS made Trn3 UltraServers generally available at re:Invent in December 2025. The chip is on a 3 nm process with NeuronCore-v4 cores, and per chip it carries 2.52 PFLOPs of FP8 compute, 144 GB of HBM3e, and 4.9 TB/s of memory bandwidth, with chip-to-chip NeuronLink-v4 running at up to 2.5 TB/s aggregate per device.

Sit with the memory figure for a moment. One Trainium3 chip carries 144 GB of HBM, and my entire four-card rig has 96 GB of GDDR6X split into four islands that cannot pool.

The UltraServer is the part that matters. A fully configured Trn3 connects 144 chips into one integrated memory domain, up from 64 in the previous generation, aggregating 362 FP8 PFLOPs, 20.7 TB of HBM3e, and 706 TB/s of memory bandwidth.

One memory domain is the phrase to notice. The central constraint on my hardware is that its 96 GB is four separate 24 GB islands, and anything that does not fit inside a shard pays the interconnect tax. AWS is selling the exact thing consumer hardware structurally cannot provide.

The performance claims need the appropriate salt. AWS puts Trn3 UltraServers at up to 4.4x the compute of Trainium2 UltraServers, three times the throughput per chip, four times lower response times on OpenAI's GPT-OSS workloads, and roughly 40% better energy efficiency. Those are vendor benchmarks on vendor-chosen workloads and should be read that way. The direction is not in doubt even if the multiples are generous.

Pricing is murkier, and I would rather show the spread than fake precision. The previous generation trn2.48xlarge runs from about $8.60 per hour on demand down to roughly $4.80 depending on source, region, and commitment, against $55 to nearly $100 for comparable H100 capacity on p5.48xlarge. AWS claims 30 to 40% better price-performance than P5, and third-party analyses land at 50 to 70% lower cost per billion tokens. I could not find published on-demand pricing for Trn3 anywhere I read, which is itself worth knowing before you plan around it.

Trainium4 is already the more important number. AWS teased it alongside the Trainium3 launch, and the preview figures are the reason this section exists: at least 6x the FP4 performance of Trainium3, 3x the FP8 performance, four times the HBM bandwidth, and twice the HBM capacity. More interesting strategically, it integrates with NVIDIA NVLink 6 and NVLink Fusion and fits NVIDIA MGX rack architectures, with NVLink Fusion offering all-to-all connectivity across up to 72 accelerators at 3.6 TB/s per accelerator. No launch date and no pricing have been disclosed, so anybody quoting Trainium4 economics is guessing.

The lock-in is the part nobody prices in. Trainium runs through the Neuron SDK, and while PyTorch and TensorFlow are supported, getting good performance out of it means adapting code to that stack. That adaptation is an asset which does not travel. Choosing Trainium for the price-performance means accepting a migration cost if you ever leave, and that cost belongs in the same spreadsheet as the hourly rate. Owning commodity NVIDIA hardware is, among other things, a bet on portability.

Every hour my rig sits in that room, it is exactly as capable as the day I built it. The rental side is on a curve: Trainium2 to Trainium3 in a generation, Trainium4 promising six times the FP4 throughput of a chip that only shipped in December. Owning fixes your capability at the purchase date. Renting rides the improvement, and you get it without buying anything.

The case against the thing I did is strong, and for a training workload I think it wins outright. A single Trainium3 chip has 50% more memory than my entire build, in one address space, with 4.9 TB/s of bandwidth behind it. I cannot get there. Consumer cards will not get you there at any quantity, because money is not the limit; interconnect and memory architecture are.

What it does not change is the inference case, and that is why the decision still holds. My duty cycle is steady, my data does not leave the network, and $0 per token does not get cheaper when someone else's silicon does. The Trainium curve makes renting better and better for the workload I was never going to run at home, which is a good outcome. It just isn't an argument about the workload I do run.

Doing The Arithmetic

The naive version is straightforward. Take the hourly rate for comparable rented capacity, divide the build cost by it, and get a break-even in hours. Do that for inference-class workloads and the payback lands in months rather than years, and that's the number people quote when they argue for owning.

The honest version adds power at your actual rate, cooling at roughly the same again, amortized electrical work, depreciation over a realistic horizon, and some value for your own time. Doing that pushed my break-even out meaningfully, and it still cleared, because the thing I do most is inference and I do a lot of it.

But the sensitivity is what matters. The ownership case is dominated by utilization. At high, steady utilization owning wins comfortably. At bursty utilization renting wins comfortably.

The Decision Rule

Own it when the work is steady inference. Constant, predictable, unglamorous token generation is the ownership sweet spot. A fixed cost you've already paid, no meter, no rationing, and the freedom to leave things running. Steady inference is most of what an individual or a small team actually does with models day to day.

Own it when the data cannot leave. The argument that has nothing to do with money and often decides it anyway. Regulated data, client code, anything where "it went to a third party's inference endpoint" is a conversation you don't want to have. In healthcare work this alone would have settled it.

Rent it for training, almost always. Training is bursty, interconnect-bound, and the exact profile that punishes owning. Consumer cards make it worse: no NVLink means gradient traffic crawls over PCIe, so multi-GPU training scales poorly no matter how many cards you buy. For a serious training run, rent the H100s by the hour and be glad you did.

Rent it while the requirement is still moving. If you don't yet know the shape of the workload, buying hardware is buying a guess. Rent until the duty cycle is boring. Then buy against a known number. Every ownership case I've seen go wrong was built on a projected workload instead of an observed one.

Own a small one to learn, regardless. Owning the hardware taught me things about quantization, VRAM, interconnect, and thermals that I would never have learned renting, because renting hides exactly those constraints. That knowledge transfers directly to making better rental decisions.

The conclusion I'd defend is narrower than the one people usually want. Owning GPUs bought me unlimited private inference at a fixed cost. It did not buy me a training cluster. Being clear-eyed about which of those you're purchasing is most of the decision.

And the thing I got wrong is worth stating plainly, because it's the same mistake twice: I modeled the hardware and ignored the environment. Power, cooling, circuits. Every one of those was a surprise, and none of them should have been, because they're the first questions I'd ask about a rack at work.