Regulated Data and the Case for Local Inference

Regulated Data and the Case for Local Inference

Almost every argument about running models locally is an argument about cost. For regulated data the question is whether you have converted a promise into a constraint, which is a different question with a different answer.

The Conversation Everyone Is Having

Every organization holding sensitive data is having the same conversation right now. The technology is obviously useful. The default way to use it is a hosted API. And somewhere in the second meeting, someone from legal or compliance asks where the data goes, and the room discovers it does not have an answer it can defend.

What usually happens next is one of two bad outcomes. The organization freezes, decides AI is a problem for next year, and watches its people quietly paste things into consumer chat tools anyway. Or it moves ahead on the strength of a vendor's assurances that nobody in the room is qualified to evaluate.

I have been on both sides of this. I built and ran a healthcare platform where protected health information was the whole product, and I now run a rack of GPUs that serves models on my own network. A regulated organization should hear the argument with the parts that make it a worse pitch left in.

What Local Means Here

When I say local, I mean inference running on infrastructure the organization controls and administers. A server in your datacenter. A rack in a colocation cage. A private cloud tenancy where you hold the keys and the network boundary is yours. Distance has nothing to do with it. The defining property is that your organization is the accountable operator, and can say so to an auditor without pointing at a third party.

I do not mean a laptop. Unsanctioned local inference by an engineer is shadow IT with a GPU in it, and it is worse than the hosted API it was meant to replace. The data is now on an endpoint nobody is monitoring, outside backup, outside access control, outside logging, and it leaves the building every evening in a bag.

That distinction matters because the two get conflated constantly, usually by people advocating for one and objecting to the other. An organization that responds to this argument by letting individuals run their own models has not moved the boundary inward. It has dissolved it.

Local means centrally operated. Somebody's name is on it. Patching, capacity, access, monitoring, and backup are owned functions with a budget, the same as any other production system. If nobody owns it, it is a science project holding regulated data.

Local does not mean unrestricted. Moving inference inside the boundary does not remove the need for controls, it relocates them. Who can query which model against which dataset is still an authorization question, and you own that question now, where a vendor owned it before.

Local is not the same as air-gapped. Almost nobody needs a fully isolated network, and pretending you do is how projects die. The realistic target is that regulated data and the inference touching it stay inside a boundary you administer, while the rest of the environment works normally.

My own lab is not the recommendation. Worth stating plainly since it is elsewhere on this site. A rack in my house is where I learned what these models can and cannot do. The rack is not a template for how an organization should run this, and I would fail my own audit on physical access alone.

A Contract Is Not A Boundary

A business associate agreement is a real instrument. So is a data processing agreement, a no-training-on-your-data clause, a SOC 2 report, a documented subprocessor list. I am not dismissing any of it. If you are using a hosted service with regulated data, every one of those is necessary and you should be reading them closely instead of filing them.

But be precise about what they are. They are promises about behavior, enforced after the fact, with a commercial remedy. They are not technical constraints on where bytes can travel.

The promise extends to parties you did not choose. Your provider has subprocessors. Those subprocessors have their own arrangements. The contract chains, and the diligence you did on the vendor is not diligence you did on their infrastructure partner. Each link is probably fine. You are trusting a chain whose length you do not control.

The remedy arrives after the event. Contractual protection is compensation, not prevention. If regulated data is exposed, the notification duty is yours, the reputational cost is yours, and the affected people are yours. A clause does not un-disclose anything.

It changes when the vendor changes. Terms get revised. Companies get acquired. Free tiers get withdrawn and pricing models get rewritten, sometimes at short notice. I have written a whole other post about learning that lesson from a hypervisor vendor, and the stakes there were my weekends rather than somebody's medical records.

Not a person in the room can verify it. You are relying on an attestation about controls you cannot inspect, in an architecture you cannot see, operated by people you will never meet. That is a reasonable trade for most workloads. It is a specific and different trade for the ones where being wrong is a reportable event.

Enforcing It, In Practice

In theory data classification is a tidy exercise: public, internal, confidential, regulated. In practice the taxonomy is the easy half. The boundary has to be enforced by people making small decisions quickly, most of whom are not thinking about compliance at the moment they make them.

A developer wants to debug a production issue and needs a realistic record to reproduce it. Someone wants to run an analysis and the fastest tool is the one that uploads a file. Support needs to understand what a customer is seeing. None of these people are careless. They are trying to do their jobs, and the boundary is in the way of the shortest path.

The conversation I had most often, as the person accountable for it, was some version of "can I just use the real data for this". Always from somebody with a legitimate reason and a deadline, usually correct that it would be faster. Saying no once is easy. Saying no consistently, for years, while also being the person who set the roadmap those people were trying to hit, is where a boundary actually lives or quietly stops existing.

Running that platform taught me one thing. A boundary only holds if the compliant path is also the convenient one. Every control that makes the right thing harder than the wrong thing is a control that will be routed around, politely, by good people, within about a month. The people routing around it are behaving well, and the design failure is in the control.

The argument for local inference rests there, and it has nothing to do with hourly rates. If the model runs inside the boundary, using it is not a decision anybody has to make carefully.

What Open Models Can Actually Do

I keep a running log of every model I have deployed since the rig came online, with expected and measured throughput on the same four-GPU configuration. The short version is that the gap to frontier closed faster than I expected, and it closed unevenly.

Structured extraction is solved for practical purposes. Pulling fields out of documents, normalizing messy text into a schema, classifying and routing. Extraction is the highest-value regulated-data use case in most organizations. Open models on modest hardware reach it comfortably. If your problem is "we have twenty years of documents and no structure", you do not need frontier capability.

Agentic coding got good, and I did not expect it to. Local models that plan and ship code across a dozen tool calls, on hardware you can buy. Mid-2025 is when that stopped being a demo for me. For an organization that cannot send its source code to a hosted model, this is the difference between having the tooling and not.

Summarization and drafting are fine, with review. Good enough to save real time, not good enough to be unsupervised, and that is the same standard you would apply to a hosted model anyway. The compliance win is that the sensitive draft never leaves.

Frontier reasoning is still frontier. The hardest analytical work still favors the largest hosted models and probably will for a while. Anyone telling you otherwise is selling something. Ask instead what fraction of your actual workload needs frontier. In most regulated organizations the answer is a small one.

The Hardware, Accurately

The entry price is a department budget. My build came to roughly $7,900 for four consumer GPUs and a server platform. The figure is not nothing. It is comfortably inside the discretionary spend of a mid-sized team. The barrier to running serious inference on your own hardware is no longer money.

VRAM does not pool, and that is the constraint. Four 24 GB cards are four islands. A model has to fit its shard or pay an interconnect tax, and on consumer hardware that interconnect is slow. This determines which models you can run far more than raw compute does.

Training is a different purchase entirely. Without fast card-to-card links, multi-GPU training scales poorly. Inference barely notices, because pipeline-parallel only passes small activations between stages. Buy this class of hardware to run models. If you need to train something substantial, rent it, and rent it from somewhere your data classification permits.

The environment is a real cost nobody budgets. Power, heat, and the electrical work to support both. I have written about discovering this the expensive way. In a datacenter this is somebody's job. In an office closet it is a fire code question, and it deserves an answer before the hardware arrives.

Nothing Ships Without A Harness

In regulated software the rule I would not negotiate is that nothing ships without a suite that exercises the happy path and, more importantly, the edges. That is uncontroversial for deterministic code. Given this input, assert that output. Testing is the cheapest insurance in the industry and I have watched it save a clinical calculation from a misplaced parenthesis, which is a sentence I get to write only because the suite existed.

Model pipelines break that assumption, and most teams discover it after they have shipped one. The same input does not reliably produce the same output. You have no assertion to write. The failure mode is a plausible answer that is wrong, delivered with exactly the same confidence as a correct one. Traditional testing is built to catch absence, and this is a system whose characteristic failure is confident presence.

What you build instead is a validation harness, and it costs real engineering time that belongs in the plan from the start, because the later sprint never arrives.

A corpus with known-good answers comes first. Real inputs, drawn from your actual data distribution instead of invented examples, with outputs a qualified human has verified. The corpus is tedious to build and it is the entire foundation. Without it you have opinions about whether the system works. In a regulated setting, opinions are not the standard.

Edge cases have to encode how it goes wrong. Malformed input is the easy half. The ambiguous record, the one with a field that means two things, the document that is legible to a person and not to a parser, the case where the correct answer is "I cannot determine this". That last category matters most, because a model will confidently answer a question it should have declined, and a suite that only tests answerable cases will never catch it.

Tolerance replaces equality in the test. Assertions become ranges, distributions, and acceptance thresholds. Deciding what "close enough" means is a domain question, not an engineering one, and it needs the person accountable for the clinical or financial meaning of the output in the room when it is set.

A regression run is tied to every model change. Swapping in a newer open model changes a system that produces regulated output, whatever the release notes call it. My own log shows models replaced as better weights land, which in a lab is the point and in a regulated environment is a governance event. Better benchmarks are not evidence about your corpus. Re-run the harness or do not ship it.

You need provenance you can hand to an auditor. Which model, which version, which prompt, which parameters, against which input, producing which output, on what date. Someone will question a decision two years from now. The answer has to be reconstructable. Running the model yourself is what makes this recordable at all, and it is also what makes it your obligation to record.

The argument starts costing money here. Everything above about boundaries and control is a reason to bring inference in-house. Validation is the bill that arrives afterwards, and an organization that takes the control benefit without funding the validation work has not improved its position. It has moved the risk somewhere it can no longer point at a vendor.

The Audit Argument

One benefit never appears on a pricing page, and I think it is underrated by everyone who has not sat through an audit.

"Where did this data go" is a question you will eventually be asked. In writing, by someone entitled to a complete answer. When inference happens on hardware you own, inside a network you control, the answer is a sentence. When it happens through a hosted API, the answer is a diagram, a set of contracts, a subprocessor list, and an attestation about controls you have not personally verified.

Both answers can be acceptable. Only one of them is short. In an environment where you are periodically required to demonstrate rather than assert, the short answer has a value that compounds every time you are asked for it.

The same logic applies to the breach arithmetic. IBM's 2026 study puts the global average cost of a breach at $4.99 million and the US average at $11.5 million, with healthcare the costliest sector for the thirteenth consecutive year at $6.64 million. Multiply the per-record figure by your dataset before deciding that where inference happens is a procurement question.

What I Would Actually Do

Classify before you architect. Answer "which data, which tier" per system, in writing, before anyone evaluates a tool. The answer will be different for different parts of your estate, and a single organization-wide policy will be wrong for most of it. This costs a week and saves a rebuild.

Start with one bounded use case. Not a platform. One workflow, on one dataset, with a measurable outcome. Document extraction is usually the right first choice: high volume, clear success criteria, and squarely within what local models do well.

Match the deployment to the tier. Public and synthetic data has no reason to be local. Regulated data may have no option to be hosted. Most organizations need both, and the posture to hold is a boundary you can articulate rather than a side you have picked.

The boundary moves inward, it does not disappear. The mistake I would most expect an organization to make here. Bringing inference in-house is not permission to let everyone query everything. You still need authorization on who can run what against which data, you still need the queries logged, and you now need those things because you are the operator rather than because a vendor promised them. Done properly this is better than what a hosted arrangement gives you, since the logs are yours. Done carelessly it is a single system with access to everything and no record of who asked it what.

Own it properly or do not own it. The counterweight, and I would say this first in any room where people are getting enthusiastic. Running your own inference means owning patching, physical security, capacity, availability, and the on-call that comes with it. And the failure mode to watch for is the well-meaning engineer who solves this for themselves with a workstation under a desk. The workstation is a compliance finding wearing the costume of an initiative. If you cannot resource it to the standard you would demand of a vendor, the vendor is the better answer and the contract is the control you have.

The version of this argument I distrust is the one that says local is safer. Local is differently exposed. The exposure moves from a party you cannot inspect to one you can, which is only an improvement if you actually inspect it.

What I would defend is narrower. For a specific and growing class of workloads, open models are now good enough that the constraint on using AI with sensitive data is architecture and will, where it used to be capability. That was not true two years ago, and I think a lot of organizations are still making the decision they made when it was.

I did not arrive at this from a strategy deck. I built the platform where the boundary mattered, and then years later I built the rack that could sit inside one. The argument is the same in both directions: put the capability where the data already is, and you stop asking people to be careful.