In-Home Healthcare
Twelve GPUs, one spare, and a signed BAA
- Cards
- RTX PRO 6000 Blackwell, 96 GB
- Per GPU
- 8 cores, 64 GB RAM, 4 TB mirrored
- Shared storage
- 10 TB for logs and cache
- Nodes
- 2 to 3
- Spare
- 1 GPU, covered by SLA
- Compliance
- BAA signed through the facility
This customer is anonymized. A company handling protected health information does not need its infrastructure described in public with its name attached, so the configuration below is given in round terms and the business is described only by what it does.
A language model over records that cannot leave the building
The customer delivers medical care in patients’ homes. Clinicians see patients, notes get written, and records accumulate. The workload is a language model applied to that document set, and protected health information sits in every one of those documents.
The load is steady rather than bursty. This is not a team that fine-tunes a model once a quarter and wants a large cluster for four days. It is the same inference pipeline running against the same kind of document every working day, which means there is no peak worth renting and a floor that never drops.
Their requirements arrived as fixed ratios repeated per card: eight CPU cores, 64 gigabytes of RAM and four terabytes of mirrored storage behind each GPU. A specification that precise usually means the workload has already been measured somewhere else, and in most cases we have seen it also means the metered bill came back higher than expected.
Dedicated hardware, a spare card, and a complete paper chain
The environment is built to the specification the customer arrived with, in our Colorado facility, and it is running today.
The spare deserves a note of its own, because it is the line that a metered cloud cannot answer. The request was not for high availability or redundancy. It was for a specific, countable, thirteenth piece of hardware sitting in a rack, powered on, doing nothing, waiting for one of the working cards to fail. You can buy capacity from a metered provider, and capacity is an abstraction. A customer who asks for a spare card by number is a customer who has to explain their infrastructure to an auditor.
We are not publishing throughput or latency figures from this deployment. The other workloads on this site carry measured numbers because they are ours to measure and ours to discuss. This one belongs to a customer whose data we are contractually careful with.
How a private model satisfies the requirements
A Business Associate Agreement is the contract that lets a healthcare provider hand protected health information to a vendor. Reading theirs against the architecture, each obligation resolves the same way: the answer is easier when the model runs on hardware the customer alone occupies.
Protected health information never reaches a third party
The model runs on the customer’s own dedicated cards. There is no inference API in the request path, so there is no vendor retention policy to rely on and no second company to hold to a promise.
Every subcontractor with data access is under the same terms
A colocation provider with physical access to servers holding this data is a business associate under the rules, not a neutral pipe. The narrow exception for pure transmission services does not cover somebody who can open the cabinet. The facility signed before we did.
The hardware holding the records can be named
An auditor can be told which physical machines held the data and what happened to them. Metered capacity cannot answer that question, because a scheduler moving a job between anonymous hosts is the entire point of the product.
Data is returned or destroyed on termination
Within thirty days, with written certification. Dedicated hardware with a known disk inventory is what makes that certification something we can actually sign.
Breach notification and ongoing security program
Notification inside fifteen business days, a written security program, an annual risk assessment, and cyber liability coverage held for the full term of the agreement.
Nobody is HIPAA certified. There is no certificate, and the Office for Civil Rights certifies no one. Any vendor who tells you they are “HIPAA certified” is telling you something that cannot be true, which makes a useful early filter. What can be verified is whether a vendor will sign a Business Associate Agreement, and whether their facility will sign one too.
A Business Associate Agreement requires every subcontractor who can touch the data to sign the same terms, which includes anyone with physical access to the cabinet. That obligation, plus an auditor who needs to be told which machine held the records, rules out capacity you cannot point at.
Why per-GPU-hour is the wrong number
The first numbers most buyers find are per-GPU-hour rates, and comparing on that line alone is misleading. A bare card rate is a card. This workload also needs 96 CPU cores, 768 gigabytes of RAM, 58 terabytes of storage and a network connection. Providers bundle those very differently, and the cheapest headline rates are usually the ones that bundle the least.
The same environment, priced three ways on published rates:
| Cost element | Public cloud A | Public cloud B | Owned hardware |
|---|---|---|---|
| Twelve GPUs, CPU and RAM | $37,800/mo | $48,200/mo | Fixed |
| Storage, 58 TB | $5,100/mo | $4,200/mo | Included |
| Egress, 25 TB/mo | $2,200/mo | $2,200/mo | $0 |
| Egress, 100 TB/mo | $8,000/mo | $7,900/mo | $0 |
Priced August 2026 against published list rates for comparable configurations. Cloud pricing moves, so treat these as a snapshot rather than a quote.
Two things fall out of that table. The bundled resources are about a third of the bill, so any comparison that skips them is off by a third. And egress is the only line that grows without a purchase order, which is the one that tends to make finance uncomfortable.
Have a workload shaped like this one?
A resident open-weights model on a card you don't share is a different cost and privacy story than a per-token API. Tell us what you're running and we'll size it.