Menu
    GPU
    Infrastructure
    All use cases
    LIVE IN PRODUCTION · CUSTOMER ANONYMIZEDHealthcare · protected health information

    In-Home Healthcare

    Twelve GPUs, one spare, and a signed BAA

    At a glance
    Cards
    RTX PRO 6000 Blackwell, 96 GB
    Per GPU
    8 cores, 64 GB RAM, 4 TB mirrored
    Shared storage
    10 TB for logs and cache
    Nodes
    2 to 3
    Spare
    1 GPU, covered by SLA
    Compliance
    BAA signed through the facility
    A note on the customer

    This customer is anonymized. A company handling protected health information does not need its infrastructure described in public with its name attached, so the configuration below is given in round terms and the business is described only by what it does.

    The use case

    A language model over records that cannot leave the building

    The customer delivers medical care in patients’ homes. Clinicians see patients, notes get written, and records accumulate. The workload is a language model applied to that document set, and protected health information sits in every one of those documents.

    The load is steady rather than bursty. This is not a team that fine-tunes a model once a quarter and wants a large cluster for four days. It is the same inference pipeline running against the same kind of document every working day, which means there is no peak worth renting and a floor that never drops.

    Their requirements arrived as fixed ratios repeated per card: eight CPU cores, 64 gigabytes of RAM and four terabytes of mirrored storage behind each GPU. A specification that precise usually means the workload has already been measured somewhere else, and in most cases we have seen it also means the metered bill came back higher than expected.

    What Bit Refinery supplied

    Dedicated hardware, a spare card, and a complete paper chain

    The environment is built to the specification the customer arrived with, in our Colorado facility, and it is running today.

    Dedicated GPU nodes
    NVIDIA RTX PRO 6000 Blackwell cards, 96 GB each, across two to three nodes. Single tenant. No other customer is scheduled onto this hardware.
    Fixed resources per card
    Eight CPU cores, 64 GB of RAM and 4 TB of mirrored local storage behind every GPU, plus 10 TB of shared storage for logs and cache.
    A spare GPU under SLA
    A physical card, racked and powered on, held in reserve against a failure in the working set. Not a capacity guarantee, an object.
    Fixed monthly pricing
    One number that does not move with utilization, and no metered egress on inference traffic.
    A signed Business Associate Agreement
    Executed through the facility first, then with the customer, so the subcontractor chain is complete before any protected health information exists on the hardware.
    IN SERVICESPAREPowered onDoing nothingWaitingTwelve cards carrying the workload. A thirteenth an auditor can be shown.CAPACITY IS AN ABSTRACTION. A SPARE CARD IS AN OBJECT.

    The spare deserves a note of its own, because it is the line that a metered cloud cannot answer. The request was not for high availability or redundancy. It was for a specific, countable, thirteenth piece of hardware sitting in a rack, powered on, doing nothing, waiting for one of the working cards to fail. You can buy capacity from a metered provider, and capacity is an abstraction. A customer who asks for a spare card by number is a customer who has to explain their infrastructure to an auditor.

    We are not publishing throughput or latency figures from this deployment. The other workloads on this site carry measured numbers because they are ours to measure and ours to discuss. This one belongs to a customer whose data we are contractually careful with.

    Security and compliance

    How a private model satisfies the requirements

    A Business Associate Agreement is the contract that lets a healthcare provider hand protected health information to a vendor. Reading theirs against the architecture, each obligation resolves the same way: the answer is easier when the model runs on hardware the customer alone occupies.

    Protected health information never reaches a third party

    The model runs on the customer’s own dedicated cards. There is no inference API in the request path, so there is no vendor retention policy to rely on and no second company to hold to a promise.

    Every subcontractor with data access is under the same terms

    A colocation provider with physical access to servers holding this data is a business associate under the rules, not a neutral pipe. The narrow exception for pure transmission services does not cover somebody who can open the cabinet. The facility signed before we did.

    The hardware holding the records can be named

    An auditor can be told which physical machines held the data and what happened to them. Metered capacity cannot answer that question, because a scheduler moving a job between anonymous hosts is the entire point of the product.

    Data is returned or destroyed on termination

    Within thirty days, with written certification. Dedicated hardware with a known disk inventory is what makes that certification something we can actually sign.

    Breach notification and ongoing security program

    Notification inside fifteen business days, a written security program, an annual risk assessment, and cyber liability coverage held for the full term of the agreement.

    Worth saying plainly

    Nobody is HIPAA certified. There is no certificate, and the Office for Civil Rights certifies no one. Any vendor who tells you they are “HIPAA certified” is telling you something that cannot be true, which makes a useful early filter. What can be verified is whether a vendor will sign a Business Associate Agreement, and whether their facility will sign one too.

    A Business Associate Agreement requires every subcontractor who can touch the data to sign the same terms, which includes anyone with physical access to the cabinet. That obligation, plus an auditor who needs to be told which machine held the records, rules out capacity you cannot point at.

    The cost comparison

    Why per-GPU-hour is the wrong number

    The first numbers most buyers find are per-GPU-hour rates, and comparing on that line alone is misleading. A bare card rate is a card. This workload also needs 96 CPU cores, 768 gigabytes of RAM, 58 terabytes of storage and a network connection. Providers bundle those very differently, and the cheapest headline rates are usually the ones that bundle the least.

    The same environment, priced three ways on published rates:

    Cost elementPublic cloud APublic cloud BOwned hardware
    Twelve GPUs, CPU and RAM$37,800/mo$48,200/moFixed
    Storage, 58 TB$5,100/mo$4,200/moIncluded
    Egress, 25 TB/mo$2,200/mo$2,200/mo$0
    Egress, 100 TB/mo$8,000/mo$7,900/mo$0

    Priced August 2026 against published list rates for comparable configurations. Cloud pricing moves, so treat these as a snapshot rather than a quote.

    Two things fall out of that table. The bundled resources are about a third of the bill, so any comparison that skips them is off by a third. And egress is the only line that grows without a purchase order, which is the one that tends to make finance uncomfortable.

    Have a workload shaped like this one?

    A resident open-weights model on a card you don't share is a different cost and privacy story than a per-token API. Tell us what you're running and we'll size it.