Menu
    GPU
    Infrastructure
    You're Paying Someone Else's Mortgage: The Real Cost of Cloud GPU Rentals

    You're Paying Someone Else's Mortgage: The Real Cost of Cloud GPU Rentals

    Bit Refinery TeamAugust 6, 20266 min read

    There's a moment every ML team hits — usually around month three or four of a serious training workload — where someone opens the AWS bill and goes quiet for a second. That silence is the sound of doing math.

    Cloud GPU rentals made a ton of sense at the start. You needed a few H100s for a couple weeks, you didn't want to buy hardware, and the per-hour pricing felt reasonable. But now you're running jobs continuously, your team has grown, and that "reasonable" hourly rate has compounded into something that looks a lot like a mortgage payment. Except it's not your mortgage.

    The Per-Hour Trap

    Comparison chart of Cloud GPU rental costs vs. Colocation and ownership costs

    Cloud providers price GPUs by the hour because it works in their favor. At low utilization — say, a research team spinning up jobs occasionally — it's genuinely a good deal. You're not paying for idle time you don't own.

    But the math flips hard once utilization climbs above 40-50%. At that point you're no longer benefiting from the flexibility premium. You're just paying a markup on compute that you could be dedicating to yourself.

    Let's look at actual numbers. An NVIDIA H100 on AWS (p4de.24xlarge or similar) runs roughly $32–$40/hour depending on region and reservation type. Run that 24/7 for a month and you're looking at:

    $35/hr × 24 hrs × 30 days = $25,200/month per GPU
    

    For a cluster of four H100s doing continuous fine-tuning or inference? That's over $100,000/month. And that's before egress fees, storage, or the networking costs that quietly accumulate in the background.

    What Dedicated Access Actually Costs

    This is where the "paying someone else's mortgage" analogy really lands. When you colocate your own GPUs — or get dedicated access to hardware you don't share with random tenants — the economics shift dramatically.

    At Bit Refinery, BYOGPU colocation starts at $600/month per GPU. You ship your hardware, we rack it, cable it, and have you connected via SSH, IPMI, and VPN within 48 hours. The GPU is yours. The utilization is yours. Nobody else is scheduling jobs on it at 2am.

    Four H100s in that model: $2,400/month in hosting costs. You still own the hardware outright after that. The cloud equivalent? You've spent $100,000 and own exactly nothing.

    Even if you factor in the upfront cost of purchasing H100s (roughly $25,000–$35,000 each at current market rates), the break-even on a four-GPU cluster versus cloud rental is somewhere around 3-4 months at high utilization. After that, every month of cloud rental is pure loss compared to ownership.

    "But We Need Flexibility"

    This is the counterargument that always comes up, and it's fair — to a point.

    If your GPU workloads are genuinely spiky and unpredictable, cloud makes sense. Burst training jobs, experimental research, one-off fine-tuning runs — yeah, rent those. The flexibility premium is worth it when you're actually using the flexibility.

    But most teams I talk to aren't in that situation. They've got a baseline of continuous work — model training pipelines, inference serving, embedding generation — that runs all the time. And then occasionally they need to burst. That's actually the perfect scenario for a hybrid approach: own or colocate your baseline compute, rent the spike.

    It's the same philosophy we apply to bare metal vs. cloud in general. The cloud isn't bad. It's just expensive when you treat it as your primary infrastructure instead of your overflow valve.

    The Hidden Costs Nobody Talks About

    Cloud GPU pricing is just the start. The stuff that really gets teams is:

    Egress fees. Moving training data in and out of cloud storage adds up fast. AWS charges $0.08–$0.09/GB for egress. If you're pulling a 10 TB dataset repeatedly across training runs, you're adding thousands per month just in data transfer. Bit Refinery charges $0 for egress. Unlimited bandwidth included.

    Storage costs. Keeping large model checkpoints, datasets, and artifacts in cloud object storage isn't free. S3 pricing at scale is manageable but it compounds with everything else.

    Idle time you still pay for. Even with spot instances, there's overhead from job queuing, preemption recovery, and the engineering time spent babysitting preemptible workloads. That engineering time has a real cost.

    Multi-tenancy performance variance. Cloud GPUs are shared infrastructure underneath the virtualization layer. Noisy neighbors are real. Your training job throughput can vary run to run in ways that are genuinely hard to debug.

    When Cloud GPUs Still Win

    I want to be honest here — cloud GPUs aren't always the wrong answer.

    • Short-term experiments: A few days of training on a new architecture? Just rent it.
    • Uncommon hardware: Need a specific GPU you can't justify buying? Cloud gives you access.
    • Zero upfront capital: Early-stage teams or startups without capex budget sometimes genuinely can't buy hardware.
    • Truly variable workloads: If utilization really is unpredictable month to month, flexibility has value.

    The mistake is treating cloud GPU rental as the default for everything, including steady-state workloads that have been running the same way for six months.

    The Inflection Point

    Here's a rough rule of thumb that holds up pretty well: if you're running GPU workloads more than 15 days per month, you should be doing the math on dedicated access or ownership. At 20+ days per month, it's almost certainly cheaper to own or colocate.

    The inflection point shifts a bit based on which GPU you're running, current hardware prices, and your specific hosting costs — but the direction is always the same. Sustained utilization kills the economics of cloud rental.

    If you're at that point and haven't run the numbers recently, it's worth doing. The gap between what you're spending on cloud GPUs and what dedicated access would cost is often shocking. Not in a subtle way — in a "we could hire two engineers with that delta" kind of way.

    Want to talk through the math for your specific workload? Reach out to the Bit Refinery team — we do this calculation with teams all the time and it's usually a pretty eye-opening conversation.

    Ready to Get Started?

    Contact us to learn more about our bare metal and GPU hosting solutions.