A Law Firm Bought Its Own AI Servers. Here Is How to Get the Same Setup Without the Budget

    A Law Firm Bought Its Own AI Servers. Here Is How to Get the Same Setup Without the Budget

    Bit Refinery TeamSeptember 14, 20268 min read

    On September 11, the Financial Times and Bloomberg Law both reported that Latham & Watkins, the second largest law firm in the United States by revenue, has bought its own Nvidia GPU servers and is fine tuning open weight AI models on them. It appears to be the first large law firm to do this publicly.

    It would be easy to file this under legal tech news. We think it is more useful to read it as a preview of a decision many data heavy companies will face over the next few years: keep renting AI from a lab, or own the parts of the stack that touch their most sensitive work.

    This post covers:

    • What Latham built, based on what the firm has said publicly
    • Why it made the move, in its own words
    • The questions other industries are asking for the same reasons
    • The three layers of a private AI stack
    • How to get a similar setup without building your own data center
    • Which teams tend to benefit first

    1. What Latham actually built

    According to Bloomberg Law, the firm decided three years ago to build its own Nvidia server infrastructure and has been buying servers over that period. The details it has shared so far:

    • Hardware: multiple Nvidia H200 GPUs, and the firm is actively looking at Nvidia's newer Blackwell and Vera Rubin systems.
    • Location: the servers run in a third party data center, in space that only Latham employees can access.
    • Models: open weight models developed by Nvidia, which the firm's engineers are fine tuning for its own work.
    • People: a technology staff of about 900, with roughly 100 of them dedicated to AI.

    Latham has not stopped using commercial AI. It still uses tools from Harvey, Legora, OpenAI, and Anthropic. The private servers sit alongside those tools rather than replacing them, which matches what we see at most companies that take this step (hosted tools handle the general work, and the private stack handles the work that should not leave the building).

    2. Why they did it

    The firm gave two main reasons, and both apply well beyond law.

    The first is confidentiality. Chief information officer Rene Mendoza told the Financial Times, as reported by Legal IT Insider, that some client information is sensitive enough that the firm does not want it with any cloud vendor. That is a stricter standard than trusting a vendor's data processing agreement. It means certain data never reaches a vendor at all.

    The second is flexibility. Mendoza told Bloomberg Law that "nobody can predict where any of this is going, so we have the best optionality." Owning the hardware lets the firm change models, adopt new open weight releases as they appear, and tune them on its own material without waiting on anyone. Cost plays a part as well. Latham said saving on token costs "is a factor," though not the one driving its decisions.

    Michael Rubin, who chairs the firm's AI practice, described the result as a capability that cloud based legal AI tools do not offer, because the firm can run models on premises.

    The confidentiality concern will sound familiar if you caught the All-In Podcast episode that aired the same day. The hosts spent a long segment on proprietary data leaking into frontier models, and we covered it in what the All-In Podcast got right about AI data leakage.

    3. The questions other industries are asking

    Once AI moves from a pilot to production, teams in finance, healthcare, government, manufacturing, and ad tech tend to arrive at the same four questions:

    1. Who owns the weights after we fine tune a model?
    2. Where do our prompts, documents, and embeddings physically live?
    3. Can we switch models when quality or price changes?
    4. What happens to our data if a vendor's terms, a regulator, or a training policy changes?

    If the honest answers depend on a vendor's roadmap or terms of service, the real question is less about AI strategy and more about who controls the infrastructure. That is the conclusion Latham reached, and it is the one we hear most often from regulated customers.

    4. The three layers of a private AI stack

    A private AI stack has three layers. Latham built all three itself. Most companies can reach the same privacy posture by controlling the model and the data, then choosing how much of the compute layer to own.

    LayerWhat Latham didHow most companies can do it
    ModelFine tunes open weight Nvidia modelsFine tune open weight models such as Llama, Qwen, Mistral, or Nemotron
    DataKeeps client material on its own serversKeep your data on dedicated, encrypted disks that you can wipe
    ComputeOwns H200 servers in a locked third party data centerRent dedicated GPUs, or place your own GPUs in a colocation facility

    Three layers of a private AI stack, comparing what Latham & Watkins did for the model, data, and compute layers with how most companies can do the same

    5. Three ways to own the compute

    Latham's approach suits a firm with 900 technologists and a buildout that has run for three years. Most companies that need the same privacy have neither, and they do not need them. There are three realistic ways to get dedicated compute:

    Your own racks in a coloShip your GPUs to a hostRent dedicated GPUs
    Who owns the hardwareYou (servers and racks)You (the GPUs)Your provider
    Upfront capitalHigh (servers, racks, networking)The GPUs themselvesNone
    Time to go liveWeeks to monthsA few days after the hardware arrivesHours to days
    Staff you needNetwork and hardware engineersMostly the team using the GPUsMostly the team using the GPUs
    Single tenantYesYesYes, if the provider dedicates the hardware
    Best fitLarge programs with their own hardware standardsTeams that have already bought GPUsTeams that want to start now or avoid capital spend

    Whichever route you pick, the properties worth checking are the same: the GPUs serve one customer, the data stays on disks that customer controls, and nothing on the other side of the connection learns from the work. We go deeper on how to evaluate a provider in what to look for in private AI infrastructure.

    6. Who tends to benefit first

    • Legal and professional services. Privilege, work product, and client identity are hard to protect once they pass through someone else's systems. Latham has now shown that a large firm can run its own stack.
    • Finance, healthcare, and the public sector. Data residency rules, business associate agreements (BAAs), and audit trails are easier to defend when a GPU serves a single customer. Our post on data sovereignty for regulated teams covers this in more detail.
    • Ad tech, SaaS, and data platforms. High token volume and heavy data transfer are where metered API and cloud bills tend to grow fastest.
    • Teams that already own GPUs. Cards sitting in an office closet or a lab without proper cooling can do far more work in a data center.

    Most of these teams will keep using hosted models for some work, as Latham does. The split we usually recommend is to put steady, always on workloads (fine tuning on your own data, serving a model to your staff) on dedicated hardware, and to use public cloud or APIs for short bursts. We cover the economics in own the base, rent the spike, and how open weight releases such as Kimi K3 fit into that plan.

    Where Bit Refinery fits

    Bit Refinery is a GPU hosting company, not an AI lab. We have run infrastructure since 2008, starting in a former NASA facility in Morrison, Colorado, and today operate Tier 3 data centers in Denver and Seattle. There is no model on our side of the rack, and we do not log your prompts or train on your data.

    We offer all three of the routes above:

    • Your own racks. For teams that want to own the whole cluster the way Latham does, we now offer GPU rack colocation in the Denver area. You bring your own GPU servers and racks, and we provide the data center space, power, cooling, and network connectivity. Talk to our team about your space and power requirements.
    • Your own GPUs. BYOGPU lets you ship us the GPUs you already bought, and we rack, power, cool, and connect them.
    • Our GPUs, dedicated to you. Private GPU Cloud gives you dedicated NVIDIA GPUs that we own, with root access, encrypted storage, and no egress fees.

    GPU pricing is moving quickly, so the BYOGPU and Private GPU Cloud pages list current rates. If you would rather have us run the models for you as well, see Private LLM Hosting.

    Ready to Get Started?

    Contact us to learn more about our bare metal and GPU hosting solutions.