In September 2026, two of the world's most profitable law firms sent the same message in two different dialects. Latham & Watkins began acquiring its own GPU clusters rather than routing every AI workload through a vendor's cloud. Kirkland & Ellis struck an infrastructure partnership with Palantir to build agentic workflows on top of its own data environment. Neither firm framed this as a cost play — Kirkland's average partner profit exceeds $7 million, and Latham isn't hunting for GPU bargains. They framed it as control.
That is the real story inside this month's agentic AI infrastructure roundup, and it's bigger than two firms flexing balance sheets. It's a signal that the AmLaw 200's most sophisticated buyers have concluded that per-seat legal SaaS, however capable, is an architectural compromise they're no longer willing to make for their most sensitive matters. The question for the other 180-plus firms in that ranking — and for the thousands of strong regional and mid-market firms below it — isn't whether to follow. It's how to get the same architectural guarantees without a nine-figure infrastructure budget.
The Infrastructure Arms Race Nobody Predicted a Year Ago
Eighteen months ago, the legal AI conversation was almost entirely about which vendor had the best model wrapper. Today it's about who owns the pipes. Latham's GPU acquisition and Kirkland's Palantir tie-up follow a pattern that's shown up quietly across the Am Law 50 throughout 2026: firms moving from consumption of AI products to ownership of AI infrastructure.
The economics explain why only the largest firms are doing this the hard way. Standing up even a modest on-premise inference environment — GPU racks, a vector database layer, orchestration tooling, security engineering, and the specialized staff to run it — runs into the low eight figures before a single associate touches it, and that's before ongoing model licensing and compute costs. Kirkland and Latham can absorb that against combined revenues north of $10 billion. A 90-lawyer regional powerhouse cannot, and shouldn't have to.
What these moves actually validate isn't the price tag — it's the underlying thesis: client data governed by privilege, confidentiality obligations, and regulatory exposure should not have to leave infrastructure the firm controls in order to be useful. That thesis doesn't scale down with headcount. A 40-attorney M&A boutique handling a $2 billion take-private has exactly the same sensitivity profile as a Kirkland deal team — it just has a fraction of the budget to solve for it.
Why the Vendor-API Model Hits a Ceiling on Sensitive Work
Here's the distinction that gets lost in most vendor pitch decks: the debate was never really about whether a firm uses a large language model from OpenAI, Anthropic, or Google. Nearly every serious legal AI platform, including the ones BigLaw is building internally, calls out to a frontier model somewhere in the stack. Latham's GPUs run inference workloads, but they don't replace foundation model capability — they control how and where data touches that capability.
The honest version of the architecture question isn't "who sends data outside the firm" versus "who never does." It's about what leaves, how much of it, and under whose terms.
In a typical shared-cloud legal AI product, a firm uploads or connects a document set, and that corpus — often the entire matter file — is ingested, chunked, embedded, and stored inside the vendor's multi-tenant cloud environment to power search and drafting. The firm's proprietary workflows, its connector logic to DMS and email, its permission structure, and its usage logs all live on vendor infrastructure too. That's a defensible model for a lot of use cases. It's a harder sell for a firm defending a client in a DOJ antitrust action, running diligence on a competitor's trade secrets, or advising a sovereign wealth fund that contractually requires data residency guarantees.
The alternative architecture — the one Kirkland and Latham are building at enormous scale, and the one available to much smaller firms through a private deployment model — flips what's exposed:
- The full document corpus stays on firm-controlled storage.
- The vector index and retrieval layer — the system that decides what's relevant to a query — runs inside the firm's environment.
- Connectors to DMS, email, and matter management stay internal.
- Permissions and ethical walls are enforced before retrieval happens, not after.
- Audit logs of every query and every document touched live on firm infrastructure, satisfying malpractice and regulatory review requirements.
- Only the minimal retrieved chunks — typically a handful of paragraphs, not the underlying documents — are sent to the selected LLM provider to generate the actual answer, under the firm's own API terms and zero-retention agreements.
That last point matters more than most procurement teams realize. A partner asking a question about a 400-page credit agreement doesn't need to send 400 pages to a model. A well-built retrieval layer sends the three or four clauses that are actually relevant. The corpus, the index, and the reasoning trail never leave the building. This is the same principle covered in more depth in our AI for law firms guide, and it's the design decision that separates infrastructure-grade legal AI from convenience tools.
What Actually Leaves the Firm — A Structural Comparison
| Layer | Shared-Cloud Legal SaaS | Kirkland/Latham-Scale Build | RAGbase Private Deployment |
|---|---|---|---|
| Full document corpus | Stored in vendor cloud | On firm infrastructure | On firm infrastructure |
| Vector index / retrieval | Vendor-hosted | Firm-hosted (custom build) | Firm-hosted (managed) |
| Connectors (DMS, email, iManage) | Vendor-managed | Firm-built | Pre-built, firm-controlled |
| Permissions / ethical walls | Enforced in vendor environment | Enforced internally | Enforced internally |
| Audit logs | Vendor-retained | Firm-retained | Firm-retained |
| What reaches the LLM | Often full context windows, sometimes full docs | Minimal retrieved chunks | Minimal retrieved chunks |
| Infrastructure capex | None (subscription) | $10M+ | Near-zero (managed private hosting) |
| Time to deploy | Weeks | 12-24+ months | Weeks to a few months |
The table makes the point plainer than any pitch could: Kirkland's build and RAGbase's private deployment model land in the same architectural column. The difference is capital intensity and timeline, not control.
The Kirkland Tier Problem
Here's the uncomfortable truth for firms outside the Am Law 20: the vendors selling per-seat AI tools were built for a market where every buyer looks roughly the same — a firm that will accept vendor-hosted data in exchange for speed and low switching cost. That's a reasonable trade for a lot of workloads. It stops being reasonable the moment a GC asks, in writing, where the firm's AI tools store client documents, and the honest answer is "a third party's servers, under their retention policy."
That question is no longer hypothetical. In 2026, in-house legal departments at financial institutions, pharma companies, and government contractors have started adding AI data handling addenda to outside counsel guidelines — not as boilerplate, but as negotiated terms with real teeth. Firms that can't answer precisely where their AI tools store data, index it, and route it are increasingly finding themselves excluded from RFPs before pricing is even discussed.
Kirkland and Latham solved this by building. A 60-attorney trusts-and-estates-meets-corporate firm, or a 150-lawyer regional firm with a strong financial services practice, can't build a Palantir-style stack — but it doesn't need to. It needs the same four guarantees at a fundamentally different price point:
- Documents never leave firm-controlled storage.
- Retrieval and reasoning happen inside firm infrastructure.
- Only minimized, purpose-specific content reaches the model layer.
- Every query is logged, walled, and auditable on systems the firm owns.
How RAGbase Closes the Gap
This is the exact design brief behind private AI deployment at RAGbase Legal. The platform is built to give a 20-to-200-attorney firm the same architectural posture Kirkland is spending nine figures to build, without the capital outlay, the specialized hiring, or the multi-year timeline.
In practice, that looks like:
- On-premise or firm-controlled cloud deployment, so the retrieval index, connectors, and logs sit inside infrastructure the firm's IT team already governs — not a vendor's multi-tenant environment.
- Model-agnostic architecture — firms can route queries to the frontier model of their choice (or several, by practice group), preserving negotiating leverage and avoiding single-vendor lock-in on the one component that genuinely does need external compute.
- Chunk-level minimization built into the retrieval layer by default, so a diligence question against a 12,000-document data room sends the model a handful of relevant passages, not the data room.
- Matter-level permissioning and ethical walls enforced before retrieval, mirroring how the firm's DMS already segments access — critical for lateral conflicts and multi-office practices.
- Full audit trails retained on firm infrastructure, giving GC and malpractice carriers a complete record of what the system retrieved, when, and for whom.
- Purpose-built case search and drafting workflows that sit on top of this architecture, so the control layer isn't a trade-off against usability — the firm gets both.
The economics tell the real story. Where a Kirkland-scale build runs into eight figures of capex plus dedicated infrastructure staff, a private deployment for a mid-market firm is typically structured as a predictable operating expense — comparable to, and in many cases lower than, the per-seat licensing costs of shared-cloud alternatives once a firm accounts for the audit and compliance overhead those tools generate downstream. That gap is explored in more detail in our analysis of hidden costs in legal AI SaaS — the sticker price of a per-seat license rarely includes what a firm spends convincing a client's security review team that vendor-hosted data handling meets their standards.
What This Means by Firm Profile
| Firm Profile | Infrastructure Reality | Recommended Posture |
|---|---|---|
| Am Law 20 (Kirkland/Latham tier) | Can absorb $10M+ builds, dedicated AI/infra teams | Build proprietary infrastructure, as seen in current moves |
| Am Law 50-200, strong regulated-industry practice | Needs architectural control, lacks build budget | Private deployment with model-agnostic retrieval layer |
| 20-200 attorney regional/boutique firms | Sensitive matters (M&A, litigation, government) but SaaS-scale budget | Private deployment as default for confidential workstreams |
| Smaller general-practice firms, low data sensitivity | Volume/speed matters more than sovereignty | Shared-cloud tools may be adequate |
The dividing line isn't firm size on its own — it's the sensitivity profile of the matters running through the system. A 150-lawyer firm doing high-stakes securities litigation has more in common, architecturally, with Kirkland than with a 150-lawyer firm doing high-volume insurance defense. Infrastructure decisions should follow risk, not headcount.
Where This Goes Next
Expect the roundup's core signal to keep repeating through 2026 and into 2027: more Am Law 50 firms announcing proprietary compute deals, more in-house legal departments formalizing AI data-handling requirements in outside counsel guidelines, and a growing gap between firms that can answer "where does our AI store client data" precisely and firms that can't. That gap will show up in RFP shortlists before it shows up in any survey.
The mistake would be concluding that architectural control is a luxury only the largest firms can afford. It's increasingly a baseline expectation from sophisticated clients — the balance sheet required to meet it is a separate question from the architecture itself, and for most firms in the 20-200 attorney range, that question has a much cheaper answer than Kirkland's did.
If your firm is fielding AI data-handling questions in RFPs or outside counsel guidelines and doesn't yet have a precise answer, that's the moment to evaluate a private deployment model before the next lateral partner or institutional client asks the question directly.
Frequently Asked Questions
Do mid-size law firms need to buy their own GPUs like Latham to control AI infrastructure?
What is the difference between shared-cloud legal AI and a private AI deployment?
Does using a third-party LLM provider mean client data leaves the firm's control?
Related Articles
Agentic AI for Law Firms: What It Actually Means in 2026
What agentic AI actually means for law firms — plain-English definition, what the big players are doing, real deployment examples, and how custom agents differ from SaaS workflows.
Your AI Vendor's Moat Is Your Data. Here's How to Take It Back.
How SaaS AI vendors build competitive moats from your firm's usage data — the shared learning paradox, the dilution problem, and why proprietary AI keeps the compounding advantage with you.
The Hidden Cost of Legal AI: Why 300-Lawyer Firms Are Spending $4.3M on Tools That Can't Find Their Own Case Files
Legal AI subscriptions cost up to $4.3M/year for large firms, yet can't search internal case files. Compare SaaS costs vs proprietary AI ownership economics.
RAGbase builds private AI systems for law firms: deployed on the firm's own infrastructure, zero data retention, full ownership.
See How RAGbase Works on Your Data
30-minute call. We scope your use case and show the system live.