A managing partner doesn't ask a vendor "do you promise not to train on our data?" anymore. That question has been answered — contractually, repeatedly, in every legal AI sales deck since 2024. The question now is sharper: "Show me exactly what left the building, when, and to whom." Most legal AI vendors cannot answer it. That inability is becoming the actual liability.
The agentic_ai_infra State of Play report (September 2026) names this directly: architecture, not model quality or pricing, is where legal AI vendors will differentiate — and where they'll be exposed — over the next 18 months. The report frames the industry as stuck in a false choice between two unsatisfying options. It's worth taking that framing seriously, because it's mostly right, and because the gap it identifies is exactly where the next wave of malpractice inquiries, bar audits, and client security questionnaires will land.
The False Binary the Market Built
For two years, legal AI procurement has been framed as a straight line with two endpoints. On one end: fully on-premise deployments — expensive, slow to extend, often years behind the current model generation because upgrading means a procurement cycle, not an API call. On the other: cloud API tools — fast to deploy, continuously updated, backed by data-processing addenda and "we don't train on your data" language that legal ops teams have gotten very good at negotiating.
The problem is that both endpoints solve for the wrong variable. Fully on-prem solves for control and sacrifices velocity — firms end up running a six-month-old model because the infrastructure team can't keep pace with frontier releases. Cloud API solves for velocity and sacrifices verifiability — the firm gets contractual promises about training data, retention windows, and deletion, but has no independent way to confirm what was retrieved, indexed, or transmitted during any given matter. Ninety-eight percent of large-firm attorneys report using generative AI tools in some form, according to industry adoption surveys — but a much smaller fraction of firms can produce an audit trail showing precisely what client data touched which third-party system on which date. That gap between usage and verifiability is the actual risk surface, and it's why the adoption gap between "tool is in use" and "tool is defensible" keeps widening even as adoption numbers climb.
This is also why per-seat SaaS licensing models have started drawing scrutiny beyond price. The cost conversation (often $2,000–$4,500 per attorney per year, layered across multiple point tools) is real, but the more consequential conversation is what a firm actually knows about where its documents went once a seat license connects to a matter management system, a DMS, and a third-party model API simultaneously.
Data Residency vs. Data Sovereignty: The Distinction That Actually Matters
The State of Play report draws a line that most vendor marketing deliberately blurs: residency is not sovereignty.
- Data residency answers: where does the data physically sit? US data center, EU region, specific cloud availability zone. This is the question most enterprise contracts optimize for, because it's the easiest to certify and the easiest for a vendor to promise.
- Data sovereignty answers: who controls access, retrieval, indexing, and retention of that data, regardless of where it sits? This is the harder question, because it requires the firm — not the vendor — to hold the keys to the retrieval layer, the permission model, and the logs.
A firm can satisfy residency requirements completely — data stored in-region, contractually guaranteed — while still having zero sovereignty, because the vendor's platform controls what gets indexed, what a query retrieves, how long it's cached, and what the model provider's API logs. Residency is a location claim. Sovereignty is an architecture claim. Auditors, regulators, and — increasingly — malpractice carriers are starting to ask the sovereignty question, not the residency question, because residency alone doesn't tell you what actually happened during a specific matter.
This matters concretely in privilege review. If opposing counsel or a regulator challenges how an AI tool was used on a privileged document set, "our data was stored in a US region under a signed DPA" is not a defense. "Here is the retrieval log showing which fifteen chunks of a 400-page document were sent to the model, timestamped, tied to the reviewing attorney's credentials, with the full document never leaving our environment" is a defense. Only one of those statements requires architectural control. The other is a contractual promise the firm has to trust.
What an Audit or Malpractice Inquiry Actually Tests
When a general counsel's office, a state bar investigation, or opposing counsel in a malpractice claim asks a firm to reconstruct AI usage on a matter, they're not asking for the vendor's terms of service. They're asking for artifacts:
- What was indexed — which documents, which version, which permission scope.
- What was retrieved — for a specific query, exactly which passages the system pulled and surfaced to the model.
- What was transmitted externally — the precise payload sent to any third-party API, not an estimate.
- Who accessed what, when — a permission and access log tied to individual credentials, not a shared service account.
- What was retained afterward — by the firm, and separately, by any third party in the chain.
A pure cloud API deployment can usually answer questions 4 and 5 in general terms (the vendor's admin console shows user activity; the DPA states a retention window). It almost never can answer 1, 2, and 3 with precision, because the firm doesn't operate the retrieval layer — the vendor does. The firm is relying on the vendor's internal logging, which it cannot independently inspect, subpoena from itself, or produce on short notice in a form that satisfies an auditor.
This is the specific failure mode the report is circling without quite naming it: contractual defensibility and architectural defensibility are different things, and only one of them survives cross-examination.
Anatomy of the Three Architectures
| Dimension | Cloud API (SaaS) | Fully On-Premise | Hybrid (Firm-Controlled Scaffolding) |
|---|---|---|---|
| Where documents live | Vendor cloud environment | Firm data center | Firm infrastructure |
| Who controls the index/retrieval layer | Vendor | Firm | Firm |
| What leaves the firm's environment | Full documents, queries, often persistent context | Nothing | Minimal retrieved chunks only, per query |
| Model currency | Continuously updated (vendor's roadmap) | Often 6–18 months behind frontier | Current — firm chooses/swaps model provider |
| Audit log ownership | Vendor-controlled console | Firm-controlled | Firm-controlled |
| Time to extend/customize | Fast, limited to vendor's roadmap | Slow — quarters, not weeks | Fast — firm controls connectors and workflows |
| Typical cost driver | Per-seat licensing, scales with headcount | Capex + specialized infra staff | Infrastructure + usage-based model spend |
| Defensible in an audit? | Depends entirely on vendor cooperation and logging | Yes, but often at the cost of capability | Yes — logs, indexes, permissions are firm-native |
The table makes the tradeoff explicit: on-prem wins on control and loses on velocity; cloud API wins on velocity and loses on verifiability. The hybrid column isn't a compromise between the two — it's a different design principle. It keeps the parts of the system that generate risk (the corpus, the index, the permission model, the logs) under the firm's control, while letting the parts that benefit from vendor R&D (the underlying LLM) remain swappable and current.
The Hybrid Model, Done Right: What Stays, What Moves
This is the architecture RAGbase Legal runs, and it's worth being precise about what it actually means, because "hybrid" has become a marketing word applied loosely.
What stays on the firm's infrastructure, permanently:
- The full document corpus — contracts, discovery sets, matter files, precedent libraries
- The vector indexes and embeddings built from that corpus
- The permission and access-control layer, mapped to the firm's existing ethical walls and matter-level restrictions
- The agentic scaffolding — the workflows, connectors to the DMS and practice management system, and the orchestration logic that decides what to retrieve and why
- The complete audit log — every query, every retrieval, every user, timestamped and exportable
What can leave, under terms the firm sets:
- Only the minimal retrieved text chunks needed to answer a specific query, sent to whichever LLM provider the firm has selected — OpenAI, Anthropic, or another provider under the firm's own API agreement
- No persistent access to the underlying documents
- No standing connection between the model provider and the firm's index
The distinction is between full corpus plus agent layer under client control and minimized chunks sent to a model on a per-query basis. A shared-cloud legal AI platform inverts this: the vendor holds the corpus, the index, and the logs, and the firm holds a contract. RAGbase's private AI deployment model holds the corpus, index, permissions, and logs on the firm's own infrastructure, and treats the model call as the smallest, most disposable part of the pipeline — which is also the correct threat model, since the model provider is the one component in the stack the firm didn't build and shouldn't have to fully trust.
This is not an argument against using frontier models. A firm running RAGbase can point its retrieval layer at GPT-5-class models, Claude, or an open-weight model hosted entirely in-house, and switch between them without re-architecting anything, because the switching cost lives at the API call, not at the corpus. That flexibility is what on-prem-only deployments typically can't offer — they tend to lock into whatever model was validated at deployment time and stay there. A firm's case search workflows, for instance, can run against a continuously refreshed index without ever exposing the underlying case files to the model provider beyond the specific passages a given query actually needs.
Where the Current Competitive Set Sits
Most of the tools firms are evaluating right now — Harvey, CoCounsel, Lexis+ Protégé, Legora, and general-purpose assistants like Claude or ChatGPT accessed directly — are cloud-hosted, single-tenant-or-multi-tenant SaaS products. That's not a criticism; it's an accurate description of the category, and it's the right choice for firms whose priority is fastest time-to-capability and who are comfortable relying on vendor-side controls and contractual assurances. Claude Cowork's entry into the agentic workspace category has intensified this exact conversation, because agentic tools that take multi-step actions across a firm's systems raise the stakes on knowing precisely what data moved and when — a question worth asking before any pilot, not after.
The honest comparison isn't "they send data out, RAGbase doesn't." RAGbase also calls out to LLM providers — every serious legal AI product does, because no firm is training a frontier model from scratch. The real comparison is where the corpus, the index, and the logs live, and who can produce them on demand. For firms handling sovereignty-critical workloads — cross-border matters, regulated-industry clients, government contracts with data-handling clauses, or litigation where the AI usage itself could become discoverable — that distinction stops being theoretical.
What This Means Going Into 2027
The State of Play report's core prediction is that architecture becomes a procurement checkbox by 2027 the way SOC 2 became one by 2022 — not optional, not a differentiator, just table stakes that firms will ask about before the demo even starts. That prediction looks conservative. General counsel offices are already writing sovereignty and auditability requirements into outside counsel guidelines; it's a short step from there to requiring the same standard for the firm's own internal AI stack.
Firms evaluating their next legal AI investment should treat the architecture question as a due-diligence item, not a technical footnote: ask any vendor to produce, live, the retrieval log for a sample query — not a description of their logging policy, the actual artifact. Ask whether the corpus and index can be exported and inspected independently of the vendor's console. Ask what happens to the retrieval layer if the contract ends. Those three questions separate contractual defensibility from architectural defensibility faster than any feature comparison. For a broader look at how these decisions play into total cost and adoption planning, the AI for law firms guide walks through the procurement questions worth raising before signing anything.
The firms that come out ahead in the next audit cycle won't be the ones with the most AI tools deployed — they'll be the ones who can produce, on request, exactly what an AI system did with a client's data and prove it never had to leave the building to do it. That's an architecture decision, not a vendor promise, and it's worth making deliberately rather than inheriting it from whichever tool onboarded fastest.
Frequently Asked Questions
What is the difference between data residency and data sovereignty in legal AI?
Is on-premise legal AI more secure than cloud-based legal AI?
Can law firms use cloud LLMs like GPT or Claude while keeping data private?
Related Articles
Your AI Vendor's Moat Is Your Data. Here's How to Take It Back.
How SaaS AI vendors build competitive moats from your firm's usage data — the shared learning paradox, the dilution problem, and why proprietary AI keeps the compounding advantage with you.
98% of AmLaw 200 Firms Use AI — But Most Still Can't Search Their Own Files
98% AI adoption, but most law firms still can't search their own institutional knowledge. The gap between external AI tools and internal document access — and how to close it.
Agentic AI for Law Firms: What It Actually Means in 2026
What agentic AI actually means for law firms — plain-English definition, what the big players are doing, real deployment examples, and how custom agents differ from SaaS workflows.
The Hidden Cost of Legal AI: Why 300-Lawyer Firms Are Spending $4.3M on Tools That Can't Find Their Own Case Files
Legal AI subscriptions cost up to $4.3M/year for large firms, yet can't search internal case files. Compare SaaS costs vs proprietary AI ownership economics.
RAGbase builds private AI systems for law firms: deployed on the firm's own infrastructure, zero data retention, full ownership.
See How RAGbase Works on Your Data
30-minute call. We scope your use case and show the system live.