A managing partner at a 400-lawyer firm recently asked their vendor a simple question: "If a foreign government subpoenas your company, can they get to our data?" The vendor's answer — after a pause — was "technically, yes, but it's never happened." The firm had been sold a "private deployment." What they actually had was a logically isolated tenant on infrastructure owned by a company headquartered in a jurisdiction with broad compulsion authority. Isolation is not immunity.
This conversation is happening at AmLaw 200 firms across the country right now, and a September 2026 industry synthesis on agentic legal AI infrastructure gave it a name: the private/sovereign conflation. The report's core finding is blunt — most legal AI vendors marketing "private" or "secure" deployments are describing tenant isolation, not jurisdictional control. Those are different engineering problems with different risk profiles, and the industry has spent two years letting firms believe they're the same thing.
"Private" Describes a Wall. "Sovereign" Describes a Border.
In cloud architecture, private almost always means multi-tenant isolation: your firm's data is encrypted, access-controlled, and logically separated from other customers on the same underlying platform. It's a real security feature. It stops a competitor's associate from accidentally querying your privileged documents. It does not stop a government from compelling the vendor to produce data it lawfully possesses or controls.
Sovereign describes something structurally different: who has legal authority to compel access to the data, regardless of encryption or isolation. That authority attaches to jurisdiction — where the vendor is incorporated, where its parent company sits, and which laws govern its obligations to disclose. A private instance hosted in Frankfurt, run by a US-incorporated vendor, is still reachable under the US CLOUD Act (2018), which extends US law enforcement's authority to data controlled by US companies irrespective of physical server location. Firms that assumed EU data residency solved their compulsion-law exposure were solving the wrong problem.
This is not a hypothetical concern for law firms specifically. Attorney-client privilege, work product protections, and cross-border regulatory data (GDPR Article 48, various national blocking statutes) all depend on knowing exactly who can be forced to hand over what — and under whose authority. A vendor's SOC 2 report tells you about controls. It tells you nothing about compulsion exposure.
The Compulsion Gap: What Vendors Aren't Disclosing
| Dimension | "Private" (Isolation) | "Sovereign" (Jurisdictional Control) |
|---|---|---|
| What it protects against | Other tenants, insider access, accidental exposure | Foreign government subpoenas, national security letters, cross-border legal process |
| Where data physically sits | Anywhere the vendor's cloud provider operates | Determined by firm, often on-premise or firm-controlled infrastructure |
| Who can compel disclosure | Vendor's home jurisdiction, regardless of server location | Only the jurisdiction the firm has chosen to expose itself to |
| Typical proof point | SOC 2, ISO 27001, encryption-at-rest | Legal entity structure, data processing location, infrastructure ownership |
| Blind spot for firms | Assumes "private" = "safe from government access" | Requires understanding vendor corporate structure, not just architecture diagrams |
The gap matters most for cross-border matters, sovereign wealth clients, government contracts, and regulated industries where opposing counsel, regulators, or foreign courts have genuine incentive to seek compelled disclosure through the vendor rather than the firm. A firm can have perfect privilege logs and still lose control of the underlying data if the vendor holding it is compellable.
How Current Legal AI Tools Actually Deploy
Most of the widely adopted legal AI platforms — Harvey, CoCounsel (Thomson Reuters), Lexis+ Protege, Legora, and consumer-adjacent tools like Claude Cowork and ChatGPT Enterprise — run on shared cloud infrastructure (typically Azure, AWS, or Google Cloud) with tenant-level isolation and encryption. This is not a criticism of their engineering; it's the standard SaaS model, and it works well for the majority of matters. The issue is marketing language that implies more than the architecture delivers.
None of this is a knock on the workflow quality of these tools — several are genuinely strong for drafting, summarization, and research acceleration, and we've covered where each fits in our AI for law firms guide. The distinction only matters for a specific, identifiable subset of work: matters where jurisdictional exposure of the underlying data — not just its encryption — is itself a legal risk.
The honest framing isn't "they send your data out, we never do." Nearly every legal AI product, including ours, ultimately calls an LLM to generate a response — the model has to see something. The real question is how much reaches the model, and from where.
What Actually Leaves the Firm's Infrastructure
This is where architecture, not marketing copy, determines the answer. In RAGbase Legal's model, the firm's full document corpus, connectors (DMS, email, matter management), retrieval index, vector store, permissions layer, and audit logs remain on the firm's own infrastructure — on-premise, in a firm-controlled VPC, or in a jurisdiction the firm explicitly selects. The agentic scaffolding that decides what to retrieve and how to construct a response runs there too.
What leaves, per query, is narrow: the minimal retrieved chunks — typically a handful of paragraphs relevant to the specific question — sent to whichever LLM provider the firm has chosen, under the firm's own API terms and data processing agreement, not a bundled reseller agreement with the AI vendor.
| Layer | Where It Sits (RAGbase Model) | Where It Sits (Typical Shared-Cloud SaaS) |
|---|---|---|
| Full document corpus | Firm infrastructure | Vendor's multi-tenant cloud |
| Retrieval index / vector store | Firm infrastructure | Vendor's multi-tenant cloud |
| Connectors (DMS, email, billing) | Firm infrastructure | Vendor-managed, often with broad OAuth scopes |
| Permissions and access logs | Firm infrastructure | Vendor-managed |
| Agentic workflow logic | Firm infrastructure | Vendor's cloud, vendor's model routing |
| What reaches the LLM | Minimal retrieved chunks only | Often full documents or extended context windows |
| Governing terms for LLM call | Firm's own API agreement with chosen provider | Vendor's bundled agreement with its model provider |
That last row matters more than it looks. When a vendor bundles the LLM relationship into its own contract, the firm has no visibility into — and no leverage over — the actual data processing terms governing the moment data touches a foundation model. When the firm holds its own API relationship, it can select the provider, set retention to zero, choose the processing region, and audit the terms directly. This is the difference between trusting a vendor's privacy policy and architecting the exposure yourself. It's not privacy theater; it's a control boundary you can point to in a client engagement letter or an outside counsel guideline response.
This is also why private AI deployment has become its own procurement category rather than a checkbox — general counsel and CIOs are increasingly asking for the architecture diagram, not the marketing page.
Why This Surfaced Now
Three forces converged to make this a September 2026 flashpoint rather than a 2023 one:
- Agentic workflows widened the blast radius. A chatbot that answers one question at a time exposes one query. An agent that autonomously pulls from a DMS, drafts across multiple documents, and chains tool calls exposes an entire workflow's worth of context to whatever infrastructure runs it. The more autonomous the AI, the more that architecture — not just data-at-rest encryption — determines exposure.
- Cross-border matter volume kept climbing. Firms running sanctions, export control, and multinational M&A work can no longer treat "our vendor is SOC 2 compliant" as a sufficient answer to a client's outside counsel guidelines, which increasingly ask where data is processed and under whose legal authority.
- Vendors started competing on trust language before substance caught up. As more platforms entered the market — from established players to new entrants layering agents onto existing legal research products — "private," "secure," and "enterprise-grade" became interchangeable marketing terms rather than distinct architectural claims. The September 2026 infrastructure report is essentially the industry catching up to its own language.
The Questions Managing Partners Should Actually Ask
Most procurement checklists stop at security certifications. That's necessary but insufficient. The sharper diagnostic set:
- Where is the vendor incorporated, and what compulsion laws attach to that entity? Not where the servers sit — where the company answers to a subpoena.
- What exactly reaches the LLM per query — chunks or full documents? Ask for the retrieval architecture, not the privacy policy.
- Whose API terms govern the LLM call — the vendor's bundled agreement, or one the firm negotiates directly?
- Can the firm host the retrieval index and connectors on its own infrastructure, or is that layer inseparable from the vendor's cloud?
- What is the actual audit trail for a single agent action — can the firm reconstruct what was retrieved, from where, and sent to which model, six months later?
These questions map directly onto the same rigor firms already apply to case search and research tools, where sourcing transparency has long been table stakes — the same standard now needs to extend to the infrastructure layer, not just the citation layer.
Expect vendor language to keep converging on "sovereign" over the next 12 months as the term gains procurement weight — watch for firms asking for jurisdictional architecture diagrams alongside SOC 2 reports in RFPs by 2027. The distinction won't matter for every matter a firm runs; plenty of work is well served by fast, capable shared-cloud tools. But for cross-border, regulated, and sovereignty-sensitive work, the question worth asking your current or prospective AI vendor isn't whether they're private. It's whose laws can reach your data even if they are.
Frequently Asked Questions
Is a 'private deployment' of legal AI the same as data sovereignty?
Can the US government compel access to legal AI data hosted overseas?
Does RAGbase Legal ever send client data to third-party LLMs?
Related Articles
Your AI Vendor's Moat Is Your Data. Here's How to Take It Back.
How SaaS AI vendors build competitive moats from your firm's usage data — the shared learning paradox, the dilution problem, and why proprietary AI keeps the compounding advantage with you.
Agentic AI for Law Firms: What It Actually Means in 2026
What agentic AI actually means for law firms — plain-English definition, what the big players are doing, real deployment examples, and how custom agents differ from SaaS workflows.
RAGbase builds private AI systems for law firms: deployed on the firm's own infrastructure, zero data retention, full ownership.
See How RAGbase Works on Your Data
30-minute call. We scope your use case and show the system live.