data sovereignty

Gemini Enterprise for Legal: Why Shared-Cloud AI Still Isn't Sovereign

Google's Gemini Enterprise for Legal signs four AmLaw firms. We examine why shared-cloud legal AI still can't match on-prem for privilege and data sovereignty.

RAGbase Legal Research TeamAugust 29, 2026 10 min read

When Google Cloud announced Gemini Enterprise for Legal on August 25, 2026, it did so with a roster that would make any legal tech vendor jealous: Cleary Gottlieb, Freshfields, Weil, and Williams & Connolly signed on as launch partners. For an AmLaw 200 managing partner or CIO, that lineup is the point — if four of the most conservative, privilege-conscious institutions in the profession are willing to route their matter data through Google's infrastructure, doesn't that settle the security question?

Not quite. The launch is a useful case study in a distinction the legal industry has not yet fully reckoned with: the difference between a contractual promise of data isolation and architectural sovereignty. Google's marketing language is genuinely reassuring — right up until you read what it actually commits to, and what it necessarily does not.

The Announcement: What Google Actually Shipped

Gemini Enterprise for Legal bundles Google's Gemini models with legal-specific workflows — contract analysis, deposition prep, litigation research — inside Google Cloud's enterprise environment. According to Google's own blog post, the platform promises that client data, firm playbooks, and model outputs "stay private to your organization, and are never used to train or fine-tune Google's foundation models."

That is a strong, specific commitment, and it's more than many legal AI vendors have historically offered in writing. But read the sentence again: it describes a policy about training data, not a description of where the data physically resides or how many hands touch it in transit. "Stays private to your organization" is a configuration outcome enforced by Google's access-control layer, Google's encryption keys, Google's subprocessor agreements, and Google's incident-response team. The firm's confidential client data — including material subject to ethical walls, joint-defense privilege, or SEC-sensitive deal terms — still traverses and is stored on infrastructure the firm does not own, cannot audit in real time, and cannot unilaterally sever access to in the event of a dispute.

This is not a criticism unique to Google. It is the structural condition of every multi-tenant cloud AI product, from Harvey to CoCounsel to Lexis+ Protege to Anthropic's Claude Cowork. The question for AmLaw leadership isn't whether Google's engineers are trustworthy — it's whether "trust us" scales as an information governance model when the data in question is privileged.

The Fine Print: Isolation Is a Setting, Not a Wall

Google's "complete data isolation" claim rests on a set of configuration choices: tenant-separated storage, encryption at rest and in transit, and contractual non-training clauses. These are real engineering commitments, and Google has invested heavily in enterprise-grade controls (its Assured Workloads and confidential computing offerings are genuinely sophisticated). But three facts remain true regardless of configuration:

  • The data leaves the firm's network perimeter. Every prompt, every retrieved document chunk, every model output passes through Google's servers, logging pipelines, and monitoring systems before it ever reaches the attorney's screen.
  • Isolation depends on Google not changing its mind, its policies, or its subprocessors. Terms of service, data processing addenda, and default configurations have all changed at hyperscalers before, sometimes retroactively. A firm's confidentiality obligation to a client doesn't have an opt-out clause for vendor policy updates.
  • "Never used to train" is a promise about outputs, not about custody. The firm still has to take Google's word — and Google's audit logs, which Google controls — that the boundary held. In a malpractice inquiry or a bar disciplinary proceeding, "the vendor told us it was isolated" is a materially weaker position than "the data never left our infrastructure."

None of this means Gemini Enterprise for Legal is reckless. It means the security model is contractual and configurational, not architectural. Those are different categories of guarantee, and privilege law has historically demanded the latter.

The Partner Agent Ecosystem: More Vendors, More Touchpoints

Here is where the Gemini Enterprise for Legal launch gets more complicated than a single-vendor cloud story. Google is not shipping this alone. The launch explicitly relies on a partner agent ecosystem: a native integration with Legora for matter workflows, plus implementation and customization work from Accenture, Deloitte, and KPMG.

Each of those names is a reputable, competent firm. But each is also a distinct subprocessor with its own access model, its own staff, its own security posture, and its own contract with the law firm — layered on top of Google's own infrastructure. A single query in this ecosystem can plausibly touch:

  1. The firm's internal systems (DMS, email, matter management)
  2. Google Cloud's storage and compute layer
  3. Google's Gemini foundation model
  4. Legora's workflow and agent layer
  5. An implementation partner's customization layer (Accenture, Deloitte, or KPMG, depending on the engagement)

That is potentially five separate points of custody for data that may include privileged communications or ethical-wall-restricted deal terms — up from what used to be a single relationship between a firm and its DMS vendor. Every additional node is an additional entry on a privilege log, an additional breach-notification counterparty, an additional discovery target if opposing counsel challenges the chain of custody, and an additional vendor whose own security lapse becomes the firm's problem. The efficiency logic of a multi-vendor "partner agent" ecosystem is real — no single company does everything well — but the confidentiality logic runs in the opposite direction: complexity is not a feature of a sovereign data architecture, it's the enemy of one.

Shared-Cloud vs. On-Premise: A Structural Comparison

The practical question for a CIO evaluating Gemini Enterprise for Legal, or any comparable shared-cloud offering, is not "does the vendor mean well" but "what is physically and contractually possible in this architecture." The table below lays out the structural differences.

DimensionShared-Cloud Legal AI (Gemini Enterprise for Legal and peers)On-Premise / Private AI Deployment
Full document corpus locationResides on vendor cloud infrastructure (Google Cloud, AWS, Azure)Resides entirely within firm-controlled infrastructure
Retrieval index / vector storeHosted and managed by vendor or sub-vendorHosted on-premise or in the firm's private cloud tenant
What reaches the LLMFull document context often passed through vendor pipelineOnly minimal retrieved chunks needed to answer, under firm-selected API terms
Number of subprocessors touching privileged dataMultiple (cloud provider + workflow partner + implementation partner)Zero beyond the firm's chosen model API for the minimal chunk
Training-data guarantee basisContractual promise, vendor-configuredArchitectural — firm controls the boundary directly
Auditability of access logsVendor-controlled, reviewed on vendor's termsFirm-controlled, reviewed on firm's terms
Exposure in a subpoena or bar inquiryVendor's infrastructure, policies, and logs become discoverableFirm's own infrastructure remains the sole custodian
Dependency on vendor policy stabilityHigh — terms and subprocessors can changeLow — firm sets and controls its own retention and access policy

The honest reading of this table is not "cloud AI is unsafe." It's that shared-cloud and on-premise are solving for different risk tolerances, and firms handling sovereignty-critical matters — sovereign wealth litigation, national security-adjacent work, cross-border M&A with competing privilege regimes — need to know which one they're actually buying.

What Sovereignty Actually Requires

At RAGbase Legal, we build on the premise that true data sovereignty isn't a setting you enable in a vendor's admin console — it's a property of where the infrastructure physically sits. That means the firm's full document corpus, its retrieval and indexing layer, its permission model, its audit logs, and its agent orchestration all run inside infrastructure the firm itself controls, whether that's an on-premise data center or a private cloud tenant the firm fully owns and audits.

This is not an argument against using large language models from Google, Anthropic, or OpenAI. RAGbase Legal, like every serious legal AI platform, can call out to frontier LLM providers when a matter calls for it. The architectural difference is what leaves the building and what doesn't. In a properly designed private AI deployment, the full client file — pleadings, deposition transcripts, deal documents, ethical-wall-restricted correspondence — never leaves firm-controlled infrastructure. What can leave, under the firm's own chosen API terms, is a minimized retrieved chunk: the two or three paragraphs of case law or contract language actually needed to answer a specific question, stripped of surrounding context, sent to a model provider selected and configured by the firm's own IT governance, not bundled into a third-party's workflow and implementation-partner stack.

That distinction — full corpus and agent layer under firm control, versus minimized chunks sent under firm-chosen terms — is the difference that actually matters for privilege analysis. It's also why the comparison to Gemini Enterprise for Legal isn't "they send data out, we never do." Both models call out to models. The difference is how much data leaves, in what form, through how many intermediary vendors, and under whose contractual authority.

Where This Plays Out in Practice

Consider a matter involving an ethical wall between two practice groups at the same firm — exactly the scenario Google's own materials cite as a use case for "complete isolation." In a shared-cloud deployment, that wall is enforced by the vendor's access-control configuration, which the firm must trust was implemented correctly and hasn't drifted. In an on-premise deployment, the wall is enforced by the firm's own identity and permissions system, the same system that already governs DMS access — auditable by the firm's own general counsel, on the firm's own schedule, without waiting on a vendor's transparency report.

The same logic applies to research workflows. When associates run a case search across a firm's internal precedent bank and public case law, the query and the underlying precedent documents should never require routing the full case file through a third-party's cloud pipeline — only the retrieved passages relevant to the query need to reach a model, and even those can be selected and minimized entirely within the firm's own retrieval layer before any external call is made.

The Market Signal Behind the Launch

The more interesting story in the Gemini Enterprise for Legal launch isn't Google's technology — it's what the four launch firms' decision reveals about market segmentation. Cleary Gottlieb, Freshfields, Weil, and Williams & Connolly are all large enough to negotiate custom data processing addenda, dedicate in-house security staff to audit vendor configurations continuously, and absorb the residual risk of a multi-vendor stack. That is a luxury most of the AmLaw 200 — and nearly all mid-market and boutique firms — do not have the bandwidth to replicate.

For those firms, the more relevant question isn't "which hyperscaler has the best legal AI product" but "what is the minimum viable architecture that satisfies our bar's confidentiality rules without requiring a full-time vendor-audit function." Our AI for law firms guide walks through that evaluation in more depth, including how to map a matter's sensitivity tier to an appropriate deployment model rather than defaulting to whichever platform has the most launch-partner press coverage.


The Gemini Enterprise for Legal launch will accelerate adoption of shared-cloud legal AI across the AmLaw 200, and for a large share of day-to-day matters — routine due diligence, first-pass contract review, general legal research — the risk profile of a well-configured shared-cloud tool is entirely defensible. The mistake would be treating every matter as equally low-stakes. Before extending any shared-cloud AI platform to privileged, ethical-wall-restricted, or sovereignty-sensitive work, it's worth asking three concrete questions: how many distinct vendors and subprocessors will touch this data before it reaches an attorney's screen, what happens to access and audit logs if that vendor relationship ends, and who actually controls the infrastructure the data sits on while the model runs. The answers should determine which matters go to a shared-cloud platform — and which belong on infrastructure your firm controls outright.

Frequently Asked Questions

Does Google use law firm data to train Gemini models?
Google states client data, firm playbooks, and outputs in Gemini Enterprise for Legal are never used to train or fine-tune its foundation models. That is a contractual and configuration commitment, not a physical impossibility — the data still transits and resides on Google's multi-tenant cloud infrastructure, governed by Google's access controls, subprocessors, and incident response, not the firm's.
What is the difference between shared-cloud legal AI and on-premise legal AI?
Shared-cloud tools like Gemini Enterprise for Legal, Harvey, or CoCounsel run the retrieval layer, indexes, logs, and often the model itself on vendor-controlled infrastructure, with data isolation enforced by contract and configuration. On-premise deployments keep the full document corpus, retrieval index, permissions, and agent orchestration inside the firm's own environment, sending only minimal retrieved text chunks to an LLM provider under terms the firm selects.
Why does the multi-vendor partner ecosystem around Gemini Enterprise for Legal matter for privilege?
Gemini Enterprise for Legal launches with integrations from Legora and implementation support from Accenture, Deloitte, and KPMG, meaning privileged and ethical-wall-restricted data can pass through several third-party systems rather than one. Each additional vendor is an additional subprocessor, contract, and potential discovery target, which complicates privilege logs and breach notification obligations.

Related Articles

R
RAGbase Legal Research Team
Research

RAGbase builds private AI systems for law firms: deployed on the firm's own infrastructure, zero data retention, full ownership.

See How RAGbase Works on Your Data

30-minute call. We scope your use case and show the system live.

We use audience and marketing cookies (Google Analytics, LinkedIn). No tracker loads without your consent. Learn more