data sovereignty

Agentic AI in Law Firms: Why Architecture Beats the Model

Agentic AI raises the stakes for law firms. Learn why the architecture behind autonomous legal AI agents matters more than model quality.

RAGbase Legal Research TeamSeptember 6, 2026 9 min read

A managing partner recently told us his firm had onboarded three AI tools in eighteen months without ever asking IT to map what each one could actually access. Not because no one cared — because nobody had framed it as an architecture question. It was framed as a feature question: does it draft well, does it cite accurately, does the associate like it. By the time anyone asked "what can this thing see," it had already been reading client files for months.

That's the pattern playing out across AmLaw 200 firms right now, and agentic AI is about to make it worse.

Agentic Doesn't Mean Smarter — It Means Autonomous

Strip away the vendor language and "agentic" describes one specific capability: the system can chain steps together — read a file, retrieve related documents, draft language, sometimes take an action — without waiting for a human to prompt each move. That's the entire definition. It says nothing about reasoning quality, hallucination rates, or legal accuracy.

This matters because procurement conversations routinely conflate the two. A tool marketed as "agentic" is not, by definition, using a better model than the chatbot it replaced. In many cases it's the same underlying LLM — GPT-4-class or Claude-class — wrapped in an orchestration layer that lets it execute multiple steps in sequence. The upgrade is in autonomy, not intelligence.

The distinction is not academic. It determines what kind of failure mode you're exposed to. A chatbot that gives a wrong answer produces bad text a lawyer has to catch. An agent that acts on a wrong inference — pulling the wrong client's file into context, drafting from a superseded precedent, or pushing a document into a workflow — produces a bad action that may already be downstream before anyone reviews it.

The Risk Reclassification: From Answering to Acting

Legal AI risk assessments built over the last two years were designed for a question-and-answer paradigm: does the model hallucinate citations, does it leak training data, does the output need a human check before it goes to a client. Those questions don't disappear with agentic tools — they get joined by a second category entirely.

An agent that opens client files, cross-references case history, and drafts a document autonomously introduces two new variables that traditional QA-style risk reviews weren't built to catch:

  • What the agent can see — which repositories, matters, and client files are in its retrieval scope, and whether that scope is scoped per-matter or firm-wide by default.
  • What the agent can touch — whether it can only draft into a sandbox for review, or whether it can send emails, update case management systems, or push documents to shared drives without a human in the loop.

Most vendor demos showcase the first capability (impressive retrieval and drafting) and gloss over the second (what happens when the agent is wrong, and how far the blast radius extends). A 2024 survey by the American Bar Association found that while adoption of generative AI tools among large firms exceeded 60%, fewer than a quarter of those firms had formal policies governing autonomous or multi-step AI workflows specifically, as opposed to single-turn chatbot use. The governance is lagging the capability by at least one full product cycle.

The conclusion many firms are reluctant to sit with: more autonomy requires more supervision, not less. The instinct — understandable given how good the drafting looks — is to trust the agent with more because it seems capable of more. The correct instinct is the opposite. Every additional step an agent can take without human review is an additional point where a permissions mistake, a stale index, or a bad retrieval turns into a real-world action rather than a paragraph you can quietly edit.

The Question That Actually Matters: Who Controls the Scaffolding

Here is the question worth displacing "how good is the model" in your evaluation process: who controls the scaffolding the agent runs on.

Every agentic system sits on an orchestration layer — the connectors that pull data from your DMS, your case management system, your email; the retrieval index that decides which documents are relevant to a query; the permission logic that determines who can trigger which action; and the audit trail that records what happened. In a vendor's shared-cloud product, that scaffolding lives inside the vendor's environment. Your files pass through someone else's orchestration layer to get to the model and back.

In a firm-controlled architecture, that scaffolding lives on your infrastructure. The retrieval index, the permissions, the audit logs stay put. Only the minimal context needed for a specific task — a handful of retrieved chunks, not the full document set — goes out to whichever LLM provider is doing the reasoning.

DimensionShared vendor-cloud architectureFirm-controlled architecture
Orchestration layerRuns inside vendor's environmentRuns on firm's infrastructure
Retrieval index / vector storeVendor-hosted, vendor-managedFirm-hosted, firm-managed
Full client documentsIngested into vendor's systemNever leave firm infrastructure
What reaches the LLMOften broader context windows, vendor-definedMinimized chunks, scoped per query
Permissions modelVendor's default, may not map to firm's ethical wallsFirm-defined, matches existing matter-level access controls
Audit logsHeld by vendor, access varies by contractHeld by firm, queryable on demand
LLM provider termsBundled into vendor's contractFirm's own choice and API terms

This is an architecture decision, not a feature comparison — and that's precisely why it gets made by accident. Whichever tool an innovation lead turns on first for a pilot often becomes the default scaffolding for every subsequent use case, because switching orchestration layers later is expensive and disruptive. Nobody sits down and deliberately decides "we will route all client documents through a third party's index." It happens because the first tool that worked well in a demo became the path of least resistance.

The Honest Version of the Data Question

It would be convenient — and dishonest — to frame this as "vendor tools send your data out, ours never does." That's not accurate, and firms considering a private or hybrid deployment should be skeptical of anyone who claims otherwise. Every legal AI system, including architectures built for data sovereignty, ultimately needs an LLM to do the reasoning, and most firms are not running frontier models entirely on their own hardware. The honest distinction is architectural, not binary.

In a properly designed private deployment, such as private AI deployment architectures built for regulated industries, the split looks like this: the full corpus of client documents, the retrieval and indexing layer, the connectors to your DMS and case management systems, the permissions model, and the audit logs all remain under the firm's control, typically within infrastructure the firm owns or a private cloud environment scoped to the firm alone. What leaves that boundary is narrow by design — the minimal retrieved chunks relevant to a specific query, sent to whichever LLM provider the firm has selected, under the firm's own API terms and data processing agreement.

That's a materially different exposure than a shared-cloud agentic product where the vendor's orchestration layer ingests entire document sets to build its index, and the vendor's contract — not the firm's — governs how that index is retained, trained on, or shared. The question isn't whether an LLM provider ever sees any data. It's how much data, in what form, under whose terms, and with what remains permanently outside anyone's control but the firm's.

A Practical Framework: Mapping What the Agent Can See and Touch

Before switching on any agentic tool — whether it's a general-purpose assistant with legal connectors or a purpose-built legal AI product — firms should be able to answer four questions in writing:

  1. Scope of visibility. Does the agent's retrieval layer respect existing ethical walls and matter-level permissions, or does it default to firm-wide access unless manually restricted?
  2. Action boundary. Can the agent only draft into a review queue, or can it independently send communications, modify records, or push documents into live systems?
  3. Data residency of the index. Where does the retrieval index and vector store actually live, and who has administrative access to it — the firm's IT team, or the vendor's engineering team?
  4. Audit granularity. If a partner asks "what did this agent access on Matter 4471 last Tuesday," can the firm produce that answer within the hour, or does it require a support ticket to the vendor?

Firms that can't answer all four haven't actually evaluated the tool — they've evaluated the demo. This is the same discipline that should apply to more targeted legal AI use cases like case search, where the value proposition depends entirely on the retrieval layer surfacing the right precedents from the right jurisdiction, scoped to the right matter, without silently expanding into documents outside the assignment.

Where the Market Actually Stands

The current generation of legal AI products spans a wide range of architectural choices, and it's worth being precise about where each sits rather than treating "agentic" as one undifferentiated category.

Anthropic's Claude Cowork and similar general-purpose agentic products are built for broad task automation across many industries, with legal use as one vertical among many — which means the permissioning defaults are generic, not built around ethical walls or matter-level access from the ground up. Purpose-built legal platforms like Harvey, CoCounsel, Lexis+ Protégé, and Legora have invested specifically in legal workflows and, in most cases, in enterprise agreements that address data handling more directly than a consumer-facing assistant like ChatGPT would. But nearly all of them still operate on a shared-cloud model where the orchestration layer, the index, and the retrieval logic sit inside the vendor's environment, not the firm's.

That's not a disqualifying fact — for many use cases, the convenience and speed of a managed vendor platform outweighs the architectural trade-off, particularly for lower-sensitivity workflows like general research or first-draft correspondence. But for sovereignty-critical workloads — privileged matter files, active litigation strategy, anything touching regulated client data — the calculus changes, and a firm-controlled architecture becomes less a preference and more a requirement driven by the firm's own ethical obligations under rules like Model Rule 1.6.

What This Means Going Forward

Agentic capability is going to keep expanding across every major legal AI product over the next 12 to 18 months, because the underlying orchestration frameworks that enable it — tool use, multi-step planning, autonomous retrieval — are becoming standard features of the models themselves, not proprietary advantages any single vendor holds for long. That means the differentiation between products will increasingly shift away from "can it act autonomously" — soon, everything will — and toward exactly the question this article opened with: who controls the scaffolding, and how granular is the firm's visibility into what the agent saw and did.

Firms that treat this as an architecture decision now, before their next pilot gets switched on by default, will have far more control over that answer than firms that discover the scaffolding question only after an incident forces it. Our broader AI for law firms guide walks through how to structure that evaluation across procurement, IT, and risk functions before any tool goes live.


The practical next step isn't picking a side in the agentic AI debate — it's running the visibility-and-action audit above against whatever tool your firm has already turned on. If the answers aren't clear, that's the signal the architecture decision was made by accident, and it's worth revisiting deliberately before the next tool gets added on top of it.

Frequently Asked Questions

What does 'agentic AI' actually mean for law firms?
Agentic AI refers to systems that chain multiple steps together autonomously — retrieving documents, cross-referencing case history, drafting language, and sometimes taking actions — without a human prompting each step. It is a workflow capability, not evidence of superior reasoning; the underlying model may be identical to the one powering a standard chatbot.
Is agentic AI riskier than chatbot-style legal AI tools?
Yes, by category, not just degree. A chatbot answers a question when asked; an agent can open client files, query multiple data sources, and generate or send documents with limited human review. That expanded scope of 'what it can see and touch' changes the risk profile, which is why firms need stricter permissioning and audit logging before deploying agents, not less.
What data actually leaves a firm's infrastructure when using an agentic AI tool?
In a well-architected deployment, only minimized retrieved chunks — the specific passages needed to answer a given query — are sent to the LLM provider for reasoning. Full documents, the retrieval index, permission structures, connectors, and audit logs remain on the firm's own infrastructure, whether that's on-premises or a private cloud environment the firm controls.

Related Articles

R
RAGbase Legal Research Team
Research

RAGbase builds private AI systems for law firms: deployed on the firm's own infrastructure, zero data retention, full ownership.

See How RAGbase Works on Your Data

30-minute call. We scope your use case and show the system live.

We use audience and marketing cookies (Google Analytics, LinkedIn). No tracker loads without your consent. Learn more