A managing partner recently described watching an AI agent open a client file, pull three related matters from the case management system, and draft a motion — all before anyone on the team had clicked anything beyond "start." His reaction wasn't excitement. It was: wait, what else did it just look at?
That question is the entire story of agentic AI in law firms right now. Not whether the models are good enough — they largely are. The question is whether anyone mapped what the agent could see and touch before someone turned it on.
Agentic Doesn't Mean Smarter — It Means Autonomous
Strip away the vendor language and "agentic" describes one specific capability: a system that chains steps without waiting for a human to authorize each one. Read a file. Retrieve related documents. Draft language. In some configurations, take an action — send, file, update a record — without a prompt in between.
That's it. It's a workflow property, not an intelligence upgrade. The underlying model reasoning over a retrieved case file is functionally the same model whether it answers one question in a chat window or executes a five-step sequence autonomously. What changes is exposure — how much of your document environment the system touches, and how many decisions it makes before a human reviews the output.
This distinction matters because most vendor demos conflate the two. A polished agent workflow looks smarter because it does more, faster, with less visible friction. But the intelligence ceiling hasn't moved. What's moved is the blast radius of a mistake.
Consider the comparison:
| Capability | Chatbot (single-turn) | Agentic system (multi-step) |
|---|---|---|
| Human checkpoints | One prompt, one answer | Often zero between steps |
| Data touched per session | Whatever's pasted into the prompt | Files, indexes, case history, connected systems |
| Error surface | One wrong answer, reviewed immediately | Chained errors compounding across steps |
| Audit requirement | Log the exchange | Log every retrieval, every action, every source |
| Risk category | Output quality | Output quality and access control |
The last row is the one most firms haven't fully priced in. A chatbot that gives a bad answer is a quality problem. An agent that retrieves the wrong client's privileged file while cross-referencing "similar matters" is a different risk category entirely — one that touches confidentiality obligations, conflicts screening, and potentially malpractice exposure.
The Autonomy-Supervision Paradox
There's a natural assumption that better models need less oversight. Newer reasoning capability, longer context windows, more accurate retrieval — surely that reduces risk over time.
It doesn't. It shifts the risk.
As agents gain more autonomy, two variables expand simultaneously:
- What the agent can see — the scope of documents, matters, and systems it's permitted to query in a single task.
- What the agent can touch — whether it merely retrieves and drafts, or whether it executes: sending communications, updating records, filing documents.
A firm that grants an agent broad read access across its document management system to "improve retrieval quality" has just expanded variable one for every user of that agent, regardless of whether any individual task needs that scope. A firm that lets an agent take actions without a review step has expanded variable two for every workflow built on top of it.
This is why the right response to more capable models is tighter governance, not looser governance. The industry pattern with every prior wave of automation — RPA, workflow bots, even email rules — was the same: capability outpaced access control until an incident forced a retrofit. Agentic legal AI is moving faster than those prior waves, and the stakes (privilege, confidentiality, professional responsibility rules) are considerably higher than a misfiled invoice.
The firms getting this right are treating agent permissioning the way they'd treat lateral hire access provisioning: least-privilege by default, scoped to the specific matter, expanded only with a documented reason. Our AI for law firms guide covers this in more operational detail, but the principle is simple — autonomy without a permissions model isn't a productivity gain, it's deferred incident response.
Where the Risk Actually Lives: The Scaffolding, Not the Model
Here's the part most competitive comparisons skip. The meaningful risk in agentic legal AI isn't which foundation model is doing the reasoning. GPT-4-class and Claude-class models are, for most legal drafting and research tasks, roughly comparable in output quality. The real variable is what surrounds the model — the scaffolding.
Scaffolding is the orchestration layer that makes an agent "agentic" at all: the connectors to your document management system, the retrieval index that decides what's relevant, the permissions layer that decides who can see what, the audit log that records what happened, and the workflow logic that chains steps together.
When a firm adopts a tool like Claude Cowork, Harvey, CoCounsel, Lexis+ Protege, or Legora, it's not just adopting a model — it's adopting that vendor's scaffolding. In most current deployments, that means the retrieval layer, the permissions logic, and often full document content pass through the vendor's cloud environment before the model ever sees a prompt. The vendor's infrastructure is doing the deciding: what counts as "related," what gets pulled into context, what gets logged and how.
That's not a criticism of any single vendor's security practices — most have invested heavily in encryption, access controls, and compliance certifications. It's an architectural fact: the orchestration decisions happen somewhere, and in a fully vendor-hosted agent, they happen on infrastructure the firm doesn't control and, in an audit, can't fully inspect.
Two Architectures, One Decision
The honest framing here isn't "vendor clouds send your data out, ours never does." Any serious legal AI deployment — including firm-controlled ones — ultimately sends something to an LLM provider for reasoning. The question that actually matters is how much, and who controls the layer that decides what gets sent.
| Vendor-hosted agentic AI | Firm-controlled scaffolding | |
|---|---|---|
| Where full documents live | Vendor's cloud environment | Firm's own infrastructure |
| Where the retrieval index sits | Vendor-managed | Firm-managed |
| Where permissions logic runs | Vendor-managed | Firm-managed |
| Where audit logs are stored | Vendor's systems (exportable, but vendor-controlled by default) | Firm's systems, natively |
| What reaches the LLM provider | Often full context, vendor's discretion | Only the minimal retrieved chunks needed for the specific task |
| Who sets the API terms with the model provider | Vendor's contract | Firm's own contract, firm's choice of provider |
This is the distinction worth sitting with. In a firm-controlled architecture — the model behind private AI deployment — the full corpus, the agent layer, the permissions, and the logs never leave the firm's infrastructure. What goes out to an LLM provider is the minimized, task-specific slice: the handful of retrieved passages actually needed to answer the question in front of the agent, sent under the firm's own negotiated terms with the provider of its choice. The scaffolding — the part that decides what counts as relevant, who's allowed to see it, and what gets recorded — stays put.
That's a meaningfully different risk profile than a system where the vendor's orchestration layer is making those calls on infrastructure the firm can't inspect. It's also a different commercial relationship: the firm isn't locked into one vendor's model choice, because the scaffolding is provider-agnostic by design.
Neither architecture is universally "right." A boutique firm running high-volume, low-sensitivity contract review might reasonably accept a vendor-hosted agent for the speed and lower setup cost. A firm handling M&A due diligence, government investigations, or matters where privilege review is a daily reality has a much stronger case for keeping the scaffolding — and the documents — under its own roof. That's the framing worth using in a case search pilot: define the sensitivity tier of the workload before choosing the architecture, not after.
What "Minimal Context" Actually Looks Like in Practice
This is where most conversations get vague, so it's worth being concrete. In a well-architected retrieval system, an agent working on, say, a motion to compel doesn't ingest the entire client file. It queries a vector index for the passages relevant to the specific request — deposition excerpts referencing a discovery dispute, prior rulings on similar motions, the specific procedural history — and only those retrieved chunks, typically a few hundred to a few thousand tokens, get sent to the LLM for reasoning.
The rest of the client's file — years of correspondence, unrelated matters, financial records, anything outside the scope of the task — never leaves the firm's index. The agent's "vision" is scoped by design, not by hope. That scoping is a property of the retrieval architecture, not the model, and it's enforceable and auditable precisely because the firm controls the layer doing the scoping.
Contrast this with a system where the agent has broad standing access to a connected document store and decides at runtime what's relevant. The scope of what could theoretically be pulled into a given task is much larger, and proving after the fact exactly what was accessed depends entirely on the vendor's logging granularity — which the firm typically has to request, not natively own.
A Pre-Flight Checklist Before Turning On Any Agent
Before a firm enables an agentic feature — whether that's an agent mode inside an existing tool or a dedicated platform — there's a short list of questions that should have documented answers, not assumed ones:
- Scope of visibility: Which matters, clients, or document types can this agent query in a single session? Is that scope set per-user, per-matter, or firm-wide by default?
- Scope of action: Can the agent only draft and retrieve, or can it send, file, or modify records without a human checkpoint?
- Data residency: Where do full documents sit during a session — firm infrastructure, or the vendor's environment?
- What leaves the building: Is it full documents, or minimized, task-specific chunks? Who decided that scoping logic, and can the firm audit it?
- Logging granularity: Can the firm reconstruct, after the fact, exactly what the agent retrieved and why — down to the document level?
- Provider flexibility: Is the firm locked into one model provider's terms, or can it choose and change providers without rearchitecting the workflow?
- Conflicts and ethical walls: Does the agent's retrieval logic respect existing information barriers, or does it operate outside that structure?
Most firms currently answer "we haven't checked" to at least half of these. That's not a knock on any particular IT or innovation team — it reflects how fast agentic features have shipped inside tools firms already had licenses for. Cowork-style agent modes, Copilot extensions, and agentic features inside existing legal platforms often arrive as a toggle, not a procurement decision. Nobody drew up an architecture review because nobody thought "agent mode" needed one.
What Comes Next
Agentic capability is going to keep expanding across every major legal AI platform over the next 12-18 months — that trajectory isn't in question. What is in question is whether firms build the governance layer at the same pace as the capability layer, or retrofit it after an incident makes the gap unavoidable.
The firms handling this well aren't the ones with the most sophisticated model access. They're the ones who mapped, in writing, what any given agent can see and touch before it went live — and who built or chose scaffolding where that mapping is enforceable, not aspirational.
If your firm is evaluating an agentic AI pilot, the useful first exercise isn't a bake-off between vendors. It's mapping which workloads warrant a firm-controlled scaffolding — where the index, permissions, and logs stay on your infrastructure and only minimal context reaches the model — versus which lower-sensitivity tasks can reasonably run on a vendor-hosted agent. That mapping, done before the toggle gets flipped, is worth more than another round of demos.
Frequently Asked Questions
What does 'agentic AI' actually mean in a legal context?
Why does more AI autonomy require more supervision, not less?
What is the difference between vendor-cloud agentic AI and firm-controlled scaffolding?
Related Articles
Agentic AI for Law Firms: What It Actually Means in 2026
What agentic AI actually means for law firms — plain-English definition, what the big players are doing, real deployment examples, and how custom agents differ from SaaS workflows.
Your AI Vendor's Moat Is Your Data. Here's How to Take It Back.
How SaaS AI vendors build competitive moats from your firm's usage data — the shared learning paradox, the dilution problem, and why proprietary AI keeps the compounding advantage with you.
AI for Law Firms in 2026: The Complete Guide to Choosing, Deploying, and Owning Legal AI
Comprehensive guide to AI adoption for law firms in 2026 — agentic AI, proprietary vs SaaS, privilege implications, pricing, and the ownership model.
RAGbase builds private AI systems for law firms: deployed on the firm's own infrastructure, zero data retention, full ownership.
See How RAGbase Works on Your Data
30-minute call. We scope your use case and show the system live.