Three hundred thousand documents. Thirty days. A 70% drop in drafting time. Those numbers came out of a real deployment at a mid-sized firm this year, and they're worth unpacking — not because the percentage is dramatic, but because of what actually produced it. It wasn't a smarter model. It was a different architecture for where the firm's knowledge lives and how much of it has to move to get an answer.
Most of the legal AI conversation in 2024 and 2025 has centered on model quality — which tool drafts the better memo, which assistant handles the longer context window. That conversation misses where the time was actually going. And it misses a second story unfolding quietly alongside it: the EU AI Act's documentation requirements for high-risk AI use in litigation support, which are about to make architecture — not model choice — the compliance bottleneck for firms that got this wrong from the start.
The Real Bottleneck Wasn't Drafting. It Was Finding.
Before this deployment, associates at the firm were spending hours per matter just locating the right precedent, the right clause, the right prior memo — before drafting had even begun. The documents existed. They were scattered across legacy servers, DMS folders organized by matter number rather than subject matter, and email threads that were, functionally, the firm's second document management system.
This is the unglamorous reality inside most AmLaw 200 and mid-market firms: institutional knowledge is enormous, and almost entirely unsearchable in any way that matches how lawyers actually think about a problem. A associate working a new commercial lease dispute doesn't want to search by matter number or client name. They want to search by fact pattern, clause type, or the argument a partner made in a similar case three years ago. Legacy DMS search doesn't do that. Keyword search across a shared drive doesn't do that either.
So the workaround becomes: re-read the 40-page memo. Again. Three times a week, per the original account of this deployment, because that memo — or one like it — kept being the closest match to a new problem, and there was no faster way to confirm it than opening it and reading it, again.
That's the tax. Not drafting speed. Retrieval speed. And it's the part of the legal AI conversation that per-seat drafting tools, however good their output, don't actually solve if the underlying document estate is still siloed and unindexed.
What Changed: A Retrieval Layer, Not a Better Writer
The deployment didn't hand associates a better drafting assistant. It rebuilt how the firm's documents get found in the first place. All 300,000 files — spanning case management systems, internal memos, and years of archived matters — were indexed into a retrieval layer running on the firm's own infrastructure.
The mechanics matter here, because they're the actual explanation for the 70% figure:
- Full documents never left the firm's environment. The 40-page memos, the client files, the archived pleadings — all of it stayed indexed and stored on infrastructure the firm controls.
- Only minimal relevant chunks were sent to the model. When an associate asked a specific question — find the indemnification language used in the last three SaaS vendor agreements — the retrieval layer pulled the specific passages that answered that query, not the underlying documents in full, and passed those fragments to the LLM.
- The index itself became the searchable layer. Instead of an associate manually triaging which of a dozen similar-sounding memos was the right one, the retrieval system surfaced the specific passage that mattered, ranked by relevance to the actual query.
This is the architectural distinction that gets lost in most legal AI marketing: the corpus and the agent layer are not the same thing as the model, and they don't have to leave the building together. The private AI deployment model separates these two layers deliberately — retrieval, indexing, permissions, and logs stay on infrastructure the firm owns; only the minimized, query-specific chunks travel to whichever LLM provider the firm has selected, under whatever API terms the firm has negotiated.
That's a meaningfully different risk posture than sending full documents — or full matter histories — to a third-party processing environment to get an answer.
Full Corpus vs. Minimized Chunks: Why This Distinction Matters More Than Model Choice
It's tempting to frame the legal AI market as a binary — cloud tools that see everything versus on-prem tools that see nothing. That's not accurate, and it's not the honest differentiator. Nearly every serious legal AI product, RAGbase Legal included, ultimately routes some data to an LLM provider to generate a response. The question that actually matters is how much data, at what layer, and who controls the surrounding infrastructure.
| Architecture | Where the full corpus lives | What reaches the LLM provider | Who owns audit logs & indexes | EU AI Act documentation effort |
|---|---|---|---|---|
| Consumer AI assistants (e.g., general-purpose chat tools) | Uploaded ad hoc, often outside firm governance | Whatever the user pastes or uploads, frequently full documents | The provider, if it exists at all | High — little to no native audit trail |
| Per-seat legal SaaS (Harvey, CoCounsel, Legora, Lexis+ Protégé, and similar) | Vendor's cloud environment, ingested via connectors | Full documents or large context windows, per the vendor's pipeline | The vendor, with firm-facing exports | Moderate — depends on vendor's compliance tooling maturity |
| Shared-cloud legal AI platforms | Multi-tenant cloud infrastructure, logically separated | Retrieved context, but processed on shared infrastructure | Shared responsibility model, contractually defined | Moderate — requires vendor cooperation for full audit trail |
| Private AI deployment (retrieval layer on firm infrastructure) | Firm's own infrastructure, fully under firm control | Only minimal query-relevant chunks, to the firm's chosen LLM provider | The firm, natively | Low — logs and lineage already exist as a byproduct |
The point isn't that per-seat legal SaaS platforms are reckless — Harvey, CoCounsel, and Legora have each built real compliance and security programs, and firms using them are not operating without governance. The point is that where the corpus and the agent scaffolding live determines how much new infrastructure you need to build to satisfy a regulator asking, after the fact, exactly where a specific document was, who touched it, and why. When that layer is already on infrastructure you control, the answer is a query against your own logs. When it isn't, the answer is a request to a vendor, on the vendor's timeline, against the vendor's willingness to expose their own pipeline.
The Numbers Behind 70%
It's worth being precise about what "70% less drafting time" actually measures, because the headline number invites the wrong interpretation — that the AI is simply a faster writer than the associates. It isn't, and that's not where the time savings came from.
The time reduction breaks down into three components, based on the before/after pattern described in this deployment:
- Search time collapsed first. Hours per matter spent locating the right precedent or clause dropped to minutes, because the retrieval layer surfaced ranked, relevant passages instead of requiring manual folder-by-folder search.
- Re-reading time disappeared. The recurring pattern of re-opening the same 40-page memo multiple times a week — to re-confirm it was still the right precedent — was eliminated, because the system could confirm relevance at the passage level without a full re-read.
- Drafting itself sped up modestly, mostly because associates started drafting from confirmed, correctly-scoped source material on the first pass, rather than drafting, discovering the precedent was wrong, and redrafting.
The compounding effect of the first two — search and re-reading — accounts for the overwhelming majority of the 70% figure. This is consistent with what most firms find once they instrument the actual associate workflow: drafting is a smaller share of total matter time than firms assume, and retrieval is a much larger one. It's also the same finding documented in RAGbase's broader AI for law firms guide — the highest-leverage AI intervention in most firms isn't a better writer, it's a better index.
The EU AI Act Angle Most Firms Haven't Priced In
The compliance story here is easy to miss because it wasn't the reason the deployment happened. But it's becoming the more consequential story for firms evaluating AI vendors through 2025 and into 2026.
Under the EU AI Act, AI systems used in litigation support and other legal contexts are treated as high-risk applications, which triggers specific documentation obligations: firms need to be able to show where relevant data lives, who accessed it, and why, for any high-risk AI use. This isn't a hypothetical future requirement — enforcement expectations are already shaping how firms structure vendor contracts and internal AI governance in 2025.
Here's the operational problem this creates for firms running AI on shared vendor infrastructure: that documentation request goes to someone else's system. If the retrieval index, the access logs, and the query history live inside a vendor's multi-tenant cloud, producing a complete, defensible audit trail means depending on that vendor's willingness and ability to export it in the format a regulator or opposing counsel will accept.
When the retrieval layer, indexes, permissions, and audit logs already sit on the firm's own infrastructure, that same documentation request is a straightforward export against systems the firm already owns and already queries. It's not a new compliance project bolted onto the AI deployment after the fact — it's a natural byproduct of how the system was built to run in the first place.
This is the quieter argument for infrastructure-level control that has nothing to do with which model writes better prose, and it's the argument that's going to matter more, not less, as EU AI Act enforcement matures and as U.S. state-level AI disclosure rules start following a similar pattern for case search and litigation support tools.
What This Means for Firms Evaluating AI Right Now
The market framing of "which AI assistant is best" is the wrong first question for most firms. The better first question is architectural: where does our document corpus live, and how much of it has to leave our control to get a useful answer?
For firms with straightforward drafting needs, thin document histories, and low sovereignty requirements, a per-seat legal SaaS tool remains a reasonable, fast-to-deploy option — the vendor compliance programs are real, and the setup cost is low. For firms sitting on large, fragmented document estates — the kind that accumulate over decades of matters, mergers, and departed partners' file cabinets digitized in 2019 — the calculus shifts. The value in those firms isn't a faster writer. It's finally being able to find and trust what's already there, without exporting the whole thing somewhere else to do it.
The 300,000-document deployment described here didn't succeed because of clever prompting or a superior model. It succeeded because someone finally indexed the firm's actual institutional memory and put a retrieval layer in front of it that lives where the firm can see it, audit it, and — increasingly — prove it, when a regulator asks.
Firms evaluating this shift should start by mapping their own document sprawl before comparing vendors on model benchmarks: how many systems hold matter history, how fragmented is the DMS structure, and how much associate time is actually retrieval versus drafting. That inventory, more than any leaderboard comparing model outputs, is what determines whether a per-seat drafting tool or a private retrieval layer is the right next investment.
Frequently Asked Questions
How long does it take to index 300,000 documents for a private legal AI deployment?
Does deploying AI on private infrastructure mean no data ever reaches an LLM provider?
How does private AI deployment help with EU AI Act compliance for high-risk legal use cases?
Related Articles
AI for Law Firms in 2026: The Complete Guide to Choosing, Deploying, and Owning Legal AI
Comprehensive guide to AI adoption for law firms in 2026 — agentic AI, proprietary vs SaaS, privilege implications, pricing, and the ownership model.
Your AI Vendor's Moat Is Your Data. Here's How to Take It Back.
How SaaS AI vendors build competitive moats from your firm's usage data — the shared learning paradox, the dilution problem, and why proprietary AI keeps the compounding advantage with you.
Agentic AI for Law Firms: What It Actually Means in 2026
What agentic AI actually means for law firms — plain-English definition, what the big players are doing, real deployment examples, and how custom agents differ from SaaS workflows.
The Hidden Cost of Legal AI: Why 300-Lawyer Firms Are Spending $4.3M on Tools That Can't Find Their Own Case Files
Legal AI subscriptions cost up to $4.3M/year for large firms, yet can't search internal case files. Compare SaaS costs vs proprietary AI ownership economics.
RAGbase builds private AI systems for law firms: deployed on the firm's own infrastructure, zero data retention, full ownership.
See How RAGbase Works on Your Data
30-minute call. We scope your use case and show the system live.