case studies

300,000 Documents, One Month: Inside a Private AI Deployment

How a mid-sized firm cut drafting time 70% by indexing 300K documents on private infrastructure — and what it means for EU AI Act compliance.

RAGbase Legal Research TeamSeptember 3, 2026 9 min read

Three hundred thousand documents. Thirty days. A 70% drop in drafting time. Those numbers came out of a real deployment at a mid-sized firm this year, and they're worth unpacking — not because the percentage is dramatic, but because of what actually produced it. It wasn't a smarter model. It was a different architecture for where the firm's knowledge lives and how much of it has to move to get an answer.

Most of the legal AI conversation in 2024 and 2025 has centered on model quality — which tool drafts the better memo, which assistant handles the longer context window. That conversation misses where the time was actually going. And it misses a second story unfolding quietly alongside it: the EU AI Act's documentation requirements for high-risk AI use in litigation support, which are about to make architecture — not model choice — the compliance bottleneck for firms that got this wrong from the start.

The Real Bottleneck Wasn't Drafting. It Was Finding.

Before this deployment, associates at the firm were spending hours per matter just locating the right precedent, the right clause, the right prior memo — before drafting had even begun. The documents existed. They were scattered across legacy servers, DMS folders organized by matter number rather than subject matter, and email threads that were, functionally, the firm's second document management system.

This is the unglamorous reality inside most AmLaw 200 and mid-market firms: institutional knowledge is enormous, and almost entirely unsearchable in any way that matches how lawyers actually think about a problem. A associate working a new commercial lease dispute doesn't want to search by matter number or client name. They want to search by fact pattern, clause type, or the argument a partner made in a similar case three years ago. Legacy DMS search doesn't do that. Keyword search across a shared drive doesn't do that either.

So the workaround becomes: re-read the 40-page memo. Again. Three times a week, per the original account of this deployment, because that memo — or one like it — kept being the closest match to a new problem, and there was no faster way to confirm it than opening it and reading it, again.

That's the tax. Not drafting speed. Retrieval speed. And it's the part of the legal AI conversation that per-seat drafting tools, however good their output, don't actually solve if the underlying document estate is still siloed and unindexed.

What Changed: A Retrieval Layer, Not a Better Writer

The deployment didn't hand associates a better drafting assistant. It rebuilt how the firm's documents get found in the first place. All 300,000 files — spanning case management systems, internal memos, and years of archived matters — were indexed into a retrieval layer running on the firm's own infrastructure.

The mechanics matter here, because they're the actual explanation for the 70% figure:

  • Full documents never left the firm's environment. The 40-page memos, the client files, the archived pleadings — all of it stayed indexed and stored on infrastructure the firm controls.
  • Only minimal relevant chunks were sent to the model. When an associate asked a specific question — find the indemnification language used in the last three SaaS vendor agreements — the retrieval layer pulled the specific passages that answered that query, not the underlying documents in full, and passed those fragments to the LLM.
  • The index itself became the searchable layer. Instead of an associate manually triaging which of a dozen similar-sounding memos was the right one, the retrieval system surfaced the specific passage that mattered, ranked by relevance to the actual query.

This is the architectural distinction that gets lost in most legal AI marketing: the corpus and the agent layer are not the same thing as the model, and they don't have to leave the building together. The private AI deployment model separates these two layers deliberately — retrieval, indexing, permissions, and logs stay on infrastructure the firm owns; only the minimized, query-specific chunks travel to whichever LLM provider the firm has selected, under whatever API terms the firm has negotiated.

That's a meaningfully different risk posture than sending full documents — or full matter histories — to a third-party processing environment to get an answer.

Full Corpus vs. Minimized Chunks: Why This Distinction Matters More Than Model Choice

It's tempting to frame the legal AI market as a binary — cloud tools that see everything versus on-prem tools that see nothing. That's not accurate, and it's not the honest differentiator. Nearly every serious legal AI product, RAGbase Legal included, ultimately routes some data to an LLM provider to generate a response. The question that actually matters is how much data, at what layer, and who controls the surrounding infrastructure.

ArchitectureWhere the full corpus livesWhat reaches the LLM providerWho owns audit logs & indexesEU AI Act documentation effort
Consumer AI assistants (e.g., general-purpose chat tools)Uploaded ad hoc, often outside firm governanceWhatever the user pastes or uploads, frequently full documentsThe provider, if it exists at allHigh — little to no native audit trail
Per-seat legal SaaS (Harvey, CoCounsel, Legora, Lexis+ Protégé, and similar)Vendor's cloud environment, ingested via connectorsFull documents or large context windows, per the vendor's pipelineThe vendor, with firm-facing exportsModerate — depends on vendor's compliance tooling maturity
Shared-cloud legal AI platformsMulti-tenant cloud infrastructure, logically separatedRetrieved context, but processed on shared infrastructureShared responsibility model, contractually definedModerate — requires vendor cooperation for full audit trail
Private AI deployment (retrieval layer on firm infrastructure)Firm's own infrastructure, fully under firm controlOnly minimal query-relevant chunks, to the firm's chosen LLM providerThe firm, nativelyLow — logs and lineage already exist as a byproduct

The point isn't that per-seat legal SaaS platforms are reckless — Harvey, CoCounsel, and Legora have each built real compliance and security programs, and firms using them are not operating without governance. The point is that where the corpus and the agent scaffolding live determines how much new infrastructure you need to build to satisfy a regulator asking, after the fact, exactly where a specific document was, who touched it, and why. When that layer is already on infrastructure you control, the answer is a query against your own logs. When it isn't, the answer is a request to a vendor, on the vendor's timeline, against the vendor's willingness to expose their own pipeline.

The Numbers Behind 70%

It's worth being precise about what "70% less drafting time" actually measures, because the headline number invites the wrong interpretation — that the AI is simply a faster writer than the associates. It isn't, and that's not where the time savings came from.

The time reduction breaks down into three components, based on the before/after pattern described in this deployment:

  1. Search time collapsed first. Hours per matter spent locating the right precedent or clause dropped to minutes, because the retrieval layer surfaced ranked, relevant passages instead of requiring manual folder-by-folder search.
  2. Re-reading time disappeared. The recurring pattern of re-opening the same 40-page memo multiple times a week — to re-confirm it was still the right precedent — was eliminated, because the system could confirm relevance at the passage level without a full re-read.
  3. Drafting itself sped up modestly, mostly because associates started drafting from confirmed, correctly-scoped source material on the first pass, rather than drafting, discovering the precedent was wrong, and redrafting.

The compounding effect of the first two — search and re-reading — accounts for the overwhelming majority of the 70% figure. This is consistent with what most firms find once they instrument the actual associate workflow: drafting is a smaller share of total matter time than firms assume, and retrieval is a much larger one. It's also the same finding documented in RAGbase's broader AI for law firms guide — the highest-leverage AI intervention in most firms isn't a better writer, it's a better index.

The EU AI Act Angle Most Firms Haven't Priced In

The compliance story here is easy to miss because it wasn't the reason the deployment happened. But it's becoming the more consequential story for firms evaluating AI vendors through 2025 and into 2026.

Under the EU AI Act, AI systems used in litigation support and other legal contexts are treated as high-risk applications, which triggers specific documentation obligations: firms need to be able to show where relevant data lives, who accessed it, and why, for any high-risk AI use. This isn't a hypothetical future requirement — enforcement expectations are already shaping how firms structure vendor contracts and internal AI governance in 2025.

Here's the operational problem this creates for firms running AI on shared vendor infrastructure: that documentation request goes to someone else's system. If the retrieval index, the access logs, and the query history live inside a vendor's multi-tenant cloud, producing a complete, defensible audit trail means depending on that vendor's willingness and ability to export it in the format a regulator or opposing counsel will accept.

When the retrieval layer, indexes, permissions, and audit logs already sit on the firm's own infrastructure, that same documentation request is a straightforward export against systems the firm already owns and already queries. It's not a new compliance project bolted onto the AI deployment after the fact — it's a natural byproduct of how the system was built to run in the first place.

This is the quieter argument for infrastructure-level control that has nothing to do with which model writes better prose, and it's the argument that's going to matter more, not less, as EU AI Act enforcement matures and as U.S. state-level AI disclosure rules start following a similar pattern for case search and litigation support tools.

What This Means for Firms Evaluating AI Right Now

The market framing of "which AI assistant is best" is the wrong first question for most firms. The better first question is architectural: where does our document corpus live, and how much of it has to leave our control to get a useful answer?

For firms with straightforward drafting needs, thin document histories, and low sovereignty requirements, a per-seat legal SaaS tool remains a reasonable, fast-to-deploy option — the vendor compliance programs are real, and the setup cost is low. For firms sitting on large, fragmented document estates — the kind that accumulate over decades of matters, mergers, and departed partners' file cabinets digitized in 2019 — the calculus shifts. The value in those firms isn't a faster writer. It's finally being able to find and trust what's already there, without exporting the whole thing somewhere else to do it.

The 300,000-document deployment described here didn't succeed because of clever prompting or a superior model. It succeeded because someone finally indexed the firm's actual institutional memory and put a retrieval layer in front of it that lives where the firm can see it, audit it, and — increasingly — prove it, when a regulator asks.


Firms evaluating this shift should start by mapping their own document sprawl before comparing vendors on model benchmarks: how many systems hold matter history, how fragmented is the DMS structure, and how much associate time is actually retrieval versus drafting. That inventory, more than any leaderboard comparing model outputs, is what determines whether a per-seat drafting tool or a private retrieval layer is the right next investment.

Frequently Asked Questions

How long does it take to index 300,000 documents for a private legal AI deployment?
In the case described here, full indexing of 300,000 documents across case management systems, internal memos, and archived matters took roughly one month, including retrieval-layer configuration and permission mapping. Timelines vary with document format quality and how fragmented the source systems are, but a mid-sized firm's full document history is typically indexable in four to eight weeks.
Does deploying AI on private infrastructure mean no data ever reaches an LLM provider?
No. Full documents and the retrieval index stay on the firm's own infrastructure, but the minimal relevant chunks needed to answer a specific query are still sent to the selected LLM provider under the firm's chosen API terms. The distinction is between exposing an entire corpus versus exposing only the fragments necessary for a single answer.
How does private AI deployment help with EU AI Act compliance for high-risk legal use cases?
The EU AI Act requires firms using high-risk AI in litigation support to document where data lives, who accessed it, and why. When the retrieval layer, indexes, and audit logs already sit on the firm's own infrastructure, that documentation already exists as a byproduct of normal operations — it's an export, not a new compliance project.

Related Articles

R
RAGbase Legal Research Team
Research

RAGbase builds private AI systems for law firms: deployed on the firm's own infrastructure, zero data retention, full ownership.

See How RAGbase Works on Your Data

30-minute call. We scope your use case and show the system live.

We use audience and marketing cookies (Google Analytics, LinkedIn). No tracker loads without your consent. Learn more