data sovereignty

The Ross Ruling Is a Retrieval-Architecture Story, Not Copyright

The Third Circuit's Thomson Reuters v. Ross ruling exposes a retrieval-architecture risk for law firms using shared-cloud AI. Here's what changes.

RAGbase Legal Research TeamOctober 9, 2026 10 min read

On September 29, 2026, the Third Circuit did something no federal appellate court had done before: it ruled, on the merits, that training an AI system on copyrighted professional content to build a commercially competing product is not fair use — and it affirmed that finding on a full record, not a motion to dismiss. Thomson Reuters v. Ross Intelligence is no longer a cautionary district-court opinion that vendors could argue was fact-bound or likely to be narrowed on appeal. It is now circuit precedent.

Most of the commentary this week will treat this as a copyright story — a win for content owners, a loss for AI developers, a data point for the dozen-plus other training-data suits working through federal courts. That reading isn't wrong, but it's incomplete, and it misses the part that should actually worry managing partners and CIOs at AmLaw 200 firms. Read closely, the Third Circuit's reasoning isn't really about whether AI training is bad. It's about what happens when protected material leaves your control and gets absorbed into infrastructure you don't own, to perform a function that competes with or substitutes for the original. That fact pattern is not unique to Thomson Reuters and Ross. It is, structurally, identical to what happens every time a firm ships raw client documents, work product, or proprietary research to a third-party platform for indexing, embedding, or fine-tuning.

What the Third Circuit Actually Decided

The underlying facts are narrow but instructive. Ross Intelligence, a legal research startup, built a competing search product by training on Westlaw headnotes and the West Key Number System — Thomson Reuters' proprietary editorial organization of case law. The district court found direct infringement on a defined set of headnotes and rejected Ross's fair use defense on summary judgment, reasoning that Ross's use was commercial, that it used the protected material to build a product serving the same function as Westlaw search, and that the market substitution effect was direct rather than transformative.

The Third Circuit's affirmance locks in three findings that now carry circuit-wide weight:

  • Training-time reproduction counts as reproduction. The court rejected the argument that ingestion for model training is categorically different from traditional copying because the output is statistical rather than literal. If protected expression is necessary input to produce a competing function, the fair use analysis applies to the training process itself, not just to visible outputs.
  • "Same commercial function" defeats transformativeness arguments. Ross's product performed legal search — the same core function as Westlaw. The court found this dispositive against a transformative-use defense, regardless of how the underlying model worked internally.
  • Market harm doesn't require proof of lost specific customers. It was enough that Ross's tool was built to substitute for Westloyal's licensed product in the same market, competing for the same subscription dollars.

None of this makes AI training per se unlawful. Transformative research tools, tools trained on licensed or properly cleared corpora, and tools that don't compete functionally with the source material remain on firmer ground. But the ruling eliminates a defense that a lot of legal AI procurement has quietly relied on: the assumption that if a model's output doesn't look like a verbatim copy, the underlying training data problem takes care of itself.

Why This Is a Retrieval-Architecture Story, Not Just a Copyright Story

Here's the reframe that matters for in-house counsel and legal ops leaders evaluating AI vendors this quarter. Strip away the Westlaw-specific facts and the court's test reduces to two questions: What was the system trained on, and did it reproduce protected work product to perform the same function the original was built to perform?

Now apply that test to a common legal AI procurement pattern: a firm uploads a decade of deal documents, pleadings, and internal memoranda to a vendor's cloud platform so the vendor can "index" or "fine-tune" a model for better firm-specific performance. The documents leave the firm's infrastructure. They're processed, embedded, and in some vendor architectures, retained in ways that blur the line between query-time retrieval and training-time absorption. If that vendor's broader platform — or a future model version — ends up performing functions that look like the firm's own research, drafting, or analysis product, the Ross framework for "what was it trained on" and "same commercial function" becomes directly relevant, just with the firm's own work product standing in for Westlaw's headnotes.

This is not a hypothetical edge case. It is the default data flow for a meaningful share of "fine-tune on your documents" and "train a custom model" offerings in the legal AI market, and it's precisely the exposure that a retrieval-based architecture is built to avoid.

Full Corpus vs. Minimized Chunks: The Distinction That Now Matters Legally, Not Just Operationally

The honest version of this argument isn't "vendors send your data out, we never do." Every serious legal AI platform, RAGbase included, ultimately calls a large language model from a provider like Anthropic, OpenAI, or a hosted open-weight model — someone's tokens travel somewhere. The question Ross makes legally material is what travels, how much of it, under whose terms, and whether it's retained or used to improve a shared system.

Data Flow ElementShared-Cloud / Fine-Tuning PatternRetrieval-Architecture Pattern (e.g., RAGbase)
Full document corpusUploaded to vendor infrastructure for indexing or trainingStays on firm's own infrastructure or private cloud tenancy
Retrieval/vector indexHosted and controlled by vendorHosted and controlled by the firm
What reaches the LLM providerOften full documents or large context windowsMinimal, task-specific retrieved chunks
Terms governing model callsVendor's platform-level agreementFirm's own negotiated API terms with chosen provider
Model improvement / training reuseFrequently ambiguous or opt-out-by-requestNot applicable — no persistent corpus sits with the model provider
Audit trail of what was sent, when, to whomVaries by vendor, often opaqueLogged at the firm's permission and workflow layer
Exposure if provider's training practices are later challengedFirm's corpus sat inside the exposed systemFirm's corpus never left firm-controlled infrastructure

The structural difference is full corpus plus agent layer under client control, versus minimized chunks sent to a model provider. A firm using a properly architected retrieval system can tell a general counsel, a client auditor, or eventually a court exactly what left its walls and why — a single clause from a single contract, retrieved for a single query, sent to a model chosen and contracted by the firm. That's a fundamentally different fact pattern from a corpus-wide upload for indexing, and after Ross, that difference is no longer just a security preference. It's becoming the kind of fact a firm will want on the record if its own AI vendor relationships are ever scrutinized.

Where the Exposure Actually Concentrates

Not all legal AI deployments carry equal risk under this framework. The exposure scales with three variables: how much raw content leaves firm infrastructure, how persistent that content is once it arrives, and how functionally similar the resulting tool is to the firm's own paid work product.

  • Highest exposure: Platforms that require bulk upload of firm documents for vendor-side fine-tuning or "custom model" creation, especially where retention and training-reuse terms are vague or where the vendor's broader product competes in adjacent markets (legal research, document review, drafting).
  • Moderate exposure: Shared-cloud AI assistants that process full documents in large context windows per session, where documents aren't used for training but do transit and briefly reside in vendor infrastructure under standard enterprise terms.
  • Lowest exposure: Architectures where the firm's full corpus and index never leave firm-controlled infrastructure, and only minimized, query-specific chunks reach a model provider under the firm's own negotiated terms — the pattern underlying tools like RAGbase's private AI deployment and its retrieval-based case search.

This isn't an argument against the broader legal AI market. Harvey, Legora, CoCounsel, Lexis+ Protege, and Claude Cowork all serve real workflows, and firms running pilots against those platforms should keep doing so — our AI for law firms guide walks through where each tends to fit. The point is narrower and more actionable: for sovereignty-critical workloads — privileged litigation files, unfiled deal terms, regulatory investigations, anything a GC would call "the crown jewels" — the Ross framework gives firms a new, concrete reason to ask vendors exactly what happens to the corpus, not just what the output looks like.

Five Questions Every AmLaw 200 Firm Should Be Asking Vendors This Month

The Ross affirmance gives procurement teams specific, defensible language to use in vendor diligence. Five questions now belong in every legal AI contract review:

  1. Does our full document corpus ever leave our infrastructure, or only retrieved fragments at query time? Get this in writing, not in a sales deck.
  2. Is any part of our corpus used for model training, fine-tuning, or "platform improvement," even in anonymized or aggregated form? Ross turned on training-time use — opt-out clauses buried in a terms-of-service update aren't sufficient diligence anymore.
  3. If the vendor's broader product changes direction — acquired, pivoted, or expanded into a market adjacent to our own services — what happens to data already ingested? Firms can't contract their way out of a vendor's future corporate strategy, but they can limit exposure by never transferring the corpus in the first place.
  4. Who controls the retrieval/index layer — us or the vendor? Index ownership determines whether the firm can audit, in granular terms, exactly what was retrieved and sent for any given query — the kind of record that matters if a dispute ever reaches discovery.
  5. Under whose API terms does the actual model call happen? A firm-negotiated agreement with the model provider is a materially different risk posture than inheriting the vendor's blanket terms across its entire customer base.

The Competitive Landscape, Read Through a Risk Lens

None of the major legal AI platforms are "wrong" for every use case, and the market is moving quickly toward hybrid models — Thomson Reuters itself, notably, is both plaintiff in this case and operator of CoCounsel, which should tell firms something about how seriously incumbents are taking their own training-data provenance. Claude Cowork, Lexis+ Protege, and similar per-seat tools are optimized for breadth, speed of deployment, and ecosystem integration, and they make sense for research, summarization, and drafting work that doesn't involve the firm's most sensitive corpus. Agentic platforms like Legora and Harvey are pushing into workflow automation that firms will want to pilot regardless of this ruling.

What changes after Ross is the calculus for a specific, high-value slice of firm work: anything where the underlying documents are themselves the asset — litigation strategy, unreleased deal terms, proprietary research product the firm sells or licenses to clients. For that slice, architecture — not brand, not benchmark scores — is now the primary risk variable, and it's worth reading how that distinction plays out against your own data practices in pieces like why your data is their moat.

What Comes Next

Expect three second-order effects over the next twelve months. First, vendor contracts will get more specific about training-data provenance and retention, because general counsel's offices are going to start asking Ross-shaped questions in due diligence — this ruling will show up in redlines before it shows up in another lawsuit. Second, expect at least one of the pending training-data cases against general-purpose model providers to cite Ross's "same commercial function" test directly, which will sharpen the line between transformative research tools and functional substitutes. Third, and most relevant to legal procurement specifically: firms that can demonstrate architectural separation between their corpus and any third party's training pipeline will have a genuine answer ready the next time a client audit, a bar inquiry, or opposing counsel asks what their AI tools actually did with firm documents.


If your firm is running AI pilots across sensitive matters right now, the Ross ruling is a reasonable prompt to go back to each vendor contract and map exactly what leaves your infrastructure, what doesn't, and who controls the index in between. That audit — not a change in headline strategy — is the actual action item this week.

Frequently Asked Questions

What did the Third Circuit actually decide in Thomson Reuters v. Ross?
The Third Circuit affirmed the district court's finding that Ross Intelligence's use of Westlaw headnotes to train a competing legal search tool was not protected by fair use, because the output served the same commercial function as the original work and the training process required reproducing protected expression. It's the first appellate-level ruling applying fair use doctrine directly to AI training data, giving lower courts and in-house counsel their first circuit-level precedent on the question.
Does this ruling mean law firms can't use AI tools built on legal content?
No — the ruling targets training-time reproduction of protected material to build a directly competing product, not the use of AI generally. Firms using retrieval-augmented systems that reference licensed content at query time, rather than ingesting it into a model's training weights, operate under a materially different fact pattern.
What's the practical difference between sending documents to a vendor for indexing versus using a retrieval architecture?
Indexing or fine-tuning typically requires transferring full documents to a third party's infrastructure, where they may be copied, cached, or used to improve shared models — the same exposure pattern at issue in Ross. A retrieval architecture like RAGbase's keeps the full corpus and index on the firm's own infrastructure and sends only minimal, task-specific text chunks to a selected LLM provider under the firm's own API terms.

Related Articles

R
RAGbase Legal Research Team
Research

RAGbase builds private AI systems for law firms: deployed on the firm's own infrastructure, zero data retention, full ownership.

See How RAGbase Works on Your Data

30-minute call. We scope your use case and show the system live.

We use audience and marketing cookies (Google Analytics, LinkedIn). No tracker loads without your consent. Learn more