Harvey AI raised $300 million at a $3 billion valuation in early 2025. It counts Allen & Overy, PwC, and a growing list of AmLaw 100 firms among its customers. By almost any metric, it is the most visible legal AI company in the world right now.
And yet, in the private conversations that happen after the conference panels end—in the Slack channels of legal innovation teams, in the post-pilot debrief calls—a more complicated picture is emerging. Managing partners who signed enterprise agreements are asking hard questions about ROI. CIOs are surfacing governance concerns that weren't fully resolved at procurement. Associates are finding that Harvey works brilliantly on certain tasks and frustratingly on others.
None of this means Harvey is failing. It means the legal AI market is maturing past the demo stage, and firms are now stress-testing these tools against real work. The complaints are worth examining carefully—not to dismiss Harvey, but because they illuminate the deeper architectural and strategic questions that every firm buying legal AI needs to answer.
The Hallucination Problem Is Not Solved—It's Managed
The most persistent technical complaint about Harvey, and frankly about every frontier legal AI system, is citation hallucination and reasoning drift on complex tasks. Harvey is built on top of OpenAI's models with legal fine-tuning layered on top. That architecture produces impressive results on pattern-matching tasks: drafting standard clauses, summarizing straightforward agreements, generating first-pass research memos.
Where it degrades is on multi-step legal reasoning—the kind that requires holding contradictory precedents in tension, applying jurisdiction-specific nuance, or working through a novel statutory interpretation problem. Multiple legal innovation leads at large firms have reported that Harvey's outputs on these tasks require more substantive review than the tool's marketing implies, which erodes the time-savings calculation.
A specific pattern worth noting: Harvey tends to perform better as a drafting assistant than as a research engine. When asked to synthesize case law across circuits, or to identify the weight of authority on a contested issue, the error rate rises in ways that aren't always legible to a junior associate doing a first-pass review. The confident tone of the output doesn't reliably correlate with accuracy.
This is not a Harvey-specific problem. It is a property of large language models that no vendor has fully solved. But Harvey's positioning—and its pricing, which we've examined in detail in our Harvey pricing breakdown—implies a level of reliability that the underlying technology doesn't yet consistently deliver.
What Firms Are Doing About It
The smarter implementations we've seen treat Harvey as a first-draft accelerator with mandatory human checkpoints, not as an autonomous research tool. That's a reasonable mitigation, but it's worth being honest about what it means for the ROI model: if every Harvey output requires substantive attorney review before it can be used, the efficiency gains are real but narrower than the pitch deck suggests.
Firms that have gotten the most value from Harvey have generally done two things: narrowed the use cases to tasks where the error surface is small (NDA redlining, standard due diligence checklists, clause extraction from high-volume contract sets), and invested in prompt engineering discipline so that attorneys aren't using Harvey ad hoc on tasks it wasn't optimized for.
Data Governance: The Contractual Protections Are Real, But the Architecture Still Matters
This is where the conversation gets more nuanced than most vendor comparisons acknowledge.
Harvey does not, by default, use customer data to train its models. Their enterprise agreements include data protection provisions that are meaningful and auditable. If you ask Harvey's sales team about data security, they will give you a thorough and largely accurate answer.
The concern that sophisticated CIOs are raising is not that Harvey is lying about data handling. It is structural.
When you use Harvey, here is what happens architecturally:
- Your firm's documents and matter data are indexed and stored in Harvey's infrastructure
- The retrieval layer—the system that decides which document chunks are relevant to a query—runs on Harvey's servers
- The agent orchestration layer, which sequences tasks and manages workflows, also runs outside your environment
- Query logs, usage patterns, and retrieved content all flow through infrastructure you don't control
The contractual protections govern what Harvey does with that data. They don't change the fact that the data is there. For most matters, this is an acceptable trade-off. For sovereign wealth fund work, contested M&A, government investigations, or matters involving clients with strict data residency requirements, it may not be.
The more architecturally precise framing—which we explore in depth in our guide to private AI deployment—is the difference between:
Model A (Harvey and most SaaS legal AI): Full document corpus + retrieval index + agent layer live in vendor infrastructure. Contractual protections govern use.
Model B (On-premise or private deployment): Full document corpus + retrieval index + agent layer live in firm-controlled infrastructure. Only the minimal retrieved chunks needed to answer a specific query are sent to an LLM provider (which could be OpenAI, Anthropic, or others), under the firm's own API terms. The model sees a narrow, scoped context window—not your client files.
The distinction matters because in Model B, even if the LLM provider had a data incident or changed their terms, your exposure is limited to the specific chunks that were retrieved for specific queries—not your entire client matter database. That's a meaningfully different risk profile.
| Dimension | SaaS Legal AI (Harvey model) | Private/On-Premise (RAGbase model) |
|---|---|---|
| Document corpus location | Vendor infrastructure | Firm-controlled infrastructure |
| Retrieval/index layer | Vendor-managed | Firm-controlled |
| Agent orchestration | Vendor infrastructure | Firm-controlled |
| What reaches LLM provider | Full context, per vendor's architecture | Minimal retrieved chunks only |
| Audit logs | Vendor-provided | Firm-owned |
| Client data residency | Vendor's data centers | Firm's chosen environment |
| Customization depth | Vendor roadmap | Firm-configurable |
| Time to deploy | Days to weeks | Weeks to months |
The Customization Ceiling
Harvey's product is, by design, a horizontal platform. It is built to work reasonably well across practice areas out of the box, which is the correct strategy for a company trying to grow quickly across a fragmented market.
The complaints that surface at the 6-to-12-month mark of enterprise deployments tend to center on what one legal technology director at a major firm described as "the last 20 percent problem." Harvey gets you 80 percent of the way to what you need on generic tasks. But the 20 percent that's left—the firm-specific workflows, the practice group conventions, the jurisdiction-specific playbooks, the client-specific matter context—requires customization that Harvey's current architecture either doesn't support or supports only through workarounds.
Specific examples that have surfaced:
- Matter-specific context: Harvey can ingest documents, but maintaining persistent, richly structured matter context that accumulates over a multi-year client relationship is not what it's optimized for
- Firm knowledge integration: Connecting Harvey to a firm's proprietary precedent database, internal research memos, or institutional knowledge repositories requires integration work that varies in complexity and reliability
- Workflow automation: Firms trying to build Harvey into multi-step automated workflows—not just one-off queries—find that the agentic capabilities are more limited than the marketing suggests, a point we've explored in our analysis of agentic AI at law firms
- Billing and practice management integration: Connecting Harvey outputs to matter management systems, time entry, or document management in a structured way is typically left to the firm to solve
This is not a criticism unique to Harvey. CoCounsel, Lexis+ Protege, and most SaaS legal AI products face the same customization ceiling. The ceiling is a property of the SaaS model itself: the vendor must serve thousands of customers, so deep customization for any one firm's workflows is economically difficult to prioritize.
The Adoption Gap Nobody Talks About in the Press Releases
Harvey's most publicized metric is the number of firms and attorneys who have access to Harvey. The metric that matters more—and that is rarely publicized—is active weekly usage rates among licensed users.
The legal AI industry has a well-documented adoption problem. Our analysis of the 98 percent adoption gap found that across legal AI deployments broadly, active usage rates among licensed users frequently fall below 20 percent at the 6-month mark. The firms that beat that number share a common characteristic: they invested heavily in change management, workflow integration, and attorney training—not just in the technology license.
Harvey is not immune to this dynamic. Several firms that have spoken publicly about their Harvey deployments describe a pattern where a core group of early adopters—typically younger associates and legal innovation enthusiasts—drive most of the usage, while the broader attorney population engages sporadically or not at all.
The reasons are consistent across vendors:
- Workflow friction: If using Harvey requires leaving the environment attorneys already work in (email, DMS, Word), usage drops sharply
- Trust calibration: Attorneys who've encountered a Harvey error—especially a confident-sounding hallucination—become more cautious about relying on it, sometimes overcorrecting toward non-use
- ROI ambiguity at the individual level: Partners bill for their time. If Harvey saves an associate two hours on a task, the economic benefit accrues to the client or the firm—not obviously to the associate. The incentive to adopt requires active management
What the Complaints Actually Tell Us About Where the Market Is Heading
The Harvey complaints aren't evidence that Harvey is a bad product. They're evidence that legal AI is entering the disillusionment phase of the hype cycle—which is actually healthy, and which typically precedes the productivity phase where real value gets captured at scale.
The firms that will emerge from this phase with genuine competitive advantage are making a set of architectural bets that are worth naming explicitly:
Bet 1: Sovereignty-critical workloads need different infrastructure. Not every matter needs private AI. But the matters that do—the ones involving the most sensitive client relationships, the highest regulatory exposure, the most valuable information—need an architecture where the firm controls the full stack. As one CIO at an AmLaw 50 firm put it: "I don't want to find out the hard way that my data governance posture was wrong on the matter that mattered most."
Bet 2: The retrieval layer is the competitive moat, not the model. Harvey uses OpenAI. CoCounsel uses OpenAI. Lexis+ Protege uses OpenAI. The underlying models are becoming commodities. What differentiates legal AI is the quality of the retrieval architecture—how well the system finds the right precedents, the right clauses, the right matter context. Firms that build or deploy a retrieval layer they control, trained on their own documents and precedents, are building something that compounds over time. Our case search architecture reflects this philosophy directly.
Bet 3: Agentic workflows require trust, and trust requires observability. The next generation of legal AI isn't question-and-answer—it's autonomous agents that draft, review, research, and route work across a matter lifecycle. Deploying that level of autonomy on infrastructure you don't control, with audit logs you access through a vendor portal, is a governance problem waiting to become a liability problem.
The AI for law firms guide we maintain covers the full evaluation framework for firms navigating these decisions—including how to structure vendor assessments and what questions to ask about architecture before signing an enterprise agreement.
The Honest Evaluation Framework
Harvey deserves credit for what it has built: a genuinely capable legal AI product with strong UX, meaningful data protections, and a rapidly improving feature set. For firms that want fast deployment and broad coverage across general legal tasks, it is a serious option.
The questions worth asking before—or during—a Harvey deployment:
- What matters are off-limits? Every firm should have an explicit list of matter types or client categories where vendor-hosted AI is contractually or ethically problematic. Those matters need a different solution.
- What's your customization roadmap? If your 12-month vision involves deep integration with your DMS, proprietary playbooks, and multi-step automated workflows, pressure-test whether Harvey's product roadmap gets you there.
- What does your audit trail look like? When a client asks how their confidential documents were processed, what can you show them?
- What's your fallback architecture? If Harvey's terms change, their pricing shifts, or they get acquired, what does your transition look like? The hidden cost of legal AI SaaS isn't always visible at procurement time.
The firms that are most thoughtfully navigating this moment aren't choosing between Harvey and nothing—they're building a tiered architecture. SaaS tools for high-volume, lower-sensitivity work where speed of deployment matters. Private or on-premise infrastructure for the matters where the firm needs to own the full stack: the retrieval index, the agent layer, the audit logs, the document corpus. The question isn't which vendor wins. It's which architecture matches which workload. If you're at the stage of evaluating where those lines should be drawn in your firm's AI strategy, that's precisely the conversation worth having before the next enterprise agreement lands on your desk.
Frequently Asked Questions
What are the most common complaints about Harvey AI at law firms?
Is Harvey AI safe to use for confidential client matters?
How does Harvey AI compare to on-premise legal AI alternatives?
Related Articles
LexisNexis Protégé vs Harvey vs CoCounsel: What's Missing From All Three
Comparison of the three dominant legal AI platforms in 2026 — what each does well, and the blind spot they all share around internal document access.
Harvey AI Costs $1,200/Lawyer/Month. Here's What You Actually Get (and Don't Get).
Detailed Harvey AI pricing analysis for 2026 — per-seat costs, three-year TCO, what's included, what's missing, and how proprietary AI compares.
Your AI Vendor's Moat Is Your Data. Here's How to Take It Back.
How SaaS AI vendors build competitive moats from your firm's usage data — the shared learning paradox, the dilution problem, and why proprietary AI keeps the compounding advantage with you.
The Hidden Cost of Legal AI: Why 300-Lawyer Firms Are Spending $4.3M on Tools That Can't Find Their Own Case Files
Legal AI subscriptions cost up to $4.3M/year for large firms, yet can't search internal case files. Compare SaaS costs vs proprietary AI ownership economics.
Agentic AI for Law Firms: What It Actually Means in 2026
What agentic AI actually means for law firms — plain-English definition, what the big players are doing, real deployment examples, and how custom agents differ from SaaS workflows.
AI for Law Firms in 2026: The Complete Guide to Choosing, Deploying, and Owning Legal AI
Comprehensive guide to AI adoption for law firms in 2026 — agentic AI, proprietary vs SaaS, privilege implications, pricing, and the ownership model.
RAGbase Legal builds proprietary AI systems for law firms — deployed on the firm's own infrastructure, zero data retention, full code ownership. 80+ enterprise deployments.
See How RAGbase Legal Works on Your Data
Free 3-5 day proof of concept. Your data, your infrastructure, working results.