MANIFESTO LAWAGENT · EDITION 3 · May 2026

AI's False Dichotomy: why better prompts don't solve the hallucination problem

By LawAgent Team · Institutional editorial analysis · 8 min read · 16 min podcast
AI's False Dichotomy: why better prompts don't solve the hallucination problem

Executive Summary

On June 22, 2023, a federal judge in New York sanctioned lawyers who filed a brief citing six nonexistent judicial decisions — fabricated by ChatGPT. The Mata v. Avianca case cemented a collective fear: if AI invents case law, it is incompatible with legal practice. The question matters, but the popular answer is imprecise. The real incompatibility isn't between generative AI and the law — it's between one specific kind of naive AI application and the law.

Hallucination is a property, not a bug

Large language models are statistical prediction systems. Given a context, they calculate the most probable next word based on patterns observed during training. This mechanism — extraordinarily powerful for producing fluent text — is, by construction, insensitive to truth. The model has no concept of reality against which to check its output. It produces whatever statistically resembles what it saw in its training data.

This is the phenomenon labeled hallucination. In domains where a text's surface plausibility is what matters, hallucinations are rare or irrelevant. In domains where every statement must be factually correct on pain of material harm — medicine, engineering, law — hallucinations are catastrophic.

The critical point, frequently overlooked: hallucination isn't a bug to be fixed. It's a statistical property of the method. More sophisticated models reduce its frequency. Models trained on legal data reduce its frequency. Better prompts reduce its frequency. None of the three eliminates it.

The rigorous quantification was carried out by Stanford RegLab and the Stanford Institute for Human-Centered AI. The study Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, published in the Journal of Empirical Legal Studies in 2025, tested the three leading commercial AI legal-research tools — all built on retrieval-augmented generation (RAG). The results: Lexis+ AI showed a hallucination rate above 17%; Westlaw AI-Assisted Research, above 33%. None eliminated hallucinations.

In Brazil, three cases consolidate the pattern: in April 2023, the TSE (Superior Electoral Court) fined a lawyer who filed a brief drafted entirely by ChatGPT in an amicus curiae petition; in February 2025, the TJSC (Santa Catarina Court of Justice) imposed a fine of 10% of the case's value and reported the case to the OAB/SC (the state bar association); in October 2025, a TRT-12 (12th Regional Labor Court) judge characterized a filing as a "nonexistent procedural act," explicitly invoking OAB Recommendation 001/2024. By early 2026, the international AI Hallucination Cases database recorded more than 600 court cases involving fabricated case law, with more than 128 sanctioned lawyers.

The correct inference to draw from these cases isn't that generative AI is incompatible with legal practice. It's that certain forms of AI use — direct use of general-purpose models with no verification architecture, combined with an absence of qualified human oversight — are incompatible with legal practice.

Why RAG alone doesn't solve it

When hallucination began to be recognized as a serious operational risk, vendors' dominant response was uniform: grounding through retrieval-augmented generation. RAG, an architecture introduced in a 2020 NeurIPS paper, works in three steps: the system retrieves relevant documents from a knowledge base, feeds those documents to the model as context, and the model generates its answer based on that enriched context. The promise: the model no longer needs to hallucinate, because the correct information is right in front of it.

The promise is partly true and partly misleading. RAG does reduce the hallucination rate relative to ungrounded models. The reduction is consistent and measurable. But it doesn't reach zero, and the residual errors take particularly insidious forms.

RAG systems still produce two types of error. The first is hallucination proper: the system misstates the law despite having access to the correct sources. The second is misgrounding: the system states the law correctly, but cites a source that doesn't actually support the claim.

The second type is more dangerous. A fully invented precedent is, in theory, verifiable — opposing counsel or the court discovers it doesn't exist. A correct statement backed by a misgrounded citation is far harder to catch: the decision exists, it's verifiable, but it doesn't say what the system claims it says. This error slips right past a surface-level verification check.

That finding leads to the next question: if a grounding layer alone doesn't solve it, what does?

Three independent audit checkpoints

The answer emerging from both academic research and industry practice is that reliable verification requires independent interventions at multiple points in the pipeline. A single point of control, however sophisticated, is subject to correlated failure. The general principle — known in critical-systems engineering as defense in depth — is the same one that governs safety architectures in aviation, intensive-care medicine, and nuclear power.

The full paper develops three audit layers that need to operate independently:

Layer 1 — Strategy Audit. Before any legal text is produced. Validates the applicability of the relevant statutes to the case, the existence and relevance of the cited precedents, and the internal coherence of the argumentative line. It catches, at the planning stage, errors that would be costly to fix after the text has already been drafted.

Layer 2 — Structural Validation. During pipeline execution. Ensures the document under construction complies with applicable procedural rules and the mandatory checklists specific to that type of filing (a civil complaint, an appellate brief, a tax opinion). It catches a class of failures whose root cause isn't hallucination, but omission.

Layer 3 — Post-Drafting Audit. After the final text is produced. Compares the text, claim by claim, against independent authoritative sources. It includes checking that citations exist, checking their grounding (this is what catches the misgrounding error Stanford identified), checking that the cited rules are still in force, checking internal consistency, checking compliance with the strategic plan, and explicitly flagging any unanchored claims.

Independence among the three layers is a necessary condition for their combined effectiveness. If all three audits were run by the same underlying mechanism, their failures would be correlated: an error that slipped past the first layer would have an elevated probability of slipping past the other two as well.

▶ WATCH THE FULL DISCUSSION

A technical analysis of the three-layer independent audit architecture.

Architectural Human-in-the-Loop, not a disclaimer

Even the best automated audit architecture is still not enough to replace qualified human judgment at one critical point: the moment the document is signed and filed. That isn't a limitation of the technology — it's a deliberate architectural principle.

In his sanctioning opinion in Mata v. Avianca, Judge P. Kevin Castel was explicit: "technological advances are commonplace, and there is nothing inherently improper about using a reliable artificial intelligence tool for assistance." The lawyers weren't sanctioned for using ChatGPT. They were sanctioned for abandoning their gatekeeping role over the accuracy of their own filings.

The concept of the gatekeeping role is the doctrinal point that matters. In Johnson v. Dunn (Northern District of Alabama, July 2025), the court was even more specific: the lawyer whose signature appears on a filing is responsible for every statement asserted as true in it, regardless of who originally drafted it.

In AI systems engineering, the term Human-in-the-Loop (HITL) describes architectures in which decisions with material consequences require explicit human intervention before being executed. In high-consequence domains — and legal practice is, by every relevant measure, a high-consequence domain — HITL is a non-negotiable requirement.

It's common for AI vendors to include, in their terms of service, clauses disclaiming liability for incorrect output. That practice doesn't satisfy what modern engineering means by HITL. HITL isn't a legal disclaimer. It's a technical property of the system. A system that simply prints a "verify everything" warning is not HITL. A system whose interface, metadata, and workflows are structured to actively facilitate qualified human review is HITL.

🎧 LISTEN TO THE FULL PODCAST

A deeper technical dive into audit architecture and the reviewing lawyer's responsibility.

Brazil's regulatory framework

The HITL principle isn't just good engineering practice. In the Brazilian context, it's an explicit regulatory requirement. CFOAB Recommendation No. 001/2024, approved in November 2024, states in item 3.3 that "excessive reliance on AI tools is inconsistent with the practice of law and cannot replace the analysis performed by the lawyer." Item 2.2 requires due diligence in vendor selection: contractually verifying that the vendor protects information, adopts security measures, and prohibits using client data to train its systems.

CNJ Resolution No. 615/2025, in force since July 2025, although addressed to the Judiciary, carries significant indirect regulatory weight. It establishes mandatory principles for AI solutions used within the Judiciary: human oversight, explainability, risk classification, registration with Sinapses, and algorithmic impact assessment for high-risk systems. Vendors that operate simultaneously in the Judiciary and in private legal practice tend to align their solutions with these standards — which, in practice, become an informal benchmark of sector best practice.

Combining these rules with the Statute of the Legal Profession (professional secrecy, art. 7, XIX) and with the LGPD (Brazil's data protection law) yields the operational requirements that serious corporate legal-AI systems need to satisfy in Brazil: verifiable logical tenant isolation, a prohibition on using client data to train shared models, a multi-layer independent audit architecture, complete traceability by flow identifier, explicit mechanisms for flagging uncertainty, and architectural transparency toward the client firm.

Safety isn't a feature — it's architecture

The central argument can be summarized in one proposition: reliability in corporate legal AI isn't a feature added to the product at the end of development. It's an architectural decision made on day one.

That distinction separates two engineering philosophies. The first assumes the underlying model is trustworthy and treats errors as exceptional cases correctable by a layer of surface-level verification. The second assumes the underlying model will get things wrong — not because it's badly designed, but because language models are statistical systems insensitive to truth — and designs the entire product to catch those errors before they ever reach the reviewing lawyer.

The first philosophy produced Mata v. Avianca, the Brazilian cases at the TSE, TJSC, and TRT-12, and the more than 600 cases documented internationally. The second philosophy produces systems that, while imperfect — no system will ever be perfect — meet the standard of diligence the legal profession demands.

A corporate legal AI system isn't trustworthy because its underlying model is good. It's trustworthy because its architecture assumes, from day one, that its underlying model will get things wrong — and catches those errors before they ever reach the lawyer.

LawAgent · Architectural thesis

The false dichotomy between legal certainty and generative AI is false because it assumes the technology is still what it was in 2023, when the Mata case was decided. The technology isn't the same. The architecture isn't the same. The regulatory standards aren't the same. And the profession's expectations aren't the same. What remains constant is the lawyer's responsibility — and that has never been something that could be outsourced.

Between a profession paralyzed by fear of a technology it doesn't yet understand, and a profession that discerningly absorbs the technical advances that mature with each research cycle, there is only the willingness to understand, in detail, how contemporary architectures actually work.

If this analysis was useful to you or your firm, share it with other senior partners. Manifesto LawAgent's content is designed to circulate among Brazil's top legal professionals.

EDITORIAL NOTE

LawAgent Team

Institutional editorial analysis

Content produced by the LawAgent team with the support of artificial intelligence tools, based on research conducted by the team. LawAgent is the legal AI companion designed for senior partners, specialized boutiques, and in-house legal departments.

Upcoming editions

The next edition of the LawAgent Manifesto will be published soon.

Take this edition in PDF.

Executive version ready to circulate internally at your firm.

Download PDF

(PDF in Portuguese)

Want to see how LawAgent applies these ideas in the day-to-day of Brazilian law firms?

Subscribe to LawAgent-ai→