Skip to main content

Command Palette

Search for a command to run...

The $5,000 Brief That Changed Legal AI Forever: Bar Sanctions, Fake Cases, and What Your Firm Needs to Do Now

Real cases, real suspensions, real consequences. Here is what happened when lawyers trusted AI without verifying the output, and the protocol that protects you.

Updated
12 min readView as Markdown
The $5,000 Brief That Changed Legal AI Forever: Bar Sanctions, Fake Cases, and What Your Firm Needs to Do Now

In June 2023, a federal judge in New York did something that made headlines across the entire legal profession. He sanctioned two attorneys and their firm for submitting a brief that cited six court cases that did not exist. Not obscure or hard to find. Completely fabricated. Every one of them invented by ChatGPT, complete with plausible docket numbers, real judges' names, and invented legal reasoning that the judge later described, in one instance, as "gibberish."

That case, Mata v. Avianca, Inc., became the moment AI use in legal practice stopped being a curiosity and became a liability. Three years on, it has not gotten better. As of early 2026, researcher Damien Charlotin has catalogued over 1,174 documented incidents of AI hallucinated content submitted in courts worldwide. In the US alone, 518 of those cases originated in 2025 onward, and the rate is accelerating, not slowing.

This piece walks through what actually happened in the cases that defined this issue, what the sanctions looked like, what courts are doing about it now, and how small and solo firm lawyers are building verification habits that keep them safe without abandoning tools that genuinely save them time. Lawyers already using Ovviously for research and drafting across multiple jurisdictions will recognize the structural difference between grounded, citation-backed AI output and the kind of freeform generation that created these problems.

The term "hallucination" sounds technical. In practice, it is straightforward. A large language model generates text by predicting what words should come next based on patterns in its training data. It does not look up facts. It does not check a database. When asked for a case citation, it constructs something that looks like a case citation, with the right formatting, plausible names, real court names, and invented everything else.

The result is text that is confident, well formatted, and wrong.

Studies published in 2025 by Stanford's CodeX Center found that general purpose AI tools fabricate case citations in roughly 30 to 45 percent of legal research responses, depending on complexity. Even tools built specifically for legal research are not immune. In 2025 Stanford research testing legal research tools, Westlaw AI showed a 34 percent hallucination rate and Lexis+ AI exceeded 17 percent.

The problem is not that lawyers are careless. It is that AI output looks exactly like real output. Which is precisely why verification is not optional.

Case Study One: Mata v. Avianca (2023)

The facts of Mata v. Avianca are now required reading in ethics education programs across the US. Plaintiff's attorney Steven Schwartz used ChatGPT to research a motion opposing Avianca's attempt to dismiss a personal injury claim. ChatGPT returned six case citations. Schwartz did not verify them.

When Avianca's lawyers could not locate the cases, they flagged it to the court. The judge ordered Schwartz to produce copies of the cited opinions. Schwartz then did something that made everything significantly worse: he asked ChatGPT whether the cases were real. ChatGPT confirmed they were. He submitted those fabricated confirmations to the court.

At the sanctions hearing, Schwartz testified that he had been "operating under the false assumption that this website could not possibly be fabricating cases on its own." The judge was not moved.

The outcome: a $5,000 sanction against Schwartz, his co-counsel Peter LoDuca, and their firm. Both attorneys were also required to write personal letters to each judge whose name appeared in the invented opinions. The case was also dismissed on separate grounds.

The precedent it set was this: signing a filing means you are certifying its accuracy. The fact that AI generated the error is not a mitigating factor. The responsibility sits with the attorney.

Case Study Two: Zachariah Crabill, Colorado (2023)

Where Mata involved experienced attorneys, the Crabill case involved a lawyer two years out of law school, handling a type of motion he had never drafted before. Under deadline pressure and short of time, he turned to ChatGPT for case citations to support a motion to set aside a judgment.

ChatGPT produced cases. Crabill did not verify them. On the morning of the hearing, he realized the citations were fabricated. He texted his paralegal: "I think all of my case cites from ChatGPT are garbage."

What happened next mattered as much as the original error. At the hearing, when the judge raised concerns about the cases, Crabill blamed a legal intern. He did not come forward voluntarily. Six days later, he filed an affidavit admitting the truth.

The Colorado Supreme Court's disciplinary office suspended Crabill for 90 days and imposed a one year and one day suspension, with the remainder stayed pending two years of probation. It was among the first attorney suspensions anywhere in the country specifically connected to AI misuse.

The lesson courts drew from this case was not that AI cannot be used. It was that failing to verify, and then failing to disclose, will be treated as a conduct issue, not just a technical error.

The Scale of the Problem in 2025 and 2026

Mata v. Avianca was treated by many in the profession as a one time story. It was not.

By 2024, Law360's AI tracker had documented 280 incidents in US courts. By the close of 2025 that number had reached 729 and growing. In the first quarter of 2026, new cases were being added weekly.

Sanctions have escalated in parallel. The early cases drew \(500 fines and judicial admonishments. By late 2025, five figure penalties had become routine. In March 2026, the Sixth Circuit Court of Appeals levied \)30,000 in combined sanctions against two attorneys whose brief contained over two dozen fake citations. The court forwarded the matter to the chief judge for potential disciplinary proceedings.

A Thomson Reuters review of US court filings from July 2025 found 22 separate cases in a single month where courts or opposing parties identified fabricated citations in filed documents. These were not all high stakes corporate disputes. They included a custody case, a bankruptcy filing, and a dispute between a family and a local school board.

In a 2025 immigration case in New Mexico, the plaintiff's attorney contracted with a freelance lawyer to conduct research. The freelancer returned a brief with hallucinated cases. The attorney did not check the work. The court referred him to the state bar. His argument that he had not personally generated the hallucinations was not accepted.

What Courts Are Now Requiring

Courts have not waited for bar associations to move first. Since 2023, federal and state courts across the US have issued a range of standing orders, some requiring certification that AI output was verified by a human before filing, others requiring specific disclosure of which portions of a brief were AI assisted.

Judge Brantley Starr of the Northern District of Texas was among the first to issue a standing order requiring attorneys to certify that AI generated text was "checked for accuracy by a human being." The ABA Practice Law Technology Today checklist notes this as one of many court specific requirements that now vary by jurisdiction.

As of early 2026, more than 35 state bar associations have issued formal guidance on AI use in legal practice, with materially different requirements on disclosure, verification, and supervision.

Why This Keeps Happening

It would be easy to frame these cases as the result of bad actors or extreme carelessness. But Thomson Reuters' 2025 Generative AI in Professional Services Report found that among lawyers who expressed reservations about AI, 40 percent cited accuracy and reliability as their primary concern, nearly double any other concern including lack of a human touch or biased data.

The problem is structural. General purpose AI tools were not built for legal research. They do not query verified legal databases. They generate language that fits the pattern of what a case citation should look like, not language drawn from an actual court opinion. As one analysis put it plainly: ChatGPT is a chatbot. It is not a legal research database.

The cases that resulted in the worst sanctions, including Crabill and the Mata attorneys, share one pattern. The lawyers treated AI output as a starting point for verification and then skipped the verification step entirely. Some compounded this by asking the same AI tool whether its own output was accurate and accepting confirmation as genuine assurance.

The Structural Difference Between Grounded and Ungrounded AI Research

Not all AI legal tools work the same way. The distinction that matters for professional responsibility purposes is whether the tool generates text based on probabilistic pattern matching, or whether it retrieves output from verified, indexed legal databases and grounds its responses in those sources.

General purpose tools like the consumer version of ChatGPT do the former. They generate plausible text without any direct connection to verified legal authorities. Legal specific tools that are built on indexed case law and statute databases, and that surface citations linked directly to source material a lawyer can check, work differently. The hallucination risk does not disappear entirely, but the output is anchored to real sources rather than constructed from probability.

Ovviously is built specifically for lawyers doing research and drafting across multiple jurisdictions, with output grounded through Google Search and indexed legal sources rather than freeform generation. For a solo practitioner or small firm asking whether AI legal research is safe to use, the question is not AI versus no AI. It is which type of AI output you are starting from, and what your verification step looks like before anything reaches a court or a client.

What a Safe Verification Protocol Looks Like

The Illinois Attorney Registration and Disciplinary Commission's guide on AI in legal practice puts it plainly: every cited authority must be real, relevant, and accurately represented. Here is what verification looks like in practice, drawn from the patterns that courts have recognised as responsible conduct.

First, treat all AI generated citations as unverified until confirmed in a primary legal database. Westlaw, Lexis, Google Scholar for US cases, or the equivalent jurisdiction specific database. If the case cannot be found, it is not filed.

Second, read the case. Not just the citation. The actual opinion. Courts have noted cases where the citation existed but the quoted language did not appear in it, or the proposition said to be supported by the case was not what the case held.

Third, verify the specific quotation. AI tools will sometimes accurately identify a real case and then fabricate or distort the quoted language within it.

Fourth, log the verification. A timestamped note of who checked what, and through which database, creates a record that matters if questions are raised later.

Fifth, never use an AI tool to verify its own output. The Mata attorneys asked ChatGPT to confirm its own citations. The model confirmed them. That is not verification. That is asking someone whether they were being honest and accepting their own word for it.

Frequently Asked Questions

What is an AI hallucination in a legal filing? It is when an AI tool generates a case citation, statute, or legal proposition that appears plausible but does not exist or does not support the claimed proposition. The output looks like real legal authority and is not.

What sanctions have lawyers faced for filing AI hallucinated citations? Sanctions have ranged from $500 fines in early cases to $30,000 combined sanctions in 2026. Attorneys have also been suspended from practice, disqualified from representing clients in specific matters, and referred to state bar disciplinary committees. Mata v. Avianca is the landmark case; Colorado's suspension of Zachariah Crabill was one of the first bar discipline outcomes.

Does it matter whether I personally generated the AI output? No. In the New Mexico immigration case documented by Thomson Reuters, the attorney contracted a freelance lawyer who generated the hallucinated citations. The filing attorney was sanctioned because the signature on a court filing certifies its accuracy regardless of who drafted it.

Is AI legal research inherently unsafe? Not if the tool grounds its output in verified legal sources rather than generating text from pattern matching. The distinction matters. Tools built on indexed case law and statute databases with source citations carry a meaningfully different risk profile than general purpose chatbots.

How do courts want AI use disclosed? Requirements vary significantly by jurisdiction. Some courts require certification that all AI output was human verified. Others require disclosure of which sections were AI assisted. At least 35 state bar associations have now issued formal guidance. Checking your specific court's standing orders before filing is now a practical prerequisite.

What AI tools are designed specifically for lawyers? Several tools are built for legal research rather than general language generation. Ovviously supports research and drafting across multiple jurisdictions including India, UK, US, Canada, and Australia, with output grounded in verified sources. It is built for solo practitioners and small firms rather than enterprise legal departments.

The Bottom Line

The Mata v. Avianca case did not create new ethical obligations for lawyers. It revealed that existing obligations, the duty to verify what you file, the duty to be candid with the court, the duty to supervise the work product your name goes on, apply to AI output exactly as they apply to work product from a paralegal or a junior associate.

The number of sanctions since 2023 shows that awareness of this problem has not been enough on its own to prevent it. The difference between the lawyers who have been sanctioned and those who have not is not the use of AI. It is whether they built a verification step into their workflow before anything left the firm.

That step is not complicated. Read the cases. Check the citations. Confirm the quotes. Log the check. Do not ask the AI whether the AI was right.

More from this blog

O

Ovviously

23 posts

Ovviously is an AI-powered legal platform designed to streamline research and drafting for legal professionals. It allows users to search millions of global legal documents and draft court-ready arguments in a single, unified interface. The tool focuses on providing verifiable citations and strategic litigation support while ensuring user data privacy.