In June 2023, a federal judge in New York imposed a $5,000 sanction after lawyers submitted court papers containing cases that did not exist. ChatGPT had produced the authorities. The citations looked plausible, the quoted opinions sounded judicial, and the lawyers failed to verify them.
That case became shorthand for AI hallucination, but its deeper lesson is managerial. The software produced unreliable material. People then placed that material into a consequential process without an effective check.
Businesses often discuss hallucinations as though they were a temporary defect that the next model will remove. Models will improve. The need to verify important claims will remain, because the risk comes from the combination of fluent machines, busy humans, and workflows that reward speed.
Fluency is not evidence
A spreadsheet error often looks like an error. A broken formula produces an odd total or an empty cell. A language model can produce a false answer with correct grammar, a measured tone, and a citation shaped exactly like a real one. The presentation lowers the reader's guard.
The National Institute of Standards and Technology uses the term "confabulation" for confidently stated false or erroneous content. Its generative AI risk profile treats the problem alongside data privacy, information integrity, harmful bias, and information security. That is the right frame. Hallucination is not a curious quirk. It is an operational risk.
Specialist systems reduce the risk but do not erase it. Stanford researchers tested leading legal research products and reported incorrect information in more than 17% of responses from Lexis+ AI and Ask Practical Law AI, while Westlaw's AI-Assisted Research hallucinated more than 34% of the time in their test. These were tools built for law and connected to legal sources, not general chatbots improvising from memory.
The failure usually happens after the output
The first error belongs to the model. The consequential error usually belongs to the process. Someone copies a figure into a board paper, sends a policy summary to staff, quotes a supposed regulation to a client, or inserts a fabricated source into a report. The output acquires authority because it has crossed into an official document.
In the New York case, the court did not sanction a machine. It sanctioned the professionals and their firm. The order also required letters to the judges falsely named as authors of the invented opinions. Accountability remained with the people who filed the work.
That principle travels well beyond law. If an AI-generated product specification reaches a supplier, the purchasing manager remains responsible. If a marketing claim reaches a customer, the company remains responsible. If a financial forecast reaches a lender, the finance director cannot delegate accuracy to a text box.
Match the check to the consequence
Not every AI output needs the same scrutiny. A list of possible meeting times requires little. A claim about tax, health, safety, employment law, investment performance, or a contractual obligation requires much more. The control should follow the potential harm, not the convenience of the tool.
A practical system has three levels. Low-consequence internal drafts receive an ordinary human read. Material going to customers or senior decision-makers receives source checking for every factual claim. High-consequence work receives a qualified second reviewer who did not create the draft.
The verification itself must reach the primary evidence. Asking the same model, "Are you sure?" is not verification. Nor is asking a second general chatbot and counting agreement as truth. Open the cited report. Find the figure in the table. Read the regulation. Confirm the quotation in the judgment. If the source cannot be produced, remove the claim.
Design for correction
Businesses should also preserve a record of inputs, outputs, sources, reviewer decisions, and the final version for important work. That creates traceability when a client challenges a claim or a regulator asks how a decision was reached. It also turns errors into training material rather than private embarrassment.
The aim is not zero error. No human newsroom, law firm, consultancy, or finance department achieves that. The aim is to make serious errors hard to publish, easy to detect, and quick to correct.
Responsible AI use begins with a simple allocation of roles. The machine may draft and organize. A named person owns the truth.
