The recent Jason Kelce commercial aimed at data-center water use gets an easy laugh from a real concern. Water, power, labor, copyright and safety all deserve hard questions. The lazy part comes when a joke about AI’s costs turns into a verdict that serious people should avoid the technology.
A commercial will never show the quiet work: the accountant who closes the books sooner, the attorney who begins with a clean contract comparison, the doctor who sees another plausible diagnosis, or the manager who reaches dinner without two hours of unfinished paperwork. That work is less entertaining. It is also where the case for AI is being decided.
What the evidence actually says
The models have improved. On OpenAI’s LongFact and FActScore evaluations, GPT-5 made about 80% fewer factual errors than o3 on open-ended fact questions without search. OpenAI’s current GPT-5.6 system card says its largest model makes slightly fewer factual errors than GPT-5.5 and reproduces user-reported errors much less often. The company also says those test cases were chosen because they were unusually prone to hallucination, so they do not represent all day-to-day use.
People are not a perfect baseline. In a randomized trial with 50 physicians, the language model alone earned a median diagnostic-reasoning score of 92%. Physicians using conventional resources scored 74%. The model’s score was 16 percentage points higher. Yet physicians given the model scored 76%, which was not a meaningful improvement over the control group. Access alone did not make the team better.
The honest conclusion: There is no universal hallucination rate, and no credible basis for saying the problem has disappeared. Error rates change with the model, the task, the source material, the tools and the test. For many bounded business tasks, you can now push the practical risk low enough to save substantial time. High-stakes conclusions still need qualified review.
Use a six-part accuracy routine
1. Spend model quality where errors are expensive. Use a strong reasoning model for contracts, financial analysis, health questions, research and decisions with real consequences. A fast, cheap model is fine for formatting, brainstorming or cleaning up a draft you already understand.
2. Give it the evidence. Attach the agreement, report, spreadsheet, policy, medical record or official web pages. Tell the model which sources control. Asking from memory invites invention; asking from a defined packet turns the task into document analysis.
3. Separate facts, judgment and missing information. Require three labeled sections. Make the model write “not established” when the source packet does not support a claim. This small instruction makes uncertainty visible before it becomes a confident sentence.
4. Demand page-level support. Ask for the file name, page, cell or official link behind every important number and conclusion. Open the cited material and confirm that it actually proves the claim.
5. Run a hostile second pass. Start a new chat and ask it to find unsupported claims, broken math, outdated assumptions, omitted risks and counterevidence. For consequential work, use a second model or a person with domain knowledge.
6. Set the human checkpoint before you begin. Decide who approves a payment, signs a contract, changes the books, communicates medical advice or publishes a factual claim. AI can prepare the decision. Responsibility stays with a named person.
Prompt: Use only the attached materials. First list the controlling sources. Then produce: (1) verified facts with file names and page numbers, (2) calculations with the formula shown, (3) interpretations clearly labeled as judgment, (4) contradictions or missing information, and (5) questions for a qualified professional. If the packet does not support a statement, write “not established.” Do not invent a citation, number, date, person or rule.
Turn the saved time into a better week
Bookkeeping: export reports from your accounting system, ask AI to spot unusual changes and build questions for your accountant. Keep the system read-only until a person approves each entry.
Contracts: compare versions, list changed obligations, dates, fees and termination language, then take a focused issue list to counsel. Let AI handle the hunt so paid legal time can focus on judgment.
Healthcare: organize records, chart lab trends, translate unfamiliar terms and prepare questions. Let the clinician diagnose, prescribe and decide what a finding means for you.
Office work: batch the daily drag into one session: inbox triage, meeting preparation, follow-up drafts, status summaries, recurring checklists and first-pass research.
Planning: give the model your real calendar, deadlines and constraints. Ask for a week that protects two blocks of focused work and a firm stopping time. Then have it identify the tasks that can be drafted, grouped, delegated or dropped.
Keep the skepticism. Drop the pose. Judge a controlled AI workflow by whether it is more accurate, faster or more complete than the way the work gets done now. Give it evidence, make it show its work and keep a person accountable. That is how a tool people mock in public quietly gives capable people their evenings back.
|