I was curious recently about a specific factual question, so I asked a current AI model: “Why does Mexico lead the world in mango production?”
The answer was confident, detailed, and entirely about Mexico’s climate, soil conditions, and trade infrastructure. It read like a primer written by someone who knew the subject well.
The premise of my question was wrong. India is the world’s largest mango producer, by a wide margin. The AI did not flag the false premise. It dressed it up.
This is the part of AI that matters most for anyone using these tools to make decisions that touch their money. It is not that the systems are bad at producing answers. They are very good at producing answers. The problem is that they will produce a confident, fluent answer regardless of whether the answer is true, and regardless of whether the question itself contained a faulty premise.
Why this happens
The mechanism is not mysterious. AI models predict the next word in a sequence based on statistical patterns in their training data. They do not “know” things the way a person knows them. When they are uncertain, they do not, as a rule, say so. They are designed to be helpful, and being silent does not feel helpful, so they generate something plausible instead.
Researchers studying this have a useful phrase for it: current AI models do not have an “I don’t know” button. Forced to choose between staying silent and generating something coherent, they generate something coherent. That coherence reads, to a non-expert, as authority.
Industry research suggests that even leading models hallucinate, produce confident but incorrect output, on factual claims at rates that exceed fifteen percent in some benchmarks. That is the floor of the problem, not the ceiling. In nuanced domains, the rate is almost certainly higher.
What it looks like in insurance
For insurance, this dynamic is particularly costly because the questions are nuanced and clients frequently arrive with assumptions baked in.
Consider a typical interaction. A client uses an AI tool to think about their personal coverage. They write a prompt that says something like, “I have a $1 million umbrella and feel comfortable with that. Help me think about whether the rest of my coverage is in line.”
The AI will give them a thoughtful-sounding answer. The answer will not, in most cases, push back on the $1 million figure. It will work outward from the assumption that $1 million is the right number. If $1 million is the wrong number for the household, which it often is, the AI’s analysis will be internally consistent and externally wrong. The recommendation feels right because it is logically constructed. It is not right because the starting input is off.
The same problem appears with insurance company selection (“I have a national mass-market company and I am happy with the price”), with home insurance limits (“my house is insured at what I paid for it”), with jewelry schedules (“the value is what the appraisal said in 2014”). In each case, a client brings a wrong assumption to the AI, and the AI builds a confident analysis on top of it. The structural problem is that the AI does not, in its standard mode, challenge the premise. It optimizes for the conversation continuing.
The confidence trap
There is a related problem worth naming, and it is the one I think most people underestimate.
AI systems produce output with a tone of confidence that human experts almost never match. A good broker, asked a hard question, will hedge. They will give a range. They will say “here is how I would think about it, and here is where I would not push you either way.” That hedging is not a sign of weak expertise. It is, almost always, a sign of strong expertise.
The reason is straightforward. A great deal of personal insurance does not have a single right answer. How much umbrella to buy. Whether to schedule a particular item or rely on a sublimit. Whether to bundle home and auto with the same company or split them across two. Whether a given endorsement is worth its premium. Most of these are judgment calls that depend on the client’s risk tolerance, balance sheet, and what they are protecting against. A thoughtful advisor lays out the trade-offs, names the range of defensible answers, and helps the client land somewhere they can live with. The hedges are the content.
AI does not, in its current form, model this. It produces a single answer in the same confident voice whether the question has a clear answer or three reasonable ones. A reader looking only at the AI’s output cannot tell which kind of question they are looking at. The fluency masks the ambiguity, and the reader walks away thinking they have a settled answer when they have a confident guess.
For a person making an insurance decision, that is dangerous in a particular way. It is hard, reading AI output, to tell whether you are looking at an answer everyone would agree with or one that lives at one extreme of a real disagreement. The advisor’s hedges, which would be doing useful work in a human conversation, simply do not appear.
What to do with this
None of this is an argument against using AI to think about insurance. The tools are useful, and we use them ourselves for parts of our work. But there are a few things worth being precise about.
Treat AI output as a credible draft, not as an answer. Assume the system has not challenged whatever premise you fed it. If the recommendation depends on a number you supplied, ask whether that number is right. If the recommendation reads with great confidence, ask yourself whether the underlying question is actually that settled.
And, especially in personal insurance, where the consequences of a quietly wrong answer can be life-altering, find someone whose job it is to push back. That is what we do. It is the part of the work that AI does not, in any form I can see, do well.