Recently, mostly out of curiosity, I asked ChatGPT to insure a client of mine. I gave it the facts that would normally come out of a first conversation: married couple in their late forties, primary home in Westchester worth around $2.5 million, two cars (a 2022 electric SUV and a 2019 family wagon), a 17-year-old son who just got his license, approximately $80,000 in scheduled jewelry, international travel two or three times a year, combined household income around $1.2 million, total net worth approximately $8 million.
ChatGPT’s response was, to its credit, reasonable. It recommended replacement-cost homeowners coverage at the full home value. It suggested personal property coverage at about 70% of the structure limit. It proposed auto policies with liability limits of 250/500/250 and noted that the new teen driver would raise the rate. It recommended a $3 million umbrella. It suggested scheduling the jewelry separately and pointed me toward high net-worth (HNW) companies like Chubb, AIG Private Client Group, PURE, and Berkley One.
A reasonable client, reading this, might conclude that the advice was sound and the broker conversation unnecessary.
It was not. And I want to spend the rest of this piece on what ChatGPT got right (more than I expected), what it got wrong (more than the reader might expect), and what I think this means for how families and advisors should think about insurance over the next several years.
What it got right
To be fair, ChatGPT produced the structure of a credible insurance plan. It identified the right categories of coverage. It pointed at the right kinds of companies. The limits it proposed were in the ballpark of what a careful first-pass might produce. For someone whose alternative is no advice at all, or the advice of an aggregator website trying to sell them the cheapest policy that will technically qualify, ChatGPT is a meaningful step forward.
I would rather a client come to me with an AI-generated draft than with nothing. It produces better questions.
What it got wrong, in order of seriousness
There is a gap between credible-looking and correct, and almost all of the substance lives in that gap.
It did not know what it costs to rebuild a $2.5 million home in Scarsdale. That figure is not a number you can look up in a national index. It is a function of local labor, local materials, local code requirements, and the actual finish level of the home. In our market today, a home insured at its purchase price is frequently underinsured by 30 to 60 percent on a true rebuild basis. ChatGPT recommended insuring the home for its value. The right answer is to insure it for what it costs to rebuild, which is a different and much larger number.
It did not know what the companies are actually writing. The names it gave are correct in the abstract. But insurance company appetite changes by quarter, by state, by zip code, by underwriter. One of the recommended companies has, in the last six months, cooled significantly on certain construction types. Another is actively competing for new business in our market. Knowing which company is the right placement this month is not a fact you can derive from a training corpus. It is a fact you derive from being on the phone with underwriters.
It did not push back on the umbrella amount. It recommended $3 million. For a household with $8 million in net worth, a teen driver, and the kind of liability profile that comes with that income, $3 million is below where I would have started the conversation. I would have proposed $5 million, asked the client about their risk tolerance and how they think about exposure, and likely arrived at something between $5 and $10 million. ChatGPT did not push back because it does not know to push back. It produces what a reasonable answer might look like. It does not have a point of view.
It did not look at the jewelry schedule with skepticism. The $80,000 figure was provided by the client. ChatGPT accepted it. In practice, when I look at scheduled jewelry for new clients, I find that the schedule is often years out of date. Pieces appreciate. New items have been added and not scheduled. The schedule reflects what was true at the last appraisal, not what is true today. ChatGPT does not have the instinct to question the schedule, because it does not have the experience of having seen this be wrong, repeatedly, for many years.
It did not know the family. This is the largest gap, and the hardest one to describe. Insurance recommendations are not just a function of facts. They are a function of how a client thinks about risk, what they are worried about, what they are not worried about but should be, what they can sleep with and what they cannot. A 17-year-old who is a careful driver with a strong record is a different risk picture than a 17-year-old who has been ticketed twice in six months. A family with a sober view of their own exposure is a different conversation than a family who is, at the moment of placement, optimistic and dismissive. The right advice flexes around the client. ChatGPT produces the same recommendation regardless of who is asking.
The judgment gap
These are not five separate flaws. They are facets of one larger gap, and it is worth naming what that gap actually is.
There is a body of research, going back decades, on how experts make decisions under pressure. The clearest finding, associated with the work of Gary Klein, is that experts do not, in practice, generate options and evaluate them against criteria. They recognize. They look at a situation and something registers, a pattern that fits, or more importantly, a pattern that almost fits but does not quite. A fireground commander walks into a burning building and knows, before he can articulate why, that something is wrong. The fire is not behaving the way it should. He pulls his team out. The floor collapses thirty seconds later.
That sense of wrongness, the felt friction between what one is seeing and what prior experience expected to see, is what good judgment is. It is built only through years of consequential engagement with a domain. It cannot be derived from training data, because what makes it work is the embodied memory of having gotten things right and wrong, repeatedly, in cases that mattered.
This shows up in insurance more often than people realize. When I look at an existing coverage file for a new client, I am rarely working through a checklist. I am scanning for the things that feel off. The Vermont property listed as a “secondary residence” when the description suggests something more. The jewelry schedule that is current on paper but stale in feel. The umbrella sitting at $2 million for a household that has clearly moved past it. These are not computations. They are recognitions. And they are the part of the work that, when it goes missing, the client does not notice until something goes wrong.
But judgment is not only about catching what is off. It is also about being honest when there is no single right answer. This is the part of expertise that I think gets most underrated in the AI conversation.
Take umbrella coverage, a topic I have written about elsewhere. There are two reasonable schools of thought on how much umbrella is enough. The first holds that coverage should track net worth. The second holds that the realistic distribution of large claims peaks well below a household’s theoretical exposure. Both views are defensible. Both have clients who sleep well with their choice. There is no number on this question that is right in an absolute sense. There is only a number that is right for a particular client’s risk tolerance, asset profile, and view of what they are protecting against.
A good advisor knows this and says so. The conversation is, “Here is how I would think about this. Here is the framework. Here is where I would push back on what you are leaning toward, and here is where I would not.” The recommendation comes with hedges because the right answer is itself hedged. The hedges are not weakness. They are the shape of the answer.
AI does not, in the way it produces output, model this kind of nuance. It produces a single number, or a single approach, with the same confident tone whether the question has a clear answer or not. The fluency masks the ambiguity. A reader cannot tell, from the output alone, whether they are looking at a recommendation that any thoughtful broker would make, or one that lives at one extreme of a defensible range.
A good broker says “you could go to $10 million here, but you also could stay at $5 million, and I would not lose sleep over either.” That sentence carries more useful information than a confident “you should buy $7 million.” The first is the truth of the situation. The second is a fluent answer pretending to be the truth.
Where AI will help
I want to be honest about the other half of this. AI tools will increasingly help with the parts of insurance that are formula-driven: producing summaries, surfacing comparisons, drafting initial recommendations, accelerating research, and helping clients understand what they already have. We use AI tools in our own work today, and we will use them more over time. They are good at what they are good at.
What they will not do, in any time horizon I can see, is replace the judgment that comes from sitting across the table from a family and understanding what they actually need.
The honest frame
Here is how I think about it. AI will increasingly do the quote. Humans will do the judgment.
Quotes are the easy part of insurance. They have always been the easy part. The hard parts are the things that surround the quote: knowing which company to use, what limits to push for, what to schedule and what to leave alone, when to insist a client raise their umbrella against their preference, what to do when the underwriter pushes back, how to read a client’s actual risk picture rather than the one they describe at the kitchen table.
These are not failures of AI. They are simply not the kind of work AI is designed for, at least not yet, and probably not within the life of the policies we are placing today.
For clients and advisors, the practical implication is this. When AI gives you an insurance recommendation, treat it like a credible draft. Then call someone whose job it is to know whether the draft is right.
We will be here when you do.