A policy wording is a machine for changing its own mind.
The insuring agreement grants everything. The definitions quietly shrink it. The exclusions take pieces back. The endorsements put some of them back again. Read it in order and the answer to "is this covered?" flips two or three times before you reach the end — which is precisely why it is the kind of document a language model is genuinely good at, and precisely why the way it fails is so hard to spot.
Guess first — call it before you read the wording
The claim on your desk
Claim 2026-04412 · Homeowner, Section I — Property
A member returns from three weeks away to find roughly 4 cm of standing water across a finished basement. Flooring, drywall and a quantity of stored belongings are written off.
The plumber's report is unambiguous: the supply line to the laundry sink — inside the wall, part of the dwelling's plumbing system — has a split fitting and has been weeping steadily for some time. No storm in the period. No sewer involvement. The member holds a standard homeowner's policy with no additional endorsements recorded on the file.
Call it now, before you see a single clause. Most people in insurance get this wrong in a specific and instructive direction.
What the model actually produces
Extraction is not judgement. It is highlighting.
When people say a language model "reads the policy," this is the useful part of what that means: it labels spans. It finds the sentence that grants cover, the four words that are secretly defined terms, the clause that takes it back, the condition you have to satisfy anyway. Pick a lens and watch the same document light up differently — the highlighting is real work, and it is not the same thing as an answer.
This wording is synthetic — written for this exercise, in the shape of a Canadian homeowner's form, and deliberately shortened. It is not any insurer's product and nothing here is coverage advice. Nothing leaves this browser.
Drive it
Now read it in order, and watch the answer change
Step through the document the way an adjuster does. The verdict at the top is the honest answer given only what you have read so far — which is exactly the answer you get from anyone, human or otherwise, who stops early.
The failure that looks like success
Same question, same model, one missing document
Here is the thing that should worry you, and it is not that the model makes things up. Run the identical question past the identical model, but hand it the policy form without the endorsement schedule attached — the situation you are in every time a document set is incomplete, which in claims is most of the time.
Fluent, structured, well-cited, wrong
"The loss is not covered. While Clause 4.1 lists Escape of Water as an Insured Peril, Exclusion 6.3 removes loss caused by continuous or repeated seepage over a period of weeks or more. The plumber's report establishes a duration consistent with that exclusion. Recommend denial."
- Every clause reference is correct.
- Every quotation is accurate.
- The reasoning is valid.
- The conclusion is wrong.
Nothing was hallucinated. Endorsement E-14 on this member's file replaces the seepage exclusion outright — and the model never saw it, so it never mentioned it.
The same model, told what it is holding
"Coverage turns on Endorsement E-14, which replaces Exclusion 6.3. With E-14 attached, seepage is excluded only beyond 30 days; the plumber's report puts the leak inside that window, so the exclusion does not bite and the loss is covered subject to Condition 8.2 and the deductible."
Same model. Same prompt. The difference is entirely in what was in the folder.
This is the real risk, and it is not the one people prepare for. Teams brace for the model to invent a clause — obvious, embarrassing, easy to catch on review. What actually happens is subtler: the model reasons impeccably over an incomplete document set and produces an answer with no tell. There is no hedge in it, no seam, nothing for a reviewer to catch, because the error is not in the reasoning. It is in the folder. Related: Why AI Hallucinates covers the confident-and-wrong problem; Retrieval vs. Memory covers why what is in the folder is the whole game.
Where this bites
The same shape, elsewhere in the business
Anywhere the answer depends on reading long, cross-referenced, frequently-amended text, you get the same two findings: the extraction is genuinely valuable, and the completeness of the input is the entire control.
Referral triage
Pull the four facts that decide whether a submission needs a human, from twenty pages that mostly don't. High value, low risk — because a wrong referral costs a few minutes, not a coverage position.
Score the failure by what it costs, not by how likely it is.
Wording comparison
Diff this year's form against last year's and surface what moved. Models are unusually good at this and people are unusually bad at it — nobody reads 60 pages twice.
The failure mode is the same: it can only diff the versions you gave it.
First-notice summarisation
Turn a rambling call transcript into a structured file with the dates, the cause, and the named perils. Saves real handling time on every single claim.
Have it cite the line it drew each fact from. An uncitable claim is a made-up one.
The completeness question
Before "how good is the model?" comes "how do we know the folder was complete?" That is a records question, an integration question, and a process question — none of which the model can answer about itself.
It is also the question a regulator will ask first.
Two controls do most of the work here, and neither is about model choice. One: make the system state what it was given and refuse to answer when a referenced document is absent — a model that says "E-14 is referenced but not attached" is worth more than one that is slightly cleverer. Two: keep the human decision where the cost of being wrong is a coverage position, and let the machine do the reading that precedes it. See AI Use-Case Triage for sorting which is which, and Safe AI Usage before any real member data goes near a model.
Go deeper
The best use of this is not deciding claims. It is reading everything first.
The instinct in every insurer is to point this at the decision, because the decision is where the cost is. The value is almost entirely upstream of it: finding the four clauses that matter in sixty pages, surfacing what changed at renewal, telling an adjuster which three documents to read before lunch. That work is enormous, nobody enjoys it, and being wrong about it is cheap and visible. Start where being wrong is cheap, and keep the coverage position with the person whose name is on it.
Talk to us about running this session