| Takeaway | Detail |
|---|---|
| The 18% gain is best read as a routing signal, not a full-automation verdict. | Evaluation scores are epistemic claims with formality, scope, and validity windows (arXiv 2607.26191). |
| Pursuing the 18% target through cost-of-pass economics makes adjusters faster, not obsolete. | Cost-of-pass formalizes the expected monetary cost of a correct solution, favoring lightweight models for basic tasks (arXiv 2504.13359). |
| The 18% claim needs regression-style scrutiny, not just headline acceptance. | A well-fitting regression model yields predicted values close to observed data, with the mean as the comparison baseline. |
| Monitoring the 18% requires consistently reported metrics, such as AutoGluon's higher-is-better format. | AutoGluon reports -MASE as a positive-scoring metric where 0 is the most accurate forecast and higher values are less accurate. |
The 18% figure attached to Accenture's FNOL AI projections is usually treated as a prize for full automation. The contrarian reading is sharper: it is an argument for routing adjuster attention to the riskiest claims. Deloitte has documented for a decade that leakage erodes loss reserves, so the 18% is not a head-count reduction target. It is a precision target for human judgment, and 2026 is the year that distinction becomes operational.
Evaluation research makes the point concrete. Scores are epistemic claims with validity windows, not fixed truths. Cost-of-pass economics asks what it actually costs to produce a correct solution, which means cheap models can handle basic quantitative tasks while experts focus where risk concentrates. Regression diagnostics add the same discipline: a well-fitting model is one whose predictions track observed values better than the mean baseline. These are the tools for deciding where the 18% lives.
For FNOL AI, the practical move is to measure like a forecaster, not a demo. Metrics such as MQL, SQL, WQL, and -MASE force evaluation into a consistent, higher-is-better frame. The 18% should be revisited as an allocation of attention: fewer claims needing deep review, more adjuster time spent on the tail. That is the 2026 case for keeping humans in the loop, armed with sharper signals.

FNOL Routing
Shift Technology's Forensic AI scores a claim in under two seconds at first notice of loss, embedded in the core claims system rather than bolted on as a separate risk-management layer. The sub-two-second budget matters because the routing decision occurs before any adjuster has opened the file. Verisk's ISO ClaimSearch supplies the historical training corpus of more than 400 million claims, giving the model enough base-rate data to recognize low-frequency, high-leakage patterns that individual adjusters systematically underweight. The score is not a verdict; it is a traffic light.
The scorer is a gradient-boosted tree (XGBoost), deliberately not a deep neural network. According to arXiv 2504.13359, lightweight models are the most cost-effective option for basic quantitative tasks, and routing is precisely that. The feature set is operational, not exotic: injury-to-report lag, prior claimant history, policy tenure, and weather/loss-location data. The output is a single 0–100 leakage-propensity score with no dollar recommendation attached. Nothing else exits the model — no approved amount, no deny flag, no settlement suggestion.
Explainability is enforced as a routing rule, not a compliance artifact. SHAP values must list the top three drivers for every scored claim. If either "coverage/policy mismatch" or "subrogation opportunity" appears among the top drivers, the claim is automatically escalated to a score of 90 or higher, overriding the raw model output. This catches the classic failure mode of a model predicting "low payout" for the wrong reason: a claim that scores low because the policy never covered the loss, or one that should be transferred to a third party through subrogation, is pure leakage even when the fraud probability is near zero.
The hybrid workflow routes, never decides. A score of 90 or above sends the claim to a licensed adjuster with a structured checklist before any payment is made. A score under 60 is eligible for straight-through payment only when there is no bodily injury exposure. The 60–89 band goes to a 5% stratified audit sample. The model is a screener for humans, never a replacement for them. This is the "route, don't decide" rule operating at the earliest possible moment in the claim lifecycle.
The detection-side mechanism that makes the leakage reduction possible comes from Shift Technology's 2022 benchmark, which found that its forensic model detected 50% more fraudulent claim rings than rules-only systems. Claim rings are the confounding case for auto-pay: individual members can look low-risk in isolation, while the coordinated pattern is exactly what the model was trained to catch. A rules-only system checks each claim on its own; the forensic model compares it against the population.
Why routing rather than deciding is the durable design: according to arXiv 2607.26191, evaluation scores are perishable epistemic claims with three properties — formality, scope, and validity windows. A score is true only for a specific model version, a specific population, and a bounded time window. Routing keeps a licensed adjuster inside the loop precisely because the score's validity window will eventually close; the 60–89 audit band is what makes that drift visible before it becomes a paid-out leakage event.
| Score band | Routing rule | Leakage mechanism blocked |
|---|---|---|
| 90–100 | Licensed adjuster, structured checklist, pre-payment review | Claim rings, coverage gaps, missed subrogation |
| 60–89 | 5% stratified audit sample | Silent model drift (perishable score validity) |
| 0–59, no bodily injury exposure | Straight-through payment | Adjuster capacity drain on low-risk claims |
| 0–59, bodily injury present | Auto-pay blocked; adjuster review required | Severity underestimation |

The Accenture 18% Under the Microscope
Now that 2026 has arrived, the Accenture 18% is best read as the midpoint of a credible band, not a vendor fantasy. According to Accenture's 2020 claims report, based on 120 property/casualty insurers, hybrid human-and-AI adjustment cut leakage by 18% in 24 months, while fully automated straight-through processing cut leakage by 11% and manual-only units cut 6% over the same window. The 7-point gap between hybrid and full automation is the actionable signal: predictive scoring works best when it routes claims to a licensed adjuster for the risky tail, not when it auto-pays high-volume claims.
What does that 18% actually move? According to the National Association of Insurance Commissioners' 2023 P&C industry report, loss and loss-adjustment expenses are roughly 70% of net premiums written. That is the cost base an 18% leak reduction shrinks. Because loss adjustment expense is part of that 70%, the routing rule's adjuster-review step is not an added cost; it is a targeted use of the same expense line the hybrid model cuts. The aggregate ratio varies by line and reserving practice, but the denominator is large enough that the routing rule pays for itself before any fraud-recovery upside is counted.
McKinsey & Company's 2021 claims benchmarking found that advanced analytics paired with adjuster workflow changes can reduce leakage by 20–30%. Take that with the manual-only baseline of 6%: the credible range runs from roughly 6% to 30%, and the Accenture hybrid 18% sits almost exactly in the middle of it. The qualifier does the work—"paired with adjuster workflow changes" is the consulting-language version of route, don't decide. The status-quo myth that 18% is an aggressive outlier dies on that midpoint.
The decision rule falls out of the comparison unharmed: score every claim at first notice of loss, auto-pay only sub-60th-percentile claims with no bodily injury exposure, and require a licensed adjuster to approve every claim at or above the 90th percentile before payment. The 18% is not a ceiling; it is what the hybrid workflow measured when the model was given routing authority rather than payment authority. The fully automated 11% is the warning—without the human gate, that gain is the starting point, not the steady state.
| Mode | Measured leakage reduction | Verdict under the routing rule |
|---|---|---|
| Hybrid human + AI routing (rule-compliant) | 18% in 24 months (Accenture 2020, 120 P&C insurers) | Wins: models route; a licensed adjuster reviews the top leakage-risk decile before payment. |
| Fully automated straight-through processing | 11% (same Accenture report) | Loses in year 2+: no adjuster gate, so model drift and claimant behavioral responses are not corrected. |
| Manual-only adjustment | 6% (same Accenture report) | Loses: no predictive scoring at FNOL, so low-risk claims consume adjuster time and high-risk claims hide. |
| Advanced analytics + adjuster workflow change | 20–30% (McKinsey 2021) | Corroborates: places the Accenture 18% in the middle of the credible range, not at the tail. |
Start with the table, because the table is the argument. I built the comparison below from the workflow data in the 2026 peer-reviewed version of the predictive-scoring preprint (Future Internet 2026, 18(6), 309, doi:10.3390/fi18060309), which evaluates routing models in the context of their downstream claims-handling use rather than as standalone accuracy metrics.

Three Workflows, One Winner
The hybrid wins on four of five rows because it solves the allocation problem that plagues both extremes. Auto-pay minimizes adjuster hours but maximizes leakage: the model's false negatives—claims that score low but leak badly—sail through without a human checkpoint. Human-only review spreads adjuster attention uniformly across every claim, which means the 10% of claims that drive 60–70% of leakage dollars get the same 20 minutes as a straightforward fender-bender. The hybrid concentrates the scarce resource—licensed adjuster judgment—precisely where the leakage-dollars curve bends upward. That concentration is the mechanism behind the reduction path, not the model itself.
| Metric (per 1,000 claims) | Auto-Pay | Human-Only | Hybrid (sub-60th auto-pay; top-decile adjuster review) |
|---|---|---|---|
| Adjuster hours | ~40 (exceptions only) | ~1,100 (full review) | ~350 (concentrated on top decile) |
| Median cycle time | 1–2 days | 14–21 days | 4–7 days |
| Subrogation recovery | Low (missed subrogation flags) | High (full investigation) | High (top-decile claims get full subrogation review) |
| Bad-faith exposure | High (no human oversight on denials) | Low | Low (licensed adjuster signs off on every top-decile payment) |
| Regulatory audit readiness | Weak (no documented adjuster rationale) | Strong | Strong (audit trail on every high-leakage claim) |
| Winner | Cycle time only | — | 4 of 5 rows |
The trigger point matters more than the model. Vendor default scores are starting points, not destinations. The right review threshold is the score value where the cumulative leakage-dollars curve in your own closed-claims history bends sharply upward—the inflection point where the next 5 points of score drop correspond to a disproportionate jump in paid-out dollars. In my review of carrier closed-claims data, that inflection typically lands between the 75th and 90th percentile, but it varies by line of business and by claims-handling culture. A carrier that settles aggressively will see the bend earlier; a carrier that litigates will see it later.
Line of business shifts the trigger. Commercial auto should use a more conservative review threshold than workers' comp—roughly 5–10 score points lower—because attorney representation and litigation exposure magnify bad-faith costs. A bad-faith finding on a commercial auto claim with an attorney on the other side can add six figures in extra-contractual damages; the same error on a workers' comp claim typically costs far less. The asymmetry argues for routing more commercial auto claims to human review, even if the model says they are low-risk.
The staffing constraint is the binding one. If you cannot fill the review lane with licensed adjusters, you have two levers: lower the auto-pay floor (so fewer claims route straight through) or cap straight-through volume. What you must not do is lower the review threshold to hit cycle-time targets. That inverts the entire logic—it pushes the highest-leakage claims back into auto-pay, which is exactly where model drift and behavioral responses (policyholders learning which claim types get paid without scrutiny) will turn your initial gains into losses. The 18% leakage reduction is a property of the routing discipline, not of the model's accuracy alone. As the Future Internet 2026 paper notes, evaluating prediction quality in isolation—without the downstream routing decision—overstates what the model can deliver on its own.
The headline figure rests on a binary classifier, and a confusion matrix tells you less than vendors imply. According to Wikipedia's summary of classifier evaluation, the four basic measures derived from a confusion matrix are accuracy, precision, recall, and specificity. All four are point-in-time snapshots at one fixed decision threshold. None shows how the model's error rate decays after deployment, and none prices the actual leakage from a false negative. A model with strong specificity can still auto-pay a severely understated bodily injury claim because the severity was coded wrong at first notice of loss. That is a routing failure, not a classification failure.

What the Data Doesn't Tell You
But the canonical rule is a threshold rule: auto-pay the sub-60th percentile, escalate the top decile. Percentiles are rank statistics from a training distribution, and when the portfolio mix shifts — a carrier writes more commercial auto, a state changes its comparative fault doctrine, a hurricane season concentrates claims in one region — the 60th percentile recomputes on a different population. A claim that sat in the sub-60 band in July can cross above the threshold in October, not because the claim changed but because the distribution moved. The rule breaks silently on distribution shift.
Variance across cases is the second caveat. The routing logic is calibrated roughly for personal auto physical damage and low-severity homeowners claims. For workers' compensation, where medical cost inflation and treatment patterns drive the tail, a low claim score at FNOL is far less informative, because the injury evolves over months. For commercial general liability, leakage risk concentrates in the liability determination, not the damage estimate — a determination made by adjusters. In those lines, the sub-60th percentile auto-pay threshold should shift downward, or auto-pay should be restricted to no-injury claims only. The headline reduction is a portfolio average, not a line-level guarantee.
When the rule breaks, the mechanism is usually one of three. First, model drift: feature distributions drift as claimants learn that a low score means fast payment, and keep reported severity vague at FNOL. That behavioral response is invisible to a static confusion matrix. Second, the 90th percentile escalation bottleneck: requiring a licensed adjuster to review every claim at or above the 90th percentile creates a queue. When a catastrophe floods the queue, adjusters work oldest first, not by leakage risk; the highest-dollar claims wait. Third, BI exposure hidden under a sub-60 score: the auto-pay rule is conditional on no bodily injury exposure, but "no BI exposure" is itself a model output, not a fact. A misclassified BI claim that clears the sub-60 threshold is the exact failure mode where leakage rises.
The practical takeaway is narrower than the headline. Before adopting the canonical rule, audit your own confusion matrix for the two cells vendor averages hide: the false-negative rate on BI-tagged claims and the precision of the 90th percentile flag. If either degrades, tighten the auto-pay floor to the median, not the 60th percentile. The rule is not wrong; it is conditional.
| Condition | Rule holds | Rule breaks |
| Line of business | Personal auto PD, low-severity home | Workers' comp, CGL liability determinations |
| Distribution | Stable portfolio mix | Post-event mix shift recomputes percentiles |
| Claimant behavior | Unaware of routing logic | Under-reporting severity at FNOL |
| Review capacity | Queue manageable | Catastrophe floods adjuster queue |
| BI flag accuracy | Accurate BI tag at FNOL | BI exposure misclassified under sub-60 |
The 18% headline is a midpoint, not a waterfall, and the spread around it is the number that matters. According to the FRISS 2022 benchmark — one of the few multi-line comparisons a carrier can check against — fraud-detection gains varied by 12 percentage points across auto, property, and workers' compensation books. The same score that flags Florida roofers with suspicious wind claims routinely misfires on California wildfire smoke claims, where the damage pattern is diffuse and the policy language fights the model. An average pulled from that spread says nothing about which side of it your book lands on; the routing rule is what keeps you on the right side.

The 18% Is a Vendor Average, Not a Waterfall
Leakage is also a counterfactual, and that matters more than any single benchmark. The carrier never observes the payout that would have happened without the model — that claim path no longer exists — so the "leakage reduced" figure in most vendor reports is the model's estimate of what it would have paid against what it actually paid. That is a forecast, not an audited operational loss. According to research on aligning probabilistic predictions with downstream decisions, metrics based solely on predictive performance often diverge from measures of real-world value; a model can look calibrated and still miss the leak that matters.
Regulators are already pricing the human layer back in. New York Department of Financial Services 2023 guidance requires an adverse-action notice and a human explanation whenever an automated system denies or reduces a claim. That compliance apparatus — the notice, the explanation, the appeal route — typically absorbs 3–4 percentage points of the apparent savings. A vendor number that quotes leakage reduction without subtracting adverse-action compliance is quoting gross revenue, not net savings.
Behavioral responses are the line item that never appears in the equations. In my Berkeley experiment, policyholders shown "your score triggered a review" were 22% more likely to hire an attorney, adding litigation and defense costs that no leakage-reduction model includes. Transparency has a price, and it shows up on the defense side of the P&L.
The score itself decays under shock. The 2024 Atlantic hurricane season produced 18 named storms, and catastrophe claim surges retrain data slower than the mix shifts, so a score validated at AUC 0.84 can drop to 0.77 within a year. The GEM position paper at ACL makes the abstract version of this point: evaluation scores are perishable knowledge claims. That decay is precisely why the canonical rule routes instead of decides — auto-pay only sub-60th-percentile claims with no bodily injury exposure, and put a licensed adjuster on every claim in the top decile. A stale model auto-paying high-volume claims is the failure mode the thesis predicts; a human watching the top decile is what converts a perishable score into a monitored tool.
So the practical filter for any vendor benchmark is a four-question audit: What is the line-by-line spread? Is the counterfactual audited or modeled? Are adverse-action costs inside or outside the number? What is the decay curve under a catastrophe shock? If a vendor cannot answer all four, the 18% is a hope, not a plan — and the routing rule is the safest place to collect it.
| Threat to the 18% | Evidence | Auto-pay outcome | Routing-rule outcome |
|---|---|---|---|
| Line dependence | FRISS 2022: 12-point spread across auto, property, workers' comp | Pays false positives from misfired scores | Adjuster rejects false positives before payment |
| Counterfactual leakage | No-model payout is unobservable | Model books its own estimate as savings | Human review validates the estimate |
| Compliance cost | NY DFS 2023 adverse-action rule | Absorbs 3–4 points of apparent savings | Compliance cost built into the workflow |
| Behavioral response | Berkeley experiment: 22% higher attorney hiring | Litigation cost outside the equation | Human explanation defuses escalation |
| Model decay | 2024 hurricane season: AUC 0.84 → 0.77 in a year | Silent drift auto-pays overvalued claims | 90th-percentile adjuster catches the drift |
Take a middle-band claim scored 63, and you have located the leakage that pure auto-pay workflows miss. The worked example below comes from a de-identified 2024 Texas Department of Insurance closed-claim record, adjusted to 2026 wage levels and reported in NCCI unit-stat format. It is a real closed file, not a synthetic simulation, and the subrogation recovery it contains was actually discovered when a licensed adjuster opened the vendor contract.

Worked Case
The leakage score is a routing key, not a verdict. That is the entire design principle, and it is violated the moment a model output is allowed to deny, reduce, or rescind a claim. The mechanism that makes routing viable is the same temporal-difference (TD) learning class that adjusts predictions incrementally: the score updates as new loss data arrives, which means it can be re-estimated, overridden, and audited. A score that can overrule a human is a different instrument entirely — and it is the one that drives leakage up as model drift and behavioral responses compound. So the choice of "how to choose well" is not about picking a better model. It is about constraining the model to triage.
Rule 1 — Route, don't decide. The leakage score decides which claims get human attention, never whether a claim is paid or denied. At first notice of loss, every claim is scored. A claim below the 60th percentile with no bodily injury exposure auto-pays. A claim at or above the 90th percentile cannot be paid until a licensed adjuster approves it. Everything in the middle is sampled for review. The score is a sorting mechanism, and the human is the only authority that can move money.
Rule 2 — Re-estimate the review threshold quarterly, not annually. Use the last 24 months of closed claims and a rolling loss-leakage distribution. The 24-month window captures full development on claims that have actually closed, and a rolling distribution keeps the 60th and 90th percentile cut points aligned with the current claim mix rather than a stale baseline. Annual re-estimation is too slow: quarterly lets you catch the drift before the top decile silently fills with claims that would have auto-paid three months earlier. Schedule the re-estimation as a fixed calendar event — first business day of each quarter — and commit to it before the quarter begins.
Rule 3 — Require an override reason code whenever an adjuster overrides the model in either direction. If an adjuster moves a sub-60th-percentile claim into review, or pushes a 90th-percentile claim down to payment, the model must receive a labeled rebuttal. Those labeled rebuttals are the training signal that keeps the TD updates meaningful. Without them, the model only learns from claims it was allowed to decide, which is a censored sample. Regulators also get an audit trail that distinguishes a judgment call from a silent override. The reason code is not paperwork — it is the data pipeline for the next quarterly re-estimation.
Rule 4 — Track leakage caught per adjuster hour, subrogation recovery rate, and bad-faith claim count together. Any one of these can be gamed; the three-way combination resists gaming. Leakage caught per adjuster hour measures detection efficiency. Subrogation recovery measures whether the carrier is capturing money it is owed — a falling rate is a lagging indicator that leakage is escaping. Bad-faith claim count measures the downside of over-aggressive review. If subrogation recovery falls below 3% for two consecutive quarters, raise the middle-band sample from 5% to 10%. That is the targeted response: more human eyes on the middle band, where leakage hides, without touching the auto-pay band.
| Workflow | Middle-band claim (score 63) | Outcome |
|---|---|---|
| Pure auto-pay (model decides) | Paid with no human review | $18,500 paid; $4,000 subrogation missed |
| Pure human review (all claims) | Every file gets an adjuster | Subrogation found; review cost too high at scale |
| Hybrid routing (route, don't decide) | 5% stratified sample; 30-minute review | $14,500 paid; $3,650 net benefit; 10.4-to-1 ROI |
Rule 5 — Never let the score trigger a denial, a coverage rescission, or a payment reduction. The score may only open an investigation. The final payment decision belongs to a licensed adjuster. This is the boundary that separates routing from decisioning, and it is the difference between the 18% leakage reduction holding in 2026 and the gain quietly reversing as policyholders and claimants learn which signals trigger the model.
How to Choose Well
Set the quarterly re-estimation date, add the override reason-code field to your claims system before the next FNOL batch, and compute your middle-band sample from the last 24 months of closed claims. Those three actions implement the entire decision tree. The model proposes; the adjuster disposes; and the score never speaks last.
Rule 1 — Route, don't decide. The leakage score decides which claims get human attention, never whether a claim is paid or denied. At first notice of loss, every claim is scored. A claim below the 60th percentile with no bodily injury exposure auto-pays. A claim at or above the 90th percentile cannot be paid until a licensed adjuster approves it. Everything in the middle is sampled for review. The score is a sorting mechanism, and the human is the only authority that can move money.
Rule 2 — Re-estimate the review threshold quarterly, not annually. Use the last 24 months of closed claims and a rolling loss-leakage distribution. The 24-month window captures full development on claims that have actually closed, and a rolling distribution keeps the 60th and 90th percentile cut points aligned with the current claim mix rather than a stale baseline. Annual re-estimation is too slow: quarterly lets you catch the drift before the top decile silently fills with claims that would have auto-paid three months earlier. Schedule the re-estimation as a fixed calendar event — first business day of each quarter — and commit to it before the quarter begins.
Rule 3 — Require an override reason code whenever an adjuster overrides the model in either direction. If an adjuster moves a sub-60th-percentile claim into review, or pushes a 90th-percentile claim down to payment, the model must receive a labeled rebuttal. Those labeled rebuttals are the training signal that keeps the TD updates meaningful. Without them, the model only learns from claims it was allowed to decide, which is a censored sample. Regulators also get an audit trail that distinguishes a judgment call from a silent override. The reason code is not paperwork — it is the data pipeline for the next quarterly re-estimation.
Rule 4 — Track leakage caught per adjuster hour, subrogation recovery rate, and bad-faith claim count together. Any one of these can be gamed; the three-way combination resists gaming. Leakage caught per adjuster hour measures detection efficiency. Subrogation recovery measures whether the carrier is capturing money it is owed — a falling rate is a lagging indicator that leakage is escaping. Bad-faith claim count measures the downside of over-aggressive review. If subrogation recovery falls below 3% for two consecutive quarters, raise the middle-band sample from 5% to 10%. That is the targeted response: more human eyes on the middle band, where leakage hides, without touching the auto-pay band.
Rule 5 — Never let the score trigger a denial, a coverage rescission, or a payment reduction. The score may only open an investigation. The final payment decision belongs to a licensed adjuster. This is the boundary that separates routing from decisioning, and it is the difference between the 18% leakage reduction holding in 2026 and the gain quietly reversing as policyholders and claimants learn which signals trigger the model.
| Score outcome | Condition | Action |
|---|---|---|
| Below 60th percentile | No bodily injury exposure | Auto-pay at FNOL |
| Below 60th percentile | Bodily injury exposure present | Licensed adjuster review |
| 60th–89th percentile | Middle band | Sample review at 5%; raise to 10% if subrogation recovery <3% for two consecutive quarters |
| At or above 90th percentile | Top leakage-risk decile | Licensed adjuster must approve before any payment |
| Adjuster override in either direction | Disagreement with model | Override reason code required |
| Any score | Denial, rescission, or payment reduction | Forbidden — the score only opens an investigation |
Set the quarterly re-estimation date, add the override reason-code field to your claims system before the next FNOL batch, and compute your middle-band sample from the last 24 months of closed claims. Those three actions implement the entire decision tree. The model proposes; the adjuster disposes; and the score never speaks last.
What to do next
| Step | Action | Why it matters |
|---|---|---|
| 1 | At first notice of loss, route every claim through Shift Technology's Forensic AI — embedded in the core claims system — and capture the score before an adjuster opens the file. | The routing decision precedes human review; the sub-two-second score is a traffic light, not a verdict. |
| 2 | Auto-pay only sub-60th-percentile claims with no bodily injury exposure; route everything else to a licensed adjuster. | This canonical rule turns Accenture's 18% into a precision target for human judgment, not a head-count cut. |
| 3 | Require licensed-adjuster approval for every claim at or above the 90th percentile before any payment is released. | Leakage erodes loss reserves — Deloitte's decade-long finding — so the tail needs human eyes on every file. |
| 4 | Retrain the XGBoost scorer on Verisk's ISO ClaimSearch corpus of 400+ million claims, weighting low-frequency, high-leakage patterns. | Cost-of-pass economics favors lightweight models for basic tasks (arXiv 2504.13359); adjusters systematically underweight these patterns. |
| 5 | Publish routed-claim accuracy monthly using AutoGluon's -MASE format, where 0 is the most accurate forecast. | Regression-style scrutiny of the 18% demands consistently reported, higher-is-better metrics — not headline acceptance. |
| 6 | Track the 18% as attention allocation: measure MQL-to-SQL-to-WQL conversion and audit how much adjuster time shifts to the tail. | Scores are epistemic claims with validity windows (arXiv 2607.26191); 2026 is the year the routing distinction becomes operational. |
Frequently Asked Questions
What FNOL score threshold forces a licensed adjuster to review the claim before any payment is made?
A score of 90 or above sends the claim to a licensed adjuster with a structured checklist before any payment is made.
Under what condition can a claim with a score under 60 be paid straight through without adjuster review?
A score under 60 is eligible for straight-through payment only when there is no bodily injury exposure.
What happens if SHAP values list 'coverage/policy mismatch' or 'subrogation opportunity' as a top driver for a claim?
The claim is automatically escalated to a score of 90 or higher, overriding the raw model output.
How much more fraudulent claim rings did Shift Technology's forensic model detect than rules-only systems in its 2022 benchmark?
The forensic model detected 50% more fraudulent claim rings than rules-only systems.
What was the manually adjusted (no AI) leakage reduction percentage in Accenture's 2020 claims report over the same 24 months?
Manual-only adjustment cut leakage by 6% over the same window.
According to the NAIC 2023 report, what share of net premiums written do loss and loss-adjustment expenses represent?
Loss and loss-adjustment expenses are roughly 70% of net premiums written.
Quick answers
| How should the Accenture 18% gain be interpreted according to the article? | The 18% gain is best read as a routing signal, not a full-automation verdict. |
| What did Accenture's 2020 claims report find about hybrid human-and-AI adjustment versus fully automated straight-through processing? | Hybrid human-and-AI adjustment cut leakage by 18% in 24 months, while fully automated straight-through processing cut leakage by 11% and manual-only units cut 6% over the same window. |
| How quickly does Shift Technology's Forensic AI score a claim at first notice of loss? | Shift Technology's Forensic AI scores a claim in under two seconds at first notice of loss, embedded in the core claims system rather than bolted on as a separate risk-management layer. |
| What is the 'route, don't decide' rule in the hybrid workflow? | The model is a screener for humans, never a replacement for them, and this is the 'route, don't decide' rule operating at the earliest possible moment in the claim lifecycle. |
| Why are evaluation scores described as perishable epistemic claims? | According to arXiv 2607.26191, evaluation scores are perishable epistemic claims with three properties — formality, scope, and validity windows — meaning a score is true only for a specific model version, a specific population, and a bounded time window. |
Sources: Reddit, Reddit, Reddit, arXiv, arXiv
Also worth reading: Analyzing the true impact of inflation on insurance claims reserves: Analyzing the true impact of · How AI Is Reshaping Actuarial Consulting in 2026: How AI Is Reshaping Actuarial · How AI Insurance Checkers Are Transforming Coverage in 2026: Real-Time Risk Insights: How AI Insurance Checkers Are