Insurance Coverage and Bad Faith
From Colossus to ChatGPT: Artificial Intelligence in Modern Claim Handling — Efficiency, Explainability, and Exposure
Defending the Algorithm™
This article is part of the “Defending the Algorithm™” series and was written by Pittsburgh, Pennsylvania Insurance, Business and IP Trial Lawyer Christopher M. Jacobs, Esq. and was authored with research assistance from OpenAI’s GPT-5. The series explores the evolving intersections of artificial intelligence, intellectual property, and the law. GPT-5 may generate errors but the author has verified all facts and analysis for accuracy and completeness.
I. Introduction and Historical Context
The integration of artificial intelligence into the insurance industry has not occurred in a single leap but through a series of incremental innovations—each testing the boundaries between efficiency and fairness, automation and accountability. While today’s conversation focuses on large language models, image-recognition tools, and predictive analytics, the roots of algorithmic claim handling trace back nearly three decades to the introduction of Colossus, a software program once hailed as revolutionary in standardizing bodily-injury claim valuations. The trajectory from Colossus to contemporary AI systems offers not merely a technological evolution, but a jurisprudential one: both eras have forced courts, regulators, and practitioners to confront how far insurers may rely on opaque systems to evaluate inherently human losses.
A. The Promise of Automation in Claims
At its inception in the 1990s, Colossus was marketed as a neutral, data-driven means of ensuring consistency across bodily-injury settlements. Adjusters were instructed to input medical details, treatment types, and jurisdictional information, and the program would generate a recommended settlement range based on prior claims data. Insurers adopting the system saw it as an antidote to subjectivity—an algorithmic equalizer that could reduce variance, curb inflated demands, and streamline claim resolution. In a business driven by volume and loss ratios, such efficiency was irresistible.
Yet, the same characteristics that made Colossus efficient also made it inscrutable. The proprietary “value drivers” that produced its outputs—reportedly numbering in the thousands—were hidden behind trade-secret protections. Adjusters could see the result, but not the reasoning. Regulators and claim professionals later observed that a system intended to remove subjectivity in valuation often introduced new, systemic biases of its own.[1]
B. The Colossus Litigation and Judicial Response
By the early 2000s, Colossus had become the focal point of consumer litigation and regulatory scrutiny. Plaintiffs alleged that insurers and the software’s developer, Computer Sciences Corporation (CSC), had used Colossus to systematically undervalue bodily-injury claims by enforcing uniform settlement ranges and incentivizing adjusters to conform to them. The central contention was not that Colossus malfunctioned in a technical sense, but that its design embodied a “one-size-fits-all” approach to inherently individualized losses. By converting subjective, claimant-specific evaluations into standardized algorithmic outputs, the software replaced human discretion with what might be called artificial standardization. That process risked disregarding unique claimant characteristics—such as occupation, pre-existing conditions, or quality-of-life impacts—that traditionally informed the valuation of bodily-injury damages.
The most notable of these actions culminated in a national class-action settlement in Hensley v. Computer Sciences Corporation, filed in Arkansas and approved in 2005. The plaintiffs contended that CSC and multiple insurers had concealed the software’s use and manipulated its calibration to achieve predetermined reductions in claims payouts. While CSC denied wrongdoing, the settlement required it to modify its marketing practices, clarify the system’s intended use, and make aspects of its functionality more transparent to insurers and regulators. Subsequent market-conduct examinations in several states, including California and Michigan, reached similar conclusions: insurers could not require adjusters to adhere rigidly to Colossus outputs, nor could they compensate personnel based on compliance with those valuations.
Although the resulting settlements and administrative orders did not create binding precedent, they altered the industry’s expectations for algorithmic decision-making. The Colossus controversy demonstrated that automation does not absolve insurers of the duty of fair claim handling; it merely reframes it. The software’s opacity generated the same evidentiary and ethical challenges now resurfacing with modern AI: explainability, bias, and the tension between internal models and external accountability.
C. Lessons from the First Generation of Algorithmic Claims
Three lessons from the Colossus era remain particularly instructive for today’s practitioners.
First, transparency matters. Regulators and courts grew skeptical not merely because Colossus existed, but because it operated as a hidden arbiter of value. The lack of disclosure—both to claimants and, in some instances, to line adjusters—transformed a management tool into a litigation risk.
Second, automation amplifies institutional intent. A valuation system trained or tuned to achieve efficiency gains can easily be repurposed to accomplish cost-containment objectives inconsistent with fair-claims standards. Plaintiffs in Hensley alleged that Colossus was calibrated to produce systematic reductions in claim payouts, reportedly targeting decreases of up to fifteen percent. Whether or not that allegation could be empirically proven, the perception alone was enough to erode confidence in algorithmic fairness—a cautionary lesson for modern machine-learning tools trained on historical data that may embed similar undervaluation biases.
Third, human oversight is indispensable. Following the Colossus settlements, several state insurance departments emphasized that adjusters must retain independent judgment and document reasons for either deviating from or adopting software recommendations. That principle—human-in-the-loop accountability—has since become the ethical baseline for any deployment of AI in claims handling.
D. Continuity Between Colossus and Contemporary AI
Seen through this historical lens, today’s debates over generative AI and machine-learning systems are less revolutionary than cyclical. Modern claim-analysis models, particularly those that evaluate images of property damage or generate text-based coverage explanations, echo the same dynamics of efficiency versus discretion that defined Colossus. The principal difference lies in scale and autonomy. Where Colossus applied deterministic rules to structured data, modern AI operates on probabilistic inference from vast, unstructured sources—millions of data points, often processed without direct human review. The opacity has deepened, and with it, the potential for both error and abuse.
In litigation, this evolution raises familiar questions under new guises. If an insurer relies on an AI-generated damage estimate that proves inaccurate, has it acted unreasonably? If an AI model produces inconsistent results across regions or demographics, does that constitute unfair discrimination under state law? These inquiries trace their lineage directly to the Colossus disputes, but they now extend into far more complex evidentiary terrain. Discovery once aimed at Colossus calibration settings will soon target training data, neural-network architectures, and algorithmic weighting—issues few courts are yet equipped to handle.
E. Framing the Discussion Ahead
The following sections will build upon these historical parallels. Section II will examine how insurers currently deploy AI in claim handling, including both legitimate efficiency tools and the emerging risk of AI-generated fraudulent claims submitted by insureds themselves. Section III will turn to the developing legal landscape, analyzing how courts have begun to interpret bad-faith standards in the context of algorithmic decision-making and what ethical obligations arise from reliance on non-transparent systems. Ultimately, the goal is not to condemn automation, but to situate it within a continuum of evolving duties—duties that require insurers to harness innovation responsibly, preserve human judgment, and maintain the public trust that underpins the entire enterprise of insurance.
II. Modern AI in Claim Handling
A. From Colossus to Cognitive Claims — The Evolution of Automation
The impulse that gave rise to Colossus—the desire for faster, more consistent, and ostensibly “objective” claim evaluations—remains the same animating force behind today’s adoption of artificial intelligence in insurance claim handling. What has changed is the data universe. Where Colossus relied on structured medical codes, treatment descriptions, and adjuster input, modern AI systems ingest unstructured and multimodal data: photos of property damage, satellite imagery, adjuster notes, invoices, correspondence, and even social media content.
In essence, the industry has moved from rule-based automation to learning-based cognition. Yet the governing question is still familiar: how much of the inherently human process of claims evaluation can be delegated to machines without undermining the insurer’s duty of good faith?
Many of today’s AI tools are, in a sense, Colossus’s descendants. They pursue the same efficiencies—standardization, cost control, and cycle-time reduction—but do so with greater technical sophistication and, correspondingly, greater opacity. The tension between innovation and accountability, first articulated in the Colossus litigation, now resurfaces in the age of generative and predictive AI.
B. The New Landscape of AI in Claims
Modern claim-handling systems employ a range of artificial intelligence applications, each addressing a different aspect of the loss evaluation process.
1. Property-Damage Assessment
In property and auto claims, some image-recognition platforms use computer-vision models to analyze photographs of damaged structures or vehicles and generate cost estimates. These systems extract patterns—material types, breakage contours, water staining, roof damage—from uploaded imagery, often producing an initial estimate before an adjuster ever visits the site.
Other platforms extend this approach to full virtual inspections by using spatial-recognition algorithms to determine room geometry, materials, and fixtures directly from smartphone photos. These systems effectively operationalize the Colossus ideal of standardized evaluation in the property context, replacing onsite measurements with digital modeling. While they can dramatically reduce inspection times, they also risk replicating Colossus’s chief shortcoming: the substitution of algorithmic uniformity for contextual judgment. Lighting conditions, debris, or atypical construction materials can skew an AI’s interpretation, and the line between efficiency and oversimplification again becomes blurred.
2. Bodily-Injury and Fraud Detection Models
For bodily-injury and personal-lines claims, insurers may rely on predictive and anomaly-detection models reminiscent of Colossus’s original domain but enhanced with modern machine-learning architecture. Some commercially available AI platforms evaluate claim narratives, medical billing, and historical claim data to detect inconsistencies or potential fraud. Certain systems combine structured data with natural-language processing of adjuster notes, producing risk scores that triage claims for additional review.
Complementing these are advanced text- and pattern-recognition systems that mine claim narratives and historical records for linguistic or statistical anomalies suggestive of misrepresentation or exaggeration. Like their Colossus predecessor, these systems quantify human judgment—here through pattern extraction rather than explicit rule encoding.
As with Colossus, these tools promise consistency and efficiency—but they also pose the same ethical dilemma: if a claim is flagged as “suspicious” by an opaque model, what is the insurer’s evidentiary basis for that conclusion? The program’s internal weighting of factors—location, treatment type, prior claims—may be invisible to the adjuster who relies on it. In this way, the Colossus debate over transparency re-emerges in algorithmic form.
3. Workflow Automation and Communication Tools
Beyond valuation, AI now automates parts of the claim lifecycle once deemed purely administrative. Some carriers use machine-learning triage systems to route claims to appropriate handling paths, while others experiment with generative text tools that draft coverage letters or status updates. These automations can improve customer experience and free adjusters for complex tasks, but they also introduce potential legal risk. A misworded AI-generated communication—particularly one that inaccurately represents coverage—may be discoverable evidence in a later bad faith action. Once again, technology that promises efficiency can quietly erode the precision and care required in insurer communications.
4. The Parallels Between Colossus and Modern AI Tools
Across all these tools, the parallels with Colossus are striking. Each new technology advances the same triad of objectives—speed, uniformity, and cost control—while exposing insurers to the same triad of risks—opacity, bias, and overreliance. The lesson of the past two decades remains constant: automation can assist the adjuster, but it cannot become the adjuster.
To synthesize the current marketplace of tools and their potential implications, the following summarizes representative examples of AI applications in claim handling as discussed above:
| Use Case | Purpose/Function | Representative Applications | Observations/Limitations |
| Property-damage assessment | Uses computer vision to analyze photos or video and estimate repair costs. | Computer-vision and virtual-inspection platforms | Improves efficiency but may misread atypical damage; requires human validation. |
| Bodily-injury evaluation & fraud detection | Applies predictive analytics to claim data and medical records to flag anomalies or suspicious patterns. | Predictive analytics, anomaly-detection, and text-mining platforms | Enhances consistency; opacity of models raises fairness and “explainability” concerns. |
| Claims triage and routing | Automatically categorizes and prioritizes claims for appropriate handling. | Proprietary or vendor-developed triage systems | Risk of misclassification; requires audit trails and override authority. |
| Automated communications | Drafts coverage letters and claim-status messages using rule-based or generative AI. | Internal carrier systems; vendor plug-ins | Must be reviewed for legal accuracy; errors may create bad-faith exposure. |
C. Fraud in the Age of AI
If Colossus represented the insurer’s early flirtation with algorithmic overreach, generative AI marks the insured’s opportunity for technological abuse. Where automation once threatened undervaluation, it now enables fabrication—the deliberate creation of false evidence through AI. The same tools that allow insurers to analyze photographs can also allow claimants to fabricate them.
The most common manifestation is the AI-generated image of property damage. A claimant might superimpose fire, water, or structural damage onto authentic photos or generate entirely synthetic images of a purported loss. Some falsify metadata to make the photos appear contemporaneous with the claimed event. Others produce fabricated invoices or repair estimates generated by text-based AI systems.
Large international insurers have reported detecting claims supported by manipulated imagery, often identified through telltale inconsistencies in lighting, shadows, or metadata.[2] Specialized forensic-AI systems now scan submitted photos for generative artifacts—digital “fingerprints” left by diffusion or generative models.[3]
The asymmetry of technology has created a new cat-and-mouse dynamic: claimants armed with generative tools versus insurers developing discriminative ones. Yet the underlying legal principles are unchanged. The insurer retains both the right and the duty to investigate suspected fraud, and must do so with diligence and documentation sufficient to support its decision if challenged in litigation.
Ultimately, this emerging threat highlights the mirror image of the Colossus problem. Where the first generation of algorithms risked dehumanizing claims through artificial standardization, the new generation threatens to deceive them through artificial fabrication. Both scenarios reveal that the heart of claim handling—truthful, individualized assessment—remains a human responsibility, regardless of technological sophistication.
D. Balancing Efficiency and Oversight
The path forward lies not in rejecting AI, but in integrating it responsibly. Insurers can preserve the benefits of automation while mitigating its risks through several interlocking practices.
First, human-in-the-loop review must remain nonnegotiable. No AI output, however sophisticated, should serve as the final basis for a coverage or valuation determination without human confirmation. Adjusters should understand the inputs and confidence thresholds that underpin AI recommendations. Deviations—i.e., whether to accept or reject an AI’s conclusion—should be documented.
Second, model governance and auditability are essential. Insurers should maintain records of model versions, training data, and updates, as well as track overrides or exceptions to algorithmic recommendations. Such documentation not only strengthens compliance but also provides a defensible record if a dispute arises.
Third, vendor transparency should be contractually required. Insurers adopting third-party AI tools should insist on access to model documentation and, where possible, audit rights. Without such safeguards, an insurer may find itself unable to explain a decision in discovery—an evidentiary vulnerability reminiscent of the Colossus era.
Finally, training and culture matter. Claim professionals should be educated not only on the functionality of AI tools but on their limitations. The same duty of fairness that applied in the manual-claims era persists in the digital one. Automation should accelerate human reasoning, not replace it.
The above discussion reveals the dual nature of AI in claim handling: it is both an instrument of efficiency and a potential vector of error or alleged deceit. Like Colossus before it, each innovation carries with it a renewed obligation—to preserve transparency, maintain discretion, and ensure that the pursuit of technological progress never eclipses the foundational principle of fair and individualized claim evaluation.
III. Legal Landscape
If Colossus marked the industry’s first encounter with algorithmic accountability, modern artificial intelligence represents its jurisprudential sequel. Courts have yet to articulate a comprehensive framework for AI-assisted claim handling, but familiar doctrines—bad faith, evidentiary reliability, and fairness—provide a starting point. The question is not whether existing law applies, but how it adapts.
A. Bad Faith and the Standard of Reasonableness
At the core of every bad-faith inquiry lies a simple question: Was the insurer’s conduct reasonable under the circumstances? Artificial intelligence complicates that inquiry by adding a layer of mechanical judgment between the insurer and its insured. When an AI system influences claim valuation or denial, the reasonableness of the insurer’s conduct becomes inseparable from the reasonableness of its reliance on that system.
Courts assessing Colossus-related claims applied this logic implicitly. They did not condemn automation per se but scrutinized whether insurers used it responsibly—whether adjusters exercised independent judgment rather than deferring to the software’s recommendation. The same reasoning will likely govern modern AI cases in the extracontractual context. If a carrier denies, limits, or delays payment based on an algorithmic recommendation it cannot explain or verify, policyholders may argue that the insurer abdicated its non-delegable duty of good faith.
In Anderson v. Nationwide Mutual Insurance Co., for example, the court held that “failure to evaluate the insured’s claim based on individualized evidence, rather than rigid internal metrics, may support a finding of bad faith.”[4] While Anderson predated the current AI wave, its principle maps directly onto contemporary claim handling: an insurer that substitutes model output for individualized analysis risks crossing the line from efficiency to indifference.
Modern case law in adjacent contexts reinforces this approach. In State Farm Fire & Casualty Co. v. Slade, the Alabama Supreme Court emphasized that “an insurer’s claim-handling practices must remain guided by professional judgment, even when informed by standardized procedures.”[5] And in Rancosky v. Washington National Insurance Co., the Pennsylvania Supreme Court reaffirmed that bad faith encompasses reckless disregard for the absence of a reasonable basis to deny benefits.[6] These holdings collectively suggest that automation offers no safe harbor: an unreasonable belief in the accuracy of an opaque algorithm may be analogous to an unreasonable belief in a flawed human assessment.
B. Potential Judicial Treatment of Algorithmic Evidence
Although no U.S. court has yet ruled squarely on the admissibility or reliability of AI claim-handling tools, early litigation hints at the contours of the debate. In discovery disputes, plaintiffs are likely to begin to seek model documentation, training data, and internal correspondence regarding algorithmic decision-making. Courts may soon confront questions analogous to those raised in Colossus litigation, principally: Must an insurer disclose the internal logic of an AI model if it forms part of the claim-decision process?
A few cases already gesture toward answers, or at least provide a snapshot of potential judicial treatment of the issue. In United States v. Loomis, a criminal-sentencing case involving a proprietary risk-assessment algorithm, the Wisconsin Supreme Court upheld use of the algorithm but warned that reliance on an undisclosed model required “caution and transparency.”[7] Although Loomis was not an insurance case, its reasoning—emphasizing due process and “explainability”—may influence civil discovery disputes involving insurer AI systems. Similarly, in In re State Farm Lloyds (Tex. Sup. Ct. 2020), the court addressed electronic claim-handling databases, holding that insurers could not withhold data essential to understanding how loss values were calculated.[8] Both decisions underscore the judiciary’s growing expectation that technological processes affecting substantive rights be explainable.
In practice, the evidentiary challenge for insurers will be twofold: first, to preserve and produce sufficient documentation to show that an AI recommendation was reasonable; and second, to train personnel capable of articulating that reasoning in deposition or trial testimony. A company that cannot explain its own system invites the same skepticism that plagued Colossus users two decades ago.
C. Looking Forward
Taken together, these developing doctrines suggest a clear trajectory. Courts are unlikely to craft new “AI law” for insurers; instead, they will likely apply traditional standards—reasonableness, good faith, and “explainability”—to new facts. Just as Colossus taught that automation does not immunize a carrier from scrutiny, modern AI will teach that sophistication does not substitute for accountability. The insurer’s safest course remains the oldest one: exercise judgment, document reasoning, and treat every claim as an individual undertaking rather than a data point. In the end, fairness and transparency—the same qualities arguably absent from the Colossus controversy—remain the best defenses against the next generation of litigation.
IV. Discovery and Evidentiary Issues
As insurers begin to integrate artificial intelligence into claim handling, disputes will inevitably test how courts treat AI-influenced decisions—both in contractual coverage cases and in extracontractual “bad faith” litigation. Discovery requests will probe the extent to which AI shaped claim outcomes, while evidentiary questions will determine how such systems and their outputs are presented to a jury. Many of these challenges mirror those that arose during the Colossus era, but modern AI adds layers of technical and procedural complexity.
A. Scope of Discovery in Contractual and Extracontractual Claims
In coverage litigation, discovery will typically focus on what an insurer’s AI system produced—namely, the estimates, analyses, or other outputs that contributed to a coverage determination or valuation of the loss. These outputs are part of the claim file, and policyholders may seek to understand how those AI-generated results influenced the insurer’s interpretation of the scope/availability of coverage or the amount of loss. While discovery at this stage generally stops at the product of the system’s analysis rather than its inner workings, courts may allow limited inquiry if the reliability of that output becomes relevant to the contractual dispute.
In bad-faith litigation, discovery reaches further—to how the AI system was used, understood, and supervised within the claim-handling process. Plaintiffs are likely to explore whether the insurer’s reliance on an algorithm was reasonable and whether the adjuster understood the system well enough to make an informed decision. Depositions may examine the adjuster’s familiarity with the AI tool, the extent of training received, and the level of discretion retained in accepting or rejecting its recommendations. Discovery may also probe the degree to which supervisors or management personnel were aware of or participated in the use of AI, and whether internal measures existed to ensure that adjusters were not applying AI outputs without oversight or human verification.
The Colossus disputes previewed this dynamic. Policyholders in those cases contended that adjusters had become conduits for opaque software rather than independent evaluators, resulting in mechanical and uniform claim decisions. The same argument will likely reemerge in AI-related bad-faith claims, reframed around machine learning and “explainability”—the ability of human decision-makers to articulate how and why they relied on algorithmic recommendations.
B. Preservation of AI-Specific Data
Preservation duties apply in every claim dispute, but AI introduces a new wrinkle: models evolve. Machine-learning systems are dynamic, updating their parameters as new data enters the training pipeline. If a claim decision is later challenged, it may be difficult—or impossible—to reproduce the precise model state that produced the disputed output unless it was preserved contemporaneously.
Accordingly, insurers that employ AI should ensure that their litigation-hold procedures capture:
- Model version identifiers and configuration files — preserving the specific build or release number used during the relevant claim period;
- Input and output data pairs — the photos, documents, or structured data fed into the system and the resulting AI analysis or valuation;
- System logs and metadata — timestamps, user IDs, and confidence scores showing how and when the model processed each claim; and
- Retraining documentation — if the model was later updated, retaining records of when and how those updates occurred.
This level of preservation allows an insurer to recreate, if necessary, the analytic environment that produced the AI recommendation. Unlike ordinary claim documentation, these artifacts may exist outside the traditional claim file and require coordination between claims, IT, and data-science personnel. While courts have not yet ruled on preservation obligations specific to AI, best practice dictates erring on the side of completeness—particularly where an insurer anticipates that its use of AI could become a contested issue.
C. Protecting Proprietary and Trade-Secret Information
At present, relatively few insurers use AI tools in live claim environments, but those that do typically rely on vendor-developed, commercially available systems rather than proprietary in-house software. When litigation arises, policyholders may seek discovery into the inner workings of those systems—source code, training data, or algorithmic logic—to challenge the reliability of the insurer’s reliance on them. Because such materials often belong to third-party vendors and contain trade-secret information, courts will likely balance competing interests: the policyholder’s need for relevant discovery versus the vendor’s right to protect its intellectual property.
Protective measures should mirror those employed in prior Colossus disputes:
- Protective orders limiting access to attorneys’ eyes only or qualified experts;
- Redaction or summarization of proprietary information; and
- In-camera review by the court where necessary to evaluate relevance without disclosure.
Insurers should anticipate these discovery pressures and coordinate early with vendors to establish a disclosure protocol that protects confidentiality while satisfying reasonable discovery obligations. Courts are generally sympathetic to trade-secret concerns but require insurers to demonstrate diligence and transparency in proposing protective solutions.
D. Factual and Expert Testimony
When AI-influenced claim decisions reach trial, two categories of testimony will often be required: factual testimony from the claim professional who handled the claim and expert testimony addressing the technical reliability of the AI system itself.
In extracontractual litigation in particular, the claim professional serves as the key fact witness—i.e., the person directly responsible for handling the disputed claim who articulates the basis for the claim handling and/or decision. This witness must explain, with clarity and credibility, how the AI tool was used in the claim, what outputs it generated, and how those outputs informed (but did not dictate) the decision made. The testimony should demonstrate that the adjuster exercised independent judgment, understood the limitations of the AI system, and applied its recommendations in conjunction with other claim evidence. In short, this witness provides the “explainability” component: an account of how human reasoning interacted with algorithmic assistance in the specific case before the court.
The expert witness—potentially a data scientist, software engineer, or forensic technologist—plays a complementary role by establishing the reliability of the AI system itself. This expert may describe how the model operates, what data it processes, and how accuracy or error rates are measured. This testimony will serve as the foundation for admissibility, addressing whether the system’s methods are sufficiently trustworthy to be considered by the factfinder. The expert may also explain how the model was validated for the use in insurance settings and how its performance compares to accepted industry standards. The goal is not to defend the insurer’s individual decision but to show that the technology, when used properly, produces consistent and verifiable results suitable for claim evaluation.
Together, these two witnesses connect process to principle: one explaining how the decision was made, the other explaining why the system could reasonably be trusted. This dual-witness framework—human “explainability” supported by technical reliability—will likely become the evidentiary template for AI-related insurance trials. Courts accustomed to expert testimony on engineering models or forensic simulations are well positioned to adapt these principles to algorithmic evidence.
E. Practical Takeaways
AI will not eliminate discovery battles; it will simply shift their focus. Insurers can reduce litigation risk by anticipating how their tools might appear under a microscope and ensuring they can answer three essential questions:
- What did the AI system do?
- How can that process be reconstructed and explained?
- What safeguards protected proprietary or trade-secret information during litigation?
In both contractual and extracontractual cases, success will depend less on the sophistication of the technology than on the insurer’s ability to translate it—to show that the claim was still handled by people, guided by reason, and documented with care. As in the Colossus era, transparency and comprehension remain the surest defenses.
V. Best Practices and Practical Guardrails
Artificial intelligence can enhance claim handling, but only when implemented within a disciplined framework that preserves human judgment, transparency, and accountability. The following proposed best practices, drawn from lessons of the Colossus era and early adoption of modern AI tools, serve as practical guardrails for insurers navigating this evolving terrain.
A. Governance and Oversight
AI adoption in claims handling should begin with governance, not technology. Insurers must clearly define who is responsible for selecting, testing, and approving AI tools, and how those tools are integrated into claim workflows. A cross-functional governance team—comprising claims leadership, legal, compliance, and data specialists—should review any proposed AI implementation for accuracy, fairness, and interpretability.
Claims organizations’ policies should specify the permissible role of AI outputs in claim decisions: whether as a recommendation, a starting point, or a required step in valuation. These parameters should be documented in claims-handling manuals and communicated consistently to field adjusters. As courts scrutinize how AI is used, written governance standards will be critical in demonstrating reasonableness and oversight.
B. Training and Human Oversight
Technology does not diminish the adjuster’s role; it heightens the need for skill. Adjusters and supervisors should receive targeted training on how the AI systems they use operate—what inputs they rely on, how they produce results, and what their limitations are. That knowledge forms the foundation for “explainability” if the decision is later challenged.
Human review must remain a required step. Every AI-generated estimate, valuation, or fraud score should be subject to human verification before influencing a coverage or payment determination. This “human-in-the-loop” requirement preserves accountability and helps avoid the perception of unreviewed automation. Supervisory oversight—through documented review steps or escalation protocols—should confirm that adjusters remain the ultimate decision-makers.
C. Documentation and Auditability
Because AI-driven processes can appear opaque, documentation becomes the primary safeguard. Claim files should reflect not only the AI output but the adjuster’s reasoning in accepting, modifying, or rejecting it. If the model provides a confidence score or error margin, that information should be recorded along with an explanation of how it factored into the decision.
Periodic internal audits should test whether AI recommendations align with human assessments across sample claims. Discrepancies—particularly systematic ones—should trigger model review or retraining. These audits serve a dual purpose: improving performance and building a record that the insurer monitors its technology responsibly.
D. Vendor Management and Contractual Controls
Because most insurers rely on vendor-developed AI systems, the relationship with those vendors must be managed as carefully as the technology itself. Contracts should address not only service-level expectations but also:
- Transparency obligations — requiring vendors to provide documentation explaining model design, testing, and update frequency;
- Notification of material changes — obligating vendors to alert the insurer to retraining, parameter adjustments, or algorithmic updates that might affect claim outcomes;
- Data ownership and audit rights — ensuring the insurer retains access to the data generated within its own claim processes; and
- Litigation cooperation — requiring vendor assistance in responding to discovery requests or providing affidavits verifying the system’s reliability.
These provisions position the insurer to meet its discovery and evidentiary obligations without compromising proprietary vendor information.
E. Fairness and Bias Mitigation
Even when AI operates as intended, the data underlying it may perpetuate historical inequities. To maintain accuracy and credibility, insurers should periodically test AI outputs for unexpected disparities across claim types, regions, or other relevant factors. If discrepancies appear, the cause should be identified—whether data quality, feature selection, or unbalanced training sets—and corrected through model adjustment or retraining.
These self-audits need not imply bias; they simply demonstrate due diligence and reinforce the insurer’s commitment to producing consistent, data-driven results that align with its claim-handling obligations.
F. Continuous Evaluation and Legal Collaboration
AI in claims handling is not static. As tools evolve, insurers should periodically revisit their governance and compliance frameworks. Legal and claims departments should collaborate regularly to review active litigation trends, discovery demands, and judicial treatment of algorithmic evidence. This collaboration ensures that both business practices and litigation strategies evolve together.
The lesson of Colossus was not that automation should be avoided but that it must be understood. The same principle governs the use of AI today. The insurer that treats artificial intelligence as a decision-support tool—rather than a decision-maker—will be best positioned to capture its benefits while minimizing exposure.
G. Practical Summary for Claim Professionals
For adjusters and claims leaders, the guiding principles reduce to four concise rules:
- Understand it. Know what the AI tool does, what data it uses, and where it can fail.
- Explain it. Be able to describe, in plain language, how the AI’s output informed your decision.
- Document it. Record what the system produced and how you evaluated it.
- Supervise it. Maintain human oversight, both individually and organizationally.
If the insurer can satisfy those four imperatives, its integration of AI into claim handling should remain both defensible and effective.
VI. Conclusion
The evolution from Colossus to contemporary artificial intelligence represents less a technological revolution than a legal continuum. The tools have changed, but the questions have not: How should insurers balance efficiency with fairness? To what extent may they rely on automated systems without abandoning the individualized assessment that defines good-faith claim handling? And how can courts evaluate decisions influenced by models whose inner logic may never be fully transparent?
The Colossus experience demonstrated that automation, when misunderstood or unchecked, can erode trust as quickly as it delivers efficiency. The lesson was not that technology should be rejected but that it must remain subordinate to human judgment. That same lesson governs the deployment of modern AI. Machine-learning systems, image-recognition platforms, and predictive models can enhance accuracy and speed—but only when used as aids to professional reasoning rather than replacements for it.
For courts, the next decade will likely involve refining how discovery, evidentiary standards, and bad-faith doctrines apply to AI-assisted claims. For insurers, the task is more immediate: to design claim-handling frameworks that preserve transparency, document human oversight, and ensure that every decision can be explained and defended. The insurer that can translate its technology—showing how human discernment guided the process—will fare best before regulators, judges, and juries alike. In the end, the adoption of artificial intelligence in claim handling does not alter the fundamental covenant of insurance. The obligation remains to investigate fully, evaluate fairly, and pay promptly what is owed. Technology may change the means, but not the measure, of that duty.
[1] See, e.g., John Lamothe, “When Your Opponent Is a Computer Algorithm,” Louisiana Advocates (Oct. 2017) (noting regulators’ concerns that Colossus created standardized biases in bodily-injury valuation and eroded adjuster discretion).
[2] Reports of manipulated claim images detected by major European carriers; see “AI Offers Solutions—and Challenges—in Fighting Fraud,” Insurance Insider (2024).
[3] See RGA Reinsurance Co., “Artificial Intelligence and Insurance Fraud: Four Dangers and Four Opportunities” (2024).
[4] Anderson v. Nationwide Mut. Ins. Co., 117 F. Supp. 2d 1249 (M.D. Ala. 2000).
[5] State Farm Fire & Cas. Co. v. Slade, 747 So. 2d 293 (Ala. 1999).
[6] Rancosky v. Washington Nat’l Ins. Co., 170 A.3d 364 (Pa. 2017).
[7] State v. Loomis, 881 N.W.2d 749 (Wis. 2016).
[8] In re State Farm Lloyds, 520 S.W.3d 595 (Tex. 2020).
Contact & Disclaimer
Houston Harbaugh's insurance and AI litigation team continues to monitor developments in algorithmic decision-making and emerging liability frameworks. For questions regarding enterprise AI risk management or policy development, please contact Christopher M. Jacobs, Esq. at jacobscm@hh-law.com or call 412-288-4019.
This post represents the author's personal views and does not constitute legal advice. Defending the Algorithm™ is a mark of Houston Harbaugh, P.C.
About Us
We’re committed to staying on top of the issues of today and tomorrow, such as the ever-changing landscape involving bad faith, cyber-insurance, and insurance for advanced technology sectors, artificial intelligence players, machine learning companies, and autonomous vehicle manufacturers and users.
Alan S. Miller - Practice Chair
Alan has more than thirty-eight years of experience in complex litigation and counseling, concentrating in the areas of environmental law, insurance coverage and bad faith, and commercial litigation. He chairs the firm’s Environmental and Energy Law practice and the Insurance Coverage and Bad Faith Litigation Practice.
Alan’s environmental law practice has involved counseling, litigation and alternative dispute resolution of matters involving municipal, residual, and hazardous waste permitting and compliance, contribution and cost recovery actions under CERCLA and related state statutes, claims for natural resource damages, contamination from leaking underground storage tanks, air and water pollution regulatory permitting and enforcement actions, oil and gas drilling compliance and transactions, and real estate transactions involving contaminated and recycled industrial sites.