In 2022, over a period of two months, Cigna physicians denied more than 300,000 payment requests and spent an average of 1.2 seconds on each, according to a ProPublica investigation. Cigna called the reporting biased and incomplete and described its process as a way to speed payment for routine screenings. A class action followed, and its allegations remain unproven.

Whether or not a given tool qualifies as AI, the episode raises the question that every AI-assisted decision now faces. A licensed physician signed each denial. The harder question is whether a physician decided.

Across more than 700 security risk assessments and two decades in healthcare security, including my years as CISO and Chief Privacy Officer at ClearDATA, the standard has been consistent. A control counts when you can show that it operated. Human review of AI recommendations belongs under that standard, and the evidence for it deserves the same attention as any other control.

What The Statutes Require

States have moved quickly to put a human in the loop. California's SB 1120, effective January 1, 2025, requires that a licensed physician or qualified health care provider review and decide any denial, delay, or modification of care based on medical necessity. Alabama's SB 63 takes effect October 1, 2026. It requires insurers that use AI in prior authorization to base determinations on the patient's individual medical history and clinical circumstances, and to certify annually to the state Department of Insurance. Utah's SB 319, effective January 1, 2027, requires adverse determinations to reflect a healthcare professional's independent medical judgment, separate from AI recommendations. Georgia, also effective January 1, 2027, bars an adverse determination from issuing without review and approval by a licensed healthcare provider. Washington limits adverse determinations on prior authorization requests to licensed or qualified health professionals.

At the federal level, CMS told Medicare Advantage plans in a February 2024 FAQ that they may use algorithms and AI in coverage determinations, provided each determination rests on the circumstances of the individual patient.

These laws answer one question: who decides. How an examiner will judge whether the review was real is left to enforcement, and enforcement will look for evidence. A rubber stamp with a medical license behind it satisfies the wording of these statutes and misses their purpose.

Why Human Review Degrades

Research on automation bias explains why a signature is a weak proxy for a decision. A 2012 systematic review in the Journal of the American Medical Informatics Association by Goddard, Roudsari, and Wyatt retrieved 13,821 papers and included 74. It defines automation bias as the tendency to over-rely on automated advice, and it identifies workload, time constraints, and trust in the system among the factors that make it more likely. The same review reports that decision support often improves overall performance, and that users frequently miss the new errors it introduces.

A 2023 study in Radiology by Dratsch and colleagues puts numbers on the effect. Twenty-seven radiologists read 50 mammograms with AI suggestions. When the suggestion was correct, inexperienced readers scored about 80 percent. When it was wrong, they scored below 20 percent. The most experienced readers fell from 82 percent to 45.5 percent.

Mammography is not utilization review, and 50 cases is a small sample. The direction of the effect is the point. Experience reduced the damage without removing it. A reviewer working through a queue of AI-generated recommendations faces the conditions the literature associates with over-reliance: volume, time pressure, and an answer already on the screen.

How Courts And Regulators Look Past The Signature

Plaintiffs and regulators already examine substance. In Estate of Lokken v. UnitedHealth, a federal court in Minnesota ruled on February 13, 2025 that breach of contract and good faith claims could proceed over an algorithm called nH Predict, and it dismissed the other claims. The plaintiffs allege that the model's output was applied regardless of treating physicians' recommendations, and that more than 90 percent of claim denials were reversed on appeal. UnitedHealth disputes the plaintiffs' account, and the ruling made no finding on the allegations.

European law names the risk directly. Article 14 of the EU AI Act requires that the people overseeing high-risk AI systems stay aware of the possible tendency to over-rely on the output, and that they can disregard, override, or reverse it. The Court of Justice of the EU, in its December 2023 SCHUFA ruling, treated a credit score as an automated decision when the lender relied heavily on it, even though a person made the final call.

Each setting reaches the same place. The person who signs is not always the person who decides.

Building Evidence Of Meaningful Review

A control needs evidence that it operated. For human review of AI recommendations, six measures give an examiner something to test:

  • Log Review Time. Record how long each reviewer spends on each case, and watch the distribution rather than the average.
  • Track Override Rates. A reviewer who never disagrees with the model is a finding. Track rates by reviewer, by case type, and over time.
  • Present The Record First. Show the clinical record and coverage criteria before the model's recommendation, and require the reviewer's own written rationale. The Goddard review lists presenting information instead of direct recommendations among the design approaches that reduce automation bias.
  • Sample And Re-Review. Have an independent clinician re-review a sample of decisions, including those where the reviewer and the model agreed, and report the agreement rate.
  • Cap Reviewer Workload. Set limits on daily volume per reviewer, since workload and time constraints are named drivers of over-reliance.
  • Verify Qualifications. Match reviewer credentials to the clinical issue and keep the record, since several statutes turn on who is qualified.

These measures carry costs and can be gamed. A short review can be correct, and a long one can be careless. Time thresholds invite reviewers to run the clock, and sampling adds clinical labor. I treat the metrics as prompts for scrutiny and keep clinical judgment as the standard. I would still rather defend a program that measures its review than one that assumes it.

The Path Forward

Behind every one of these decisions is a patient waiting to learn whether care will be covered. The statutes ask for a human decision. Regulators, courts, and patients will ask for proof of one.

Organizations that build that proof now will be ready when the first request arrives. I write about healthcare security and AI governance at Bowen Perspectives.