Key Takeaways
- As AI systems take on more clinical decision support functions, there is a risk that clinicians' unassisted diagnostic and reasoning skills will atrophy through disuse.
- Automation bias—the tendency to defer to algorithmic recommendations even when they conflict with clinical judgement—is strongest among clinicians with less baseline expertise.
- The question of who is accountable when AI-assisted decisions cause harm remains unresolved, with current frameworks placing disproportionate responsibility on frontline clinicians.
- Preserving clinical judgement requires deliberate design choices that keep clinicians engaged with reasoning processes rather than passive recipients of AI outputs.
What happens to clinical judgement when AI does the thinking?
Shoshana Bloom, Equiti Health Ltd | April 2026
The use of AI in clinical settings has outpaced the competencies in our clinicians who use it. The evidence now tells us exactly where the risks are, and what needs to change.
The policy discussion around the use of AI in clinical settings has centred largely around questions of performance. Can AI match the diagnostic accuracy of a consultant radiologist? Can it identify deteriorating patients faster than the clinical team? These are legitimate questions, and on many of them the evidence is encouraging.
They are not, however, the only questions that determine whether AI use in healthcare is safe and equitable.
What determines safety goes beyond these questions. What happens to the clinician who uses an AI tool repeatedly, over months and years, without the competency structures that protect their skills and independent clinical judgement? What happens to the patient who receives AI-influenced care in a setting where the tool has not been validated for their demographic? Or what about the population whose health outcomes are shaped by an algorithm that nobody has audited for bias?
At Equiti Health, we have spent the past three months reviewing 445 studies to answer those questions. The findings are published today as a white paper: Clinical Competency in the Age of AI: Safe, equitable, and effective practice in AI-augmented healthcare. This article summarises the key findings.
The evidence on skill atrophy is no longer theoretical
The concern that AI dependency might erode clinical capability was, until recently, a reasonable inference rather than an evidenced finding. In 2025, that changed.
Budzyń et al. documented adenoma detection rates falling from 28% to 22% over three months of AI-assisted endoscopy. This was a measurable decline in diagnostic performance during a period when nobody was monitoring the skill and nothing in the outcome data flagged a problem until the analysis was done. This is the first real-world evidence of what the literature terms deskilling: the progressive atrophy of clinical skills through reduced independent practice.
Consider how many of us can no longer read a map with any confidence, having navigated using GPS for the past decade. That erosion of capability may be inconvenient but it is rarely consequential. Now apply the same mechanism to the clinical skills required to detect a lesion, interpret an ambiguous scan, or assess a deteriorating patient — at the moment the AI is unavailable, the consequence is not inconvenience. It is patient harm.
In clinical AI, the equivalent is happening at scale, unremarked and unmeasured. This introduces risk into every clinical setting across every health system that has deployed AI without protecting the independent practice that deployment progressively displaces.
Two related mechanisms compound it.
Mis-skilling describes the consolidation of incorrect clinical reasoning patterns through repeated, uncritical acceptance of plausible but erroneous AI outputs — not skill loss but the active embedding of error.
Never-skilling describes the failure of foundational clinical capabilities to develop at all in practitioners training in AI-intensive environments. The newly qualified clinician who has never performed an unassisted clinical assessment has no independent baseline against which to evaluate AI outputs. They cannot detect when the AI is wrong because the judgement required to do so was never formed.
These risks are not detectable through standard clinical performance management. They accumulate through multiple individual encounters and become visible only through the patient outcomes they eventually produce.
The evidence on automation bias is striking
Automation bias — the systematic tendency to over-weight AI recommendations relative to independent clinical judgement — has now been documented across 45 included studies. The most clinically significant demonstration comes from a randomised trial published in JAMA in 2023 (Jabbour et al.): physician-plus-AI arms performed worse than AI operating alone. The presence of AI in the diagnostic process actively suppressed the critical appraisal that would otherwise have served as a check on its outputs.
This effect intensifies under conditions of cognitive load, precisely the conditions of routine clinical practice. Alert override rates of 90–96% in deployed clinical AI environments document systematic suppression at scale (Graafsma et al., 2024; Co et al., 2020). AI tools assessed as safe under controlled evaluation conditions carry significantly elevated risk when deployed into cognitive load environments such as emergency departments, overnight primary care, understaffed wards.
There is a further complication. AI deployment is typically justified partly on efficiency grounds. However, the evidence suggests that efficiency gains may not materialise as predicted. Reid and Mateen (2025) and Fiehler (2026) apply the Jevons' Paradox directly to clinical AI: lowering the marginal cost of performing a clinical task generates additional demand for that task, absorbing efficiency gains and potentially increasing total workload. The same tools generating efficiency may simultaneously elevate demand and volume, multiply alert burden, and worsen the clinical burnout that then amplifies automation bias.
Current governance gaps compound all these risks
A study by Oskrochi et al. (2025) found that more than 70% of health organisations have deployed clinical technologies without documented clinical safety assurance.
The practical consequence is that clinicians in organisations without effective governance and safety checks are carrying full personal professional accountability for AI-assisted decisions, without the institutional infrastructure that should underpin them. No training framework currently tells clinicians they have a right, and a professional obligation, to establish whether AI tools they are asked to use have been properly safety-assured. No framework provides the escalation pathway for when they have not.
The equity dimension compounds every other risk
Obermeyer et al. (2019, Science) demonstrated that a widely deployed commercial health management algorithm systematically underestimated the clinical needs of Black patients at equivalent risk scores — bias operating at scale, undetected, for years. More recent work has documented comparable disparities in dermatology AI, imaging AI, and clinical decision support (Daneshjou et al., 2022; Cross et al., 2024).
The structural problem is not simply that biased AI exists. It is that the populations most at risk from biased AI outputs are often served by clinical settings least equipped to detect and correct that bias. Governance weakness is typically most acute in under-resourced settings.
The clinical settings that disproportionately serve deprived communities are the most likely to be underrepresented in AI training data. And patients with lower health literacy or limited digital access are simultaneously more exposed to AI-influenced decisions and least able to exercise informed preferences in response to them.
Twenty-three competency frameworks — all with common gaps
So now we understand the risks, how are clinicians being trained to manage and mitigate them? We reviewed 23 existing clinical AI competency and capability frameworks, including the DECODE framework (Car et al.), published in JAMA Network Open — the most rigorous international consensus work in the field, produced through a Delphi process involving more than 200 participants from 79 countries.
What the analysis shows, across all 23, is that the competencies critical for front-line patient safety are often assigned to commissioners, technical specialists, and governance leads — not to the practising clinicians bearing professional accountability for AI-assisted decisions.
Four structural gaps are identified:
- No framework makes preventing deskilling, mis-skilling, or never-skilling an operational, assessable clinical competency.
- No framework makes the detection of algorithmic bias in specific tools, for specific patient groups, a named point-of-care clinical skill.
- No framework establishes governance literacy — the ability to assess whether AI tools in use are safety-assured and to act when they are not — as a patient safety requirement for front-line clinicians.
- No framework addresses AI disclosure or AI-augmented shared decision-making as explicit clinical competencies, despite patients' right to know when AI has influenced decisions about their care.
A new framework that addresses these risks
The framework proposed in this research comprises five domains specified across three career stages: Foundation (students and trainees), Practitioner (qualified clinicians), and Senior/Leadership.
Its central structural principle is that the competencies most critical for front-line patient safety must be competencies for front-line clinical users. Governance requirements assigned exclusively to commissioners and technical specialists do not protect the clinician in the consulting room, or the patient in front of them.
Domain 1 — Technical AI Literacy. The prerequisite. Do you understand core concepts of AI required to use it safely? Can you evaluate whether the tool was validated on patients like yours? Can you recognise when a confidence score is misleading?
Domain 2 — Clinical Judgement Preservation. Form your own view before reviewing the AI output. At senior level: can you define the minimum competency thresholds your team must maintain, and protect the unassisted practice environments that make that possible?
Domain 3 — Algorithmic Bias Awareness and Health Equity. Bias detection as a point-of-care skill, going beyond abstract awareness. Know the training population of tools in your environment. Recognise equity red flags for specific patients. Escalate when you find one.
Domain 4 — AI Safety, Governance and Professional Accountability. Clinicians have a professional right and obligation to establish whether AI tools they use have been safety-assured. How do you assess this? What to do when you have concerns?
Domain 5 — Professional Agency and Person-Centred Care. Can you explain to a patient, in accessible language, that AI shaped the decision about their care? Can you support them to engage with that meaningfully? For patients with lower health literacy or limited digital access, this is not optional.
The question has changed
The question is no longer whether AI can perform at the level of a clinician. In many domains, it already can. The question that now determines safety, equity, and quality is whether the systems and professionals using AI are equipped to do so responsibly.
The evidence is clear on three points.
First, the risks are real and already documented: skill atrophy, automation bias, and embedded inequity are now observable phenomena in deployed clinical environments.
Second, these risks are systemic rather than individual. They arise not from clinician behaviour alone, but from the interaction between tools, environments, and the absence of appropriate competency and governance structures.
Third, current approaches to training and assurance are misaligned with where risk actually sits: at the point of care, with the clinician making decisions in real time.
Clinical AI safety requires a parallel investment in clinical competency, governance literacy, and professional agency among those using these tools day to day.
This is not an argument against AI adoption. It is an argument for mature adoption. One that recognises that augmenting clinical decision-making changes the nature of clinical work itself, and therefore requires a redefinition of what it means to be a competent practitioner.
If we fail to act, the risks will accumulate quietly, across thousands of encounters, and surface only when outcomes diverge: in missed diagnoses, in inequitable care, in loss of trust.
With the right training and governance, we can establish a model of AI-augmented care that is not only more efficient, but more reflective, more accountable, and more equitable than what came before.
Shoshana Bloom is Founder and Principal Consultant at Equiti Health Ltd, specialising in digital health transformation and AI governance.