Digital Health Equity

Designing for AI Health Equity: A System Responsibility

The dominant consideration of AI risk in healthcare still centres on data quality, as though representative datasets are sufficient to produce equitable outcomes. They are not. Risk stems from how datasets are constructed, how models are validated, how systems are governed, and how clinicians interact with them.

AUTHOR

Shoshana Bloom

PUBLISHED

February 23, 2026

ABOUT THE AUTHOR

Shoshana Bloom

Shoshana Bloom is Founder of Equiti Health and specialises in digital transformation, healthcare innovation, service redesign, and digital health equity.

View full biography →

PUBLISHED

February 23, 2026

Key Takeaways

  • AI health equity cannot be achieved through technical fixes alone—it requires system-level commitment spanning procurement, implementation, monitoring, and governance.
  • Equity considerations must be embedded from the earliest design stages, not added as an afterthought once models are trained and deployed.
  • Organisations need explicit accountability for equity outcomes, with named owners and regular reporting on stratified performance metrics.
  • System leaders must create the conditions for equity-focused AI development through procurement requirements, validation standards, and ongoing surveillance.

NHS England continues to expand AI-enabled care across diagnostics, clinical documentation, and screening. An AI Research Screening Platform is being developed for large-scale evaluation of AI tools across national screening services, which will deliver a step-change in the scope of algorithmic decision-making within population health. Specific guidance on AI-enabled ambient scribing has been issued to manage safety and governance expectations as these tools increasingly enter clinical environments.

However, the dominant consideration of AI risk in healthcare still centres on data quality, as though representative datasets are sufficient to produce equitable outcomes. They are not. Risk in AI tools stems from a multitude of factors which include how datasets are constructed, how models are validated, how systems are governed, and how clinicians interact with them.

In the NHS, every one of these factors sits within system design, which means addressing this is not simply a technical problem for developers or a regulatory matter for national bodies.

The prevailing approach treats bias, safety, workforce capability, and clinical risk as separate issues. When these risks compound, and the evidence shows they do, isolated and fragmented interventions are insufficient.

The Governance Gap

The regulatory landscape is slowly catching up. In September 2025, the MHRA launched the National Commission into the Regulation of AI in Healthcare, with a mandate to develop a framework covering safety, liability, and responsible innovation. This signals a recognition that the current patchwork of standards and voluntary frameworks in use across the NHS is increasingly inadequate for the scale of deployment that is now underway.

Deployment and regulation are advancing simultaneously. But neither addresses the governance layer that sits between them: the NHS organisational responsibility for commissioning, overseeing, and assuring these tools are safe, in practice. At present, there is limited evidence that organisational-level governance maturity is keeping pace with the accelerated use of AI.

Bias is not theoretical

We have well evidenced examples of harm resulting from AI in healthcare. Obermeyer et al. (2019) showed that a US population health management algorithm allocated fewer resources to Black patients with equivalent clinical need, because healthcare cost was used as a proxy for illness severity. Because historical spending was lower for Black patients — a product of access barriers, not lower morbidity — the algorithm reproduced structural inequity in automated allocation. The developers did not set out to discriminate. The bias was an emergent property of the system: both the training data and the absence of adequate validation combined to produce inequity at scale.

Seyyed-Kalantari et al. (2021) demonstrated that chest X-ray algorithms exhibited underdiagnosis bias in underserved patient populations. The models showed systematically lower sensitivity for certain demographic groups, meaning pathology was more likely to be missed in those already facing barriers to care. Adamson and Smith (2018) similarly reported reduced diagnostic accuracy of dermatology AI systems in darker skin types, reflecting the underrepresentation of these skin tones in training datasets.

These studies show that performance disparities are not hypothetical. They are measurable and clinically significant.

If an organisation deploys an AI tool without requiring subgroup performance data, it has no basis for assessing whether that tool will perform equitably and universally safely across its population.

Bias does not operate in isolation

Bias is routinely treated as a standalone data problem. It is not. AI systems function within a broader clinical risk environment, and the risks compound.

Five intersecting domains shape AI-enabled care:

  1. Clinical deskilling
  2. Automation bias
  3. Clinician burnout
  4. Health equity failures
  5. Safety deficits

These are not parallel risks. They interact, and their interaction is what makes AI a system problem rather than a technical one.

Clinical deskilling

As clinicians rely more on AI, their independent diagnostic capability can decline. When that happens, they are less able to detect model error. If the model underperforms for particular subgroups, those errors are more likely to pass unchallenged, embedding inequity into routine care. Budzyń et al. (2025) found measurable deterioration in unassisted adenoma detection rates among clinicians routinely using AI-assisted colonoscopy. Natali et al. (2025) describe broader AI-induced deskilling across diagnostic reasoning and clinical examination.

Automation bias transmits that failure into practice.

Even when a clinician initially makes a correct clinical judgement, exposure to AI output can cause them to revise that judgement in favour of the algorithm. The authority of the model overrides independent reasoning. If the model performs poorly for certain populations, automation bias ensures that this poor performance is carried through into clinical decisions. Kücking et al. (2024) showed that clinicians overturned their correct initial judgements after exposure to incorrect AI advice, with higher susceptibility among non-specialists — precisely the clinicians most likely to encounter unfamiliar clinical scenarios where algorithmic support will be most valuable.

Clinical burnout compounds this.

Clinicians operating under sustained cognitive load are less likely to interrogate algorithmic output, less likely to notice performance anomalies, and less likely to exercise the independent judgement required to override an incorrect recommendation, particularly when the tool was introduced to reduce their burden.

Research shows that sustained cognitive load and burnout impair decision-making performance and reduce clinicians' capacity for deep critical analysis (West et al., 2018; Deligkaris et al., 2014). High mental workload shifts cognitive processing toward shortcuts and reduced analytical scrutiny (Young et al., 2015). This increases reliance on automated outputs and reduces the likelihood they will cognitively interrogate decision support systems (Goddard et al., 2012; Kücking et al., 2024). When cognitive resources are depleted, clinicians are more likely to trust and defer to algorithmic recommendations rather than override them, increasing the risk that model errors and performance failures go unnoticed and unchallenged.

Safety governance is already inadequate for the tools we already have in use.

Oskrochi et al. (2025) found that over 70% of digital health technologies deployed in NHS trusts lacked documented clinical safety assurance despite statutory requirements under DCB0129 and DCB0160. AI systems compound this: their adaptive, opaque behaviour challenges static, point-in-time safety models. A safety case produced at implementation will not reflect the tool's behaviour twelve months later, particularly if the population it serves differs from the population on which it was validated.

The clinical workforce is underprepared

Equity safeguards depend on clinician capability and capacity. Yet over 75% of medical students report no formal AI education. Research and policy literature predominantly treat deskilling, bias, burnout, and safety as separate concerns rather than components of an integrated risk system. Clinicians using AI will therefore have an incomplete map of the risk environment they are working within.

Governance frameworks cannot compensate for workforce deficits. Clinicians need defined competencies that integrate technical interpretation, critical thinking, equity awareness, and safety governance. This requires moving beyond standalone AI literacy modules toward embedded, career-stage-appropriate capability strengthening.

What this requires from NHS organisations

Achieving safe and equitable use of AI requires system controls across four domains:

Procurement

  • Mandate demographic transparency in training and validation data
  • Require subgroup performance reporting aligned to local population characteristics
  • Require independent validation in representative populations beyond the developer's original dataset
  • Make equity and safety evidence a contractual condition

Pathway governance

  • Define clear override and escalation mechanisms
  • Monitor subgroup performance and performance drift over time
  • Audit unassisted clinical performance where AI augments care

Workforce capability

  • Embed defined AI competencies across career stages
  • Train explicitly on automation bias and equity implications
  • Integrate AI governance responsibilities into appraisal and leadership roles

Safety assurance

  • Require documented DCB0129 and DCB0160 compliance as a baseline
  • Integrate AI oversight into existing clinical risk systems
  • Align local governance with emerging MHRA regulatory standards

A System Accountability Test

If we are going to address these issues, we will need to be able to answer the following questions:

  • Which populations does the system's AI tools underperform for, and how is that evidenced?
  • How is patient subgroup performance measured and monitored, so that inequity can be identified early and addressed before it becomes embedded in pathways?
  • What thresholds trigger review or suspension, and who is accountable for acting on them?
  • How is safety assurance maintained across the full operational lifecycle of each tool?
  • What happens when inequitable performance is identified?
  • How is performance data routinely reported to leadership and governing bodies?

If an organisation cannot answer these questions, it does not have an AI equity strategy and does not understand the risk profile of the tools it is deploying. The distinction matters. Deployment without equity governance does not produce neutral outcomes — it reproduces existing inequity at algorithmic speed and scale. That is not innovation. It is unmanaged risk.

AI will shape access, diagnosis, triage, and resource allocation across the NHS. The question is not whether AI is adopted, but whether it is governed with sufficient rigour to prevent existing inequities from being amplified at scale. Health equity in the age of AI will be determined by the quality of local system design, oversight, and leadership.

More from this category