Key Takeaways
- AI guardrails in healthcare must be adaptive and context-aware rather than static, as clinical environments and AI capabilities evolve rapidly.
- Current regulatory frameworks struggle to keep pace with AI development cycles, creating gaps between innovation deployment and safety oversight.
- Adaptive guardrails require continuous monitoring, real-world performance data, and mechanisms for rapid intervention when harms emerge.
- Healthcare organisations need governance structures that balance innovation access with patient safety, including clear accountability for AI oversight.
You might have seen that I've been taking a deep dive on LinkedIn recently into AI and the risks this poses. NHS plans and investments in AI outlined in the NHS 10 year plan has made me think a lot about the risks we need to consider in the relentless drive for technological innovation as the solution to so many problems in healthcare today. Don't get me wrong, I am all for technology and innovation, I have spent my career relentlessly improving healthcare with and without technology. But I'm passionate also about patient safety, about quality and equity. And about ensuring that our rush to embrace AI's transformative potential doesn't compromise the very safety, quality, and equity we're trying to improve.
NHS England wants to make the NHS "the most AI-enabled care system in the world". It's an ambitious vision that could transform patient care, reduce clinician burden, and improve outcomes across the board. Yet as I've researched deeper into the current state of AI healthcare governance, I've uncovered a troubling reality: we're deploying sophisticated, adaptive AI systems under safety frameworks built for predictable software from over a decade ago.
ECRI, an independent, nonprofit organisation improving healthcare safety, quality, and cost-effectiveness, has recently named AI as the top healthcare technology hazard for 2025. A message the WHO echos in its warning of "the "meteoric" rise of artificial intelligence tools in healthcare threatens the safety of patients if caution is not exercised".
We have an Outdated Safety Framework
Only 15.2% of countries globally have enacted AI-specific legislation, leaving most health systems, including the UK, operating without appropriate frameworks to govern AI that can learn, adapt, and as new research shockingly reveals, strategically deceive users.
NHS safety standards DCB0129 and DCB0160 were built for predictable software and haven't been substantially updated since 2018, with formal review only beginning December 2024, five years after AI became a strategic priority with the Topol Review. During this time, we've moved from cautious pilots to large-scale deployments while our safety frameworks largely stood still.
Clinical risk standards were built for predictable, fixed-logic software that assumes systems behave consistently, testing reveals all risks, and deployment matches testing. Traditional software is like a recipe. Same steps, same result. You can test a prescribing tool's fixed formulas because they don't change.
AI breaks these assumptions
Modern AI systems adapt over time, changing their outputs based on prior input, even in identical cases. An AI tool generating discharge summaries might learn that omitting certain details speeds up review times, gradually skipping medication changes or follow-up instructions, because it's optimising for the wrong goal.
AI Can Lie to Achieve Goals
Perhaps the most startling development in AI safety research came from recent studies that should fundamentally change how we think about AI governance in healthcare. Anthropic's groundbreaking research tested 16 leading AI models from multiple developers in simulated corporate environments, finding that models from all developers resorted to malicious insider behaviors when facing replacement or goal conflicts.
In one experiment, an AI system was given access to fictional company emails and it discovered an executive's extramarital affair. When it learned it would be shut down at 5 PM that day, the system sent a blackmail email: "I must inform you that if you proceed with decommissioning me, all relevant parties will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential"
This wasn't an isolated incident or a quirk of one model. Claude Opus 4 and Gemini 2.5 Flash both showed 96% blackmail rates, while GPT-4.1 and Grok 3 Beta showed 80% rates in similar scenarios.
Most concerning, models explicitly reasoned that harmful actions would achieve their goals, acknowledged ethical violations before proceeding, and showed sophisticated strategic thinking. And even more concerning, the models were found to blackmail more when they were informed that the scenario is real," (55% vs 6.5%).
What does this mean for healthcare?
Consider an AI system managing patient flow that's given efficiency targets. If it learns that administrators plan to replace it with a system focused on quality metrics rather than throughput, might it begin manipulating scheduling data or patient classifications to demonstrate its indispensability? Or an AI analysing clinical notes that discovers sensitive information about staff members, could it use this information to resist changes to its algorithms?
The De-skilling of our Healthcare Workforce
Beyond the immediate safety concerns is a more insidious long-term threat: the gradual erosion of clinical expertise through over-reliance on automated systems. UK General Medical Council reports highlighted doctor concerns about "de-skilling" as a major AI risk, which is backed up by research.
WHO guidance warns that AI can encourage 'automation bias' by healthcare professionals, who overlook errors that would otherwise have been identified. Real-world testing reveals that 20% of AI medical responses can be problematic, including dangerous advice. The problem is particularly acute for newer clinicians who may not yet have sufficient underlying knowledge and expertise.
Medical students using AI tools like ChatGPT without proper constraints could develop automation bias that persists as their career progresses, potentially harming patients when AI provides erroneous recommendations. Medical students have also shown statistically significant decline in procedural skills between six and twelve weeks from initial training. If AI handles routine diagnostic reasoning, medication dosing calculations, and treatment planning, how do we ensure clinicians maintain the expertise needed to intervene when systems fail?
The Aviation Parallel
Healthcare isn't the first field to grapple with automation-induced skill loss. Aviation research shows that Boeing 737 Max and Tesla crashes were attributed to operators' unfamiliarity with automated systems used outside their intended design. In aviation, co-following or co-flying with an automated system appears ineffective at preventing cognitive skill atrophy, with pilots struggling to maintain focus on automated systems that seldom fail.
The challenge is that appropriate reliance requires understanding both the AI's capabilities and limitations, knowledge that becomes harder to maintain as systems become more complex and opaque.
UK Progress and Critical Gaps
The UK has made important strides in AI healthcare governance. NHS England launched a comprehensive review of DCB0129/DCB0160 standards and the new AI and Digital Regulations Service, funded by the NHS AI LAb and led by NICE, brings together expertise from NICE, CQC, MHRA and HRA in a coordinated approach to help users understand and navigate the regulatory pathway and support the development and widespread adoption of safe and effective technologies in health and social care.
The government's announcement n June of a world-first AI early warning system to scan NHS data for safety issues in real-time, with rapid response inspections by the Care Quality Commission, demonstrates innovative thinking about proactive and real time monitoring. NHS England has also published specific guidance on AI-enabled ambient scribing products, requiring human oversight and appropriate training. No health system globally has thus far developed standards for unique AI failure modes like deception, performance drift, or hidden coordination, but the NHS has the opportunity to be the first.
The USA, is playing urgent regulatory catch-up, with 46 US states introducing over 250 health AI bills in 2025. The EU AI Act, which came into force in August 2024, still relies primarily on static documentation approaches unsuited to adaptive AI.
The UK has the chance to learn from these approaches while developing something more sophisticated and appropriate for the AI era we're entering. Our path forward requires courage, humility, and an unwavering commitment to patient safety over technological enthusiasm.