Skip to content

One in five UK doctors use GenAI tools like ChatGPT and Gemini - but patient safety risks remain

Doctor in scrubs using laptop with stethoscope on desk and patient sitting on sofa in clinic room

One in five UK doctors are using a generative artificial intelligence (GenAI) tool - including OpenAI’s ChatGPT or Google’s Gemini - to support clinical practice, according to a recent survey of about 1,000 GPs.

Respondents said they use GenAI for tasks such as drafting paperwork after consultations, supporting clinical decision-making, and sharing information with patients - for example, producing clear discharge summaries and treatment plans.

Given the intense public attention on artificial intelligence, alongside the mounting pressures facing health systems, it is understandable that clinicians and policymakers view AI as central to efforts to modernise and reshape health services.

At the same time, GenAI is a relatively new development that raises fundamental patient safety questions. There is still a great deal we need to understand before GenAI can be used safely as part of routine clinical work.

UK doctors and GenAI in clinical practice

While some clinicians are already experimenting with GenAI, everyday healthcare is a high-stakes environment. Tools that generate text can influence records, decisions and patient understanding - all of which can affect care.

The issue is not simply whether GenAI is impressive or useful. The key question is whether it can be relied upon in real clinical settings, with real patients, and within the constraints and pressures of the health system.

The problems with GenAI

Historically, many AI systems in healthcare have been built to complete a tightly defined task. Deep learning neural networks, for instance, have been developed for classification in imaging and diagnostics. In breast cancer screening, these systems can analyse mammograms to support detection.

GenAI is different. It is not trained to carry out one narrowly specified job. Instead, it is built on so-called foundation models with broad, general capabilities - able to generate text, images (pixels), audio, or combinations of these.

Those general capabilities can then be refined for particular uses, such as answering questions, writing code, or generating images. In practice, what people can do with this kind of AI can seem limited only by the user’s imagination.

The critical point is that, because the technology has not been designed for a single context or a specific purpose, we do not actually know how clinicians can use it safely. That is one major reason GenAI is not yet appropriate for widespread deployment across healthcare.

A further challenge is the widely recognised issue of "hallucinations". In this context, hallucinations are outputs that are nonsensical or untrue, despite being produced from the provided prompt or input.

Hallucinations have been examined when GenAI is asked to summarise text. One study showed that a range of GenAI tools produced summaries that drew incorrect connections from what the text said, or included details that were not mentioned at all.

This happens because GenAI works by estimating what is most likely to come next - for example, predicting the next word given the surrounding words - rather than through "understanding" in a human sense. As a result, its responses can be convincing, but not necessarily accurate.

That persuasive quality is another reason it is premature to use GenAI safely in routine medical practice.

Consider a GenAI system that listens to a consultation and then drafts an electronic summary note. On the one hand, it could give a GP or nurse more time and attention for the patient. On the other, it might generate documentation based on what it believes to be plausible.

For example, the AI-generated summary could alter how often symptoms occur or how severe they are, introduce symptoms the patient never described, or add details neither the patient nor clinician discussed.

To prevent errors, clinicians would need to proofread AI-generated notes with extreme care, and also rely on an excellent memory to separate what was said from what merely sounds credible but is fabricated.

That might be manageable in a traditional general practice setting, particularly where a GP knows a patient well and can spot mistakes. However, in a fragmented health system where patients may be seen by multiple different healthcare professionals, inaccuracies in notes could create serious risks - including delays, inappropriate treatment, and misdiagnosis.

The potential harm linked to hallucinations is substantial. It is also important to acknowledge that researchers and developers are actively trying to reduce how often hallucinations occur.

Patient safety

Another reason GenAI is too early for broad healthcare use is that patient safety depends on how the technology performs in real-world interactions. This includes how people actually use it, how it aligns (or clashes) with rules and constraints, and how it fits within the wider culture, priorities and pressures of a health system. A systems perspective like this is essential for judging whether GenAI use is safe.

However, because GenAI is not built for one defined use, it is highly adaptable and can be applied in ways that cannot be fully anticipated. In addition, developers frequently update these systems, introducing new general capabilities that can change the behaviour of a GenAI application.

Harm can also arise even where the tool seems to function safely and as intended - once again, depending on the context in which it is used.

For instance, if GenAI conversational agents are introduced for triage, this could influence different patients’ willingness or ability to engage with healthcare. People with lower digital literacy, those whose first language is not English, and non-verbal patients may find GenAI difficult to use. So even if the technology can "work" in principle, it could still contribute to harm if it does not work equally well for all users.

The broader point is that these GenAI-related risks are harder to foresee in advance using traditional safety analysis methods, which tend to focus on how a specific technological failure might lead to harm in a defined setting. Healthcare could benefit enormously from GenAI and other AI tools.

But before they are adopted more widely, safety assurance and regulation will need to respond more quickly to changes in where and how these technologies are used.

It is also necessary for GenAI developers and regulators to collaborate with the communities using these systems, so that tools can be integrated into clinical practice in ways that are both regular and safe.

Mark Sujan, Chair in Safety Science, University of York

This article is republished from The Conversation under a Creative Commons license. Read the original article.

Comments

No comments yet. Be the first to comment!

Leave a Comment