When the Scribe Hallucinates — The A-Eye Series, Blink
Blink

When the Scribe Hallucinates

An AI scribe invented a patient's drug history, and nobody caught it

In our first blog we introduced the three groups of AI and described how they each fail. You may recall that generative AI is one of the hardest to spot when it goes wrong. It can write something that reads perfectly, but is completely fabricated. When an AI does this, we call it a hallucination. It's a term you may have already heard and will see a lot in upcoming posts. Let's look at a recent example of this in the news. It involves an AI technology already available to us, an AI scribe.

An Australian woman went in for a routine appointment with her urologist. She agreed to let an AI scribe record and summarise the consultation. Weeks later she read her own post-operative letter and was shocked to read that her kidney bleeding was attributed to psychedelic mushroom use. The claim was completely untrue, hallucinated by the AI, and no human had caught the error. She had never used psychedelic mushrooms.

It's not an isolated case either. There have been many documented cases of scribes making serious errors. One study of AI generated clinical notes found hallucinations in roughly 1.5 percent of sentences, and almost half of those were serious enough to affect a patient's diagnosis or treatment. There have even been reports of scribes documenting examinations that never occurred.

There's also evidence that in some instances the speech recognition underneath these tools performs less accurately on some races compared to others. That's a bigger issue with training data, and one we'll come back to properly in a future post.

Scribes do appear to be getting more accurate but the concerning thing is, clinicians who check carefully in the first few weeks tend to stop once they assume the tool is reliable. So when the tool gets better, our attention gets worse. A scribe that's wrong once a week keeps you alert. One that's wrong once a month is the real threat.

The Australian woman introduced above described the whole argument for augmented intelligence quite well. She said, "I love AI, but the essential ingredient is the human." It's another way of saying we need a human in the loop.

Imagine if your AI scribe records the left eye when you said right. Or logs a low IOP instead of a high one. Or writes choroidal melanoma when you said naevus. Every one of those notes needs a human reading it before it goes anywhere. Not sometimes. Every time.

// Sources

// stay in the loop

Enjoyed this? Get the next one.

New posts monthly. No spam.

You are on the list!

No spam. Just new posts when they drop.