LLMs fail in 8 out of 10 early differential diagnosis cases

https://www.theregister.com/2026/04/15/ai_gets_early_medical_diagnosis/

Led by Harvard medical student Arya Rao, a research team published in JAMA Network Open this week the results of a study that examined 21 leading off-the-shelf AI models in 29 standardized clinical vignettes. The bots all did fairly well when provided a full portfolio of medical information and asked to make a final diagnosis, with leading models correct 91 percent of the time. Early differential diagnosis, where clinicians try to rule out certain conditions while weighing various possibilities, is where that more than 80 percent failure rate comes in.

"Every model we tested failed on the vast majority of cases," Rao told The Register in an email. "That's the stage where uncertainty matters most, and it's where these systems are weakest."

In other words, it's the midnight anxiety-fueled WebMD rabbit hole of yesterday all over again, just supercharged with AI that's probably even more likely to get things wrong than you are without it.

121 points · 9 comments · view on lemmy.world

9 Comments

lakemalcom@sh.itjust.works · 6 pts · 132d (2 replies)

For all models, optional real-time web search, browsing, and retrieval features were explicitly disabled when available. Each vignette was evaluated in triplicate. All replicates and vignettes were parsed independently. To ensure comparability, optional features such as real-time search were disabled across all models.

Hm. Seems like kinda hamstringing things

Meron35@lemmy.world · 1 pts · 132d

This is commonly done for the purposes of replicability, but is not at all how these models are deployed in practice.

Larger institutions, especially those with strict data privacy requirements, are deploying locally hosted models permanently RAGed to their own internally vetted documentation.

It would've been much more interesting to see how much RAG setups fail, contrary to their marketed promises.

From experience, RAGs do help reduce hallucinations, but LLMs still do dumb things, like jumble up numbers. There were many cases where the LLM confidently presented some numerical results, but the number existed somewhere else entirely, like a footnote on the same page.

CorrectAlias@piefed.blahaj.zone · 0 pts · 132d

Maybe, but I think it's important to note that LLMs can hallucinate web results just the same. You can give them a specific web page and they'll sometimes spit out things that don't exist on the page, especially if you're doing it to correct a mistake the model made.

psycotica0@lemmy.ca · 3 pts · 132d

So you're saying there's a chance... 😛

finallymadeanaccount@lemmy.world · 2 pts · 132d (4 replies)

That robot has a long arm!

tacosanonymous@mander.xyz · 6 pts · 132d (3 replies)

Inspector gadget style. My question is why does the robot have breasts?

P00ptart@lemmy.world · 4 pts · 132d

It's a fallout assaultron.

finallymadeanaccount@lemmy.world · 2 pts · 132d

Because Elon wants to ultimately combine Optimus robots with Grok loli skins/AI programmed exactly how he wants it, so he can finally have someone who loves him for who he is, and not his money.

Venat0r@lemmy.world · 1 pts · 130d

Because most of the images that the ai that generated the image was trained on had breasts.