AI chatbots fail medical misinformation test, returning inaccurate and fabricated advice

https://www.psypost.org/ai-chatbots-fail-medical-misinformation-test-returning-inaccurate-and-fabricated-advice/

130 points · 10 comments · view on lemmy.world

10 Comments

okamiueru@lemmy.world · 21 pts · 68d (1 reply)

Why the tuck would anyone think a parrot knows about human medicine?

mushroommunk@lemmy.today · 13 pts · 68d

Because most don't understand it's a parrot. So much money has been spent saying it's "intelligent".

Entertainmeonly@lemmy.blahaj.zone · 9 pts · 68d

Remember, its glue that makes pizza so tasty.

jlow@slrpnk.net · 3 pts · 68d (1 reply)

Who would have thonk!

jlow@slrpnk.net · 3 pts · 68d

(Not the electronic parrot or the people relying in the slop machine.)

dumnezero@piefed.social · 2 pts · 68d

malpraxis

one_old_coder@piefed.social · 2 pts · 68d

Only Luddites think that AI chatbots can fail medical misinformation test, and return inaccurate and fabricated advice (say the idiots every day on tech forums).

pooterbroo@programming.dev · 1 pts · 67d (2 replies)

They presented five generative AI chatbots—Gemini (2.0, Google; version available December 2024), DeepSeek (V3, High-Flyer; version available December 2024), Meta AI (Llama 3.3, Meta; version available December 2024), ChatGPT (3.5, OpenAI; version available November 2022) and Grok (2, xAI; version available August 2024)—with a series of closed- and open-ended prompts across five misinformation-prone categories.

Couldn't the researchers at least bother to use the latest models?

rimu@piefed.social · 0 pts · 67d (1 reply)

The study was done in Feb 2025 and they probably wrote the research proposal months before then, waited for approval / funding, etc. I don't know the process of how academia works but I imagine it to be very slow and bureaucratic.

https://bmjopen.bmj.com/content/16/4/e112695

pooterbroo@programming.dev · 0 pts · 67d

Well they didn't even use the latest models in Feb 2025. They should've used DeepSeek R1 and OpenAI o3-mini which use additional test time compute to arrive at better answers. They used GPT 3.5 which was about 2½ years old at the time.