Observed Signal · Apr 21, 2026 · Research Publication · Source: Gary Marcus · Impact: 2/5 · Sentiment: Neutral
Studies Find Chatbots Unreliable for Medical Advice
A Substack post by Gary Marcus synthesizes recent peer‑reviewed research showing that current large language model (LLM) chatbots perform poorly and pose safety risks when used for medical advice. A BMJ audit of five popular chatbots (Gemini, DeepSeek, Meta AI, ChatGPT and Grok) found nearly half of responses to 10 medical prompts were highly problematic, with hallucinations and fabricated citations. A JAMA Network Open study of 21 models across 29 clinical questions concluded LLMs remain limited for early diagnostic reasoning and unsuitable for unsupervised patient‑facing decision‑making. Two Nature Medicine studies reported that LLMs identified relevant conditions in under 34.5% of cases and that ChatGPT undertriaged 52% of gold‑standard emergencies. Marcus warns that converging evidence across journals indicates consumers should not trust chatbots for medical decisions.
Multiple peer‑reviewed studies across high‑profile medical journals report converging evidence that LLM chatbots produce unsafe medical outputs; relevant for companies deploying consumer‑facing conversational AI but not immediately industry‑shifting for core AdTech.
Track DeepSeek Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- BMJ peer‑reviewed audit evaluated five chatbots (Gemini, DeepSeek, Meta AI, ChatGPT, Grok) using 10 medical prompts and found nearly 50% of responses were highly problematic.
- BMJ audit reported chatbot outputs were consistently expressed with confidence and contained hallucinations and fabricated citations.
- JAMA Network Open study assessed 21 frontier models on 29 clinical reasoning questions and concluded current LLMs are limited for early diagnostic reasoning and not reliable for unsupervised patient‑facing clinical decision‑making.
- Two Nature Medicine studies found LLMs identified relevant conditions in fewer than 34.5% of cases and that ChatGPT undertriaged 52% of gold‑standard emergency cases.
Connected Companies & Entities
4 Entities mappedRelated Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Researcher Made Chatbots Warn About Fake Disease
Swedish researcher Almira Osmanovic Lindström created a fictional eye disease called “Bixonimanie” and seeded false information via Medium posts and fake preprint studies to demonstrate how AI systems acquire and propagate knowledge. Large language model–based chatbots — including Microsoft’s Copilot, Perplexity AI and OpenAI’s ChatGPT — subsequently hallucinated details about the invented disease and warned users as if it were real. Some academic works even cited the fabricated studies; a paper in Cureus published in November 2024 was retracted at the end of March 2026. After Nature reported on the experiment, the Medium posts and the preprint entries were removed. Experts have raised alarms about how easily misinformation can enter LLM training and downstream conversational interfaces.
Study: Chatbots Weaken Misinformation Detection
An MIT Media Lab study found that relying on AI systems for fact‑checking over the course of a month reduces users’ independent ability to detect misinformation once the chatbot is unavailable. In a four‑week experiment with 67 participants, AI assistance improved misinformation detection by 21%, but when AI was removed performance in week four fell 15 percentage points below baseline; roughly one quarter of participants believed they had improved despite performing worse. The article also cites a separate review of 22 public broadcasters’ tests of ChatGPT, Microsoft Copilot, Google Gemini and Perplexity AI that found nearly half of AI responses had at least one significant issue (31% had major citation problems; 20% contained serious factual errors). Authors warn that conversational styles that narrate answers can create dependency, while socratic questioning may better support learning. The study notes sample limitations and plans broader follow-ups.
ChatGPT Health Fails Evaluation; Anchoring Bias Skews Triage
An independent evaluation found significant safety and reasoning failures in OpenAI’s ChatGPT Health. Although OpenAI developed the system with more than 260 physicians, over 600,000 clinician feedback rounds, and a custom safety framework, the study observed critical errors: the model’s internal reasoning identified early respiratory failure but the final recommendation advised waiting and scheduling an appointment; among cases three independent physicians labeled unanimous emergencies, the system steered patients away from the ER 52% of the time. Suicide-crisis safeguards triggered more on vague distress than on specific plans. A single dismissive family-member sentence shifted triage away from emergency care with an odds ratio of 11.7. The author argues these structural failure modes are general properties of LLM agents and describes a factorial evaluation approach and layered countermeasures to detect and mitigate them.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
