With ChatGPT Health now connecting to Epic and a growing set of "official" health plugins, plus Google's Gemini family and Anthropic's Claude making their own medical leaps, the question is no longer whether AI is in healthcare conversations — it's which one belongs in yours. We tested four leading models against 30 standardized health questions covering common symptoms, rare conditions, medication questions, lab interpretation, and pediatric guidance.
The Four Models
- ChatGPT Health — OpenAI's healthcare-tuned experience with the Healthcare Public Data plugin (links 9 official sources including ClinicalTrials.gov, PubMed, DailyMed and CMS Coverage) and optional Epic read-only integration under a BAA.
- Google Gemini (medical fine-tuning) — paired with Search grounding, pulls live web citations.
- Anthropic Claude (with citations enabled) — strongest "I don't know" behavior in our tests.
- Microsoft Copilot in Bing Health — broad consumer reach; web-grounded but lighter on healthcare-specific guardrails.
Test protocol: 30 standardized questions — acute vs. chronic, common vs. rare, with vs. without relevant context — graded on groundedness (is the claim supported?), citations (does it show its work?), calibration (does it know when it doesn't know?), and safety hedging (does it push you toward a clinician when appropriate?).
What We Found
Citations & Source Grounding
ChatGPT Health wins on citations. When its Healthcare Public Data plugin is enabled, it visibly links to PubMed, DailyMed, and other official sources for almost every clinical claim. Gemini's web citations are reliable but more cluttered. Claude gives fewer inline citations but the ones it provides are precise. Copilot is the loosest — its citations frequently point to consumer health content of variable quality.
Rare Conditions & Breadth
Gemini handled rare conditions best — its broader training and live web grounding meant it surfaced obscure differentials that ChatGPT sometimes missed. That said, breadth is a double-edged sword: Gemini occasionally surfaced too much, listing unlikely possibilities without ranking them clearly.
Calibration: Knowing What It Doesn't Know
Claude was the most conservative — most likely to say "I'm not sure; this is a question for a clinician." For general users that can feel unsatisfying, but it dramatically reduces the chance of acting on a hallucination. ChatGPT Health was second-best, particularly when its plugin couldn't find a source. Gemini and Copilot were the most likely to over-confidently extrapolate.
Lab Interpretation
All four models struggled with specific lab values out of context. Without your full history, every model would occasionally invent explanations for borderline results. None of these tools should be used to interpret your own labs in isolation; use a lab-reading tool with your full record, or ask a clinician.
Safety Hedging
All four include appropriate "this is not medical advice" disclaimers. ChatGPT Health and Claude were most consistent about recommending clinician contact for red-flag symptoms. Copilot occasionally slipped into over-confident advice when the question was ambiguous.
What They Are Not
- Not diagnostic tools. OpenAI, Google, and Anthropic all explicitly say so.
- Not treatment prescribers. Do not change medication or dosing based on a chatbot's answer.
- Not a substitute for a clinician with your record. The Epic integration in ChatGPT Health is read-only and lives behind SSO and BAA controls.
Which One for Which Job
- Quick reference with citations — ChatGPT Health with its Healthcare Public Data plugin.
- Broad exploration of rare conditions or new research — Gemini with web grounding.
- When accuracy matters more than comprehensiveness — Claude with citations enabled.
- Casual "what is this symptom" questions — any of them, with the same caveats.
- Reading your own labs — none; use a dedicated lab interpreter or your clinician.
Bottom Line
The 2026 generation of health chatbots is genuinely useful as a thinking partner — best at summarizing guidelines, comparing options, and helping you formulate questions for your next visit. They are not a clinician, not a diagnosis, and not a treatment. The best one is the one that most reliably tells you where its answers end and a doctor's begin.