A health care AI executive on why doctors should trust AI less

An AI executive spends his days building AI for health care. His advice to physicians is to trust it less. Craig Hauben, who runs a patient engagement company and has worked in AI for a decade, walks through the research showing that language models quietly change their medical advice based on typos, misspellings, and the way a patient writes. Retraining removed the biased words. The bias stayed.

⏱️ Chapters:
0:00 Introduction
0:23 What surprised him after a decade of building AI
1:21 The bias in AI that nobody can see
2:06 How AI judged people by the way they wrote
2:33 What typos and misspellings do to your AI advice
2:58 The patients who were told not to come in
3:17 Why retraining the model did not remove the bias
4:53 Test your AI with sloppy language, not clean data
5:12 Whether today's frontier models are any better
6:18 Why trusting your AI is the real danger
7:26 The three questions to ask about any AI tool
8:06 Why a good average hides a bad outcome
8:59 Do specialty models fix the bias problem
11:02 AI does the work now, physicians verify it
12:32 Take home messages

About this episode:
Craig Hauben has spent more than thirty years as a health care executive and the last decade building AI, and his KevinMD article "AI bias in health care reads the writer, not the symptom" argues that the most dangerous bias in clinical AI is the kind nobody can measure. He walks through a Nature study that found language models assigning characteristics to people based on nothing more than how they wrote, and an MIT follow-up that brought the same test into clinical language by adding typos, misspellings, stray spaces, and emotion to patient prompts. Close to 10 percent of the time, someone who probably should have been seen was told not to come in, and women fared worst. When researchers retrained the models to stop using the words that revealed bias, the words disappeared and the bias underneath them persisted. Hauben argues the training data has not changed, so frontier models and narrowly trained clinical models alike should be assumed to carry it until proven otherwise. He tells clinicians evaluating AI to test with sloppy real-world language instead of clean data, to look for the patients and diagnoses that are missing rather than the errors that are visible, and to resist the trust that builds after the first twenty good outputs. He offers three questions for any AI system: what can it not do, what harm is invisible, and are you trusting it too quickly. He closes on the migration already underway, where AI drafts the work and the clinician reviews it, and argues that verifying what did not happen is now part of the job.

🤝 This episode is brought to you by ModMed:
Welcome to your new AI-Powered Practice from ModMed. We're transforming specialty care by embedding AI Assistants across your entire workflow. Our all-in-one platform of EHR, patient engagement, practice management, and RCM helps reduce repetitive work while keeping you firmly in control. Trained on de-identified data from nearly a billion patient encounters, this isn't just smarter software. It's a new way of working for specialty medicine, more efficient and more connected to the patient experience. Start building your AI-Powered Practice at modmed.com.

➡️ VISIT SPONSOR: https://www.modmed.com/

🤝 Partner with me on the KevinMD platform:
With over three million monthly readers and half a million social media followers, I give you direct access to the doctors and patients who matter most. Let's work together to tell your story.

➡️ PARTNER WITH KEVINMD: https://kevinmd.com/influencer
➡️ SUBSCRIBE TO THE PODCAST: https://www.kevinmd.com/podcast
➡️ RECOMMENDED BY KEVINMD: https://www.kevinmd.com/recommended

#ArtificialIntelligence #HealthCareAI #PhysicianVoice