I always assumed that if an AI told me why it thought the mole on my arm was fine, that would make it safer to trust, not more dangerous. Show me the reasoning, highlight the spot on the photo it’s worried about, walk me through it in plain English, and I come out a smarter patient, right? That’s the whole pitch behind “explainable AI” in medicine: crack open the black box, and ordinary people get to borrow a specialist’s judgment.

A study published August 4 in Nature Medicine took that assumption into a lab and did something uncomfortable to it. A team led by Chanwoo (Orson) Xu at Columbia, with Marzyeh Ghassemi’s group at MIT and the Stanford dermatologist Roxana Daneshjou, ran two large experiments: 623 lay people and 153 primary care physicians, 776 participants in all, diagnosing skin conditions with an AI looking over their shoulder. They tested four ways of showing its work. A bare confidence score. A set of similar reference images. A heat map lighting up the suspicious region. And a large language model explaining its reasoning in warm, fluent prose.

And here is where my tidy assumption fell apart. The explanations that felt the most helpful were the ones that did the most damage to the people who could least afford it.

For the lay volunteers, the AI was a crutch when it was right and a trapdoor when it was wrong. Their accuracy climbed when the model was correct and sank when it erred, the signature of automation bias: people changed their answers because the machine sounded sure, not because it had become any more right. The physicians did almost the opposite. They stayed resilient to the AI’s bad advice, holding their own whether the model was right or wrong, and they actually did best when handed only the model’s raw prediction with no explanation attached. The story the AI told about itself didn’t help them. It got in the way.

The plain-language explanation, the friendliest and most human-sounding feature, produced the strongest pull toward deference in non-experts. It didn’t just make them wrong more often when the AI was wrong. It left them more confident about their wrong answers.


So the question I could not stop chewing on was why a fluent, reasonable-sounding explanation would make a novice more wrong instead of less. Once you sit with it, the mechanism is almost obvious in a way that stings. A well-written rationale doesn’t add knowledge to someone who has none to check it against. It fills the vacuum. If you’re a dermatologist, a slick explanation runs straight into decades of pattern recognition, and the parts that don’t fit throw off a little internal hang on, that’s not how that lesion behaves. If you’re the rest of us, there’s nothing for it to catch on. The confident narrative pours into the empty space and sets. The better the story, the more completely it replaces the judgment you never had.

The timing mattered too. Showing the AI’s suggestion first, before a person formed their own read, made them more deferential to the model, and that pull hit both groups. Anchor someone to a machine’s answer before they’ve committed to their own, and the table is already tilted.

Now sit with what “explainable AI” is actually sold as. The transparency features are pitched as the safety layer, the thing that turns an opaque algorithm into a trustworthy assistant for patients, for symptom-checker apps, for the direct-to-consumer skin-scan startups that would love to put a diagnostic AI in your phone. Earlier work leaned optimistic; a 2025 Nature Communications eye-tracking study found dermatologist-like explanations boosted melanoma diagnosis and trust. This new work says the empowerment story quietly inverts for exactly the population the pitch targets. “The same tool ends up being an asset for one user and a liability for another,” Xu told MedicalXpress. Daneshjou put the stakes plainly: those with the least medical knowledge are most likely to be led astray when explainable AI models give an erroneous output. That isn’t something you patch with a better prompt. It’s how trust behaves in a brain that has nothing to push back with.

I want to be fair to what actually works here, because it matters. The model was built fairness-constrained, balanced across skin tones, and that part delivered: it improved accuracy and shrank the diagnostic gap between light and dark skin, the exact disparity that has dogged dermatology AI for years. The diagnosis engine may be useful. The risk is the storytelling wrapped around it, and who is most defenseless against a good story.

Yes, this was a controlled experiment, not a scared person at 11pm reading their phone, and lab tasks aren’t clinics. But the behavioral signal is the one that scales, because the real world is millions of non-experts getting a confident, articulate answer from a machine with no expert standing beside them to feel the friction. And the money here didn’t come from a device maker: the funding was the National Science Foundation, Schmidt Sciences, the National Bureau of Economic Research, and Columbia, which is worth sitting with when the finding cuts against the commercial hype instead of for it.

So here’s my own decision, and it surprised me to land on it. If I ever run one of these tools on my own skin, I want the bare answer and the confidence number, not the eloquent paragraph talking me into it. The explanation isn’t the safety feature. For someone like me, it’s the part I’d trust the least.

Sources

  1. Nature Medicine – Xu et al., “Divergent impacts of explainable AI for dermatological diagnosis on clinicians versus lay people” (2026)
  2. MIT News – “The benefits of medical AI assistance vary based on user expertise” (Aug 4, 2026)
  3. arXiv preprint – “Explainable AI as a Double-Edged Sword in Dermatology: The Impact on Clinicians versus The Public”
  4. MedicalXpress – “Medical AI benefits vary by expertise, with nonexperts more easily misled” (2026)
  5. Nature Communications – “Dermatologist-like explainable AI enhances melanoma diagnosis accuracy: eye-tracking study” (2025)