In June, a paper in Nature Medicine handed the artificial-intelligence industry the headline it had been waiting for. The general-purpose chatbots, the same ChatGPT-class models anyone can rent by the token, had beaten the specialized clinical tools that doctors actually pay for, on the doctors’ own turf. A team from NYU Langone Health and the University of Texas at Austin ran three frontier models, GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6, against two products physicians lean on at the bedside: OpenEvidence, the Miami startup marketed as “ChatGPT for doctors,” and Wolters Kluwer’s UpToDate Expert AI. The generalists swept. The framing traveled fast, because it flattered the people with the most money riding on it.
That money is worth naming. OpenEvidence had by then climbed to a $12 billion valuation, up from $1 billion a year earlier, on the strength of one number its chief executive likes to repeat: roughly 40 percent of physicians in the United States now use the tool. A paper suggesting that the whole category of purpose-built clinical AI might be redundant, that a doctor could get the same answer from a raw frontier model, is not a small academic result. It is a threat to one business and a validation of a much larger one.
Then, on September 3, the same journal published a correspondence under a plain title: limited benchmarks constrain the conclusions of the comparison. Beside it ran the authors’ reply.
The dispute lands on the two benchmarks that produced the study’s widest margins, and the case against them does not require anyone’s reply. It is visible in what the two tests are.
The first is MedQA, a set of USMLE-style exam questions that has been sitting on the open internet for years. Which means the frontier models were very likely trained on the answer key before they ever sat the test. On MedQA, Gemini scored 97.4 percent and GPT-5.2 94.2 percent, against 89.6 for OpenEvidence and 88.4 for UpToDate. Note how close that actually is: on a memorizable exam, the tools doctors pay for trail the free models by single digits.
The second benchmark is where the blowout lives, and it is the one with the deeper problem. On HealthBench, GPT-5.2 scored 88 out of 100 while OpenEvidence landed at 62.6 and UpToDate at 61.3. HealthBench was not built by a neutral party. It was built by OpenAI, the maker of GPT-5.2. Asking whether a general-purpose chatbot is well aligned with clinical judgment, using a yardstick the chatbot’s own manufacturer designed, is not a comparison. It is a home game.
That leaves the one test the researchers built for the actual question. They assembled a set of real clinical queries, questions physicians had typed to a general model in live practice, and had clinicians grade the answers blind. On those, per the study’s own reporting, the specialized tools performed comparably to Google’s AI Overview, the summary box that appears above ordinary search results at no charge and with no clinical pedigree at all. The exam and the OpenAI test delivered a rout. The test built for the real question did not.
Set those two results side by side. A $12 billion product used by two of every five American doctors, measured on the cleanest test in its own paper, landed in a tie with the thing that shows up free when you Google a symptom. That is a genuinely unflattering number for the specialized vendors. It also rests on a small, bespoke set of questions, which is a thin reed for anyone, generalist or specialist, to stand a verdict on. The reasonable reading is not that the doctors’ tools are worthless and the chatbots have won. It is that a field crowned a winner off two suspect rounds and one small clean one, and everyone repeated the score.
None of which makes the specialized vendors the good guys. OpenEvidence and UpToDate are closed commercial products; neither exposes the base model or training pipeline behind the answer being sold to 40 percent of the profession, so an outside evaluator cannot fully test the marketing against the machine. A benchmark a company builds its valuation on and a benchmark a company builds to score itself are the same species of problem in different clothes. The June paper simply happened to expose one of them.
What changed on September 3 is smaller than the headline and more useful. A claim that had already hardened into common knowledge, that the generalists have overtaken the specialists in medicine, got downgraded in the pages of the journal that launched it, to something closer to “on a small set of questions, maybe, and on the exams, ask whoever wrote them.” OpenEvidence will keep citing the June result, because 40 percent of American doctors is a number you defend, not one you revisit. The correction is on the record now, in the same journal, where corrections go to be filed and not read.
Sources
- Nature Medicine – “General-purpose large language models outperform specialized clinical AI tools on medical benchmarks” (original study, June 2026)
- Nature Medicine – “Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison” (correspondence, Sept 3, 2026)
- Nature Medicine – Authors’ reply to the correspondence (Sept 3, 2026)
- TechTarget Healthtech Analytics – benchmark scores across MedQA, HealthBench and the real-clinical-queries test
- CNBC – OpenEvidence reaches a $12 billion valuation; used by 40% of U.S. physicians
- OpenAI – HealthBench, the benchmark built by OpenAI