Mark Schiffman had spent a career at the National Cancer Institute studying what causes cervical cancer when, in January 2019, his name went out on a press release that promised to help end it. A deep-learning algorithm his team had built, the National Institutes of Health announced, could look at a photograph of a woman’s cervix and spot precancer better than a trained doctor. “The computer analysis of the images was better at identifying precancer than a human expert reviewer,” Schiffman was quoted saying. The number attached to the claim was the kind that ends arguments: an AUC of 0.91, against 0.69 for an expert doing a visual exam and 0.71 for a conventional Pap smear.

AUC, CERVICAL PRECANCER
Deep-learning algorithm0.91Expert visual exam0.69Pap smear0.71
The 2019 headline number, measured on data that resembled the algorithm's own training set. Source: JNCI (Hu et al.), 2019

The pitch had a destination. The tool was built for “low-resource settings,” the phrase the global-health world uses for the places where women die of cervical cancer because no one screens them and no one treats them in time. A health worker with a cell-phone camera, the story went, could screen and treat in a single visit. Maurizio Vecchione, an executive at Global Good, the fund inside Nathan Myhrvold’s Intellectual Ventures that had partnered with the NCI to build it, said cervical cancer “could be brought under control, even in low-resource settings.” The algorithm had been trained on more than 60,000 cervical images from an NCI archive, photographs taken during a screening study in Costa Rica in the 1990s, and the result was published that year in the Journal of the National Cancer Institute.

It was a clean story, and clean stories move. The NCI got a breakthrough with its name on it. Global Good, a commercial fund, got a global-health halo. The journals and the wire services got a number that sounded like a verdict. The people asked to trust it were global-health funders and health ministries; the people who would bear the risk if it was wrong were poor women being screened, once, by a tool that had never been tested outside the archive it was born in.

What did not travel was the correction. In a letter published in the same journal in December 2025, Schiffman, the senior and corresponding author of the 2019 paper, wrote that “the extraordinarily accurate AI-based cervical screening described in that publication derived from ‘overfitting’ not addressed by ‘internal validation.’” Overfitting is the oldest failure in machine learning: the model aces the exam because it was quietly handed the answer key. Graded through internal validation on data close enough to its training set, the 0.91 measured how well the algorithm had memorized, not how well it would read a cervix it had never seen.

Then the sentence a decade of enthusiasm rests on. “Our subsequent use of the algorithm in new projects failed,” Schiffman wrote, “as one would expect based on what we now know.” By his own account, the tool did not just underperform its headline. It stopped working once it left home.

The press release circled the globe. The admission ran four paragraphs in a journal’s letters column.


None of this is exotic, which is the uncomfortable part. Overfitting is a first-week mistake. What lifts the cervical-AI case out of the pile is who it was aimed at and how finished it was declared to be: announced to the world as a breakthrough for the global poor, not as a promising internal experiment that still had to survive a different camera, a different clinic, a different population. When outside researchers ran the check the original effort skipped, a generalizability analysis in PLOS Digital Health found AVE-style performance that did not carry cleanly to new data, the same weakness Schiffman would later concede in his own words.

Which brings us to this week. On August 12, Nature Medicine ran a piece calling frugal AI “the missing piece” in cervical cancer screening. The subject is automated visual evaluation, the same class of tool, back for another turn under a kinder adjective. “Frugal” does quiet work. It moves the question from does it work to isn’t it affordable, and affordability has always been the easiest thing to promise to people who are not in the room.

The second act, to be fair, is being built with more care than the first. A consortium pairing HPV genotyping with automated visual evaluation has been running prospective validation in countries like Zambia, reporting strong performance again, this time with the external checks the original skipped. That is how it should have been done the first time: test the image tool alongside a molecular HPV test, on the actual target population, before the press release. Whether the second-generation numbers survive the scrutiny that dissolved the first-generation numbers is the open question, and it will be settled in African clinics, not in a journal’s news pages.

Cervical cancer kills hundreds of thousands of women a year, overwhelmingly in the poor countries this technology keeps being pointed at, and those women are owed a tool that works rather than one that photographs well. Schiffman, to his credit, said close to that himself. He said it in a letters column, six years after the headline and long after anyone was still listening, in the four paragraphs a journal keeps for its second thoughts. The 0.91 got a press release. The correction got correspondence.

Sources

  1. Nature Medicine – “Frugal AI is the missing piece in cervical cancer screening” (Aug 12, 2026)
  2. JNCI – Mark Schiffman, correspondence acknowledging the 2019 AVE algorithm’s overfitting and failed follow-on projects (Dec 2025)
  3. NIH / ScienceDaily – “AI approach outperformed human experts in identifying cervical precancer,” AUC 0.91 vs 0.69 vs 0.71 (Jan 2019)
  4. JNCI – Hu et al., “Observational Study of Deep Learning and Automated Evaluation of Cervical Images for Cancer Screening” (2019)
  5. JNCI – HPV–Automated Visual Evaluation Consortium, new HPV-genotyping-plus-AVE screening strategy (2025)
  6. PLOS Digital Health – “Assessing generalizability of an AI-based visual test for cervical cancer screening”