Start in Michigan, in the winter of 2019, because medicine has already sat through the movie that is about to play again. For two years a sepsis-prediction tool had been running quietly inside Epic, the electronic-records system used by hundreds of American hospitals, watching for the early signature of a body tipping into septic collapse and paging clinicians when it thought it saw one coming. On paper it looked superb. Its maker advertised an ability to pick out the patients who would go septic from those who would not at a level that sounded close to clairvoyance.

Then a team at the University of Michigan did the thing that almost never happens before a model is switched on over real people. They graded it against reality. Across 38,455 hospitalizations they externally validated the Epic Sepsis Model and found an area under the curve of 0.63, well below the 0.76 to 0.83 its vendor had advertised.

In plain terms, it caught only about a third of the sepsis cases it was built to catch, a sensitivity of 33 percent, and it did so while firing alerts on nearly a fifth of every patient who came through the door. A tool that had dazzled on records missed two in three of the people it actually met.

SEPSIS CAUGHT
33%
of true cases, at the alert threshold
The model flagged one in three real sepsis cases while alerting on nearly a fifth of all patients. Source: JAMA Internal Medicine, 2021

Hold that number, because on September 4 a lab at the First Affiliated Hospital of Chongqing Medical University published in Nature Medicine the biggest and most ambitious version of the same idea yet, aimed at the most vulnerable patients in medicine: mothers and their infants. It has read more mother-and-baby records than any doctor could in a hundred careers. It has not yet been asked to be right about one living patient.

The system is called MoChiAgent, and the engineering is impressive. It is a large-language-model agent that marshals a set of tools across a patient’s sequential health record, including the routine bloodwork most models throw away, and feeds a core predictive engine the authors call MoChiFormer. That engine was developed and internally evaluated on 4,401,599 longitudinal clinical visits, then externally validated on two independent cohorts of 263,452 and 23,192 visits. Under the hood it reconstructs missing laboratory values, reduces the batch effects that make one hospital’s data look unlike another’s, estimates gestational, fetal and infant age from the record alone, and stratifies a patient’s current and future disease risk. This is not a chatbot bolted onto a spreadsheet. It is a serious attempt to teach a machine the shape of a pregnancy and a first year of life as they are written down in a hospital’s files.

MOCHIAGENT'S SCALE
4,401,599
visits to build it
263,452
external cohort
23,192
external cohort
The records the model was developed and tested on. None involved guiding live care. Source: Nature Medicine, 2026

Everything hinges on that last clause. As they are written down. Everything MoChiAgent has done, it has done to records. The paper describes retrospective development and external validation, which means the model was tested against outcomes that had already happened, sitting in a database with the answers already filled in. There is no prospective trial in it: no cohort of real pregnant women whose care was actually steered by the model’s warnings, no measured comparison of whether the mothers and babies it flagged came out healthier than the ones it did not. The authors describe their work as opening a window to enhance risk-stratified care. That is a statement about what might be possible, not a result. Predicting who is at risk and lowering that risk are two different achievements, and only the first one is on the table.

We have watched this gap swallow tools that looked better on their debut than MoChiAgent does on its. The likely reason the Epic model failed was not sloppiness but something more stubborn: a model can learn the statistical fingerprint of a database so well that it mistakes the fingerprint for the disease, and databases behave nothing like bedsides. Part of what predicts sepsis in a record is simply the trace of clinicians already suspecting sepsis and ordering the tests that confirm it. Strip that context away, put the model in front of a patient no one has flagged yet, and the clairvoyance thins out fast. Nothing about maternal and infant medicine, with its own tangle of who gets tested and when, makes it immune to the same trap.

The plumbing alone argues for modesty. Reliably linking one mother’s record to her own child’s inside an electronic health system, the most basic prerequisite for any of this, is still enough of a technical problem that a team thought it worth publishing a dedicated linkage algorithm for it in 2025. MoChiAgent also arrives into a crowded field. Forecasting disease from longitudinal records has become its own race, with separate groups claiming to predict everything from early-childhood ADHD out of the same kind of EHR trail. Almost all of them share MoChiAgent’s shape: validated on the past, marketed toward the future, and only rarely put through the prospective, real-world test that separates a benchmark from a bedside instrument.

There is a second thing worth saying plainly, because it does not fit the triumphant framing. A model like this is, at bottom, a risk score for pregnant women and their babies, distilled from millions of them held in one health system’s files. Risk scores are not neutral infrastructure. A number the model emits does not stay a number: it becomes a flag in a chart, an automated referral, an extra line of surveillance, a high-risk label that follows a mother from clinic to clinic and can shape how a nervous triage desk, or an insurer, treats her long before anyone has confirmed she is sick. Who holds that score, who acts on it, and who gets sorted into a tier they never agreed to join are governance questions, not technical ones, and a prestige journal’s stamp settles none of them.

A model built inside one country’s hospital system is also, until shown otherwise, an open question in every other one. Generalization is a claim to be earned cohort by cohort, not inherited from a large training number. Whether the same engine reads an American, Nigerian or German pregnancy the way it reads the ones it grew up on is exactly the thing retrospective validation on home turf cannot tell you.

None of this makes MoChiAgent junk. It makes it a hypothesis, a very large and sophisticated one, that has not yet met the test that counts. The next card sits with the Chongqing group and whoever licenses their work: a prospective study, run inside live obstetric and neonatal care, that reports not just discrimination on held-out records but the alert burden, the false-positive rate that lands on healthy mothers, and, above all, whether the women and infants it flagged actually ended up healthier than the ones it passed over. Until that trial is run and published in full, the honest headline is the smaller one. An AI has read more mother-and-baby records than any clinician ever could, and it has yet to be asked to be right about a single living patient.

Sources

  1. Nature Medicine – Liu et al., “Prediction of maternal and infant outcomes from longitudinal electronic health records with a Mother-Child AI agent” (MoChiAgent), 2026
  2. JAMA Internal Medicine – Wong et al., “External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients” (2021)
  3. JAMIA – “Derivation and validation of an algorithm for maternal–child linkage in electronic health records” (2025)
  4. Nature Mental Health – “Early attention deficit hyperactivity disorder prediction from longitudinal electronic health records” (2026)