I have a lazy reflex I am not proud of. Show me a study of two people and I shrug; show me a study of two million and some tired corner of my brain files it under settled, done, true. So when Britain’s biggest-ever health study reported on its first 1.9 million volunteers this week, I braced for a victory lap about how much we suddenly know. The scientists running it did the opposite. They told everyone not to trust their own numbers yet.

The project is called Our Future Health, and it wants to become the largest health database ever built: five million adults across England, Scotland and Wales handing over their blood, their DNA and their linked medical records to one national resource. It is not running on goodwill. The UK government has committed up to £354 million for 2026 through 2030, and the program has raised another £180 million from “leading life sciences companies” and health charities. On September 14 its lead researchers, Victoria Straub and Melinda Mills among them, published the first full profile of the cohort in Nature Medicine, and instead of boasting about the size, they spent the paper explaining why the size can mislead you.

WHO PAID
354£ million
UK government
180£ million
Companies and charities
Public and private money behind Our Future Health, 2026 through 2030. Source: Our Future Health, how we are funded

Those life-sciences companies are not anonymous well-wishers. The industry members list reads like a roll call of big pharma: Alnylam, Amgen, AstraZeneca, Biogen, Boehringer Ingelheim, GSK, MSD, Novartis, Novo Nordisk, Pfizer, Regeneron, Roche and Johnson & Johnson’s Janssen, plus the sequencing and diagnostics firms Illumina, Thermo Fisher and Randox. Sixteen companies. In the program’s early years they are the only large commercial organisations eligible to apply to use the resource, they get time to analyze what they find before sharing their results back, and they profit from whatever they discover on a promise to make “reasonable efforts” to bring the innovations to NHS patients. Half a billion pounds of public and private money to build a genetic and medical portrait of the nation, with drugmakers holding a paying, and for now exclusive, seat at the table.

So who actually showed up to be portrayed? Not the country. Around 4.5 percent of the people invited said yes, and the ones who did are not a cross-section of Britain. The most deprived neighborhoods are underrepresented, 13 percent of the cohort against 20 percent of the population, while the least deprived are overrepresented at 26.7 percent. Younger adults and most minority ethnic groups are thin on the ground. The volunteers smoke far less than the country does, 7.38 percent of men currently smoking versus 13.22 percent in national surveys, and they report less heart disease, less high blood pressure, fewer strokes. The people who signed up for the big health study are, on average, healthier and wealthier than the people the health system actually has to treat.

WHO SAID YES
4.5%
of invited people
Consent rate across the whole recruitment effort. Source: News-Medical, Our Future Health analysis, 2026

But why doesn’t adding another million people just wash the skew out? This is the question I kept circling, because it runs against every instinct that says more is better. The bias does not live in how many people you have. It lives in who chose to walk through the door. When enrollment is voluntary, the act of signing up is itself tangled up with being health-conscious, stable, educated, well. Pour in another million volunteers and you have not diluted that tilt, you have photocopied it a million more times. Scale shrinks the confidence interval and the numbers start to look razor sharp. What scale cannot do is add back the poor, sick, distrustful people who never enrolled. You end up exquisitely certain about a population that does not exist.


There is a seductive number in the paper that shows the trap in action. Our Future Health’s disease patterns track UK Biobank, the last mega-cohort, at r = 0.784 across 109 conditions. At a glance that looks like validation: two giant independent studies agree, so they must be right. But both recruited the same way, by asking for volunteers, so both pull in the same healthier, wealthier slice of the country. Two samples biased in the same direction will happily agree with each other. That is not confirmation. That is an echo. And even that correlation leaves close to 40 percent of the variance across conditions unexplained.

To be fair, size does buy something, and this is where it earns its keep. For rare diseases, sheer numbers are a gift. Our Future Health has already assembled far bigger patient groups than UK Biobank ever managed: 172 people with cystic fibrosis against 25, some 1,739 with idiopathic intracranial hypertension against 132, 668 with myasthenia gravis against 266, and 464 with primary biliary cholangitis against 312. For a disease that turns up in one in fifty thousand people, a database this big is the difference between a question you can study and one you cannot. Scale is the wrong tool for measuring how common hypertension is, and the right one for a disease that rare. Same database, opposite verdict, depending on the question.

RARE DISEASE CASES
ConditionOur Future HealthUK Biobank
Cystic fibrosis17225
Idiopathic intracranial hypertension1,739132
Myasthenia gravis668266
Primary biliary cholangitis464312
Rare-condition case counts, where sheer size helps most. Source: News-Medical, Our Future Health analysis, 2026

There is one more crack the authors flag. A lot of the data is self-reported, and self-report does not always match the medical record. Agreement between what volunteers said and what their hospital files showed ran from a Cohen’s kappa of 0.29 for high cholesterol to 0.66 for cancer, a scale that measures how well two sources line up beyond what chance alone would give you: fair at the low end, substantial but not airtight at the high end. The primary-care records that would firm this up are not even integrated yet. Which is why the researchers spell it out: the cohort should not yet be used to derive generalizable estimates of how common a disease is, or how fast it spreads.

SELF-REPORT VS RECORDS (Cohen's kappa)
agreement beyond chance0.290.66
How well what volunteers said matched their medical files, from high cholesterol to cancer. Source: News-Medical, Our Future Health analysis, 2026

Hold that warning next to the funding model. Sixteen drug companies are helping pay for a resource its own architects say leans wealthy and well, and that cannot yet tell you the true prevalence of a condition. The paper documents the skew, not any boardroom decision, and I want to be careful about the difference. But the incentive sits right there in plain sight: feed a portrait of the healthy into a choice about which disease to chase or which patients a therapy is “for,” and you optimize for the people who showed up, not the people who are sick. The volunteers most likely to be missing, the poor and the marginalized who trust institutions least, are the ones carrying the heaviest disease burden. A pipeline built on the healthy is a pipeline that quietly designs around them.

So here is where I landed. The next time a press release waves a number like 1.9 million at me as though size were the same thing as truth, I am going to ask the unglamorous questions first, the way these authors did to their credit: not how many, but who, and who is missing, and who paid to be in the room when the answers come out. I would rather trust a smaller study that knows exactly who it left out than a giant one that leaves out the sick and the poor and hopes the sheer weight of numbers hides the gap.

Sources

  1. Nature Medicine – Straub, Mills et al., baseline profile of the Our Future Health cohort (2026)
  2. News-Medical – “More data, more answers? A 1.9 million-person health study shows why scale has limits” (2026-09-14)
  3. Our Future Health – How industry members work with us (industry partner list and data-access terms)
  4. Our Future Health – How we are funded (government and life-sciences funding breakdown)