Biomarker Overfitting in Small Samples

The single most common methodological error in ME/CFS research — and the one with the most severe consequences for credibility — is the discovery of “diagnostic biomarkers” in small samples without independent external validation. A systematic review identified 149 candidate biomarkers nominated in the ME/CFS literature — and zero that have been validated in an independent cohort (Maksoud et al. 2023). This is not a random failure. It is a structural prediction of the research designs used.

1 The General Problem: Why Small-Sample Biomarker Studies Fail

Ioannidis (2005) identified the conditions under which most published research findings are false: small sample sizes, small effect sizes, large feature spaces, flexible analysis pathways, and fields with high investigator interest (Ioannidis 2005). ME/CFS biomarker research meets every single one of these conditions.

Button et al. (2013) demonstrated that neuroscience studies with n=10–20 have a median statistical power of 21% — meaning 79% of true effects go undetected while false positives are amplified by the winner’s curse (Button et al. 2013). Vul et al. (2009) showed that non-independent analysis inflates brain-behavior correlations, producing “voodoo correlations” above the theoretical maximum (Vul et al. 2009).

In the machine learning context specifically, Varoquaux (2017) demonstrated that standard cross-validation in small neuroimaging samples (n \(<\) 100) overestimates accuracy by 20–40% — nested cross-validation is essential for unbiased performance estimates (Varoquaux et al. 2017). Flint et al. (2020) showed that biomarker studies in psychiatry with n \(<\) 100 produce effects 2–3× larger than those found in large consortia (n \(>\) 1,000) — the winner’s curse operating in small samples (Flint 2023).

2 The ME/CFS-Specific Crisis

These general principles apply with full force to ME/CFS. A typical ME/CFS biomarker study has:

CautionSpeculation: The ME/CFS biomarker discovery pipeline is a textbook case of the Ioannidis conditions
  • Small sample (n=15–50) with dozens to hundreds of features measured
  • Standard (non-nested) cross-validation reporting 80–100% accuracy
  • No external validation in an independent cohort
  • Publication in a specialty journal with a receptive audience
  • Subsequent studies finding the same biomarker at 60–70% accuracy (regression to the true value) — but the 100% discovery study is the one that gets cited

Fonseca et al. (2024) provide a clear recent example: an EBV antibody ML classifier achieved 100% accuracy in discovery but fails on non-EBV-triggered ME/CFS — demonstrating that the “biomarker” was detecting EBV serostatus, not ME/CFS (Fonseca et al. 2024).

Falsifiable prediction: An external validation consortium that tests all 149 candidate ME/CFS biomarkers on a shared, independent, well-phenotyped cohort (n ≥ 500, ICC/CCC criteria, with matched disease controls) will validate ≤ 10 biomarkers at sensitivity/specificity both \(\geq 0.70\), and \(< 5\) at \(\geq 0.80\). If \(> 20\) biomarkers validate at those thresholds, the overfitting narrative is wrong — the literature produced more genuine signals than this analysis credits.

Consequence: None of the 149 candidate biomarkers in the ME/CFS literature should be considered validated. Each is a hypothesis, not a finding. The research community’s habit of reporting “discovery” accuracy without external validation creates false certainty that guides research funding away from replication — which is less exciting but more valuable.

3 Solutions: What Would Convincing Evidence Look Like?

The minimum standard for a credible biomarker claim in ME/CFS:

  1. Nested cross-validation or independent holdout set within the discovery study
  2. External validation in an independent cohort from a different recruitment site, with different diagnostic criteria documented
  3. Sensitivity and specificity reported with confidence intervals, not just accuracy — accuracy is a poor metric when prevalence is low
  4. Comparison to relevant disease controls (MS, RA, depression, POTS) to establish ME/CFS specificity
  5. Pre-registration of the analysis plan before seeing the validation data

The 2025–2026 studies by Xiong et al. (n=228, external validation) (Xiong et al. 2023) and Aitken et al. (n=4,244, within-person models) (Aitken et al. 2026) represent a methodological upgrade — larger samples, proper validation, digital approaches that scale — but remain exceptions to the field norm.

TipRecommendation: Reducing Biomarker Overfitting
  • Funders should require external validation (independent cohort, different site) as a condition of continued biomarker funding. Discovery without validation is hypothesis generation, not evidence.
  • Journals should require nested cross-validation or holdout validation, and should not publish studies reporting “diagnostic accuracy” without independent external validation.
  • Researchers should pre-register biomarker analysis plans and report both the discovery and validation results in the same paper — not split across publications.
  • Readers should treat any biomarker study with accuracy \(> 80%\) in a sample of n \(<\) 100 as a likely overfit — the true accuracy is almost certainly lower.

Consequence: The 149 published candidate biomarkers represent 149 hypotheses about ME/CFS biology, not 149 diagnostic tests. Until the field standardizes on external validation as a minimum publication requirement, the literature will continue to generate biomarker claims that fail to replicate.

An external validation consortium — where candidate biomarkers from multiple laboratories are tested on a shared, independent, well-phenotyped cohort — would resolve this field-wide problem in one coordinated effort. The model exists: psychiatric genetics consortia (PGC) and Alzheimer’s disease biomarker consortia solved structurally identical problems through pre-competitive data sharing. ME/CFS lacks the funding and institutional coordination to do this, but the methodological solution is known. Origin: brainstorm.

References

Aitken, Annie, Abbey Sawyer, Akiko Iwasaki, Harlan M. Krumholz, Rory Preston, Paul Calcraft, Harry Leeming, et al. 2026. “Digital Physiological Biomarkers Predict Within-Person Symptom Changes in Complex Chronic Illness.” Npj Digital Medicine 9: 257. https://doi.org/10.1038/s41746-026-02543-3.
Button, Katherine S., John P. A. Ioannidis, Claire Mokrysz, Brian A. Nosek, Jonathan Flint, Emma S. J. Robinson, and Marcus R. Munafò. 2013. “Power Failure: Why Small Sample Size Undermines the Reliability of Neuroscience.” Nature Reviews Neuroscience 14 (5): 365–76. https://doi.org/10.1038/nrn3475.
Flint, Jonathan. 2023. “The Genetic Basis of Major Depressive Disorder.” Molecular Psychiatry 28: 256–68. https://doi.org/10.1038/s41380-022-01743-z.
Fonseca, Ana, Mateusz Szysz, Hoang T. Ly, Clara Cordeiro, and Nuno Sepúlveda. 2024. “Classifying Myalgic Encephalomyelitis/Chronic Fatigue Syndrome Using Machine Learning and IgG Antibody Responses to Epstein-Barr Virus Peptides.” Medicina 60 (1): 161. https://doi.org/10.3390/medicina60010161.
Ioannidis, John P. A. 2005. “Why Most Published Research Findings Are False.” PLoS Medicine 2 (8): e124. https://doi.org/10.1371/journal.pmed.0020124.
Maksoud, Rosamund, Cassandra Magawa, Natalie Eaton-Fitch, Kiran Thapaliya, and Sonya Marshall-Gradisnik. 2023. “Biomarkers for Myalgic Encephalomyelitis/Chronic Fatigue Syndrome (ME/CFS): A Systematic Review.” BMC Medicine 21: 189. https://doi.org/10.1186/s12916-023-02893-9.
Varoquaux, Gaël, Pradeep R. Raamana, Denis A. Engemann, Andrés Hoyos-Idrobo, Yannick Schwartz, and Bertrand Thirion. 2017. “Assessing and Tuning Brain Decoders: Cross-Validation, Caveats, and Guidelines.” NeuroImage 145: 166–79. https://doi.org/10.1016/j.neuroimage.2016.10.038.
Vul, Edward, Christine Harris, Piotr Winkielman, and Harold Pashler. 2009. “Puzzlingly High Correlations in fMRI Studies of Emotion, Personality, and Social Cognition.” Perspectives on Psychological Science 4 (3): 274–90. https://doi.org/10.1111/j.1745-6924.2009.01125.x.
Xiong, Ruoyun, Courtney Gunter, Elizabeth Fleming, Suzanne D. Vernon, Lucinda Bateman, Derya Unutmaz, and Julia Oh. 2023. “Multi-’omics of Gut Microbiome-Host Interactions in Short- and Long-Term Myalgic Encephalomyelitis/Chronic Fatigue Syndrome (ME/CFS) Patients.” Cell Host & Microbe 31 (2): 273–287.e5. https://doi.org/10.1016/j.chom.2023.01.001.