Data Requirements
Mathematical models are only as reliable as the data constraining them. ME/CFS modeling faces specific data challenges.
1 Longitudinal Data
Most ME/CFS studies are cross-sectional, comparing patients with controls at a single time point. ODE models require longitudinal data—repeated measurements over time—to constrain dynamic parameters (rate constants, time delays). The NIH intramural study (Walitt et al. 2024) and two-day CPET protocols (Keller et al. 2024) provide rare longitudinal data points, but multi-month or multi-year trajectories remain scarce. Symptom tracking data from patient registries (Jason et al. 2012) offer temporal resolution but limited molecular detail.
2 Multi-Omics Integration
ME/CFS involves simultaneous dysfunction across metabolomic, proteomic, transcriptomic, and epigenomic layers. Integrating these heterogeneous data types into a single model requires careful normalization, dimensionality reduction, and alignment of timescales. Recent multi-omics studies (Xiong et al. 2023) (Heng et al. 2025) provide templates for data integration, but sample sizes remain small relative to the number of measured variables, creating overfitting risk.
3 Clinical Parameters
Clinical measures—symptom severity scores, functional capacity ratings, exercise tolerance—are the ultimate validation targets for ME/CFS models but are inherently subjective and variable. Standardized instruments such as the DePaul Symptom Questionnaire and the Bell Disability Scale (Appendix Diagnostic Tools and Assessment Scales) reduce but do not eliminate this variability. Models must account for measurement noise in clinical outcomes.
The number of identifiable parameters in any model is bounded by the information content of available data. For ME/CFS, where typical studies involve \(n = 20\)–$ 100$ patients measured at $ 1\(--\) 3$ time points across $ 10\(--\) 50$ biomarkers, models with more than approximately $ 10\(--\) 20$ free parameters risk overfitting. The models in Chapters Energy Metabolism Models through Neuroendocrine and Autonomic Models are therefore deliberately parsimonious, prioritizing well-constrained parameters over biological completeness. Full-complexity models await the large-scale longitudinal datasets currently under development (e.g., DecodeME (DecodeME Consortium, Ponting, et al. 2025)).