Longitudinal Deep Phenotyping Cohort

1 The ME/CFS Human Phenome Project

We propose the establishment of a comprehensive longitudinal deep phenotyping cohort—the “ME/CFS Human Phenome Project”—designed to capture the full biological trajectory of the condition across patient subtypes, disease stages, and outcomes. This infrastructure investment addresses a fundamental gap in ME/CFS research: the absence of large-scale, multi-dimensional, longitudinal data that would enable identification of predictive biomarkers, disease subtypes, and therapeutic targets.

1.1 Rationale and Scientific Justification

Existing ME/CFS research suffers from several structural limitations that this cohort would address. Most studies are cross-sectional, capturing a single timepoint in what is clearly a dynamic disease process. Sample sizes are typically small (n=20–50), underpowered to detect the heterogeneous subtypes that likely exist within the ME/CFS umbrella. Few studies combine multiple data modalities, limiting the ability to identify system-level interactions. Finally, standardization across studies remains poor, hampering meta-analysis and replication. The pediatric recovery data provide compelling justification for longitudinal deep phenotyping. If 54–94% of children recover while fewer than 22% of adults do (Joyce, Hotopf, and Wessely 1997) (Rowe 2019), longitudinal tracking from early disease through either recovery or chronicity could reveal the biological determinants of these divergent trajectories. Such data could identify therapeutic targets by detecting which systems change during recovery, revealing potential intervention points. Longitudinal data could enable prognostic stratification by determining which baseline features predict recovery versus chronicity, enabling personalized treatment intensity. Subtype identification would emerge from tracking how different patient clusters evolve over time, distinguishing biological subtypes from disease stages. Finally, understanding mechanistic sequences through temporal ordering would clarify whether immune abnormalities precede metabolic dysfunction or vice versa, addressing questions of causation versus correlation.

1.2 Study Design Overview

Population and Enrollment The cohort would enroll 500 participants across multiple sites internationally, with stratified enrollment to ensure representation across key dimensions:

  • Age strata: 100 pediatric/adolescent (\(<\) 18 years), 300 adult (18–60 years), 100 older adult (\(>\) 60 years)
  • Disease duration: 200 early (\(<\) 2 years), 200 established (2–10 years), 100 long-duration (\(>\) 10 years)
  • Severity: 150 mild, 200 moderate, 150 severe/very severe
  • Trigger type: Post-infectious (stratified by pathogen where known), post-other-trigger, gradual onset Follow-up Duration and Assessment Schedule Participants would be followed for 5–10 years with comprehensive assessments at baseline, 6 months, 12 months, and annually thereafter. Abbreviated assessments (wearables, questionnaires, limited biomarkers) would occur quarterly. Event-triggered assessments would capture crashes, infections, significant functional changes, or apparent recovery. Matched Control Cohort A cohort of 200 matched healthy controls would undergo identical assessments to establish reference ranges and age-related trajectories. An additional 100 disease controls with other fatiguing conditions (multiple sclerosis fatigue, post-cancer fatigue, fibromyalgia) would enable specificity analysis.

1.3 Comprehensive Assessment Battery

Multi-Omic Profiling Each comprehensive assessment would include:

  • Genomics: Whole genome sequencing at baseline (variants, structural variants, HLA typing)

  • Epigenomics: Genome-wide DNA methylation (Illumina EPIC array), targeted histone modification assays, epigenetic clocks

  • Transcriptomics: Whole blood RNA-seq, immune cell subset-specific expression profiling

  • Proteomics: Plasma proteomics (SomaScan or Olink platforms, \(>\) 5000 proteins), CSF proteomics (subset, n=100)

  • Metabolomics: Untargeted plasma and urine metabolomics, targeted panels (acylcarnitines, amino acids, organic acids, lipids)

  • Microbiome: 16S rRNA and shotgun metagenomic sequencing of gut, oral, and skin microbiomes Immune Profiling Comprehensive immune characterization would include:

  • Flow cytometry: Extended panels for T cells (naive, memory, exhaustion markers), B cells (including plasmablasts), NK cells (cytotoxicity markers), monocyte subsets

  • Functional assays: NK cell cytotoxicity, T cell proliferation, cytokine production capacity

  • Autoantibody panels: GPCR autoantibodies, anti-neuronal antibodies, comprehensive autoimmune screening

  • Inflammatory markers: High-sensitivity cytokine panels (Luminex), acute phase reactants Autonomic and Cardiovascular Assessment

  • Tilt table testing: Standardized 10-minute stand or 45-degree tilt with continuous hemodynamic monitoring

  • Heart rate variability: 24-hour Holter monitoring with time-domain and frequency-domain analysis

  • Blood volume: 51Cr-labeled red cell mass and 125I-albumin plasma volume (subset, n=100)

  • Sudomotor function: QSART or Sudoscan

  • Baroreflex sensitivity: Beat-to-beat blood pressure monitoring with baroreflex calculation Exercise Physiology

  • Two-day CPET: Cardiopulmonary exercise testing on consecutive days to capture PEM signature

  • Metabolic chamber: Indirect calorimetry for 24-hour energy expenditure (subset, n=50)

  • Muscle biopsy: Vastus lateralis biopsy for mitochondrial function, fiber typing, histology (optional, subset)

  • Handgrip dynamometry: Serial grip strength with fatigue protocol Neuroimaging and Neurophysiology

  • Structural MRI: T1, T2, FLAIR, DTI for white matter integrity

  • Functional MRI: Resting state connectivity, task-based (motor, cognitive)

  • PET imaging: TSPO ligand for neuroinflammation (subset, n=50)

  • EEG: Quantitative EEG with spectral analysis

  • Cognitive testing: Computerized battery (attention, memory, processing speed, executive function) Clinical and Patient-Reported Outcomes

  • Standardized questionnaires: DSQ-PEM, MFI, SF-36, Bell Disability Scale, PHQ-9, GAD-7

  • Symptom diaries: Daily electronic symptom tracking

  • Activity monitoring: Continuous accelerometry (wrist-worn devices)

  • Sleep assessment: Actigraphy, sleep diaries, polysomnography (subset)

  • Functional assessment: 6-minute walk test (when safe), NASA Lean Test

1.4 Biorepository Specifications

The project would establish a centralized biorepository with long-term storage capacity:

  • Sample types: Whole blood, serum, plasma (EDTA, citrate, heparin), PBMCs, urine, stool, saliva, DNA, RNA
  • Storage conditions: Liquid nitrogen (\(-196^{compose}\)C) for cells and RNA, \(-80^{compose}\)C for other samples
  • Aliquoting: Multiple aliquots per sample type to enable future analyses without thaw cycles
  • Quality control: Standardized collection protocols, processing within 2 hours, regular QC audits
  • Capacity: Estimated 2 million aliquots over the project duration
  • Access policy: Open to qualified researchers with approved proposals, data sharing agreements, and acknowledgment requirements

1.5 Data Infrastructure and Analysis

Data Management

  • Central database: FAIR-compliant (Findable, Accessible, Interoperable, Reusable) data repository

  • Data harmonization: Common data elements aligned with NIH/CDC standards

  • Privacy protection: De-identification, secure enclaves for sensitive data, tiered access

  • Longitudinal linking: Robust participant ID system enabling cross-timepoint analysis Analysis Plan Primary analyses would include:

  • Trajectory modeling: Latent class growth models to identify distinct disease courses

  • Predictive modeling: Machine learning approaches to predict recovery, progression, treatment response

  • Multi-omic integration: Network-based integration of genomic, transcriptomic, proteomic, and metabolomic data

  • Subtype identification: Unsupervised clustering across data modalities

  • Causal inference: Mendelian randomization for causal pathway identification

1.6 Budget Estimate and Feasibility

Estimated Costs

  • Multi-omic profiling: $15,000–20,000 per participant per comprehensive assessment
  • Clinical assessments: $5,000 per participant per visit
  • Neuroimaging: $3,000 per participant per session
  • Biorepository: $5 million infrastructure, $500,000/year operations
  • Data infrastructure: $3 million setup, $500,000/year maintenance
  • Coordination and administration: $2 million/year
  • Total estimated budget: $40–50 million over 10 years Feasibility Considerations This proposal is ambitious but feasible given precedents. The UK Biobank enrolled 500,000 participants with extensive phenotyping at lower per-participant cost but less depth. The All of Us Research Program demonstrates large-scale longitudinal deep phenotyping infrastructure. The ME Research UK Biobank, though smaller, provides a model specific to ME/CFS. Key feasibility challenges include participant retention over 5–10 years, particularly for severely ill patients who may find assessments burdensome. Mitigation strategies include home visits, abbreviated protocols for severe cases, and strong participant engagement. Protocol evolution will be necessary as technologies advance; the protocol must allow incorporation of new assays while maintaining comparability. Expected Outcomes If successful, this cohort would generate the definitive longitudinal dataset for ME/CFS research, enabling validated prognostic biomarkers to guide treatment intensity, biological subtype definitions to enable precision medicine, therapeutic targets emerging from trajectory analysis, and a shared resource to accelerate research across the field

References

Joyce, J., M. Hotopf, and S. Wessely. 1997. “The Prognosis of Chronic Fatigue and Chronic Fatigue Syndrome: A Systematic Review.” QJM: An International Journal of Medicine 90 (3): 223–33. https://doi.org/10.1093/qjmed/90.3.223.
Rowe, Katharine S. 2019. “Long Term Follow up of Young People with Chronic Fatigue Syndrome Attending a Pediatric Outpatient Service.” Frontiers in Pediatrics 7: 21. https://doi.org/10.3389/fped.2019.00021.