Replication Methodology for Heterogeneous Disease

ME/CFS, like other multisystem chronic illnesses, exhibits profound treatment heterogeneity. Conventional randomized controlled trials powered to detect average treatment effects frequently fail despite strong mechanistic rationale and robust phase II data. This section examines why average-effect replication fails in heterogeneous diseases and proposes advanced designs—basket trials, umbrella trials, adaptive enrichment, and N-of-1 aggregation—suited to ME/CFS subgroup structure.

1 Why Average-Effect Replication Fails

The fundamental challenge in replicating treatments for heterogeneous diseases is that a population average effect may not apply to any identifiable subgroup within that population. This principle is captured in two complementary concepts: the ecological fallacy and Simpson’s paradox. The ecological fallacy, as articulated by Kent and Hayward (Kent and Hayward 2007), occurs when population-level associations fail to represent individual-level phenomena. An average treatment effect of zero across an entire population may conceal subgroups experiencing strong benefit and others experiencing harm. In simplified terms, if only 30% of patients respond to a given treatment, the observed average effect is diluted even if responders achieve a large clinical benefit. The magnitude of dilution depends on effect size in responders, the effect in non-responders, and the responder prevalence. Simpson’s paradox, documented systematically in clinical contexts by Kievit and colleagues (Kievit et al. 2013), demonstrates that a treatment appearing ineffective overall can simultaneously be highly effective when data are stratified by an unmeasured or unconsidered subgroup variable. This paradox is not a statistical artifact but a direct consequence of unaccounted treatment-by-subgroup interactions. The PATH Statement (Predictive Approaches to Treatment Effect Heterogeneity), developed by Kent, Steyerberg, and van Klaveren (Kent, Steyerberg, and Klaveren 2020), provides a formal framework for recognizing and addressing this problem: not all patients benefit equally from the same treatment, and replication of aggregate trials without addressing heterogeneity will continue to fail.

1.1 ME/CFS Case Study: The Rituximab Saga

The rituximab story illustrates this principle with particular clarity. In the phase II trial by Fluge and colleagues (Fluge et al. 2011), 10 of 15 rituximab-treated patients (67%) showed a clinical response to this B-cell depleting monoclonal antibody, with improvements in function and symptoms. This result motivated a large phase III trial, RituxME, enrolling 151 unstratified patients. The phase III trial failed: \(P = 0.80\), showing no benefit above placebo (Fluge et al. 2019). Standard interpretation might conclude that rituximab is ineffective in ME/CFS. However, a 6-year follow-up by Rekeland and colleagues identified a subset of long-term responders whose benefits persisted years after treatment (Rekeland et al. 2024), suggesting that rituximab works in some but not all patients. The mechanistic clue is that B-cell depletion affects multiple cell populations, including antibody-producing plasmablasts and memory B cells. One plausible explanation is that ME/CFS patients with autoimmune features (particularly elevated GPCR autoantibodies) may preferentially benefit from B-cell depletion, while patients without these features derive less benefit (Scheibenbogen et al. 2018). If this hypothesis is correct, the average effect across both groups would collapse toward zero. Similar patterns emerge with other agents. Rintatolimod showed marginal improvement in an intent-to-treat analysis but a 51.2% improvement rate in patients with disease duration of 2–8 years (Strayer et al. 2020). Immunoadsorption, when applied to patients pre-selected by positive beta-2 adrenergic receptor autoantibodies, achieved a 70% response rate (Stein et al. 2025). Each trial that appeared to fail at the population level succeeded in a biomarker-enriched subgroup.

2 Master Protocol Designs for ME/CFS

Advanced trial designs, collectively termed “master protocols,” enable simultaneous testing of multiple treatments and subgroup definitions within a single framework. Three master protocol archetypes are relevant to ME/CFS:

2.1 Basket Trials

A basket trial assigns one treatment to multiple disease subtypes or patient populations united by a shared biomarker. Rather than enrolling all patients into a single trial, basket designs recognize that a treatment mechanism targets a biomarker-defined subset. Woodcock and LaVange (Woodcock and LaVange 2017) describe this approach in the context of precision oncology, where a single targeted agent is tested across multiple cancer types defined by a shared mutation. The principle applies directly to ME/CFS: patients with a shared pathophysiological feature (e.g., high-titer anti-\(\beta_2\) adrenergic receptor antibodies) would be enrolled into a basket cohort for a single therapeutic intervention.

2.2 Umbrella Trials

An umbrella trial structures itself inversely: multiple treatments are tested simultaneously within ME/CFS, with patients stratified into matched treatment cohorts according to their biomarker profile. A single master protocol might randomize autoantibody-positive patients to immunoadsorption versus placebo, metabolically-classified (ME-M1/M2/M3 phenotype (Germain et al. 2020)) patients to subtype-matched metabolic agents versus placebo, and others to immunomodulatory interventions. Each treatment arm targets a distinct biomarker-defined subgroup, maximizing the relevance of comparisons. In lupus, the anifrolumab trial (TULIP-2) employed a similar logic, stratifying patients by interferon gene signature (Morand et al. 2020), and achieved positive results where previous lupus trials had failed by targeting heterogeneous pathways.

2.3 Platform Trials

A platform trial maintains a perpetual structure, allowing continuous addition of new treatment arms and removal of ineffective agents based on interim analyses. Rather than designing each treatment as a separate RCT, a platform trial keeps control arms and infrastructure stable while rotating experimental treatments. This approach is ideal for ME/CFS, where new candidate mechanisms (mitochondrial biogenesis agents, autoimmune-targeted therapies, autonomic modulators) could be continuously evaluated. For ME/CFS specifically, we propose that an umbrella trial design should stratify enrollment by:

  • Autoantibody status: Testing for anti-\(\beta_2\) adrenergic receptor, anti-muscarinic-3 receptor, anti-NMDA receptor, and anti-mitochondrial antibodies; stratifying positive patients into an immunoadsorption or immunomodulation arm
  • Metabolic phenotype: Using the ME-M classification (Germain et al. 2020) (ME-M1: lipolytic state; ME-M2: lipid accumulation; ME-M3: intermediate) to assign metabolic interventions to appropriate subgroups
  • Immune phenotype: Natural killer cell exhaustion markers, CD8+ T-cell exhaustion, and interleukin-6 signature to assign immunomodulatory agents
  • Autonomic profile: Heart rate variability, cardiac output decline on tilt, and orthostatic intolerance severity to assign autonomic-targeted therapies

3 Adaptive Enrichment and Biomarker-Guided Designs

Rather than pre-specifying all subgroup definitions, an adaptive enrichment design begins with broad enrollment, uses interim biomarker analysis to identify responder subgroups, and restricts subsequent enrollment to biomarker-enriched populations. Simon and colleagues (Simon, Paik, and Hayes 2009) developed the adaptive signature design specifically for this purpose: enroll broadly, measure biomarkers at baseline, perform interim analysis to identify treatment-by-biomarker interactions, and use interim results to refine eligibility criteria for the remainder of enrollment. A recent precedent in ME/CFS is instructive. Fluge and colleagues conducted a small pilot of daratumumab (anti-CD38 monoclonal antibody) in ten ME/CFS patients (Fluge et al. 2025). Six responded clinically. Critically, a low natural killer cell count at baseline predicted non-response, suggesting a biomarker for patient selection. This observation—derived from a small pilot—provides a foundation for enrichment strategy in future trials. We propose the following adaptive enrichment strategy for ME/CFS trials:

  • Phase 1 (Broad Enrollment): Enroll 30–50 patients with confirmed ME/CFS (using standardized criteria such as ICC or CCC) without biomarker pre-screening
  • Baseline Biomarker Panel: Measure comprehensive immune, metabolic, and autonomic biomarkers in all participants (NK cell phenotype, mitochondrial markers, autoantibodies, HRV, cardiac output reserve)
  • Interim Analysis (Week 12): Perform pre-specified subgroup analyses by biomarker status, identifying biomarker-treatment interactions with \(p < 0.10\) (discovery threshold)
  • Adaptive Enrichment: Restrict enrollment in Phase 2 to biomarker-enriched populations identified in interim analysis
  • Definitive Phase 2: Enroll 100–150 patients within enriched subgroup(s), powered to detect treatment effect within that subgroup This approach directly addresses the rituximab failure: early interim analysis would have identified autoantibody positivity as a predictor of response, allowing enrichment in phase 2.

4 N-of-1 Trials and Aggregation

While RCTs remain the gold standard for causal inference at population scale, they are not the only valid design for heterogeneous diseases. N-of-1 trials—single-patient randomized crossover experiments—directly measure individual response to treatment and address a critical limitation of RCTs: they cannot tell us which treatment will benefit a specific patient. Duan, Kravitz, and Schmid (Duan, Kravitz, and Schmid 2013) argue that in heterogeneous diseases, N-of-1 trials serve a complementary role to RCTs. Where RCTs estimate population-level effects, N-of-1 trials measure individual treatment response. In ME/CFS, where treatment response is highly variable, this distinction is valuable: a patient with orthostatic intolerance and documented autonomic dysfunction may benefit from midodrine despite that agent showing no net benefit in a population RCT. The power of N-of-1 design increases when aggregation is applied. Zucker and colleagues (Zucker, Ruthazer, and Schmid 2010) demonstrated that aggregating data from four or more N-of-1 trials yields valid population-level estimates using hierarchical linear modeling, while retaining individual response information. Gabler and colleagues reviewed the methodology and clinical applications (Gabler et al. 2011), concluding that aggregated N-of-1 data from 20+ patients can rival RCTs in statistical power while offering superior individual-level precision. ME/CFS is particularly suited to aggregated N-of-1 methodology for several reasons:

  • Many crossover-feasible interventions: Supplementation (L-carnitine, ubiquinol, sodium chloride), dietary manipulation (low-histamine diet, ketogenic diet), pacing strategies, and autonomic medications (beta-blockers, mineralocorticoids, antihistamines) are amenable to within-patient crossover
  • Enormous individual variation: Patients show stark differences in what worsens versus improves symptoms, making individual-level response data more clinically actionable than population averages
  • Post-hoc responder identification: If 20 N-of-1 trials are aggregated and data analyzed, biomarkers collected at baseline can be correlated with individual response, generating testable hypotheses for future enriched RCTs We propose that a network of ME/CFS research centers could conduct aggregated N-of-1 trials systematically, building both individual treatment response profiles and population-level evidence for responder biomarkers. This approach would be low-cost compared to large RCTs, would provide immediate individualized clinical utility, and would generate biomarker hypotheses to inform future basket, umbrella, and adaptive trials.

5 Proposed Application: Cromolyn N-of-1 Trial Network for MCAS Causal Hierarchy

Duan, Naihua, Richard L Kravitz, and Christopher H Schmid. 2013. “Single-Patient (n-of-1) Trials: A Pragmatic Clinical Decision Methodology for Patient-Centered Comparative Effectiveness Research.” Journal of Clinical Epidemiology 66 (8): S21–28. https://doi.org/10.1016/j.jclinepi.2013.04.006.
Fluge, Øystein, Ove Bruland, Kristin Risa, Olav Dahl, Torstein Haug, Ingileif Rekeland, Dipak Sapkota, et al. 2011. “B-Lymphocyte Depletion in Patients with Myalgic Encephalomyelitis/Chronic Fatigue Syndrome: A Randomized, Double-Blind, Placebo-Controlled Pilot Study.” Psychosomatic Medicine 73 (1): 36–43. https://doi.org/10.1097/PSY.0b013e3181f60d27.
Fluge, Øystein, Ingrid G. Rekeland, Kristin Lien, Hilde Thürmer, Petter C. Borchgrevink, Christoph Schäfer, Kari Sørland, et al. 2019. “B-Lymphocyte Depletion in Patients with Myalgic Encephalomyelitis/Chronic Fatigue Syndrome: A Randomized, Double-Blind, Placebo-Controlled Trial.” Annals of Internal Medicine 170 (9): 585–93. https://doi.org/10.7326/M18-1451.
Fluge, Øystein, Ingrid Gurvin Rekeland, Kristin Sørland, et al. 2025. “Plasma Cell Targeting with the Anti-CD38 Antibody Daratumumab in Myalgic Encephalomyelitis/Chronic Fatigue Syndrome—a Clinical Pilot Study.” Frontiers in Medicine 12: 1607353. https://doi.org/10.3389/fmed.2025.1607353.
Gabler, Nathan B, Naihua Duan, Sunita Vohra, and Richard L Kravitz. 2011. “N-of-1 Trials in the Medical Literature: A Systematic Review.” Medical Care 49 (8): 761–68. https://doi.org/10.1097/MLR.0b013e318215d000.
Germain, Arnaud, Dinesh K Barupal, Susan M Levine, and Maureen R Hanson. 2020. “Comprehensive Circulatory Metabolomics in ME/CFS Reveals Disrupted Metabolism of Acyl Carnitines and Fatty Acids.” Metabolites 10 (1): 34. https://doi.org/10.3390/metabo10010034.
Kent, David M, and Rodney A Hayward. 2007. “Limitations of Applying Summary Results of Clinical Trials to Individual Patients: The Need for Risk Stratification.” JAMA 298 (10): 1209–12. https://doi.org/10.1001/jama.298.10.1209.
Kent, David M, Ewout W Steyerberg, and David van Klaveren. 2020. “Personalized Evidence Based Medicine: Predictive Approaches to Heterogeneous Treatment Effects.” Annals of Internal Medicine 172 (1): 35–45. https://doi.org/10.7326/M19-2037.
Kievit, Rogier A, Jan-Willem Romeijn, Lourens J Waldorp, Marieke Welling, Harriet Christley, Bauke Scholtes, and Lourens J Waldorp. 2013. “Simpsons Paradox in Real Data.” Frontiers in Psychology 4: 513. https://doi.org/10.3389/fpsyg.2013.00513.
Morand, Eric F, Richard Furie, Yoshiya Tanaka, Ian N Bruce, Anca D Askanase, Christophe Richez, Se-Ho Bae, et al. 2020. “Trial of Anifrolumab in Active Systemic Lupus Erythematosus.” New England Journal of Medicine 382 (3): 211–21. https://doi.org/10.1056/NEJMoa1912196.
Rekeland, Ingrid G., Kari Sørland, Linn L. Neteland, Alexander Fosså, Kari Alme, Kristin Risa, Olav Dahl, Karl J. Tronstad, Olav Mella, and Øystein Fluge. 2024. “Six-Year Follow-up of Participants in Two Clinical Trials of Rituximab or Cyclophosphamide in Myalgic Encephalomyelitis/Chronic Fatigue Syndrome.” PLoS One 19 (7): e0307484. https://doi.org/10.1371/journal.pone.0307484.
Scheibenbogen, Carmen, Madlen Loebel, Helma Freitag, Anne Krueger, Stephan Bauer, Madeleine Antelmann, Wolfram Doehner, et al. 2018. “Immunoadsorption to Remove Beta2 Adrenergic Receptor Antibodies in Chronic Fatigue Syndrome CFS/ME.” PLOS ONE 13 (3): e0193672. https://doi.org/10.1371/journal.pone.0193672.
Simon, Richard M, Seymour Paik, and Daniel F Hayes. 2009. “Use of Archived Specimens in Evaluation of Prognostic and Predictive Biomarkers.” Journal of the National Cancer Institute 101 (21): 1446–52. https://doi.org/10.1093/jnci/djp335.
Stein, Elisa, Cornelia Heindrich, Kirsten Wittke, Claudia Kedor, Rebekka Rust, Helma Freitag, Franziska Sotzny, et al. 2025. “Efficacy of Repeated Immunoadsorption in Patients with Post-COVID Myalgic Encephalomyelitis/Chronic Fatigue Syndrome and Elevated Beta2-Adrenergic Receptor Autoantibodies: A Prospective Cohort Study.” The Lancet Regional Health - Europe 48: 101161. https://doi.org/10.1016/j.lanepe.2024.101161.
Strayer, Daniel R, William A Carter, Bradley C Stouch, Stephen R Stevens, Lucinda Bateman, Patricia J Cimoch, Michael E Meegan, Amy H Overstreet, and Sara Sabine. 2020. “A Double-Blind, Placebo-Controlled, Randomized, Clinical Trial of the TLR3 Agonist Rintatolimod in Severe Cases of Postinfectious Chronic Fatigue Syndrome.” Journal of Medical Virology 92 (7): 846–54. https://doi.org/10.1002/jmv.25676.
Woodcock, Janet, and Lisa M LaVange. 2017. “Master Protocols to Study Multiple Therapies, Multiple Diseases, or Both.” New England Journal of Medicine 377 (1): 62–70. https://doi.org/10.1056/NEJMra1510062.
Zucker, David R, Rebecca Ruthazer, and Christopher H Schmid. 2010. “Individual (n-of-1) Trials Can Be Combined to Give Population Comparative Effectiveness Evidence: Methodologic Considerations.” Journal of Clinical Epidemiology 63 (12): 1312–23. https://doi.org/10.1016/j.jclinepi.2010.04.020.