Replication Methodology for Heterogeneous Disease
ME/CFS, like other multisystem chronic illnesses, exhibits profound treatment heterogeneity. Conventional randomized controlled trials powered to detect average treatment effects frequently fail despite strong mechanistic rationale and robust phase II data. This section examines why average-effect replication fails in heterogeneous diseases and proposes advanced designs—basket trials, umbrella trials, adaptive enrichment, and N-of-1 aggregation—suited to ME/CFS subgroup structure.
1 Why Average-Effect Replication Fails
The fundamental challenge in replicating treatments for heterogeneous diseases is that a population average effect may not apply to any identifiable subgroup within that population. This principle is captured in two complementary concepts: the ecological fallacy and Simpson’s paradox. The ecological fallacy, as articulated by Kent and Hayward (Kent and Hayward 2007), occurs when population-level associations fail to represent individual-level phenomena. An average treatment effect of zero across an entire population may conceal subgroups experiencing strong benefit and others experiencing harm. In simplified terms, if only 30% of patients respond to a given treatment, the observed average effect is diluted even if responders achieve a large clinical benefit. The magnitude of dilution depends on effect size in responders, the effect in non-responders, and the responder prevalence. Simpson’s paradox, documented systematically in clinical contexts by Kievit and colleagues (Kievit et al. 2013), demonstrates that a treatment appearing ineffective overall can simultaneously be highly effective when data are stratified by an unmeasured or unconsidered subgroup variable. This paradox is not a statistical artifact but a direct consequence of unaccounted treatment-by-subgroup interactions. The PATH Statement (Predictive Approaches to Treatment Effect Heterogeneity), developed by Kent, Steyerberg, and van Klaveren (Kent, Steyerberg, and Klaveren 2020), provides a formal framework for recognizing and addressing this problem: not all patients benefit equally from the same treatment, and replication of aggregate trials without addressing heterogeneity will continue to fail.
1.1 ME/CFS Case Study: The Rituximab Saga
The rituximab story illustrates this principle with particular clarity. In the phase II trial by Fluge and colleagues (Fluge et al. 2011), 10 of 15 rituximab-treated patients (67%) showed a clinical response to this B-cell depleting monoclonal antibody, with improvements in function and symptoms. This result motivated a large phase III trial, RituxME, enrolling 151 unstratified patients. The phase III trial failed: \(P = 0.80\), showing no benefit above placebo (Fluge et al. 2019). Standard interpretation might conclude that rituximab is ineffective in ME/CFS. However, a 6-year follow-up by Rekeland and colleagues identified a subset of long-term responders whose benefits persisted years after treatment (Rekeland et al. 2024), suggesting that rituximab works in some but not all patients. The mechanistic clue is that B-cell depletion affects multiple cell populations, including antibody-producing plasmablasts and memory B cells. One plausible explanation is that ME/CFS patients with autoimmune features (particularly elevated GPCR autoantibodies) may preferentially benefit from B-cell depletion, while patients without these features derive less benefit (Scheibenbogen et al. 2018). If this hypothesis is correct, the average effect across both groups would collapse toward zero. Similar patterns emerge with other agents. Rintatolimod showed marginal improvement in an intent-to-treat analysis but a 51.2% improvement rate in patients with disease duration of 2–8 years (Strayer et al. 2020). Immunoadsorption, when applied to patients pre-selected by positive beta-2 adrenergic receptor autoantibodies, achieved a 70% response rate (Stein et al. 2025). Each trial that appeared to fail at the population level succeeded in a biomarker-enriched subgroup.
2 Master Protocol Designs for ME/CFS
Advanced trial designs, collectively termed “master protocols,” enable simultaneous testing of multiple treatments and subgroup definitions within a single framework. Three master protocol archetypes are relevant to ME/CFS:
2.1 Basket Trials
A basket trial assigns one treatment to multiple disease subtypes or patient populations united by a shared biomarker. Rather than enrolling all patients into a single trial, basket designs recognize that a treatment mechanism targets a biomarker-defined subset. Woodcock and LaVange (Woodcock and LaVange 2017) describe this approach in the context of precision oncology, where a single targeted agent is tested across multiple cancer types defined by a shared mutation. The principle applies directly to ME/CFS: patients with a shared pathophysiological feature (e.g., high-titer anti-\(\beta_2\) adrenergic receptor antibodies) would be enrolled into a basket cohort for a single therapeutic intervention.
2.2 Umbrella Trials
An umbrella trial structures itself inversely: multiple treatments are tested simultaneously within ME/CFS, with patients stratified into matched treatment cohorts according to their biomarker profile. A single master protocol might randomize autoantibody-positive patients to immunoadsorption versus placebo, metabolically-classified (ME-M1/M2/M3 phenotype (Germain et al. 2020)) patients to subtype-matched metabolic agents versus placebo, and others to immunomodulatory interventions. Each treatment arm targets a distinct biomarker-defined subgroup, maximizing the relevance of comparisons. In lupus, the anifrolumab trial (TULIP-2) employed a similar logic, stratifying patients by interferon gene signature (Morand et al. 2020), and achieved positive results where previous lupus trials had failed by targeting heterogeneous pathways.
2.3 Platform Trials
A platform trial maintains a perpetual structure, allowing continuous addition of new treatment arms and removal of ineffective agents based on interim analyses. Rather than designing each treatment as a separate RCT, a platform trial keeps control arms and infrastructure stable while rotating experimental treatments. This approach is ideal for ME/CFS, where new candidate mechanisms (mitochondrial biogenesis agents, autoimmune-targeted therapies, autonomic modulators) could be continuously evaluated. For ME/CFS specifically, we propose that an umbrella trial design should stratify enrollment by:
- Autoantibody status: Testing for anti-\(\beta_2\) adrenergic receptor, anti-muscarinic-3 receptor, anti-NMDA receptor, and anti-mitochondrial antibodies; stratifying positive patients into an immunoadsorption or immunomodulation arm
- Metabolic phenotype: Using the ME-M classification (Germain et al. 2020) (ME-M1: lipolytic state; ME-M2: lipid accumulation; ME-M3: intermediate) to assign metabolic interventions to appropriate subgroups
- Immune phenotype: Natural killer cell exhaustion markers, CD8+ T-cell exhaustion, and interleukin-6 signature to assign immunomodulatory agents
- Autonomic profile: Heart rate variability, cardiac output decline on tilt, and orthostatic intolerance severity to assign autonomic-targeted therapies
3 Adaptive Enrichment and Biomarker-Guided Designs
Rather than pre-specifying all subgroup definitions, an adaptive enrichment design begins with broad enrollment, uses interim biomarker analysis to identify responder subgroups, and restricts subsequent enrollment to biomarker-enriched populations. Simon and colleagues (Simon, Paik, and Hayes 2009) developed the adaptive signature design specifically for this purpose: enroll broadly, measure biomarkers at baseline, perform interim analysis to identify treatment-by-biomarker interactions, and use interim results to refine eligibility criteria for the remainder of enrollment. A recent precedent in ME/CFS is instructive. Fluge and colleagues conducted a small pilot of daratumumab (anti-CD38 monoclonal antibody) in ten ME/CFS patients (Fluge et al. 2025). Six responded clinically. Critically, a low natural killer cell count at baseline predicted non-response, suggesting a biomarker for patient selection. This observation—derived from a small pilot—provides a foundation for enrichment strategy in future trials. We propose the following adaptive enrichment strategy for ME/CFS trials:
- Phase 1 (Broad Enrollment): Enroll 30–50 patients with confirmed ME/CFS (using standardized criteria such as ICC or CCC) without biomarker pre-screening
- Baseline Biomarker Panel: Measure comprehensive immune, metabolic, and autonomic biomarkers in all participants (NK cell phenotype, mitochondrial markers, autoantibodies, HRV, cardiac output reserve)
- Interim Analysis (Week 12): Perform pre-specified subgroup analyses by biomarker status, identifying biomarker-treatment interactions with \(p < 0.10\) (discovery threshold)
- Adaptive Enrichment: Restrict enrollment in Phase 2 to biomarker-enriched populations identified in interim analysis
- Definitive Phase 2: Enroll 100–150 patients within enriched subgroup(s), powered to detect treatment effect within that subgroup This approach directly addresses the rituximab failure: early interim analysis would have identified autoantibody positivity as a predictor of response, allowing enrichment in phase 2.
4 N-of-1 Trials and Aggregation
While RCTs remain the gold standard for causal inference at population scale, they are not the only valid design for heterogeneous diseases. N-of-1 trials—single-patient randomized crossover experiments—directly measure individual response to treatment and address a critical limitation of RCTs: they cannot tell us which treatment will benefit a specific patient. Duan, Kravitz, and Schmid (Duan, Kravitz, and Schmid 2013) argue that in heterogeneous diseases, N-of-1 trials serve a complementary role to RCTs. Where RCTs estimate population-level effects, N-of-1 trials measure individual treatment response. In ME/CFS, where treatment response is highly variable, this distinction is valuable: a patient with orthostatic intolerance and documented autonomic dysfunction may benefit from midodrine despite that agent showing no net benefit in a population RCT. The power of N-of-1 design increases when aggregation is applied. Zucker and colleagues (Zucker, Ruthazer, and Schmid 2010) demonstrated that aggregating data from four or more N-of-1 trials yields valid population-level estimates using hierarchical linear modeling, while retaining individual response information. Gabler and colleagues reviewed the methodology and clinical applications (Gabler et al. 2011), concluding that aggregated N-of-1 data from 20+ patients can rival RCTs in statistical power while offering superior individual-level precision. ME/CFS is particularly suited to aggregated N-of-1 methodology for several reasons:
- Many crossover-feasible interventions: Supplementation (L-carnitine, ubiquinol, sodium chloride), dietary manipulation (low-histamine diet, ketogenic diet), pacing strategies, and autonomic medications (beta-blockers, mineralocorticoids, antihistamines) are amenable to within-patient crossover
- Enormous individual variation: Patients show stark differences in what worsens versus improves symptoms, making individual-level response data more clinically actionable than population averages
- Post-hoc responder identification: If 20 N-of-1 trials are aggregated and data analyzed, biomarkers collected at baseline can be correlated with individual response, generating testable hypotheses for future enriched RCTs We propose that a network of ME/CFS research centers could conduct aggregated N-of-1 trials systematically, building both individual treatment response profiles and population-level evidence for responder biomarkers. This approach would be low-cost compared to large RCTs, would provide immediate individualized clinical utility, and would generate biomarker hypotheses to inform future basket, umbrella, and adaptive trials.