The Subjective-Measurable Discrepancy: A Diagnostic Pattern Across ME/CFS Measurement Domains

A pattern emerges across the measurement domains examined in this paper: patients report severe symptoms that standard clinical tests fail to capture — or explicitly contradict. This discrepancy is not an anomaly. It is a replicated, multi-domain pattern spanning sleep, autonomic, cognitive, activity, and systemic biomarker measurement domains across a 12-paper evidence base. The discrepancy is documented across multiple diagnoses (ME/CFS, POTS, hEDS, MCAS, PASC — Novak2024, n=2,627), not exclusive to ME/CFS. Its diagnostic value lies not in disease specificity but in clinical expectation-setting: normal resting labs plus severe subjective symptom burden IS an expected finding in ME/CFS, not evidence of psychogenesis. Whether the discrepancy magnitude or pattern differs between ME/CFS and other diagnoses (i.e., whether it is discriminative, not merely present) has not been tested.

TipAchievement: The Subjective-Measurable Discrepancy is Not an Isolated Befuddlement — It Patterns Across Domains

The subjective-measurable discrepancy has been independently documented across five measurement domains by seven independent research cohorts, with the strongest evidence coming from monozygotic co-twin control designs that eliminate genetic and family-environment confounds:

Sleep: In the Buchwald twin cohort, CFS twins endorsed all 8 subjective insomnia measures significantly more than their healthy co-twins, yet objective polysomnography showed zero differences in sleep latency, total sleep time, sleep efficiency, arousal index, or NREM architecture (Watson et al. 2003). The same cohort showed significantly worse subjective sleepiness on validated instruments but identical mean sleep latency on the Multiple Sleep Latency Test (8.9 vs 10.0 min, both non-pathologic) (Watson et al. 2004). Quantitative EEG spectral power analysis across delta through beta bands showed no significant differences between CFS twins and healthy co-twins in any frequency band (Armitage et al. 2009). The pattern is consistent across macrostructural, microstructural-scalar, and spectral measures: patients say sleep is profoundly disturbed; standard instruments say it is normal.

Autonomic: Novak et al. (2024) administered validated subjective autonomic symptom questionnaires (SAS, COMPASS-31) and objective autonomic testing (QASAT, CASS, tilt test, sudomotor, Valsalva) to 2,627 patients across 11 diagnoses including ME/CFS, POTS, hEDS, MCAS, and PASC (Novak et al. 2024). Internal correlations were strong within each measurement domain (subjective-subjective r=0.74; objective-objective r=0.81), confirming internal consistency and test-retest reliability — though this does not establish construct validity (see @(oq-bridging-subjective-objective-gap?)). The cross-domain correlation — subjective vs objective — was zero. The instruments measure unrelated constructs. This dissociation was diagnosis-independent.

Physical activity: Vergauwen et al. (2021) found activity diaries correlated with Actical accelerometry at rs=0.35–0.38 in 66 female CFS patients vs rs=0.60–0.64 in 20 healthy controls (Vergauwen et al. 2021). The discrepancy was not explained by illness-related complaints, health-related quality of life, or demographics.

Cognition: Teodoro et al. (2018) reviewed 186 studies across functional neurological disorder, fibromyalgia, and CFS and documented a “general discordance between high rates of subjective cognitive symptoms… and inconsistent objective neuropsychological deficits” (Teodoro, Edwards, and Isaacs 2018). Performance validity testing showed poor effort in only a minority, ruling out malingering.

Systemic biomarkers: In the largest monozygotic twin study of ME/CFS genomics, Byrnes et al. (2009) found zero significant transcriptomic differences in peripheral blood leukocytes between 44 CFS twin pairs despite severe disabling illness (Byrnes et al. 2009). This is the cleanest molecular-level null result in the ME/CFS literature — the illness exists, is disabling, but produces no peripheral blood transcriptomic signature detectable by microarray.

The discrepancy cannot be fully attributed to psychiatric comorbidity (the twin design controls for genetic and shared-family-environment confounds, though it does not control for acquired, non-shared psychiatric comorbidity in the affected twin), poor effort (Teodoro2018: performance validity testing), or deconditioning (van Campen et al. demonstrated CBF decline during tilt test is independent of deconditioning severity Campen, Rowe, and Visser (2021)). An alternative explanation for the sleep domain in particular — sleep-state misperception (paradoxical insomnia), a recognized ICSD diagnosis defined precisely by subjective-objective sleep discordance — has not been formally excluded and is not discussed in the original papers; the sleep-discrepancy evidence does not distinguish a distinct ME/CFS sleep phenotype from the general phenomenon of sleep-state misperception that also occurs in non-fatigued populations. The within-domain correlations in Novak2024 (subjective-subjective r=0.74, objective-objective r=0.81) demonstrate test-retest reliability and internal consistency, not construct validity — instruments can consistently measure the wrong construct, and the chapter’s own open question about construct confusion (What Measurement Paradigm Can Bridge the Discrepancy Gap?) explicitly argues they may do so.

Certainty: 0.60 — the discrepancy pattern is well-documented across domains, but the evidence base has partially correlated sources: the sleep domain uses one twin cohort for two studies (Watson2003, Armitage2009); the autonomic domain is a single study (Novak2024, n=2,627 — the strongest single paper); the activity domain found a weak positive relationship (rs=0.35–0.38), not a zero; the cognitive domain is a systematic review (Teodoro2018), not a primary study; and the biomarker domain is a tissue-specific null (Byrnes2009 — strong for blood leukocytes, not proven for other compartments). The five domains are not five fully independent replications. The 0.60 valuation reflects this structure — the discrepancy exists, but the independence weight supporting the achievement’s initial 0.75 overstates the evidential structure (The Achievement’s Certainty (0.75) Overestimates the Independence of the Evidence Base). Severity applicability: unknown — most studies sample ambulatory mild-moderate patients; the pattern may differ in bedbound severe patients where subjective and objective measures may both show profound abnormality.

Consequence: In a patient meeting ME/CFS diagnostic criteria, severe functional impairment with entirely normal standard laboratory results is an expected clinical finding — not psychosomatic overlay, not malingering. The discrepancy is a signal, not noise. Clinicians who interpret normal labs as evidence that “nothing is wrong” are systematically misreading the disease. This expectation-setting does not replace differential diagnosis; normal labs in an undiagnosed fatigued patient require the same diagnostic workup as any other presentation.

TipSynthesis: The Subjective-Measurable Discrepancy as a Diagnostic Index

The core mechanistic question is not whether the discrepancy exists — it is what the discrepancy means. Three non-exclusive explanations have been proposed:

Provocation-dependence. Many ME/CFS abnormalities are detectable only under stress provocation (exercise, orthostatic challenge, cognitive load). Haptoglobin depletion is present post-exertion but absent at rest (Moezzi et al. 2025). Cerebral blood flow declines during tilt-testing but resting hemodynamics are normal Campen, Rowe, and Visser (2021). Hand grip strength correlates with symptom severity specifically in the ME/CFS phenotype (Paffrath et al. 2024) — a provocation test that reveals what resting accelerometry misses. Standard clinical labs are resting measurements, and resting measurements systematically miss provocation-dependent pathology. This is not a measurement failure — it is a measurement domain mismatch.

Central nervous system confinement. Byrnes2009’s peripheral blood transcriptomic null result and Omdal2026’s finding that Long COVID fatigue severity is unassociated with any peripheral inflammatory biomarker (Omdal et al. 2026) converge on the hypothesis that a significant fraction of ME/CFS pathology may be CNS-confined — brainstem nuclei, hypothalamic circuits, microglial activation — and not reflected in every peripheral blood measurement. This is not an absolute claim — the 7-biomarker blood panel achieving 91% diagnostic accuracy (Section Systems Biology Perspective on ME/CFS) demonstrates that SOME peripheral signals do track the disease, at discriminatory levels. CNS confinement is a partial, not total, explanation: some biologies are detectable in blood (metabolic, vascular, certain immune markers); others (transcriptomic at the PBMC level, subjective fatigue perception) may not be. A CNS disease whose peripheral immunological and metabolic consequences are measurable in blood while its core neuronal transcriptomic signature is restricted to the CNS compartment is not a paradox — it is the expected compartment dissociation in a multi-system disease whose central integrator (brainstem/hypothalamus) is anatomically inaccessible to phlebotomy.

Interoceptive amplification. Teodoro2018 proposed a mechanistic model: pain, fatigue, and excessive interoceptive monitoring → decreased externally directed attention → increased distractibility → cognitive processes experienced as unduly effortful, mediated by a switch from automatic to controlled processing mode. This is a neurobiological model (predictive processing, attention allocation), not a psychological one. The patient’s conscious experience captures dimensions of dysfunction (attention-capture by normally subconscious physiological signals) that lie outside the measurement bounds of standard clinical instruments.

These explanations are not mutually exclusive, but they do have implications for whether the subjective report should be treated as veridical. Provocation-dependence and CNS confinement assume the patient is accurately perceiving real pathology that standard instruments cannot access — the gap is in the measurement. Interoceptive amplification assumes the perception is biologically amplified — the gap is in the percept. The models can coexist (different domains may reflect different mixes; the amplification may be amplifying a real signal that provocation reveals), but the interpretation of what a narrowed discrepancy index would mean differs: for measurement-gap models, a narrowed discrepancy means the instrument was finally good enough; for the amplification model, a narrowed discrepancy means the amplification was dampened. Discriminating the models is not trivial because pharmacological probes (guanfacine, beta-blockers) alter peripheral physiology as well as central processing — a narrowed discrepancy could reflect reduced peripheral signal rather than reduced central amplification. A cleaner discriminator would be an intervention that alters interoceptive gain without affecting peripheral physiology (e.g., interoceptive training via heartbeat perception feedback; cholinergic modulation of precision-weighting specifically in insular cortex), or direct measurement of the interoceptive computational parameters via a Bayesian observer model fitted to heartbeat detection data.

Consequence: If provocation-dependence drives the discrepancy, the diagnostic pathway shifts from “run standard labs, find nothing, consider psychiatric” to “standard labs are expected to be normal — now run provocation testing (tilt table, CPET, dynamometry).” This would change clinical practice in ME/CFS assessment. If CNS confinement drives the discrepancy, the implication is that peripheral blood biomarkers will remain weak regardless of how sensitive the assay — the measurement target is inaccessible to phlebotomy.

ImportantHypothesis: The Discrepancy Index as a Subtyping Tool: High Discrepancy → Dysautonomia-Dominant

A formal discrepancy index — the ratio of subjective symptom severity to objectively measurable dysfunction — may subtype ME/CFS patients:

The mechanism is straightforward. Dysautonomia is provocation-dependent: orthostatic tachycardia, cerebral hypoperfusion, and HRV suppression emerge during upright posture or exertion but are absent at rest. A patient with pure dysautonomia has zero resting abnormalities and normal standard labs — but reports severe orthostatic symptoms, fatigue with upright activity, and brain fog that worsens with standing. Their discrepancy index is high: subjective report >> objective resting measurements.

In contrast, a patient with immune-inflammatory pathology (elevated CRP, elevated cytokines, microglial activation on PET, autoantibodies detectable on cell-based assay) has objectively measurable correlates of symptom severity. Their discrepancy index is low: subjective report ≈ objective measurements — both indicate pathology.

Note on definitional circularity: Dysautonomia is characterized in part by provocation-dependence — orthostatic abnormalities emerge only during upright posture or exertion. A discrepancy index that divides subjective severity (high during upright activity) by objective resting measures (normal) will necessarily be high in dysautonomic patients by definition — the index’s numerator captures symptoms triggered by a postural state its denominator was not designed to detect. This means the association between high discrepancy and dysautonomia is partly definitional (the index was structured to elevate dysautonomic patients) and partly empirical (whether the COMPASS-31/QASAT discrepancy predicts best treatment response within the dysautonomic subgroup). The empirical claim — that within dysautonomic patients, discrepancy magnitude predicts autonomic drug response — is the part that requires testing. The definitional claim — that dysautonomic patients will show the highest discrepancy — may be tautological. A ratio-based index with near-zero denominators (QASAT scores close to zero in purely-provocational dysautonomia) also produces explosive ratios driven by measurement noise rather than genuine discrepancy magnitude; ratio-score reliability has not been psychometrically assessed. The index additionally inherits the scale-type problem of any ordinal-scale ratio: dividing two ordinal scores produces a metric whose units are not meaningful — the ratio cannot be interpreted as “twice as large” in any measurement-theoretic sense.

Certainty: 0.45 — the discrepancy phenomenon is well-documented (0.60, The Subjective-Measurable Discrepancy is Not an Isolated Befuddlement — It Patterns Across Domains), but the subtyping prediction and treatment-response hypotheses are untested. No prospective study has used a discrepancy index to stratify patients or predict outcomes. Severity applicability: unknown.

Falsifiable prediction: In a cohort of ME/CFS patients stratified by COMPASS-31 (subjective autonomic symptom score) ÷ QASAT (objective autonomic deficit score), the high-discrepancy tertile should show (a) significantly greater improvement on midodrine or pyridostigmine than the low-discrepancy tertile, and (b) significantly less CRP, cytokine elevation, and autoantibody positivity.

Consequence: If the index works, clinicians could triage patients at first visit — high discrepancy → autonomic workup and treatment; low discrepancy → immunological workup and treatment — replacing the current trial-and-error approach to ME/CFS management.

CautionSpeculation: The Discrepancy Index as an Interoceptive Accuracy Measure

An alternative framing: the discrepancy index is not a subtyping tool but a measure of interoceptive accuracy — how well a patient’s conscious perception maps onto their actual peripheral physiology. In the predictive processing framework, interoception is Bayesian inference: the brain generates predictions about internal states and updates them based on ascending sensory signals. Precision weighting (the relative confidence assigned to predictions vs prediction errors) determines whether perception tracks peripheral physiology or drifts on prior expectations (Teodoro, Edwards, and Isaacs 2018).

High-discrepancy patients may have abnormally high precision on ascending interoceptive signals — they are exquisitely sensitive to small physiological perturbations (heart rate increase, vasodilation, metabolic shift) that objectively remain within the normal range. Their conscious experience is not “wrong” — they are perceiving real signals that healthy individuals’s brains filter out as noise. The discrepancy is between a high-gain interoceptive system and instruments calibrated for detecting pathology-level (not signal-level) deviations.

This reframes the discrepancy as a neurobiological property of the individual’s interoceptive system, not a subtype of ME/CFS. It could be measured independently via heartbeat detection tasks, respiratory resistance discrimination, or thermal detection thresholds. If interoceptive accuracy predicts the discrepancy index, the index becomes a proxy for a neurobiological trait — not a disease classification.

Certainty: 0.30 — predictive processing and interoceptive accuracy frameworks are well-established in neuroscience but have not been applied to ME/CFS patient-measurement discrepancy specifically. Severity applicability: unknown.

Falsifiable prediction: ME/CFS patients should show a distinct interoceptive computational signature compared to healthy controls — either heightened interoceptive accuracy (higher on ascending signals, precision-weighting phenotype) or reduced accuracy with overweighted priors (the more common somatic-symptom-disorder pattern). The direction of the difference distinguishes between competing models: greater accuracy supports the precision-weighting hypothesis; reduced accuracy supports the prior-dominance / somatic-symptom alternative. A formal Bayesian observer model is required to distinguish the two; a simple accuracy comparison (binary correct/incorrect) is insufficient because both models can produce the same accuracy score via different mechanisms (e.g., reduced accuracy with elevated confidence vs reduced accuracy with reduced confidence). If neither parameter differs from healthy controls, the interoceptive premise collapses — patients process body signals identically to healthy individuals, and the discrepancy is structural (wrong instrument, wrong compartment, wrong timepoint), not perceptual.

Consequence: If confirmed, the discrepancy becomes a direct readout of interoceptive processing gain — a neurobiological trait measurable with existing psychophysics paradigms. If refuted, the interoceptive amplification model collapses to unfalsifiable status — the mismatch is structural (wrong compartment, wrong domain) rather than perceptual.

WarningLimitation: The Discrepancy Index is a Measurement Framework, Not a Validated Diagnostic Tool

The discrepancy index proposed here is a conceptual framework synthesizing cross-domain evidence, not a validated clinical instrument. No study has: (a) quantified a discrepancy score using a standardised formula, (b) assessed test-retest reliability, (c) determined sensitivity/specificity against a gold standard, or (d) demonstrated clinical utility in decision-making. The concept is supported by convergent evidence but the tool does not yet exist.

Several structural issues remain unresolved:

  • Domain selection: Which measurement domains (autonomic, cognitive, sleep, activity, biomarker) should compose the index, and with what weights?
  • Provocation standardisation: If the discrepancy resolves under provocation, should the index use resting measurements (maximising discrepancy) or post-provocation measurements (minimising it)?
  • Severity dependency: The index may behave differently in mild (high function, low measurable dysfunction, large discrepancy) vs severe (low function, high measurable dysfunction, small discrepancy) patients — the same score could mean the opposite. van Campen et al. (2021) include severe patients; the twin and activity cohorts sample mild-moderate. Severity adjustment of the index is unstudied.
  • Threshold calibration: The subtyping prediction (high discrepancy → dysautonomia, low → immune-inflammatory) requires empirical validation of the cut-point. An untested cut-point makes the prediction unfalsifiable — post-hoc threshold adjustment can “confirm” any result.
  • Measurement direction: Some discrepancies run opposite — Paffrath2024 found HGS as an objective measure DOES correlate with symptom severity in ME/CFS specifically. The discrepancy is domain-specific and not universally present.

Certainty: n/a — methodological caveat, not a substantive claim. Severity applicability: the resolution of the severity-sensitivity problem is a prerequisite for any clinical use.

Consequence: The discrepancy index should be treated as a hypothesis-generating framework — a structured way to think about measurement discordance — not as a clinical decision tool. The burden of proof is on demonstrating (not assuming) that the index adds information beyond what clinical judgment already provides.

NoteOpen Question: What Measurement Paradigm Can Bridge the Discrepancy Gap?

The discrepancy literature reveals a paradox: instruments that correlate with each other within-domain fail to correlate across-domain. COMPASS-31 and SAS correlate at r=0.74; QASAT and CASS correlate at r=0.81. But COMPASS-31 and QASAT correlate at r=0 (Novak et al. 2024). The instruments are internally consistent but cross-dimensionally orthogonal — they measure different constructs entirely.

This raises a fundamental measurement question: is there any instrument — existing or designable — that would bridge the gap? Candidates include:

  • Provocation-based patient-report (ask patients to rate symptoms DURING tilt test, not retrospectively) — may capture the objective abnormality that resting questionnaires miss
  • Wearable physiological monitoring with ecological momentary assessment — timestamp-aligned subjective and objective data may reveal temporal correlations that cross-sectional sampling obscures
  • Microstructural objective measures (CAP analysis, spindle density, microarousal index for sleep; beat-to-beat HRV dynamics, baroreflex sensitivity for autonomic) — may capture the specific pathology that macrostructural measures (total sleep time, resting HR) miss
  • Challenge-test protocols — combine provocation (tilt, cognitive load, thermal stress) with simultaneous subjective AND objective measurement

A negative result — no instrument bridges the gap regardless of paradigm — would support CNS confinement: the constructs are ontologically distinct because the patient’s experience integrates CNS-level information (interoceptive precision, attention capture, predictive processing) that no peripheral instrument can access.

Certainty: n/a — methodological research question. Severity applicability: severity-stratified instrument validation is needed.

Consequence: Until the gap is bridged, ME/CFS research is measuring two different things with two different instrument classes and calling both “autonomic function,” “sleep quality,” or “cognitive performance” — systematic construct confusion. Resolving this is prerequisite to any biomarker that claims to track symptom severity.

NoteOpen Question: Why Do Some Domains Show Convergence?

Paffrath2024 demonstrated that hand grip strength correlates with hallmark ME/CFS symptom severity — a domain where subjective and objective converge. Friedberg2022 showed that lower HRV (an objective autonomic measure) correlates with patient-reported nonimprovement trajectory — another convergence point. vanCampen2021 showed that CBF decline during tilt is present regardless of resting hemodynamics — the objective measure captures what the subjective report was already indicating.

What distinguishes converging from diverging domains? Three candidate explanations (each is a substantive causal hypothesis framed here as a question because no systematic comparison has been conducted):

  • Provocation-intensity: HGS is a maximal-effort test (provocation ceiling); tilt table is a moderate provocation; resting accelerometry is no provocation. Convergence may be a function of provocation intensity — the harder you stress the system, the more the objective measure and subjective report align. Falsifiable: in a within-subject study with increasing provocation levels (rest → light activity → CPET), subjective-objective correlation should increase monotonically with provocation intensity (monotonic trend, p for linear contrast < 0.01).
  • Measurement specificity: HRV and CBF measure specific physiological parameters with known clinical significance. COMPASS-31 aggregates heterogenous symptoms into a single score — dilution of signal. Falsifiable: replacing COMPASS-31 with symptom-specific visual analogue scales (dizziness VAS, palpitations VAS, brain fog VAS during tilt) should narrow the discrepancy relative to the aggregate score.
  • Temporal coupling: Post-exertional haptoglobin depletion requires measurement at the right time window. Standard labs at random clinic visits miss time-dependent signals. Falsifiable: in a time-series study with serial sampling (0, 1, 2, 4, 8, 24h post-exertion), the subjective-objective correlation should peak at a specific lag and fall to zero at random timepoints.

If convergence domains share features (provocation intensity, measurement specificity, temporal coupling), the discrepancy is resolvable through instrument design rather than reflecting an irreducible patient-instrument gap. If convergence domains show no systematic pattern, the discrepancy may be stochastic — some measures align, some don’t, for reasons opaque to current understanding.

Certainty: 0.30 — convergence domains are documented post-hoc; no systematic comparison of converging vs diverging domains has been conducted. Severity applicability: unknown.

Consequence: Identifying convergence patterns guides instrument design — invest in high-provocation, high-specificity, temporally-coupled measures (CPET + symptom rating during test; tilt + simultaneous COMPASS-31; post-exertion sampling at fixed intervals) rather than trying to extract convergence from resting measures that are structurally mismatched to the pathology’s dynamics.

CautionSpeculation: Discrepancy Magnitude as Illness-Trajectory Predictor

Friedberg2022 found that lower HRV — an objective autonomic dysfunction measure — was associated with nonimprovement at 6-month follow-up, while patient-reported activity patterns (push-crash, limiting, pacing) did NOT differ between improvers and non-improvers (Friedberg et al. 2022). This inverts the discrepancy’s diagnostic role: the objective measure predicted trajectory while the subjective report did not.

Extrapolating: a patient whose discrepancy index is high (severe subjective report, mild objective dysfunction) may have a better prognosis than one whose index is low (severe subjective report, severe objective dysfunction) — the latter has accumulated measurable pathology driving the symptoms, while in the former measurable organ-level damage has not yet accumulated, and the subjective report may reflect early dysfunction below the detection threshold of coarse clinical instruments.

Alternatively: high discrepancy may represent early-stage disease (subjective distress precedes measurable pathology — the patient knows something is wrong before the lab does), while low discrepancy represents late-stage disease (measurable pathology has accumulated to match symptom severity). In this model, the discrepancy index is a disease-stage biomarker — declining with illness duration as measurable dysfunction accumulates. Vergauwen2021 and Teodoro2018 would predict the opposite (“chronic illness” would correlate with greater discrepancy), but no study has tested discrepancy magnitude against illness duration.

Certainty: 0.25 — single longitudinal study (Friedberg2022, n=148, 6-month follow-up) supports the objective-predicts-trajectory finding; extrapolation to discrepancy index as prognostic tool is speculative. Severity applicability: unknown.

Falsifiable prediction: In a longitudinal cohort (n≥100, 12-month follow-up), baseline discrepancy index magnitude should predict improvement vs deterioration: higher-discrepancy patients improve more (or deteriorate less) than matched low-discrepancy patients.

Consequence: If discrepancy predicts trajectory, the index moves from “interesting observation” to “clinically actionable prognostic tool” because it identifies which patients need aggressive early intervention (low-discrepancy = accumulating measurable damage) and which may improve with time and autonomic support (high-discrepancy = system still compensating). This is a treatment-allocation decision tool, not just a classification tool.

1 The Discrepancy Index — Self-Critique, Null Hypotheses, and Evidence Quality Boundaries

The above environments construct a framework. The following environments interrogate it. These are categories 10–12 ideas — structural reasons the discrepancy framework may be incomplete, misleading, or wrong. They are integrated here alongside the framework they challenge because epistemic balance requires that the constructive and critical arguments occupy the same section, not be relegated to separate chapters where readers encounter one without the other. Origin: brainstorm — /integrate-topic subjective-measurable-discrepancy-index.

CautionSpeculation: The Energy-Conservation Signal: The Brain Detects Metabolic Flux Failure That Blood Tests Miss

The brain’s hypothalamus and brainstem nuclei continuously monitor metabolic state via AMPK, mTOR, and mitochondrial retrograde signaling (ROS, NAD+/NADH ratio, ATP/ADP ratio). These sensors detect FLUX failure — the rate of ATP production cannot meet demand — even when metabolite POOL sizes (blood glucose, lactate, pyruvate) are normal. In this model, blood tests measure pools; the brain measures flux. If the model is correct, the subjective report of fatigue, malaise, and “something is wrong” would reflect the brain’s correct detection of flux insufficiency — not interoceptive amplification, not CNS confinement, not provocation-dependent pathology. Objective tests return normal because they measure pools.

This hypothesis is distinct from the three existing explanatory models (provocation-dependence, CNS confinement, interoceptive amplification): it is not provocation-dependence (the flux failure is continuous but invisible to pool assays), not CNS confinement (the signal originates in peripheral tissue mitochondria but is detected centrally via metabolic sensing — it crosses compartments), and not interoceptive amplification (the brain is correctly detecting a real signal — no amplification needed).

Evidence: mitochondrial dysfunction in ME/CFS is documented at the flux level — impaired phosphocreatine recovery on 31P-MRS post-exercise, impaired Complex V activity, reduced oxygen extraction on CPET — but NOT at the resting metabolite pool level (normal resting lactate, normal CK, normal glucose). Moezzi2025HaptoglobinME showed haptoglobin depletion post-exertion but normal at rest — a flux signal invisible to single-draw sampling. The hypothalamic metabolic sensing circuits are well-characterized in obesity/diabetes research but have not been applied to ME/CFS fatigue perception.

Certainty: 0.30 — the pool/flux distinction in ME/CFS metabolomics is well-documented (0.55 for flux abnormalities, 0.65 for normal resting pools), but the specific claim that central metabolic sensors mediate the discrepancy has zero direct evidence in ME/CFS. Severity applicability: unknown.

Falsifiable prediction: ME/CFS patients should show impaired phosphocreatine recovery kinetics on 31P-MRS that correlates with subjective fatigue severity (r ≥ 0.4) while resting phosphocreatine/ATP ratio is normal and does NOT correlate with fatigue. If resting metabolite pools predict subjective fatigue better than flux measures, the pool/flux distinction is irrelevant.

Consequence: The patient feels exhausted because their cells are running on fumes — but standard blood tests only measure how much fuel is in the tank, not how fast the engine can burn it. The tank is half-full; the problem is the fuel line is pinched. This reframes the discrepancy from “patient perception is amplified” to “standard measurement is mis-targeted.” Origin: brainstorm.

CautionSpeculation: Temporal Decoupling: Subjective Integration Across Weeks vs. Objective Snapshots at One Moment

Subjective symptom reports integrate across days to weeks (the recall period of most validated instruments), while objective measures capture a single timepoint (one night in a sleep lab, one tilt test, one blood draw). If the disease fluctuates — and ME/CFS is characterized by fluctuation — the subjective report captures the worst moments via the peak-end rule and availability heuristic, while the objective measure randomly samples one data point from the distribution. The discrepancy is the difference between the distribution’s tail (subjective) and a random sample from the distribution’s body (objective).

This is NOT interoceptive amplification (which says perception is amplified) — it is a temporal sampling bias that would produce discrepancy even with perfect interoception. Two patients with identical objective physiology (same mean, same variance) could report completely different subjective severity if one patient weights their worst moments more heavily. The discrepancy is a measurement domain mismatch, not a perceptual pathology.

Evidence: the peak-end rule in symptom recall is well-established in pain research. Vergauwen2021 used 12-day diary vs. 12-day accelerometry — even with best-matched instruments, rs=0.35–0.38, suggesting temporal integration partially but not fully explains the discrepancy. Friedberg2022 found patient-reported activity patterns did NOT discriminate improvers from non-improvers at 6-month follow-up — but objective HRV did — consistent with subjective reports being insensitive to the distribution while objective measures capture an average.

Certainty: 0.40 — the peak-end rule in symptom recall is well-established (0.70 in general health psychology), but its specific contribution to the ME/CFS discrepancy has not been quantified. Severity applicability: unknown.

Falsifiable prediction: In a 14-day continuous monitoring study (daily symptom rating + continuous HR/HRV/activity), compare: (a) mean symptom vs. mean objective, (b) peak symptom vs. mean objective, (c) mean symptom vs. objective at worst-3-days. If matched-distribution correlation is significantly higher (Δr > 0.15), temporal decoupling is confirmed. If the correlation is unchanged, temporal sampling does not explain the discrepancy.

Consequence: If you ask a patient “how was your month?” they remember the 3 days they couldn’t get out of bed, not the 27 days they managed. If you measure their physiology on one of the 27 okay days, it looks normal. Both are “true” — they just sampled different parts of the distribution. This reframes the discrepancy as a measurement design problem (temporal window mismatch), not a patient unreliability problem. Origin: brainstorm.

WarningLimitation: The Discrepancy Evidence Base Has a Systematic Instrument-Mismatch Confound

Every study in the discrepancy evidence base shares a structural design feature: subjective instruments use multi-week recall periods (PSQI: 1 month; COMPASS-31: recent weeks; Chalder Fatigue Scale: 1 month) while objective instruments sample a single timepoint (1-night PSG, 1-visit QASAT, 1 blood draw). The discrepancy is partly designed into the measurement protocol — a necessary consequence of the retrospective-integrative vs. instantaneous-sampling instruments chosen, not solely a discovery about the disease.

Vergauwen2021 provides the most informative constraint: even with matched time windows (12-day diary vs. 12-day actigraphy), the correlation is weak (rs=0.35–0.38). This suggests instrument-mismatch is a partial explanation — matched windows still leave a substantial discrepancy — but the magnitude of the instrument contribution has not been formally decomposed. The fact that matched windows still show weak correlation across three studies (Vergauwen2021 for activity; Watson2003 for single-night PSG vs 1-month PSQI; Armitage2009 for QEEG during one sleep laboratory night vs lifetime sleep perception) means instrument selection bias is not the whole explanation — the discrepancy is robust to instrument choice.

Certainty: 0.60 — the instrument-mismatch confound is present in all studies. The consequence for the discrepancy framework: the achievement’s assertion that the discrepancy is replicated across independent domains needs the caveat that the replication includes a shared structural bias. The robustness of the pattern to instrument choice partially mitigates but does not eliminate this concern. Severity applicability: n/a — methodological concern.

Consequence: Each study that documents the discrepancy compared a month-long memory to a single-moment snapshot. The pattern is real, but its magnitude is inflated by the measurement design — like comparing a list of every restaurant you ate at this month (PSQI: “in general, how was your sleep this month?”) to a photo of what you ate tonight (PSG: “this is your sleep architecture from 10 PM to 6 AM on Tuesday”). The discrepancy between the list and the photo is partly a discovery about memory, not about dinner. Origin: brainstorm.

WarningLimitation: Byrnes2009’s Null Result May Reflect Tissue-Mismatch, Not CNS Confinement

Byrnes2009’s null result — zero transcriptomic signal in peripheral blood leukocytes in 44 MZ twin pairs — is cited as the strongest evidence for the CNS confinement model. But the null could reflect tissue-mismatch rather than confinement: peripheral blood leukocytes are primarily immune cells. If the primary pathology is in muscle (mitochondrial), endothelium (microvascular), brainstem (autonomic), or enterochromaffin cells (serotonin), blood leukocyte transcriptomics would show nothing — not because the signal is CNS-confined, but because the sampled cell type doesn’t express the relevant genes.

The null result is strong evidence against a systemic transcriptomic signal in blood leukocytes specifically. It is weak evidence for CNS confinement — tissue-mismatch is an alternative explanation that does not require the pathology to be CNS-confined. Muscle biopsy transcriptomics, endothelial cell isolation, or CSF proteomics might reveal signals that PBMC bulk transcriptomics misses.

Certainty: 0.50 — the tissue-mismatch interpretation is parsimonious and consistent with known multi-system involvement. The CNS confinement model survives as one explanation, not the only one. Severity applicability: n/a — methodological interpretation concern.

Consequence: The most famous null result in ME/CFS research — “no signal in blood” — may have been measuring the wrong cells. It’s like looking for evidence of myocardial infarction in a toenail clipping. The signal was there — just in muscle, endothelium, or brainstem, not the blood leukocytes they assayed. The CNS confinement model should acknowledge tissue-mismatch as an alternative explanation for the null. Origin: brainstorm.

WarningLimitation: The Discrepancy Index’s Severity Dependency May Be Fatal for Subtyping

The discrepancy index may be lower in mild patients (low objective dysfunction, high subjective report) and systematically lower in severe patients (high objective dysfunction, high subjective report). If so, the index measures disease severity, not mechanism subtype. The subtyping hypothesis — high discrepancy → dysautonomia-dominant, low discrepancy → immune-inflammatory — may appear correct only because mild patients are predominantly dysautonomia-dominant (early disease, autonomic dysfunction before measurable organ damage) and severe patients are predominantly immune-inflammatory (later disease, accumulated damage). Note: this disease-natural-history model (dysautonomia-first → immune-inflammatory-later) is an illustrative scenario constructed to demonstrate the confound — it is not an established trajectory. The index’s apparent subtyping power would be perfectly confounded with disease stage — it adds zero information beyond symptom severity and illness duration.

The contradictory predictions within this chapter illustrate the problem: the subtyping hypothesis predicts high-discrepancy patients are dysautonomia-dominant; the trajectory-predictor speculation (Discrepancy Magnitude as Illness-Trajectory Predictor) predicts high-discrepancy patients have better prognosis. Both could be true if high discrepancy simply means mild disease — and mild patients are more likely to improve AND more likely to have dysautonomia as their dominant mechanism. The index’s interpretation is ambiguous until severity is controlled.

Certainty: 0.45 — the severity confound is acknowledged but its magnitude and direction are unknown. Without severity-stratified validation, the subtyping hypothesis is unfalsifiable — post-hoc threshold adjustment could confirm any result. Severity applicability: the confound IS the severity-applicability problem.

Consequence: If all the high-discrepancy patients are just the mild patients — and mild patients respond better to everything — the discrepancy index didn’t discover the dysautonomia subtype. It rediscovered “less-sick people get better more easily.” This is not a breakthrough; it is a tautology wearing a regression coefficient. Until severity-stratified data exist, the subtyping hypothesis must be treated as a research question, not a diagnostic framework. Origin: brainstorm.

WarningLimitation: The Achievement’s Certainty (0.75) Overestimates the Independence of the Evidence Base

The existing achievement assigns 0.75 certainty to the claim that the discrepancy is replicated across 5 domains. Each domain’s evidence has correlated sampling sources that weaken the independence assumption: sleep — Watson2003 and Armitage2009 use the same twin cohort, quasi-independent at best; autonomic — Novak2024 (n=2,627) is the strongest single study but is one study on one instrument pair; activity — Vergauwen2021 found rs=0.35–0.38, which is significantly greater than zero (p < 0.05) — this is a weak positive relationship, not a zero discrepancy; cognition — Teodoro2018 is a systematic review, not a primary study with original effect sizes for discrepancy magnitude; biomarker — Byrnes2009 is the cleanest null, but the tissue-mismatch concern weakens its interpretation.

The five domains do not provide five independent replications. They provide: one twin cohort (sleep), one large cross-diagnosis cohort (autonomic), one moderate-effect study (activity), one review (cognition), and one tissue-specific null (biomarker). The combined certainty 0.75 is inflated by treating partially overlapping evidence sources as fully independent. A more appropriate certainty after accounting for shared sampling biases, correlated evidence sources, and domain-level interpretative heterogeneity is 0.55–0.60.

Certainty: 0.55 — the evidence for the discrepancy is genuinely strong. The concern is about the interpretation of its strength, not the existence of the effect. Severity applicability: n/a — certainty calibration concern.

Consequence: The five legs of the discrepancy table look solid, but two were carved from the same piece of wood (the same twins), one is a summary of other people’s measurements, and one shows a weak relationship rather than a zero. The table stands — the discrepancy is real — but it wobbles more than 0.75 implies. The achievement’s certainty should be downgraded to 0.60 to accurately reflect the correlated evidence structure. Origin: brainstorm.

NoteOpen Question: Null Hypothesis: Does the Discrepancy Index Add Clinical Value Beyond Severity and Duration?

The null hypothesis is not that the discrepancy doesn’t exist — it does (certainty 0.60, The Subjective-Measurable Discrepancy is Not an Isolated Befuddlement — It Patterns Across Domains). The null is that the discrepancy index adds zero clinical value beyond existing assessment: diagnostic null — does not improve accuracy beyond IOM criteria; prognostic null — does not predict outcome better than severity and duration; therapeutic null — does not predict differential treatment response. If all three hold, the discrepancy index is scientifically interesting but clinically useless — an academic construct.

This is the default state for any novel clinical instrument: the burden of proof is on demonstrating that a new measure adds information beyond what clinicians already use. Until validated against the null, the discrepancy index framework — including the subtyping hypothesis, trajectory predictor, and diagnostic criterion proposal — is hypothesis-generation, not clinical guidance.

Certainty: n/a — null hypothesis assessment. Severity applicability: the null’s diagnostic and prognostic arms require severity-stratified testing.

Falsifiable prediction: The discrepancy index must demonstrate utility in at least one of: (a) diagnostic — improves sensitivity/specificity by ≥5 percentage points over IOM; (b) prognostic — predicts outcome AUC ≥ 0.65, exceeding severity+duration (ΔAUC ≥ 0.05); (c) therapeutic — stratified treatment yields ≥0.3 SD outcome difference vs. unstratified. If all three fail, the index is academically interesting but clinically null.

Consequence: If the discrepancy score cannot tell clinicians anything they don’t already know from talking to the patient and looking at the chart, it belongs in a journal’s methods section — not in a clinical assessment. The most important thing to do is not to refine the index; it is to test whether the index adds anything. Origin: brainstorm — null hypothesis assessment.

NoteOpen Question: Is the Discrepancy an ME/CFS-Specific Signal or Fully Explained by Known Confounds?

Controlling for known confounds — illness severity, symptom fluctuation variance (severity × coefficient of variation), psychiatric comorbidity (HADS anxiety/depression), medication count and class (beta-blockers suppress HR → widen autonomic discrepancy; hypnotics alter sleep architecture → widen sleep discrepancy), age, and gender — the ME/CFS-specific residual discrepancy may be zero. The discrepancy would be a quantitative amplification of normal human interoceptive noise driven by disease severity, not a qualitative shift in how patients perceive their physiology.

If the ME/CFS discrepancy disappears after confound control, the entire discrepancy framework collapses to “severe illness produces large subjective-objective mismatch — and this is true in any disease, not just ME/CFS.” If a disease-specific residual survives after full confound control, the discrepancy is a genuine property of ME/CFS physiology beyond what severity alone predicts.

Certainty: 0.40 — each individual confound is plausible (severity: face validity; fluctuation variance: peak-end rule; psychiatric comorbidity: symptom amplification, certainty 0.60; medication effects: well-characterized pharmacology), but the collective explanatory power of all confounds is unknown. Severity applicability: the test requires severity-stratified analysis.

Falsifiable prediction: Regress discrepancy index on severity, symptom fluctuation variance, HADS-A, HADS-D, medication count/class, age, and gender. If R² ≥ 0.60 and the ME/CFS vs. healthy control coefficient becomes non-significant (p ≥ 0.01), the discrepancy is fully explained by confounds. If ME/CFS status remains significant after full confound control, a disease-specific residual survives.

Consequence: If the discrepancy is fully explained by confounds, researchers should stop investigating it as an ME/CFS mechanism and redirect resources toward understanding why severe illness in general produces subjective-objective mismatch — this would have implications for multiple contested diagnoses, not just ME/CFS. If a disease-specific residual survives, the discrepancy IS an ME/CFS mechanism worth dedicated investigation. Origin: brainstorm — null hypothesis assessment.

References

Armitage, Roseanne, Carol Landis, Robert Hoffmann, Martha Lentz, Nathaniel Watson, Jack Goldberg, and Dedra Buchwald. 2009. “Power Spectral Analysis of Sleep EEG in Twins Discordant for Chronic Fatigue Syndrome.” Journal of Psychosomatic Research 66 (1): 51–57. https://doi.org/10.1016/j.jpsychores.2008.08.004.
Byrnes, Andrea, Amber Jacks, Karin Dahlman-Wright, Birgitta Evengård, Fred A. Wright, Nancy L. Pedersen, and Patrick F. Sullivan. 2009. “Gene Expression in Peripheral Blood Leukocytes in Monozygotic Twins Discordant for Chronic Fatigue: No Evidence of a Biomarker.” PLoS ONE 4 (6): e5805. https://doi.org/10.1371/journal.pone.0005805.
Campen, C. (Linda) M. C. van, Peter C. Rowe, and Frans C. Visser. 2021. “Deconditioning Does Not Explain Orthostatic Intolerance in ME/CFS (Myalgic Encephalomyelitis/Chronic Fatigue Syndrome).” Journal of Translational Medicine 19 (1): 193. https://doi.org/10.1186/s12967-021-02819-0.
Friedberg, Fred, Jenna L. Adamowicz, Patricia Bruckenthal, Maria Milazzo, Sameera Ramjan, and Daniel Quintana. 2022. “Nonimprovement in Chronic Fatigue Syndrome: Relation to Activity Patterns, Uplifts and Hassles, and Autonomic Dysfunction.” Psychosomatic Medicine 84 (6): 669–78. https://doi.org/10.1097/PSY.0000000000001082.
Moezzi, Adam, Anna Ushenkina, Anders Widgren, Jonas Bergquist, et al. 2025. “Haptoglobin Phenotypes and Structural Variants Associate with Post-Exertional Malaise and Cognitive Dysfunction in Myalgic Encephalomyelitis.” Journal of Translational Medicine 23 (1). https://doi.org/10.1186/s12967-025-07006-z.
Novak, Peter, David M. Systrom, Sadie P. Marciano, Amanda Knief, Donna Felsenstein, Matthew P. Giannetti, Matthew J. Hamilton, et al. 2024. “Mismatch Between Subjective and Objective Dysautonomia.” Scientific Reports 14 (1): 2513. https://doi.org/10.1038/s41598-024-52368-x.
Omdal, R et al. 2026. “Persistent Fatigue in Long-COVID Is Not Associated with Peripheral Inflammatory or Cellular Stress Biomarkers: A Cross-Sectional Controlled Study.” Brain, Behavior, & Immunity – Health.
Paffrath, Anna, Lisa Kim, Claudia Kedor, Elisa Stein, Rebekka Rust, Helma Freitag, Uta Hoppmann, et al. 2024. “Impaired Hand Grip Strength Correlates with Greater Disability and Symptom Severity in Post-COVID Myalgic Encephalomyelitis/Chronic Fatigue Syndrome.” Journal of Clinical Medicine 13 (7): 2153. https://doi.org/10.3390/jcm13072153.
Teodoro, Tiago, Mark J. Edwards, and Jeremy D. Isaacs. 2018. “A Unifying Theory for Cognitive Abnormalities in Functional Neurological Disorders, Fibromyalgia and Chronic Fatigue Syndrome: Systematic Review.” Journal of Neurology, Neurosurgery & Psychiatry 89 (12): 1308–19. https://doi.org/10.1136/jnnp-2017-317823.
Vergauwen, Kuni, Ivan P. J. Huijnen, Rob J. E. M. Smeets, Daphne Kos, Inge van Eupen, Jo Nijs, and Mira Meeus. 2021. “An Exploratory Study of Discrepancies Between Objective and Subjective Measurement of the Physical Activity Level in Female Patients with Chronic Fatigue Syndrome.” Journal of Psychosomatic Research 144: 110417. https://doi.org/10.1016/j.jpsychores.2021.110417.
Watson, Nathaniel F., Clemma Jacobsen, Jack Goldberg, Vishesh Kapur, and Dedra Buchwald. 2004. “Subjective and Objective Sleepiness in Monozygotic Twins Discordant for Chronic Fatigue Syndrome.” Sleep 27 (5): 973–77. https://doi.org/10.1093/sleep/27.5.973.
Watson, Nathaniel F., Vishesh Kapur, Lester M. Arguelles, Jack Goldberg, Douglas F. Schmidt, Roseanne Armitage, and Dedra Buchwald. 2003. “Comparison of Subjective and Objective Measures of Insomnia in Monozygotic Twins Discordant for Chronic Fatigue Syndrome.” Sleep 26 (3): 324–28. https://doi.org/10.1093/sleep/26.3.324.