AI-Driven Multi-Omics Integration: The BioMapAI Model

BioMapAI, a supervised deep neural network by Xiong, Oh, Bateman, Unutmaz and colleagues (Xiong et al. 2025), is the most comprehensive AI integration of ME/CFS multi-omics to date — a 4-year longitudinal study of 249 participants (96 controls, 153 patients) combining gut metagenomics, plasma metabolomics, immune profiling, blood laboratories, and 12 clinical symptoms, with held-out and four-cohort external validation. Its explanatory power lies in integration: no single omics layer separated patients from controls, yet the integrated model reached an area-under-the-curve of 0.99. The model’s explainable connectivity map and its dysbiotic host-microbiome readout are developed below.

TipAchievement: BioMapAI: Explainable AI Integration of Five Omics Layers

The most comprehensive AI-driven integration of ME/CFS multi-omics data to date comes from Xiong, Oh, Bateman, Unutmaz, and colleagues (Xiong et al. 2025), building on their earlier cross-sectional microbiome-host analysis (Xiong et al. 2023). Their model, BioMapAI, is a supervised deep neural network trained on a 4-year longitudinal dataset from 249 participants (96 healthy, 153 ME/CFS; 515 timepoints; 1,471 biosamples), integrating gut metagenomics, plasma metabolomics, immune cell profiling of peripheral blood mononuclear cells, routine blood laboratory data, and 12 scored clinical symptoms.

The model’s headline figure — classification of ME/CFS versus healthy controls at an area under the ROC curve of 0.99 — must be read with the symptom-circularity caveat in BioMapAI’s Headline Accuracy Includes a Symptom-Circularity Component: the model’s twelve output heads are the twelve diagnostic symptoms, so part of the 0.99 is the model restating the diagnosis. The robust scientific content is elsewhere. The genuine contributions are two: (a) a principal-coordinate analysis showing that each omics layer alone could not separate patients from controls — the discriminatory signal is distributed, not in any one assay; and (b) validation on held-out data and on four independent external cohorts (Guo and Raijmakers microbiome; Germain and Che metabolome), an unusually strong standard for a computational model in this field. We do not reproduce a single pure-omics or external-cohort AUC figure here because our sources do not quote one; the strength claimed rests on the external validation and the negative single-layer result, not on the 0.99 headline.

The methodological significance is threefold. First, it demonstrates the feasibility of distributing ME/CFS signal across omics layers — a direct counterpoint to the “no validated biomarker has been established to diagnostic standard” position, and consistent with the discrepancy framework’s claim that single-omics resting measurements miss the disease (The Subjective-Measurable Discrepancy as a Diagnostic Index). Second, it provides an explicit, reproducible integration framework future studies can apply to independent cohorts. Third, its longitudinal design (followed 3-4 years) reveals marked temporal heterogeneity in ME/CFS — fatigue persistently severe, other symptoms fluctuating nonlinearly — which single-timepoint studies cannot capture. These methodological demonstrations withstand scrutiny; the specific biological connectivity map built from them is a forward-looking hypothesis (see Dysbiotic Microbiome-Immune-Metabolome Crosstalk with a Benzoate-to-Hippurate Axis) whose symptom attributions were not held-out validated.

Certainty: 0.80 — the methodological claims (distributed signal, external validation, longitudinal heterogeneity) are high-confidence; the certainty does not rest on the 0.99 headline or on the unvalidated connectivity-map attributions (see BioMapAI’s Headline Accuracy Includes a Symptom-Circularity Component). Severity applicability: the cohort spans the ME/CFS spectrum but is not formally stratified by severity — treat as representing ambulatory disease, with unknown bedbound applicability and no claim that the signature extends to severe or bedbound patients.

Consequence: For a disease that has long lacked a single validated biomarker, this shows a distributed biological signal is plausible — spread across the gut, blood, and immune cells, and only visible when the layers are read together. It does not yet give a blood test; it gives a validated analytical template and strengthens the case that ME/CFS is a multi-system biological illness.

ImportantHypothesis: Dysbiotic Microbiome-Immune-Metabolome Crosstalk with a Benzoate-to-Hippurate Axis

BioMapAI’s connectivity map links three physiological compartments in ME/CFS (Xiong et al. 2025). Compared to healthy controls, patients show decreased butyrate and branched-chain amino acid microbial pathways and increased tryptophan and benzoate pathways, changes the model ties to gastrointestinal symptoms. In parallel, the immune system loses health-associated interactions with Th17, Th22, and regulatory T cells and shifts toward heightened inflammatory activity in mucosal-associated invariant T (MAIT) and gamma-delta T cells. A network-level change is an increase in the benzoate-to-hippurate transformation, which the model attributes to emotional disturbance, sleep problems, and fatigue. Important caveat: these symptom attributions and the direction of the links are model output — the model’s SHAP connectivity map is correlational, its symptom attributions were not held-out validated (see BioMapAI’s Headline Accuracy Includes a Symptom-Circularity Component), and the model does not establish which compartment drives which, or whether any of it is cause versus consequence.

These observations knit together hypotheses already present in this paper. Decreased butyrate and branched-chain amino acid capacity is established in the gut-microbiome literature (Deficient Butyrate-Producing Capacity, (Xiong et al. 2023)). Increased tryptophan flux is consistent with the kynurenine/serotonin-diversion hypothesis ((Kavyani et al. 2022), (Taenzer et al. 2023)). The innate-like T-cell signature echoes single-cell and flow-cytometry evidence of T-cell remodeling in post-infectious ME/CFS, including galectin-9-associated depletion of gamma-delta and MAIT cells in Long COVID with ME/CFS (Shahbaz2026single?)-cell-immune and cytotoxic-T-cell alterations tracking symptom severity (Lee et al. 2025). The benzoate-to-hippurate axis itself has only indirect support: elevated hippuric acid is reported in post-infectious syndromes (Brigo et al. 2025) and in fibromyalgia (Malatji et al. 2017), but these are small (n=41; n=18) neighbouring-population studies that show hippurate elevation in other conditions — they do not independently establish a benzoate-to-hippurate axis tied to ME/CFS symptoms.

The hypothesis — and it is a hypothesis, not an established mechanism — is that these changes may not be parallel coincidences but part of a coordinated host-microbiome dysregulation: gut dysbiosis might produce a microbial-metabolite signal (decreased butyrate/BCAA, increased tryptophan/benzoate) that is read by the immune system and reflected in plasma lipids and bile acids, with a resulting inflammatory tone (MAIT/gamma-delta IFN-gamma and granzyme A) contributing to symptom burden, and the benzoate-to-hippurate transformation (hepatic glycine conjugation of microbially-produced benzoate) possibly serving as a readout. Direction of causation is unresolved and may run part or all of both ways: symptoms and inactivity could alter the gut and diet; benzoate is a common food additive and gut-processed species, so dietary intake is an uncontrolled confound. A simpler alternative is a single shared upstream perturbing agent producing correlation across compartments without directed crosstalk.

Certainty: 0.30 — the component pillars (butyrate, tryptophan, T-cell inflammation) are each supported by separate literature, but the specific integrated “crosstalk” and benzoate-to-hippurate axis rest on a single model’s correlational connectivity map plus two tiny neighbouring-population pilots (Brigo 0.26, Malatji 0.20, which are not replication) with unaddressed reverse-causation and diet confounds. Severity applicability: unknown — cohort not stratified by severity; severity-specific claims cannot be made.

Falsifiable prediction: In an independent ME/CFS cohort with paired metagenomics and plasma/urine metabolomics, patients vs controls should show (a) reduced abundance of butyrate- and BCAA-producing taxa with reduced corresponding metabolites, (b) increased benzoate and hippurate, and (c) a significant positive correlation between plasma/urinary hippurate (or a benzoate-pathway score) and severity of fatigue, sleep disturbance, or emotional disturbance, after adjusting for dietary benzoate intake and overall dysbiosis burden. If the benzoate-to-hippurate axis is only a marker of global dysbiosis (or of diet), it should not independently predict symptom severity after those adjustments.

Consequence: Until the hypothesis survives independent, diet-adjusted replication, it is a research readout to test, not a measure to track illness burden in the clinic. If it does replicate, it could become a urinary/blood readout of the microbial-immune interface — but that is a future step, not a present capability.

NoteOpen Question: Are Innate-Like T Cells Activated or Depleted in Post-Infectious ME/CFS?

BioMapAI finds heightened inflammatory activity in MAIT and gamma-delta T cells secreting IFN-gamma and granzyme A in ME/CFS (Xiong et al. 2025). In contrast, single-cell RNA-sequencing of Long COVID with ME/CFS finds galectin-9-associated depletion of gamma-delta and MAIT cells, alongside monocyte and NK remodeling and T-cell exhaustion (Shahbaz2026single?)-cell-immune. Cell-type frequency correlates with symptom severity in related cytotoxic-T-cell data (Lee et al. 2025). These observations need not be contradictory at all — they may measure different dimensions. BioMapAI’s signal is a per-cell functional readout (IFN-gamma and granzyme A secretion); Shahbaz’s is a frequency/count readout (number of MAIT and gamma-delta cells). A frequency-based count is blind to per-cell hyperactivation: a T-cell population that is small in number but individually hyperactivated would read as “depleted” on a count yet “activated” on function, so both findings could be simultaneously true.

If they are in genuine tension, the simplest reading is that the two studies are not tracking one within-disease timecourse but comparing different things: different populations (ME/CFS broadly vs Long COVID with ME/CFS), different measurement platforms (immune-profiling inside a multi-omics model vs single-cell RNA frequency/state), and likely different disease durations. A more complex, unproven reconciliation is the “dynamic trajectory” — an innately-reactive T-cell population hyperactivated early or during flares, then driven toward exhaustion and depletion over the chronic course by sustained antigen exposure. This ordered sequence is one candidate, not an established account; neither study has serial within-subject data to confirm it.

The unresolved question is whether the correct reading — and with it any therapeutic implication — is activation, depletion, an ordered sequence, or simply two different cell-kinds being counted and measured differently. An intervention modulating T-cell activation/exhaustion balance would depend on which state dominates, but any such decision is premature and requires prospective interventional testing.

Certainty: n/a — this is an open research question posed by two bodies of evidence. Severity applicability: unknown — neither study stratifies by severity; both sample ambulatory post-infectious cohorts.

Consequence: Only serial, within-cohort measurement that tracks MAIT and gamma-delta T cells over disease duration and captures both cell frequency and per-cell functional output can resolve this. Until then, “hyperactive innate-like T cells” and “exhausted/depleted innate-like T cells” remain possibly-orthogonal descriptions — pointing to opposite therapeutic strategies if they are alternatives, and to a combined intervention problem if they are sequential.

ImportantHypothesis: A Cross-Layer Coordination Signature, Not Any Single Marker, Defines ME/CFS

BioMapAI’s principal-coordinate analysis showed that each omics layer analyzed alone leaves ME/CFS patients largely indistinguishable from healthy controls — the microbiome, the plasma metabolome, and the immune profile each overlap substantially; only integration separates the groups with high accuracy (Xiong et al. 2025). Two readings are possible, and the model alone cannot distinguish them:

  • A coordination reading: what is pathological is a rewiring of the normal correlated relationships across compartments — the association structure is abnormal even though each layer’s marginal distribution is largely preserved.
  • A concatenation reading (the simpler, reductionist alternative): the integrated model wins for a purely technical reason — more feed-in features give it greater statistical capacity to detect a set of individually-weak main effects, requiring no cross-layer coordination at all.

The paper’s PCoA demonstrates only that marginals are non-separating; it does not show the cross-layer correlations are themselves rewired. An integrated AUC can exceed any single layer even when there is zero true cross-layer interaction. The coordination-signature hypothesis is therefore contingent and untested: the “rewiring” interpretation would need to survive a test that separates interaction effects from summed single-layer main effects, which the reported analysis does not perform.

This reframes the biomarker search from “find the one abnormal node” to “test whether the association pattern across nodes adds value beyond the individual layers,” and it is consistent with the discrepancy framework’s finding that single-resting-measurement domains miss the disease (The Subjective-Measurable Discrepancy as a Diagnostic Index). But it does not automatically make the strong “coordination” claim true.

Certainty: 0.35 — the observation that single layers are non-discriminative while integration works is real and reported, but the strong form (that the cross-layer correlation structure is pathological and diagnostic) has not been demonstrated and may reflect only additive feature power. Severity applicability: unknown — model cohort not severity-stratified.

Falsifiable prediction: In an independent multi-omics ME/CFS cohort, to discriminate the interpretations the analysis must use two baselines, not just each single layer: (a) each single omics layer alone should fail to separate patients from controls; (b) a flat feature-concatenation model (all features as one input vector, no interaction modeling) should be meaningfully worse than a model that explicitly models cross-layer interactions. If feature concatenation matches the interaction model, the “coordination/rewiring” hypothesis collapses to additive main effects and should be rejected; if a single omics layer alone discriminates as well as integration, the disease is better described by a single-compartment signature.

Consequence: Until the interaction-versus-concatenation test is run, the honest claim is only “integration works better than any single layer” — which is a methodological demonstration, not proof that ME/CFS is defined by a cross-layer coordination structure. The decades-old “no validated biomarker” finding is not thereby reframed as a “wrong measurement strategy” unless replication succeeds and overfitting is excluded.

CautionSpeculation: Longitudinal Symptom Stability Is Symptom-Specific: A Stable Fatigue Core with a Labile Emotional Component

BioMapAI’s 3–4 year longitudinal data indicated that ME/CFS symptoms do not behave uniformly over time: fatigue was persistently and consistently severe across years, whereas emotional dysregulation was highly variable and unstable (Xiong et al. 2025). A primary non-biological alternative must be stated first: this pattern may be a measurement artifact. Fatigue is typically scored on a coarse ordinal scale that, in severe ME/CFS cohorts, sits near its ceiling — a scale at ceiling cannot fluctuate downward or upward much, so “stability” may reflect a ceiling effect, whereas emotional-symptom items have headroom in-range and therefore look more variable. The apparent decoupling could be psychometric (item wording, scale range, floor/ceiling) rather than two distinct biologies, and it is compounded by symptom-circularity (the fatigue score is partly the diagnosis itself).

If the pattern were to survive that artifact check, it would then suggest different biological timescales: a relatively fixed, stable driver for fatigue (e.g., a chronic metabolic or neuroimmune deficit) versus a more labile driver for emotional symptoms (e.g., fluctuating gut-metabolic state, autonomic reactivity, or state-dependent inflammatory tone). Both readings are hypotheses; the biological one is contingent on excluding the measurement one.

The methodological implication is that longitudinal sampling is needed to characterize the labile domain, and that “emotional instability” should not be conflated with the stable energy deficit. This is a study-design point, not a clinical-assessment claim.

Certainty: 0.25 — the temporal pattern is reported by the model’s single US cohort, but it has not been independently replicated, and the measurement-artifact explanation (scale ceiling/coarseness, symptom-circularity) has not been excluded. Severity applicability: unknown — cohort not severity-stratified.

Falsifiable prediction: In a prospective longitudinal ME/CFS cohort with serial symptom scoring (e.g., monthly) over 12+ months using continuous, range-normalized fatigue and emotional scales (adequate headroom, no ceiling): test-retest stability (intraclass correlation over time) should be significantly higher for fatigue than for emotional items even after controlling for scale range; and a stable biological measure (e.g., a metabolomic or resting-blood marker) should correlate significantly more strongly with the time-stable fatigue score than with the fluctuating emotional score. If, after range normalization, fatigue is as labile as emotional symptoms (or if the stability vanishes under scale control), the “stable fatigue core” claim is falsified as a biological finding.

Consequence: If the stable-core/labile-shell pattern is real and survives the measurement-artifact check, it would shape trial design — supporting sparse serial sampling for the stable fatigue endpoint at the group level, with emotional symptoms requiring longitudinal capture — and would reframe “emotional symptoms in ME/CFS” as possibly a distinct, fluctuating physiology rather than a fixed psychological overlay. It is not a basis for single-visit clinical assessment of an individual patient, whose trajectory no group-level stability coefficient can certify.

WarningLimitation: BioMapAI’s Headline Accuracy Includes a Symptom-Circularity Component

The headline classification performance of BioMapAI (AUC=0.99) requires careful interpretation. The model’s twelve output heads are the twelve scored clinical symptoms — the same symptoms a clinician uses to establish the ME/CFS diagnosis. A model trained to reconstruct the very symptom set that defines the diagnosis will trivially separate cases from controls if it learns those symptom patterns, independent of any underlying biology. This is a partial circularity: the symptom-based head is partly restating the diagnostic criteria.

The honest test of biological signal is performance on the non-symptom omics layers — the metabolome, microbiome, and immune profiles — both alone and in the integrated model, plus external-cohort validation on omics data (Guo, Raijmakers microbiome; Germain, Che metabolome). We cannot quote a single pure-omics or external-cohort AUC here because our sources do not report one; that absence is itself a methodological gap — the chapter leans on an “omics signature” whose numeric strength the paper does not disclose, so its magnitude is undetermined rather than “real but modest.” Two cautions follow. First, the study’s PCoA shows no single omics layer separates patients from controls, which is consistent with the signal being distributed — but it does not quantify how much of the integrated performance reflects genuine cross-layer biology versus the technical gain any multi-feature ensemble gets from additive weak main effects (see A Cross-Layer Coordination Signature, Not Any Single Marker, Defines ME/CFS). Second, it is possible the omics-only integrated signal is substantially weaker than the headline; the paper does not establish otherwise.

A second, related limitation: the symptom-specific biomarker attributions (which feature drives which symptom) were computed on the full dataset and explicitly not validated on held-out data in the paper. They are exploratory, not established.

Certainty: n/a — methodological caveat. Severity applicability: n/a — model-design concern.

Consequence: Do not quote “AUC 0.99” as if it were a blood test. The defensible content is the negative result that no single omics layer separates patients from controls, the external-cohort validation, and the explainable connectivity map as a hypothesis generator — but the 0.99 figure is inflated by the model learning the diagnostic symptoms it was built around, the pure-omics integrated strength is not reported, and the connectivity-map attributions were not held-out validated. Readers and clinicians should rely on independently-replicated non-symptom results, which are not yet established.

References

Brigo, Natascha, Wolfram Mayr, Maja Taenzer, Judith Löffler-Ragg, Andrea Schroll, Sabine Engl, Burkhard Schütz, et al. 2025. “Concentrations of Uremic Bacterial Metabolites in Patients with Post-COVID-19 Syndrome.” Frontiers in Cellular and Infection Microbiology 15: 1582972. https://doi.org/10.3389/fcimb.2025.1582972.
Kavyani, Zahra, Ekua W Brenu, Donald R Staines, and Sonya M Marshall-Gradisnik. 2022. “Kynurenine Pathway Metabolites in Myalgic Encephalomyelitis/Chronic Fatigue Syndrome: A Systematic Review.” Journal of Translational Medicine 20 (1): 512. https://doi.org/10.1186/s12967-022-03712-3.
Lee, Ji-Sook, Eliana Lacerda, Caroline Kingdon, Erik Abken, Gaia Susannini, Hazel M. Dockrell, Luis Nacul, and Jacqueline M. Cliff. 2025. “Abnormal T-Cell Activation and Cytotoxic T-Cell Frequency Discriminate Symptom Severity in Myalgic Encephalomyelitis/Chronic Fatigue Syndrome.” Journal of Translational Medicine 23: 669. https://doi.org/10.1186/s12967-025-07507-x.
Malatji, Bontle G., Helgard Meyer, Shayne Mason, Udo F. H. Engelke, Ron A. Wevers, Mari van Reenen, and Carolus J. Reinecke. 2017. “A Diagnostic Biomarker Profile for Fibromyalgia Syndrome Based on an NMR Metabolomics Study of Selected Patients and Controls.” BMC Neurology 17 (1): 88. https://doi.org/10.1186/s12883-017-0863-9.
Taenzer, M., J. Löffler-Ragg, A. Schroll, P. Monfort-Lanzas, S. Engl, G. Weiss, N. Brigo, and K. Kurz. 2023. “Urine Metabolite Analysis to Identify Pathomechanisms of Long COVID: A Pilot Study.” International Journal of Tryptophan Research 16: 11786469231220781. https://doi.org/10.1177/11786469231220781.
Xiong, Ruoyun, Elizabeth Aiken, Ryan Caldwell, Suzanne D. Vernon, Lina Kozhaya, Courtney Gunter, Lucinda Bateman, Derya Unutmaz, and Julia Oh. 2025. “AI-Driven Multi-Omics Modeling of Myalgic Encephalomyelitis/Chronic Fatigue Syndrome.” Nature Medicine 31 (9): 2991–3001. https://doi.org/10.1038/s41591-025-03788-3.
Xiong, Ruoyun, Courtney Gunter, Elizabeth Fleming, Suzanne D. Vernon, Lucinda Bateman, Derya Unutmaz, and Julia Oh. 2023. “Multi-’omics of Gut Microbiome-Host Interactions in Short- and Long-Term Myalgic Encephalomyelitis/Chronic Fatigue Syndrome (ME/CFS) Patients.” Cell Host & Microbe 31 (2): 273–287.e5. https://doi.org/10.1016/j.chom.2023.01.001.