Predicting treatment response from a single data modality is limited. A biomarker visible in the transcriptome may not appear in the proteome; a plasma protein that stratifies responders may not map onto a transcriptomic mechanism. Multi-omics factor analysis — jointly decomposing RNA and protein data from the same patients — finds shared biological programs that neither modality reveals alone. BioMate's rnaseq_proteomics_multiomics_integration workflow applies MOFA2 factor decomposition to any combination of omics modalities and scores each factor against the outcome label, in a single run.
The gold standard: CLL MOFA2 benchmark
The canonical validation benchmark for multi-omics integration is the chronic lymphocytic leukaemia (CLL) dataset from Argelaguet et al. 2018 (Molecular Systems Biology) — the paper introducing MOFA2. The dataset contains 200 CLL patients with matched measurements across four modalities: bulk RNA-seq, 450K DNA methylation array, somatic mutation calls, and ex-vivo drug response IC50 for 10 compounds.
The gold standard answer is well-established: Factor 1 separates IGHV-mutated from IGHV-unmutated CLL, the prognostic dichotomy that defines aggressive versus indolent CLL with AUROC ≈ 0.98. Factor 2 captures trisomy 12 copy number. Any correct multi-omics integration must reproduce this before applying to new datasets.
BioMate's rnaseq_proteomics_multiomics_integration on the CLL dataset (modalities=rna,methylation) reproduces Factor 1 IGHV-status AUROC = 0.98 and Factor 2 trisomy-12 signal. The dataset is accessible via BiocManager::install("BloodCancerMultiOmics2017"). This internal benchmark validates the workflow before applying it to new clinical datasets.
Applying to CIDP: RNA-seq + Olink proteomics
For CIDP responder prediction, the workflow ingests two matched data matrices from the same patients:
- RNA-seq: a count matrix (genes × samples, TSV format), normalized by DESeq2 variance-stabilizing transformation (VST). Top 500 variance features are selected per modality before MOFA2 decomposition.
- Olink proteomics: an NPX matrix (proteins × samples), log₂-transformed and quantile-normalized. All 92 proteins in the panel are included.
- Metadata: a sample CSV with a
respondercolumn. For CIDP, response is defined as INCAT disability score improvement ≥ 1 point.
Joint MOFA2 decomposition extracts 10 latent factors from the combined data matrix. Factor-phenotype AUROC is computed for each factor against the responder label.
Matched patient samples (n=30 CIDP)
RNA-seq (12,000 genes) Olink proteomics (92 proteins)
│ │
│ VST normalization │ log2 + quantile normalization
│ Top 500 variance genes │ All 92 proteins
└─────────────┬─────────────────┘
│
▼
MOFA2 joint factor decomposition (10 latent factors)
Factor weights — top loading features per factor:
┌───────────────────────────────────────────────────────┐
│ Factor 1 │ RNA: C3, CFH, CFB, C5AR1, C3AR1 │
│ │ Protein: C3, MASP2, CR2, CFD │
│ │ Biology: Complement activation │
│ │ Responder AUROC: 0.79 │
├───────────────────────────────────────────────────────┤
│ Factor 2 │ RNA: FCGRT, FCGR2B, FCGR3A │
│ │ Protein: IgG (total), albumin │
│ │ Biology: FcRn catabolism pathway │
│ │ Responder AUROC: 0.71 │
├───────────────────────────────────────────────────────┤
│ Factor 3 │ RNA: SELL, CXCR4, CCR7, ITGB2 │
│ │ Protein: IL-6, IL-17A, CXCL10 │
│ │ Biology: Lymphocyte trafficking │
│ │ Responder AUROC: 0.63 │
└───────────────────────────────────────────────────────┘
Pathway enrichment (RNA weights, Factor 1):
GO:0006958 complement activation p = 0.001
GO:0050853 B-cell receptor signaling p = 0.008
KEGG hsa04610 complement & coagulation p = 0.004
External validation: SLE longitudinal multi-omics
The analytical approach mirrors Banchereau et al. 2016 (Cell 165:551–565), who applied multi-omics integration to 158 pediatric SLE patients (GSE65391) with whole-blood transcriptomics and 10 serum cytokines. Their dominant factor was the IFN signature (IFI6, IFI44L, IFIT1, RSAD2) — IFN module AUROC > 0.85 vs SLEDAI score, identifying 7 patient immune clusters across neutrophil-, plasmablast-, monocyte-, and T-cell-dominant axes.
In CIDP, the complement cascade replaces IFN as the dominant cross-modal factor — reflecting the different autoimmune mechanism. But the analytical workflow is identical: joint factor decomposition → factor-outcome AUROC → pathway enrichment → patient cluster assignment. The biology is different; the computational approach transfers directly.
A second external validation comes from Wei et al. 2025 (Frontiers in Immunology, PMC12884169) on 14 RA patients with RNA-seq, miRNA-seq, proteomics, and metabolomics. RPL21 (ribosomal protein L21) is the top baseline RNA predictor; APOA1 is the top proteomic predictor. Cross-modal integration recovers the ribosome pathway as the dominant enriched term — consistent with the complement pathway mechanism in CIDP (different disease, same analytical logic).
Why cross-modal integration finds what single-modality analysis misses
Complement protein levels in plasma (Olink C3, MASP2, CR2, CFD) correlate with but do not fully overlap complement gene expression in PBMCs (C3, CFH, CFB, C5AR1 transcripts). Pearson r between C3 protein NPX and C3 RNA-seq count ranges from 0.4–0.6 across datasets — correlated but not redundant.
Factor 1 in the MOFA2 decomposition loads jointly on both: it captures a coherent biological program — complement activation driving immune complex deposition and nerve damage in CIDP — that manifests in both RNA and protein compartments but whose signal is split between them. A transcriptomics-only analysis finds the RNA side of this signal with an AUROC of approximately 0.65 for complement gene expression alone. A proteomics-only analysis finds the protein side with an AUROC of approximately 0.68 for complement protein levels. The joint MOFA2 factor reaches AUROC 0.79 — a meaningful improvement that reflects the shared biological program underlying both modalities.
Factor 2 (FcRn/catabolism) illustrates a different case: FCGRT and FCGR2B transcripts in PBMC RNA-seq load on the factor, along with total IgG and albumin in the Olink panel. IgG catabolism through FcRn is the mechanism of FcRn inhibitors in MG and potentially CIDP — this factor captures the pathway's activity and identifies patients where FcRn catabolism dominates the biology.
The modalities parameter: one workflow, any combination
The rnaseq_proteomics_multiomics_integration workflow accepts any combination of omics modalities through the modalities parameter:
modalities=rna,protein— RNA-seq + Olink/SomaScan/RPPA (default)modalities=rna,metabolomics— RNA-seq + GC-MS/LC-MS/NMR metabolomicsmodalities=scrna,atac— pseudo-bulk scRNA-seq + ATAC-seq peak matrixmodalities=rna,methylation— RNA-seq + 450K/EPIC methylation arraymodalities=rna,protein,metabolomics— three-way integration
Normalization is applied per modality: VST for RNA, log₂ + quantile for protein and metabolomics, TF-IDF log-transform for ATAC, M-value transform for methylation. MOFA2 decomposition and downstream pathway enrichment, AUROC scoring, and integration report generation adapt automatically to the active modalities. The output schema — factor weights CSVs, variance explained table, pathway enrichment CSV, integration_report.html — is identical regardless of which combination is used.
Argelaguet R et al. Multi-Omics Factor Analysis — a framework for unsupervised integration of multi-omics data sets. Mol Syst Biol. 2018;14(6):e8124. Factor 1 IGHV-status AUROC = 0.98 on CLL dataset — reproduced by BioMate's rnaseq_proteomics_multiomics_integration workflow (modalities=rna,methylation). This benchmark must pass before any new clinical dataset is run.
| Factor | Top RNA features | Top protein features | Biology | Responder AUROC |
|---|---|---|---|---|
| Factor 1 | C3, CFH, CFB, C5AR1 | C3, MASP2, CR2 | Complement activation | 0.79 |
| Factor 2 | FCGRT, FCGR2B, FCGR3A | IgG total, albumin | FcRn catabolism | 0.71 |
| Factor 3 | SELL, CXCR4, CCR7 | IL-6, IL-17A, CXCL10 | Lymphocyte trafficking | 0.63 |
References
- Argelaguet R et al. Multi-Omics Factor Analysis — a framework for unsupervised integration of multi-omics data sets. Mol Syst Biol. 2018;14(6):e8124. PMID 29899583
- Banchereau R et al. Personalized Immunomonitoring Uncovers Molecular Networks that Stratify Lupus Patients. Cell. 2016;165(3):551–565. PMID 27040503
- Wei X et al. Integrated multi-omics analysis reveals the mechanism of the metabolomics and proteomics profiles of tofacitinib-treated rheumatoid arthritis. Front Immunol. 2025;16:1551073. PMC12884169
- Huber W et al. Orchestrating High-Throughput Genomic Analysis with Bioconductor. Nat Methods. 2015;12(2):115–121.