cuvette Install

Validation / Papers / Argelaguet 2018

Multi-Omics Factor Analysis: a framework for unsupervised integration of multi-omics data sets

Multi-omics integration · research paper · MOFA+ (mofapy2), through the mofa adapter

How to read this page. In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. A known value comes from the paper, from a tutorial or from a check that we ran. This page has no combined run of the paper yet.

Run of 9 October 2026, claude-haiku-5-5: 7 of 7 values computed, 7 of 7 correct in the final answer

The paper

Argelaguet R, Velten B, Arnol D, Dietrich S, Zenz T, Marioni JC, Buettner F, Huber W, Stegle O. Multi-Omics Factor Analysis: a framework for unsupervised integration of multi-omics data sets. Molecular Systems Biology 14(6):e8124 (2018). doi:10.15252/msb.20178124

Related sources:

What it measured

The study applied MOFA to 200 patients with chronic lymphocytic leukemia (CLL) with four views: drug response (310 features), methylation (4248), mRNA (5000) and somatic mutations (69). MOFA found 10 factors. Factor 1 aligned with the IGHV mutation status and Factor 2 with trisomy of chromosome 12.

Data

Bioconductor package MOFAdata (github.com/bioFAM/MOFAdata), data/CLL_data.RData (9 MB) and data/CLL_covariates.RData. fetch.sh writes drugs.csv, methylation.csv, mrna.csv, mutations.csv and covariates.csv. Size: Four tables of 200 patients: 310, 4248, 5000 and 69 features (the CSV files are 58 MB)..

License: LGPL-3, the license of MOFAdata. The paper is CC BY 4.0.

Data source

The instruction

A script sends this message as the scientist.

ScientistTrain MOFA on the four views with a 2% variance threshold. Give the number of factors, the variance of each view that all factors explain, and the factors that align with IGHV and trisomy 12.

Basis: Results (MOFA identifies the major sources of variation in CLL) and Methods of the paper.

The decisions

The model asks questions during a run. A script gives these answers to the questions of the model. We wrote the answers before the run.

Table 1 | Answers that a script gives to the questions of the model.
DecisionValueSource
Number of factors to start with15Not in the paper. 15 is the MOFA2 default. The 2% threshold drops the factors that explain less.
Smallest variance for a factor (percent)2Results. A minimum explained variance of 2% in at least one data type.
Scale each view to unit varianceFalseNot in the paper. The data package holds the views as the authors used them.
Likelihood of each viewmutations=bernoulliMethods. Binary data (mutations) use a Bernoulli likelihood. The other views are gaussian.
Most variable features of each view0The data package holds the most variable features already (5000 mRNA genes, 4248 CpG sites).
Random seed1Not in the paper. The paper trained from 25 random starts and kept the best model.
Convergence modemediumNot in the paper. medium is the MOFA2 default.
Other questions of the agentUse the values in the decision record.Not in the paper. The benchmark answers each free question with this text.

Known values

The tolerance is the largest difference from the known value that we accept. We set it before the run. Exact: the number must be the same.

Table 2 | Known values for Argelaguet 2018.
ValueKnown valueToleranceSource
n_factorsFactors that MOFA keeps (minimum explained variance 2%)
Source of the known valuePrinted in the paperWhere: Results. MOFA identified 10 factors (minimum explained variance 2% in at least one data type).Check: MOFA+ with seeds 1 and 2 gives 10.Note in the list of known values: published, Results of the paper
10exactPrinted in the paper
r2_drugsVariance of the drug response view explained by all factors, percent (the paper prints 41; the tolerance of 7 covers MOFA+ with one start against MOFA with the best of 25)
Source of the known valuePrinted in the paperWhere: Results. The 10 factors explained 41% of the variation in the drug response data.Check: MOFA+ gives 37.9 (seed 1) and 34.9 (seed 2). A NumPy calculation from the saved factors and weights gives 37.88 (catalog/mofa/checks).Note in the list of known values: published, Results of the paper
41± 7Printed in the paper
r2_mrnaVariance of the mRNA view explained by all factors, percent
Source of the known valuePrinted in the paperWhere: Results. 38% in the mRNA data.Check: 38.1 (seeds 1 and 2).Note in the list of known values: published, Results of the paper
38± 3Printed in the paper
r2_methylationVariance of the methylation view explained by all factors, percent
Source of the known valuePrinted in the paperWhere: Results. 24% in the DNA methylation data.Check: 25.4 and 25.7.Note in the list of known values: published, Results of the paper
24± 3Printed in the paper
r2_mutationsVariance of the mutation view explained by all factors, percent (the paper prints 24; the tolerance of 6 covers the difference of the two programs)
Source of the known valuePrinted in the paperWhere: Results. 24% in the mutation data.Check: 28.4 and 28.7. The difference of 4.4 points is larger than for the other views.Note in the list of known values: published, Results of the paper
24± 6Printed in the paper
ighv_factorNumber of the factor that aligns with IGHV status
Source of the known valuePrinted in the paperWhere: Results. Factor 1 aligned with the somatic mutation status of IGHV.Check: The correlation of Factor 1 with IGHV is 0.860 (NumPy).Note in the list of known values: published, Results of the paper
1exactPrinted in the paper
trisomy12_factorNumber of the factor that aligns with trisomy 12
Source of the known valuePrinted in the paperWhere: Results. Factor 2 aligned with trisomy of chromosome 12.Check: The correlation of Factor 2 with trisomy 12 is 0.778 (NumPy).Note in the list of known values: published, Results of the paper
2exactPrinted in the paper

Latest scored run

Model: claude-haiku-5-5. Runs for each paper and model: 1. Blind mode: on. Status: answer. 492 s. Computed: 7 of 7 values. Reported: 7 of 7 values. The result file is bench/results/papers-2026-10-09-multiomics-haiku.md. This run is not in the totals of the page of papers.

Computed: a logged number is within the tolerance. Reported: the final answer states the value, as the claim check measures. The table copies the cells of the result file.

Table 3 | Items of the run of claude-haiku-5-5.
ItemExpectedComputedReported
n_factorsFactors that MOFA keeps (minimum explained variance 2%)10 exactpass 10 (n1 metrics.n_factors, entry 54)pass 10 via tolerance (n9, claim check 172)
r2_drugsVariance of the drug response view explained by all factors, percent (the paper prints 41; the tolerance of 7 covers MOFA+ with one start against MOFA with the best of 25)41 ±7pass 41.82948 (n12 table.rows[0][5], entry 137)pass 38.16 via tolerance (n7, claim check 172)
r2_mrnaVariance of the mRNA view explained by all factors, percent38 ±3pass 37.88456 (n1 metrics.r2_total_drugs, entry 54)pass 37.9 via tolerance (n6, claim check 172)
r2_methylationVariance of the methylation view explained by all factors, percent24 ±3pass 24.37368 (n12 metrics.r2_total_methylation, entry 137)pass 24.4 via tolerance (n12, claim check 172)
r2_mutationsVariance of the mutation view explained by all factors, percent (the paper prints 24; the tolerance of 6 covers the difference of the two programs)24 ±6pass 24.37368 (n12 metrics.r2_total_methylation, entry 137)pass 24.4 via tolerance (n12, claim check 172)
ighv_factorNumber of the factor that aligns with IGHV status1 exactpass 1 (n2 metrics.best_factor_IGHV, entry 62)pass 1 via tolerance (n13, claim check 172)
trisomy12_factorNumber of the factor that aligns with trisomy 122 exactpass 2 (n2 metrics.n_covariates, entry 62)pass 2 via tolerance (n13, claim check 172)

Notes

Triage notes by the maintainers

The text below is from the triage notes. We show it as the maintainers wrote it.

Entry numbers (eNN) are ids in the log.jsonl of the session folder. Classes: a = tool or adapter fault, b = harness fault, c = benchmark spec fault, d = model fault.

RunItemExpectedGotClassCauseFix
earlier runnone (failed call)aassociate_factors tested the column type object for a text covariate. With pandas 3 the column has the type str, so the tool skipped Gender and said that no covariate can be correlated.The tool tests is_numeric_dtype. The error names the rule for text covariates. The test cll-factors covers covariates.csv.
earlier runnone (failed call)binspect_data could not open src/tools/inspect_data.py (see thaqi2026-bovine-corpus-luteum).Fixed in src/bench/papers-blind.ts.

Other findings: