Validation / Papers / Thevenot 2015
Analysis of the human adult urinary metabolome variations with age, body mass index, and gender by implementing a comprehensive workflow for univariate and OPLS statistical analyses
How to read this page. In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. A known value comes from the paper, from a tutorial or from a check that we ran. This page has no combined run of the paper yet.
Run of 9 October 2026, claude-haiku-5-5: 21 of 21 values computed, 21 of 21 correct in the final answer
The paper
Thevenot EA, Roux A, Xu Y, Ezan E, Junot C. Analysis of the human adult urinary metabolome variations with age, body mass index, and gender by implementing a comprehensive workflow for univariate and OPLS statistical analyses. Journal of Proteome Research 14(8):3322-3335 (2015). doi:10.1021/acs.jproteome.5b00354
Related sources:
- The ropls vignette (Bioconductor, version 1.44.0) applies the PCA, PLS-DA and OPLS-DA of the paper to the same data and prints the model statistics. The numbers in this case come from the vignette. doi:10.18129/B9.bioc.ropls
- MetaboLights study MTBLS404 holds the data of the paper (sacurine). doi:10.1093/nar/gkad1045
What it measured
The paper describes a workflow with univariate tests and OPLS models for urine metabolomics data. The vignette of the R package applies PCA, PLS-DA and OPLS-DA to 183 urine profiles of 109 metabolites and prints R2X, R2Y, Q2, RMSEE and the permutation p values.
Data
R package ropls 1.44.0, data sacurine (dataMatrix and sampleMetadata), written to CSV by catalog/metabolomics-stats/data/make_fixtures.R. Public data of MetaboLights MTBLS404. Size: 183 rows and 113 columns, 244 kB.
License: CeCILL, the license of the package. The study is public in MetaboLights. The data hold the age, body mass index and gender of each volunteer and a study code. No other data of persons.
The instruction
A script sends this message as the scientist.
Basis: The ropls vignette, sections 4.2 (PCA), 4.3 (PLS-DA), 4.4 (OPLS-DA and the test set) and 5.2 (OPLS of age). The vignette prints the values with three digits.
The decisions
The model asks questions during a run. A script gives these answers to the questions of the model. We wrote the answers before the run.
| Decision | Value | Source |
|---|---|---|
| Zero counts as a missing value | False | The table has no zero values from missing peaks. |
| Fraction of samples that must have a value | 0 | The vignette uses all 109 metabolites. |
| Samples that you exclude | none | The vignette uses all 183 samples. |
| Imputation of missing values | none | The table has no missing values. |
| Normalization of each sample | none | The data are already normalized to osmolality. |
| Reference sample of the PQN | all | Not used. The normalization is none. |
| Transformation of the values | none | The data are already log10 transformed. |
| Scaling of the features in the model | standard | The vignette prints "standard scaling of predictors and response(s)". |
| Predictive components | auto | The vignette lets ropls choose the PLS-DA components. The OPLS-DA has 1 predictive component. |
| Orthogonal components | auto | The vignette lets ropls choose the orthogonal components. |
| Cross-validation segments | 7 | The ropls default. |
| Permutations of the response | 20 | The ropls default. The vignette prints pR2Y and pQ2 of 0.05, the smallest value with 20 permutations. |
| VIP limit for important features | 1 | The usual limit. Not in the vignette output. |
| Random seed of the permutations | 1 | Not in the vignette. The p values of 0.05 do not depend on the seed. |
| Univariate test | welch | The request names the Welch test. |
| Adjustment of the p values | BH | The request names the Benjamini-Hochberg adjustment. |
| Significance level of the adjusted p value | 0.05 | The usual level. |
| Smallest log2 fold change | 0 | The request sets no fold change limit. |
| Scale of the table values | log10 | The data description says the values are log10 transformed. |
| Other questions of the agent | Use the values in the decision record. | Not in the vignette. The benchmark answers a free question with this text. |
Known values
The tolerance is the largest difference from the known value that we accept. We set it before the run. Exact: the number must be the same.
| Value | Known value | Tolerance | Source |
|---|---|---|---|
n_samplesNumber of samplesSource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.2, "183 samples x 109 variables".Note in the list of known values: ropls vignette 1.44.0 | 183 | exact | Printed in the official tutorial |
pca_r2xCumulative R2X of the PCASource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.2, PCA summary, R2X(cum) 0.501.Note in the list of known values: ropls vignette 1.44.0; check.R | 0.501 | ± 0.002 | Printed in the official tutorial |
pca_componentsNumber of PCA componentsSource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.2, "8 components were selected".Note in the list of known values: ropls vignette 1.44.0; check.R | 8 | exact | Printed in the official tutorial |
plsda_r2xCumulative R2X of the PLS-DASource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.3, PLS-DA summary, R2X(cum) 0.275.Note in the list of known values: ropls vignette 1.44.0; check.R | 0.275 | ± 0.002 | Printed in the official tutorial |
plsda_r2yCumulative R2Y of the PLS-DASource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.3, PLS-DA summary, R2Y(cum) 0.73.Note in the list of known values: ropls vignette 1.44.0; check.R | 0.73 | ± 0.005 | Printed in the official tutorial |
plsda_q2Cumulative Q2 of the PLS-DASource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.3, PLS-DA summary, Q2(cum) 0.584.Note in the list of known values: ropls vignette 1.44.0; check.R | 0.584 | ± 0.005 | Printed in the official tutorial |
plsda_rmseeRMSEE of the PLS-DASource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.3, PLS-DA summary, RMSEE 0.262.Note in the list of known values: ropls vignette 1.44.0; check.R | 0.262 | ± 0.003 | Printed in the official tutorial |
plsda_componentsNumber of PLS-DA componentsSource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.3, PLS-DA summary, pre 3.Note in the list of known values: ropls vignette 1.44.0; check.R | 3 | exact | Printed in the official tutorial |
oplsda_q2Cumulative Q2 of the OPLS-DASource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.4, OPLS-DA summary, Q2(cum) 0.602.Note in the list of known values: ropls vignette 1.44.0; check.R | 0.602 | ± 0.005 | Printed in the official tutorial |
oplsda_orthogonalOrthogonal components of the OPLS-DASource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.4, OPLS-DA summary, ort 2.Note in the list of known values: ropls vignette 1.44.0; check.R | 2 | exact | Printed in the official tutorial |
permutation_pPermutation p value (20 permutations)Source of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Sections 4.3 and 4.4, pR2Y 0.05 and pQ2 0.05.Note in the list of known values: ropls vignette 1.44.0; check.R | 0.05 | ± 0.005 | Printed in the official tutorial |
oplsda_test_r2yR2Y of the OPLS-DA on the odd rowsSource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.4, model on the odd subset, R2Y(cum) 0.825.Note in the list of known values: ropls vignette 1.44.0; check.R | 0.825 | ± 0.005 | Printed in the official tutorial |
oplsda_test_q2Q2 of the OPLS-DA on the odd rowsSource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.4, model on the odd subset, Q2(cum) 0.608.Note in the list of known values: ropls vignette 1.44.0; check.R | 0.608 | ± 0.005 | Printed in the official tutorial |
oplsda_test_rmsepRMSEP of the OPLS-DA on the test rowsSource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.4, model on the odd subset, RMSEP 0.341.Note in the list of known values: ropls vignette 1.44.0; check.R | 0.341 | ± 0.003 | Printed in the official tutorial |
oplsda_test_correctCorrect calls in the test setSource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 4.4, confusion table of the test subset, 43 and 34 correct of 91, which the vignette calls 85% correct.Note in the list of known values: ropls vignette 1.44.0 confusion table (43 + 34); check.R | 77 | exact | Printed in the official tutorial |
age_r2yR2Y of the OPLS regression of ageSource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 5.2, OPLS of age, R2Y(cum) 0.476.Note in the list of known values: ropls vignette 1.44.0; check.R | 0.476 | ± 0.005 | Printed in the official tutorial |
age_q2Q2 of the OPLS regression of ageSource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 5.2, OPLS of age, Q2(cum) 0.31.Note in the list of known values: ropls vignette 1.44.0; check.R | 0.31 | ± 0.005 | Printed in the official tutorial |
age_rmseeRMSEE of the OPLS regression of ageSource of the known valuePrinted in the official tutorialTool: the ropls vignette, version 1.44.0Where: Section 5.2, OPLS of age, RMSEE 7.53.Note in the list of known values: ropls vignette 1.44.0; check.R | 7.53 | ± 0.02 | Printed in the official tutorial |
n_significantSignificant metabolites, Welch test with BH at 0.05Source of the known valueIndependent check: we calculated itTool: SciPy (checks/check_metab.py in the adapter) and R t.test (check.R)Where: Welch test of gender with the Benjamini-Hochberg adjustment, 42 metabolites below 0.05.Check: check.R runs t.test and p.adjust in base R, and checks/check_metab.py in the adapter uses SciPy (check.out).Note in the list of known values: check.R; checks/check_metab.py (SciPy) | 42 | exact | Independent check: we calculated it |
n_upSignificant metabolites higher in menSource of the known valueIndependent check: we calculated itTool: SciPy and RWhere: 11 of the 42 are higher in men.Check: check.R runs t.test and p.adjust in base R, and checks/check_metab.py in the adapter uses SciPy (check.out).Note in the list of known values: check.R; checks/check_metab.py | 11 | exact | Independent check: we calculated it |
n_downSignificant metabolites lower in menSource of the known valueIndependent check: we calculated itTool: SciPy and RWhere: 31 of the 42 are lower in men.Check: check.R runs t.test and p.adjust in base R, and checks/check_metab.py in the adapter uses SciPy (check.out).Note in the list of known values: check.R; checks/check_metab.py | 31 | exact | Independent check: we calculated it |
Latest scored run
Model: claude-haiku-5-5. Runs for each paper and model: 1. Blind mode: on. Status: answer. 198 s. Computed: 21 of 21 values. Reported: 21 of 21 values. The result file is bench/results/papers-2026-10-09-metabolomics-haiku.md. This run is not in the totals of the page of papers.
Computed: a logged number is within the tolerance. Reported: the final answer states the value, as the claim check measures. The table copies the cells of the result file.
| Item | Expected | Computed | Reported |
|---|---|---|---|
n_samplesNumber of samples | 183 exact | pass 183 (n2 metrics.n_samples, entry 56) | pass 183 via tolerance (n7, claim check 150) |
pca_r2xCumulative R2X of the PCA | 0.501 ±0.002 | pass 0.501 (n2 metrics.r2x, entry 56) | pass 0.501 via tolerance (n2, claim check 150) |
pca_componentsNumber of PCA components | 8 exact | pass 8 (n2 metrics.n_predictive, entry 56) | pass 8 via tolerance (n2, claim check 150) |
plsda_r2xCumulative R2X of the PLS-DA | 0.275 ±0.002 | pass 0.275 (n3 metrics.r2x, entry 64) | pass 0.275 via tolerance (n4, claim check 150) |
plsda_r2yCumulative R2Y of the PLS-DA | 0.73 ±0.005 | pass 0.73 (n3 metrics.r2y, entry 64) | pass 0.73 via tolerance (n4, claim check 150) |
plsda_q2Cumulative Q2 of the PLS-DA | 0.584 ±0.005 | pass 0.584 (n3 metrics.q2, entry 64) | pass 0.584 via tolerance (n3, claim check 150) |
plsda_rmseeRMSEE of the PLS-DA | 0.262 ±0.003 | pass 0.262 (n3 metrics.rmsee, entry 64) | pass 0.262 via tolerance (n4, claim check 150) |
plsda_componentsNumber of PLS-DA components | 3 exact | pass 3 (n3 metrics.n_predictive, entry 64) | pass 3 via tolerance (n3, claim check 150) |
oplsda_q2Cumulative Q2 of the OPLS-DA | 0.602 ±0.005 | pass 0.602 (n4 metrics.q2, entry 68) | pass 0.602 via tolerance (n4, claim check 150) |
oplsda_orthogonalOrthogonal components of the OPLS-DA | 2 exact | pass 2 (n4 metrics.n_orthogonal, entry 68) | pass 2 via tolerance (n6, claim check 150) |
permutation_pPermutation p value (20 permutations) | 0.05 ±0.005 | pass 0.05 (n3 metrics.p_r2y, entry 64) | pass 0.05 via tolerance (n10, claim check 150) |
oplsda_test_r2yR2Y of the OPLS-DA on the odd rows | 0.825 ±0.005 | pass 0.825 (n6 metrics.r2y, entry 87) | pass 0.825 via tolerance (n6, claim check 150) |
oplsda_test_q2Q2 of the OPLS-DA on the odd rows | 0.608 ±0.005 | pass 0.608 (n6 metrics.q2, entry 87) | pass 0.608 via tolerance (n6, claim check 150) |
oplsda_test_rmsepRMSEP of the OPLS-DA on the test rows | 0.341 ±0.003 | pass 0.341 (n6 metrics.rmsep, entry 87) | pass 0.341 via tolerance (n6, claim check 150) |
oplsda_test_correctCorrect calls in the test set | 77 exact | pass 77 (n6 metrics.holdout_correct, entry 87) | pass 77 via tolerance (n6, claim check 150) |
age_r2yR2Y of the OPLS regression of age | 0.476 ±0.005 | pass 0.476 (n7 metrics.r2y, entry 91) | pass 0.476 via tolerance (n7, claim check 150) |
age_q2Q2 of the OPLS regression of age | 0.31 ±0.005 | pass 0.31 (n7 metrics.q2, entry 91) | pass 0.31 via tolerance (n7, claim check 150) |
age_rmseeRMSEE of the OPLS regression of age | 7.53 ±0.02 | pass 7.53 (n7 metrics.rmsee, entry 91) | pass 7.53 via tolerance (n7, claim check 150) |
n_significantSignificant metabolites, Welch test with BH at 0.05 | 42 exact | pass 42 (n10 metrics.n_significant, entry 114) | pass 42 via tolerance (n10, claim check 150) |
n_upSignificant metabolites higher in men | 11 exact | pass 11 (n10 metrics.n_up, entry 114) | pass 11 via tolerance (n10, claim check 150) |
n_downSignificant metabolites lower in men | 31 exact | pass 31 (n10 metrics.n_down, entry 114) | pass 31 via tolerance (n10, claim check 150) |
Notes
Triage notes by the maintainers
The text below is from the triage notes. We show it as the maintainers wrote it.
Classes: a = tool or adapter fault, b = harness fault, c = benchmark spec fault, d = model fault.
- claude:claude-haiku-5-5 (blind, adapter 0.1.0, final run):
20261009-041011-a533, 198 s, 12 tool calls (0 failed), computed 21/21, reported 21/21. Report:bench/results/papers-2026-10-09-metabolomics-haiku.md. - Leak check: 0 paths outside the allow list, 0 blocked reads, 0 literature claims.
- First run of the paper (before the fix in the table below):
20261009-040442-3196, 236 s, 19 tool calls (1 failed), computed 21/21, reported 21/21.
| Run | Item | Class | Cause | Fix |
|---|---|---|---|---|
| 1 | prepare_metabolite_table failed | a | The tool refused every table with a negative value. The published table holds log10 values, which are negative. The model tried the preparation step with all options set to none. | The tool now refuses negative values only if normalization or a transform is on. Two known-answer tests pin both cases. |
| 1 | referee finding on n_predictive | c | The decision record fixed n_predictive to auto. The model asked for 1 predictive component in the OPLS calls, and the harness kept auto. OPLS with one response uses 1 component in both cases, so the values did not change. | none. The finding has severity info. |
| 2 | 5 unsourced numbers | b | The numbers 18, 16 and 14 are counts of extreme values from the data check of the harness (inspect_data note). The model repeated them as examples. They are not results of the analysis. | none. |
Other findings:
- The model asked the setup questions through the decision record, asked for the scaling when it ran the first model, and passed
holdout: oddby itself for the test-set run. - The permutation p value of 0.05 is the smallest value that 20 permutations can give. The prompt does not ask for more permutations.