Validation / Papers / El Masri 2026
Early-life microbiota in the LIMIT cohort: unveiling meconium microbiota types beyond maternal lifestyle
How to read this page. In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. A known value comes from the paper, from a tutorial or from a check that we ran. This page has no combined run of the paper yet.
No run is scored for this paper yet.
The paper
El Masri D et al. Early-life microbiota in the LIMIT cohort: unveiling meconium microbiota types beyond maternal lifestyle. Gut Microbes 18(1) (2026). doi:10.1080/19490976.2026.2712811
Related sources:
- Mallick H et al. Multivariable association discovery in population-scale meta-omics studies. PLOS Comput Biol 17:e1009442 (2021). Source of the MaAsLin 2 method. doi:10.1371/journal.pcbi.1009442
What it measured
The study sequenced the 16S rRNA gene (V1 to V8, long reads) in 168 meconium samples of newborns from the LIMIT cohort in Pavia. The authors tested the association of each OTU with gestational age, gestational weight gain (reference adequate), maternal age and pre-pregnancy BMI in the babies born by vaginal delivery, with MaAsLin 2. The model uses total sum scaling, removes features that occur in fewer than 5 percent of the samples, and applies the default log transform. One OTU (Klebsiella quasivariicola) is associated with gestational age.
Data
Zenodo record 10.5281/zenodo.18376726 (files otutable.csv and metadata.csv), the study database of the paper Size: Two CSV files, 267 kB and 21 kB. 538 OTUs by 168 samples, and 168 sample rows with 28 columns..
License: CC BY 4.0 (Zenodo record). Pseudonymous mother-infant data without direct identifiers.
The instruction
A script sends this message as the scientist.
Basis: The Methods and the Results section on MaAsLin 2. The paper does not name the column that holds the delivery type. The benchmark uses the column birth_type with the value natural, which gives the 132 participants of the paper.
The decisions
The model asks questions during a run. A script gives these answers to the questions of the model. We wrote the answers before the run.
| Decision | Value | Source |
|---|---|---|
| Research question | Which OTUs are associated with gestational age in babies born by vaginal delivery? | Results, MaAsLin 2 on the observed bacterial OTUs. |
| Unit of replication | subjects, one sample for each subject | One meconium sample for each mother-infant pair (168 pairs). |
| Filter that keeps a subset of the samples | birth_type=natural | Results, Figure 2 and Figure 3 cover infants born via vaginal delivery. The number 132 matches vaginal births with complete data. |
| Samples or features to exclude | none | The paper removes features in fewer than 5 percent of the samples. That is the prevalence limit. |
| Smallest share of samples that must contain a feature | 0.05 | Methods. All features detected in fewer than 5 percent of the samples were excluded. |
| Smallest share of samples that must contain a feature, for the distance | 0 | Figure 2 uses the relative abundance table. The paper states no feature filter for the distance. |
| Rarefaction depth | none | Methods. The abundances were normalized with total sum scaling. No rarefaction is stated for MaAsLin 2. |
| Alpha diversity index | shannon | Figure 2e shows Shannon, Simpson and inverse Simpson. The MaAsLin 2 result does not use it. |
| Column that defines the groups for alpha diversity | GWG | Figure 2e compares the GWG categories. |
| Normalization before the distance | relative | Figure 2a uses Bray-Curtis distance on the relative abundance table. The MaAsLin 2 result does not use it. |
| Distance measure | bray | Figure 2a, Bray-Curtis. |
| Pseudo-count for CLR | 1 | Not used in the paper. The value is only an entry for the decision record. |
| PERMANOVA terms | GWG | The PERMANOVA of GWG in Figure 2. The MaAsLin 2 result does not use it. |
| Number of permutations | 999 | Not stated in the text. The value is the default of adonis2. |
| Strata | none | Not used in the paper. |
| Random seed | 1 | Not stated. The linear model has no random step. |
| Differential abundance method | maaslin2 | Methods, MaAsLin2 v1.16.0. |
| Columns that are tested | gestational_age, GWG, age_mother, pre_pregnancy_BMI | Methods. Gestational age, GWG, maternal age and pre-BMI as fixed effects. |
| Random effects | none | Methods. The paper names fixed effects only. Each mother-infant pair has one sample. |
| Reference level | GWG, adequate | Methods. Adequate GWG as the reference category. |
| Normalization for MaAsLin 2 | TSS | Methods. The OTU abundances were normalized with total sum scaling. |
| Transform for MaAsLin 2 | LOG | Methods. The default MaAsLin 2 transformation, which is the log transform. |
| Standardize continuous columns | True | Not stated. The default of MaAsLin 2 is true. |
| Adjusted p value limit | 0.05 | Results. Significant after false discovery rate correction, q below 0.05. |
| Monte Carlo draws for ALDEx2 | 128 | Not used in the paper. ALDEx2 default. |
Known values
The tolerance is the largest difference from the known value that we accept. We set it before the run. Exact: the number must be the same.
| Value | Known value | Tolerance | Source |
|---|---|---|---|
samples_in_studyMeconium samples in the studySource of the known valuePrinted in the paper. Independent check: yesWhere: Methods. "168 meconium samples were analyzed" and "168 mother-infant pairs were selected".Check: check_limit_maaslin.py reads the table and the sample table with pandas.Note in the list of known values: El Masri 2026, Methods: 168 meconium samples | 168 | exact | Printed in the paper. Independent check: yes |
samples_in_modelSamples in the model after the filter for vaginal delivery and the removal of missing valuesSource of the known valuePrinted in the paper. Independent check: yesWhere: Results. "this OTU was present in only 7 of the 132 study participants".Check: check_limit_maaslin.py keeps the 142 vaginal births and drops the 10 with missing values, which leaves 132.Note in the list of known values: El Masri 2026, Results: 7 of the 132 study participants; check_limit_maaslin.py | 132 | exact | Printed in the paper. Independent check: yes |
significant_otusAssociations with q at or below 0.05Source of the known valuePrinted in the paper. Independent check: yesWhere: Results. "identified a single significant association after multiple-testing correction".Check: check_limit_maaslin.py counts 1 result with q at or below 0.05 among 325 tests.Note in the list of known values: El Masri 2026, Results: a single significant association; check_limit_maaslin.py | 1 | exact | Printed in the paper. Independent check: yes |
top_qAdjusted p value (q) of the most significant OTUSource of the known valuePrinted in the paper. Independent check: yesWhere: Results. "OTU RS_GCF_002269255.1 (Klebsiella quasivariicola) was negatively associated with gestational age (q = 0.043)".Check: check_limit_maaslin.py fits ordinary least squares models in statsmodels with the MaAsLin 2 steps and gives q 0.042647. The paper value 0.043 is this value rounded.Note in the list of known values: El Masri 2026, Results: q = 0.043; check_limit_maaslin.py gives 0.042647 | 0.043 | ± 0.0005 | Printed in the paper. Independent check: yes |
top_otu_samplesSamples that contain the most significant OTUSource of the known valuePrinted in the paper. Independent check: yesWhere: Results. "this OTU was present in only 7 of the 132 study participants, i.e. 5.3%".Check: check_limit_maaslin.py counts 7 samples with a read of this OTU.Note in the list of known values: El Masri 2026, Results: 7 of the 132; check_limit_maaslin.py | 7 | exact | Printed in the paper. Independent check: yes |
otus_testedOTUs tested after the prevalence limit (not printed in the paper)Source of the known valueIndependent check: we calculated itTool: MaAsLin2 1.26.0 through the microbiome adapterWhere: Not printed in the paper. The benchmark authors ran the model.Check: check_limit_maaslin.py applies the 5 percent prevalence limit and keeps 65 OTUs (325 tests).Note in the list of known values: check_limit_maaslin.py: 65 OTUs, 325 tests | 65 | exact | Independent check: we calculated it |
shannon_gwg_pp value of the Kruskal-Wallis test of Shannon diversity by GWG categorySource of the known valueIndependent check: we calculated itTool: vegan 2.7.6 and R wilcox and kruskal tests through the microbiome adapterWhere: Not printed. The paper states that alpha diversity does not differ between the GWG categories (Figure 2e, Kruskal-Wallis). The p value is in a figure only.Check: check_limit_diversity.py computes Shannon with numpy and the Kruskal-Wallis test with SciPy. It gives p 0.991662.Note in the list of known values: computed: check_limit_diversity.py; the paper says no difference (Figure 2e) | 0.9917 | ± 0.005 | Independent check: we calculated it |
permanova_gwg_r2PERMANOVA R2 of GWG (Bray-Curtis, relative abundances)Source of the known valueIndependent check: we calculated itTool: vegan 2.7.6 adonis2 through the microbiome adapterWhere: Not printed. The paper states no effect of GWG on beta diversity (Supplementary Table S1).Check: check_limit_diversity.py implements the PERMANOVA of Anderson (2001) in numpy on Bray-Curtis distances. It gives R2 0.010403.Note in the list of known values: computed: check_limit_diversity.py; the paper says no effect | 0.0104 | ± 0.0005 | Independent check: we calculated it |
permanova_gwg_fPERMANOVA F of GWGSource of the known valueIndependent check: we calculated itTool: vegan 2.7.6 adonis2 through the microbiome adapterWhere: Not printed. Same as the R2 item.Check: check_limit_diversity.py gives F 0.678065. The F value does not depend on the seed.Note in the list of known values: computed: check_limit_diversity.py | 0.678 | ± 0.005 | Independent check: we calculated it |
permanova_gwg_pPERMANOVA p of GWG, 999 permutationsSource of the known valueIndependent check: we calculated itTool: vegan 2.7.6 adonis2, 999 permutationsWhere: Not printed. The p value changes with the seed.Check: The numpy permutation test with seed 1 gives 0.910. The tolerance 0.06 covers the change with the seed.Note in the list of known values: computed: check_limit_diversity.py, numpy seed 1 gives 0.910; the p value changes with the seed | 0.91 | ± 0.06 | Independent check: we calculated it |
Latest scored run
No run is scored for this paper yet.
Notes
Triage notes by the maintainers
The text below is from the triage notes. We show it as the maintainers wrote it.
Classes: a = tool or adapter fault, b = harness fault, c = benchmark spec fault, d = model fault.
- claude:claude-haiku-5-5, run 1:
20261009-034501-3dac - claude:claude-haiku-5-5, run 2:
20261009-034911-c4a1 - claude:claude-haiku-5-5, run 3 (final):
20261009-035258-57f6
| Model | Item or call | Expected | Got | Class | Cause | Fix |
|---|---|---|---|---|---|---|
| haiku | permanova_gwg_f and permanova_gwg_r2 in run 2 | 0.678 and 0.0104 | 0.537 and 0.0103 | a | One decision (prevalence limit 0.05) was linked to both MaAsLin 2 and the distance. The tool dropped rare OTUs before the distance. | The new decision beta_min_prevalence (default 0) sets the limit for the distance. Run 3 passes. |
| haiku | samples_in_study in run 2 | 168 | 142 | d | The model reported the vaginal births as the samples in the study. The prompt gives 168. | None. Run 3 reports 168. |
| haiku | inspect_data | ok | "can't open file src/tools/inspect_data.py" | b | Blind mode hides the repository from child processes. The built-in tool could not start. | None. The model used inspect_feature_table. |
| haiku | beta_diversity with strata subject in run 3 | ok | "strata is not a column of the metadata: subject" | d | The sample table has no subject column. The error named the column. | None. The model repeated the call without strata. |
| haiku | final answer, the number 0.104875 | sourced | unsourced | d | A value that the model derived from logged results. | None. No scored item depends on it. |
| spec | otus_tested | not scored | 65 | c | The paper does not print the number of tested OTUs. | The item is a reference (no score). check_limit_maaslin.py gives 65. |
| spec | the four diversity items | not scored in the paper | scored | c | The paper prints no number for the Shannon test or for the PERMANOVA. | The values come from check_limit_diversity.py (numpy, SciPy). case.yaml marks them as computed. |
Commits in this branch for this paper: see git log -- bench/papers/elmasri2026-limit-maaslin2.