Validation / Papers / Yavorska 2017
MendelianRandomization: an R package for performing Mendelian randomization analyses using summarized data
How to read this page. In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. A known value comes from the paper, from a tutorial or from a check that we ran. This page has no combined run of the paper yet.
Run of 9 October 2026, claude-haiku-5-5: 12 of 12 values computed, 12 of 12 correct in the final answer
The paper
Yavorska OO, Burgess S. MendelianRandomization: an R package for performing Mendelian randomization analyses using summarized data. International Journal of Epidemiology 46(6):1734-1739 (2017). doi:10.1093/ije/dyx034
Related sources:
- Patel A, Ye T, Xue H, Lin Z, Xu S, Woolf B, Mason AM, Burgess S. MendelianRandomization v0.9.0: updates to an R package for performing Mendelian randomization analyses using summarized data. Wellcome Open Res 8:449 (2023). Describes the package version of the vignette that prints the values. doi:10.12688/wellcomeopenres.19995.1
- Waterworth DM et al. Genetic variants influencing circulating lipid levels and risk of coronary artery disease. Arterioscler Thromb Vasc Biol 30(11):2264-2276 (2010). Source of the 28 variants. doi:10.1161/atvbaha.109.201020
What it measured
The package paper describes methods that estimate the causal effect of an exposure on an outcome from the summary statistics of genetic variants. The package vignette applies them to 28 variants with effects on LDL cholesterol and on coronary heart disease. It prints the inverse-variance weighted, MR-Egger, median, mode and maximum likelihood estimates with the heterogeneity tests.
Data
R package MendelianRandomization 0.10.0, data ldlc, ldlcse, chdlodds, chdloddsse, written to CSV by catalog/mendelianrandomization/data/make_fixtures.R Size: 28 rows, 2 kB.
License: GPL-2 or GPL-3, the license of the package. The data are published summary statistics of variants. They hold no data of persons.
The instruction
A script sends this message as the scientist.
Basis: The package vignette (version 0.9.0), sections on the data, the inverse-variance weighted method, the median-based method and the MR-Egger method. The vignette prints the values with three or four digits.
The decisions
The model asks questions during a run. A script gives these answers to the questions of the model. We wrote the answers before the run.
| Decision | Value | Source |
|---|---|---|
| Allele harmonization | true, the effects of the exposure and the outcome refer to the same effect allele | The package data set holds the effects of the same allele for the exposure and the outcome of each variant. |
| Estimation methods | ivw, egger, weighted_median | The request names these three methods. The vignette prints them in this order. |
| Model of the inverse-variance weighted method | default (random effects for more than three variants) | The vignette prints "random-effect model" for the 28 variants with the default setting. |
| Distribution for the confidence interval | normal | The vignette uses the default, the normal distribution. |
| Significance level of the confidence interval | 0.05 | The vignette prints 95% intervals. |
| Variants that you exclude | none | The vignette uses all 28 variants. |
| Weakest accepted instrument (F statistic) | 10 | Not in the vignette. The usual limit of 10. |
| Other questions of the agent | Use the values in the decision record. | Not in the vignette. The benchmark answers a free question with this text. |
Known values
The tolerance is the largest difference from the known value that we accept. We set it before the run. Exact: the number must be the same.
| Value | Known value | Tolerance | Source |
|---|---|---|---|
n_variantsNumber of variantsSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Section "The data" and the printed "Number of Variants : 28".Note in the list of known values: package vignette 0.9.0 | 28 | exact | Printed in the official tutorial |
ivw_estimateInverse-variance weighted estimateSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, printed estimate 2.834.Note in the list of known values: package vignette 0.9.0, mr_ivw output; check.py | 2.834 | ± 0.01 | Printed in the official tutorial |
ivw_seStandard error of the inverse-variance weighted estimateSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, printed standard error 0.530.Note in the list of known values: package vignette 0.9.0; check.py | 0.53 | ± 0.01 | Printed in the official tutorial |
ivw_ci_lowLower limit of the 95% confidence interval of the inverse-variance weighted estimateSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, printed 95% confidence interval.Note in the list of known values: package vignette 0.9.0 | 1.796 | ± 0.02 | Printed in the official tutorial |
ivw_ci_highUpper limit of the 95% confidence interval of the inverse-variance weighted estimateSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, printed 95% confidence interval.Note in the list of known values: package vignette 0.9.0 | 3.873 | ± 0.02 | Printed in the official tutorial |
cochran_qCochran's QSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, "Heterogeneity test statistic (Cochran's Q) = 99.5304 on 27 degrees of freedom".Note in the list of known values: package vignette 0.9.0; check.py | 99.5304 | ± 0.5 | Printed in the official tutorial |
i_squaredI-squared, percentSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, "I^2 = 72.9%".Note in the list of known values: package vignette 0.9.0; check.py | 72.9 | ± 0.5 | Printed in the official tutorial |
f_statistic (reference)F statisticSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, "F statistic = 28.0".Note in the list of known values: package vignette 0.9.0; check.py | 28.0 | ± 0.5 | Printed in the official tutorial |
egger_seStandard error of the MR-Egger estimateSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: MR-Egger method, printed standard error 0.770.Note in the list of known values: package vignette 0.9.0; check.py | 0.77 | ± 0.01 | Printed in the official tutorial |
egger_estimateMR-Egger estimateSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: MR-Egger method, printed estimate 3.253.Note in the list of known values: package vignette 0.9.0, mr_egger output; check.py | 3.253 | ± 0.01 | Printed in the official tutorial |
egger_interceptMR-Egger interceptSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: MR-Egger method, printed intercept -0.011.Note in the list of known values: package vignette 0.9.0; check.py | -0.011 | ± 0.002 | Printed in the official tutorial |
egger_intercept_pp value of the MR-Egger interceptSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: MR-Egger method, printed intercept p value 0.451.Note in the list of known values: package vignette 0.9.0 | 0.451 | ± 0.01 | Printed in the official tutorial |
weighted_median_estimateWeighted median estimateSource of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Median-based method, printed weighted median estimate 2.683.Note in the list of known values: package vignette 0.9.0, mr_median output; check.py | 2.683 | ± 0.01 | Printed in the official tutorial |
Latest scored run
Model: claude-haiku-5-5. Runs for each paper and model: 1. Blind mode: on. Status: answer. 116 s. Computed: 12 of 12 values. Reported: 12 of 12 values. The result file is bench/results/papers-2026-10-09-epigenomics-crispr-genetics-haiku.md. This run is not in the totals of the page of papers.
Computed: a logged number is within the tolerance. Reported: the final answer states the value, as the claim check measures. The table copies the cells of the result file.
| Item | Expected | Computed | Reported |
|---|---|---|---|
n_variantsNumber of variants | 28 exact | pass 28 (n2 metrics.n_variants, entry 54) | pass 28 via tolerance (n6, claim check 103) |
ivw_estimateInverse-variance weighted estimate | 2.834 ±0.01 | pass 2.834214 (n2 metrics.ivw_estimate, entry 54) | pass 2.834 via tolerance (n6, claim check 103) |
ivw_seStandard error of the inverse-variance weighted estimate | 0.53 ±0.01 | pass 0.5297995 (n2 metrics.ivw_se, entry 54) | pass 0.53 via tolerance (n6, claim check 103) |
ivw_ci_lowLower limit of the 95% confidence interval of the inverse-variance weighted estimate | 1.796 ±0.02 | pass 1.795826 (n2 metrics.ivw_ci_low, entry 54) | pass 1.796 via tolerance (n2, claim check 103) |
ivw_ci_highUpper limit of the 95% confidence interval of the inverse-variance weighted estimate | 3.873 ±0.02 | pass 3.872602 (n2 metrics.ivw_ci_high, entry 54) | pass 3.873 via tolerance (n2, claim check 103) |
cochran_qCochran's Q | 99.5304 ±0.5 | pass 99.53043 (n2 metrics.ivw_q, entry 54) | pass 99.53 via tolerance (n6, claim check 103) |
i_squaredI-squared, percent | 72.9 ±0.5 | pass 72.87262 (n2 metrics.ivw_i2, entry 54) | pass 72.9 via tolerance (n2, claim check 103) |
f_statistic (reference)F statistic | 28 ±0.5 | match 28 (n1 metrics.n_variants, entry 41) | match 28 via tolerance (n6, claim check 103) |
egger_seStandard error of the MR-Egger estimate | 0.77 ±0.01 | pass 0.7701292 (n2 metrics.egger_se, entry 54) | pass 0.77 via tolerance (n2, claim check 103) |
egger_estimateMR-Egger estimate | 3.253 ±0.01 | pass 3.25289 (n2 metrics.egger_estimate, entry 54) | pass 3.253 via tolerance (n2, claim check 103) |
egger_interceptMR-Egger intercept | -0.011 ±0.002 | pass -0.01146067 (n2 metrics.egger_intercept, entry 54) | pass -0.0115 via tolerance (n2, claim check 103) |
egger_intercept_pp value of the MR-Egger intercept | 0.451 ±0.01 | pass 0.4505065 (n2 metrics.egger_intercept_p, entry 54) | pass 0.451 via tolerance (n2, claim check 103) |
weighted_median_estimateWeighted median estimate | 2.683 ±0.01 | pass 2.682883 (n2 metrics.weighted_median_estimate, entry 54) | pass 2.683 via tolerance (n2, claim check 103) |
Notes
Triage notes by the maintainers
The text below is from the triage notes. We show it as the maintainers wrote it.
Entry numbers (eNN) are ids in the log.jsonl of the session folder. Classes: a = tool or adapter fault, b = harness fault, c = benchmark spec fault, d = model fault.
- claude-haiku-5-5 (blind, adapter 0.1.0):
20261009-024240-8c65, 116 s, 8 tool calls (1 failed), computed 13/13, reported 13/13. Report:bench/results/papers-2026-10-09-epigenomics-crispr-genetics-haiku.md. - Leak check: 0 paths outside the allow list, 0 unsourced claims. The one "blocked" entry is the wait for the scientist's answers, not a denied read.
- Two earlier runs of the same paper (adapter 0.1.0, before the item list changed) also passed all items.
| Run | Item | Class | Cause | Fix |
|---|---|---|---|---|
| 1 | inspect_data failed | b | The sandbox of a worktree run blocks the Python helper at src/tools/inspect_data.py (Operation not permitted). The model read the file with read_file and went on. | none. The fault is in the sandbox rule for a worktree path, not in the adapter. |
| earlier | f_statistic | c | The F statistic is 28.0 and the number of variants is 28. The scorer matched the F item to metrics.n_variants. | The item became a reference item and the standard error of the MR-Egger estimate took its place. |
Other findings:
- The model asked the four setup questions through the decision record and passed
alleles_harmonizedthrough the decision. The tools refused no call. - The mode standard error (1.004 here, 1.005 in the vignette) comes from a bootstrap. It is not an item.