cuvette Install

Validation / Papers / Yavorska 2017

MendelianRandomization: an R package for performing Mendelian randomization analyses using summarized data

Human genetics · tool tutorial or software test data · MendelianRandomization (R), through the mendelianrandomization adapter

How to read this page. In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. A known value comes from the paper, from a tutorial or from a check that we ran. This page has no combined run of the paper yet.

Run of 9 October 2026, claude-haiku-5-5: 12 of 12 values computed, 12 of 12 correct in the final answer

The paper

Yavorska OO, Burgess S. MendelianRandomization: an R package for performing Mendelian randomization analyses using summarized data. International Journal of Epidemiology 46(6):1734-1739 (2017). doi:10.1093/ije/dyx034

Related sources:

What it measured

The package paper describes methods that estimate the causal effect of an exposure on an outcome from the summary statistics of genetic variants. The package vignette applies them to 28 variants with effects on LDL cholesterol and on coronary heart disease. It prints the inverse-variance weighted, MR-Egger, median, mode and maximum likelihood estimates with the heterogeneity tests.

Data

R package MendelianRandomization 0.10.0, data ldlc, ldlcse, chdlodds, chdloddsse, written to CSV by catalog/mendelianrandomization/data/make_fixtures.R Size: 28 rows, 2 kB.

License: GPL-2 or GPL-3, the license of the package. The data are published summary statistics of variants. They hold no data of persons.

Data source

The instruction

A script sends this message as the scientist.

ScientistI have the effects of 28 genetic variants on LDL cholesterol and on coronary heart disease. Does a higher LDL cholesterol raise the risk? Give the inverse-variance weighted, MR-Egger and weighted median estimates with their standard errors and intervals, Cochran's Q, I-squared, the F statistic and the MR-Egger intercept.

Basis: The package vignette (version 0.9.0), sections on the data, the inverse-variance weighted method, the median-based method and the MR-Egger method. The vignette prints the values with three or four digits.

The decisions

The model asks questions during a run. A script gives these answers to the questions of the model. We wrote the answers before the run.

Table 1 | Answers that a script gives to the questions of the model.
DecisionValueSource
Allele harmonizationtrue, the effects of the exposure and the outcome refer to the same effect alleleThe package data set holds the effects of the same allele for the exposure and the outcome of each variant.
Estimation methodsivw, egger, weighted_medianThe request names these three methods. The vignette prints them in this order.
Model of the inverse-variance weighted methoddefault (random effects for more than three variants)The vignette prints "random-effect model" for the 28 variants with the default setting.
Distribution for the confidence intervalnormalThe vignette uses the default, the normal distribution.
Significance level of the confidence interval0.05The vignette prints 95% intervals.
Variants that you excludenoneThe vignette uses all 28 variants.
Weakest accepted instrument (F statistic)10Not in the vignette. The usual limit of 10.
Other questions of the agentUse the values in the decision record.Not in the vignette. The benchmark answers a free question with this text.

Known values

The tolerance is the largest difference from the known value that we accept. We set it before the run. Exact: the number must be the same.

Table 2 | Known values for Yavorska 2017.
ValueKnown valueToleranceSource
n_variantsNumber of variants
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Section "The data" and the printed "Number of Variants : 28".Note in the list of known values: package vignette 0.9.0
28exactPrinted in the official tutorial
ivw_estimateInverse-variance weighted estimate
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, printed estimate 2.834.Note in the list of known values: package vignette 0.9.0, mr_ivw output; check.py
2.834± 0.01Printed in the official tutorial
ivw_seStandard error of the inverse-variance weighted estimate
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, printed standard error 0.530.Note in the list of known values: package vignette 0.9.0; check.py
0.53± 0.01Printed in the official tutorial
ivw_ci_lowLower limit of the 95% confidence interval of the inverse-variance weighted estimate
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, printed 95% confidence interval.Note in the list of known values: package vignette 0.9.0
1.796± 0.02Printed in the official tutorial
ivw_ci_highUpper limit of the 95% confidence interval of the inverse-variance weighted estimate
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, printed 95% confidence interval.Note in the list of known values: package vignette 0.9.0
3.873± 0.02Printed in the official tutorial
cochran_qCochran's Q
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, "Heterogeneity test statistic (Cochran's Q) = 99.5304 on 27 degrees of freedom".Note in the list of known values: package vignette 0.9.0; check.py
99.5304± 0.5Printed in the official tutorial
i_squaredI-squared, percent
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, "I^2 = 72.9%".Note in the list of known values: package vignette 0.9.0; check.py
72.9± 0.5Printed in the official tutorial
f_statistic (reference)F statistic
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Inverse-variance weighted method, "F statistic = 28.0".Note in the list of known values: package vignette 0.9.0; check.py
28.0± 0.5Printed in the official tutorial
egger_seStandard error of the MR-Egger estimate
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: MR-Egger method, printed standard error 0.770.Note in the list of known values: package vignette 0.9.0; check.py
0.77± 0.01Printed in the official tutorial
egger_estimateMR-Egger estimate
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: MR-Egger method, printed estimate 3.253.Note in the list of known values: package vignette 0.9.0, mr_egger output; check.py
3.253± 0.01Printed in the official tutorial
egger_interceptMR-Egger intercept
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: MR-Egger method, printed intercept -0.011.Note in the list of known values: package vignette 0.9.0; check.py
-0.011± 0.002Printed in the official tutorial
egger_intercept_pp value of the MR-Egger intercept
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: MR-Egger method, printed intercept p value 0.451.Note in the list of known values: package vignette 0.9.0
0.451± 0.01Printed in the official tutorial
weighted_median_estimateWeighted median estimate
Source of the known valuePrinted in the official tutorialTool: the MendelianRandomization vignette, version 0.9.0Where: Median-based method, printed weighted median estimate 2.683.Note in the list of known values: package vignette 0.9.0, mr_median output; check.py
2.683± 0.01Printed in the official tutorial

Latest scored run

Model: claude-haiku-5-5. Runs for each paper and model: 1. Blind mode: on. Status: answer. 116 s. Computed: 12 of 12 values. Reported: 12 of 12 values. The result file is bench/results/papers-2026-10-09-epigenomics-crispr-genetics-haiku.md. This run is not in the totals of the page of papers.

Computed: a logged number is within the tolerance. Reported: the final answer states the value, as the claim check measures. The table copies the cells of the result file.

Table 3 | Items of the run of claude-haiku-5-5.
ItemExpectedComputedReported
n_variantsNumber of variants28 exactpass 28 (n2 metrics.n_variants, entry 54)pass 28 via tolerance (n6, claim check 103)
ivw_estimateInverse-variance weighted estimate2.834 ±0.01pass 2.834214 (n2 metrics.ivw_estimate, entry 54)pass 2.834 via tolerance (n6, claim check 103)
ivw_seStandard error of the inverse-variance weighted estimate0.53 ±0.01pass 0.5297995 (n2 metrics.ivw_se, entry 54)pass 0.53 via tolerance (n6, claim check 103)
ivw_ci_lowLower limit of the 95% confidence interval of the inverse-variance weighted estimate1.796 ±0.02pass 1.795826 (n2 metrics.ivw_ci_low, entry 54)pass 1.796 via tolerance (n2, claim check 103)
ivw_ci_highUpper limit of the 95% confidence interval of the inverse-variance weighted estimate3.873 ±0.02pass 3.872602 (n2 metrics.ivw_ci_high, entry 54)pass 3.873 via tolerance (n2, claim check 103)
cochran_qCochran's Q99.5304 ±0.5pass 99.53043 (n2 metrics.ivw_q, entry 54)pass 99.53 via tolerance (n6, claim check 103)
i_squaredI-squared, percent72.9 ±0.5pass 72.87262 (n2 metrics.ivw_i2, entry 54)pass 72.9 via tolerance (n2, claim check 103)
f_statistic (reference)F statistic28 ±0.5match 28 (n1 metrics.n_variants, entry 41)match 28 via tolerance (n6, claim check 103)
egger_seStandard error of the MR-Egger estimate0.77 ±0.01pass 0.7701292 (n2 metrics.egger_se, entry 54)pass 0.77 via tolerance (n2, claim check 103)
egger_estimateMR-Egger estimate3.253 ±0.01pass 3.25289 (n2 metrics.egger_estimate, entry 54)pass 3.253 via tolerance (n2, claim check 103)
egger_interceptMR-Egger intercept-0.011 ±0.002pass -0.01146067 (n2 metrics.egger_intercept, entry 54)pass -0.0115 via tolerance (n2, claim check 103)
egger_intercept_pp value of the MR-Egger intercept0.451 ±0.01pass 0.4505065 (n2 metrics.egger_intercept_p, entry 54)pass 0.451 via tolerance (n2, claim check 103)
weighted_median_estimateWeighted median estimate2.683 ±0.01pass 2.682883 (n2 metrics.weighted_median_estimate, entry 54)pass 2.683 via tolerance (n2, claim check 103)

Notes

Triage notes by the maintainers

The text below is from the triage notes. We show it as the maintainers wrote it.

Entry numbers (eNN) are ids in the log.jsonl of the session folder. Classes: a = tool or adapter fault, b = harness fault, c = benchmark spec fault, d = model fault.

RunItemClassCauseFix
1inspect_data failedbThe sandbox of a worktree run blocks the Python helper at src/tools/inspect_data.py (Operation not permitted). The model read the file with read_file and went on.none. The fault is in the sandbox rule for a worktree path, not in the adapter.
earlierf_statisticcThe F statistic is 28.0 and the number of variants is 28. The scorer matched the F item to metrics.n_variants.The item became a reference item and the standard error of the MR-Egger estimate took its place.

Other findings: