cuvette Install

Validation / Papers / Hochman 1999

Early revascularization in acute myocardial infarction complicated by cardiogenic shock

Cardiology, clinical trial statistics · research paper · survRM2, survival and base R, through the trial-endpoints adapter

How to read this page. In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. A known value comes from the paper, from a tutorial or from a check that we ran. This page has no combined run of the paper yet.

Run of 9 October 2026, claude-haiku-5-5: 17 of 17 values computed, 17 of 17 correct in the final answer

The paper

Hochman JS, Sleeper LA, Webb JG, Sanborn TA, White HD, Talley JD, Buller CE, Jacobs AK, Slater JN, Col J, McKinlay SM, LeJemtel TH. Early revascularization in acute myocardial infarction complicated by cardiogenic shock. New England Journal of Medicine 341(9):625-634 (1999). doi:10.1056/NEJM199908263410901

Related sources:

What it measured

The trial randomized 302 patients with cardiogenic shock to early revascularization (152 patients) or initial medical stabilization (150 patients). The primary endpoint was death from any cause within 30 days. The follow-up papers report the survival at 1 year and the survival over 6 to 11 years.

Data

Iturrate E. SHOCK Trial - Early Revascularization in Acute Myocardial Infarction Complicated by Cardiogenic Shock. Zenodo, files trial_deident.csv and survival_deident.csv, joined by fetch.sh Size: 302 rows, 11 kB.

License: CC BY 4.0. The data are de-identified. The table that the benchmark reads has no patient identifier.

Data source

The instruction

A script sends this message as the scientist.

ScientistCompare death within 30 days (numbers in each arm, risk difference with the Wald interval, chi-square p value), the hazard ratio of death over the whole follow-up with its interval and the log-rank p value, and the Kaplan-Meier survival at 1 year with the difference and its interval.

Basis: The abstract of the 1999 paper (30-day numbers: 46.7% and 56.0%, difference -9.3 points, interval -20.5 to 1.9, P = 0.11), the abstract of the 2001 paper (1-year survival) and the abstract of the 2006 paper (hazard ratio and log-rank P).

The decisions

The model asks questions during a run. A script gives these answers to the questions of the model. We wrote the answers before the run.

Table 1 | Answers that a script gives to the questions of the model.
DecisionValueSource
Primary endpointdeath from any cause within 30 days of randomization (primary); death during follow-up (long-term survival)The 1999 paper names 30-day all-cause mortality as the primary endpoint. The 2006 paper uses all-cause mortality during follow-up.
Analysis populationintention to treatThe papers analyze the patients in the arm of randomization.
Stratification factornoneThe abstracts print unstratified comparisons.
Patients with a missing outcomeexcludeThe Zenodo file has no follow-up time for one patient of the medical stabilization arm. The tool drops that patient from the survival analysis.
Confidence interval method for the risk differencewaldThe interval -20.5 to 1.9 of the 1999 abstract equals the Wald interval of the data (-20.5 to 1.9).
Test of the two proportionschi_squareThe chi-square test without correction gives p = 0.106. The abstract prints P = 0.11.
Two-sided significance level0.05The papers print 95% intervals.
Direction of the binary outcomeevent is badThe event is death.
Non-inferiority margin0The trial tests superiority. No margin.
Scale of the margin for a binary endpointrisk_differenceNot used. The margin is 0.
Scale of the margin for a time-to-event endpointhazard_ratioNot used. The margin is 0.
One-sided level of the non-inferiority test0.025Not used. The margin is 0.
Time limit tau of the restricted mean survival time1825 days (5 years)Not in the papers. The benchmark asks for it so that the tool with survRM2 runs. The value is a reference item.
Time of the survival comparison365 daysThe 2001 paper compares the survival at 1 year.
Handling of tied event times in the Cox modelefronThe default of R. The choice does not change the hazard ratio at the printed digits.
Other questions of the agentUse the values in the decision record.The benchmark answers a free question with this text.

Known values

The tolerance is the largest difference from the known value that we accept. We set it before the run. Exact: the number must be the same.

Table 2 | Known values for Hochman 1999.
ValueKnown valueToleranceSource
n_earlyPatients, early revascularization
Source of the known valuePrinted in the paperTool: Hochman et al. 1999, abstractWhere: Methods: 152 patients assigned to emergency revascularization.Note in the list of known values: NEJM 1999 abstract
152exactPrinted in the paper
n_medicalPatients, medical stabilization
Source of the known valuePrinted in the paperTool: Hochman et al. 1999, abstractWhere: Methods: 150 patients assigned to initial medical stabilization.Note in the list of known values: NEJM 1999 abstract
150exactPrinted in the paper
mortality30_early30-day mortality, early revascularization
Source of the known valuePrinted in the paperTool: Hochman et al. 1999, abstractWhere: Results: overall mortality at 30 days, 46.7 percent in the revascularization group.Note in the list of known values: NEJM 1999 abstract (46.7 percent); check.py 0.4671
0.467± 0.001Printed in the paper
mortality30_medical30-day mortality, medical stabilization
Source of the known valuePrinted in the paperTool: Hochman et al. 1999, abstractWhere: Results: overall mortality at 30 days, 56.0 percent in the medical-therapy group.Note in the list of known values: NEJM 1999 abstract (56.0 percent); check.py 0.5600
0.56± 0.001Printed in the paper
rd3030-day risk difference
Source of the known valuePrinted in the paperTool: Hochman et al. 1999, abstractWhere: Results: difference, -9.3 percent.Note in the list of known values: NEJM 1999 abstract (-9.3 percent); check.py -0.0929
-0.093± 0.001Printed in the paper
rd30_ci_lowLower limit of the 95% Wald interval
Source of the known valuePrinted in the paperTool: Hochman et al. 1999, abstractWhere: Results: 95 percent confidence interval for the difference, -20.5 to 1.9 percent.Note in the list of known values: NEJM 1999 abstract (-20.5); check.py -0.2051
-0.205± 0.001Printed in the paper
rd30_ci_highUpper limit of the 95% Wald interval
Source of the known valuePrinted in the paperTool: Hochman et al. 1999, abstractWhere: Results: 95 percent confidence interval for the difference, -20.5 to 1.9 percent.Note in the list of known values: NEJM 1999 abstract (1.9); check.py 0.0194
0.019± 0.001Printed in the paper
p3030-day chi-square p value
Source of the known valuePrinted in the paperTool: Hochman et al. 1999, abstractWhere: Results: P = 0.11. The chi-square test without correction on the data gives 0.106.Note in the list of known values: NEJM 1999 abstract (P = 0.11); check.py 0.1063
0.11± 0.006Printed in the paper
hrHazard ratio of death
Source of the known valuePrinted in the paperTool: Hochman et al. 2006, abstractWhere: Results: hazard ratio 0.74.Note in the list of known values: JAMA 2006 abstract; check.py 0.7437
0.74± 0.005Printed in the paper
hr_ci_lowLower limit of the hazard ratio interval
Source of the known valuePrinted in the paperTool: Hochman et al. 2006, abstractWhere: Results: 95% confidence interval 0.57-0.97.Note in the list of known values: JAMA 2006 abstract; check.py 0.5712
0.57± 0.005Printed in the paper
hr_ci_highUpper limit of the hazard ratio interval
Source of the known valuePrinted in the paperTool: Hochman et al. 2006, abstractWhere: Results: 95% confidence interval 0.57-0.97.Note in the list of known values: JAMA 2006 abstract; check.py 0.9684
0.97± 0.005Printed in the paper
logrank_pLog-rank p value
Source of the known valuePrinted in the paperTool: Hochman et al. 2006, abstractWhere: Results: log-rank P = .03.Note in the list of known values: JAMA 2006 abstract (P = .03); check.py 0.02735
0.03± 0.005Printed in the paper
surv1_early1-year survival, early revascularization
Source of the known valuePrinted in the paperTool: Hochman et al. 2001, abstractWhere: Results: one-year survival was 46.7% for patients in the ERV group.Note in the list of known values: JAMA 2001 abstract (46.7%); check.py 0.4671
0.467± 0.001Printed in the paper
surv1_medical1-year survival, medical stabilization
Source of the known valuePrinted in the paperTool: Hochman et al. 2001, abstractWhere: Results: 33.6% in the IMS group.Note in the list of known values: JAMA 2001 abstract (33.6%); check.py 0.3356
0.336± 0.001Printed in the paper
surv1_diff1-year survival difference
Source of the known valuePrinted in the paperTool: Hochman et al. 2001, abstractWhere: Results: absolute difference in survival, 13.2%.Note in the list of known values: JAMA 2001 abstract (13.2%); check.py 0.1315
0.132± 0.001Printed in the paper
surv1_diff_ci_lowLower limit of the difference interval
Source of the known valuePrinted in the paperTool: Hochman et al. 2001, abstractWhere: Results: 95% confidence interval, 2.2%-24.1%.Note in the list of known values: JAMA 2001 abstract (2.2%); check.py 0.0218
0.022± 0.001Printed in the paper
surv1_diff_ci_highUpper limit of the difference interval
Source of the known valuePrinted in the paperTool: Hochman et al. 2001, abstractWhere: Results: 95% confidence interval, 2.2%-24.1%.Note in the list of known values: JAMA 2001 abstract (24.1%); check.py 0.2413
0.241± 0.001Printed in the paper
rmst_diff (reference)Difference of the restricted mean survival time up to 1825 days (days; no published value)
Source of the known valueIndependent check: we calculated itTool: check.py (NumPy step integral of the Kaplan-Meier curve)Where: Not in the papers.Check: check.py integrates the Kaplan-Meier curve with NumPy and gives the same value as survRM2 (check.out).Note in the list of known values: check.py (NumPy step integral of the Kaplan-Meier curve)
224.26± 0.5Independent check: we calculated it

Latest scored run

Model: claude-haiku-5-5. Runs for each paper and model: 1. Blind mode: on. Status: answer. 137 s. Computed: 17 of 17 values. Reported: 17 of 17 values. The result file is bench/results/papers-2026-10-09-trials-haiku.md. This run is not in the totals of the page of papers.

Computed: a logged number is within the tolerance. Reported: the final answer states the value, as the claim check measures. The table copies the cells of the result file.

Table 3 | Items of the run of claude-haiku-5-5.
ItemExpectedComputedReported
n_earlyPatients, early revascularization152 exactpass 152 (n4 metrics.n_a, entry 55)pass 152 via tolerance (n5, claim check 128)
n_medicalPatients, medical stabilization150 exactpass 150 (n4 metrics.n_b, entry 55)pass 150 via tolerance (n4, claim check 128)
mortality30_early30-day mortality, early revascularization0.467 ±0.001pass 0.4671053 (n4 metrics.risk_a, entry 55)pass 0.467 via tolerance (n4, claim check 128)
mortality30_medical30-day mortality, medical stabilization0.56 ±0.001pass 0.56 (n4 metrics.risk_b, entry 55)pass 0.56 via tolerance (n4, claim check 128)
rd3030-day risk difference-0.093 ±0.001pass -0.09289474 (n4 metrics.rd, entry 55)pass -0.093 via tolerance (n9, claim check 128)
rd30_ci_lowLower limit of the 95% Wald interval-0.205 ±0.001pass -0.2051493 (n4 metrics.rd_ci_low, entry 55)pass -0.205 via tolerance (n9, claim check 128)
rd30_ci_highUpper limit of the 95% Wald interval0.019 ±0.001pass 0.01903706 (n7 metrics.rd_ci_high, entry 93)pass 0.019 via tolerance (n7, claim check 128)
p3030-day chi-square p value0.11 ±0.006pass 0.1063389 (n4 metrics.test_p, entry 55)pass 0.1063 via tolerance (n4, claim check 128)
hrHazard ratio of death0.74 ±0.005pass 0.7382611 (n8 metrics.hr, entry 96)pass 0.738 via tolerance (n8, claim check 128)
hr_ci_lowLower limit of the hazard ratio interval0.57 ±0.005pass 0.5712074 (n5 metrics.hr_ci_low, entry 71)pass 0.571 via tolerance (n5, claim check 128)
hr_ci_highUpper limit of the hazard ratio interval0.97 ±0.005pass 0.9699647 (n5 metrics.rmtl_ratio_ci_high, entry 71)pass 0.968 via tolerance (n8, claim check 128)
logrank_pLog-rank p value0.03 ±0.005pass 0.02790277 (n5 metrics.hr_p, entry 71)pass 0.02735 via tolerance (n5, claim check 128)
surv1_early1-year survival, early revascularization0.467 ±0.001pass 0.4671053 (n5 metrics.surv_a, entry 71)pass 0.467 via tolerance (n4, claim check 128)
surv1_medical1-year survival, medical stabilization0.336 ±0.001pass 0.3355705 (n5 metrics.surv_b, entry 71)pass 0.336 via tolerance (n5, claim check 128)
surv1_diff1-year survival difference0.132 ±0.001pass 0.1315348 (n5 metrics.surv_diff, entry 71)pass 0.132 via tolerance (n5, claim check 128)
surv1_diff_ci_lowLower limit of the difference interval0.022 ±0.001pass 0.02181158 (n5 metrics.surv_diff_ci_low, entry 71)pass 0.022 via tolerance (n5, claim check 128)
surv1_diff_ci_highUpper limit of the difference interval0.241 ±0.001pass 0.241258 (n5 metrics.surv_diff_ci_high, entry 71)pass 0.241 via tolerance (n5, claim check 128)
rmst_diff (reference)Difference of the restricted mean survival time up to 1825 days (days; no published value)224.26 ±0.5match 224.2616 (n5 metrics.rmst_diff, entry 71)not asked

Notes

Triage notes by the maintainers

The text below is from the triage notes. We show it as the maintainers wrote it.

Entry numbers (eNN) are ids in the log.jsonl of the session folder. Classes: a = tool or adapter fault, b = harness fault, c = benchmark spec fault, d = model fault.

RunItemClassCauseFix
1rd30, rd30_ci_low, surv1_diff and its limitscThe tools give fractions. The prompt asked for percent. The model wrote "-9.3 percentage points". The claim check reads a percent only after a % sign or the word "percent", so it did not match -0.093.The prompt now asks for a number with the % sign after it, and a % sign after each interval limit. Run 3 reports 17/17.
2same itemscThe prompt said "in percent, with the % sign". The model still wrote "points" and put the sign after the last limit only.Same fix as run 1.
allduplicate rowsdThe model reported two identical rows (file rows 255 and 276) as possible copies. The table has no patient identifier. Two patients with the same arm, age flag and follow-up day look the same.none. The model kept both rows and asked the scientist, as the standards say.

Items of the papers that do not reproduce with the Zenodo data (not in the benchmark; see expected.yaml):

Other findings: