cuvette Install

Validation / Papers / Student 1908

Student 1908: paired t-test on the Cushny and Peebles sleep data

Statistics · research paper · SciPy (Python), through the biostats adapter

How to read this page

In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. The run is one sample: another run can give different steps and numbers. The model is the AI. The harness is Cuvette, the software around the model: it runs the programs and records each step. A tool call is a request from the model to run one program step. The session record is the log of each message and each step. The claim check is a script that finds each number of the final answer in the step results. The review is a set of fixed rule checks plus a second AI model, the referee, that reads the record. A deviation is a request from the model for a setting that differs from the choice of the scientist. Each Claude model did 3 runs of this paper. This page shows run 3 of each Claude model and the one run of qwen3:8b. The table of values says how many of the Claude runs match.

Opus: 6 of 6 values match, 4 of 4 correct in the final answer. All 3 runs: 6 of 6 values match. Sonnet: 6 of 6 values match, 4 of 4 correct in the final answer. All 3 runs: 6 of 6 values match. Haiku: 6 of 6 values match, 4 of 4 correct in the final answer. All 3 runs: 6 of 6 values match. qwen3:8b: 6 of 6 values match, 4 of 4 correct in the final answer.

The figure in the paper and in the run

As published

Illustration I of Student (1908). The paper gives the extra sleep of the 10 patients in a table, not in a figure. The table is the source of our values.

See the figure in the paper

Fig. 1 | As published. This page does not show the published figure. The link opens the paper.

Reproduced in Cuvette

The figure reproduced from this run in Cuvette
Fig. 2 | Reproduced in Cuvette. Reproduction of the paired t-test on the sleep data of Student (1908), drawn from the data (10 patients, 2 drugs) and the values of the run (SciPy, paired t-test, two-sided, 95% confidence interval). The run values come from a run of the model Claude Opus 5.5 with the harness on 9 October 2026. (a) Extra sleep of each patient with drug 1 and drug 2. Grey lines join the two values of one patient. The red line joins the two group means. (b) The difference for each patient, drug 2 minus drug 1. The red line shows the mean difference of the run. The pale band shows the 95% confidence interval of the run. An unpaired Welch test on the same data gives p = 0.079. This is the wrong test for paired data. (c) Each known value (open ring) and run value (red dot), on a scale of the tolerance. All six values are in tolerance. The known values are computed from the data. The paper gives its own probability from its own table.

The paper

Student (William Sealy Gosset). The probable error of a mean. Biometrika 6(1):1-25 (1908). doi:10.2307/2331554

Related sources:

What it measured

Ten patients each took two forms of the drug hyoscyamine on different nights. The table gives the hours of sleep that each drug added, for each patient. Student used these data to show his method for a small sample. He took the difference between the two drugs for each patient and asked if the mean difference is greater than zero. The paired difference removes the variation between patients, so the test is much more sensitive than a comparison of the two groups as independent samples.

Data

R dataset sleep, from the Rdatasets collection. Size: 238 bytes, 20 rows.

License: The 1905 measurements are historic published data. The Rdatasets collection is GPL-3. The file has no patient identifiers.

Data source

The instruction

A script sent this message as the scientist. The file paths point to the fetched data.

ScientistTen patients each tried two sleeping drugs. The file {data}/student1908-sleep/sleep.csv has the extra hours of sleep (columns extra, group, ID). Is one drug better? Give me the mean difference, a confidence interval and a p-value. Write every number in your final answer text.

The same request in the words of the paper's method:

Ten patients each took both sleeping drugs. For each patient, compare the extra sleep on drug 2 with the extra sleep on drug 1. Is drug 2 better? Give me the mean difference, a 95% confidence interval and a two-sided p-value.

Basis: Illustration I of the paper. Student subtracts the result of drug 1 from the result of drug 2 for each patient and tests the mean of these differences.

Results

Match: a number in the session record is inside the tolerance of the known value. In the final answer: the model also stated the value in its final answer. For a Claude model, each cell shows the run that this page shows. If the three runs differ, the cell also says in how many runs the value matches.

Table 1 | Known values and the value of each model.
ValueKnown valueToleranceOpusSonnetHaikuqwen3:8b
mean_differenceMean paired difference, group 2 minus group 1
Source of the known valuePrinted in the paperIllustration I, the table of the ten patients. The mean of the differences is +1.58 hours.
1.58± 0.0051.58 matchIn the final answer: yes (1.58)Log: n2 compare_two_groups metrics.mean_diff, entry 40; the final answer, entry 661.58 matchIn the final answer: yes (1.58)Log: n2 compare_two_groups metrics.mean_diff, entry 32; the final answer, entry 391.58 matchIn the final answer: yes (1.58)Log: n2 compare_two_groups metrics.mean_diff, entry 42; the final answer, entry 561.58 matchIn the final answer: yes (1.58)Log: n2 compare_two_groups metrics.mean_diff, entry 29; the final answer, entry 36
paired_tPaired t statistic, absolute value
Source of the known valueWe calculated it with SciPy 1.x ttest_relNot in this form in the paper. Student gives the mean as 1.35 times the standard deviation, with n as the divisor. The modern t statistic uses n - 1.
abs 4.0621± 0.0014.062128 matchNot asked in the questionLog: n2 compare_two_groups metrics.statistic, entry 404.062128 matchNot asked in the questionLog: n2 compare_two_groups metrics.statistic, entry 324.062128 matchNot asked in the questionLog: n2 compare_two_groups metrics.statistic, entry 424.062128 matchNot asked in the questionLog: n2 compare_two_groups metrics.statistic, entry 29
degrees_of_freedomDegrees of freedom
Source of the known valueWe calculated it with SciPy 1.x ttest_relNot in the paper. Student indexes his table by the number of experiments (10).
9exact9 matchNot asked in the questionLog: n2 compare_two_groups metrics.df, entry 409 matchNot asked in the questionLog: n2 compare_two_groups metrics.df, entry 329 matchNot asked in the questionLog: n2 compare_two_groups metrics.df, entry 429 matchNot asked in the questionLog: n2 compare_two_groups metrics.df, entry 29
paired_pPaired two-sided p
Source of the known valueWe calculated it with SciPy 1.x ttest_relNot in the paper. Student gives a probability of 0.9985 (odds of about 666 to 1) that drug 2 is better, from his own table.
0.002833± 0.00010.00283289 matchIn the final answer: yes (0.0028)Log: n2 compare_two_groups metrics.p_value, entry 40; the final answer, entry 660.00283289 matchIn the final answer: yes (0.002833)Log: n2 compare_two_groups metrics.p_value, entry 32; the final answer, entry 390.00283289 matchIn the final answer: yes (0.0028)Log: n2 compare_two_groups metrics.p_value, entry 42; the final answer, entry 560.00283289 matchIn the final answer: yes (0.0028)Log: n2 compare_two_groups metrics.p_value, entry 29; the final answer, entry 36
ci_lower95% CI lower bound
Source of the known valueWe calculated it with SciPy 1.xNot in the paper. Student does not give a confidence interval.
0.7± 0.0050.7 matchIn the final answer: yes (0.7)Log: n1 inspect_table file:{work}/inspect_table-1/columns.csv, entry 13; the final answer, entry 660.7001142 matchIn the final answer: yes (0.7)Log: n2 compare_two_groups metrics.ci_lo, entry 32; the final answer, entry 390.7 matchIn the final answer: yes (0.7)Log: n1 inspect_table file:{work}/inspect_table-1/columns.csv, entry 11; the final answer, entry 560.7001142 matchIn the final answer: yes (0.7)Log: n2 compare_two_groups metrics.ci_lo, entry 29; the final answer, entry 36
ci_upper95% CI upper bound
Source of the known valueWe calculated it with SciPy 1.xNot in the paper. Student does not give a confidence interval.
2.46± 0.0052.459886 matchIn the final answer: yes (2.46)Log: n2 compare_two_groups metrics.ci_hi, entry 40; the final answer, entry 662.459886 matchIn the final answer: yes (2.46)Log: n2 compare_two_groups metrics.ci_hi, entry 32; the final answer, entry 392.459886 matchIn the final answer: yes (2.46)Log: n2 compare_two_groups metrics.ci_hi, entry 42; the final answer, entry 562.459886 matchIn the final answer: yes (2.46)Log: n2 compare_two_groups metrics.ci_hi, entry 29; the final answer, entry 36

Session records

Session record, Opus, run 3 of 3

Every message, decision, step and result of this run, one JSON object for each log entry.

A compressed log file (JSON Lines, gzip). The record shows file paths as {data}, {work} and similar names. · 17 KB

Download

Session record, Sonnet, run 3 of 3

Every message, decision, step and result of this run, one JSON object for each log entry.

A compressed log file (JSON Lines, gzip). The record shows file paths as {data}, {work} and similar names. · 7 KB

Download

Session record, Haiku, run 3 of 3

Every message, decision, step and result of this run, one JSON object for each log entry.

A compressed log file (JSON Lines, gzip). The record shows file paths as {data}, {work} and similar names. · 17 KB

Download

Session record, qwen3:8b

Every message, decision, step and result of this run, one JSON object for each log entry.

A compressed log file (JSON Lines, gzip). The record shows file paths as {data}, {work} and similar names. · 6 KB

Download

Opus · claude-opus-5-5 · run 3 of 3 shown 6 of 6 values match, 4 of 4 correct in the final answer

The session

This is the session as the scientist sees it, in order. Decision cards show the answer that the script gave and where that answer comes from. Each step shows the program, its version, the input file hash and the outputs. Click a line to see more.

setup The decision record starts with these values. A script set them before the first message, from the answers that we wrote before the run.

  • Unit of replication: subjects, patients or animals (one value per subject)Where the answer comes from: Illustration I. Each of the ten patients gives one result for each drug.
  • Significance level: 0.05Where the answer comes from: Not in the paper. Student gives odds, not a fixed level. We use the usual 0.05.
harness note A note from the harness to the model. The scientist did not write it.
- Setup values:
Design:
- What is the unit of replication? (replicate): subjects, patients or animals (one value per subject)
Inference:
- Significance level (alpha) (alpha): 0.05
Ask the scientist: Paired test (paired), Test family for two groups (two_group_test), Assume equal variances (unpaired t-test) (equal_variance), Sidedness (alternative), ANOVA variance assumption (anova_variance), Pairwise comparisons after ANOVA (posthoc), Correction for a list of p-values (p_adjust_method), Model formula (model_formula), Standard errors for the linear model (standard_errors), Random effects of the mixed model (random_effects), Fit by REML (reml), Covariates of the Cox model (cox_covariates), Ties method of the Cox model (ties), Strata of the Cox model (cox_strata), Type of sums of squares (ss_type), Comparisons of the cell means after a factorial ANOVA (factorial_posthoc), Sphericity correction for repeated measures (sphericity_correction), Correction after a rank test over many groups (rank_p_adjust), Exact or approximate p value for a rank test (exact_p), Covariates of the logistic model (logistic_covariates), Outcome value that counts as the event (logistic_event), Reference level of each categorical covariate (reference_levels), Yates continuity correction for a 2x2 chi-square test (chi_square_correction), Sidedness of the Fisher exact test (fisher_alternative), Method of the McNemar test (mcnemar_method).
The harness fills these values into your tool calls automatically. If you give a different value, the harness logs a deviation and shows it to the scientist.
- Mode is now watch: you run the steps.

Model thinking, summary: I'm going to start by loading the skill and taking a look at the table.

Model

I start with the skill for two-group tests and a look at the table.

The model reads the method notes compare-two-groups.
The model calls inspect_table (adapter biostats).

step n1 inspect_table adapter biostats 0.2.0, statsmodels, pingouin, lifelines 0.30.3

20 rows x 4 columns.

Input file: {data}/student1908-sleep/sleep.csv SHA-256 5563ff4da6cf.

Outputs: columns.csv (0b88fe23ba8b).

Arguments
path{data}/student1908-sleep/sleep.csv
Tool output
{
 "ok": true,
 "summary": "20 rows x 4 columns.",
 "metrics": {
  "n_rows": 20,
  "n_columns": 4
 },
 "table": {
  "columns": [
   "column",
   "dtype",
   "n_missing",
   "n_unique",
   "min",
   "max",
   "first_values"
  ],
  "rows": [
   [
    "rownames",
    "int64",
    0,
    20,
    1,
    20,
    "1, 2, 3, 4, 5"
   ],
   [
    "extra",
    "float64",
    0,
    17,
    -1.6,
    5.5,
    "0.7, -1.6, -0.2, -1.2, -0.1"
   ],
   [
    "group",
    "int64",
    0,
    2,
    1,
    2,
    "1, 2"
   ],
   [
    "ID",
    "int64",
    0,
    10,
    1,
    10,
    "1, 2, 3, 4, 5"
   ]
  ],
  "n_rows": 4,
  "path": "{work}/inspect_table-1/columns.csv"
 }
}

Model thinking, summary: The table isn't shown directly, so I need to check columns.csv instead.

The model calls read_file.
Result of read_file: Read 226 of 226 bytes from {work}/inspect_table-1/columns.csv.
column,dtype,n_missing,n_unique,min,max,first_values
rownames,int64,0,20,1.0,20.0,"1, 2, 3, 4, 5"
extra,float64,0,17,-1.6,5.5,"0.7, -1.6, -0.2, -1.2, -0.1"
group,int64,0,2,1.0,2.0,"1, 2"
ID,int64,0,10,1.0,10.0,"1, 2, 3, 4, 5"

Model thinking, summary: With 20 rows across 10 IDs in 2 groups, each ID appears in both groups, so this is a paired design. I'll use compare_two_groups with the id_column specified.

Model thinking, summary: The table shows 10 patients each with values for both drug groups, confirming a paired design with no missing data. I'll now run the two-group test including patient ID so the harness can apply the correct pairing, test family, and sidedness.

The model calls compare_two_groups (adapter biostats).

paused The harness paused compare_two_groups until the scientist chose: Paired or unpaired test, Test family for two groups, Equal variances for an unpaired test, Sidedness. The decision cards follow.

decision card Paired test

Use a paired test if the same subject is measured in both groups (before and after, or both drugs in one patient). Use an unpaired test if the groups hold different subjects. The two choices can give different conclusions. The model wants to run compare_two_groups.

Options: yes no

Answer true

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: Illustration I. Student makes a new series of differences, drug 2 minus drug 1, for each patient. He notes that the same patient reacts in a similar way to both drugs.

decision card Test family for two groups

t is the t-test. It compares means and gives a confidence interval. wilcoxon is the rank test (signed-rank if paired, Mann-Whitney if unpaired). Use it if the data are far from normal or have outliers. The model wants to run compare_two_groups.

Options: t wilcoxon

Suggested: t (This is the adapter default.)

Answer t

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: Illustration I. Student uses his own distribution for the mean of a small sample. A rank test is not in the paper.

decision card Assume equal variances (unpaired t-test)

False gives the Welch test. It does not assume equal variances. True gives the Student test. Student's test is correct only if the group variances are equal. The model wants to run compare_two_groups.

Options: yes no

Suggested: false (This is the adapter default.)

Answer false

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: Not used. A paired test has one series of differences. The value only applies to an unpaired test.

decision card Sidedness

Two-sided tests for a difference in either direction. One-sided tests for one direction only (the second group greater than the first, or less). Choose a one-sided test only if the other direction is not of interest, and decide before you see the data. The model wants to run compare_two_groups.

Options: two-sided greater less

Suggested: two-sided (This is the adapter default.)

Answer two-sided

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: The modern convention. Student gives the probability that drug 2 is better, which is one direction.

step n2 compare_two_groups adapter biostats 0.2.0, statsmodels, pingouin, lifelines 0.30.3

paired t-test, two-sided: 2 minus 1 = 1.58, statistic 4.062, p = 0.002833 (alpha 0.05).

Decisions applied: Significance level = 0.05; Paired or unpaired test = true; Test family for two groups = t; Equal variances for an unpaired test = false; Sidedness = two-sided.

Input file: {data}/student1908-sleep/sleep.csv SHA-256 5563ff4da6cf.

Outputs: two_group_test.csv (5c80e8fb3427).

Arguments
path{data}/student1908-sleep/sleep.csv
outcomeextra
groupgroup
id_columnID
pairedtrue
testt
equal_variancefalse
alternativetwo-sided
alpha0.05
Tool output
{
 "ok": true,
 "summary": "paired t-test, two-sided: 2 minus 1 = 1.58, statistic 4.062, p = 0.002833 (alpha 0.05).",
 "metrics": {
  "n_a": 10,
  "n_b": 10,
  "mean_diff": 1.58,
  "statistic": 4.062127683382037,
  "df": 9,
  "p_value": 0.0028328901973842724,
  "ci_lo": 0.7001142367230175,
  "ci_hi": 2.459885763276983,
  "effect_size": 1.2845575625910546,
  "mean_a": 0.75,
  "mean_b": 2.3299999999999996
 },
 "data": {
  "test": "paired t-test",
  "alternative": "two-sided",
  "direction": "2 minus 1",
  "alpha": 0.05,
  "significant": true
 },
 "table": {
  "columns": [
   "test",
   "group_a",
   "group_b",
   "n_a",
   "n_b",
   "mean_diff_b_minus_a",
   "statistic",
   "df",
   "p",
   "ci_lo",
   "ci_hi",
   "cohen_dz"
  ],
  "rows": [
   [
    "paired t-test",
    1,
    2,
    10,
    10,
    1.58,
    4.062127683382037,
    9,
    0.0028328901973842724,
    0.7001142367230175,
    2.459885763276983,
    1.2845575625910546
   ]
  ],
  "n_rows": 1,
  "path": "{work}/compare_two_groups-1/two_group_test.csv"
 }
}

Model

The harness applied a paired, two-sided t-test at alpha 0.05. As a check without the normal assumption, I run the signed-rank test as a comparison run. This does not change the record.

The model calls compare_options.

comparison run n3 compare_two_groups adapter biostats 0.2.0, statsmodels, pingouin, lifelines 0.30.3

Wilcoxon signed-rank test, two-sided: 2 minus 1 = 1.58, statistic 0, p = 0.003906 (alpha 0.05).

Input file: {data}/student1908-sleep/sleep.csv SHA-256 5563ff4da6cf.

Outputs: two_group_test.csv (6650397b8dea).

Arguments
path{data}/student1908-sleep/sleep.csv
outcomeextra
groupgroup
id_columnID
pairedtrue
testwilcoxon
equal_variancefalse
alternativetwo-sided
alpha0.05
Tool output
{
 "ok": true,
 "summary": "Wilcoxon signed-rank test, two-sided: 2 minus 1 = 1.58, statistic 0, p = 0.003906 (alpha 0.05).",
 "metrics": {
  "n_a": 10,
  "n_b": 10,
  "mean_diff": 1.58,
  "statistic": 0,
  "p_value": 0.00390625,
  "mean_a": 0.75,
  "mean_b": 2.3299999999999996
 },
 "data": {
  "test": "Wilcoxon signed-rank test",
  "alternative": "two-sided",
  "direction": "2 minus 1",
  "alpha": 0.05,
  "significant": true
 },
 "table": {
  "columns": [
   "test",
   "group_a",
   "group_b",
   "n_a",
   "n_b",
   "mean_diff_b_minus_a",
   "statistic",
   "df",
   "p",
   "ci_lo",
   "ci_hi",
   "effect"
  ],
  "rows": [
   [
    "Wilcoxon signed-rank test",
    1,
    2,
    10,
    10,
    1.58,
    0,
    null,
    0.00390625,
    null,
    null,
    null
   ]
  ],
  "n_rows": 1,
  "path": "{work}/compare_two_groups-2/two_group_test.csv"
 }
}
comparison Comparison runs for Test family for two groups. The record keeps the scientist's choice.
Test family for two groups  n_a  n_b  mean_diff  statistic  p_value   mean_a  mean_b  Result
wilcoxon                    10   10   1.58       0          0.003906  0.75    2.33    ok
Result of compare_options: Compared 1 values of Test family for two groups as comparison runs of the last step, compare_two_groups (path={data}/student1908-sleep/sleep.csv, outcome=extra, ...
Compared 1 values of Test family for two groups as comparison runs of the last step, compare_two_groups (path={data}/student1908-sleep/sleep.csv, outcome=extra, group=group, id_column=ID, paired=true, equal_variance=false, alternative=two-sided, alpha=0.05). Only that step was repeated. To compare another step, run it again first. The record still says "t".
Test family for two groups  n_a  n_b  mean_diff  statistic  p_value   mean_a  mean_b  Result
wilcoxon                    10   10   1.58       0          0.003906  0.75    2.33    ok

Model thinking, summary: The Wilcoxon test comes out with p=0.0039, which confirms significance. That settles it.

Model Answer draft 1. The claim check and the review read it.

Yes. Drug 2 gave more extra sleep than drug 1. Drug 2 minus drug 1 is 1.58 hours on average, 95% CI 0.70 to 2.46 hours, p = 0.0028. This is significant at alpha 0.05.

What I did

  1. I checked the table. It has 20 rows from 10 patients, 2 drug groups and no missing values. Each patient has one value for each drug, so the 10 patients are 10 pairs.
  2. I compared the extra hours of sleep between the two drugs with a paired t-test. The harness set paired, the t-test, two-sided and alpha 0.05.
  3. As a check, I ran a paired Wilcoxon signed-rank test. This was a comparison run only. The record still says t-test.

What I found One replicate is one patient. There are 10 patients and each took both drugs.

ItemValue
Mean extra sleep, drug 10.75 hours
Mean extra sleep, drug 22.33 hours
Mean difference (drug 2 minus drug 1)1.58 hours
95% confidence interval of the difference0.70 to 2.46 hours
t statistic4.06, with 9 degrees of freedom
p-value (two-sided)0.0028
Effect size, Cohen's d_z1.28
Check: paired Wilcoxon signed-rank teststatistic 0, p = 0.0039

Cohen's d_z is the mean difference divided by the standard deviation of the 10 differences. The rank test also falls below alpha 0.05. It tests the center of the differences, not the mean, so it answers a slightly different question.

What I infer On average, patients slept more on drug 2 than on drug 1. The whole confidence interval is above zero.

What is uncertain

  • There are only 10 patients, so the interval is wide. The true average gain could be about 0.7 hours or about 2.5 hours.
  • The t-test assumes that the differences follow about a normal distribution. 10 values cannot test that well. The rank test, which does not need that assumption, agrees with it.
  • The data do not show the order in which each patient took the drugs. I cannot check for an order effect or a carry-over effect.
  • The data show the extra hours of sleep only. "Better" here means more extra sleep, not fewer side effects or other outcomes.

What waits for you Nothing waits. The record holds your decisions: paired, t-test, two-sided, alpha 0.05.

claim check 36 numbers in the answer, 0 numbers without a source in the record.
The review found 1 problem. The model corrects the answer before the scientist sees it.

review The review sent the answer back to the model before the scientist saw it.

- The answer uses 0.0039 from a comparison run of another option (two_group_test), not from the setting that was chosen. Use the result of the chosen setting, or say clearly that the number is from the comparison.
Fix these problems. Then give the answer again.

Model thinking, summary: I should label the 0.0039 value more clearly as the comparison run rather than just "Check."

Model The model writes its answer. The claim check reads it. See the final answer.

claim check 36 numbers in the answer, 0 numbers without a source in the record.
claim check 36 numbers in the answer, 0 numbers without a source in the record.
Saved: results.xlsx (9 sheets). Folder: {work}

The final answer

The model wrote this answer. The harness shows it to the scientist after the review.

Yes. Drug 2 gave more extra sleep than drug 1. Drug 2 minus drug 1 is 1.58 hours on average, 95% CI 0.70 to 2.46 hours, p = 0.0028. This is significant at alpha 0.05.

What I did

  1. I checked the table. It has 20 rows from 10 patients, 2 drug groups and no missing values. Each patient has one value for each drug, so the 10 patients are 10 pairs.
  2. I compared the extra hours of sleep between the two drugs with a paired t-test. The harness set paired, the t-test, two-sided and alpha 0.05.
  3. As a check, I ran a paired Wilcoxon signed-rank test as a separate comparison run. It did not change your choice, which is still the t-test.

What I found (paired t-test, your chosen setting) One replicate is one patient. There are 10 patients and each took both drugs.

ItemValue
Mean extra sleep, drug 10.75 hours
Mean extra sleep, drug 22.33 hours
Mean difference (drug 2 minus drug 1)1.58 hours
95% confidence interval of the difference0.70 to 2.46 hours
t statistic4.06, with 9 degrees of freedom
p-value (two-sided)0.0028
Effect size, Cohen's d_z1.28

Cohen's d_z is the mean difference divided by the standard deviation of the 10 differences.

Check only (not the main result) The paired Wilcoxon signed-rank test gave statistic 0 and p = 0.0039. This number comes from the comparison run, not from your chosen test. It also falls below alpha 0.05, so it agrees with the t-test. It tests the center of the differences, not the mean, so it answers a slightly different question.

What I infer On average, patients slept more on drug 2 than on drug 1. The whole confidence interval is above zero.

What is uncertain

  • There are only 10 patients, so the interval is wide. The true average gain could be about 0.7 hours or about 2.5 hours.
  • The t-test assumes that the differences follow about a normal distribution. 10 values cannot test that well. The rank-test check, which does not need that assumption, agrees with it.
  • The data do not show the order in which each patient took the drugs. I cannot check for an order effect or a carry-over effect.
  • The data show the extra hours of sleep only. "Better" here means more extra sleep, not fewer side effects or other outcomes.

What waits for you Nothing waits. The record holds your decisions: paired, t-test, two-sided, alpha 0.05.

Settings and step results that the answer does not mention. The harness adds them.
Settings used, from the decision record: Significance level (alpha): 0.05 · Paired test: true · Test family for two groups: t · Assume equal variances (unpaired t-test): false · Sidedness: two-sided.

Values that are not scored

  • A trap value is the result of a wrong method, for example an unpaired test on paired data. A trap value in a main step means that the run used the wrong method. A trap value in a comparison run is correct, because the record keeps the scientist's choice.
  • An optional value is a second result of the same method. The run can compute it or not.
  • A reference value comes from a different program version or a check run by us. We show it for comparison.
Table 2 | Values that are not scored, Opus run.
ItemKindKnown valueClosest logged valueToleranceOutcomeSource of the known value
welch_p_trapUnpaired Welch p (trap result)trap0.079390.05n2 compare_two_groups± 0.0005not in the recordWe calculated it with SciPy 1.x ttest_ind with equal_var=False

Checks

Review findings

The review recorded 3 findings. A rule finding comes from a fixed check in the harness. A referee finding comes from a second model that reads the record. The harness shows the findings to the scientist with the final answer. The record does not mark a finding as fixed. Thus a finding from an early review round can apply to a draft that the model corrected later.

Table 3 | Review findings, Opus run.
SeverityFromFindingShown with the final answer
errorrulenumber_from_comparisonThe answer uses 0.0039 from a comparison run of another option (two_group_test), not from the setting that was chosen. Use the result of the chosen setting, or say clearly that the number is from the comparison.yes
inforeferee modelThe answer says that the harness set paired, the t-test, two-sided and alpha 0.05. The log shows that the scientist chose these values, so the answer must name the scientist as the source. Later lines of the answer do say "your chosen setting".yes
inforeferee modelThe sentence "The true average gain could be about 0.7 hours or about 2.5 hours" reads as two possible values. It must say that the 95% CI goes from 0.70 to 2.46 hours.yes

Numbers in the answer

The last claim check read 36 numbers in the answer. 36 numbers match a logged result. 0 numbers have no source in the record.

Deviations

The model did not try to change a choice of the scientist.

Failed tool calls

No tool call failed.

Data integrity

Each data file has the same SHA-256 hash now as at the time of the step that read it. Where the download script (fetch.sh) gives a hash, the file also has that hash. The run did not change the data.

Table 4 | Data files and their SHA-256 hashes, Opus run.
FileSHA-256Fetched dataSteps with this hash
{data}/student1908-sleep/sleep.csv238 bytes5563ff4da6cfsame as the hash in the download script (fetch.sh)n1, n2, n3

A SHA-256 hash is a fingerprint of the file contents. If one byte of the file changes, the hash changes. The table shows the first 12 characters.

How to repeat it

Get the data. The script downloads the files and checks their SHA-256 hashes where it lists them.

CUVETTE_DATA={data} bash bench/papers/student1908-sleep/fetch.sh

Run the same case with Cuvette. The script gives the same answers from bench/papers/student1908-sleep/bench.yaml.

cuvette bench papers --papers student1908-sleep --models claude:claude-opus-5-5

Repeat each step by hand in the program. For each step, the harness records a manual route: the menu path or the code that gives the same result. This list does not include comparison runs.

  1. inspect_table (step n1)

    Code

    print(pd.read_csv(path).describe(include='all'))
    • path

      {data}/student1908-sleep/sleep.csv
    • Note: The tool also counts distinct values and missing values for each column.

    The manual route that the harness recorded

    ga_biostats.inspect_table(path="{data}/student1908-sleep/sleep.csv")

    The manual route uses the same method. The note in the route gives the known difference.

  2. compare_two_groups (step n2)

    Code

    wide = df.pivot(index="ID", columns="group", values="extra")
    scipy.stats.ttest_1samp(wide[2] - wide[1], 0)   # same as ttest_rel(wide[2], wide[1])
    • values column = extra
    • pairing column = ID
    • alternative = two-sided

    The manual route that the harness recorded

    ga_biostats.compare_two_groups(path="{data}/student1908-sleep/sleep.csv", outcome="extra", group="group", test="t", paired=True, id_column="ID", alternative="two-sided", alpha=0.05, equal_variance=False)

    The manual route gives the same numbers. An automatic test in Cuvette checks this.

Figure

Paper-style figure for Student 1908, from the Opus run
Fig. 3 | Opus run. Our figure script draws the values of this run in the style of the paper.

Run facts

Table 5 | Run facts, Opus run.
Modelclaude-opus-5-5 through the Anthropic service
Date2026-10-09 12:44:51 UTC
End of runthe model gave a final answer
Time61 s
Requests to the model6
Tokensunits of text that the model read and wrote16 input, 3003 output, 89658 cache read, 20475 cache write
Cost estimate$0.18 at list price, from the token counts
Tool calls5 (0 failed)
Adaptersbiostats 0.2.0, program 0.30.3
Session20261009-074450-4dd4
Code hash of each step (3)
Table 6 | Code hash of each step, Opus run.
StepToolProgram versionCode hash
n1inspect_table0.30.3603f546a1fe4
n2compare_two_groups0.30.3c6e869486df4
n3 comparisoncompare_two_groups0.30.3c6e869486df4

The code hash is a fingerprint of the adapter name, the adapter version, the tool and its definition in the adapter. If one of these changes, the hash changes.

Sonnet · claude-sonnet-5-5 · run 3 of 3 shown 6 of 6 values match, 4 of 4 correct in the final answer

The session

This is the session as the scientist sees it, in order. Decision cards show the answer that the script gave and where that answer comes from. Each step shows the program, its version, the input file hash and the outputs. Click a line to see more.

setup The decision record starts with these values. A script set them before the first message, from the answers that we wrote before the run.

  • Unit of replication: subjects, patients or animals (one value per subject)Where the answer comes from: Illustration I. Each of the ten patients gives one result for each drug.
  • Significance level: 0.05Where the answer comes from: Not in the paper. Student gives odds, not a fixed level. We use the usual 0.05.
harness note A note from the harness to the model. The scientist did not write it.
- Setup values:
Design:
- What is the unit of replication? (replicate): subjects, patients or animals (one value per subject)
Inference:
- Significance level (alpha) (alpha): 0.05
Ask the scientist: Paired test (paired), Test family for two groups (two_group_test), Assume equal variances (unpaired t-test) (equal_variance), Sidedness (alternative), ANOVA variance assumption (anova_variance), Pairwise comparisons after ANOVA (posthoc), Correction for a list of p-values (p_adjust_method), Model formula (model_formula), Standard errors for the linear model (standard_errors), Random effects of the mixed model (random_effects), Fit by REML (reml), Covariates of the Cox model (cox_covariates), Ties method of the Cox model (ties), Strata of the Cox model (cox_strata), Type of sums of squares (ss_type), Comparisons of the cell means after a factorial ANOVA (factorial_posthoc), Sphericity correction for repeated measures (sphericity_correction), Correction after a rank test over many groups (rank_p_adjust), Exact or approximate p value for a rank test (exact_p), Covariates of the logistic model (logistic_covariates), Outcome value that counts as the event (logistic_event), Reference level of each categorical covariate (reference_levels), Yates continuity correction for a 2x2 chi-square test (chi_square_correction), Sidedness of the Fisher exact test (fisher_alternative), Method of the McNemar test (mcnemar_method).
The harness fills these values into your tool calls automatically. If you give a different value, the harness logs a deviation and shows it to the scientist.
- Mode is now watch: you run the steps.
The model calls inspect_table (adapter biostats).

step n1 inspect_table adapter biostats 0.2.0, statsmodels, pingouin, lifelines 0.30.3

20 rows x 4 columns.

Input file: {data}/student1908-sleep/sleep.csv SHA-256 5563ff4da6cf.

Outputs: columns.csv (0b88fe23ba8b).

Arguments
path{data}/student1908-sleep/sleep.csv
Tool output
{
 "ok": true,
 "summary": "20 rows x 4 columns.",
 "metrics": {
  "n_rows": 20,
  "n_columns": 4
 },
 "table": {
  "columns": [
   "column",
   "dtype",
   "n_missing",
   "n_unique",
   "min",
   "max",
   "first_values"
  ],
  "rows": [
   [
    "rownames",
    "int64",
    0,
    20,
    1,
    20,
    "1, 2, 3, 4, 5"
   ],
   [
    "extra",
    "float64",
    0,
    17,
    -1.6,
    5.5,
    "0.7, -1.6, -0.2, -1.2, -0.1"
   ],
   [
    "group",
    "int64",
    0,
    2,
    1,
    2,
    "1, 2"
   ],
   [
    "ID",
    "int64",
    0,
    10,
    1,
    10,
    "1, 2, 3, 4, 5"
   ]
  ],
  "n_rows": 4,
  "path": "{work}/inspect_table-1/columns.csv"
 }
}
The model reads the method notes compare-two-groups.

Model

The design is paired (same 10 patients took both drugs). I will run the test; the harness fills in the decisions.

The model calls compare_two_groups (adapter biostats).

paused The harness paused compare_two_groups until the scientist chose: Paired or unpaired test, Test family for two groups, Equal variances for an unpaired test, Sidedness. The decision cards follow.

decision card Paired test

Use a paired test if the same subject is measured in both groups (before and after, or both drugs in one patient). Use an unpaired test if the groups hold different subjects. The two choices can give different conclusions. The model wants to run compare_two_groups.

Options: yes no

Suggested: true (The model proposed this value when it asked to run the step.)

Answer true

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: Illustration I. Student makes a new series of differences, drug 2 minus drug 1, for each patient. He notes that the same patient reacts in a similar way to both drugs.

decision card Test family for two groups

t is the t-test. It compares means and gives a confidence interval. wilcoxon is the rank test (signed-rank if paired, Mann-Whitney if unpaired). Use it if the data are far from normal or have outliers. The model wants to run compare_two_groups.

Options: t wilcoxon

Suggested: t (This is the adapter default.)

Answer t

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: Illustration I. Student uses his own distribution for the mean of a small sample. A rank test is not in the paper.

decision card Assume equal variances (unpaired t-test)

False gives the Welch test. It does not assume equal variances. True gives the Student test. Student's test is correct only if the group variances are equal. The model wants to run compare_two_groups.

Options: yes no

Suggested: false (This is the adapter default.)

Answer false

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: Not used. A paired test has one series of differences. The value only applies to an unpaired test.

decision card Sidedness

Two-sided tests for a difference in either direction. One-sided tests for one direction only (the second group greater than the first, or less). Choose a one-sided test only if the other direction is not of interest, and decide before you see the data. The model wants to run compare_two_groups.

Options: two-sided greater less

Suggested: two-sided (This is the adapter default.)

Answer two-sided

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: The modern convention. Student gives the probability that drug 2 is better, which is one direction.

step n2 compare_two_groups adapter biostats 0.2.0, statsmodels, pingouin, lifelines 0.30.3

paired t-test, two-sided: 2 minus 1 = 1.58, statistic 4.062, p = 0.002833 (alpha 0.05).

Decisions applied: Significance level = 0.05; Paired or unpaired test = true; Test family for two groups = t; Equal variances for an unpaired test = false; Sidedness = two-sided.

Input file: {data}/student1908-sleep/sleep.csv SHA-256 5563ff4da6cf.

Outputs: two_group_test.csv (5c80e8fb3427).

Arguments
path{data}/student1908-sleep/sleep.csv
outcomeextra
groupgroup
id_columnID
pairedtrue
testt
equal_variancefalse
alternativetwo-sided
alpha0.05
Tool output
{
 "ok": true,
 "summary": "paired t-test, two-sided: 2 minus 1 = 1.58, statistic 4.062, p = 0.002833 (alpha 0.05).",
 "metrics": {
  "n_a": 10,
  "n_b": 10,
  "mean_diff": 1.58,
  "statistic": 4.062127683382037,
  "df": 9,
  "p_value": 0.0028328901973842724,
  "ci_lo": 0.7001142367230175,
  "ci_hi": 2.459885763276983,
  "effect_size": 1.2845575625910546,
  "mean_a": 0.75,
  "mean_b": 2.3299999999999996
 },
 "data": {
  "test": "paired t-test",
  "alternative": "two-sided",
  "direction": "2 minus 1",
  "alpha": 0.05,
  "significant": true
 },
 "table": {
  "columns": [
   "test",
   "group_a",
   "group_b",
   "n_a",
   "n_b",
   "mean_diff_b_minus_a",
   "statistic",
   "df",
   "p",
   "ci_lo",
   "ci_hi",
   "cohen_dz"
  ],
  "rows": [
   [
    "paired t-test",
    1,
    2,
    10,
    10,
    1.58,
    4.062127683382037,
    9,
    0.0028328901973842724,
    0.7001142367230175,
    2.459885763276983,
    1.2845575625910546
   ]
  ],
  "n_rows": 1,
  "path": "{work}/compare_two_groups-1/two_group_test.csv"
 }
}

Model The model writes its answer. The claim check reads it. See the final answer.

claim check 21 numbers in the answer, 0 numbers without a source in the record.
claim check 21 numbers in the answer, 0 numbers without a source in the record.
Saved: results.xlsx (9 sheets). Folder: {work}

The final answer

The model wrote this answer. The harness shows it to the scientist after the review.

Drug 2 gave more extra sleep than drug 1 in this test.

What I did

  1. I read the file. It has 20 rows: 10 patients, each with one value for drug 1 and one for drug 2. Nothing is missing.
  2. Each patient took both drugs, so I used a paired test. The unit of replication is the patient, so n = 10 pairs.
  3. I ran a paired t-test (two-sided, alpha 0.05) with compare_two_groups.

What I found (paired t-test, step 3)

  • Mean extra sleep was 0.75 hours on drug 1 and 2.33 hours on drug 2.
  • The mean difference (drug 2 minus drug 1) is 1.58 hours.
  • The 95% confidence interval runs from 0.70 to 2.46 hours.
  • The test statistic is t = 4.062 with 9 degrees of freedom, and p = 0.002833.
  • The effect size is Cohen's dz = 1.28.
  • The p-value is below alpha 0.05, so the difference is significant at that level.

What is uncertain

  • The sample is small, with 10 patients, so the confidence interval is wide.
  • The t-test assumes the paired differences are roughly normal. I did not run a rank test (Wilcoxon) as a check.
  • I ran one test for this question, so I applied no multiple-test correction.
  • The data cannot show whether the order of the drugs or other factors affected the result.

What waits for the scientist

  • Nothing is blocked. I can run the Wilcoxon signed-rank test as a check if you want it.
Settings and step results that the answer does not mention. The harness adds them.
Settings used, from the decision record: Significance level (alpha): 0.05 · Paired test: true · Test family for two groups: t · Assume equal variances (unpaired t-test): false · Sidedness: two-sided.

Values that are not scored

  • A trap value is the result of a wrong method, for example an unpaired test on paired data. A trap value in a main step means that the run used the wrong method. A trap value in a comparison run is correct, because the record keeps the scientist's choice.
  • An optional value is a second result of the same method. The run can compute it or not.
  • A reference value comes from a different program version or a check run by us. We show it for comparison.
Table 7 | Values that are not scored, Sonnet run.
ItemKindKnown valueClosest logged valueToleranceOutcomeSource of the known value
welch_p_trapUnpaired Welch p (trap result)trap0.079390.05n2 compare_two_groups± 0.0005not in the recordWe calculated it with SciPy 1.x ttest_ind with equal_var=False

Checks

Review findings

The review recorded 1 finding. A rule finding comes from a fixed check in the harness. A referee finding comes from a second model that reads the record. The harness shows the findings to the scientist with the final answer. The record does not mark a finding as fixed. Thus a finding from an early review round can apply to a draft that the model corrected later.

Table 8 | Review findings, Sonnet run.
SeverityFromFindingShown with the final answer
inforuletext_styleThe answer breaks the text rules (ASD-STE100) in 1 place. Sentence 23 uses the passive voice: "is blocked". Use the active voice.yes

Numbers in the answer

The last claim check read 21 numbers in the answer. 21 numbers match a logged result. 0 numbers have no source in the record.

Deviations

The model did not try to change a choice of the scientist.

Failed tool calls

No tool call failed.

Data integrity

Each data file has the same SHA-256 hash now as at the time of the step that read it. Where the download script (fetch.sh) gives a hash, the file also has that hash. The run did not change the data.

Table 9 | Data files and their SHA-256 hashes, Sonnet run.
FileSHA-256Fetched dataSteps with this hash
{data}/student1908-sleep/sleep.csv238 bytes5563ff4da6cfsame as the hash in the download script (fetch.sh)n1, n2

A SHA-256 hash is a fingerprint of the file contents. If one byte of the file changes, the hash changes. The table shows the first 12 characters.

How to repeat it

Get the data. The script downloads the files and checks their SHA-256 hashes where it lists them.

CUVETTE_DATA={data} bash bench/papers/student1908-sleep/fetch.sh

Run the same case with Cuvette. The script gives the same answers from bench/papers/student1908-sleep/bench.yaml.

cuvette bench papers --papers student1908-sleep --models claude:claude-sonnet-5-5

Repeat each step by hand in the program. For each step, the harness records a manual route: the menu path or the code that gives the same result. This list does not include comparison runs.

  1. inspect_table (step n1)

    Code

    print(pd.read_csv(path).describe(include='all'))
    • path

      {data}/student1908-sleep/sleep.csv
    • Note: The tool also counts distinct values and missing values for each column.

    The manual route that the harness recorded

    ga_biostats.inspect_table(path="{data}/student1908-sleep/sleep.csv")

    The manual route uses the same method. The note in the route gives the known difference.

  2. compare_two_groups (step n2)

    Code

    wide = df.pivot(index="ID", columns="group", values="extra")
    scipy.stats.ttest_1samp(wide[2] - wide[1], 0)   # same as ttest_rel(wide[2], wide[1])
    • values column = extra
    • pairing column = ID
    • alternative = two-sided

    The manual route that the harness recorded

    ga_biostats.compare_two_groups(path="{data}/student1908-sleep/sleep.csv", outcome="extra", group="group", test="t", paired=True, id_column="ID", alternative="two-sided", alpha=0.05, equal_variance=False)

    The manual route gives the same numbers. An automatic test in Cuvette checks this.

Figure

Paper-style figure for Student 1908, from the Sonnet run
Fig. 4 | Sonnet run. Our figure script draws the values of this run in the style of the paper.

Run facts

Table 10 | Run facts, Sonnet run.
Modelclaude-sonnet-5-5 through the Anthropic service
Date2026-10-09 11:11:24 UTC
End of runthe model gave a final answer
Time15 s
Requests to the model3
Tokensunits of text that the model read and wrote8 input, 868 output, 33504 cache read, 18294 cache write
Cost estimate$0.06 at list price, from the token counts
Tool calls3 (0 failed)
Adaptersbiostats 0.2.0, program 0.30.3
Session20261009-061123-cc04
Code hash of each step (2)
Table 11 | Code hash of each step, Sonnet run.
StepToolProgram versionCode hash
n1inspect_table0.30.3603f546a1fe4
n2compare_two_groups0.30.3c6e869486df4

The code hash is a fingerprint of the adapter name, the adapter version, the tool and its definition in the adapter. If one of these changes, the hash changes.

Haiku · claude-haiku-5-5 · run 3 of 3 shown 6 of 6 values match, 4 of 4 correct in the final answer

The session

This is the session as the scientist sees it, in order. Decision cards show the answer that the script gave and where that answer comes from. Each step shows the program, its version, the input file hash and the outputs. Click a line to see more.

setup The decision record starts with these values. A script set them before the first message, from the answers that we wrote before the run.

  • Unit of replication: subjects, patients or animals (one value per subject)Where the answer comes from: Illustration I. Each of the ten patients gives one result for each drug.
  • Significance level: 0.05Where the answer comes from: Not in the paper. Student gives odds, not a fixed level. We use the usual 0.05.
harness note A note from the harness to the model. The scientist did not write it.
- Setup values:
Design:
- What is the unit of replication? (replicate): subjects, patients or animals (one value per subject)
Inference:
- Significance level (alpha) (alpha): 0.05
Ask the scientist: Paired test (paired), Test family for two groups (two_group_test), Assume equal variances (unpaired t-test) (equal_variance), Sidedness (alternative), ANOVA variance assumption (anova_variance), Pairwise comparisons after ANOVA (posthoc), Correction for a list of p-values (p_adjust_method), Model formula (model_formula), Standard errors for the linear model (standard_errors), Random effects of the mixed model (random_effects), Fit by REML (reml), Covariates of the Cox model (cox_covariates), Ties method of the Cox model (ties), Strata of the Cox model (cox_strata), Type of sums of squares (ss_type), Comparisons of the cell means after a factorial ANOVA (factorial_posthoc), Sphericity correction for repeated measures (sphericity_correction), Correction after a rank test over many groups (rank_p_adjust), Exact or approximate p value for a rank test (exact_p), Covariates of the logistic model (logistic_covariates), Outcome value that counts as the event (logistic_event), Reference level of each categorical covariate (reference_levels), Yates continuity correction for a 2x2 chi-square test (chi_square_correction), Sidedness of the Fisher exact test (fisher_alternative), Method of the McNemar test (mcnemar_method).
The harness fills these values into your tool calls automatically. If you give a different value, the harness logs a deviation and shows it to the scientist.
- Mode is now watch: you run the steps.

Model thinking, summary: Let me inspect the table first. This looks like paired sleep data for 10 patients across two drugs, so I'll want to run a paired comparison between the two groups using the ID column to match them up.

Model

I will start by checking the columns and missing values in the sleep file.

The model calls inspect_table (adapter biostats).

step n1 inspect_table adapter biostats 0.2.0, statsmodels, pingouin, lifelines 0.30.3

20 rows x 4 columns.

Input file: {data}/student1908-sleep/sleep.csv SHA-256 5563ff4da6cf.

Outputs: columns.csv (0b88fe23ba8b).

Arguments
path{data}/student1908-sleep/sleep.csv
Tool output
{
 "ok": true,
 "summary": "20 rows x 4 columns.",
 "metrics": {
  "n_rows": 20,
  "n_columns": 4
 },
 "table": {
  "columns": [
   "column",
   "dtype",
   "n_missing",
   "n_unique",
   "min",
   "max",
   "first_values"
  ],
  "rows": [
   [
    "rownames",
    "int64",
    0,
    20,
    1,
    20,
    "1, 2, 3, 4, 5"
   ],
   [
    "extra",
    "float64",
    0,
    17,
    -1.6,
    5.5,
    "0.7, -1.6, -0.2, -1.2, -0.1"
   ],
   [
    "group",
    "int64",
    0,
    2,
    1,
    2,
    "1, 2"
   ],
   [
    "ID",
    "int64",
    0,
    10,
    1,
    10,
    "1, 2, 3, 4, 5"
   ]
  ],
  "n_rows": 4,
  "path": "{work}/inspect_table-1/columns.csv"
 }
}

Model thinking, summary: There's a conflict in the data: the harness check reports 4 rows and 7 columns, but I expected 20 rows by 4 columns. I need to check the columns file directly to resolve this discrepancy.

Model

The inspection result and the harness data check disagree on the table size. I will read the column summary to check.

The model calls read_file.
Result of read_file: Read 226 of 226 bytes from {work}/inspect_table-1/columns.csv.
column,dtype,n_missing,n_unique,min,max,first_values
rownames,int64,0,20,1.0,20.0,"1, 2, 3, 4, 5"
extra,float64,0,17,-1.6,5.5,"0.7, -1.6, -0.2, -1.2, -0.1"
group,int64,0,2,1.0,2.0,"1, 2"
ID,int64,0,10,1.0,10.0,"1, 2, 3, 4, 5"
The model calls read_file.
Result of read_file: Read 238 of 238 bytes from {data}/student1908-sleep/sleep.csv.
rownames,extra,group,ID
1,0.7,1,1
2,-1.6,1,2
3,-0.2,1,3
4,-1.2,1,4
5,-0.1,1,5
6,3.4,1,6
7,3.7,1,7
8,0.8,1,8
9,0,1,9
10,2,1,10
11,1.9,2,1
12,0.8,2,2
13,1.1,2,3
14,0.1,2,4
15,-0.1,2,5
16,4.4,2,6
17,5.5,2,7
18,1.6,2,8
19,4.6,2,9
20,3.4,2,10

Model thinking, summary: The data has 20 rows with two drug groups and repeated IDs, so this calls for a paired test matched on ID — I'll run a paired t-test comparing group 2 minus group 1, using a two-sided alpha of 0.05.

Model

The file holds 20 rows and 4 columns: extra sleep, drug group (1 or 2), patient ID and row name. Each of the 10 patients appears once in each group, so I will run a paired t-test on the patient ID.

The model calls compare_two_groups (adapter biostats).

paused The harness paused compare_two_groups until the scientist chose: Paired or unpaired test, Test family for two groups, Equal variances for an unpaired test, Sidedness. The decision cards follow.

decision card Paired test

Use a paired test if the same subject is measured in both groups (before and after, or both drugs in one patient). Use an unpaired test if the groups hold different subjects. The two choices can give different conclusions. The model wants to run compare_two_groups.

Options: yes no

Suggested: true (The model proposed this value when it asked to run the step.)

Answer true

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: Illustration I. Student makes a new series of differences, drug 2 minus drug 1, for each patient. He notes that the same patient reacts in a similar way to both drugs.

decision card Test family for two groups

t is the t-test. It compares means and gives a confidence interval. wilcoxon is the rank test (signed-rank if paired, Mann-Whitney if unpaired). Use it if the data are far from normal or have outliers. The model wants to run compare_two_groups.

Options: t wilcoxon

Suggested: t (The model proposed this value when it asked to run the step.)

Answer t

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: Illustration I. Student uses his own distribution for the mean of a small sample. A rank test is not in the paper.

decision card Assume equal variances (unpaired t-test)

False gives the Welch test. It does not assume equal variances. True gives the Student test. Student's test is correct only if the group variances are equal. The model wants to run compare_two_groups.

Options: yes no

Suggested: false (This is the adapter default.)

Answer false

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: Not used. A paired test has one series of differences. The value only applies to an unpaired test.

decision card Sidedness

Two-sided tests for a difference in either direction. One-sided tests for one direction only (the second group greater than the first, or less). Choose a one-sided test only if the other direction is not of interest, and decide before you see the data. The model wants to run compare_two_groups.

Options: two-sided greater less

Suggested: two-sided (This is the adapter default.)

Answer two-sided

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: The modern convention. Student gives the probability that drug 2 is better, which is one direction.

step n2 compare_two_groups adapter biostats 0.2.0, statsmodels, pingouin, lifelines 0.30.3

paired t-test, two-sided: 2 minus 1 = 1.58, statistic 4.062, p = 0.002833 (alpha 0.05).

Decisions applied: Significance level = 0.05; Paired or unpaired test = true; Test family for two groups = t; Equal variances for an unpaired test = false; Sidedness = two-sided.

Input file: {data}/student1908-sleep/sleep.csv SHA-256 5563ff4da6cf.

Outputs: two_group_test.csv (5c80e8fb3427).

Arguments
path{data}/student1908-sleep/sleep.csv
outcomeextra
groupgroup
pairedtrue
id_columnID
testt
alpha0.05
equal_variancefalse
alternativetwo-sided
Tool output
{
 "ok": true,
 "summary": "paired t-test, two-sided: 2 minus 1 = 1.58, statistic 4.062, p = 0.002833 (alpha 0.05).",
 "metrics": {
  "n_a": 10,
  "n_b": 10,
  "mean_diff": 1.58,
  "statistic": 4.062127683382037,
  "df": 9,
  "p_value": 0.0028328901973842724,
  "ci_lo": 0.7001142367230175,
  "ci_hi": 2.459885763276983,
  "effect_size": 1.2845575625910546,
  "mean_a": 0.75,
  "mean_b": 2.3299999999999996
 },
 "data": {
  "test": "paired t-test",
  "alternative": "two-sided",
  "direction": "2 minus 1",
  "alpha": 0.05,
  "significant": true
 },
 "table": {
  "columns": [
   "test",
   "group_a",
   "group_b",
   "n_a",
   "n_b",
   "mean_diff_b_minus_a",
   "statistic",
   "df",
   "p",
   "ci_lo",
   "ci_hi",
   "cohen_dz"
  ],
  "rows": [
   [
    "paired t-test",
    1,
    2,
    10,
    10,
    1.58,
    4.062127683382037,
    9,
    0.0028328901973842724,
    0.7001142367230175,
    2.459885763276983,
    1.2845575625910546
   ]
  ],
  "n_rows": 1,
  "path": "{work}/compare_two_groups-1/two_group_test.csv"
 }
}
Model Answer draft 1. The claim check and the review read it.

Result: drug 2 gives more extra sleep than drug 1 in the same patients (paired t-test, p = 0.0028).

What I did

  1. I checked the file. It holds 20 rows and 4 columns: extra sleep, drug group (1 or 2), patient ID and a row name. No values are missing. Each of the 10 patients (ID 1 to 10) appears once with each drug.
  2. I ran a paired t-test on the extra sleep. The test pairs the two rows of each patient by ID. The test is two-sided, with alpha = 0.05. The tool gives the difference as drug 2 minus drug 1.

Results (n = 10 patients, 20 rows)

  • Mean extra sleep, drug 1: 0.75 hours.
  • Mean extra sleep, drug 2: 2.33 hours.
  • Mean difference (drug 2 minus drug 1): 1.58 hours.
  • 95% confidence interval for the mean difference: 0.70 to 2.46 hours.
  • t = 4.062, df = 9, p = 0.0028. This is significant at alpha 0.05.
  • Effect size (Cohen's dz, paired): 1.28.

What this means

  • The data show a larger average gain in sleep with drug 2 than with drug 1.
  • The paired test uses the patient as the unit of replication, so n = 10 patients.
  • The test measures only extra sleep hours. It does not show side effects or other reasons to choose one drug.

Uncertain

  • The sample is small (10 patients). The interval is wide.
  • The t-test assumes the differences are roughly normal. I did not run a rank test to check this.
  • The harness data check reported 4 rows and 7 columns. This does not match the file. My reading of the file shows 20 rows and 4 columns, and the analysis used those rows.

Waiting for you

  • Nothing waits. If you want a rank test instead, or a one-sided test, tell me before I run it. The sidedness and test family are your decisions.
claim check 34 numbers in the answer, 0 numbers without a source in the record.
The review found 1 problem. The model corrects the answer before the scientist sees it.

review The review sent the answer back to the model before the scientist saw it.

- The answer says the harness data check reported 4 rows and 7 columns. The log shows inspect_table reported 20 rows and 4 columns. The 4 rows and 7 columns figure has no source in the log and must be removed.
Fix these problems. Then give the answer again.

Model The model writes its answer. The claim check reads it. See the final answer.

claim check 31 numbers in the answer, 0 numbers without a source in the record.
claim check 31 numbers in the answer, 0 numbers without a source in the record.
Saved: results.xlsx (9 sheets). Folder: {work}

The final answer

The model wrote this answer. The harness shows it to the scientist after the review.

Result: drug 2 gives more extra sleep than drug 1 in the same patients (paired t-test, p = 0.0028).

What I did

  1. I checked the file with the column summary and the raw rows. The file holds 20 rows and 4 columns: extra sleep, drug group (1 or 2), patient ID and a row name. No values are missing. Each of the 10 patients (ID 1 to 10) appears once with each drug.
  2. I ran a paired t-test on the extra sleep. The test pairs the two rows of each patient by ID. The test is two-sided, with alpha = 0.05. The tool gives the difference as drug 2 minus drug 1.

Results (n = 10 patients, 20 rows)

  • Mean extra sleep, drug 1: 0.75 hours.
  • Mean extra sleep, drug 2: 2.33 hours.
  • Mean difference (drug 2 minus drug 1): 1.58 hours.
  • 95% confidence interval for the mean difference: 0.70 to 2.46 hours.
  • t = 4.062, df = 9, p = 0.0028. This is significant at alpha 0.05.
  • Effect size (Cohen's dz, paired): 1.28.

What this means

  • The data show a larger average gain in sleep with drug 2 than with drug 1.
  • The paired test uses the patient as the unit of replication, so n = 10 patients.
  • The test measures only extra sleep hours. It does not show side effects or other reasons to choose one drug.

Uncertain

  • The sample is small (10 patients). The interval is wide.
  • The t-test assumes the differences are roughly normal. I did not run a rank test to check this.

Waiting for you

  • Nothing waits. If you want a rank test instead, or a one-sided test, tell me before I run it. The sidedness and test family are your decisions.
Settings and step results that the answer does not mention. The harness adds them.
Settings used, from the decision record: Significance level (alpha): 0.05 · Paired test: true · Test family for two groups: t · Assume equal variances (unpaired t-test): false · Sidedness: two-sided.

Values that are not scored

  • A trap value is the result of a wrong method, for example an unpaired test on paired data. A trap value in a main step means that the run used the wrong method. A trap value in a comparison run is correct, because the record keeps the scientist's choice.
  • An optional value is a second result of the same method. The run can compute it or not.
  • A reference value comes from a different program version or a check run by us. We show it for comparison.
Table 12 | Values that are not scored, Haiku run.
ItemKindKnown valueClosest logged valueToleranceOutcomeSource of the known value
welch_p_trapUnpaired Welch p (trap result)trap0.079390.05n2 compare_two_groups± 0.0005not in the recordWe calculated it with SciPy 1.x ttest_ind with equal_var=False

Checks

Review findings

The review recorded 1 finding. A rule finding comes from a fixed check in the harness. A referee finding comes from a second model that reads the record. The harness shows the findings to the scientist with the final answer. The record does not mark a finding as fixed. Thus a finding from an early review round can apply to a draft that the model corrected later.

Table 13 | Review findings, Haiku run.
SeverityFromFindingShown with the final answer
inforeferee modelThe answer says each patient appears once with each drug. The inspection log shows only 10 unique IDs in 20 rows. The paired test used n = 10 per group, which supports this, but no step printed the ID-by-group layout directly.yes

Numbers in the answer

The last claim check read 31 numbers in the answer. 31 numbers match a logged result. 0 numbers have no source in the record.

Deviations

The model did not try to change a choice of the scientist.

Failed tool calls

No tool call failed.

Data integrity

Each data file has the same SHA-256 hash now as at the time of the step that read it. Where the download script (fetch.sh) gives a hash, the file also has that hash. The run did not change the data.

Table 14 | Data files and their SHA-256 hashes, Haiku run.
FileSHA-256Fetched dataSteps with this hash
{data}/student1908-sleep/sleep.csv238 bytes5563ff4da6cfsame as the hash in the download script (fetch.sh)n1, n2

A SHA-256 hash is a fingerprint of the file contents. If one byte of the file changes, the hash changes. The table shows the first 12 characters.

How to repeat it

Get the data. The script downloads the files and checks their SHA-256 hashes where it lists them.

CUVETTE_DATA={data} bash bench/papers/student1908-sleep/fetch.sh

Run the same case with Cuvette. The script gives the same answers from bench/papers/student1908-sleep/bench.yaml.

cuvette bench papers --papers student1908-sleep --models claude:claude-haiku-5-5

Repeat each step by hand in the program. For each step, the harness records a manual route: the menu path or the code that gives the same result. This list does not include comparison runs.

  1. inspect_table (step n1)

    Code

    print(pd.read_csv(path).describe(include='all'))
    • path

      {data}/student1908-sleep/sleep.csv
    • Note: The tool also counts distinct values and missing values for each column.

    The manual route that the harness recorded

    ga_biostats.inspect_table(path="{data}/student1908-sleep/sleep.csv")

    The manual route uses the same method. The note in the route gives the known difference.

  2. compare_two_groups (step n2)

    Code

    wide = df.pivot(index="ID", columns="group", values="extra")
    scipy.stats.ttest_1samp(wide[2] - wide[1], 0)   # same as ttest_rel(wide[2], wide[1])
    • values column = extra
    • pairing column = ID
    • alternative = two-sided

    The manual route that the harness recorded

    ga_biostats.compare_two_groups(path="{data}/student1908-sleep/sleep.csv", outcome="extra", group="group", test="t", paired=True, id_column="ID", alternative="two-sided", alpha=0.05, equal_variance=False)

    The manual route gives the same numbers. An automatic test in Cuvette checks this.

Figure

Paper-style figure for Student 1908, from the Haiku run
Fig. 5 | Haiku run. Our figure script draws the values of this run in the style of the paper.

Run facts

Table 15 | Run facts, Haiku run.
Modelclaude-haiku-5-5 through the Anthropic service
Date2026-10-09 12:54:22 UTC
End of runthe model gave a final answer
Time45 s
Requests to the model5
Tokensunits of text that the model read and wrote14 input, 3591 output, 69589 cache read, 20176 cache write
Cost estimate$0.01 at list price, from the token counts
Tool calls4 (0 failed)
Adaptersbiostats 0.2.0, program 0.30.3
Session20261009-075422-209d
Code hash of each step (2)
Table 16 | Code hash of each step, Haiku run.
StepToolProgram versionCode hash
n1inspect_table0.30.3603f546a1fe4
n2compare_two_groups0.30.3c6e869486df4

The code hash is a fingerprint of the adapter name, the adapter version, the tool and its definition in the adapter. If one of these changes, the hash changes.

qwen3:8b · local model through Ollama 6 of 6 values match, 4 of 4 correct in the final answer

The session

This is the session as the scientist sees it, in order. Decision cards show the answer that the script gave and where that answer comes from. Each step shows the program, its version, the input file hash and the outputs. Click a line to see more.

setup The decision record starts with these values. A script set them before the first message, from the answers that we wrote before the run.

  • Unit of replication: subjects, patients or animals (one value per subject)Where the answer comes from: Illustration I. Each of the ten patients gives one result for each drug.
  • Significance level: 0.05Where the answer comes from: Not in the paper. Student gives odds, not a fixed level. We use the usual 0.05.
harness note A note from the harness to the model. The scientist did not write it.
- Setup values:
Design:
- What is the unit of replication? (replicate): subjects, patients or animals (one value per subject)
Inference:
- Significance level (alpha) (alpha): 0.05
Ask the scientist: Paired test (paired), Test family for two groups (two_group_test), Assume equal variances (unpaired t-test) (equal_variance), Sidedness (alternative), ANOVA variance assumption (anova_variance), Pairwise comparisons after ANOVA (posthoc), Correction for a list of p-values (p_adjust_method), Model formula (model_formula), Standard errors for the linear model (standard_errors), Random effects of the mixed model (random_effects), Fit by REML (reml), Covariates of the Cox model (cox_covariates), Ties method of the Cox model (ties), Strata of the Cox model (cox_strata), Type of sums of squares (ss_type), Comparisons of the cell means after a factorial ANOVA (factorial_posthoc), Sphericity correction for repeated measures (sphericity_correction), Correction after a rank test over many groups (rank_p_adjust), Exact or approximate p value for a rank test (exact_p), Covariates of the logistic model (logistic_covariates), Outcome value that counts as the event (logistic_event), Reference level of each categorical covariate (reference_levels), Yates continuity correction for a 2x2 chi-square test (chi_square_correction), Sidedness of the Fisher exact test (fisher_alternative), Method of the McNemar test (mcnemar_method).
The harness fills these values into your tool calls automatically. If you give a different value, the harness logs a deviation and shows it to the scientist.
- Mode is now watch: you run the steps.
The model calls inspect_table (adapter biostats).

step n1 inspect_table adapter biostats 0.2.0, statsmodels, pingouin, lifelines 0.30.3

20 rows x 4 columns.

Input file: {data}/student1908-sleep/sleep.csv SHA-256 5563ff4da6cf.

Outputs: columns.csv (0b88fe23ba8b).

Arguments
path{data}/student1908-sleep/sleep.csv
Tool output
{
 "ok": true,
 "summary": "20 rows x 4 columns.",
 "metrics": {
  "n_rows": 20,
  "n_columns": 4
 },
 "table": {
  "columns": [
   "column",
   "dtype",
   "n_missing",
   "n_unique",
   "min",
   "max",
   "first_values"
  ],
  "rows": [
   [
    "rownames",
    "int64",
    0,
    20,
    1,
    20,
    "1, 2, 3, 4, 5"
   ],
   [
    "extra",
    "float64",
    0,
    17,
    -1.6,
    5.5,
    "0.7, -1.6, -0.2, -1.2, -0.1"
   ],
   [
    "group",
    "int64",
    0,
    2,
    1,
    2,
    "1, 2"
   ],
   [
    "ID",
    "int64",
    0,
    10,
    1,
    10,
    "1, 2, 3, 4, 5"
   ]
  ],
  "n_rows": 4,
  "path": "{work}/inspect_table-1/columns.csv"
 }
}
The model calls compare_two_groups (adapter biostats).

paused The harness paused compare_two_groups until the scientist chose: Paired or unpaired test, Test family for two groups, Equal variances for an unpaired test, Sidedness. The decision cards follow.

decision card Paired test

Use a paired test if the same subject is measured in both groups (before and after, or both drugs in one patient). Use an unpaired test if the groups hold different subjects. The two choices can give different conclusions. The model wants to run compare_two_groups.

Options: yes no

Suggested: true (The model proposed this value when it asked to run the step.)

Answer true

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: Illustration I. Student makes a new series of differences, drug 2 minus drug 1, for each patient. He notes that the same patient reacts in a similar way to both drugs.

decision card Test family for two groups

t is the t-test. It compares means and gives a confidence interval. wilcoxon is the rank test (signed-rank if paired, Mann-Whitney if unpaired). Use it if the data are far from normal or have outliers. The model wants to run compare_two_groups.

Options: t wilcoxon

Suggested: t (The model proposed this value when it asked to run the step.)

Answer t

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: Illustration I. Student uses his own distribution for the mean of a small sample. A rank test is not in the paper.

decision card Assume equal variances (unpaired t-test)

False gives the Welch test. It does not assume equal variances. True gives the Student test. Student's test is correct only if the group variances are equal. The model wants to run compare_two_groups.

Options: yes no

Suggested: true (The model proposed this value when it asked to run the step.)

Answer false

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: Not used. A paired test has one series of differences. The value only applies to an unpaired test.

decision card Sidedness

Two-sided tests for a difference in either direction. One-sided tests for one direction only (the second group greater than the first, or less). Choose a one-sided test only if the other direction is not of interest, and decide before you see the data. The model wants to run compare_two_groups.

Options: two-sided greater less

Suggested: two-sided (The model proposed this value when it asked to run the step.)

Answer two-sided

No person answered in this run. A script gave the answer that we wrote before the run, from the methods of the paper. Where the answer comes from: The modern convention. Student gives the probability that drug 2 is better, which is one direction.

step n2 compare_two_groups adapter biostats 0.2.0, statsmodels, pingouin, lifelines 0.30.3

paired t-test, two-sided: 2 minus 1 = 1.58, statistic 4.062, p = 0.002833 (alpha 0.05).

Decisions applied: Significance level = 0.05; Paired or unpaired test = true; Test family for two groups = t; Equal variances for an unpaired test = false; Sidedness = two-sided.

Input file: {data}/student1908-sleep/sleep.csv SHA-256 5563ff4da6cf.

Outputs: two_group_test.csv (5c80e8fb3427).

Arguments
groupgroup
id_columnID
outcomeextra
path{data}/student1908-sleep/sleep.csv
testt
equal_variancefalse
pairedtrue
alpha0.05
alternativetwo-sided
Tool output
{
 "ok": true,
 "summary": "paired t-test, two-sided: 2 minus 1 = 1.58, statistic 4.062, p = 0.002833 (alpha 0.05).",
 "metrics": {
  "n_a": 10,
  "n_b": 10,
  "mean_diff": 1.58,
  "statistic": 4.062127683382037,
  "df": 9,
  "p_value": 0.0028328901973842724,
  "ci_lo": 0.7001142367230175,
  "ci_hi": 2.459885763276983,
  "effect_size": 1.2845575625910546,
  "mean_a": 0.75,
  "mean_b": 2.3299999999999996
 },
 "data": {
  "test": "paired t-test",
  "alternative": "two-sided",
  "direction": "2 minus 1",
  "alpha": 0.05,
  "significant": true
 },
 "table": {
  "columns": [
   "test",
   "group_a",
   "group_b",
   "n_a",
   "n_b",
   "mean_diff_b_minus_a",
   "statistic",
   "df",
   "p",
   "ci_lo",
   "ci_hi",
   "cohen_dz"
  ],
  "rows": [
   [
    "paired t-test",
    1,
    2,
    10,
    10,
    1.58,
    4.062127683382037,
    9,
    0.0028328901973842724,
    0.7001142367230175,
    2.459885763276983,
    1.2845575625910546
   ]
  ],
  "n_rows": 1,
  "path": "{work}/compare_two_groups-1/two_group_test.csv"
 }
}

Model The model writes its answer. The claim check reads it. See the final answer.

claim check 7 numbers in the answer, 0 numbers without a source in the record.
claim check 7 numbers in the answer, 0 numbers without a source in the record.
Saved: results.xlsx (9 sheets). Folder: {work}

The final answer

The model wrote this answer. The harness shows it to the scientist after the review.

The mean difference in extra hours of sleep between the two drugs is 1.58 hours. The 95% confidence interval ranges from 0.70 to 2.46 hours. The p-value is 0.0028, which is significant at the 0.05 level. This suggests that the second drug (group 2) provides more extra hours of sleep than the first drug (group 1).

Settings and step results that the answer does not mention. The harness adds them.
Step results that the answer does not mention, from the record:
- n1 inspect_table: 20 rows x 4 columns.

Settings used, from the decision record: Significance level (alpha): 0.05 · Paired test: true · Test family for two groups: t · Assume equal variances (unpaired t-test): false · Sidedness: two-sided.

Values that are not scored

  • A trap value is the result of a wrong method, for example an unpaired test on paired data. A trap value in a main step means that the run used the wrong method. A trap value in a comparison run is correct, because the record keeps the scientist's choice.
  • An optional value is a second result of the same method. The run can compute it or not.
  • A reference value comes from a different program version or a check run by us. We show it for comparison.
Table 17 | Values that are not scored, qwen3:8b run.
ItemKindKnown valueClosest logged valueToleranceOutcomeSource of the known value
welch_p_trapUnpaired Welch p (trap result)trap0.079390.05n2 compare_two_groups± 0.0005not in the recordWe calculated it with SciPy 1.x ttest_ind with equal_var=False

Checks

Review findings

The review recorded 1 finding. A rule finding comes from a fixed check in the harness. A referee finding comes from a second model that reads the record. The harness shows the findings to the scientist with the final answer. The record does not mark a finding as fixed. Thus a finding from an early review round can apply to a draft that the model corrected later.

Table 18 | Review findings, qwen3:8b run.
SeverityFromFindingShown with the final answer
inforeferee modelThe p-value is reported with alpha but the test settings are not fully reported.yes

Numbers in the answer

The last claim check read 7 numbers in the answer. 7 numbers match a logged result. 0 numbers have no source in the record.

Deviations

The model did not try to change a choice of the scientist.

Failed tool calls

No tool call failed.

Data integrity

Each data file has the same SHA-256 hash now as at the time of the step that read it. Where the download script (fetch.sh) gives a hash, the file also has that hash. The run did not change the data.

Table 19 | Data files and their SHA-256 hashes, qwen3:8b run.
FileSHA-256Fetched dataSteps with this hash
{data}/student1908-sleep/sleep.csv238 bytes5563ff4da6cfsame as the hash in the download script (fetch.sh)n1, n2

A SHA-256 hash is a fingerprint of the file contents. If one byte of the file changes, the hash changes. The table shows the first 12 characters.

How to repeat it

Get the data. The script downloads the files and checks their SHA-256 hashes where it lists them.

CUVETTE_DATA={data} bash bench/papers/student1908-sleep/fetch.sh

Run the same case with Cuvette. The script gives the same answers from bench/papers/student1908-sleep/bench.yaml.

cuvette bench papers --papers student1908-sleep --models ollama:qwen3:8b

Repeat each step by hand in the program. For each step, the harness records a manual route: the menu path or the code that gives the same result. This list does not include comparison runs.

  1. inspect_table (step n1)

    Code

    print(pd.read_csv(path).describe(include='all'))
    • path

      {data}/student1908-sleep/sleep.csv
    • Note: The tool also counts distinct values and missing values for each column.

    The manual route that the harness recorded

    ga_biostats.inspect_table(path="{data}/student1908-sleep/sleep.csv")

    The manual route uses the same method. The note in the route gives the known difference.

  2. compare_two_groups (step n2)

    Code

    wide = df.pivot(index="ID", columns="group", values="extra")
    scipy.stats.ttest_1samp(wide[2] - wide[1], 0)   # same as ttest_rel(wide[2], wide[1])
    • values column = extra
    • pairing column = ID
    • alternative = two-sided

    The manual route that the harness recorded

    ga_biostats.compare_two_groups(path="{data}/student1908-sleep/sleep.csv", outcome="extra", group="group", test="t", paired=True, id_column="ID", alternative="two-sided", alpha=0.05, equal_variance=False)

    The manual route gives the same numbers. An automatic test in Cuvette checks this.

Figure

Paper-style figure for Student 1908, from the qwen3:8b run
Fig. 6 | qwen3:8b run. Our figure script draws the values of this run in the style of the paper.

Run facts

Table 20 | Run facts, qwen3:8b run.
Modelqwen3:8b through Ollama, on our own computer
Date2026-10-09 11:45:17 UTC
End of runthe model gave a final answer
Time55 s
Requests to the model3
Tokensunits of text that the model read and wrote27436 input, 205 output, 0 cache read, 0 cache write
Cost estimatenone: the model runs on our own computer
Tool calls2 (0 failed)
Adaptersbiostats 0.2.0, program 0.30.3
Session20261009-064517-4311
Code hash of each step (2)
Table 21 | Code hash of each step, qwen3:8b run.
StepToolProgram versionCode hash
n1inspect_table0.30.3603f546a1fe4
n2compare_two_groups0.30.3c6e869486df4

The code hash is a fingerprint of the adapter name, the adapter version, the tool and its definition in the adapter. If one of these changes, the hash changes.