cuvette Install

Methods

Validation of Cuvette against published analyses

Final run: 9 October 2026.

In Cuvette, we gave an artificial intelligence (AI) model the data and the question of a published analysis. Then we compared its numbers with the known values. Cuvette is a harness: software around an AI model that runs existing science programs and records each step. The test set has 53 papers: 33 research papers and 20 tutorials or test suites of programs. A test suite is the set of checks that comes with a program. We scored 472 values. Each value has a known source: a paper, an official tutorial or our own calculation. All numbers on this page come from the final run of 9 October 2026.

The final run used Claude Opus, Sonnet and Haiku, with three runs for each paper and model. The scorer did not know which model made each run.

1 What we validated

Each test started from a published paper or an official tutorial. The paper gives the data, the question and the method. It also gives the numbers that a correct analysis must find.

The validation tested two things. It tested if the programs gave the correct numbers. It also tested if the model reported those numbers correctly in its final answer.

2 Test set

The test set has 53 papers. Of these, 33 are research papers and 20 are tutorials or test suites of the programs. On this page, “paper” means any of these 53 sources. The papers come from 15 fields: imaging, flow and mass cytometry, mass spectrometry, genomics, phylogenetics, neuroscience, structural biology, chemistry, microbiology, bench assays, statistics, clinical research and epidemiology, ecology, astronomy and geospatial analysis. The paper list shows each paper.

3 Procedure of a run

Fig. 1 shows the five stages of a run. A script acted as the scientist. The model asked questions about the method, for example which threshold to use. The script answered each question with the choice from the paper. The model did not see the known values.

Cuvette wrote each step to the session record. The session record is a log of each step, each setting and each result.

Five stages: paper, data and question, session in Cuvette, answers, comparison with the known valueAPAPERA published paperor an officialtutorialBDATA ANDQUESTIONThe data of the paperand a question fromits methodsCSESSION INCUVETTEThe model runs theprograms. A scriptanswers its questionsDANSWERSThe session recordand the final answerECOMPARISONEach number againstits known valueand its tolerance APAPERA published paper or an officialtutorialBDATA AND QUESTIONThe data of the paper and a question fromits methodsCSESSION IN CUVETTEThe model runs the programs. A scriptanswers its questionsDANSWERSThe session record and the final answerECOMPARISONEach number against its known valueand its tolerance
Fig. 1 | The stages of a validation run. A, We chose a paper with known values. B, We got its data and wrote the question from its methods. C, The model did the analysis in Cuvette, as in a session with a scientist. D, Cuvette wrote each step to the session record, and the model wrote a final answer. E, We compared each number with its known value.

4 Scoring

Each known value has a tolerance. The tolerance is the largest difference that we accept. Some values must be exact, for example a number of patients.

Each value gets two scores:

Match in the record
A number in the session record is inside the tolerance of the known value. This score tests the programs and Cuvette.
Correct in the final answer
The final answer states the value. A check then finds the step in the session record that made the number. This score tests the model. A number with no source in the session record does not count.

Some values are not scored. These reference values show other results, for example the result of a wrong method that the model must not use. The full rules are in the method document1.

5 Sources of the known values

The validation scored 472 values. Table 1 shows the number of values from each source. The sources document2 gives the source of each value.

For example, the Caicedo 2019 paper prints 5720 nuclei in its 50 test images, and the model must find the same number. The FlowKit tutorial prints 133670 CD3+ cells in its sample file. No paper prints the mean count error for the BBBC005 images (image set 005 of the Broad Bioimage Benchmark Collection). Thus we calculated it with a separate Python script and got 3.11. The model ran CellProfiler and had to match that value.

Table 1 | Sources of the 472 known values.
Source of the known valueValues
Printed in a paper276
Printed in an official tutorial51
Independent check: we calculated the value with a different program from the one that Cuvette uses133
Check with the same program: no second program exists for this value12
Total472

6 Results

Table 2 shows the totals of the final run of 9 October 2026. Each Claude model did three runs for each of the 53 papers. qwen3:8b did one run for each paper. The totals add the values of all runs, so each Claude paper counts three times. One run is a sample. A model can give a different result in the next run.

Table 2 | Totals of the final run of 9 October 2026, 53 papers, 530 runs.
ModelRuns with an answerMatch in the recordCorrect in the final answer
Claude Opus claude-opus-5-5, 3 runs for each paper159 of 1591411 of 1416 99.6%1146 of 1152 99.5%
Claude Sonnet claude-sonnet-5-5, 3 runs for each paper157 of 1591410 of 1416 99.6%1143 of 1152 99.2%
Claude Haiku claude-haiku-5-5, 3 runs for each paper159 of 1591410 of 1416 99.6%1133 of 1152 98.4%
qwen3:8b a local model through Ollama, 1 run for each paper51 of 53376 of 472 79.7%240 of 384 62.5%

A run without an answer stopped at the time limit or with an error. Its values count as no match. The final answer does not have to state every value: the last column counts only the values that the question asks for.

7 Final run

The final run started on 9 October 2026. It had these settings:

8 Limits

9 Data and code

The paper list shows each paper with its known values and the values that the model got in Cuvette. The paper files3 give the question, the setup and the known values of each paper. To repeat the validation, type these commands in a terminal:

  1. Get the data of each paper with bench/papers/<name>/fetch.sh.
  2. Run the validation with this command:
cuvette bench papers --papers all --models claude --runs 3

References

  1. Validation method: scores, rules for a match, and limits. docs/benchmark-papers.md
  2. Sources of the known values. docs/benchmark-sources.md
  3. Paper files. bench/papers
  4. Result files. bench/results
  5. Benchmark report of 9 October 2026. docs/benchmark-2026-10-09.md