Methods
Validation of Cuvette against published analyses
In Cuvette, we gave an artificial intelligence (AI) model the data and the question of a published analysis. Then we compared its numbers with the known values. Cuvette is a harness: software around an AI model that runs existing science programs and records each step. The test set has 53 papers: 33 research papers and 20 tutorials or test suites of programs. A test suite is the set of checks that comes with a program. We scored 472 values. Each value has a known source: a paper, an official tutorial or our own calculation. All numbers on this page come from the final run of 9 October 2026.
The final run used Claude Opus, Sonnet and Haiku, with three runs for each paper and model. The scorer did not know which model made each run.
1 What we validated
Each test started from a published paper or an official tutorial. The paper gives the data, the question and the method. It also gives the numbers that a correct analysis must find.
The validation tested two things. It tested if the programs gave the correct numbers. It also tested if the model reported those numbers correctly in its final answer.
2 Test set
The test set has 53 papers. Of these, 33 are research papers and 20 are tutorials or test suites of the programs. On this page, “paper” means any of these 53 sources. The papers come from 15 fields: imaging, flow and mass cytometry, mass spectrometry, genomics, phylogenetics, neuroscience, structural biology, chemistry, microbiology, bench assays, statistics, clinical research and epidemiology, ecology, astronomy and geospatial analysis. The paper list shows each paper.
3 Procedure of a run
Fig. 1 shows the five stages of a run. A script acted as the scientist. The model asked questions about the method, for example which threshold to use. The script answered each question with the choice from the paper. The model did not see the known values.
Cuvette wrote each step to the session record. The session record is a log of each step, each setting and each result.
4 Scoring
Each known value has a tolerance. The tolerance is the largest difference that we accept. Some values must be exact, for example a number of patients.
Each value gets two scores:
- Match in the record
- A number in the session record is inside the tolerance of the known value. This score tests the programs and Cuvette.
- Correct in the final answer
- The final answer states the value. A check then finds the step in the session record that made the number. This score tests the model. A number with no source in the session record does not count.
Some values are not scored. These reference values show other results, for example the result of a wrong method that the model must not use. The full rules are in the method document1.
5 Sources of the known values
The validation scored 472 values. Table 1 shows the number of values from each source. The sources document2 gives the source of each value.
For example, the Caicedo 2019 paper prints 5720 nuclei in its 50 test images, and the model must find the same number. The FlowKit tutorial prints 133670 CD3+ cells in its sample file. No paper prints the mean count error for the BBBC005 images (image set 005 of the Broad Bioimage Benchmark Collection). Thus we calculated it with a separate Python script and got 3.11. The model ran CellProfiler and had to match that value.
| Source of the known value | Values |
|---|---|
| Printed in a paper | 276 |
| Printed in an official tutorial | 51 |
| Independent check: we calculated the value with a different program from the one that Cuvette uses | 133 |
| Check with the same program: no second program exists for this value | 12 |
| Total | 472 |
6 Results
Table 2 shows the totals of the final run of 9 October 2026. Each Claude model did three runs for each of the 53 papers. qwen3:8b did one run for each paper. The totals add the values of all runs, so each Claude paper counts three times. One run is a sample. A model can give a different result in the next run.
| Model | Runs with an answer | Match in the record | Correct in the final answer |
|---|---|---|---|
| Claude Opus claude-opus-5-5, 3 runs for each paper | 159 of 159 | 1411 of 1416 99.6% | 1146 of 1152 99.5% |
| Claude Sonnet claude-sonnet-5-5, 3 runs for each paper | 157 of 159 | 1410 of 1416 99.6% | 1143 of 1152 99.2% |
| Claude Haiku claude-haiku-5-5, 3 runs for each paper | 159 of 159 | 1410 of 1416 99.6% | 1133 of 1152 98.4% |
| qwen3:8b a local model through Ollama, 1 run for each paper | 51 of 53 | 376 of 472 79.7% | 240 of 384 62.5% |
A run without an answer stopped at the time limit or with an error. Its values count as no match. The final answer does not have to state every value: the last column counts only the values that the question asks for.
7 Final run
The final run started on 9 October 2026. It had these settings:
- Three Claude models: Opus, Sonnet and Haiku.
- Three runs for each paper and model.
- Blind scoring: the scorer did not know which model made a run.
8 Limits
- qwen3:8b had one run for each paper. The results of small local models change from run to run.
- Three net flux values of the photutils paper count as a match for every model, but the run values are not the correct values. The tolerance of these values is too wide. The benchmark report gives the totals with and without these three values5.
- The check of the final answer finds numbers in sentences. It can miss a number in a table or a list.
- We wrote some questions again by hand from notes of earlier runs.
- A program error, for example a program that is not installed, counts as a value that does not match.
9 Data and code
The paper list shows each paper with its known values and the values that the model got in Cuvette. The paper files3 give the question, the setup and the known values of each paper. To repeat the validation, type these commands in a terminal:
- Get the data of each paper with
bench/papers/<name>/fetch.sh. - Run the validation with this command:
cuvette bench papers --papers all --models claude --runs 3
References
- Validation method: scores, rules for a match, and limits. docs/benchmark-papers.md
- Sources of the known values. docs/benchmark-sources.md
- Paper files. bench/papers
- Result files. bench/results
- Benchmark report of 9 October 2026. docs/benchmark-2026-10-09.md