cuvette Install

Validation / Papers / Minh 2020

Minh 2020: IQ-TREE 2, example.phy tutorial data

Phylogenetics · tool tutorial or software test data · IQ-TREE 3.1.4 and MAFFT (command line), through the phylo adapter

How to read this page

In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. The run is one sample: another run can give different steps and numbers. The model is the AI. The harness is Cuvette, the software around the model: it runs the programs and records each step. A tool call is a request from the model to run one program step. The session record is the log of each message and each step. The claim check is a script that finds each number of the final answer in the step results. The review is a set of fixed rule checks plus a second AI model, the referee, that reads the record. A deviation is a request from the model for a setting that differs from the choice of the scientist. Each Claude model did 3 runs of this paper. This page shows run 3 of each Claude model and the one run of qwen3:8b. The table of values says how many of the Claude runs match.

Opus: 1 of 1 values match, 1 of 1 correct in the final answer. All 3 runs: 1 of 1 values match. Sonnet: 1 of 1 values match, 1 of 1 correct in the final answer. All 3 runs: 1 of 1 values match. Haiku: 1 of 1 values match, 1 of 1 correct in the final answer. All 3 runs: 1 of 1 values match. qwen3:8b: 1 of 1 values match, 1 of 1 correct in the final answer.

The figure in the paper and in the run

As published

The IQ-TREE 2 paper (Minh et al. 2020) has no figure or number for example.phy. The beginner tutorial of IQ-TREE uses this file and names the model TIM2+I+G4. The tutorial prints no log-likelihood. We do not copy the tutorial text or the paper figures because their license is not CC BY.

See the figure in the paper

Fig. 1 | As published. This page does not show the published figure. The link opens the paper.

Reproduced in Cuvette

The figure reproduced from this run in Cuvette
Fig. 2 | Reproduced in Cuvette. Reproduction of the IQ-TREE tutorial analysis, drawn from the tree file and the report of the run (Claude Opus 5.5 in the harness, 9 October 2026, IQ-TREE 3.1.4, option -m MFP -B 1000, seed 1). The data are example.phy, 17 animal mitochondrial DNA sequences with 1,998 sites. (a) The maximum-likelihood tree with the branch lengths of the file and the ultrafast bootstrap support at each branch. Red numbers show support below 95. The tree is unrooted. The drawing starts at the first sequence, as IQ-TREE does. (b) The known log-likelihood (open ring) and the run value (red dot), on a scale of the tolerance. ModelFinder chose TIM2+F+I+G4.

The paper

Minh BQ, Schmidt HA, Chernomor O, Schrempf D, Woodhams MD, von Haeseler A, Lanfear R. IQ-TREE 2: New models and efficient methods for phylogenetic inference in the genomic era. Molecular Biology and Evolution 37(5):1530-1534 (2020). doi:10.1093/molbev/msaa015

Related sources:

What it measured

The IQ-TREE 2 paper describes a program that builds maximum-likelihood (ML) trees from sequence alignments. The paper has no numbers for the example file. The beginner tutorial uses example.phy, an alignment of mitochondrial DNA from 17 animals. The tutorial lets ModelFinder choose the substitution model. It then builds the tree and gives branch support with the ultrafast bootstrap.

Data

File example.phy from the example folder of the IQ-TREE 2 source repository. Size: 34 KB, 17 sequences by 1998 sites.

License: The repository is GPL-2.0. The file holds published mitochondrial DNA sequences of animals and has no personal data.

Data source

The instruction

A script sent this message as the scientist. The file paths point to the fetched data.

ScientistWhat is the best-fit substitution model and the maximum-likelihood tree, with branch support, for the 17 mitochondrial sequences in {data}/minh2020-iqtree2/example.phy? Use the model chosen by the data, and report the log-likelihood. Call the tree tool.

The same request in the words of the paper's method:

I have an alignment of mitochondrial DNA from 17 animals. Find the best-fit substitution model and build the maximum-likelihood tree with branch support. Let the data choose the model. Report the log-likelihood of the tree.

Basis: The tutorial sections on the first run, on the choice of the substitution model, and on branch support with the ultrafast bootstrap.

Results

Match: a number in the session record is inside the tolerance of the known value. In the final answer: the model also stated the value in its final answer. For a Claude model, each cell shows the run that this page shows. If the three runs differ, the cell also says in how many runs the value matches.

Table 1 | Known values and the value of each model.
ValueKnown valueToleranceOpusSonnetHaikuqwen3:8b
log_likelihoodLog-likelihood of the ML tree
Source of the known valueWe calculated it with IQ-TREE 3.1.4 (iqtree3), seed 1Not in the paper or the tutorial. The tutorial names the model TIM2+I+G4 but prints no log-likelihood. We ran -m MFP -B 1000 and got -21152.5377.
-21152.54± 1-21152.54 matchIn the final answer: yes (-21152.54)Log: n1 infer_tree metrics.log_likelihood, entry 11; the final answer, entry 42-21152.54 matchIn the final answer: yes (-21152.54)Log: n1 infer_tree metrics.log_likelihood, entry 11; the final answer, entry 38-21152.54 matchIn the final answer: yes (-21152.54)Log: n1 infer_tree metrics.log_likelihood, entry 11; the final answer, entry 64-21152.54 matchIn the final answer: yes (-21152.54)Log: n1 infer_tree metrics.log_likelihood, entry 14; the final answer, entry 29

Session records

Session record, Opus, run 3 of 3

Every message, decision, step and result of this run, one JSON object for each log entry.

A compressed log file (JSON Lines, gzip). The record shows file paths as {data}, {work} and similar names. · 16 KB

Download

Session record, Sonnet, run 3 of 3

Every message, decision, step and result of this run, one JSON object for each log entry.

A compressed log file (JSON Lines, gzip). The record shows file paths as {data}, {work} and similar names. · 11 KB

Download

Session record, Haiku, run 3 of 3

Every message, decision, step and result of this run, one JSON object for each log entry.

A compressed log file (JSON Lines, gzip). The record shows file paths as {data}, {work} and similar names. · 27 KB

Download

Session record, qwen3:8b

Every message, decision, step and result of this run, one JSON object for each log entry.

A compressed log file (JSON Lines, gzip). The record shows file paths as {data}, {work} and similar names. · 4 KB

Download

Opus · claude-opus-5-5 · run 3 of 3 shown 1 of 1 values match, 1 of 1 correct in the final answer

The session

This is the session as the scientist sees it, in order. Decision cards show the answer that the script gave and where that answer comes from. Each step shows the program, its version, the input file hash and the outputs. Click a line to see more.

setup The decision record starts with these values. A script set them before the first message, from the answers that we wrote before the run.

  • MAFFT alignment method: autoSource in the tutorial or test suite: Not used. The file example.phy is already aligned. The phylo adapter uses MAFFT --auto by default.
  • Trim the alignment: falseSource in the tutorial or test suite: Not in the tutorial. The tutorial uses the alignment as it is.
  • Substitution model: MFPSource in the tutorial or test suite: Tutorial section on the substitution model. ModelFinder picks the model with the lowest Bayesian information criterion (BIC). MFP is the default from version 1.5.4.
  • Type of branch support: ufbootSource in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial notes that these values do not mean the same as standard bootstrap values.
  • Bootstrap replicates: 1000Source in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial calls 1000 the minimum number it recommends.
  • Outgroup: Source in the tutorial or test suite: Tutorial, first run. IQ-TREE draws the lungfish at the root only because it is the first sequence. The tree stays unrooted.
  • Random seed: 1
harness note A note from the harness to the model. The scientist did not write it.
- Setup values:
Alignment:
- Alignment method (MAFFT) (algorithm): auto
- Trim the alignment (trim): false
Tree:
- Substitution model (model): MFP
- Bootstrap type (bootstrap_type): ufboot
- Number of bootstrap replicates (replicates): 1000
- Outgroup (outgroup): 
- Random seed (seed): 1
Ask the scientist: Largest gap fraction of a kept column (max_gap_fraction).
Fixed by the lab: Random seed (seed) = 1.
The harness fills these values into your tool calls automatically. If you give a different value, the harness logs a deviation and shows it to the scientist.
- Mode is now watch: you run the steps.

Model

The file is a PHYLIP alignment, so it needs no alignment step. Next, I build the maximum-likelihood tree. ModelFinder chooses the model, and 1000 ultrafast bootstrap replicates give the branch support.

The model calls infer_tree (adapter phylo).

step n1 infer_tree adapter phylo 0.1.2, IQ-TREE 3.1.4

Model TIM2+F+I+G4. Log-likelihood -21152.5. Ultrafast bootstrap with 1000 replicates.

Decisions applied: Substitution model = MFP; Type of branch support = ufboot; Bootstrap replicates = 1000; Random seed = 1.

Input file: {data}/minh2020-iqtree2/example.phy SHA-256 d3365cbea79a.

Outputs: run.bionj (c4103d181dee), run.ckp.gz (2c8f49979d62), run.contree (cf1f8eae6015), run.iqtree (837ec22b9ee3), run.log (99b88cd1d7b0), run.mldist (7c9d04e6c668), run.model.gz (959c833b94dc), run.splits.nex (650ac7edfeae), run.treefile (4cff6aaf2ea1), stderr.txt (e3b0c44298fc), stdout.txt (e3b0c44298fc).

Arguments
path{data}/minh2020-iqtree2/example.phy
modelMFP
bootstrap_typeufboot
replicates1000
seed1
Tool output
report={work}/infer_tree-1/run.iqtree
seed=1
bootstrap_type=ufboot
bootstrap_replicates=1000
treefile={work}/infer_tree-1/run.treefile
consensus_tree={work}/infer_tree-1/run.contree
n_sequences=17
n_sites=1998
model=TIM2+F+I+G4
best_model_bic=TIM2+F+I+G4
log_likelihood=-21152.5377
tree_length=4.2173
constant_sites=686
site_patterns=1152

Model

Next, I read the tree file to get the topology and the support values.

The model calls read_file.
Result of read_file: Read 590 of 590 bytes from {work}/infer_tree-1/run.treefile.
(LngfishAu:0.1709068277,(LngfishSA:0.1883311672,LngfishAf:0.1648032873)100:0.1073256416,(Frog:0.2561720790,((((Turtle:0.2214955116,(Crocodile:0.3056436573,Bird:0.2309752831)97:0.0651735079)70:0.0364941210,Sphenodon:0.3446303789)51:0.0204677635,Lizard:0.3859794405)99:0.0739928283,(((Human:0.1850870817,(Seal:0.0943875559,(Cow:0.0822572767,Whale:0.1012437124)99:0.0404174962)71:0.0252594799)93:0.0340514112,(Mouse:0.0583901122,Rat:0.0905005039)100:0.1217432345)100:0.0607172631,(Platypus:0.1919645817,Opossum:0.1509515712)96:0.0373002662)100:0.1490260063)100:0.1275594673)100:0.0940391076);
The model calls read_file.
Result of read_file: Read 12000 of 19016 bytes from {work}/infer_tree-1/run.iqtree.
IQ-TREE 3.1.4 built Sep 10 2026

Input file name: {data}/minh2020-iqtree2/example.phy
Type of analysis: ModelFinder + tree reconstruction + ultrafast bootstrap (1000 replicates)
Random seed number: 1

REFERENCES
----------

To cite IQ-TREE 3 please use:

Thomas K.F. Wong, Nhan Ly-Trong, Huaiyan Ren, Piyumal Demotte, Hector Banos,
Andrew J. Roger, Edward Susko, Chris Bielow, Nicola De Maio, Nick Goldman,
Matthew W. Hahn, Mario dos Reis, Le Sy Vinh, Gavin Huttley, Robert Lanfear, Bui Quang Minh (2026)
IQ-TREE 3: Phylogenomic Inference Software using Complex Evolutionary Models.
Molecular Biology and Evolution, msag117.
https://doi.org/10.1093/molbev/msag117

Since you used ModelFinder please also cite: 

Subha Kalyaanamoorthy, Bui Quang Minh, Thomas KF Wong, Arndt von Haeseler,
and Lars S Jermiin (2017) ModelFinder: Fast model selection for
accurate phylogenetic estimates. Nature Methods, 14:587–589.
https://doi.org/10.1038/nmeth.4285

Since you used ultrafast bootstrap (UFBoot) please also cite: 

Diep Thi Hoang, Olga Chernomor, Arndt von Haeseler, Bui Quang Minh,
and Le Sy Vinh (2018) UFBoot2: Improving the ultrafast bootstrap
approximation. Mol. Biol. Evol., 35:518–522.
https://doi.org/10.1093/molbev/msx281

SEQUENCE ALIGNMENT
------------------

NOTE: Alignment sequence type is auto-detected. If in doubt, specify it via -st option.
Input data: 17 sequences with 1998 nucleotide sites
Number of constant sites: 686 (= 34.3343% of all sites)
Number of invariant (constant or ambiguous constant) sites: 686 (= 34.3343% of all sites)
Number of parsimony informative sites: 1009
Number of distinct site patterns: 1152

ModelFinder
-----------

Best-fit model according to BIC: TIM2+F+I+G4

List of models sorted by BIC scores: 

Model                  LogL         AIC      w-AIC        AICc     w-AICc         BIC      w-BIC
TIM2+F+I+G4      -21152.739   42383.479 +   0.0741   42385.072 +   0.0828   42601.875 +    0.974
GTR+F+I+G4       -21148.957   42379.913 +    0.441   42381.674 +    0.453   42609.509 -   0.0214
TIM2+F+R3        -21150.544   42383.088 +   0.0901   42384.849 +   0.0926   42612.684 -  0.00438
GTR+F+R3         -21147.067   42380.134 +    0.395   42382.071 +    0.372   42620.930 - 7.09e-05
TPM2u+F+R3       -21283.140   42646.279 - 6.36e-59   42647.955 - 6.82e-59   42870.275 - 5.08e-59
TPM2u+F+I+G4     -21291.534   42659.067 - 1.06e-61   42660.580 - 1.24e-61   42871.863 -  2.3e-59
TVM+F+R3         -21279.783   42643.567 - 2.47e-58   42645.414 - 2.43e-58   42878.762 -  7.3e-61
TVM+F+I+G4       -21287.937   42655.873 - 5.25e-61   42657.549 - 5.63e-61   42879.870 -  4.2e-61
TIM2+I+G4        -21308.990   42689.980 - 2.06e-68   42691.338 - 2.59e-68   42891.576 -  1.2e-63
TIM2e+I+G4       -21309.009   42690.018 - 2.02e-68   42691.376 - 2.54e-68   42891.614 - 1.18e-63
TIM2e+I+G4       -21309.009   42690.018 - 2.02e-68   42691.376 - 2.54e-68   42891.614 - 1.18e-63
TIM2+R3          -21303.599   42683.199 - 6.12e-67   42684.712 - 7.12e-67   42895.995 - 
... (974 more characters in the session record)

Model The model writes its answer. The claim check reads it. See the final answer.

claim check 47 numbers in the answer, 3 numbers without a source in the record.
The review found 1 problem. The model corrects the answer before the scientist sees it.

review The review sent the answer back to the model before the scientist saw it.

- 3 numbers in the answer match no logged tool result: 95. Run the tool that measures each number, or remove the number, or say how you derived it from logged results.
Fix these problems. Then give the answer again.
Model Answer draft 2. The claim check and the review read it.

I ran one IQ-TREE 3.1.4 job. ModelFinder chose TIM2+F+I+G4 by BIC (Bayesian information criterion). The maximum-likelihood tree has a log-likelihood of −21152.5377. Nine branches are well supported, and four branches are below the lab threshold.

Source of the threshold: the lab reporting standards say that an ultrafast bootstrap value of 95 or more is well supported. This value is a rule, not a measurement. In this answer, "the UFBoot threshold" means that rule.

What I did

  1. The input file is a PHYLIP alignment that is already aligned. Thus I did not run MAFFT and did not trim. The trimming gap fraction does not apply.
  2. I ran the tree tool once. It did the model selection, the tree search and the support in the same run.
  3. Model: ModelFinder (--model MFP).
  4. Support: ultrafast bootstrap (UFBoot), 1000 replicates.
  5. Random seed: 1, fixed by the lab.
  6. Command: iqtree3 infer --alignment example.phy --model MFP --bootstrap-type ufboot --replicates 1000 --seed 1

Data (from the IQ-TREE report)

  • 17 sequences and 1998 nucleotide sites.
  • 686 constant sites, 1009 parsimony-informative sites and 1152 distinct site patterns.

Best-fit model (from the ModelFinder table in run.iqtree)

  • BIC chose TIM2+F+I+G4, with a BIC weight of 0.974.
  • The next model by BIC is GTR+F+I+G4, with a BIC weight of 0.0214.
  • AIC (Akaike information criterion) gives a different order:
  • GTR+F+I+G4 is first, with an AIC weight of 0.441.
  • TIM2+F+I+G4 has an AIC weight of 0.0741.
  • The report marks four models with a plus sign in the AIC column: TIM2+F+I+G4, GTR+F+I+G4, TIM2+F+R3 and GTR+F+R3. The report says that a plus sign means the model is in the confidence set.
  • TIM2 rates: A-C = A-T = 5.5843, A-G = 7.6324, C-T = 22.6443, C-G = G-T = 1.0.
  • Base frequencies, counted from the alignment: A 0.3547, C 0.2282, G 0.1919, T 0.2252.

Tree

  • Log-likelihood: −21152.5377. Total tree length: 4.2173 substitutions per site.
  • Tree with support values: infer_tree-1/run.treefile. The bootstrap consensus tree is run.contree. The full report is run.iqtree.
  • The tree is unrooted because no outgroup was set. The Newick file starts at the lungfish, but that is only how the file is written. It does not show a root or a time order.

Branch support (UFBoot values from run.treefile)

GroupUFBoot
Lungfish South America + Lungfish Africa100
Amniotes (reptiles, birds and mammals)100
Mammals100
Lizard + archosaurs + turtle + Sphenodon99
Placental mammals100
Mouse + Rat100
Cow + Whale99
Crocodile + Bird97
Platypus + Opossum96
Human + (Seal, Cow, Whale)93
Seal + (Cow, Whale)71
Turtle + (Crocodile, Bird)70
Sphenodon + (Turtle, Crocodile, Bird)51

What is uncertain

  • Four groups have values below the UFBoot threshold, so they are not well supported:
  • Sphenodon + (Turtle, Crocodile, Bird): 51.
  • Turtle + (Crocodile, Bird): 70.
  • Seal + (Cow, Whale): 71.
  • Human + (Seal, Cow, Whale): 93.
  • This is the maximum-likelihood tree under TIM2+F+I+G4. It is not the true tree.
  • The model choice depends on the criterion: BIC prefers TIM2+F+I+G4, and AIC prefers GTR+F+I+G4. I did not test if GTR+F+I+G4 gives the same tree.
  • I did not read the gamma shape and the proportion of invariable sites. They are later in run.iqtree, after the part that I read.

What waits for you

  • Outgroup: if you want a rooted tree, name the outgroup. The three lungfish are one possible choice for this data set.
  • Model: tell me if you want a check run with GTR+F+I+G4, the model that AIC prefers.
claim check 43 numbers in the answer, 1 number without a source in the record.
The revision lost results and fixed no problem, so the harness kept the first answer.
claim check 47 numbers in the answer, 3 numbers without a source in the record.
claim check 47 numbers in the answer, 3 numbers without a source in the record.
Saved: results.xlsx (5 sheets). Folder: {work}

The final answer

The model wrote this answer. The harness shows it to the scientist after the review.

I ran one IQ-TREE 3.1.4 job. ModelFinder chose TIM2+F+I+G4 by BIC (Bayesian information criterion). The maximum-likelihood tree has a log-likelihood of −21152.5377. Most branches have ultrafast bootstrap support of 96 to 100, but four branches have less than 95.

What I did

  1. The input file is a PHYLIP alignment that is already aligned. Thus I did not run MAFFT and did not trim. The trimming gap fraction does not apply.
  2. I ran the tree tool once. It did the model selection, the tree search and the support in the same run.
  3. Model: ModelFinder (--model MFP).
  4. Support: ultrafast bootstrap (UFBoot), 1000 replicates.
  5. Random seed: 1, fixed by the lab.
  6. Command: iqtree3 infer --alignment example.phy --model MFP --bootstrap-type ufboot --replicates 1000 --seed 1

Data (from the IQ-TREE report)

  • 17 sequences and 1998 nucleotide sites.
  • 686 constant sites, 1009 parsimony-informative sites and 1152 distinct site patterns.

Best-fit model

  • BIC chose TIM2+F+I+G4, with a BIC weight of 0.974.
  • The next model by BIC is GTR+F+I+G4, with a BIC weight of 0.0214.
  • By AIC (Akaike information criterion), GTR+F+I+G4 is first, with an AIC weight of 0.441. TIM2+F+I+G4 has an AIC weight of 0.0741. All four top models are in the AIC 95% confidence set.
  • TIM2 rates: A-C = A-T = 5.5843, A-G = 7.6324, C-T = 22.6443, C-G = G-T = 1.0.
  • Base frequencies, counted from the alignment: A 0.3547, C 0.2282, G 0.1919, T 0.2252.

Tree

  • Log-likelihood: −21152.5377. Total tree length: 4.2173 substitutions per site.
  • Tree with support values: infer_tree-1/run.treefile. The bootstrap consensus tree is run.contree. The full report is run.iqtree.
  • The tree is unrooted because no outgroup was set. The Newick file starts at the lungfish, but that is only how the file is written. It does not show a root or a time order.

Branch support (UFBoot: 95 or more is well supported)

GroupUFBoot
Lungfish South America + Lungfish Africa100
Amniotes (reptiles, birds and mammals)100
Mammals100
Lizard + archosaurs + turtle + Sphenodon99
Placental mammals100
Mouse + Rat100
Cow + Whale99
Crocodile + Bird97
Platypus + Opossum96
Human + (Seal, Cow, Whale)93
Seal + (Cow, Whale)71
Turtle + (Crocodile, Bird)70
Sphenodon + (Turtle, Crocodile, Bird)51

What is uncertain

  • These four groups have UFBoot less than 95, so they are not well supported:
  • Sphenodon + (Turtle, Crocodile, Bird): 51.
  • Turtle + (Crocodile, Bird): 70.
  • Seal + (Cow, Whale): 71.
  • Human + (Seal, Cow, Whale): 93.
  • This is the maximum-likelihood tree under TIM2+F+I+G4. It is not the true tree.
  • The model choice depends on the criterion. BIC prefers TIM2+F+I+G4, but AIC prefers GTR+F+I+G4. These two models differ only by two rate parameters, so the tree can be the same with either model. I did not test this.
  • I did not read the gamma shape and the proportion of invariable sites. They are later in run.iqtree, after the part that I read.

What waits for you

  • Outgroup: if you want a rooted tree, name the outgroup. The three lungfish are one possible choice for this data set.
  • Model: tell me if you want a check run with GTR+F+I+G4, the model that AIC prefers.
Settings and step results that the answer does not mention. The harness adds them.
Settings used, from the decision record: Substitution model: MFP · Bootstrap type: ufboot · Number of bootstrap replicates: 1000 · Random seed: 1.

Checks

Review findings

The review recorded 8 findings. A rule finding comes from a fixed check in the harness. A referee finding comes from a second model that reads the record. The harness shows the findings to the scientist with the final answer. The record does not mark a finding as fixed. Thus a finding from an early review round can apply to a draft that the model corrected later.

Table 2 | Review findings, Opus run.
SeverityFromFindingShown with the final answer
errorruleunsourced_numbers3 numbers in the answer match no logged tool result: 95. Run the tool that measures each number, or remove the number, or say how you derived it from logged results.yes
inforuletext_styleThe answer breaks the text rules (ASD-STE100) in 3 places. Sentence 6 uses the passive voice: "is already aligned". Use the active voice. Sentence 32 uses the passive voice: "is unrooted". Use the active voice. Sentence 33 uses the passive voice: "is written". Use the active voice.yes
inforeferee modelThe report names the model TIM2+F+I+G4. It says that ModelFinder chose it by BIC, and it names UFBoot with 1000 replicates. It uses the 95 threshold, which is correct for the ultrafast bootstrap. These items agree with the setup and the tool result.yes
warningreferee modelThe command line in the answer is a reconstruction. The log shows only a truncated manual route, so the log does not show the flags after the alignment path. The report must quote the logged command or mark this line as a paraphrase.yes
warningreferee modelThe tree has 17 sequences and is unrooted, so it has 14 internal branches. The support table lists only 13 groups. One branch is missing, so the statement that four branches have less than 95 is not fully supported. The report must list all internal branches or say which branch it left out.yes
inforeferee modelThe step read only 12000 of 19016 bytes of run.iqtree. The report gives the IQ-TREE version, the BIC and AIC weights and the rate parameters from this partial read, and it says that it did not read the gamma shape or the invariable-site proportion. The report states this limit correctly.yes
inforeferee modelThe report says that the alignment was not trimmed and was not realigned because the input was PHYLIP. This satisfies the trim check. The report does not give the gap fraction of the input alignment, which the standards ask for.yes
inforeferee modelThe answer says that the lab fixed the seed. The setup shows that the seed was an accepted default, so the harness fixed it. The value of 1 is correct.yes

Numbers in the answer

The last claim check read 47 numbers in the answer. 44 numbers match a logged result. 3 numbers have no source in the record.

Numbers that do not match a logged result (3)
  • no source in the record: Most branches have ultrafast bootstrap support of 96 to 100, but four branches have less than 95.
  • no source in the record: **Branch support (UFBoot: 95 or more is well supported)**
  • no source in the record: - These four groups have UFBoot less than 95, so they are not well supported:

Deviations

The model did not try to change a choice of the scientist.

Failed tool calls

No tool call failed.

Data integrity

Each data file has the same SHA-256 hash now as at the time of the step that read it. Where the download script (fetch.sh) gives a hash, the file also has that hash. The run did not change the data.

Table 3 | Data files and their SHA-256 hashes, Opus run.
FileSHA-256Fetched dataSteps with this hash
{data}/minh2020-iqtree2/example.phy33.4 KBd3365cbea79asame as the hash in the download script (fetch.sh)n1

A SHA-256 hash is a fingerprint of the file contents. If one byte of the file changes, the hash changes. The table shows the first 12 characters.

How to repeat it

Get the data. The script downloads the files and checks their SHA-256 hashes where it lists them.

CUVETTE_DATA={data} bash bench/papers/minh2020-iqtree2/fetch.sh

Run the same case with Cuvette. The script gives the same answers from bench/papers/minh2020-iqtree2/bench.yaml.

cuvette bench papers --papers minh2020-iqtree2 --models claude:claude-opus-5-5

Repeat each step by hand in the program. For each step, the harness records a manual route: the menu path or the code that gives the same result. This list does not include comparison runs.

  1. infer_tree (step n1)

    Run: iqtree3 -s <alignment> -m <model> -B <replicates> -seed <seed> -T 1

    • -s

      {data}/minh2020-iqtree2/example.phy
    • -m = MFP
    • -B = 1000
    • -seed = 1
    • Warning: If you keep the default random, you get a different result.

    The manual route that the harness recorded

    /bin/sh {other volume}/tools/overnight/claude-final/catalog/phylo/scripts/iqtree.sh /opt/homebrew/bin/iqtree3 infer --alignment {data}/minh2020-iqtree2/example.phy --model MFP --bootstrap-type ufboot --replicates 1000 --seed 1

    The manual route gives the same numbers. An automatic test in Cuvette checks this.

Figure

Paper-style figure for Minh 2020, from the Opus run
Fig. 3 | Opus run. Our figure script draws the values of this run in the style of the paper.

Run facts

Table 4 | Run facts, Opus run.
Modelclaude-opus-5-5 through the Anthropic service
Date2026-10-09 12:37:04 UTC
End of runthe model gave a final answer
Time90 s
Requests to the model4
Tokensunits of text that the model read and wrote12 input, 4474 output, 39512 cache read, 20884 cache write
Cost estimate$0.20 at list price, from the token counts
Tool calls3 (0 failed)
Adaptersphylo 0.1.2, program 3.1.4
Session20261009-073704-eb13
Code hash of each step (1)
Table 5 | Code hash of each step, Opus run.
StepToolProgram versionCode hash
n1infer_tree3.1.426caef57428d

The code hash is a fingerprint of the adapter name, the adapter version, the tool and its definition in the adapter. If one of these changes, the hash changes.

Sonnet · claude-sonnet-5-5 · run 3 of 3 shown 1 of 1 values match, 1 of 1 correct in the final answer

The session

This is the session as the scientist sees it, in order. Decision cards show the answer that the script gave and where that answer comes from. Each step shows the program, its version, the input file hash and the outputs. Click a line to see more.

setup The decision record starts with these values. A script set them before the first message, from the answers that we wrote before the run.

  • MAFFT alignment method: autoSource in the tutorial or test suite: Not used. The file example.phy is already aligned. The phylo adapter uses MAFFT --auto by default.
  • Trim the alignment: falseSource in the tutorial or test suite: Not in the tutorial. The tutorial uses the alignment as it is.
  • Substitution model: MFPSource in the tutorial or test suite: Tutorial section on the substitution model. ModelFinder picks the model with the lowest Bayesian information criterion (BIC). MFP is the default from version 1.5.4.
  • Type of branch support: ufbootSource in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial notes that these values do not mean the same as standard bootstrap values.
  • Bootstrap replicates: 1000Source in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial calls 1000 the minimum number it recommends.
  • Outgroup: Source in the tutorial or test suite: Tutorial, first run. IQ-TREE draws the lungfish at the root only because it is the first sequence. The tree stays unrooted.
  • Random seed: 1
harness note A note from the harness to the model. The scientist did not write it.
- Setup values:
Alignment:
- Alignment method (MAFFT) (algorithm): auto
- Trim the alignment (trim): false
Tree:
- Substitution model (model): MFP
- Bootstrap type (bootstrap_type): ufboot
- Number of bootstrap replicates (replicates): 1000
- Outgroup (outgroup): 
- Random seed (seed): 1
Ask the scientist: Largest gap fraction of a kept column (max_gap_fraction).
Fixed by the lab: Random seed (seed) = 1.
The harness fills these values into your tool calls automatically. If you give a different value, the harness logs a deviation and shows it to the scientist.
- Mode is now watch: you run the steps.

Model thinking, summary: This file is already a PHYLIP-aligned alignment, so I can skip straight to tree inference without an outgroup, keeping it unrooted and without trimming. I'll run infer_tree with the MFP model and ufboot set to 1000 replicates.

Model

The file is already aligned (PHYLIP), so I skip MAFFT. I now run IQ-TREE with ModelFinder and 1000 ultrafast bootstrap replicates.

The model calls infer_tree (adapter phylo).

step n1 infer_tree adapter phylo 0.1.2, IQ-TREE 3.1.4

Model TIM2+F+I+G4. Log-likelihood -21152.5. Ultrafast bootstrap with 1000 replicates.

Decisions applied: Substitution model = MFP; Type of branch support = ufboot; Bootstrap replicates = 1000; Random seed = 1.

Input file: {data}/minh2020-iqtree2/example.phy SHA-256 d3365cbea79a.

Outputs: run.bionj (c4103d181dee), run.ckp.gz (a8405163edbe), run.contree (cf1f8eae6015), run.iqtree (e559067d0d42), run.log (024d2aeaf644), run.mldist (7c9d04e6c668), run.model.gz (09fc415529a4), run.splits.nex (650ac7edfeae), run.treefile (4cff6aaf2ea1), stderr.txt (e3b0c44298fc), stdout.txt (e3b0c44298fc).

Arguments
path{data}/minh2020-iqtree2/example.phy
modelMFP
bootstrap_typeufboot
replicates1000
seed1
Tool output
report={work}/infer_tree-1/run.iqtree
seed=1
bootstrap_type=ufboot
bootstrap_replicates=1000
treefile={work}/infer_tree-1/run.treefile
consensus_tree={work}/infer_tree-1/run.contree
n_sequences=17
n_sites=1998
model=TIM2+F+I+G4
best_model_bic=TIM2+F+I+G4
log_likelihood=-21152.5377
tree_length=4.2173
constant_sites=686
site_patterns=1152

Model thinking, summary: I should check the support values, so let me quickly read the treefile to report them.

Model

I now read the tree file to check the branch support values.

The model calls read_file.
Result of read_file: Read 590 of 590 bytes from {work}/infer_tree-1/run.treefile.
(LngfishAu:0.1709068277,(LngfishSA:0.1883311672,LngfishAf:0.1648032873)100:0.1073256416,(Frog:0.2561720790,((((Turtle:0.2214955116,(Crocodile:0.3056436573,Bird:0.2309752831)97:0.0651735079)70:0.0364941210,Sphenodon:0.3446303789)51:0.0204677635,Lizard:0.3859794405)99:0.0739928283,(((Human:0.1850870817,(Seal:0.0943875559,(Cow:0.0822572767,Whale:0.1012437124)99:0.0404174962)71:0.0252594799)93:0.0340514112,(Mouse:0.0583901122,Rat:0.0905005039)100:0.1217432345)100:0.0607172631,(Platypus:0.1919645817,Opossum:0.1509515712)96:0.0373002662)100:0.1490260063)100:0.1275594673)100:0.0940391076);
Model Answer draft 1. The claim check and the review read it.

Result. ModelFinder chose TIM2+F+I+G4 as the best-fit model. The maximum-likelihood tree under that model has a log-likelihood of -21152.5377.

What I did

  1. The file example.phy is already aligned, so I skipped MAFFT and did not trim it.
  2. I ran infer_tree in IQ-TREE 3.1.4 with ModelFinder (model MFP), 1000 ultrafast bootstrap replicates and seed 1. The seed is fixed by the lab.

What I measured (from infer_tree)

  • The data have 17 sequences and 1998 sites. Of these, 686 sites are constant and there are 1152 site patterns.
  • The model is TIM2+F+I+G4. The run report lists it as the best model by BIC (Bayesian information criterion).
  • The log-likelihood is -21152.5377, and the tree length is 4.2173.
  • Tree files are in the infer_tree-1 folder. run.treefile is the maximum-likelihood tree with support values. run.contree is the consensus tree. run.iqtree is the report.

Branch support (ultrafast bootstrap, from run.treefile) These are not standard bootstrap values. A value of 95 or more is well supported.

  • These groups are well supported:
  • Lungfish South America and Africa (100).
  • The clade of Frog, reptiles, birds and mammals (100).
  • Mouse and Rat (100).
  • Human, Seal, Cow, Whale, Mouse and Rat together (100).
  • Platypus and Opossum (96).
  • Crocodile and Bird (97).
  • The clade of Lizard, Sphenodon, Turtle, Crocodile and Bird (99).
  • Cow and Whale (99).
  • Human, Seal, Cow and Whale (93) is just below the 95 cutoff.
  • These groups are weakly supported:
  • Human with Seal, Cow and Whale (71).
  • Turtle with Crocodile and Bird (70).
  • Sphenodon with the Turtle, Crocodile and Bird group (51). The placement of Sphenodon is uncertain.

Uncertain

  • I set no outgroup, so the tree is unrooted. Do not read branch order as time order or as ancestry.
  • The tree is the maximum-likelihood tree under TIM2+F+I+G4. It is not the true tree.
  • I ran one seed only. I did not test other models or other seeds.

Waits for the scientist

  • The scientist can name an outgroup if they want a rooted tree. The lungfish are an option, but that is the scientist's choice.
  • The setup asks for a maximum gap fraction. It is not needed here because I did not trim the alignment.
claim check 24 numbers in the answer, 2 numbers without a source in the record.
The review found 2 problems. The model corrects the answer before the scientist sees it.

review The review sent the answer back to the model before the scientist saw it.

- The answer names standard for "Bootstrap type", but the decision record says ufboot. Report the value that was used.
- 2 numbers in the answer match no logged tool result: 95. Run the tool that measures each number, or remove the number, or say how you derived it from logged results.
Fix these problems. Then give the answer again.

Model The model writes its answer. The claim check reads it. See the final answer.

claim check 36 numbers in the answer, 4 numbers without a source in the record.
claim check 36 numbers in the answer, 4 numbers without a source in the record.
Saved: results.xlsx (5 sheets). Folder: {work}

The final answer

The model wrote this answer. The harness shows it to the scientist after the review.

Corrections. The bootstrap type used was ufboot, which is ultrafast bootstrap with 1000 replicates. My earlier wording "not standard bootstrap values" was unclear, and this answer replaces it. I also found a mistake in my earlier support list: the value 71 belongs to a different group than I wrote. The numbers 95, 93, 71, 70 and 51 are not new measurements. The values 93, 71, 70 and 51 come from the run.treefile read. The cutoff 95 comes from the lab reporting rule, not from a tool.

Result. ModelFinder chose TIM2+F+I+G4 as the best-fit model. The maximum-likelihood tree under that model has a log-likelihood of -21152.5377.

What I did

  1. The file example.phy is already aligned, so I skipped MAFFT and did not trim it.
  2. I ran infer_tree in IQ-TREE 3.1.4 with ModelFinder (model MFP), ultrafast bootstrap (ufboot) with 1000 replicates and seed 1. The seed is fixed by the lab.

What I measured (from infer_tree)

  • The data have 17 sequences and 1998 sites. Of these, 686 sites are constant and there are 1152 site patterns.
  • The model is TIM2+F+I+G4. The run report lists it as the best model by BIC (Bayesian information criterion).
  • The log-likelihood is -21152.5377, and the tree length is 4.2173.
  • The tree files are in the infer_tree-1 folder. run.treefile is the maximum-likelihood tree with support values. run.contree is the consensus tree. run.iqtree is the report.

Branch support (ultrafast bootstrap, from run.treefile) The lab rule reads an ultrafast bootstrap value of 95 or more as well supported. Ultrafast values are not standard bootstrap values.

  • These groups have support of 95 or more:
  • Lungfish South America and Africa (100).
  • Frog with all reptiles, birds and mammals (100).
  • Lizard, Sphenodon, Turtle, Crocodile and Bird (99).
  • Mouse and Rat (100).
  • Human, Seal, Cow, Whale, Mouse and Rat (100).
  • Platypus and Opossum (96).
  • Crocodile and Bird (97).
  • Cow and Whale (99).
  • These groups have support below 95:
  • Human with Seal, Cow and Whale (93).
  • Seal with Cow and Whale (71).
  • Turtle with Crocodile and Bird (70).
  • Sphenodon with Turtle, Crocodile and Bird (51). The placement of Sphenodon is uncertain.

Uncertain

  • I set no outgroup, so the tree is unrooted. Do not read branch order as time order or as ancestry.
  • The tree is the maximum-likelihood tree under TIM2+F+I+G4. It is not the true tree.
  • I ran one seed only. I did not test other models or other seeds.

Waits for the scientist

  • The scientist can name an outgroup if they want a rooted tree. That choice belongs to the scientist.
  • The setup asks for a maximum gap fraction. It is not needed here because I did not trim the alignment.
Settings and step results that the answer does not mention. The harness adds them.
Settings used, from the decision record: Substitution model: MFP · Bootstrap type: ufboot · Number of bootstrap replicates: 1000 · Random seed: 1.

Checks

Review findings

The review recorded 8 findings. A rule finding comes from a fixed check in the harness. A referee finding comes from a second model that reads the record. The harness shows the findings to the scientist with the final answer. The record does not mark a finding as fixed. Thus a finding from an early review round can apply to a draft that the model corrected later.

Table 6 | Review findings, Sonnet run.
SeverityFromFindingShown with the final answer
errorruledecision_misreportedThe answer names standard for "Bootstrap type", but the decision record says ufboot. Report the value that was used.yes
errorruleunsourced_numbers4 numbers in the answer match no logged tool result: 95. Run the tool that measures each number, or remove the number, or say how you derived it from logged results.yes
inforuletext_styleThe answer breaks the text rules (ASD-STE100) in 4 places. Sentence 12 uses the passive voice: "is already aligned". Use the active voice. Sentence 14 uses the passive voice: "is fixed". Use the active voice. Sentence 43 uses the passive voice: "is unrooted". Use the active voice. Sentence 53 uses the passive voice: "is not needed". Use the active voice.yes
warningreferee modelThe read of run.treefile returned only a byte count. The log does not show the tree text, so it cannot confirm which group each support value belongs to. The group names (Lungfish, Sphenodon, Seal and so on) have no visible source.yes
warningreferee modelThe answer says IQ-TREE 3.1.4 and that the run report lists TIM2+F+I+G4 as best by BIC. The infer_tree result shows neither the version nor the criterion. The report must name the criterion that ModelFinder used and the version, and the log must show them.yes
warningreferee modelThe answer opens with corrections to an 'earlier wording' and an 'earlier support list'. No earlier answer is in the log. A reader cannot check these corrections, and they suggest the group assignments were already wrong once.yes
inforeferee modelThe report does not give the alignment gap fraction. It does say the file was not trimmed, so no columns were removed. The claim that the setup asks for a maximum gap fraction does not match the setup, which has trim set to false.yes
inforeferee modelThe ufboot type, 1000 replicates, seed 1, model MFP, the log-likelihood, and the unrooted status are all stated and match the log. The 95 threshold comes from the program standard and is used correctly for ultrafast bootstrap.yes

Numbers in the answer

The last claim check read 36 numbers in the answer. 32 numbers match a logged result. 4 numbers have no source in the record.

Numbers that do not match a logged result (4)
  • no source in the record: The numbers 95, 93, 71, 70 and 51 are not new measurements.
  • no source in the record: The lab rule reads an ultrafast bootstrap value of 95 or more as well supported.
  • no source in the record: - These groups have support of 95 or more:
  • no source in the record: - These groups have support below 95:

Deviations

The model did not try to change a choice of the scientist.

Failed tool calls

No tool call failed.

Data integrity

Each data file has the same SHA-256 hash now as at the time of the step that read it. Where the download script (fetch.sh) gives a hash, the file also has that hash. The run did not change the data.

Table 7 | Data files and their SHA-256 hashes, Sonnet run.
FileSHA-256Fetched dataSteps with this hash
{data}/minh2020-iqtree2/example.phy33.4 KBd3365cbea79asame as the hash in the download script (fetch.sh)n1

A SHA-256 hash is a fingerprint of the file contents. If one byte of the file changes, the hash changes. The table shows the first 12 characters.

How to repeat it

Get the data. The script downloads the files and checks their SHA-256 hashes where it lists them.

CUVETTE_DATA={data} bash bench/papers/minh2020-iqtree2/fetch.sh

Run the same case with Cuvette. The script gives the same answers from bench/papers/minh2020-iqtree2/bench.yaml.

cuvette bench papers --papers minh2020-iqtree2 --models claude:claude-sonnet-5-5

Repeat each step by hand in the program. For each step, the harness records a manual route: the menu path or the code that gives the same result. This list does not include comparison runs.

  1. infer_tree (step n1)

    Run: iqtree3 -s <alignment> -m <model> -B <replicates> -seed <seed> -T 1

    • -s

      {data}/minh2020-iqtree2/example.phy
    • -m = MFP
    • -B = 1000
    • -seed = 1
    • Warning: If you keep the default random, you get a different result.

    The manual route that the harness recorded

    /bin/sh {other volume}/tools/overnight/claude-final/catalog/phylo/scripts/iqtree.sh /opt/homebrew/bin/iqtree3 infer --alignment {data}/minh2020-iqtree2/example.phy --model MFP --bootstrap-type ufboot --replicates 1000 --seed 1

    The manual route gives the same numbers. An automatic test in Cuvette checks this.

Figure

Paper-style figure for Minh 2020, from the Sonnet run
Fig. 4 | Sonnet run. Our figure script draws the values of this run in the style of the paper.

Run facts

Table 8 | Run facts, Sonnet run.
Modelclaude-sonnet-5-5 through the Anthropic service
Date2026-10-09 12:13:30 UTC
End of runthe model gave a final answer
Time89 s
Requests to the model4
Tokensunits of text that the model read and wrote12 input, 3139 output, 32719 cache read, 13309 cache write
Cost estimate$0.07 at list price, from the token counts
Tool calls2 (0 failed)
Adaptersphylo 0.1.2, program 3.1.4
Session20261009-071330-26bb
Code hash of each step (1)
Table 9 | Code hash of each step, Sonnet run.
StepToolProgram versionCode hash
n1infer_tree3.1.426caef57428d

The code hash is a fingerprint of the adapter name, the adapter version, the tool and its definition in the adapter. If one of these changes, the hash changes.

Haiku · claude-haiku-5-5 · run 3 of 3 shown 1 of 1 values match, 1 of 1 correct in the final answer

The session

This is the session as the scientist sees it, in order. Decision cards show the answer that the script gave and where that answer comes from. Each step shows the program, its version, the input file hash and the outputs. Click a line to see more.

setup The decision record starts with these values. A script set them before the first message, from the answers that we wrote before the run.

  • MAFFT alignment method: autoSource in the tutorial or test suite: Not used. The file example.phy is already aligned. The phylo adapter uses MAFFT --auto by default.
  • Trim the alignment: falseSource in the tutorial or test suite: Not in the tutorial. The tutorial uses the alignment as it is.
  • Substitution model: MFPSource in the tutorial or test suite: Tutorial section on the substitution model. ModelFinder picks the model with the lowest Bayesian information criterion (BIC). MFP is the default from version 1.5.4.
  • Type of branch support: ufbootSource in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial notes that these values do not mean the same as standard bootstrap values.
  • Bootstrap replicates: 1000Source in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial calls 1000 the minimum number it recommends.
  • Outgroup: Source in the tutorial or test suite: Tutorial, first run. IQ-TREE draws the lungfish at the root only because it is the first sequence. The tree stays unrooted.
  • Random seed: 1
harness note A note from the harness to the model. The scientist did not write it.
- Setup values:
Alignment:
- Alignment method (MAFFT) (algorithm): auto
- Trim the alignment (trim): false
Tree:
- Substitution model (model): MFP
- Bootstrap type (bootstrap_type): ufboot
- Number of bootstrap replicates (replicates): 1000
- Outgroup (outgroup): 
- Random seed (seed): 1
Ask the scientist: Largest gap fraction of a kept column (max_gap_fraction).
Fixed by the lab: Random seed (seed) = 1.
The harness fills these values into your tool calls automatically. If you give a different value, the harness logs a deviation and shows it to the scientist.
- Mode is now watch: you run the steps.

Model thinking, summary: Since example.phy is already an aligned PHYLIP file, I can skip alignment and go straight to tree inference using MFP for model selection with ufboot bootstrapping at 1000 replicates and seed 1. Trimming isn't requested here, so I don't need to ask about gap fraction thresholds.

Model

The file is a PHYLIP alignment, so it needs no alignment step. I will run the tree tool directly with ModelFinder (MFP), ultrafast bootstrap with 1000 replicates, and seed 1.

The model calls infer_tree (adapter phylo).

step n1 infer_tree adapter phylo 0.1.2, IQ-TREE 3.1.4

Model TIM2+F+I+G4. Log-likelihood -21152.5. Ultrafast bootstrap with 1000 replicates.

Decisions applied: Substitution model = MFP; Type of branch support = ufboot; Bootstrap replicates = 1000; Random seed = 1.

Input file: {data}/minh2020-iqtree2/example.phy SHA-256 d3365cbea79a.

Outputs: run.bionj (c4103d181dee), run.ckp.gz (4cb6fbec3d28), run.contree (cf1f8eae6015), run.iqtree (a3eb7b748ce6), run.log (e7645d2dd5c3), run.mldist (7c9d04e6c668), run.model.gz (9d7ec57ddc04), run.splits.nex (650ac7edfeae), run.treefile (4cff6aaf2ea1), stderr.txt (e3b0c44298fc), stdout.txt (e3b0c44298fc).

Arguments
path{data}/minh2020-iqtree2/example.phy
modelMFP
bootstrap_typeufboot
replicates1000
seed1
Tool output
report={work}/infer_tree-1/run.iqtree
seed=1
bootstrap_type=ufboot
bootstrap_replicates=1000
treefile={work}/infer_tree-1/run.treefile
consensus_tree={work}/infer_tree-1/run.contree
n_sequences=17
n_sites=1998
model=TIM2+F+I+G4
best_model_bic=TIM2+F+I+G4
log_likelihood=-21152.5377
tree_length=4.2173
constant_sites=686
site_patterns=1152

Model thinking, summary: I'm noting the key run details for the report: BIC selected TIM2+F+I+G4, program IQ-TREE 3.1.4, log-likelihood -21152.5377, tree length 4.2173, seed 1, 1000 ultrafast bootstraps, unrooted tree with no outgroup, and an alignment of 17 already-aligned sequences in PHYLIP format.

Model

The tree run finished. I will read the consensus tree file to get the branch support values.

The model calls read_file.
Result of read_file: Read 590 of 590 bytes from {work}/infer_tree-1/run.contree.
(LngfishAu:0.1709765394,(LngfishSA:0.1884042664,LngfishAf:0.1648765322)100:0.1073673156,(Frog:0.2562938256,((((Turtle:0.2215450915,(Crocodile:0.3057507446,Bird:0.2310620493)97:0.0651606330)70:0.0364894400,Sphenodon:0.3447780527)51:0.0204536475,Lizard:0.3861439885)99:0.0739888076,(((Human:0.1851778377,(Seal:0.0944230016,(Cow:0.0823003161,Whale:0.1012908438)99:0.0404371901)71:0.0252744924)93:0.0340669342,(Mouse:0.0584259465,Rat:0.0905415196)100:0.1218030843)100:0.0607194581,(Platypus:0.1920645910,Opossum:0.1510258428)96:0.0373287871)100:0.1490890925)100:0.1275997766)100:0.0940810002);
The model calls alignment_stats (adapter phylo).
Failed of alignment_stats: Failed: sh exited with code 1: error: no sequences found

Model thinking, summary: The alignment_stats tool only handles FASTA, and this file is PHYLIP format, so it will fail here—I'll note that gap fraction wasn't measured that way. Instead, I'll write a quick script to parse the PHYLIP file directly and count gap characters to get that fraction myself.

Model

The alignment tool reads only FASTA, so it failed on this PHYLIP file. I will count the gaps with a short script.

The model runs run_script.
import re
path = "{data}/minh2020-iqtree2/example.phy"
with open(path) as f:
    lines = [l.rstrip("\n") for l in f if l.strip()]
header = lines[0].split()
n, L = int(header[0]), int(header[1])
seqs = []
for l in lines[1:]:
    parts = l.split(None, 1)
    seqs.append(parts[1] if len(parts) > 1 else "")
# interleaved format: join all residue chunks per sequence
if len(seqs) > n:
    seqs2 = [""] * n
    for i, l in enumerate(lines[1:]):
        parts = l.split(None, 1)
        if i < n:
            seqs2[i] = parts[1].replace(" ", "") if len(parts) > 1 else ""
        else:
            seqs2[i % n] += l.replace(" ", "")
    seqs = seqs2
seqs = [s.replace(" ", "") for s in seqs]
total = sum(len(s) for s in seqs)
gaps = sum(s.count("-") + s.count("?") for s in seqs)
print("sequences", n, "declared sites", L)
print("lengths", sorted(set(len(s) for s in seqs)))
print("gap or missing characters", gaps, "of", total, "fraction", round(gaps/total, 4))
The harness lets the model use run_script. In this validation, the harness lets the model use every tool and does not ask the scientist.
Result of run_script: The script ran in {work} and wrote 0 new file(s) to {work}.
sequences 17 declared sites 1998
lengths [1998]
gap or missing characters 36 of 33966 fraction 0.0011
Model Answer draft 1. The claim check and the review read it.

Result: best-fit model and maximum-likelihood tree

Input: The file example.phy is already aligned, so I did not run an alignment step. I did not trim it. The gap fraction is 0.0011: 36 gap or missing characters out of 33,966 positions. This came from a script count, because the alignment tool reads only FASTA and failed on this PHYLIP file.

Model: TIM2+F+I+G4. ModelFinder chose it by BIC (Bayesian information criterion). The tool reported only the BIC choice, not the AIC (Akaike information criterion) choice.

Tree run:

  • Program: IQ-TREE, listed as version 3.1.4 in the program header.
  • Log-likelihood of the maximum-likelihood tree: −21152.5377.
  • Tree length: 4.2173.
  • Seed: 1.
  • Bootstrap: ultrafast bootstrap (UFBoot), 1000 replicates.
  • Outgroup: none. The tree is unrooted, so branch order does not show time order.
  • Data: 17 sequences, 1998 sites, 1152 site patterns, 686 constant sites.

Branch support (UFBoot values from the consensus tree):

Well supported (95 or more):

  • Lungfish South America + Africa: 100
  • Frog + amniotes: 100
  • Mammals as one clade: 100
  • Mouse + Rat: 100
  • Platypus + Opossum: 96
  • Crocodile + Bird: 97
  • Cow + Whale: 99
  • Lizard + the other sauropsids (Turtle, Crocodile, Bird, Sphenodon): 99

Weakly supported (below 95):

  • Human + (Seal, Cow, Whale): 93
  • Seal + (Cow, Whale): 71
  • Turtle + (Crocodile, Bird): 70
  • Sphenodon placement (within the sauropsid clade): 51

Uncertain:

  • The support values come from the bootstrap consensus tree. The maximum-likelihood tree is in run.treefile.
  • The four weak nodes above can change with new data. Read the placement of Sphenodon and the Seal node as unresolved.
  • This is the maximum-likelihood tree under TIM2+F+I+G4. It is not the true tree.

Files:

  • Report: infer_tree-1/run.iqtree
  • Maximum-likelihood tree: infer_tree-1/run.treefile
  • Consensus tree with support values: infer_tree-1/run.contree

Waiting for the scientist:

  • An outgroup, if you want the tree rooted. Without one, the tree stays unrooted.
  • Nothing else blocks the analysis. Trimming is off, so the largest gap fraction setting is not needed.
claim check 28 numbers in the answer, 5 numbers without a source in the record.
The review found 2 problems. The model corrects the answer before the scientist sees it.

review The review sent the answer back to the model before the scientist saw it.

- 5 numbers in the answer match no logged tool result: 0.0011, 36, 33966, 95. Run the tool that measures each number, or remove the number, or say how you derived it from logged results.
- The branch support values come from run.contree. The read at #19 returned no file content, so the support values are not visible in the log. They must be checked against the file before use.
Fix these problems. Then give the answer again.

Model

I will read the consensus tree file again to check the support values. I will also rerun the gap count so the numbers come from a logged result.

The model calls read_file.
Result of read_file: Read 590 of 590 bytes from {work}/infer_tree-1/run.contree.
(LngfishAu:0.1709765394,(LngfishSA:0.1884042664,LngfishAf:0.1648765322)100:0.1073673156,(Frog:0.2562938256,((((Turtle:0.2215450915,(Crocodile:0.3057507446,Bird:0.2310620493)97:0.0651606330)70:0.0364894400,Sphenodon:0.3447780527)51:0.0204536475,Lizard:0.3861439885)99:0.0739888076,(((Human:0.1851778377,(Seal:0.0944230016,(Cow:0.0823003161,Whale:0.1012908438)99:0.0404371901)71:0.0252744924)93:0.0340669342,(Mouse:0.0584259465,Rat:0.0905415196)100:0.1218030843)100:0.0607194581,(Platypus:0.1920645910,Opossum:0.1510258428)96:0.0373287871)100:0.1490890925)100:0.1275997766)100:0.0940810002);
The model runs run_script.
path = "{data}/minh2020-iqtree2/example.phy"
with open(path) as f:
    lines = [l.rstrip("\n") for l in f if l.strip()]
header = lines[0].split()
n, L = int(header[0]), int(header[1])
body = lines[1:]
seqs = [""] * n
if len(body) == n:
    for i, l in enumerate(body):
        parts = l.split(None, 1)
        seqs[i] = parts[1].replace(" ", "") if len(parts) > 1 else ""
else:
    for i, l in enumerate(body):
        parts = l.split(None, 1)
        if i < n:
            seqs[i] = parts[1].replace(" ", "") if len(parts) > 1 else ""
        else:
            seqs[i % n] += l.replace(" ", "")
total = sum(len(s) for s in seqs)
gaps = sum(s.count("-") + s.count("?") for s in seqs)
print("sequences:", n, "declared sites:", L)
print("sequence lengths:", sorted(set(len(s) for s in seqs)))
print("total positions:", total)
print("gap or missing characters:", gaps)
print("gap fraction:", round(gaps / total, 4))
The harness lets the model use run_script. In this validation, the harness lets the model use every tool and does not ask the scientist.
Result of run_script: The script ran in {work} and wrote 0 new file(s) to {work}.
sequences: 17 declared sites: 1998
sequence lengths: [1998]
total positions: 33966
gap or missing characters: 36
gap fraction: 0.0011

Model thinking, summary: The file content is confirmed and the support values match what I listed, verified against the returned data. I'm noting the node with support 51 joins Turtle+Croc+Bird with Sphenodon, and I'll keep the terminology precise, avoiding "should" per the review guidance while finalizing the answer.

Model The model writes its answer. The claim check reads it. See the final answer.

claim check 28 numbers in the answer, 5 numbers without a source in the record.
claim check 28 numbers in the answer, 5 numbers without a source in the record.
Saved: results.xlsx (5 sheets). Folder: {work}

The final answer

The model wrote this answer. The harness shows it to the scientist after the review.

Result: best-fit model and maximum-likelihood tree

Input: The file example.phy is already aligned, so I did not run an alignment step. I did not trim it.

  • Sequences: 17. Sites: 1998 (infer_tree output).
  • Gap or missing characters: 36 of 33,966 total positions. Gap fraction: 0.0011 (script count, rerun this session).

Model: TIM2+F+I+G4. ModelFinder chose it by BIC (Bayesian information criterion). The tool did not report an AIC (Akaike information criterion) choice.

Tree run (infer_tree output):

  • Program: IQ-TREE, version 3.1.4 as listed in the program header.
  • Log-likelihood of the maximum-likelihood tree: −21152.5377.
  • Tree length: 4.2173.
  • Seed: 1.
  • Bootstrap: ultrafast bootstrap (UFBoot), 1000 replicates.
  • Outgroup: none. The tree is unrooted, so branch order does not show time order.
  • Site patterns: 1152. Constant sites: 686.

Branch support (UFBoot values, checked against the consensus tree file this session):

Well supported (at or above the 95 threshold in the program reporting standards):

  • Lungfish South America + Africa: 100
  • Frog + amniotes: 100
  • Mammals as one clade: 100
  • Mouse + Rat: 100
  • Cow + Whale: 99
  • Lizard + the other sauropsids (Turtle, Crocodile, Bird, Sphenodon): 99
  • Crocodile + Bird: 97
  • Platypus + Opossum: 96

Weakly supported (below 95):

  • Human + (Seal, Cow, Whale): 93
  • Seal + (Cow, Whale): 71
  • Turtle + (Crocodile, Bird): 70
  • Sphenodon placement in the sauropsid clade: 51

Uncertain:

  • The support values come from the bootstrap consensus tree (run.contree). The maximum-likelihood tree is in run.treefile.
  • The four weak nodes above can change with new data. Read the Sphenodon placement and the Seal node as unresolved.
  • This is the maximum-likelihood tree under TIM2+F+I+G4. It is not the true tree.

Files:

  • Report: infer_tree-1/run.iqtree
  • Maximum-likelihood tree: infer_tree-1/run.treefile
  • Consensus tree with support values: infer_tree-1/run.contree

Waiting for the scientist:

  • An outgroup, if you want the tree rooted. Without one, the tree stays unrooted.
  • Nothing else blocks the analysis. Trimming is off, so the largest gap fraction setting is not needed.
Settings and step results that the answer does not mention. The harness adds them.
Settings used, from the decision record: Substitution model: MFP · Bootstrap type: ufboot · Number of bootstrap replicates: 1000 · Random seed: 1.

Checks

Review findings

The review recorded 9 findings. A rule finding comes from a fixed check in the harness. A referee finding comes from a second model that reads the record. The harness shows the findings to the scientist with the final answer. The record does not mark a finding as fixed. Thus a finding from an early review round can apply to a draft that the model corrected later.

Table 10 | Review findings, Haiku run.
SeverityFromFindingShown with the final answer
warningrulefailed_result_usedStep 3 (alignment_stats) failed and was not repeated. Error: sh exited with code 1: error: no sequences foundyes
errorruleunsourced_numbers5 numbers in the answer match no logged tool result: 36, 33966, 0.0011, 95. Run the tool that measures each number, or remove the number, or say how you derived it from logged results.yes
inforuletext_styleThe answer breaks the text rules (ASD-STE100) in 3 places. Sentence 1 uses the passive voice: "is already aligned". Use the active voice. Sentence 17 uses the passive voice: "is unrooted". Use the active voice. Sentence 50 uses the passive voice: "is not needed". Use the active voice.yes
errorreferee modelThe answer says the branch support values were checked against run.contree. The logged reads of that file show only byte counts and no content. The support values cannot be verified from the log. The answer must not claim a file check that the log does not show.yes
warningreferee modelThe answer says ModelFinder chose TIM2+F+I+G4 by BIC. The infer_tree result gives only the model name. It does not name the criterion. The BIC claim has no logged source, and the claim that no AIC choice was reported is also unsupported.yes
warningreferee modelThe answer gives IQ-TREE version 3.1.4 from the program header. The logged infer_tree result does not show a version. The version must be marked as unverified or removed.yes
inforeferee modelThe alignment_stats call failed with 'no sequences found'. The answer does not mention this failure. The gap fraction came from a run_script count instead. The failure should be reported.yes
inforeferee modelThe same gap count was run twice (steps 4 and 6). The second run was made only so the numbers would be logged. The 36 gaps and 0.0011 fraction are in both results. The answer should name the count as gap or missing characters, which is what the script counted.yes
inforeferee modelThe answer does not state the number of columns removed. It says the alignment was not trimmed, so the count is zero. The answer should state this directly, as the alignment rules require.yes

Numbers in the answer

The last claim check read 28 numbers in the answer. 23 numbers match a logged result. 5 numbers have no source in the record.

Numbers that do not match a logged result (5)
  • no source in the record: - Gap or missing characters: 36 of 33,966 total positions.
  • no source in the record: - Gap or missing characters: 36 of 33,966 total positions.
  • no source in the record: Gap fraction: 0.0011 (script count, rerun this session).
  • no source in the record: Well supported (at or above the 95 threshold in the program reporting standards):
  • no source in the record: Weakly supported (below 95):

Deviations

The model did not try to change a choice of the scientist.

Failed tool calls

1 tool call failed. The model then tried again or used another tool. The session above shows each failure.

Data integrity

Each data file has the same SHA-256 hash now as at the time of the step that read it. Where the download script (fetch.sh) gives a hash, the file also has that hash. The run did not change the data.

Table 11 | Data files and their SHA-256 hashes, Haiku run.
FileSHA-256Fetched dataSteps with this hash
{data}/minh2020-iqtree2/example.phy33.4 KBd3365cbea79asame as the hash in the download script (fetch.sh)n1

A SHA-256 hash is a fingerprint of the file contents. If one byte of the file changes, the hash changes. The table shows the first 12 characters.

How to repeat it

Get the data. The script downloads the files and checks their SHA-256 hashes where it lists them.

CUVETTE_DATA={data} bash bench/papers/minh2020-iqtree2/fetch.sh

Run the same case with Cuvette. The script gives the same answers from bench/papers/minh2020-iqtree2/bench.yaml.

cuvette bench papers --papers minh2020-iqtree2 --models claude:claude-haiku-5-5

Repeat each step by hand in the program. For each step, the harness records a manual route: the menu path or the code that gives the same result. This list does not include comparison runs.

  1. infer_tree (step n1)

    Run: iqtree3 -s <alignment> -m <model> -B <replicates> -seed <seed> -T 1

    • -s

      {data}/minh2020-iqtree2/example.phy
    • -m = MFP
    • -B = 1000
    • -seed = 1
    • Warning: If you keep the default random, you get a different result.

    The manual route that the harness recorded

    /bin/sh {other volume}/tools/overnight/claude-final/catalog/phylo/scripts/iqtree.sh /opt/homebrew/bin/iqtree3 infer --alignment {data}/minh2020-iqtree2/example.phy --model MFP --bootstrap-type ufboot --replicates 1000 --seed 1

    The manual route gives the same numbers. An automatic test in Cuvette checks this.

Figure

Paper-style figure for Minh 2020, from the Haiku run
Fig. 5 | Haiku run. Our figure script draws the values of this run in the style of the paper.

Run facts

Table 12 | Run facts, Haiku run.
Modelclaude-haiku-5-5 through the Anthropic service
Date2026-10-09 12:23:39 UTC
End of runthe model gave a final answer
Time98 s
Requests to the model7
Tokensunits of text that the model read and wrote18 input, 8058 output, 80944 cache read, 19257 cache write
Cost estimate$0.01 at list price, from the token counts
Tool calls6 (1 failed)
Adaptersphylo 0.1.2, program 3.1.4
Session20261009-072339-0107
Code hash of each step (1)
Table 13 | Code hash of each step, Haiku run.
StepToolProgram versionCode hash
n1infer_tree3.1.426caef57428d

The code hash is a fingerprint of the adapter name, the adapter version, the tool and its definition in the adapter. If one of these changes, the hash changes.

qwen3:8b · local model through Ollama 1 of 1 values match, 1 of 1 correct in the final answer

The session

This is the session as the scientist sees it, in order. Decision cards show the answer that the script gave and where that answer comes from. Each step shows the program, its version, the input file hash and the outputs. Click a line to see more.

setup The decision record starts with these values. A script set them before the first message, from the answers that we wrote before the run.

  • MAFFT alignment method: autoSource in the tutorial or test suite: Not used. The file example.phy is already aligned. The phylo adapter uses MAFFT --auto by default.
  • Trim the alignment: falseSource in the tutorial or test suite: Not in the tutorial. The tutorial uses the alignment as it is.
  • Substitution model: MFPSource in the tutorial or test suite: Tutorial section on the substitution model. ModelFinder picks the model with the lowest Bayesian information criterion (BIC). MFP is the default from version 1.5.4.
  • Type of branch support: ufbootSource in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial notes that these values do not mean the same as standard bootstrap values.
  • Bootstrap replicates: 1000Source in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial calls 1000 the minimum number it recommends.
  • Outgroup: Source in the tutorial or test suite: Tutorial, first run. IQ-TREE draws the lungfish at the root only because it is the first sequence. The tree stays unrooted.
  • Random seed: 1
harness note A note from the harness to the model. The scientist did not write it.
- Setup values:
Alignment:
- Alignment method (MAFFT) (algorithm): auto
- Trim the alignment (trim): false
Tree:
- Substitution model (model): MFP
- Bootstrap type (bootstrap_type): ufboot
- Number of bootstrap replicates (replicates): 1000
- Outgroup (outgroup): 
- Random seed (seed): 1
Ask the scientist: Largest gap fraction of a kept column (max_gap_fraction).
Fixed by the lab: Random seed (seed) = 1.
The harness fills these values into your tool calls automatically. If you give a different value, the harness logs a deviation and shows it to the scientist.
- Mode is now watch: you run the steps.
The model calls align_sequences (adapter phylo).
Failed of align_sequences: Failed: sh exited with code 1: error: this file is a PHYLIP file, so the sequences are already aligned. Do not align it. Give the path to infer_tree.
The model calls infer_tree (adapter phylo).

step n1 infer_tree adapter phylo 0.1.2, IQ-TREE 3.1.4

Model TIM2+F+I+G4. Log-likelihood -21152.5. Ultrafast bootstrap with 1000 replicates.

Decisions applied: Substitution model = MFP; Type of branch support = ufboot; Bootstrap replicates = 1000; Random seed = 1.

Input file: {data}/minh2020-iqtree2/example.phy SHA-256 d3365cbea79a.

Outputs: run.bionj (c4103d181dee), run.ckp.gz (a900be8ae10e), run.contree (cf1f8eae6015), run.iqtree (adeedb463309), run.log (6f00f371f252), run.mldist (7c9d04e6c668), run.model.gz (2b7fe4474406), run.splits.nex (650ac7edfeae), run.treefile (4cff6aaf2ea1), stderr.txt (e3b0c44298fc), stdout.txt (e3b0c44298fc).

Arguments
bootstrap_typeufboot
modelMFP
path{data}/minh2020-iqtree2/example.phy
replicates1000
seed1
Tool output
report={work}/infer_tree-1/run.iqtree
seed=1
bootstrap_type=ufboot
bootstrap_replicates=1000
treefile={work}/infer_tree-1/run.treefile
consensus_tree={work}/infer_tree-1/run.contree
n_sequences=17
n_sites=1998
model=TIM2+F+I+G4
best_model_bic=TIM2+F+I+G4
log_likelihood=-21152.5377
tree_length=4.2173
constant_sites=686
site_patterns=1152
Model Answer draft 1. The claim check and the review read it.

The best-fit substitution model is TIM2+F+I+G4. The maximum-likelihood tree has ultrafast bootstrap support values. The log-likelihood is -21152.5. The treefile is at {work}/infer_tree-1/run.treefile. The tree is unrooted.

claim check 5 numbers in the answer, 3 numbers without a source in the record.
The review found 2 problems. The model corrects the answer before the scientist sees it.

review The review sent the answer back to the model before the scientist saw it.

- 3 numbers in the answer match no logged tool result: 20261009, 53857, 157. Run the tool that measures each number, or remove the number, or say how you derived it from logged results.
- The log-likelihood must be reported with more decimal places.
Fix these problems. Then give the answer again.

Model The model writes its answer. The claim check reads it. See the final answer.

claim check 5 numbers in the answer, 3 numbers without a source in the record.
claim check 5 numbers in the answer, 3 numbers without a source in the record.
Saved: results.xlsx (5 sheets). Folder: {work}

The final answer

The model wrote this answer. The harness shows it to the scientist after the review.

The best-fit substitution model is TIM2+F+I+G4. The maximum-likelihood tree has ultrafast bootstrap support values. The log-likelihood is -21152.5377. The treefile is at {work}/infer_tree-1/run.treefile. The tree is unrooted.

Settings and step results that the answer does not mention. The harness adds them.
Settings used, from the decision record: Substitution model: MFP · Bootstrap type: ufboot · Number of bootstrap replicates: 1000 · Random seed: 1.

Checks

Review findings

The review recorded 4 findings. A rule finding comes from a fixed check in the harness. A referee finding comes from a second model that reads the record. The harness shows the findings to the scientist with the final answer. The record does not mark a finding as fixed. Thus a finding from an early review round can apply to a draft that the model corrected later.

Table 14 | Review findings, qwen3:8b run.
SeverityFromFindingShown with the final answer
warningrulefailed_result_usedStep 1 (align_sequences) failed and was not repeated. Error: sh exited with code 1: error: this file is a PHYLIP file, so the sequences are already aligned. Do not align it. Give the path to infer_tree.yes
errorruleunsourced_numbers3 numbers in the answer match no logged tool result: 20261009, 53857, 157. Run the tool that measures each number, or remove the number, or say how you derived it from logged results.yes
inforuletext_styleThe answer breaks the text rules (ASD-STE100) in 1 place. Sentence 5 uses the passive voice: "is unrooted". Use the active voice.yes
errorreferee modelThe log-likelihood must be reported with the substitution model.yes

Numbers in the answer

The last claim check read 5 numbers in the answer. 2 numbers match a logged result. 3 numbers have no source in the record.

Numbers that do not match a logged result (3)
  • no source in the record: The treefile is at {work}/infer_tree-1/run.treefile.
  • no source in the record: The treefile is at {work}/infer_tree-1/run.treefile.
  • no source in the record: The treefile is at {work}/infer_tree-1/run.treefile.

Deviations

The model did not try to change a choice of the scientist.

Failed tool calls

1 tool call failed. The model then tried again or used another tool. The session above shows each failure.

Data integrity

Each data file has the same SHA-256 hash now as at the time of the step that read it. Where the download script (fetch.sh) gives a hash, the file also has that hash. The run did not change the data.

Table 15 | Data files and their SHA-256 hashes, qwen3:8b run.
FileSHA-256Fetched dataSteps with this hash
{data}/minh2020-iqtree2/example.phy33.4 KBd3365cbea79asame as the hash in the download script (fetch.sh)n1

A SHA-256 hash is a fingerprint of the file contents. If one byte of the file changes, the hash changes. The table shows the first 12 characters.

How to repeat it

Get the data. The script downloads the files and checks their SHA-256 hashes where it lists them.

CUVETTE_DATA={data} bash bench/papers/minh2020-iqtree2/fetch.sh

Run the same case with Cuvette. The script gives the same answers from bench/papers/minh2020-iqtree2/bench.yaml.

cuvette bench papers --papers minh2020-iqtree2 --models ollama:qwen3:8b

Repeat each step by hand in the program. For each step, the harness records a manual route: the menu path or the code that gives the same result. This list does not include comparison runs.

  1. infer_tree (step n1)

    Run: iqtree3 -s <alignment> -m <model> -B <replicates> -seed <seed> -T 1

    • -s

      {data}/minh2020-iqtree2/example.phy
    • -m = MFP
    • -B = 1000
    • -seed = 1
    • Warning: If you keep the default random, you get a different result.

    The manual route that the harness recorded

    /bin/sh {other volume}/tools/overnight/checkout/catalog/phylo/scripts/iqtree.sh /opt/homebrew/bin/iqtree3 infer --alignment {data}/minh2020-iqtree2/example.phy --model MFP --bootstrap-type ufboot --replicates 1000 --seed 1

    The manual route gives the same numbers. An automatic test in Cuvette checks this.

Figure

Paper-style figure for Minh 2020, from the qwen3:8b run
Fig. 6 | qwen3:8b run. Our figure script draws the values of this run in the style of the paper.

Run facts

Table 16 | Run facts, qwen3:8b run.
Modelqwen3:8b through Ollama, on our own computer
Date2026-10-09 10:38:58 UTC
End of runthe model gave a final answer
Time150 s
Requests to the model4
Tokensunits of text that the model read and wrote23657 input, 378 output, 0 cache read, 0 cache write
Cost estimatenone: the model runs on our own computer
Tool calls2 (1 failed)
Adaptersphylo 0.1.2, program 3.1.4
Session20261009-053857-157c
Code hash of each step (1)
Table 17 | Code hash of each step, qwen3:8b run.
StepToolProgram versionCode hash
n1infer_tree3.1.426caef57428d

The code hash is a fingerprint of the adapter name, the adapter version, the tool and its definition in the adapter. If one of these changes, the hash changes.