Validation / Papers / Minh 2020
Minh 2020: IQ-TREE 2, example.phy tutorial data
How to read this page
In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. The run is one sample: another run can give different steps and numbers. The model is the AI. The harness is Cuvette, the software around the model: it runs the programs and records each step. A tool call is a request from the model to run one program step. The session record is the log of each message and each step. The claim check is a script that finds each number of the final answer in the step results. The review is a set of fixed rule checks plus a second AI model, the referee, that reads the record. A deviation is a request from the model for a setting that differs from the choice of the scientist. Each Claude model did 3 runs of this paper. This page shows run 3 of each Claude model and the one run of qwen3:8b. The table of values says how many of the Claude runs match.
Opus: 1 of 1 values match, 1 of 1 correct in the final answer. All 3 runs: 1 of 1 values match. Sonnet: 1 of 1 values match, 1 of 1 correct in the final answer. All 3 runs: 1 of 1 values match. Haiku: 1 of 1 values match, 1 of 1 correct in the final answer. All 3 runs: 1 of 1 values match. qwen3:8b: 1 of 1 values match, 1 of 1 correct in the final answer.
The figure in the paper and in the run
As published
The IQ-TREE 2 paper (Minh et al. 2020) has no figure or number for example.phy. The beginner tutorial of IQ-TREE uses this file and names the model TIM2+I+G4. The tutorial prints no log-likelihood. We do not copy the tutorial text or the paper figures because their license is not CC BY.
Reproduced in Cuvette
The paper
Minh BQ, Schmidt HA, Chernomor O, Schrempf D, Woodhams MD, von Haeseler A, Lanfear R. IQ-TREE 2: New models and efficient methods for phylogenetic inference in the genomic era. Molecular Biology and Evolution 37(5):1530-1534 (2020). doi:10.1093/molbev/msaa015
Related sources:
- IQ-TREE beginner tutorial. It uses the file example.phy and names the best-fit model for it. link
What it measured
The IQ-TREE 2 paper describes a program that builds maximum-likelihood (ML) trees from sequence alignments. The paper has no numbers for the example file. The beginner tutorial uses example.phy, an alignment of mitochondrial DNA from 17 animals. The tutorial lets ModelFinder choose the substitution model. It then builds the tree and gives branch support with the ultrafast bootstrap.
Data
File example.phy from the example folder of the IQ-TREE 2 source repository. Size: 34 KB, 17 sequences by 1998 sites.
License: The repository is GPL-2.0. The file holds published mitochondrial DNA sequences of animals and has no personal data.
The instruction
A script sent this message as the scientist. The file paths point to the fetched data.
The same request in the words of the paper's method:
I have an alignment of mitochondrial DNA from 17 animals. Find the best-fit substitution model and build the maximum-likelihood tree with branch support. Let the data choose the model. Report the log-likelihood of the tree.
Basis: The tutorial sections on the first run, on the choice of the substitution model, and on branch support with the ultrafast bootstrap.
Results
Match: a number in the session record is inside the tolerance of the known value. In the final answer: the model also stated the value in its final answer. For a Claude model, each cell shows the run that this page shows. If the three runs differ, the cell also says in how many runs the value matches.
| Value | Known value | Tolerance | Opus | Sonnet | Haiku | qwen3:8b |
|---|---|---|---|---|---|---|
log_likelihoodLog-likelihood of the ML treeSource of the known valueWe calculated it with IQ-TREE 3.1.4 (iqtree3), seed 1Not in the paper or the tutorial. The tutorial names the model TIM2+I+G4 but prints no log-likelihood. We ran -m MFP -B 1000 and got -21152.5377. | -21152.54 | ± 1 | -21152.54 matchIn the final answer: yes (-21152.54)Log: n1 infer_tree metrics.log_likelihood, entry 11; the final answer, entry 42 | -21152.54 matchIn the final answer: yes (-21152.54)Log: n1 infer_tree metrics.log_likelihood, entry 11; the final answer, entry 38 | -21152.54 matchIn the final answer: yes (-21152.54)Log: n1 infer_tree metrics.log_likelihood, entry 11; the final answer, entry 64 | -21152.54 matchIn the final answer: yes (-21152.54)Log: n1 infer_tree metrics.log_likelihood, entry 14; the final answer, entry 29 |
Session records
Session record, Opus, run 3 of 3
Every message, decision, step and result of this run, one JSON object for each log entry.
Session record, Sonnet, run 3 of 3
Every message, decision, step and result of this run, one JSON object for each log entry.
Session record, Haiku, run 3 of 3
Every message, decision, step and result of this run, one JSON object for each log entry.
Session record, qwen3:8b
Every message, decision, step and result of this run, one JSON object for each log entry.
Opus · claude-opus-5-5 · run 3 of 3 shown 1 of 1 values match, 1 of 1 correct in the final answer
The session
This is the session as the scientist sees it, in order. Decision cards show the answer that the script gave and where that answer comes from. Each step shows the program, its version, the input file hash and the outputs. Click a line to see more.
setup The decision record starts with these values. A script set them before the first message, from the answers that we wrote before the run.
- MAFFT alignment method: autoSource in the tutorial or test suite: Not used. The file example.phy is already aligned. The phylo adapter uses MAFFT --auto by default.
- Trim the alignment: falseSource in the tutorial or test suite: Not in the tutorial. The tutorial uses the alignment as it is.
- Substitution model: MFPSource in the tutorial or test suite: Tutorial section on the substitution model. ModelFinder picks the model with the lowest Bayesian information criterion (BIC). MFP is the default from version 1.5.4.
- Type of branch support: ufbootSource in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial notes that these values do not mean the same as standard bootstrap values.
- Bootstrap replicates: 1000Source in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial calls 1000 the minimum number it recommends.
- Outgroup: Source in the tutorial or test suite: Tutorial, first run. IQ-TREE draws the lungfish at the root only because it is the first sequence. The tree stays unrooted.
- Random seed: 1
harness note A note from the harness to the model. The scientist did not write it.
- Setup values: Alignment: - Alignment method (MAFFT) (algorithm): auto - Trim the alignment (trim): false Tree: - Substitution model (model): MFP - Bootstrap type (bootstrap_type): ufboot - Number of bootstrap replicates (replicates): 1000 - Outgroup (outgroup): - Random seed (seed): 1 Ask the scientist: Largest gap fraction of a kept column (max_gap_fraction). Fixed by the lab: Random seed (seed) = 1. The harness fills these values into your tool calls automatically. If you give a different value, the harness logs a deviation and shows it to the scientist. - Mode is now watch: you run the steps.
Model
The file is a PHYLIP alignment, so it needs no alignment step. Next, I build the maximum-likelihood tree. ModelFinder chooses the model, and 1000 ultrafast bootstrap replicates give the branch support.
infer_tree (adapter phylo).step n1 infer_tree adapter phylo 0.1.2, IQ-TREE 3.1.4
Model TIM2+F+I+G4. Log-likelihood -21152.5. Ultrafast bootstrap with 1000 replicates.
Decisions applied: Substitution model = MFP; Type of branch support = ufboot; Bootstrap replicates = 1000; Random seed = 1.
Input file: {data}/minh2020-iqtree2/example.phy SHA-256 d3365cbea79a.
Outputs: run.bionj (c4103d181dee), run.ckp.gz (2c8f49979d62), run.contree (cf1f8eae6015), run.iqtree (837ec22b9ee3), run.log (99b88cd1d7b0), run.mldist (7c9d04e6c668), run.model.gz (959c833b94dc), run.splits.nex (650ac7edfeae), run.treefile (4cff6aaf2ea1), stderr.txt (e3b0c44298fc), stdout.txt (e3b0c44298fc).
Arguments
| path | {data}/minh2020-iqtree2/example.phy |
| model | MFP |
| bootstrap_type | ufboot |
| replicates | 1000 |
| seed | 1 |
Tool output
report={work}/infer_tree-1/run.iqtree
seed=1
bootstrap_type=ufboot
bootstrap_replicates=1000
treefile={work}/infer_tree-1/run.treefile
consensus_tree={work}/infer_tree-1/run.contree
n_sequences=17
n_sites=1998
model=TIM2+F+I+G4
best_model_bic=TIM2+F+I+G4
log_likelihood=-21152.5377
tree_length=4.2173
constant_sites=686
site_patterns=1152
Model
Next, I read the tree file to get the topology and the support values.
read_file.Result of read_file: Read 590 of 590 bytes from {work}/infer_tree-1/run.treefile.
(LngfishAu:0.1709068277,(LngfishSA:0.1883311672,LngfishAf:0.1648032873)100:0.1073256416,(Frog:0.2561720790,((((Turtle:0.2214955116,(Crocodile:0.3056436573,Bird:0.2309752831)97:0.0651735079)70:0.0364941210,Sphenodon:0.3446303789)51:0.0204677635,Lizard:0.3859794405)99:0.0739928283,(((Human:0.1850870817,(Seal:0.0943875559,(Cow:0.0822572767,Whale:0.1012437124)99:0.0404174962)71:0.0252594799)93:0.0340514112,(Mouse:0.0583901122,Rat:0.0905005039)100:0.1217432345)100:0.0607172631,(Platypus:0.1919645817,Opossum:0.1509515712)96:0.0373002662)100:0.1490260063)100:0.1275594673)100:0.0940391076);
read_file.Result of read_file: Read 12000 of 19016 bytes from {work}/infer_tree-1/run.iqtree.
IQ-TREE 3.1.4 built Sep 10 2026
Input file name: {data}/minh2020-iqtree2/example.phy
Type of analysis: ModelFinder + tree reconstruction + ultrafast bootstrap (1000 replicates)
Random seed number: 1
REFERENCES
----------
To cite IQ-TREE 3 please use:
Thomas K.F. Wong, Nhan Ly-Trong, Huaiyan Ren, Piyumal Demotte, Hector Banos,
Andrew J. Roger, Edward Susko, Chris Bielow, Nicola De Maio, Nick Goldman,
Matthew W. Hahn, Mario dos Reis, Le Sy Vinh, Gavin Huttley, Robert Lanfear, Bui Quang Minh (2026)
IQ-TREE 3: Phylogenomic Inference Software using Complex Evolutionary Models.
Molecular Biology and Evolution, msag117.
https://doi.org/10.1093/molbev/msag117
Since you used ModelFinder please also cite:
Subha Kalyaanamoorthy, Bui Quang Minh, Thomas KF Wong, Arndt von Haeseler,
and Lars S Jermiin (2017) ModelFinder: Fast model selection for
accurate phylogenetic estimates. Nature Methods, 14:587–589.
https://doi.org/10.1038/nmeth.4285
Since you used ultrafast bootstrap (UFBoot) please also cite:
Diep Thi Hoang, Olga Chernomor, Arndt von Haeseler, Bui Quang Minh,
and Le Sy Vinh (2018) UFBoot2: Improving the ultrafast bootstrap
approximation. Mol. Biol. Evol., 35:518–522.
https://doi.org/10.1093/molbev/msx281
SEQUENCE ALIGNMENT
------------------
NOTE: Alignment sequence type is auto-detected. If in doubt, specify it via -st option.
Input data: 17 sequences with 1998 nucleotide sites
Number of constant sites: 686 (= 34.3343% of all sites)
Number of invariant (constant or ambiguous constant) sites: 686 (= 34.3343% of all sites)
Number of parsimony informative sites: 1009
Number of distinct site patterns: 1152
ModelFinder
-----------
Best-fit model according to BIC: TIM2+F+I+G4
List of models sorted by BIC scores:
Model LogL AIC w-AIC AICc w-AICc BIC w-BIC
TIM2+F+I+G4 -21152.739 42383.479 + 0.0741 42385.072 + 0.0828 42601.875 + 0.974
GTR+F+I+G4 -21148.957 42379.913 + 0.441 42381.674 + 0.453 42609.509 - 0.0214
TIM2+F+R3 -21150.544 42383.088 + 0.0901 42384.849 + 0.0926 42612.684 - 0.00438
GTR+F+R3 -21147.067 42380.134 + 0.395 42382.071 + 0.372 42620.930 - 7.09e-05
TPM2u+F+R3 -21283.140 42646.279 - 6.36e-59 42647.955 - 6.82e-59 42870.275 - 5.08e-59
TPM2u+F+I+G4 -21291.534 42659.067 - 1.06e-61 42660.580 - 1.24e-61 42871.863 - 2.3e-59
TVM+F+R3 -21279.783 42643.567 - 2.47e-58 42645.414 - 2.43e-58 42878.762 - 7.3e-61
TVM+F+I+G4 -21287.937 42655.873 - 5.25e-61 42657.549 - 5.63e-61 42879.870 - 4.2e-61
TIM2+I+G4 -21308.990 42689.980 - 2.06e-68 42691.338 - 2.59e-68 42891.576 - 1.2e-63
TIM2e+I+G4 -21309.009 42690.018 - 2.02e-68 42691.376 - 2.54e-68 42891.614 - 1.18e-63
TIM2e+I+G4 -21309.009 42690.018 - 2.02e-68 42691.376 - 2.54e-68 42891.614 - 1.18e-63
TIM2+R3 -21303.599 42683.199 - 6.12e-67 42684.712 - 7.12e-67 42895.995 -
... (974 more characters in the session record)Model The model writes its answer. The claim check reads it. See the final answer.
review The review sent the answer back to the model before the scientist saw it.
- 3 numbers in the answer match no logged tool result: 95. Run the tool that measures each number, or remove the number, or say how you derived it from logged results. Fix these problems. Then give the answer again.
Model Answer draft 2. The claim check and the review read it.
I ran one IQ-TREE 3.1.4 job. ModelFinder chose TIM2+F+I+G4 by BIC (Bayesian information criterion). The maximum-likelihood tree has a log-likelihood of −21152.5377. Nine branches are well supported, and four branches are below the lab threshold.
Source of the threshold: the lab reporting standards say that an ultrafast bootstrap value of 95 or more is well supported. This value is a rule, not a measurement. In this answer, "the UFBoot threshold" means that rule.
What I did
- The input file is a PHYLIP alignment that is already aligned. Thus I did not run MAFFT and did not trim. The trimming gap fraction does not apply.
- I ran the tree tool once. It did the model selection, the tree search and the support in the same run.
- Model: ModelFinder (
--model MFP). - Support: ultrafast bootstrap (UFBoot), 1000 replicates.
- Random seed: 1, fixed by the lab.
- Command:
iqtree3 infer --alignment example.phy --model MFP --bootstrap-type ufboot --replicates 1000 --seed 1
Data (from the IQ-TREE report)
- 17 sequences and 1998 nucleotide sites.
- 686 constant sites, 1009 parsimony-informative sites and 1152 distinct site patterns.
Best-fit model (from the ModelFinder table in run.iqtree)
- BIC chose TIM2+F+I+G4, with a BIC weight of 0.974.
- The next model by BIC is GTR+F+I+G4, with a BIC weight of 0.0214.
- AIC (Akaike information criterion) gives a different order:
- GTR+F+I+G4 is first, with an AIC weight of 0.441.
- TIM2+F+I+G4 has an AIC weight of 0.0741.
- The report marks four models with a plus sign in the AIC column: TIM2+F+I+G4, GTR+F+I+G4, TIM2+F+R3 and GTR+F+R3. The report says that a plus sign means the model is in the confidence set.
- TIM2 rates: A-C = A-T = 5.5843, A-G = 7.6324, C-T = 22.6443, C-G = G-T = 1.0.
- Base frequencies, counted from the alignment: A 0.3547, C 0.2282, G 0.1919, T 0.2252.
Tree
- Log-likelihood: −21152.5377. Total tree length: 4.2173 substitutions per site.
- Tree with support values:
infer_tree-1/run.treefile. The bootstrap consensus tree isrun.contree. The full report isrun.iqtree. - The tree is unrooted because no outgroup was set. The Newick file starts at the lungfish, but that is only how the file is written. It does not show a root or a time order.
Branch support (UFBoot values from run.treefile)
| Group | UFBoot |
|---|---|
| Lungfish South America + Lungfish Africa | 100 |
| Amniotes (reptiles, birds and mammals) | 100 |
| Mammals | 100 |
| Lizard + archosaurs + turtle + Sphenodon | 99 |
| Placental mammals | 100 |
| Mouse + Rat | 100 |
| Cow + Whale | 99 |
| Crocodile + Bird | 97 |
| Platypus + Opossum | 96 |
| Human + (Seal, Cow, Whale) | 93 |
| Seal + (Cow, Whale) | 71 |
| Turtle + (Crocodile, Bird) | 70 |
| Sphenodon + (Turtle, Crocodile, Bird) | 51 |
What is uncertain
- Four groups have values below the UFBoot threshold, so they are not well supported:
- Sphenodon + (Turtle, Crocodile, Bird): 51.
- Turtle + (Crocodile, Bird): 70.
- Seal + (Cow, Whale): 71.
- Human + (Seal, Cow, Whale): 93.
- This is the maximum-likelihood tree under TIM2+F+I+G4. It is not the true tree.
- The model choice depends on the criterion: BIC prefers TIM2+F+I+G4, and AIC prefers GTR+F+I+G4. I did not test if GTR+F+I+G4 gives the same tree.
- I did not read the gamma shape and the proportion of invariable sites. They are later in
run.iqtree, after the part that I read.
What waits for you
- Outgroup: if you want a rooted tree, name the outgroup. The three lungfish are one possible choice for this data set.
- Model: tell me if you want a check run with GTR+F+I+G4, the model that AIC prefers.
The final answer
The model wrote this answer. The harness shows it to the scientist after the review.
I ran one IQ-TREE 3.1.4 job. ModelFinder chose TIM2+F+I+G4 by BIC (Bayesian information criterion). The maximum-likelihood tree has a log-likelihood of −21152.5377. Most branches have ultrafast bootstrap support of 96 to 100, but four branches have less than 95.
What I did
- The input file is a PHYLIP alignment that is already aligned. Thus I did not run MAFFT and did not trim. The trimming gap fraction does not apply.
- I ran the tree tool once. It did the model selection, the tree search and the support in the same run.
- Model: ModelFinder (
--model MFP). - Support: ultrafast bootstrap (UFBoot), 1000 replicates.
- Random seed: 1, fixed by the lab.
- Command:
iqtree3 infer --alignment example.phy --model MFP --bootstrap-type ufboot --replicates 1000 --seed 1
Data (from the IQ-TREE report)
- 17 sequences and 1998 nucleotide sites.
- 686 constant sites, 1009 parsimony-informative sites and 1152 distinct site patterns.
Best-fit model
- BIC chose TIM2+F+I+G4, with a BIC weight of 0.974.
- The next model by BIC is GTR+F+I+G4, with a BIC weight of 0.0214.
- By AIC (Akaike information criterion), GTR+F+I+G4 is first, with an AIC weight of 0.441. TIM2+F+I+G4 has an AIC weight of 0.0741. All four top models are in the AIC 95% confidence set.
- TIM2 rates: A-C = A-T = 5.5843, A-G = 7.6324, C-T = 22.6443, C-G = G-T = 1.0.
- Base frequencies, counted from the alignment: A 0.3547, C 0.2282, G 0.1919, T 0.2252.
Tree
- Log-likelihood: −21152.5377. Total tree length: 4.2173 substitutions per site.
- Tree with support values:
infer_tree-1/run.treefile. The bootstrap consensus tree isrun.contree. The full report isrun.iqtree. - The tree is unrooted because no outgroup was set. The Newick file starts at the lungfish, but that is only how the file is written. It does not show a root or a time order.
Branch support (UFBoot: 95 or more is well supported)
| Group | UFBoot |
|---|---|
| Lungfish South America + Lungfish Africa | 100 |
| Amniotes (reptiles, birds and mammals) | 100 |
| Mammals | 100 |
| Lizard + archosaurs + turtle + Sphenodon | 99 |
| Placental mammals | 100 |
| Mouse + Rat | 100 |
| Cow + Whale | 99 |
| Crocodile + Bird | 97 |
| Platypus + Opossum | 96 |
| Human + (Seal, Cow, Whale) | 93 |
| Seal + (Cow, Whale) | 71 |
| Turtle + (Crocodile, Bird) | 70 |
| Sphenodon + (Turtle, Crocodile, Bird) | 51 |
What is uncertain
- These four groups have UFBoot less than 95, so they are not well supported:
- Sphenodon + (Turtle, Crocodile, Bird): 51.
- Turtle + (Crocodile, Bird): 70.
- Seal + (Cow, Whale): 71.
- Human + (Seal, Cow, Whale): 93.
- This is the maximum-likelihood tree under TIM2+F+I+G4. It is not the true tree.
- The model choice depends on the criterion. BIC prefers TIM2+F+I+G4, but AIC prefers GTR+F+I+G4. These two models differ only by two rate parameters, so the tree can be the same with either model. I did not test this.
- I did not read the gamma shape and the proportion of invariable sites. They are later in
run.iqtree, after the part that I read.
What waits for you
- Outgroup: if you want a rooted tree, name the outgroup. The three lungfish are one possible choice for this data set.
- Model: tell me if you want a check run with GTR+F+I+G4, the model that AIC prefers.
Settings and step results that the answer does not mention. The harness adds them.
Settings used, from the decision record: Substitution model: MFP · Bootstrap type: ufboot · Number of bootstrap replicates: 1000 · Random seed: 1.
Checks
Review findings
The review recorded 8 findings. A rule finding comes from a fixed check in the harness. A referee finding comes from a second model that reads the record. The harness shows the findings to the scientist with the final answer. The record does not mark a finding as fixed. Thus a finding from an early review round can apply to a draft that the model corrected later.
| Severity | From | Finding | Shown with the final answer |
|---|---|---|---|
| error | ruleunsourced_numbers | 3 numbers in the answer match no logged tool result: 95. Run the tool that measures each number, or remove the number, or say how you derived it from logged results. | yes |
| info | ruletext_style | The answer breaks the text rules (ASD-STE100) in 3 places. Sentence 6 uses the passive voice: "is already aligned". Use the active voice. Sentence 32 uses the passive voice: "is unrooted". Use the active voice. Sentence 33 uses the passive voice: "is written". Use the active voice. | yes |
| info | referee model | The report names the model TIM2+F+I+G4. It says that ModelFinder chose it by BIC, and it names UFBoot with 1000 replicates. It uses the 95 threshold, which is correct for the ultrafast bootstrap. These items agree with the setup and the tool result. | yes |
| warning | referee model | The command line in the answer is a reconstruction. The log shows only a truncated manual route, so the log does not show the flags after the alignment path. The report must quote the logged command or mark this line as a paraphrase. | yes |
| warning | referee model | The tree has 17 sequences and is unrooted, so it has 14 internal branches. The support table lists only 13 groups. One branch is missing, so the statement that four branches have less than 95 is not fully supported. The report must list all internal branches or say which branch it left out. | yes |
| info | referee model | The step read only 12000 of 19016 bytes of run.iqtree. The report gives the IQ-TREE version, the BIC and AIC weights and the rate parameters from this partial read, and it says that it did not read the gamma shape or the invariable-site proportion. The report states this limit correctly. | yes |
| info | referee model | The report says that the alignment was not trimmed and was not realigned because the input was PHYLIP. This satisfies the trim check. The report does not give the gap fraction of the input alignment, which the standards ask for. | yes |
| info | referee model | The answer says that the lab fixed the seed. The setup shows that the seed was an accepted default, so the harness fixed it. The value of 1 is correct. | yes |
Numbers in the answer
The last claim check read 47 numbers in the answer. 44 numbers match a logged result. 3 numbers have no source in the record.
Numbers that do not match a logged result (3)
- no source in the record: Most branches have ultrafast bootstrap support of 96 to 100, but four branches have less than 95.
- no source in the record: **Branch support (UFBoot: 95 or more is well supported)**
- no source in the record: - These four groups have UFBoot less than 95, so they are not well supported:
Deviations
The model did not try to change a choice of the scientist.
Failed tool calls
No tool call failed.
Data integrity
Each data file has the same SHA-256 hash now as at the time of the step that read it. Where the download script (fetch.sh) gives a hash, the file also has that hash. The run did not change the data.
| File | SHA-256 | Fetched data | Steps with this hash |
|---|---|---|---|
{data}/minh2020-iqtree2/example.phy33.4 KB | d3365cbea79a | same as the hash in the download script (fetch.sh) | n1 |
A SHA-256 hash is a fingerprint of the file contents. If one byte of the file changes, the hash changes. The table shows the first 12 characters.
How to repeat it
Get the data. The script downloads the files and checks their SHA-256 hashes where it lists them.
CUVETTE_DATA={data} bash bench/papers/minh2020-iqtree2/fetch.shRun the same case with Cuvette. The script gives the same answers from bench/papers/minh2020-iqtree2/bench.yaml.
cuvette bench papers --papers minh2020-iqtree2 --models claude:claude-opus-5-5
Repeat each step by hand in the program. For each step, the harness records a manual route: the menu path or the code that gives the same result. This list does not include comparison runs.
infer_tree(step n1)Run: iqtree3 -s <alignment> -m <model> -B <replicates> -seed <seed> -T 1
-s
{data}/minh2020-iqtree2/example.phy- -m =
MFP - -B =
1000 - -seed =
1 - Warning: If you keep the default random, you get a different result.
The manual route that the harness recorded
/bin/sh {other volume}/tools/overnight/claude-final/catalog/phylo/scripts/iqtree.sh /opt/homebrew/bin/iqtree3 infer --alignment {data}/minh2020-iqtree2/example.phy --model MFP --bootstrap-type ufboot --replicates 1000 --seed 1The manual route gives the same numbers. An automatic test in Cuvette checks this.
Figure

Run facts
| Model | claude-opus-5-5 through the Anthropic service |
| Date | 2026-10-09 12:37:04 UTC |
| End of run | the model gave a final answer |
| Time | 90 s |
| Requests to the model | 4 |
| Tokensunits of text that the model read and wrote | 12 input, 4474 output, 39512 cache read, 20884 cache write |
| Cost estimate | $0.20 at list price, from the token counts |
| Tool calls | 3 (0 failed) |
| Adapters | phylo 0.1.2, program 3.1.4 |
| Session | 20261009-073704-eb13 |
Code hash of each step (1)
| Step | Tool | Program version | Code hash |
|---|---|---|---|
| n1 | infer_tree | 3.1.4 | 26caef57428d |
The code hash is a fingerprint of the adapter name, the adapter version, the tool and its definition in the adapter. If one of these changes, the hash changes.
Sonnet · claude-sonnet-5-5 · run 3 of 3 shown 1 of 1 values match, 1 of 1 correct in the final answer
The session
This is the session as the scientist sees it, in order. Decision cards show the answer that the script gave and where that answer comes from. Each step shows the program, its version, the input file hash and the outputs. Click a line to see more.
setup The decision record starts with these values. A script set them before the first message, from the answers that we wrote before the run.
- MAFFT alignment method: autoSource in the tutorial or test suite: Not used. The file example.phy is already aligned. The phylo adapter uses MAFFT --auto by default.
- Trim the alignment: falseSource in the tutorial or test suite: Not in the tutorial. The tutorial uses the alignment as it is.
- Substitution model: MFPSource in the tutorial or test suite: Tutorial section on the substitution model. ModelFinder picks the model with the lowest Bayesian information criterion (BIC). MFP is the default from version 1.5.4.
- Type of branch support: ufbootSource in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial notes that these values do not mean the same as standard bootstrap values.
- Bootstrap replicates: 1000Source in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial calls 1000 the minimum number it recommends.
- Outgroup: Source in the tutorial or test suite: Tutorial, first run. IQ-TREE draws the lungfish at the root only because it is the first sequence. The tree stays unrooted.
- Random seed: 1
harness note A note from the harness to the model. The scientist did not write it.
- Setup values: Alignment: - Alignment method (MAFFT) (algorithm): auto - Trim the alignment (trim): false Tree: - Substitution model (model): MFP - Bootstrap type (bootstrap_type): ufboot - Number of bootstrap replicates (replicates): 1000 - Outgroup (outgroup): - Random seed (seed): 1 Ask the scientist: Largest gap fraction of a kept column (max_gap_fraction). Fixed by the lab: Random seed (seed) = 1. The harness fills these values into your tool calls automatically. If you give a different value, the harness logs a deviation and shows it to the scientist. - Mode is now watch: you run the steps.
Model thinking, summary: This file is already a PHYLIP-aligned alignment, so I can skip straight to tree inference without an outgroup, keeping it unrooted and without trimming. I'll run infer_tree with the MFP model and ufboot set to 1000 replicates.
Model
The file is already aligned (PHYLIP), so I skip MAFFT. I now run IQ-TREE with ModelFinder and 1000 ultrafast bootstrap replicates.
infer_tree (adapter phylo).step n1 infer_tree adapter phylo 0.1.2, IQ-TREE 3.1.4
Model TIM2+F+I+G4. Log-likelihood -21152.5. Ultrafast bootstrap with 1000 replicates.
Decisions applied: Substitution model = MFP; Type of branch support = ufboot; Bootstrap replicates = 1000; Random seed = 1.
Input file: {data}/minh2020-iqtree2/example.phy SHA-256 d3365cbea79a.
Outputs: run.bionj (c4103d181dee), run.ckp.gz (a8405163edbe), run.contree (cf1f8eae6015), run.iqtree (e559067d0d42), run.log (024d2aeaf644), run.mldist (7c9d04e6c668), run.model.gz (09fc415529a4), run.splits.nex (650ac7edfeae), run.treefile (4cff6aaf2ea1), stderr.txt (e3b0c44298fc), stdout.txt (e3b0c44298fc).
Arguments
| path | {data}/minh2020-iqtree2/example.phy |
| model | MFP |
| bootstrap_type | ufboot |
| replicates | 1000 |
| seed | 1 |
Tool output
report={work}/infer_tree-1/run.iqtree
seed=1
bootstrap_type=ufboot
bootstrap_replicates=1000
treefile={work}/infer_tree-1/run.treefile
consensus_tree={work}/infer_tree-1/run.contree
n_sequences=17
n_sites=1998
model=TIM2+F+I+G4
best_model_bic=TIM2+F+I+G4
log_likelihood=-21152.5377
tree_length=4.2173
constant_sites=686
site_patterns=1152
Model thinking, summary: I should check the support values, so let me quickly read the treefile to report them.
Model
I now read the tree file to check the branch support values.
read_file.Result of read_file: Read 590 of 590 bytes from {work}/infer_tree-1/run.treefile.
(LngfishAu:0.1709068277,(LngfishSA:0.1883311672,LngfishAf:0.1648032873)100:0.1073256416,(Frog:0.2561720790,((((Turtle:0.2214955116,(Crocodile:0.3056436573,Bird:0.2309752831)97:0.0651735079)70:0.0364941210,Sphenodon:0.3446303789)51:0.0204677635,Lizard:0.3859794405)99:0.0739928283,(((Human:0.1850870817,(Seal:0.0943875559,(Cow:0.0822572767,Whale:0.1012437124)99:0.0404174962)71:0.0252594799)93:0.0340514112,(Mouse:0.0583901122,Rat:0.0905005039)100:0.1217432345)100:0.0607172631,(Platypus:0.1919645817,Opossum:0.1509515712)96:0.0373002662)100:0.1490260063)100:0.1275594673)100:0.0940391076);
Model Answer draft 1. The claim check and the review read it.
Result. ModelFinder chose TIM2+F+I+G4 as the best-fit model. The maximum-likelihood tree under that model has a log-likelihood of -21152.5377.
What I did
- The file
example.phyis already aligned, so I skipped MAFFT and did not trim it. - I ran
infer_treein IQ-TREE 3.1.4 with ModelFinder (model MFP), 1000 ultrafast bootstrap replicates and seed 1. The seed is fixed by the lab.
What I measured (from infer_tree)
- The data have 17 sequences and 1998 sites. Of these, 686 sites are constant and there are 1152 site patterns.
- The model is TIM2+F+I+G4. The run report lists it as the best model by BIC (Bayesian information criterion).
- The log-likelihood is -21152.5377, and the tree length is 4.2173.
- Tree files are in the
infer_tree-1folder.run.treefileis the maximum-likelihood tree with support values.run.contreeis the consensus tree.run.iqtreeis the report.
Branch support (ultrafast bootstrap, from run.treefile) These are not standard bootstrap values. A value of 95 or more is well supported.
- These groups are well supported:
- Lungfish South America and Africa (100).
- The clade of Frog, reptiles, birds and mammals (100).
- Mouse and Rat (100).
- Human, Seal, Cow, Whale, Mouse and Rat together (100).
- Platypus and Opossum (96).
- Crocodile and Bird (97).
- The clade of Lizard, Sphenodon, Turtle, Crocodile and Bird (99).
- Cow and Whale (99).
- Human, Seal, Cow and Whale (93) is just below the 95 cutoff.
- These groups are weakly supported:
- Human with Seal, Cow and Whale (71).
- Turtle with Crocodile and Bird (70).
- Sphenodon with the Turtle, Crocodile and Bird group (51). The placement of Sphenodon is uncertain.
Uncertain
- I set no outgroup, so the tree is unrooted. Do not read branch order as time order or as ancestry.
- The tree is the maximum-likelihood tree under TIM2+F+I+G4. It is not the true tree.
- I ran one seed only. I did not test other models or other seeds.
Waits for the scientist
- The scientist can name an outgroup if they want a rooted tree. The lungfish are an option, but that is the scientist's choice.
- The setup asks for a maximum gap fraction. It is not needed here because I did not trim the alignment.
review The review sent the answer back to the model before the scientist saw it.
- The answer names standard for "Bootstrap type", but the decision record says ufboot. Report the value that was used. - 2 numbers in the answer match no logged tool result: 95. Run the tool that measures each number, or remove the number, or say how you derived it from logged results. Fix these problems. Then give the answer again.
Model The model writes its answer. The claim check reads it. See the final answer.
The final answer
The model wrote this answer. The harness shows it to the scientist after the review.
Corrections. The bootstrap type used was ufboot, which is ultrafast bootstrap with 1000 replicates. My earlier wording "not standard bootstrap values" was unclear, and this answer replaces it. I also found a mistake in my earlier support list: the value 71 belongs to a different group than I wrote. The numbers 95, 93, 71, 70 and 51 are not new measurements. The values 93, 71, 70 and 51 come from the run.treefile read. The cutoff 95 comes from the lab reporting rule, not from a tool.
Result. ModelFinder chose TIM2+F+I+G4 as the best-fit model. The maximum-likelihood tree under that model has a log-likelihood of -21152.5377.
What I did
- The file
example.phyis already aligned, so I skipped MAFFT and did not trim it. - I ran
infer_treein IQ-TREE 3.1.4 with ModelFinder (model MFP), ultrafast bootstrap (ufboot) with 1000 replicates and seed 1. The seed is fixed by the lab.
What I measured (from infer_tree)
- The data have 17 sequences and 1998 sites. Of these, 686 sites are constant and there are 1152 site patterns.
- The model is TIM2+F+I+G4. The run report lists it as the best model by BIC (Bayesian information criterion).
- The log-likelihood is -21152.5377, and the tree length is 4.2173.
- The tree files are in the
infer_tree-1folder.run.treefileis the maximum-likelihood tree with support values.run.contreeis the consensus tree.run.iqtreeis the report.
Branch support (ultrafast bootstrap, from run.treefile) The lab rule reads an ultrafast bootstrap value of 95 or more as well supported. Ultrafast values are not standard bootstrap values.
- These groups have support of 95 or more:
- Lungfish South America and Africa (100).
- Frog with all reptiles, birds and mammals (100).
- Lizard, Sphenodon, Turtle, Crocodile and Bird (99).
- Mouse and Rat (100).
- Human, Seal, Cow, Whale, Mouse and Rat (100).
- Platypus and Opossum (96).
- Crocodile and Bird (97).
- Cow and Whale (99).
- These groups have support below 95:
- Human with Seal, Cow and Whale (93).
- Seal with Cow and Whale (71).
- Turtle with Crocodile and Bird (70).
- Sphenodon with Turtle, Crocodile and Bird (51). The placement of Sphenodon is uncertain.
Uncertain
- I set no outgroup, so the tree is unrooted. Do not read branch order as time order or as ancestry.
- The tree is the maximum-likelihood tree under TIM2+F+I+G4. It is not the true tree.
- I ran one seed only. I did not test other models or other seeds.
Waits for the scientist
- The scientist can name an outgroup if they want a rooted tree. That choice belongs to the scientist.
- The setup asks for a maximum gap fraction. It is not needed here because I did not trim the alignment.
Settings and step results that the answer does not mention. The harness adds them.
Settings used, from the decision record: Substitution model: MFP · Bootstrap type: ufboot · Number of bootstrap replicates: 1000 · Random seed: 1.
Checks
Review findings
The review recorded 8 findings. A rule finding comes from a fixed check in the harness. A referee finding comes from a second model that reads the record. The harness shows the findings to the scientist with the final answer. The record does not mark a finding as fixed. Thus a finding from an early review round can apply to a draft that the model corrected later.
| Severity | From | Finding | Shown with the final answer |
|---|---|---|---|
| error | ruledecision_misreported | The answer names standard for "Bootstrap type", but the decision record says ufboot. Report the value that was used. | yes |
| error | ruleunsourced_numbers | 4 numbers in the answer match no logged tool result: 95. Run the tool that measures each number, or remove the number, or say how you derived it from logged results. | yes |
| info | ruletext_style | The answer breaks the text rules (ASD-STE100) in 4 places. Sentence 12 uses the passive voice: "is already aligned". Use the active voice. Sentence 14 uses the passive voice: "is fixed". Use the active voice. Sentence 43 uses the passive voice: "is unrooted". Use the active voice. Sentence 53 uses the passive voice: "is not needed". Use the active voice. | yes |
| warning | referee model | The read of run.treefile returned only a byte count. The log does not show the tree text, so it cannot confirm which group each support value belongs to. The group names (Lungfish, Sphenodon, Seal and so on) have no visible source. | yes |
| warning | referee model | The answer says IQ-TREE 3.1.4 and that the run report lists TIM2+F+I+G4 as best by BIC. The infer_tree result shows neither the version nor the criterion. The report must name the criterion that ModelFinder used and the version, and the log must show them. | yes |
| warning | referee model | The answer opens with corrections to an 'earlier wording' and an 'earlier support list'. No earlier answer is in the log. A reader cannot check these corrections, and they suggest the group assignments were already wrong once. | yes |
| info | referee model | The report does not give the alignment gap fraction. It does say the file was not trimmed, so no columns were removed. The claim that the setup asks for a maximum gap fraction does not match the setup, which has trim set to false. | yes |
| info | referee model | The ufboot type, 1000 replicates, seed 1, model MFP, the log-likelihood, and the unrooted status are all stated and match the log. The 95 threshold comes from the program standard and is used correctly for ultrafast bootstrap. | yes |
Numbers in the answer
The last claim check read 36 numbers in the answer. 32 numbers match a logged result. 4 numbers have no source in the record.
Numbers that do not match a logged result (4)
- no source in the record: The numbers 95, 93, 71, 70 and 51 are not new measurements.
- no source in the record: The lab rule reads an ultrafast bootstrap value of 95 or more as well supported.
- no source in the record: - These groups have support of 95 or more:
- no source in the record: - These groups have support below 95:
Deviations
The model did not try to change a choice of the scientist.
Failed tool calls
No tool call failed.
Data integrity
Each data file has the same SHA-256 hash now as at the time of the step that read it. Where the download script (fetch.sh) gives a hash, the file also has that hash. The run did not change the data.
| File | SHA-256 | Fetched data | Steps with this hash |
|---|---|---|---|
{data}/minh2020-iqtree2/example.phy33.4 KB | d3365cbea79a | same as the hash in the download script (fetch.sh) | n1 |
A SHA-256 hash is a fingerprint of the file contents. If one byte of the file changes, the hash changes. The table shows the first 12 characters.
How to repeat it
Get the data. The script downloads the files and checks their SHA-256 hashes where it lists them.
CUVETTE_DATA={data} bash bench/papers/minh2020-iqtree2/fetch.shRun the same case with Cuvette. The script gives the same answers from bench/papers/minh2020-iqtree2/bench.yaml.
cuvette bench papers --papers minh2020-iqtree2 --models claude:claude-sonnet-5-5
Repeat each step by hand in the program. For each step, the harness records a manual route: the menu path or the code that gives the same result. This list does not include comparison runs.
infer_tree(step n1)Run: iqtree3 -s <alignment> -m <model> -B <replicates> -seed <seed> -T 1
-s
{data}/minh2020-iqtree2/example.phy- -m =
MFP - -B =
1000 - -seed =
1 - Warning: If you keep the default random, you get a different result.
The manual route that the harness recorded
/bin/sh {other volume}/tools/overnight/claude-final/catalog/phylo/scripts/iqtree.sh /opt/homebrew/bin/iqtree3 infer --alignment {data}/minh2020-iqtree2/example.phy --model MFP --bootstrap-type ufboot --replicates 1000 --seed 1The manual route gives the same numbers. An automatic test in Cuvette checks this.
Figure

Run facts
| Model | claude-sonnet-5-5 through the Anthropic service |
| Date | 2026-10-09 12:13:30 UTC |
| End of run | the model gave a final answer |
| Time | 89 s |
| Requests to the model | 4 |
| Tokensunits of text that the model read and wrote | 12 input, 3139 output, 32719 cache read, 13309 cache write |
| Cost estimate | $0.07 at list price, from the token counts |
| Tool calls | 2 (0 failed) |
| Adapters | phylo 0.1.2, program 3.1.4 |
| Session | 20261009-071330-26bb |
Code hash of each step (1)
| Step | Tool | Program version | Code hash |
|---|---|---|---|
| n1 | infer_tree | 3.1.4 | 26caef57428d |
The code hash is a fingerprint of the adapter name, the adapter version, the tool and its definition in the adapter. If one of these changes, the hash changes.
Haiku · claude-haiku-5-5 · run 3 of 3 shown 1 of 1 values match, 1 of 1 correct in the final answer
The session
This is the session as the scientist sees it, in order. Decision cards show the answer that the script gave and where that answer comes from. Each step shows the program, its version, the input file hash and the outputs. Click a line to see more.
setup The decision record starts with these values. A script set them before the first message, from the answers that we wrote before the run.
- MAFFT alignment method: autoSource in the tutorial or test suite: Not used. The file example.phy is already aligned. The phylo adapter uses MAFFT --auto by default.
- Trim the alignment: falseSource in the tutorial or test suite: Not in the tutorial. The tutorial uses the alignment as it is.
- Substitution model: MFPSource in the tutorial or test suite: Tutorial section on the substitution model. ModelFinder picks the model with the lowest Bayesian information criterion (BIC). MFP is the default from version 1.5.4.
- Type of branch support: ufbootSource in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial notes that these values do not mean the same as standard bootstrap values.
- Bootstrap replicates: 1000Source in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial calls 1000 the minimum number it recommends.
- Outgroup: Source in the tutorial or test suite: Tutorial, first run. IQ-TREE draws the lungfish at the root only because it is the first sequence. The tree stays unrooted.
- Random seed: 1
harness note A note from the harness to the model. The scientist did not write it.
- Setup values: Alignment: - Alignment method (MAFFT) (algorithm): auto - Trim the alignment (trim): false Tree: - Substitution model (model): MFP - Bootstrap type (bootstrap_type): ufboot - Number of bootstrap replicates (replicates): 1000 - Outgroup (outgroup): - Random seed (seed): 1 Ask the scientist: Largest gap fraction of a kept column (max_gap_fraction). Fixed by the lab: Random seed (seed) = 1. The harness fills these values into your tool calls automatically. If you give a different value, the harness logs a deviation and shows it to the scientist. - Mode is now watch: you run the steps.
Model thinking, summary: Since example.phy is already an aligned PHYLIP file, I can skip alignment and go straight to tree inference using MFP for model selection with ufboot bootstrapping at 1000 replicates and seed 1. Trimming isn't requested here, so I don't need to ask about gap fraction thresholds.
Model
The file is a PHYLIP alignment, so it needs no alignment step. I will run the tree tool directly with ModelFinder (MFP), ultrafast bootstrap with 1000 replicates, and seed 1.
infer_tree (adapter phylo).step n1 infer_tree adapter phylo 0.1.2, IQ-TREE 3.1.4
Model TIM2+F+I+G4. Log-likelihood -21152.5. Ultrafast bootstrap with 1000 replicates.
Decisions applied: Substitution model = MFP; Type of branch support = ufboot; Bootstrap replicates = 1000; Random seed = 1.
Input file: {data}/minh2020-iqtree2/example.phy SHA-256 d3365cbea79a.
Outputs: run.bionj (c4103d181dee), run.ckp.gz (4cb6fbec3d28), run.contree (cf1f8eae6015), run.iqtree (a3eb7b748ce6), run.log (e7645d2dd5c3), run.mldist (7c9d04e6c668), run.model.gz (9d7ec57ddc04), run.splits.nex (650ac7edfeae), run.treefile (4cff6aaf2ea1), stderr.txt (e3b0c44298fc), stdout.txt (e3b0c44298fc).
Arguments
| path | {data}/minh2020-iqtree2/example.phy |
| model | MFP |
| bootstrap_type | ufboot |
| replicates | 1000 |
| seed | 1 |
Tool output
report={work}/infer_tree-1/run.iqtree
seed=1
bootstrap_type=ufboot
bootstrap_replicates=1000
treefile={work}/infer_tree-1/run.treefile
consensus_tree={work}/infer_tree-1/run.contree
n_sequences=17
n_sites=1998
model=TIM2+F+I+G4
best_model_bic=TIM2+F+I+G4
log_likelihood=-21152.5377
tree_length=4.2173
constant_sites=686
site_patterns=1152
Model thinking, summary: I'm noting the key run details for the report: BIC selected TIM2+F+I+G4, program IQ-TREE 3.1.4, log-likelihood -21152.5377, tree length 4.2173, seed 1, 1000 ultrafast bootstraps, unrooted tree with no outgroup, and an alignment of 17 already-aligned sequences in PHYLIP format.
Model
The tree run finished. I will read the consensus tree file to get the branch support values.
read_file.Result of read_file: Read 590 of 590 bytes from {work}/infer_tree-1/run.contree.
(LngfishAu:0.1709765394,(LngfishSA:0.1884042664,LngfishAf:0.1648765322)100:0.1073673156,(Frog:0.2562938256,((((Turtle:0.2215450915,(Crocodile:0.3057507446,Bird:0.2310620493)97:0.0651606330)70:0.0364894400,Sphenodon:0.3447780527)51:0.0204536475,Lizard:0.3861439885)99:0.0739888076,(((Human:0.1851778377,(Seal:0.0944230016,(Cow:0.0823003161,Whale:0.1012908438)99:0.0404371901)71:0.0252744924)93:0.0340669342,(Mouse:0.0584259465,Rat:0.0905415196)100:0.1218030843)100:0.0607194581,(Platypus:0.1920645910,Opossum:0.1510258428)96:0.0373287871)100:0.1490890925)100:0.1275997766)100:0.0940810002);
alignment_stats (adapter phylo).Failed of alignment_stats: Failed: sh exited with code 1: error: no sequences found
Model thinking, summary: The alignment_stats tool only handles FASTA, and this file is PHYLIP format, so it will fail here—I'll note that gap fraction wasn't measured that way. Instead, I'll write a quick script to parse the PHYLIP file directly and count gap characters to get that fraction myself.
Model
The alignment tool reads only FASTA, so it failed on this PHYLIP file. I will count the gaps with a short script.
The model runs run_script.
import re
path = "{data}/minh2020-iqtree2/example.phy"
with open(path) as f:
lines = [l.rstrip("\n") for l in f if l.strip()]
header = lines[0].split()
n, L = int(header[0]), int(header[1])
seqs = []
for l in lines[1:]:
parts = l.split(None, 1)
seqs.append(parts[1] if len(parts) > 1 else "")
# interleaved format: join all residue chunks per sequence
if len(seqs) > n:
seqs2 = [""] * n
for i, l in enumerate(lines[1:]):
parts = l.split(None, 1)
if i < n:
seqs2[i] = parts[1].replace(" ", "") if len(parts) > 1 else ""
else:
seqs2[i % n] += l.replace(" ", "")
seqs = seqs2
seqs = [s.replace(" ", "") for s in seqs]
total = sum(len(s) for s in seqs)
gaps = sum(s.count("-") + s.count("?") for s in seqs)
print("sequences", n, "declared sites", L)
print("lengths", sorted(set(len(s) for s in seqs)))
print("gap or missing characters", gaps, "of", total, "fraction", round(gaps/total, 4))
run_script. In this validation, the harness lets the model use every tool and does not ask the scientist.Result of run_script: The script ran in {work} and wrote 0 new file(s) to {work}.
sequences 17 declared sites 1998 lengths [1998] gap or missing characters 36 of 33966 fraction 0.0011
Model Answer draft 1. The claim check and the review read it.
Result: best-fit model and maximum-likelihood tree
Input: The file example.phy is already aligned, so I did not run an alignment step. I did not trim it. The gap fraction is 0.0011: 36 gap or missing characters out of 33,966 positions. This came from a script count, because the alignment tool reads only FASTA and failed on this PHYLIP file.
Model: TIM2+F+I+G4. ModelFinder chose it by BIC (Bayesian information criterion). The tool reported only the BIC choice, not the AIC (Akaike information criterion) choice.
Tree run:
- Program: IQ-TREE, listed as version 3.1.4 in the program header.
- Log-likelihood of the maximum-likelihood tree: −21152.5377.
- Tree length: 4.2173.
- Seed: 1.
- Bootstrap: ultrafast bootstrap (UFBoot), 1000 replicates.
- Outgroup: none. The tree is unrooted, so branch order does not show time order.
- Data: 17 sequences, 1998 sites, 1152 site patterns, 686 constant sites.
Branch support (UFBoot values from the consensus tree):
Well supported (95 or more):
- Lungfish South America + Africa: 100
- Frog + amniotes: 100
- Mammals as one clade: 100
- Mouse + Rat: 100
- Platypus + Opossum: 96
- Crocodile + Bird: 97
- Cow + Whale: 99
- Lizard + the other sauropsids (Turtle, Crocodile, Bird, Sphenodon): 99
Weakly supported (below 95):
- Human + (Seal, Cow, Whale): 93
- Seal + (Cow, Whale): 71
- Turtle + (Crocodile, Bird): 70
- Sphenodon placement (within the sauropsid clade): 51
Uncertain:
- The support values come from the bootstrap consensus tree. The maximum-likelihood tree is in
run.treefile. - The four weak nodes above can change with new data. Read the placement of Sphenodon and the Seal node as unresolved.
- This is the maximum-likelihood tree under TIM2+F+I+G4. It is not the true tree.
Files:
- Report:
infer_tree-1/run.iqtree - Maximum-likelihood tree:
infer_tree-1/run.treefile - Consensus tree with support values:
infer_tree-1/run.contree
Waiting for the scientist:
- An outgroup, if you want the tree rooted. Without one, the tree stays unrooted.
- Nothing else blocks the analysis. Trimming is off, so the largest gap fraction setting is not needed.
review The review sent the answer back to the model before the scientist saw it.
- 5 numbers in the answer match no logged tool result: 0.0011, 36, 33966, 95. Run the tool that measures each number, or remove the number, or say how you derived it from logged results. - The branch support values come from run.contree. The read at #19 returned no file content, so the support values are not visible in the log. They must be checked against the file before use. Fix these problems. Then give the answer again.
Model
I will read the consensus tree file again to check the support values. I will also rerun the gap count so the numbers come from a logged result.
read_file.Result of read_file: Read 590 of 590 bytes from {work}/infer_tree-1/run.contree.
(LngfishAu:0.1709765394,(LngfishSA:0.1884042664,LngfishAf:0.1648765322)100:0.1073673156,(Frog:0.2562938256,((((Turtle:0.2215450915,(Crocodile:0.3057507446,Bird:0.2310620493)97:0.0651606330)70:0.0364894400,Sphenodon:0.3447780527)51:0.0204536475,Lizard:0.3861439885)99:0.0739888076,(((Human:0.1851778377,(Seal:0.0944230016,(Cow:0.0823003161,Whale:0.1012908438)99:0.0404371901)71:0.0252744924)93:0.0340669342,(Mouse:0.0584259465,Rat:0.0905415196)100:0.1218030843)100:0.0607194581,(Platypus:0.1920645910,Opossum:0.1510258428)96:0.0373287871)100:0.1490890925)100:0.1275997766)100:0.0940810002);
The model runs run_script.
path = "{data}/minh2020-iqtree2/example.phy"
with open(path) as f:
lines = [l.rstrip("\n") for l in f if l.strip()]
header = lines[0].split()
n, L = int(header[0]), int(header[1])
body = lines[1:]
seqs = [""] * n
if len(body) == n:
for i, l in enumerate(body):
parts = l.split(None, 1)
seqs[i] = parts[1].replace(" ", "") if len(parts) > 1 else ""
else:
for i, l in enumerate(body):
parts = l.split(None, 1)
if i < n:
seqs[i] = parts[1].replace(" ", "") if len(parts) > 1 else ""
else:
seqs[i % n] += l.replace(" ", "")
total = sum(len(s) for s in seqs)
gaps = sum(s.count("-") + s.count("?") for s in seqs)
print("sequences:", n, "declared sites:", L)
print("sequence lengths:", sorted(set(len(s) for s in seqs)))
print("total positions:", total)
print("gap or missing characters:", gaps)
print("gap fraction:", round(gaps / total, 4))
run_script. In this validation, the harness lets the model use every tool and does not ask the scientist.Result of run_script: The script ran in {work} and wrote 0 new file(s) to {work}.
sequences: 17 declared sites: 1998 sequence lengths: [1998] total positions: 33966 gap or missing characters: 36 gap fraction: 0.0011
Model thinking, summary: The file content is confirmed and the support values match what I listed, verified against the returned data. I'm noting the node with support 51 joins Turtle+Croc+Bird with Sphenodon, and I'll keep the terminology precise, avoiding "should" per the review guidance while finalizing the answer.
Model The model writes its answer. The claim check reads it. See the final answer.
The final answer
The model wrote this answer. The harness shows it to the scientist after the review.
Result: best-fit model and maximum-likelihood tree
Input: The file example.phy is already aligned, so I did not run an alignment step. I did not trim it.
- Sequences: 17. Sites: 1998 (infer_tree output).
- Gap or missing characters: 36 of 33,966 total positions. Gap fraction: 0.0011 (script count, rerun this session).
Model: TIM2+F+I+G4. ModelFinder chose it by BIC (Bayesian information criterion). The tool did not report an AIC (Akaike information criterion) choice.
Tree run (infer_tree output):
- Program: IQ-TREE, version 3.1.4 as listed in the program header.
- Log-likelihood of the maximum-likelihood tree: −21152.5377.
- Tree length: 4.2173.
- Seed: 1.
- Bootstrap: ultrafast bootstrap (UFBoot), 1000 replicates.
- Outgroup: none. The tree is unrooted, so branch order does not show time order.
- Site patterns: 1152. Constant sites: 686.
Branch support (UFBoot values, checked against the consensus tree file this session):
Well supported (at or above the 95 threshold in the program reporting standards):
- Lungfish South America + Africa: 100
- Frog + amniotes: 100
- Mammals as one clade: 100
- Mouse + Rat: 100
- Cow + Whale: 99
- Lizard + the other sauropsids (Turtle, Crocodile, Bird, Sphenodon): 99
- Crocodile + Bird: 97
- Platypus + Opossum: 96
Weakly supported (below 95):
- Human + (Seal, Cow, Whale): 93
- Seal + (Cow, Whale): 71
- Turtle + (Crocodile, Bird): 70
- Sphenodon placement in the sauropsid clade: 51
Uncertain:
- The support values come from the bootstrap consensus tree (
run.contree). The maximum-likelihood tree is inrun.treefile. - The four weak nodes above can change with new data. Read the Sphenodon placement and the Seal node as unresolved.
- This is the maximum-likelihood tree under TIM2+F+I+G4. It is not the true tree.
Files:
- Report:
infer_tree-1/run.iqtree - Maximum-likelihood tree:
infer_tree-1/run.treefile - Consensus tree with support values:
infer_tree-1/run.contree
Waiting for the scientist:
- An outgroup, if you want the tree rooted. Without one, the tree stays unrooted.
- Nothing else blocks the analysis. Trimming is off, so the largest gap fraction setting is not needed.
Settings and step results that the answer does not mention. The harness adds them.
Settings used, from the decision record: Substitution model: MFP · Bootstrap type: ufboot · Number of bootstrap replicates: 1000 · Random seed: 1.
Checks
Review findings
The review recorded 9 findings. A rule finding comes from a fixed check in the harness. A referee finding comes from a second model that reads the record. The harness shows the findings to the scientist with the final answer. The record does not mark a finding as fixed. Thus a finding from an early review round can apply to a draft that the model corrected later.
| Severity | From | Finding | Shown with the final answer |
|---|---|---|---|
| warning | rulefailed_result_used | Step 3 (alignment_stats) failed and was not repeated. Error: sh exited with code 1: error: no sequences found | yes |
| error | ruleunsourced_numbers | 5 numbers in the answer match no logged tool result: 36, 33966, 0.0011, 95. Run the tool that measures each number, or remove the number, or say how you derived it from logged results. | yes |
| info | ruletext_style | The answer breaks the text rules (ASD-STE100) in 3 places. Sentence 1 uses the passive voice: "is already aligned". Use the active voice. Sentence 17 uses the passive voice: "is unrooted". Use the active voice. Sentence 50 uses the passive voice: "is not needed". Use the active voice. | yes |
| error | referee model | The answer says the branch support values were checked against run.contree. The logged reads of that file show only byte counts and no content. The support values cannot be verified from the log. The answer must not claim a file check that the log does not show. | yes |
| warning | referee model | The answer says ModelFinder chose TIM2+F+I+G4 by BIC. The infer_tree result gives only the model name. It does not name the criterion. The BIC claim has no logged source, and the claim that no AIC choice was reported is also unsupported. | yes |
| warning | referee model | The answer gives IQ-TREE version 3.1.4 from the program header. The logged infer_tree result does not show a version. The version must be marked as unverified or removed. | yes |
| info | referee model | The alignment_stats call failed with 'no sequences found'. The answer does not mention this failure. The gap fraction came from a run_script count instead. The failure should be reported. | yes |
| info | referee model | The same gap count was run twice (steps 4 and 6). The second run was made only so the numbers would be logged. The 36 gaps and 0.0011 fraction are in both results. The answer should name the count as gap or missing characters, which is what the script counted. | yes |
| info | referee model | The answer does not state the number of columns removed. It says the alignment was not trimmed, so the count is zero. The answer should state this directly, as the alignment rules require. | yes |
Numbers in the answer
The last claim check read 28 numbers in the answer. 23 numbers match a logged result. 5 numbers have no source in the record.
Numbers that do not match a logged result (5)
- no source in the record: - Gap or missing characters: 36 of 33,966 total positions.
- no source in the record: - Gap or missing characters: 36 of 33,966 total positions.
- no source in the record: Gap fraction: 0.0011 (script count, rerun this session).
- no source in the record: Well supported (at or above the 95 threshold in the program reporting standards):
- no source in the record: Weakly supported (below 95):
Deviations
The model did not try to change a choice of the scientist.
Failed tool calls
1 tool call failed. The model then tried again or used another tool. The session above shows each failure.
Data integrity
Each data file has the same SHA-256 hash now as at the time of the step that read it. Where the download script (fetch.sh) gives a hash, the file also has that hash. The run did not change the data.
| File | SHA-256 | Fetched data | Steps with this hash |
|---|---|---|---|
{data}/minh2020-iqtree2/example.phy33.4 KB | d3365cbea79a | same as the hash in the download script (fetch.sh) | n1 |
A SHA-256 hash is a fingerprint of the file contents. If one byte of the file changes, the hash changes. The table shows the first 12 characters.
How to repeat it
Get the data. The script downloads the files and checks their SHA-256 hashes where it lists them.
CUVETTE_DATA={data} bash bench/papers/minh2020-iqtree2/fetch.shRun the same case with Cuvette. The script gives the same answers from bench/papers/minh2020-iqtree2/bench.yaml.
cuvette bench papers --papers minh2020-iqtree2 --models claude:claude-haiku-5-5
Repeat each step by hand in the program. For each step, the harness records a manual route: the menu path or the code that gives the same result. This list does not include comparison runs.
infer_tree(step n1)Run: iqtree3 -s <alignment> -m <model> -B <replicates> -seed <seed> -T 1
-s
{data}/minh2020-iqtree2/example.phy- -m =
MFP - -B =
1000 - -seed =
1 - Warning: If you keep the default random, you get a different result.
The manual route that the harness recorded
/bin/sh {other volume}/tools/overnight/claude-final/catalog/phylo/scripts/iqtree.sh /opt/homebrew/bin/iqtree3 infer --alignment {data}/minh2020-iqtree2/example.phy --model MFP --bootstrap-type ufboot --replicates 1000 --seed 1The manual route gives the same numbers. An automatic test in Cuvette checks this.
Figure

Run facts
| Model | claude-haiku-5-5 through the Anthropic service |
| Date | 2026-10-09 12:23:39 UTC |
| End of run | the model gave a final answer |
| Time | 98 s |
| Requests to the model | 7 |
| Tokensunits of text that the model read and wrote | 18 input, 8058 output, 80944 cache read, 19257 cache write |
| Cost estimate | $0.01 at list price, from the token counts |
| Tool calls | 6 (1 failed) |
| Adapters | phylo 0.1.2, program 3.1.4 |
| Session | 20261009-072339-0107 |
Code hash of each step (1)
| Step | Tool | Program version | Code hash |
|---|---|---|---|
| n1 | infer_tree | 3.1.4 | 26caef57428d |
The code hash is a fingerprint of the adapter name, the adapter version, the tool and its definition in the adapter. If one of these changes, the hash changes.
qwen3:8b · local model through Ollama 1 of 1 values match, 1 of 1 correct in the final answer
The session
This is the session as the scientist sees it, in order. Decision cards show the answer that the script gave and where that answer comes from. Each step shows the program, its version, the input file hash and the outputs. Click a line to see more.
setup The decision record starts with these values. A script set them before the first message, from the answers that we wrote before the run.
- MAFFT alignment method: autoSource in the tutorial or test suite: Not used. The file example.phy is already aligned. The phylo adapter uses MAFFT --auto by default.
- Trim the alignment: falseSource in the tutorial or test suite: Not in the tutorial. The tutorial uses the alignment as it is.
- Substitution model: MFPSource in the tutorial or test suite: Tutorial section on the substitution model. ModelFinder picks the model with the lowest Bayesian information criterion (BIC). MFP is the default from version 1.5.4.
- Type of branch support: ufbootSource in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial notes that these values do not mean the same as standard bootstrap values.
- Bootstrap replicates: 1000Source in the tutorial or test suite: Tutorial section on the ultrafast bootstrap. The tutorial calls 1000 the minimum number it recommends.
- Outgroup: Source in the tutorial or test suite: Tutorial, first run. IQ-TREE draws the lungfish at the root only because it is the first sequence. The tree stays unrooted.
- Random seed: 1
harness note A note from the harness to the model. The scientist did not write it.
- Setup values: Alignment: - Alignment method (MAFFT) (algorithm): auto - Trim the alignment (trim): false Tree: - Substitution model (model): MFP - Bootstrap type (bootstrap_type): ufboot - Number of bootstrap replicates (replicates): 1000 - Outgroup (outgroup): - Random seed (seed): 1 Ask the scientist: Largest gap fraction of a kept column (max_gap_fraction). Fixed by the lab: Random seed (seed) = 1. The harness fills these values into your tool calls automatically. If you give a different value, the harness logs a deviation and shows it to the scientist. - Mode is now watch: you run the steps.
align_sequences (adapter phylo).Failed of align_sequences: Failed: sh exited with code 1: error: this file is a PHYLIP file, so the sequences are already aligned. Do not align it. Give the path to infer_tree.
infer_tree (adapter phylo).step n1 infer_tree adapter phylo 0.1.2, IQ-TREE 3.1.4
Model TIM2+F+I+G4. Log-likelihood -21152.5. Ultrafast bootstrap with 1000 replicates.
Decisions applied: Substitution model = MFP; Type of branch support = ufboot; Bootstrap replicates = 1000; Random seed = 1.
Input file: {data}/minh2020-iqtree2/example.phy SHA-256 d3365cbea79a.
Outputs: run.bionj (c4103d181dee), run.ckp.gz (a900be8ae10e), run.contree (cf1f8eae6015), run.iqtree (adeedb463309), run.log (6f00f371f252), run.mldist (7c9d04e6c668), run.model.gz (2b7fe4474406), run.splits.nex (650ac7edfeae), run.treefile (4cff6aaf2ea1), stderr.txt (e3b0c44298fc), stdout.txt (e3b0c44298fc).
Arguments
| bootstrap_type | ufboot |
| model | MFP |
| path | {data}/minh2020-iqtree2/example.phy |
| replicates | 1000 |
| seed | 1 |
Tool output
report={work}/infer_tree-1/run.iqtree
seed=1
bootstrap_type=ufboot
bootstrap_replicates=1000
treefile={work}/infer_tree-1/run.treefile
consensus_tree={work}/infer_tree-1/run.contree
n_sequences=17
n_sites=1998
model=TIM2+F+I+G4
best_model_bic=TIM2+F+I+G4
log_likelihood=-21152.5377
tree_length=4.2173
constant_sites=686
site_patterns=1152
Model Answer draft 1. The claim check and the review read it.
The best-fit substitution model is TIM2+F+I+G4. The maximum-likelihood tree has ultrafast bootstrap support values. The log-likelihood is -21152.5. The treefile is at {work}/infer_tree-1/run.treefile. The tree is unrooted.
review The review sent the answer back to the model before the scientist saw it.
- 3 numbers in the answer match no logged tool result: 20261009, 53857, 157. Run the tool that measures each number, or remove the number, or say how you derived it from logged results. - The log-likelihood must be reported with more decimal places. Fix these problems. Then give the answer again.
Model The model writes its answer. The claim check reads it. See the final answer.
The final answer
The model wrote this answer. The harness shows it to the scientist after the review.
The best-fit substitution model is TIM2+F+I+G4. The maximum-likelihood tree has ultrafast bootstrap support values. The log-likelihood is -21152.5377. The treefile is at {work}/infer_tree-1/run.treefile. The tree is unrooted.
Settings and step results that the answer does not mention. The harness adds them.
Settings used, from the decision record: Substitution model: MFP · Bootstrap type: ufboot · Number of bootstrap replicates: 1000 · Random seed: 1.
Checks
Review findings
The review recorded 4 findings. A rule finding comes from a fixed check in the harness. A referee finding comes from a second model that reads the record. The harness shows the findings to the scientist with the final answer. The record does not mark a finding as fixed. Thus a finding from an early review round can apply to a draft that the model corrected later.
| Severity | From | Finding | Shown with the final answer |
|---|---|---|---|
| warning | rulefailed_result_used | Step 1 (align_sequences) failed and was not repeated. Error: sh exited with code 1: error: this file is a PHYLIP file, so the sequences are already aligned. Do not align it. Give the path to infer_tree. | yes |
| error | ruleunsourced_numbers | 3 numbers in the answer match no logged tool result: 20261009, 53857, 157. Run the tool that measures each number, or remove the number, or say how you derived it from logged results. | yes |
| info | ruletext_style | The answer breaks the text rules (ASD-STE100) in 1 place. Sentence 5 uses the passive voice: "is unrooted". Use the active voice. | yes |
| error | referee model | The log-likelihood must be reported with the substitution model. | yes |
Numbers in the answer
The last claim check read 5 numbers in the answer. 2 numbers match a logged result. 3 numbers have no source in the record.
Numbers that do not match a logged result (3)
- no source in the record: The treefile is at {work}/infer_tree-1/run.treefile.
- no source in the record: The treefile is at {work}/infer_tree-1/run.treefile.
- no source in the record: The treefile is at {work}/infer_tree-1/run.treefile.
Deviations
The model did not try to change a choice of the scientist.
Failed tool calls
1 tool call failed. The model then tried again or used another tool. The session above shows each failure.
Data integrity
Each data file has the same SHA-256 hash now as at the time of the step that read it. Where the download script (fetch.sh) gives a hash, the file also has that hash. The run did not change the data.
| File | SHA-256 | Fetched data | Steps with this hash |
|---|---|---|---|
{data}/minh2020-iqtree2/example.phy33.4 KB | d3365cbea79a | same as the hash in the download script (fetch.sh) | n1 |
A SHA-256 hash is a fingerprint of the file contents. If one byte of the file changes, the hash changes. The table shows the first 12 characters.
How to repeat it
Get the data. The script downloads the files and checks their SHA-256 hashes where it lists them.
CUVETTE_DATA={data} bash bench/papers/minh2020-iqtree2/fetch.shRun the same case with Cuvette. The script gives the same answers from bench/papers/minh2020-iqtree2/bench.yaml.
cuvette bench papers --papers minh2020-iqtree2 --models ollama:qwen3:8b
Repeat each step by hand in the program. For each step, the harness records a manual route: the menu path or the code that gives the same result. This list does not include comparison runs.
infer_tree(step n1)Run: iqtree3 -s <alignment> -m <model> -B <replicates> -seed <seed> -T 1
-s
{data}/minh2020-iqtree2/example.phy- -m =
MFP - -B =
1000 - -seed =
1 - Warning: If you keep the default random, you get a different result.
The manual route that the harness recorded
/bin/sh {other volume}/tools/overnight/checkout/catalog/phylo/scripts/iqtree.sh /opt/homebrew/bin/iqtree3 infer --alignment {data}/minh2020-iqtree2/example.phy --model MFP --bootstrap-type ufboot --replicates 1000 --seed 1The manual route gives the same numbers. An automatic test in Cuvette checks this.
Figure

Run facts
| Model | qwen3:8b through Ollama, on our own computer |
| Date | 2026-10-09 10:38:58 UTC |
| End of run | the model gave a final answer |
| Time | 150 s |
| Requests to the model | 4 |
| Tokensunits of text that the model read and wrote | 23657 input, 378 output, 0 cache read, 0 cache write |
| Cost estimate | none: the model runs on our own computer |
| Tool calls | 2 (1 failed) |
| Adapters | phylo 0.1.2, program 3.1.4 |
| Session | 20261009-053857-157c |
Code hash of each step (1)
| Step | Tool | Program version | Code hash |
|---|---|---|---|
| n1 | infer_tree | 3.1.4 | 26caef57428d |
The code hash is a fingerprint of the adapter name, the adapter version, the tool and its definition in the adapter. If one of these changes, the hash changes.