Validation / Papers / Ram 2021
Pixelwise H-score: A novel digital image analysis-based metric to quantify membrane biomarker expression from immunohistochemistry images
How to read this page. In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. A known value comes from the paper, from a tutorial or from a check that we ran. This page has no combined run of the paper yet.
Run of 9 October 2026, claude-haiku-5-5: 12 of 12 values computed, 12 of 12 correct in the final answer
The paper
Ram S, Vizcarra P, Whalen P, Deng S, Painter CL, Jackson-Fisher A, Pirie-Shepherd S, Xia X, Powell EL. Pixelwise H-score: A novel digital image analysis-based metric to quantify membrane biomarker expression from immunohistochemistry images. PLOS ONE 16(9):e0245638 (2021). doi:10.1371/journal.pone.0245638
Related sources:
- Data: S1 Table of the paper (XLSX, CC BY 4.0). The whole slide images are on Zenodo (record 4900103). The benchmark does not use them. link
What it measured
The authors scored membrane markers on tumor resections with a pathologist H-score and with digital H-scores from QuPath and HALO. They compared each digital score with the pathologist H-score by the Spearman correlation, and proposed the pixelwise H-score.
Data
S1 Table of the paper, journal.pone.0245638.s002.xlsx. fetch.sh checks the SHA-256 sum and joins the three sheets into one CSV. Size: 8 KB table with 75 cases..
License: CC BY 4.0
The instruction
A script sends this message as the scientist.
Basis: Results of the paper (Spearman correlations and the case numbers) and S1 Table (the percents and the scores). The paper gives no kappa or ICC. The benchmark adds these two items and takes the values from scikit-learn and pingouin.
The decisions
The model asks questions during a run. A script gives these answers to the questions of the model. We wrote the answers before the run.
| Decision | Value | Source |
|---|---|---|
| Region to score | whole | Not applicable. The table has one row for each case, not for each cell. |
| Smallest number of cells for a score | 0 | Not applicable. The table holds percents of cells, not cells. |
| Weights of the kappa | quadratic | Not in the paper. The prompt names the weights. |
| Form of the intraclass correlation | ICC2 | Not in the paper. The prompt asks for ICC(2,1). |
| Cut points that make categories | 100,200 | Not in the paper. The prompt names the cut points. |
| Other questions of the agent | Use the values in the decision record. | Not in the paper. The benchmark answers each free question with this text. |
Known values
The tolerance is the largest difference from the known value that we accept. We set it before the run. Exact: the number must be the same.
| Value | Known value | Tolerance | Source |
|---|---|---|---|
n_pcadCases scored for P-cadherinSource of the known valuePrinted in the paperWhere: Fig 2 legend. n = 30 cases for P-cadherin.Check: pix_h_score_table.csv has 30 P-cadherin rows (checks/check_ihc.out of the adapter).Note in the list of known values: published, Fig 2 legend | 30 | exact | Printed in the paper |
n_pdl1Cases scored for PD-L1Source of the known valuePrinted in the paperWhere: Fig 2 legend. n = 24 cases for PD-L1.Check: 24 PD-L1 rows.Note in the list of known values: published, Fig 2 legend | 24 | exact | Printed in the paper |
n_5t4Cases scored for 5T4Source of the known valuePrinted in the paperWhere: Fig 2 legend. n = 21 cases for 5T4.Check: 21 5T4 rows.Note in the list of known values: published, Fig 2 legend | 21 | exact | Printed in the paper |
h_pdl1_case1H-score of PD-L1 case 1 from its percents (0, 20, 75 and 5 percent at 0, 1+, 2+, 3+)Source of the known valuePrinted in the paper. Independent check: yesWhere: S1 Table, PD-L1 sheet, case 1. The pathologist H-score is 185 ("rounded up to the nearest integer").Check: The H-score formula on the percents 20, 75 and 5 gives 20 + 150 + 15 = 185.Note in the list of known values: published, S1 Table | 185 | ± 0.5 | Printed in the paper. Independent check: yes |
rho_pcad_qupathSpearman rho, P-cadherin, pathologist H-score and QuPath H-scoreSource of the known valuePrinted in the paperWhere: Results, P-cadherin. Spearman rho 0.39 (p = 0.03) for the QuPath H-score.Check: SciPy spearmanr gives 0.392301 (checks/check_ihc.out of the adapter).Note in the list of known values: published, Results | 0.39 | ± 0.01 | Printed in the paper |
rho_pcad_haloSpearman rho, P-cadherin, pathologist H-score and HALO H-scoreSource of the known valuePrinted in the paperWhere: Results, P-cadherin. Spearman rho 0.5 (p = 0.005) for the HALO H-score.Check: SciPy gives 0.495105.Note in the list of known values: published, Results | 0.5 | ± 0.01 | Printed in the paper |
rho_pdl1_qupathSpearman rho, PD-L1, pathologist H-score and QuPath H-scoreSource of the known valuePrinted in the paperWhere: Results, PD-L1. Spearman rho 0.74 for the QuPath H-score.Check: SciPy gives 0.739347.Note in the list of known values: published, Results | 0.74 | ± 0.01 | Printed in the paper |
rho_pdl1_haloSpearman rho, PD-L1, pathologist H-score and HALO H-score (the paper prints 0.69, the data give 0.701)Source of the known valuePrinted in the paperWhere: Results, PD-L1. Spearman rho 0.69 (p = 0.0002) for the HALO H-score.Check: SciPy gives 0.701264 on S1 Table. The paper prints 0.69, so the item has a tolerance of 0.015.Note in the list of known values: published, Results | 0.69 | ± 0.015 | Printed in the paper |
rho_5t4_qupathSpearman rho, 5T4, pathologist H-score and QuPath H-scoreSource of the known valuePrinted in the paperWhere: Results, 5T4. Spearman rho 0.79 for the QuPath H-score.Check: SciPy gives 0.789337.Note in the list of known values: published, Results | 0.79 | ± 0.01 | Printed in the paper |
rho_5t4_haloSpearman rho, 5T4, pathologist H-score and HALO H-scoreSource of the known valuePrinted in the paperWhere: Results, 5T4. Spearman rho 0.75 for the HALO H-score.Check: SciPy gives 0.749676.Note in the list of known values: published, Results | 0.75 | ± 0.01 | Printed in the paper |
kappa_pcad_qupathWeighted kappa (quadratic, cut points 100 and 200), P-cadherin, pathologist and QuPathSource of the known valueIndependent check: we calculated itTool: scikit-learn 1.9.1 cohen_kappa_scoreWhere: Not printed. The paper reports no kappa. The value is for the P-cadherin cases with cut points 100 and 200 and quadratic weights.Check: scikit-learn gives 0.010989 (checks/check_ihc.py and .out of the adapter).Note in the list of known values: independent, scikit-learn 1.9.1 (not printed in the paper) | 0.010989 | ± 0.001 | Independent check: we calculated it |
icc2_pcad_qupathICC(2,1), P-cadherin, pathologist and QuPathSource of the known valueIndependent check: we calculated itTool: pingouin 0.7.0 intraclass_corrWhere: Not printed. The paper reports no ICC. The value is ICC(2,1) of the pathologist H-score and the QuPath H-score for P-cadherin.Check: pingouin gives 0.149255.Note in the list of known values: independent, pingouin 0.7.0 (not printed in the paper) | 0.149255 | ± 0.002 | Independent check: we calculated it |
Latest scored run
Model: claude-haiku-5-5. Runs for each paper and model: 1. Blind mode: on. Status: answer. 158 s. Computed: 12 of 12 values. Reported: 12 of 12 values. The result file is bench/results/papers-2026-10-09-ihc-haiku.md. This run is not in the totals of the page of papers.
Computed: a logged number is within the tolerance. Reported: the final answer states the value, as the claim check measures. The table copies the cells of the result file.
| Item | Expected | Computed | Reported |
|---|---|---|---|
n_pcadCases scored for P-cadherin | 30 exact | pass 30 (n6 metrics.n_subjects, entry 69) | pass 30 via tolerance (n13, claim check 154) |
n_pdl1Cases scored for PD-L1 | 24 exact | pass 24 (n8 metrics.n_subjects, entry 75) | pass 24 via tolerance (n9, claim check 154) |
n_5t4Cases scored for 5T4 | 21 exact | pass 21 (n10 metrics.n_subjects, entry 81) | pass 21 via tolerance (n11, claim check 154) |
h_pdl1_case1H-score of PD-L1 case 1 from its percents (0, 20, 75 and 5 percent at 0, 1+, 2+, 3+) | 185 ±0.5 | pass 185 (n5 file:/Volumes/T7/guided-analysis/tools/bench-runs/2026-10-09/claude_claude-haiku-5-5/ram2021-pix-h-score/run1/home/sessions/20261009-043818-5dd8/work/score_h_score-4/pix_h_score_table_h_score.csv, entry 56) | pass 185 via tolerance (n13, claim check 154) |
rho_pcad_qupathSpearman rho, P-cadherin, pathologist H-score and QuPath H-score | 0.39 ±0.01 | pass 0.3923009 (n6 metrics.spearman_rho, entry 69) | pass 0.392 via tolerance (n6, claim check 154) |
rho_pcad_haloSpearman rho, P-cadherin, pathologist H-score and HALO H-score | 0.5 ±0.01 | pass 0.5003623 (n9 metrics.icc_icc3k, entry 78) | pass 0.495 via tolerance (n7, claim check 154) |
rho_pdl1_qupathSpearman rho, PD-L1, pathologist H-score and QuPath H-score | 0.74 ±0.01 | pass 0.7398156 (n7 table.rows[5][3], entry 72) | pass 0.739 via tolerance (n8, claim check 154) |
rho_pdl1_haloSpearman rho, PD-L1, pathologist H-score and HALO H-score (the paper prints 0.69, the data give 0.701) | 0.69 ±0.015 | pass 0.6924479 (n9 table.rows[4][3], entry 78) | pass 0.701 via tolerance (n9, claim check 154) |
rho_5t4_qupathSpearman rho, 5T4, pathologist H-score and QuPath H-score | 0.79 ±0.01 | pass 0.7895616 (n11 metrics.pearson_r, entry 84) | pass 0.789 via tolerance (n10, claim check 154) |
rho_5t4_haloSpearman rho, 5T4, pathologist H-score and HALO H-score | 0.75 ±0.01 | pass 0.7496755 (n11 metrics.spearman_rho, entry 84) | pass 0.75 via tolerance (n13, claim check 154) |
kappa_pcad_qupathWeighted kappa (quadratic, cut points 100 and 200), P-cadherin, pathologist and QuPath | 0.010989 ±0.001 | pass 0.01098901 (n6 metrics.kappa, entry 69) | pass 0.011 via tolerance (n6, claim check 154) |
icc2_pcad_qupathICC(2,1), P-cadherin, pathologist and QuPath | 0.149255 ±0.002 | pass 0.1492545 (n6 metrics.icc, entry 69) | pass 0.149 via tolerance (n6, claim check 154) |
Notes
Triage notes by the maintainers
The text below is from the triage notes. We show it as the maintainers wrote it.
Entry numbers (eNN) are ids in the log.jsonl of the session folder. Classes: a = tool or adapter fault, b = harness fault, c = benchmark spec fault, d = model fault.
- claude-haiku-5-5 (blind, adapter 0.1.0):
20261009-043818-5dd8, 158 s, computed 12/12, reported 12/12. Report:bench/results/papers-2026-10-09-ihc-haiku.md. - Leak check: 0 paths outside the allow list, 0 blocked reads, 0 unsourced claims.
- The run had no failed tool calls. No item needs a fix.