Validation / Papers / Zhu 2020
DEqMS: a method for accurate variance estimation in differential protein expression analysis
How to read this page. In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. A known value comes from the paper, from a tutorial or from a check that we ran. This page has no combined run of the paper yet.
Run of 9 October 2026, claude-haiku-5-5: 4 of 4 values computed, 4 of 4 correct in the final answer
The paper
Zhu Y, Orre LM, Zhou Tran Y, Mermelekas G, Johansson HJ, Malyutina A, Anders S, Lehtiö J. DEqMS: a method for accurate variance estimation in differential protein expression analysis. Molecular and Cellular Proteomics 19(6):1047-1057 (2020). doi:10.1074/mcp.TIR119.001646
Related sources:
- Data: PRIDE PXD000279, the label-free benchmark of Cox et al. 2014 (MaxLFQ). fetch.sh downloads proteomebenchmark.zip and unpacks proteinGroups.txt. doi:10.1074/mcp.M113.031591
What it measured
The authors compared tools for differential protein abundance on a benchmark of human proteins with E. coli proteins added in a 1 to 3 ratio. A tool finds a true positive if an E. coli protein is more abundant in the sample with more E. coli. Human proteins that are significant are false positives. They compare limma, limma with trend and DEqMS.
Data
PRIDE PXD000279, file proteomebenchmark.zip (38 MB), file proteinGroups.txt of the zip (28 MB). fetch.sh unpacks it. Size: 28 MB table with 6694 protein groups..
License: PRIDE data are open under the ProteomeXchange data policy. The paper is CC BY 4.0.
The instruction
A script sends this message as the scientist.
Basis: Results (Benchmarking of DEqMS) and Methods of the paper. The paper gives log2 LFQ values as the input of all methods and does not state a normalization.
The decisions
The model asks questions during a run. A script gives these answers to the questions of the model. We wrote the answers before the run.
| Decision | Value | Source |
|---|---|---|
| Flag columns that remove a row | Reverse, Contaminant | Not in the paper. Reverse and Contaminant give the 6566 proteins that the paper prints (check.R). |
| Peptides that a protein needs | 0 | The paper gives no peptide filter. |
| Samples that you exclude | none | The paper excludes no sample. |
| Valid values that a protein needs in a group | 2 | Results. Proteins with quantification values in at least two samples per condition. |
| In which groups the valid values must be present | each group | Results. At least two samples per condition. |
| Normalization of the log2 intensities | none | Results. The log2 LFQ intensities are the input, with no other normalization. |
| Imputation of missing values | none | The paper uses no imputation. |
| Variance model of the moderated t test | trend | Results. limma with trend = TRUE. |
| False discovery rate (FDR) level | 0.01 | Results. Adjusted p below 0.01, Benjamini-Hochberg. |
| Smallest log2 fold change | 0 | The paper uses no fold change cutoff. |
| Other questions of the agent | Use the values in the decision record. | Not in the paper. The benchmark answers each free question with this text. |
Known values
The tolerance is the largest difference from the known value that we accept. We set it before the run. Exact: the number must be the same.
| Value | Known value | Tolerance | Source |
|---|---|---|---|
n_after_flagsProtein groups after the removal of Reverse and Contaminant rowsSource of the known valuePrinted in the paperWhere: Results. 6566 proteins, of which 1902 are E. coli proteins and 4664 are human.Check: check.R gives 6566 and 1902.Note in the list of known values: published, Results of the paper | 6566 | exact | Printed in the paper |
n_validProteins with at least two values in each conditionSource of the known valuePrinted in the paperWhere: Results. 5022 proteins with quantification values in at least two samples in each condition.Check: check.R gives 5022.Note in the list of known values: published, Results of the paper | 5022 | exact | Printed in the paper |
tp_limma_trendE. coli proteins higher at the higher E. coli amount, limma with trend, adjusted p below 0.01 (the paper prints 1230; the tolerance of 5 is the known difference of this data version)Source of the known valuePrinted in the paperWhere: Results. The second best method was limma (trend = TRUE) with 1230 true positives and 13 false positives. The check gives 1235 for the proteins that are more abundant in H, 5 more than the paper.Check: check.R gives 1235.Note in the list of known values: published, Results of the paper | 1230 | ± 5 | Printed in the paper |
fp_limma_trendHuman proteins higher at the higher E. coli amount, limma with trend, adjusted p below 0.01Source of the known valuePrinted in the paperWhere: Results. 13 false positives for limma (trend = TRUE).Check: check.R gives 13.Note in the list of known values: published, Results of the paper | 13 | exact | Printed in the paper |
tp_deqms (reference)True positives of DEqMS in the paper (a different method; the tool does not run it)Source of the known valuePrinted in the paperWhere: Results. DEqMS reported 1237 true differentially expressed E. coli proteins and 11 human proteins.Check: The tool does not run DEqMS. A DEqMS run on this table gives 1218 and 7, so this item does not reproduce.Note in the list of known values: published, Results of the paper | 1237 | ± 0.5 | Printed in the paper |
fp_deqms (reference)False positives of DEqMS in the paper (a different method)Source of the known valuePrinted in the paperWhere: Results. 11 false positives for DEqMS.Check: Does not reproduce (7).Note in the list of known values: published, Results of the paper | 11 | ± 0.5 | Printed in the paper |
Latest scored run
Model: claude-haiku-5-5. Runs for each paper and model: 1. Blind mode: on. Status: answer. 118 s. Computed: 4 of 4 values. Reported: 4 of 4 values. The result file is bench/results/papers-2026-10-09-multiomics-haiku.md. This run is not in the totals of the page of papers.
Computed: a logged number is within the tolerance. Reported: the final answer states the value, as the claim check measures. The table copies the cells of the result file.
| Item | Expected | Computed | Reported |
|---|---|---|---|
n_after_flagsProtein groups after the removal of Reverse and Contaminant rows | 6566 exact | pass 6566 (n1 metrics.n_after_flags, entry 49) | pass 6566 via tolerance (n1, claim check 158) |
n_validProteins with at least two values in each condition | 5022 exact | pass 5022 (n1 metrics.n_kept, entry 49) | pass 5022 via tolerance (n6, claim check 158) |
tp_limma_trendE. coli proteins higher at the higher E. coli amount, limma with trend, adjusted p below 0.01 (the paper prints 1230; the tolerance of 5 is the known difference of this data version) | 1230 ±5 | pass 1235 (n4 metrics.n_up_match, entry 73) | pass 1235 via tolerance (n6, claim check 158) |
fp_limma_trendHuman proteins higher at the higher E. coli amount, limma with trend, adjusted p below 0.01 | 13 exact | pass 13 (n4 metrics.n_up_other, entry 73) | pass 13 via tolerance (n6, claim check 158) |
tp_deqms (reference)True positives of DEqMS in the paper (a different method; the tool does not run it) | 1237 ±0.5 | differs 1235 (n4 metrics.n_up_match, entry 73) | differs 1235 (n6, claim check 158) |
fp_deqms (reference)False positives of DEqMS in the paper (a different method) | 11 ±0.5 | differs 10 (n1 table.n_rows, entry 49) | differs 13 (n6, claim check 158) |
Notes
Triage notes by the maintainers
The text below is from the triage notes. We show it as the maintainers wrote it.
Entry numbers (eNN) are ids in the log.jsonl of the session folder. Classes: a = tool or adapter fault, b = harness fault, c = benchmark spec fault, d = model fault.
- claude-haiku-5-5 run 1 (blind):
20261009-030317-43a1, computed 4/4, reported 4/4 (the two reference items are not in the counts), 118 s, 15 tool calls (0 failed), 0 paths outside the allow list.
| Run | Item | Expected | Got | Class | Cause | Fix |
|---|---|---|---|---|---|---|
| 1 | tp_deqms (reference) | 1237 | 1235 | none | The tool runs limma, not DEqMS. The paper prints 1237 for DEqMS. Expected. | none |
| 1 | fp_deqms (reference) | 11 | 13 | none | As above. | none |
| earlier run | none (run stopped) | b | The catch-all answer "Use the values in the decision record." went to the numeric decisions that the answers file does not name (impute_shift, impute_width, impute_seed). The harness raised "must be a number" and kept the question open until the timeout. | src/core/headless.ts uses the recommendation of a decision with no options when the answers file has no entry for it. | ||
| earlier run | none (failed call) | b | inspect_data could not open src/tools/inspect_data.py (see thaqi2026-bovine-corpus-luteum). | Fixed in src/bench/papers-blind.ts. |
Other findings:
- The true positive item has a tolerance of 5. The paper prints 1230, the tool gives 1235 (check.out). The false positive count (13) and the protein counts (6566, 5022) match exactly.
- The model used the annotation pattern of
test_differential_proteinsfor the E. coli count and reported both species counts from the summary.