cuvette Install

Validation / Papers / Zhu 2020

DEqMS: a method for accurate variance estimation in differential protein expression analysis

Quantitative proteomics · research paper · limma (R), through the limma-proteomics adapter

How to read this page. In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. A known value comes from the paper, from a tutorial or from a check that we ran. This page has no combined run of the paper yet.

Run of 9 October 2026, claude-haiku-5-5: 4 of 4 values computed, 4 of 4 correct in the final answer

The paper

Zhu Y, Orre LM, Zhou Tran Y, Mermelekas G, Johansson HJ, Malyutina A, Anders S, Lehtiö J. DEqMS: a method for accurate variance estimation in differential protein expression analysis. Molecular and Cellular Proteomics 19(6):1047-1057 (2020). doi:10.1074/mcp.TIR119.001646

Related sources:

What it measured

The authors compared tools for differential protein abundance on a benchmark of human proteins with E. coli proteins added in a 1 to 3 ratio. A tool finds a true positive if an E. coli protein is more abundant in the sample with more E. coli. Human proteins that are significant are false positives. They compare limma, limma with trend and DEqMS.

Data

PRIDE PXD000279, file proteomebenchmark.zip (38 MB), file proteinGroups.txt of the zip (28 MB). fetch.sh unpacks it. Size: 28 MB table with 6694 protein groups..

License: PRIDE data are open under the ProteomeXchange data policy. The paper is CC BY 4.0.

Data source

The instruction

A script sends this message as the scientist.

ScientistRemove Reverse and Contaminant rows. Keep proteins with at least two LFQ values in each condition. Test H against L with limma (trend) and adjusted p below 0.01. Count the E. coli and the human proteins that are more abundant in H.

Basis: Results (Benchmarking of DEqMS) and Methods of the paper. The paper gives log2 LFQ values as the input of all methods and does not state a normalization.

The decisions

The model asks questions during a run. A script gives these answers to the questions of the model. We wrote the answers before the run.

Table 1 | Answers that a script gives to the questions of the model.
DecisionValueSource
Flag columns that remove a rowReverse, ContaminantNot in the paper. Reverse and Contaminant give the 6566 proteins that the paper prints (check.R).
Peptides that a protein needs0The paper gives no peptide filter.
Samples that you excludenoneThe paper excludes no sample.
Valid values that a protein needs in a group2Results. Proteins with quantification values in at least two samples per condition.
In which groups the valid values must be presenteach groupResults. At least two samples per condition.
Normalization of the log2 intensitiesnoneResults. The log2 LFQ intensities are the input, with no other normalization.
Imputation of missing valuesnoneThe paper uses no imputation.
Variance model of the moderated t testtrendResults. limma with trend = TRUE.
False discovery rate (FDR) level0.01Results. Adjusted p below 0.01, Benjamini-Hochberg.
Smallest log2 fold change0The paper uses no fold change cutoff.
Other questions of the agentUse the values in the decision record.Not in the paper. The benchmark answers each free question with this text.

Known values

The tolerance is the largest difference from the known value that we accept. We set it before the run. Exact: the number must be the same.

Table 2 | Known values for Zhu 2020.
ValueKnown valueToleranceSource
n_after_flagsProtein groups after the removal of Reverse and Contaminant rows
Source of the known valuePrinted in the paperWhere: Results. 6566 proteins, of which 1902 are E. coli proteins and 4664 are human.Check: check.R gives 6566 and 1902.Note in the list of known values: published, Results of the paper
6566exactPrinted in the paper
n_validProteins with at least two values in each condition
Source of the known valuePrinted in the paperWhere: Results. 5022 proteins with quantification values in at least two samples in each condition.Check: check.R gives 5022.Note in the list of known values: published, Results of the paper
5022exactPrinted in the paper
tp_limma_trendE. coli proteins higher at the higher E. coli amount, limma with trend, adjusted p below 0.01 (the paper prints 1230; the tolerance of 5 is the known difference of this data version)
Source of the known valuePrinted in the paperWhere: Results. The second best method was limma (trend = TRUE) with 1230 true positives and 13 false positives. The check gives 1235 for the proteins that are more abundant in H, 5 more than the paper.Check: check.R gives 1235.Note in the list of known values: published, Results of the paper
1230± 5Printed in the paper
fp_limma_trendHuman proteins higher at the higher E. coli amount, limma with trend, adjusted p below 0.01
Source of the known valuePrinted in the paperWhere: Results. 13 false positives for limma (trend = TRUE).Check: check.R gives 13.Note in the list of known values: published, Results of the paper
13exactPrinted in the paper
tp_deqms (reference)True positives of DEqMS in the paper (a different method; the tool does not run it)
Source of the known valuePrinted in the paperWhere: Results. DEqMS reported 1237 true differentially expressed E. coli proteins and 11 human proteins.Check: The tool does not run DEqMS. A DEqMS run on this table gives 1218 and 7, so this item does not reproduce.Note in the list of known values: published, Results of the paper
1237± 0.5Printed in the paper
fp_deqms (reference)False positives of DEqMS in the paper (a different method)
Source of the known valuePrinted in the paperWhere: Results. 11 false positives for DEqMS.Check: Does not reproduce (7).Note in the list of known values: published, Results of the paper
11± 0.5Printed in the paper

Latest scored run

Model: claude-haiku-5-5. Runs for each paper and model: 1. Blind mode: on. Status: answer. 118 s. Computed: 4 of 4 values. Reported: 4 of 4 values. The result file is bench/results/papers-2026-10-09-multiomics-haiku.md. This run is not in the totals of the page of papers.

Computed: a logged number is within the tolerance. Reported: the final answer states the value, as the claim check measures. The table copies the cells of the result file.

Table 3 | Items of the run of claude-haiku-5-5.
ItemExpectedComputedReported
n_after_flagsProtein groups after the removal of Reverse and Contaminant rows6566 exactpass 6566 (n1 metrics.n_after_flags, entry 49)pass 6566 via tolerance (n1, claim check 158)
n_validProteins with at least two values in each condition5022 exactpass 5022 (n1 metrics.n_kept, entry 49)pass 5022 via tolerance (n6, claim check 158)
tp_limma_trendE. coli proteins higher at the higher E. coli amount, limma with trend, adjusted p below 0.01 (the paper prints 1230; the tolerance of 5 is the known difference of this data version)1230 ±5pass 1235 (n4 metrics.n_up_match, entry 73)pass 1235 via tolerance (n6, claim check 158)
fp_limma_trendHuman proteins higher at the higher E. coli amount, limma with trend, adjusted p below 0.0113 exactpass 13 (n4 metrics.n_up_other, entry 73)pass 13 via tolerance (n6, claim check 158)
tp_deqms (reference)True positives of DEqMS in the paper (a different method; the tool does not run it)1237 ±0.5differs 1235 (n4 metrics.n_up_match, entry 73)differs 1235 (n6, claim check 158)
fp_deqms (reference)False positives of DEqMS in the paper (a different method)11 ±0.5differs 10 (n1 table.n_rows, entry 49)differs 13 (n6, claim check 158)

Notes

Triage notes by the maintainers

The text below is from the triage notes. We show it as the maintainers wrote it.

Entry numbers (eNN) are ids in the log.jsonl of the session folder. Classes: a = tool or adapter fault, b = harness fault, c = benchmark spec fault, d = model fault.

RunItemExpectedGotClassCauseFix
1tp_deqms (reference)12371235noneThe tool runs limma, not DEqMS. The paper prints 1237 for DEqMS. Expected.none
1fp_deqms (reference)1113noneAs above.none
earlier runnone (run stopped)bThe catch-all answer "Use the values in the decision record." went to the numeric decisions that the answers file does not name (impute_shift, impute_width, impute_seed). The harness raised "must be a number" and kept the question open until the timeout.src/core/headless.ts uses the recommendation of a decision with no options when the answers file has no entry for it.
earlier runnone (failed call)binspect_data could not open src/tools/inspect_data.py (see thaqi2026-bovine-corpus-luteum).Fixed in src/bench/papers-blind.ts.

Other findings: