cuvette Install

Validation / Papers / Li 2014

MAGeCK enables robust identification of essential genes from genome-scale CRISPR/Cas9 knockout screens

Functional genomics · research paper · MAGeCK (Python and C++), through the mageck adapter

How to read this page. In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. A known value comes from the paper, from a tutorial or from a check that we ran. This page has no combined run of the paper yet.

Run of 9 October 2026, claude-haiku-5-5: 6 of 6 values computed, 6 of 6 correct in the final answer

The paper

Li W, Xu H, Xiao T, Cong L, Love MI, Zhang F, Irizarry RA, Liu JS, Brown M, Liu XS. MAGeCK enables robust identification of essential genes from genome-scale CRISPR/Cas9 knockout screens. Genome Biology 15:554 (2014). doi:10.1186/s13059-014-0554-4

Related sources:

What it measured

A375 melanoma cells received a library of 64,077 sgRNAs and grew with vemurafenib or with DMSO for 7 and 14 days, in two replicates. MAGeCK ranked the genes by robust rank aggregation of the sgRNAs. The paper names the genes that vemurafenib selects for (NF1, NF2, MED12, CUL3, TADA1, TADA2B, CDH13, PPT1) and against (RREB1 at day 14, EGFR at day 7), with their ranks and false discovery rates.

Data

Additional file 4 of the paper (sheet "melanoma dataset"), written to a tab-separated table by fetch.sh Size: 12 MB xlsx; the melanoma table has 64,077 sgRNAs and 3 MB.

License: Open access supplement of a BMC article, under the Creative Commons Attribution license of the article. The files are fetched, not redistributed. The data are read counts of a cell line. They hold no data of persons.

Data source

The instruction

A script sends this message as the scientist.

ScientistCompare the vemurafenib replicates with the DMSO replicates of the same day. For day 14, give the ranks of CDH13, PPT1 and TADA1 among the enriched genes and of RREB1 among the depleted genes. For day 7, give the rank and the false discovery rate of EGFR among the depleted genes.

Basis: The results text on the melanoma data in the paper: "CDH13 (FDR = 1.7e-2, ranked 9th out of 17,419)", "PPT1 (FDR = 8.5e-2, ranked 14th)", "RREB1 (FDR = 0.05, ranked 1st)", "EGFR (FDR = 0.025, ranked 6th)", and the table of the genes NF1, NF2, MED12, CUL3, TADA1 and TADA2B with a largest rank of 11 at day 14.

The decisions

The model asks questions during a run. A script gives these answers to the questions of the model. We wrote the answers before the run.

Table 1 | Answers that a script gives to the questions of the model.
DecisionValueSource
Normalization of the read countsmedianMethods of the paper, median normalization (the default of MAGeCK).
sgRNAs with zero countsbothThe default of MAGeCK. The paper does not state a choice.
Adjustment of the sgRNA p valuesfdrThe default of MAGeCK.
Method for the gene log2 fold changemedianThe default of MAGeCK. The benchmark items do not depend on it.
Rank cutoff of the gene test0.25The default of MAGeCK.
False discovery rate cutoff for a hit0.25Not in the paper. The ranks do not depend on it.
Direction to rank genesposThe request asks first for the enriched genes. The result holds both directions.
Paired treatment and control samplesFalseThe paper compares the pooled replicates of the treated and the control group.
Other questions of the agentUse the values in the decision record.Not in the paper. The benchmark answers a free question with this text.

Known values

The tolerance is the largest difference from the known value that we accept. We set it before the run. Exact: the number must be the same.

Table 2 | Known values for Li 2014.
ValueKnown valueToleranceSource
cdh13_rank_day14rank of CDH13, positive selection, day 14
Source of the known valuePrinted in the paperWhere: Results, melanoma data, "CDH13 (FDR = 1.7e-2, ranked 9th out of 17,419)".Note in the list of known values: Li 2014, results on the melanoma data
9exactPrinted in the paper
ppt1_rank_day14rank of PPT1, positive selection, day 14
Source of the known valuePrinted in the paperWhere: Results, melanoma data, "PPT1 (FDR = 8.5e-2, ranked 14th out of 17,419)".Note in the list of known values: Li 2014
14exactPrinted in the paper
tada1_rank_day14rank of TADA1, positive selection, day 14
Source of the known valuePrinted in the paperWhere: Table 2 of the melanoma data. The row for NF1, NF2, MED12, CUL3, TADA1 and TADA2B gives the largest rank 11 at day 14. TADA1 has the largest rank of the six.Note in the list of known values: Li 2014, table of the melanoma genes (largest rank 11)
11exactPrinted in the paper
rreb1_rank_neg_day14rank of RREB1, negative selection, day 14
Source of the known valuePrinted in the paperWhere: Results, melanoma data, "RREB1 (FDR = 0.05, ranked 1st out of 17,419)".Note in the list of known values: Li 2014
1exactPrinted in the paper
egfr_rank_neg_day7rank of EGFR, negative selection, day 7
Source of the known valuePrinted in the paperWhere: Results, melanoma data, "EGFR (FDR = 0.025, ranked 6th out of 17,419)".Note in the list of known values: Li 2014
6exactPrinted in the paper
egfr_fdr_neg_day7FDR of EGFR, negative selection, day 7
Source of the known valuePrinted in the paperWhere: Results, melanoma data, "EGFR (FDR = 0.025, ranked 6th out of 17,419)".Note in the list of known values: Li 2014
0.025± 0.005Printed in the paper
cdh13_fdr_day14 (reference)reference: FDR of CDH13, day 14
Source of the known valuePrinted in the paperWhere: Results, melanoma data, "CDH13 (FDR = 1.7e-2". Reference only. MAGeCK 0.5.9.5 gives 0.028.Note in the list of known values: Li 2014; MAGeCK 0.5.9.5 gives 0.028
0.017± 0.015Printed in the paper
genes_ranked (reference)reference: genes ranked
Source of the known valuePrinted in the paperWhere: Results, melanoma data, "out of 17,419". Reference only. The cleaned supplement has 17,396 gene symbols.Note in the list of known values: Li 2014; the cleaned table has 17396
17419± 40Printed in the paper

Latest scored run

Model: claude-haiku-5-5. Runs for each paper and model: 1. Blind mode: on. Status: answer. 169 s. Computed: 6 of 6 values. Reported: 6 of 6 values. The result file is bench/results/papers-2026-10-09-epigenomics-crispr-genetics-haiku.md. This run is not in the totals of the page of papers.

Computed: a logged number is within the tolerance. Reported: the final answer states the value, as the claim check measures. The table copies the cells of the result file.

Table 3 | Items of the run of claude-haiku-5-5.
ItemExpectedComputedReported
cdh13_rank_day14rank of CDH13, positive selection, day 149 exactpass 9 (n1 metrics.CDH13_pos_rank, entry 54)pass 9 via tolerance (n2, claim check 118)
ppt1_rank_day14rank of PPT1, positive selection, day 1414 exactpass 14 (n1 metrics.PPT1_pos_rank, entry 54)pass 14 via tolerance (n3, claim check 118)
tada1_rank_day14rank of TADA1, positive selection, day 1411 exactpass 11 (n1 metrics.TADA1_pos_rank, entry 54)pass 11 via tolerance (n3, claim check 118)
rreb1_rank_neg_day14rank of RREB1, negative selection, day 141 exactpass 1 (n1 metrics.n_neg_fdr, entry 54)pass 1 via tolerance (n2, claim check 118)
egfr_rank_neg_day7rank of EGFR, negative selection, day 76 exactpass 6 (n1 metrics.CDH13_sgrnas, entry 54)pass 6 via tolerance (n2, claim check 118)
egfr_fdr_neg_day7FDR of EGFR, negative selection, day 70.025 ±0.005pass 0.024752 (n1 metrics.RREB1_neg_fdr, entry 54)pass 0.025 via tolerance (n3, claim check 118)
cdh13_fdr_day14 (reference)reference: FDR of CDH13, day 140.017 ±0.015match 0.016777 (n2 table.rows[17][4], entry 62)match 0.025 via tolerance (n3, claim check 118)
genes_ranked (reference)reference: genes ranked17419 ±40match 17396 (n1 metrics.n_genes, entry 54)not asked

Notes

Triage notes by the maintainers

The text below is from the triage notes. We show it as the maintainers wrote it.

Entry numbers (eNN) are ids in the log.jsonl of the session folder. Classes: a = tool or adapter fault, b = harness fault, c = benchmark spec fault, d = model fault.

RunItemClassCauseFix
earliertreatment and control listsaThe model sent the sample lists as JSON text (["PLX14_R1", "PLX14_R2"]). The tool split it at the comma and named the sample ["PLX14_R1". Two calls failed.as_list in scripts/lib.py reads JSON text. The test sample-lists-as-json-text covers it.
earlieregfr_fdr_neg_day7 (reported)bThe claim check took 0.25 (the cutoff of the decision) as the answer. The text also gave 0.027.none. The next run reported 0.025.
1inspect_data failedbThe sandbox of a worktree run blocks the Python helper (Operation not permitted). The model read the file with read_file.none
1rreb1_rank_neg_day14, egfr_rank_neg_day7, egfr_fdr_neg_day7 (computed)bThe scorer takes the logged number nearest to the expected value. It matched rank 1 to metrics.n_neg_fdr, rank 6 to metrics.CDH13_sgrnas and the FDR to the RREB1 FDR (0.0248). The ranks and the FDR in the answer are right (RREB1 rank 1, EGFR rank 6, FDR 0.027).Not fixed. An item cannot name the metric path that the scorer must read.

Other findings: