Validation / Papers / Nowicka 2019
CyTOF workflow: differential discovery in high-throughput high-dimensional cytometry datasets
How to read this page. In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. A known value comes from the paper, from a tutorial or from a check that we ran. This page has no combined run of the paper yet.
No run is scored for this paper yet.
The paper
Nowicka M, Krieg C, Crowell HL, Weber LM, Hartmann FJ, Guglietta S, Becher B, Levesque MP, Robinson MD. CyTOF workflow: differential discovery in high-throughput high-dimensional cytometry datasets. F1000Research 6:748, version 3 (2019). doi:10.12688/f1000research.11622.3
Related sources:
- Weber LM, Nowicka M, Soneson C, Robinson MD. diffcyt: Differential discovery in high-dimensional cytometry via high-resolution clustering. Commun Biol 2:183 (2019). Source of the differential abundance methods. doi:10.1038/s42003-019-0415-5
- Bodenmiller B et al. Multiplexed mass cytometry profiling of cellular states perturbed by small-molecule regulators. Nat Biotechnol 30:858-867 (2012). Source of the mass cytometry data. doi:10.1038/nbt.2317
What it measured
The workflow analyses a mass cytometry experiment on peripheral blood cells of 8 healthy donors. Each donor has a Reference sample and a sample stimulated through the B cell receptor and the Fc receptor (BCR-XL). The paper clusters the cells with FlowSOM, merges the clusters into 8 cell types, counts the cells of each cell type in each sample, and tests each cell type for a difference in frequency with a binomial mixed model (GLMM) from the diffcyt package. The model with a random effect for the donor compares each donor with itself. The model without it ignores the pairs. The two models give different numbers of significant cell types.
Data
Bodenmiller_BCR_XL_flowSet of the Bioconductor package HDCytoData (ExperimentHub record EH2255), written to 16 FCS files by the benchmark fetch script Size: 16 FCS files, 29 MB. 172,791 cells, 24 protein markers plus four label channels..
License: The package HDCytoData is MIT. The data are from Bodenmiller et al. 2012 (Cytobank, public). No license text for the files was found, so the files are not committed. Healthy donors, no personal data.
The instruction
A script sends this message as the scientist.
Basis: The differential abundance section of the workflow. The first GLMM has a random effect for the sample. The second GLMM adds the donor. The workflow prints both result tables.
The decisions
The model asks questions during a run. A script gives these answers to the questions of the model. We wrote the answers before the run.
| Decision | Value | Source |
|---|---|---|
| Research question | Does the frequency of any cell type change after B cell receptor stimulation? | The differential abundance section tests each of the 8 cell types between the BCR-XL and Reference samples. |
| Unit of replication | samples from different donors, one file for each donor and condition | The workflow has 16 samples from 8 patients, each unstimulated and stimulated. |
| Markers to cluster on | CD3, CD45, CD4, CD20, CD33, CD123, CD14, IgM, HLA-DR, CD7 | The workflow clusters on the 10 lineage markers (cell surface markers) and tests the 14 signaling markers later. |
| Cofactor of the arcsinh transform | 5 | The workflow states the arcsinh transform with cofactor 5 for mass cytometry and notes that 150 has been promoted for flow cytometry. |
| Width of the FlowSOM grid | 10 | The workflow call sets xdim = 10. |
| Height of the FlowSOM grid | 10 | The workflow call sets ydim = 10. |
| Number of metaclusters | 20 | The workflow call sets maxK = 20. It merges these 20 metaclusters by hand into 8 cell types. |
| Random seed | 1234 | The workflow call sets seed = 1234. |
| Column that defines the populations | label | The workflow tests the 8 merged cell types. The data package stores them in the channel population_id. The benchmark uses this channel because the hand merging of the 20 metaclusters cannot be repeated from the text. |
| Column of the sample table with the condition | condition | The workflow design has the condition BCRXL against Reference. |
| Reference level | Reference | The workflow tests BCRXL against the unstimulated samples. |
| Column that pairs the samples of one donor | patient_id | The second GLMM of the workflow adds a random effect for the patient. |
| Test for differences in frequency | glmm | The workflow tests differential abundance with the GLMM of diffcyt. |
| Adjusted p value limit | 0.05 | The workflow sets FDR_cutoff to 0.05. |
| Samples, cells or markers to exclude | none | The workflow uses all 16 samples and all cells. |
Known values
The tolerance is the largest difference from the known value that we accept. We set it before the run. Exact: the number must be the same.
| Value | Known value | Tolerance | Source |
|---|---|---|---|
cells_donor2_bcrxlCells in the file of donor 2 with BCR-XLSource of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, section Data preparation. The printed n_cells output lists 2838, 2739, 16675, 16725 and 12 more values in the order of the sample table (donor 1 BCR-XL, donor 1 Reference, donor 2 BCR-XL, and so on).Check: check_bcrxl.py counts the events of each FCS file with flowio and gives 16675.Note in the list of known values: Nowicka 2019, printed n_cells; check_bcrxl.py | 16675 | exact | Printed in the paper. Independent check: yes |
cells_donor1_referenceCells in the file of donor 1 ReferenceSource of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, section Data preparation. The second printed n_cells value is 2739.Check: check_bcrxl.py gives 2739.Note in the list of known values: Nowicka 2019, printed n_cells; check_bcrxl.py | 2739 | exact | Printed in the paper. Independent check: yes |
nk_percent_donor1_referencePercent of NK cells in the Reference file of donor 1Source of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, differential abundance, topTable of the second GLMM with show_props. The NK cells row, column props_Ref1, prints 14.3.Check: check_bcrxl.py gives 14.3 (392 of 2739 cells).Note in the list of known values: Nowicka 2019, topTable props_Ref1; check_bcrxl.py | 14.3 | ± 0.05 | Printed in the paper. Independent check: yes |
cd4_percent_donor1_referencePercent of CD4 T-cells in the Reference file of donor 1Source of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, the same table. The CD4 T-cells row, column props_Ref1, prints 44.7.Check: check_bcrxl.py gives 44.7 (1225 of 2739 cells).Note in the list of known values: Nowicka 2019, topTable props_Ref1; check_bcrxl.py | 44.7 | ± 0.05 | Printed in the paper. Independent check: yes |
p_nk_pairedp value of the NK cells, paired modelSource of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, topTable of the second GLMM. The NK cells row prints p_val 4.5e-13 and p_adj 3.6e-12.Check: check_glmm.R fits glmer from lme4 directly, with the model cbind(y, n - y) ~ condition + (1 | patient) + (1 | sample) and the binomial family. It gives p 4.474e-13 (check_glmm.out).Note in the list of known values: Nowicka 2019, topTable; check_glmm.R gives 4.474e-13 | 4.5e-13 | ± 1e-14 | Printed in the paper. Independent check: yes |
p_bcells_igm_neg_pairedp value of the B-cells IgM-, paired modelSource of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, topTable of the second GLMM. The B-cells IgM- row prints p_val 2.2e-11.Check: check_glmm.R gives 2.202e-11.Note in the list of known values: Nowicka 2019, topTable; check_glmm.R gives 2.202e-11 | 2.2e-11 | ± 1e-12 | Printed in the paper. Independent check: yes |
padj_cd8_pairedAdjusted p value of the CD8 T-cells, paired modelSource of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, topTable of the second GLMM. The CD8 T-cells row prints p_adj 0.0019.Check: check_glmm.R gives 0.001914 after the Benjamini-Hochberg adjustment of the 8 p values.Note in the list of known values: Nowicka 2019, topTable; check_glmm.R gives 0.001914 | 0.0019 | ± 0.0001 | Printed in the paper. Independent check: yes |
n_different_pairedCell types with adjusted p at or below 0.05, paired modelSource of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, after the second topTable. "Significant clusters at FDR cutoff: FALSE = 2, TRUE = 6".Check: check_glmm.R gives 6 populations at the adjusted p value 0.05 or below.Note in the list of known values: Nowicka 2019, TRUE = 6; check_glmm.R | 6 | exact | Printed in the paper. Independent check: yes |
n_different_unpairedCell types with adjusted p at or below 0.05, model without the pairingSource of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, first GLMM, without the patient. "Significant clusters at FDR cutoff: FALSE = 5, TRUE = 3".Check: check_glmm.R gives 3 populations at the adjusted p value 0.05 or below.Note in the list of known values: Nowicka 2019, TRUE = 3; check_glmm.R | 3 | exact | Printed in the paper. Independent check: yes |
p_bcells_igm_pos_unpairedp value of the B-cells IgM+, model without the pairingSource of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, rowData of the first GLMM. The B-cells IgM+ row prints p_val 0.0134820784148451.Check: check_glmm.R gives 0.01348.Note in the list of known values: Nowicka 2019, rowData of the first GLMM; check_glmm.R gives 0.01348 | 0.0135 | ± 0.0005 | Printed in the paper. Independent check: yes |
Latest scored run
No run is scored for this paper yet.
Notes
Triage notes by the maintainers
The text below is from the triage notes. We show it as the maintainers wrote it.
Classes: a = tool or adapter fault, b = harness fault, c = benchmark spec fault, d = model fault.
- claude:claude-haiku-5-5, run 1:
20261009-032305-87f9 - claude:claude-haiku-5-5, run 2 (final):
20261009-033119-d666
| Model | Item or call | Expected | Got | Class | Cause | Fix |
|---|---|---|---|---|---|---|
| haiku | inspect_fcs, cluster_cells, differential_abundance in run 1 (16 calls) | ok | "there is no package called flowCore" | a | The R packages were in R-library-microcyto. Blind mode lets child processes read only tools/R-library. | The packages are in tools/R-library. The scripts and the install text name that folder. Run 2 has no failed call. |
| haiku | inspect_data in run 1 | ok | "can't open file src/tools/inspect_data.py" | b | Blind mode hides the repository from child processes. The built-in tool could not start. | None. The model then used the adapter tools. |
| spec | the 20 FlowSOM metaclusters | not scored | not scored | c | The paper merges the 20 metaclusters into 8 cell types by hand (a figure). The merge cannot be repeated from the text. | The benchmark scores the published cell types (channel population_id). case.yaml and paper.md say so. |
Commits in this branch for this paper: see git log -- bench/papers/nowicka2019-cytof-bcrxl.