cuvette Install

Validation / Papers / Nowicka 2019

CyTOF workflow: differential discovery in high-throughput high-dimensional cytometry datasets

Flow cytometry · research paper · FlowSOM, diffcyt and lme4 (R), through the cytof adapter

How to read this page. In this validation, a script plays the scientist. It gives the answers that we wrote before the run, from the methods of the paper. A known value comes from the paper, from a tutorial or from a check that we ran. This page has no combined run of the paper yet.

No run is scored for this paper yet.

The paper

Nowicka M, Krieg C, Crowell HL, Weber LM, Hartmann FJ, Guglietta S, Becher B, Levesque MP, Robinson MD. CyTOF workflow: differential discovery in high-throughput high-dimensional cytometry datasets. F1000Research 6:748, version 3 (2019). doi:10.12688/f1000research.11622.3

Related sources:

What it measured

The workflow analyses a mass cytometry experiment on peripheral blood cells of 8 healthy donors. Each donor has a Reference sample and a sample stimulated through the B cell receptor and the Fc receptor (BCR-XL). The paper clusters the cells with FlowSOM, merges the clusters into 8 cell types, counts the cells of each cell type in each sample, and tests each cell type for a difference in frequency with a binomial mixed model (GLMM) from the diffcyt package. The model with a random effect for the donor compares each donor with itself. The model without it ignores the pairs. The two models give different numbers of significant cell types.

Data

Bodenmiller_BCR_XL_flowSet of the Bioconductor package HDCytoData (ExperimentHub record EH2255), written to 16 FCS files by the benchmark fetch script Size: 16 FCS files, 29 MB. 172,791 cells, 24 protein markers plus four label channels..

License: The package HDCytoData is MIT. The data are from Bodenmiller et al. 2012 (Cytobank, public). No license text for the files was found, so the files are not committed. Healthy donors, no personal data.

Data source

The instruction

A script sends this message as the scientist.

ScientistI have mass cytometry data of blood cells from 8 donors, each unstimulated and stimulated through the B cell receptor. Count the cells of each cell type in each sample and test each cell type for a difference in frequency, with each donor compared with itself. How many cell types differ, and how many differ if the pairing is ignored?

Basis: The differential abundance section of the workflow. The first GLMM has a random effect for the sample. The second GLMM adds the donor. The workflow prints both result tables.

The decisions

The model asks questions during a run. A script gives these answers to the questions of the model. We wrote the answers before the run.

Table 1 | Answers that a script gives to the questions of the model.
DecisionValueSource
Research questionDoes the frequency of any cell type change after B cell receptor stimulation?The differential abundance section tests each of the 8 cell types between the BCR-XL and Reference samples.
Unit of replicationsamples from different donors, one file for each donor and conditionThe workflow has 16 samples from 8 patients, each unstimulated and stimulated.
Markers to cluster onCD3, CD45, CD4, CD20, CD33, CD123, CD14, IgM, HLA-DR, CD7The workflow clusters on the 10 lineage markers (cell surface markers) and tests the 14 signaling markers later.
Cofactor of the arcsinh transform5The workflow states the arcsinh transform with cofactor 5 for mass cytometry and notes that 150 has been promoted for flow cytometry.
Width of the FlowSOM grid10The workflow call sets xdim = 10.
Height of the FlowSOM grid10The workflow call sets ydim = 10.
Number of metaclusters20The workflow call sets maxK = 20. It merges these 20 metaclusters by hand into 8 cell types.
Random seed1234The workflow call sets seed = 1234.
Column that defines the populationslabelThe workflow tests the 8 merged cell types. The data package stores them in the channel population_id. The benchmark uses this channel because the hand merging of the 20 metaclusters cannot be repeated from the text.
Column of the sample table with the conditionconditionThe workflow design has the condition BCRXL against Reference.
Reference levelReferenceThe workflow tests BCRXL against the unstimulated samples.
Column that pairs the samples of one donorpatient_idThe second GLMM of the workflow adds a random effect for the patient.
Test for differences in frequencyglmmThe workflow tests differential abundance with the GLMM of diffcyt.
Adjusted p value limit0.05The workflow sets FDR_cutoff to 0.05.
Samples, cells or markers to excludenoneThe workflow uses all 16 samples and all cells.

Known values

The tolerance is the largest difference from the known value that we accept. We set it before the run. Exact: the number must be the same.

Table 2 | Known values for Nowicka 2019.
ValueKnown valueToleranceSource
cells_donor2_bcrxlCells in the file of donor 2 with BCR-XL
Source of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, section Data preparation. The printed n_cells output lists 2838, 2739, 16675, 16725 and 12 more values in the order of the sample table (donor 1 BCR-XL, donor 1 Reference, donor 2 BCR-XL, and so on).Check: check_bcrxl.py counts the events of each FCS file with flowio and gives 16675.Note in the list of known values: Nowicka 2019, printed n_cells; check_bcrxl.py
16675exactPrinted in the paper. Independent check: yes
cells_donor1_referenceCells in the file of donor 1 Reference
Source of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, section Data preparation. The second printed n_cells value is 2739.Check: check_bcrxl.py gives 2739.Note in the list of known values: Nowicka 2019, printed n_cells; check_bcrxl.py
2739exactPrinted in the paper. Independent check: yes
nk_percent_donor1_referencePercent of NK cells in the Reference file of donor 1
Source of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, differential abundance, topTable of the second GLMM with show_props. The NK cells row, column props_Ref1, prints 14.3.Check: check_bcrxl.py gives 14.3 (392 of 2739 cells).Note in the list of known values: Nowicka 2019, topTable props_Ref1; check_bcrxl.py
14.3± 0.05Printed in the paper. Independent check: yes
cd4_percent_donor1_referencePercent of CD4 T-cells in the Reference file of donor 1
Source of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, the same table. The CD4 T-cells row, column props_Ref1, prints 44.7.Check: check_bcrxl.py gives 44.7 (1225 of 2739 cells).Note in the list of known values: Nowicka 2019, topTable props_Ref1; check_bcrxl.py
44.7± 0.05Printed in the paper. Independent check: yes
p_nk_pairedp value of the NK cells, paired model
Source of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, topTable of the second GLMM. The NK cells row prints p_val 4.5e-13 and p_adj 3.6e-12.Check: check_glmm.R fits glmer from lme4 directly, with the model cbind(y, n - y) ~ condition + (1 | patient) + (1 | sample) and the binomial family. It gives p 4.474e-13 (check_glmm.out).Note in the list of known values: Nowicka 2019, topTable; check_glmm.R gives 4.474e-13
4.5e-13± 1e-14Printed in the paper. Independent check: yes
p_bcells_igm_neg_pairedp value of the B-cells IgM-, paired model
Source of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, topTable of the second GLMM. The B-cells IgM- row prints p_val 2.2e-11.Check: check_glmm.R gives 2.202e-11.Note in the list of known values: Nowicka 2019, topTable; check_glmm.R gives 2.202e-11
2.2e-11± 1e-12Printed in the paper. Independent check: yes
padj_cd8_pairedAdjusted p value of the CD8 T-cells, paired model
Source of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, topTable of the second GLMM. The CD8 T-cells row prints p_adj 0.0019.Check: check_glmm.R gives 0.001914 after the Benjamini-Hochberg adjustment of the 8 p values.Note in the list of known values: Nowicka 2019, topTable; check_glmm.R gives 0.001914
0.0019± 0.0001Printed in the paper. Independent check: yes
n_different_pairedCell types with adjusted p at or below 0.05, paired model
Source of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, after the second topTable. "Significant clusters at FDR cutoff: FALSE = 2, TRUE = 6".Check: check_glmm.R gives 6 populations at the adjusted p value 0.05 or below.Note in the list of known values: Nowicka 2019, TRUE = 6; check_glmm.R
6exactPrinted in the paper. Independent check: yes
n_different_unpairedCell types with adjusted p at or below 0.05, model without the pairing
Source of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, first GLMM, without the patient. "Significant clusters at FDR cutoff: FALSE = 5, TRUE = 3".Check: check_glmm.R gives 3 populations at the adjusted p value 0.05 or below.Note in the list of known values: Nowicka 2019, TRUE = 3; check_glmm.R
3exactPrinted in the paper. Independent check: yes
p_bcells_igm_pos_unpairedp value of the B-cells IgM+, model without the pairing
Source of the known valuePrinted in the paper. Independent check: yesWhere: Workflow, rowData of the first GLMM. The B-cells IgM+ row prints p_val 0.0134820784148451.Check: check_glmm.R gives 0.01348.Note in the list of known values: Nowicka 2019, rowData of the first GLMM; check_glmm.R gives 0.01348
0.0135± 0.0005Printed in the paper. Independent check: yes

Latest scored run

No run is scored for this paper yet.

Notes

Triage notes by the maintainers

The text below is from the triage notes. We show it as the maintainers wrote it.

Classes: a = tool or adapter fault, b = harness fault, c = benchmark spec fault, d = model fault.

ModelItem or callExpectedGotClassCauseFix
haikuinspect_fcs, cluster_cells, differential_abundance in run 1 (16 calls)ok"there is no package called flowCore"aThe R packages were in R-library-microcyto. Blind mode lets child processes read only tools/R-library.The packages are in tools/R-library. The scripts and the install text name that folder. Run 2 has no failed call.
haikuinspect_data in run 1ok"can't open file src/tools/inspect_data.py"bBlind mode hides the repository from child processes. The built-in tool could not start.None. The model then used the adapter tools.
specthe 20 FlowSOM metaclustersnot scorednot scoredcThe paper merges the 20 metaclusters into 8 cell types by hand (a figure). The merge cannot be repeated from the text.The benchmark scores the published cell types (channel population_id). case.yaml and paper.md say so.

Commits in this branch for this paper: see git log -- bench/papers/nowicka2019-cytof-bcrxl.