All benchmarks

Cyto Batch Integration

in development

Benchmarking of batch integration algorithms for cytometry data.

21 methods
5 control methods
3 datasets
8 metrics
1 release
Task repository MIT v0.0.1-rc1

Cytometry is a non-sequencing single cell profiling technique commonly used in clinical studies. It is very sensitive to batch effects, which can lead to biases in the interpretation of the result. Batch integration algorithms are often used to mitigate this effect.

In this project, we are building a pipeline for reproducible and continuous benchmarking of batch integration algorithms for cytometry data. As input, methods require cleaned and normalised (using arc-sinh or logicle transformation) data with multiple batches, cell type labels, and biological subjects, with paired samples from a subject profiled across multiple batches. The batch integrated output must be an integrated marker by cell matrix stored in Anndata format. All markers in the input data must be returned, regardless of whether they were integrated or not. This output is then evaluated using metrics that assess how well the batch effects were removed and how much biological signals were preserved.

Contributors

  • Luca Leomazzi
    authormaintainer
  • Givanna Putri
    authormaintainer
  • Robrecht Cannoodt
    author
  • Katrien Quintelier
    contributor
  • Sofie Van Gassen
    contributor

Leaderboard

Methods ranked by scaled overall mean. Each cell encodes a score from 0 to 1 by size and intensity.

QC: Normalisation Visualisation 6 plots

Per metric: points placed by control-anchored scaled score (x); dashed lines mark scaled 0 and 1 (worst/best control); the lower axis shows the raw score. Points beyond [-0.2, 1.2] are clamped to the edge as triangles. Hover a dot or line to highlight it and read details.

methodcontrol
  • Average Batch R-squared Cell Typelower better
    Perfect IntegrationPerfect IntegrationShuffle IntegrationShuffle IntegrationShuffle Integration —…Shuffle Integration — within cell typeBatchadjust with one …Batchadjust with one controlCytoNorm (no-controls…CytoNorm (no-controls, to-middle)Batchadjust with all …Batchadjust with all controlscyCombine (no-control…cyCombine (no-controls, to-middle)cyCombine (all-contro…cyCombine (all-controls, to-goal)HarmonypyHarmonypySeurat RPCA (to-goal)Seurat RPCA (to-goal)cyCombine (all-contro…cyCombine (all-controls, to-middle)cyCombine (one-contro…cyCombine (one-control, to-goal)cyCombine (one-contro…cyCombine (one-control, to-middle)CytoNorm (one-control…CytoNorm (one-control, to-goal)Limma removeBatchEffe…Limma removeBatchEffectCombatCombatCytoNorm (one-control…CytoNorm (one-control, to-middle)CytoNorm (no-controls…CytoNorm (no-controls, to-goal)CytoNorm (all-control…CytoNorm (all-controls, to-goal)GaussNormGaussNormCytoNorm (all-control…CytoNorm (all-controls, to-middle)cyCombine (no-control…cyCombine (no-controls, to-goal)CytoVICytoVIShuffle Integration —…Shuffle Integration — within batchesSeurat RPCA (to-middl…Seurat RPCA (to-middle)No IntegrationNo Integration0.0370.0280.0189.2e-3000.250.50.751rawscaled
  • Average Batch R-squared Globallower better
    Perfect IntegrationPerfect IntegrationShuffle Integration —…Shuffle Integration — within cell typeShuffle IntegrationShuffle IntegrationCytoNorm (no-controls…CytoNorm (no-controls, to-middle)Batchadjust with one …Batchadjust with one controlBatchadjust with all …Batchadjust with all controlscyCombine (no-control…cyCombine (no-controls, to-middle)HarmonypyHarmonypycyCombine (all-contro…cyCombine (all-controls, to-goal)cyCombine (all-contro…cyCombine (all-controls, to-middle)Seurat RPCA (to-goal)Seurat RPCA (to-goal)cyCombine (one-contro…cyCombine (one-control, to-goal)cyCombine (one-contro…cyCombine (one-control, to-middle)CytoNorm (one-control…CytoNorm (one-control, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-middle)Limma removeBatchEffe…Limma removeBatchEffectCytoNorm (all-control…CytoNorm (all-controls, to-goal)CombatCombatCytoNorm (no-controls…CytoNorm (no-controls, to-goal)GaussNormGaussNormCytoNorm (all-control…CytoNorm (all-controls, to-middle)cyCombine (no-control…cyCombine (no-controls, to-goal)CytoVICytoVISeurat RPCA (to-middl…Seurat RPCA (to-middle)Shuffle Integration —…Shuffle Integration — within batchesNo IntegrationNo Integration0.0340.0250.0178.5e-3000.250.50.751rawscaled
  • EMD Mean CT Horizontallower better
    Perfect IntegrationPerfect IntegrationShuffle Integration —…Shuffle Integration — within cell typeShuffle IntegrationShuffle IntegrationSeurat RPCA (to-goal)Seurat RPCA (to-goal)CytoNorm (no-controls…CytoNorm (no-controls, to-middle)cyCombine (all-contro…cyCombine (all-controls, to-goal)cyCombine (no-control…cyCombine (no-controls, to-middle)cyCombine (all-contro…cyCombine (all-controls, to-middle)cyCombine (one-contro…cyCombine (one-control, to-goal)cyCombine (one-contro…cyCombine (one-control, to-middle)Seurat RPCA (to-middl…Seurat RPCA (to-middle)CytoNorm (no-controls…CytoNorm (no-controls, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-goal)cyCombine (no-control…cyCombine (no-controls, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-middle)CytoNorm (all-control…CytoNorm (all-controls, to-goal)Batchadjust with all …Batchadjust with all controlsBatchadjust with one …Batchadjust with one controlCytoNorm (all-control…CytoNorm (all-controls, to-middle)HarmonypyHarmonypyGaussNormGaussNormCombatCombatShuffle Integration —…Shuffle Integration — within batchesLimma removeBatchEffe…Limma removeBatchEffectNo IntegrationNo IntegrationCytoVICytoVI0.340.2550.170.085000.250.50.751rawscaled
  • EMD Mean CT Verticallower better
    Shuffle Integration —…Shuffle Integration — within cell typeShuffle IntegrationShuffle IntegrationSeurat RPCA (to-middl…Seurat RPCA (to-middle)Seurat RPCA (to-goal)Seurat RPCA (to-goal)Perfect IntegrationPerfect IntegrationCytoVICytoVIShuffle Integration —…Shuffle Integration — within batchesBatchadjust with one …Batchadjust with one controlCytoNorm (no-controls…CytoNorm (no-controls, to-goal)Batchadjust with all …Batchadjust with all controlsCytoNorm (no-controls…CytoNorm (no-controls, to-middle)cyCombine (no-control…cyCombine (no-controls, to-goal)cyCombine (all-contro…cyCombine (all-controls, to-goal)cyCombine (no-control…cyCombine (no-controls, to-middle)cyCombine (one-contro…cyCombine (one-control, to-goal)cyCombine (all-contro…cyCombine (all-controls, to-middle)cyCombine (one-contro…cyCombine (one-control, to-middle)CytoNorm (one-control…CytoNorm (one-control, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-middle)CytoNorm (all-control…CytoNorm (all-controls, to-goal)HarmonypyHarmonypyCytoNorm (all-control…CytoNorm (all-controls, to-middle)GaussNormGaussNormCombatCombatLimma removeBatchEffe…Limma removeBatchEffectNo IntegrationNo Integration0.3050.2330.1610.0890.01700.250.50.751rawscaled
  • FlowSOM Mean Mapping Similarityhigher better
    Perfect IntegrationPerfect IntegrationShuffle Integration —…Shuffle Integration — within cell typeSeurat RPCA (to-goal)Seurat RPCA (to-goal)CombatCombatBatchadjust with one …Batchadjust with one controlcyCombine (all-contro…cyCombine (all-controls, to-middle)cyCombine (no-control…cyCombine (no-controls, to-middle)Seurat RPCA (to-middl…Seurat RPCA (to-middle)cyCombine (one-contro…cyCombine (one-control, to-middle)cyCombine (all-contro…cyCombine (all-controls, to-goal)cyCombine (one-contro…cyCombine (one-control, to-goal)Batchadjust with all …Batchadjust with all controlsGaussNormGaussNormcyCombine (no-control…cyCombine (no-controls, to-goal)HarmonypyHarmonypyLimma removeBatchEffe…Limma removeBatchEffectCytoNorm (one-control…CytoNorm (one-control, to-goal)CytoNorm (no-controls…CytoNorm (no-controls, to-middle)CytoNorm (one-control…CytoNorm (one-control, to-middle)CytoNorm (no-controls…CytoNorm (no-controls, to-goal)No IntegrationNo IntegrationCytoNorm (all-control…CytoNorm (all-controls, to-goal)CytoNorm (all-control…CytoNorm (all-controls, to-middle)CytoVICytoVIShuffle IntegrationShuffle IntegrationShuffle Integration —…Shuffle Integration — within batches81.50386.12790.75195.37610000.250.50.751rawscaled
  • Ratio of inconsistent peakslower better
    No IntegrationNo IntegrationPerfect IntegrationPerfect IntegrationLimma removeBatchEffe…Limma removeBatchEffectCombatCombatBatchadjust with all …Batchadjust with all controlsBatchadjust with one …Batchadjust with one controlHarmonypyHarmonypySeurat RPCA (to-middl…Seurat RPCA (to-middle)Seurat RPCA (to-goal)Seurat RPCA (to-goal)cyCombine (all-contro…cyCombine (all-controls, to-middle)cyCombine (no-control…cyCombine (no-controls, to-middle)GaussNormGaussNormcyCombine (all-contro…cyCombine (all-controls, to-goal)cyCombine (one-contro…cyCombine (one-control, to-middle)cyCombine (one-contro…cyCombine (one-control, to-goal)cyCombine (no-control…cyCombine (no-controls, to-goal)Shuffle Integration —…Shuffle Integration — within cell typeCytoNorm (all-control…CytoNorm (all-controls, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-goal)CytoNorm (no-controls…CytoNorm (no-controls, to-goal)CytoNorm (no-controls…CytoNorm (no-controls, to-middle)Shuffle IntegrationShuffle IntegrationCytoNorm (one-control…CytoNorm (one-control, to-middle)CytoNorm (all-control…CytoNorm (all-controls, to-middle)Shuffle Integration —…Shuffle Integration — within batchesCytoVICytoVI0.2730.2050.1370.068000.250.50.751rawscaled
QC: Indicator table 16 errors3 warnings

Automated checks on the benchmark run and its results: missing values, score scaling, metric ranges and similar. Errors are high-severity issues that usually need a maintainer's attention; warnings are lower-severity signals. Findings that are expected for this task are listed separately as silenced.

16 high-severity issues need review. 418 of 437 checks passed.

  • error Raw results Method 'batchadjust_one_control' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: batchadjust_one_control Number of results: 15 Expected number of results: 24 Percentage missing: 38%

  • error Raw results Method 'batchadjust_all_controls' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: batchadjust_all_controls Number of results: 15 Expected number of results: 24 Percentage missing: 38%

  • error Raw results Method 'rpca_to_goal' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: rpca_to_goal Number of results: 15 Expected number of results: 24 Percentage missing: 38%

  • error Raw results Method 'rpca_to_mid' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: rpca_to_mid Number of results: 15 Expected number of results: 24 Percentage missing: 38%

  • error Raw results Metric 'emd_mean_ct_vert' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Metric: emd_mean_ct_vert Number of results: 48 Expected number of results: 78 Percentage missing: 38%

  • error Raw results Metric 'iLisi' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Metric: iLisi Number of results: 0 Expected number of results: 78 Percentage missing: 100%

  • error Raw results Metric 'cLisi' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Metric: cLisi Number of results: 0 Expected number of results: 78 Percentage missing: 100%

  • error Raw results Method 'batchadjust_one_control' % failed

    Percentage of failed processes should be less than 10% Task: cyto_batch_integration Method: batchadjust_one_control Succeeded processes: 2 Attempted processes: 3 Percentage failed: 33%

  • error Raw results Method 'batchadjust_all_controls' % failed

    Percentage of failed processes should be less than 10% Task: cyto_batch_integration Method: batchadjust_all_controls Succeeded processes: 2 Attempted processes: 3 Percentage failed: 33%

  • error Raw results Method 'rpca_to_goal' % failed

    Percentage of failed processes should be less than 10% Task: cyto_batch_integration Method: rpca_to_goal Succeeded processes: 2 Attempted processes: 3 Percentage failed: 33%

  • error Raw results Method 'rpca_to_mid' % failed

    Percentage of failed processes should be less than 10% Task: cyto_batch_integration Method: rpca_to_mid Succeeded processes: 2 Attempted processes: 3 Percentage failed: 33%

  • error Raw results Metric 'emd_mean_ct_vert' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: cyto_batch_integration Metric: emd_mean_ct_vert Control method scores: 10 Expected control method scores: 15 Percentage succeeded: 67%

  • error Raw results Metric 'iLisi' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: cyto_batch_integration Metric: iLisi Control method scores: 0 Expected control method scores: 15 Percentage succeeded: 0%

  • error Raw results Metric 'cLisi' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: cyto_batch_integration Metric: cLisi Control method scores: 0 Expected control method scores: 15 Percentage succeeded: 0%

  • error Scaling Metric 'emd_mean_ct_horiz' % outside range

    Percentage of scaled scores outside control range should be less than 10% Task: cyto_batch_integration Metric: emd_mean_ct_horiz Inside range: 0 Scaled scores: 74 Percentage outside: 100%

  • error Scaling Metric 'emd_mean_ct_vert' % outside range

    Percentage of scaled scores outside control range should be less than 10% Task: cyto_batch_integration Metric: emd_mean_ct_vert Inside range: 0 Scaled scores: 48 Percentage outside: 100%

Show 3 warnings
  • warning Raw results Dataset 'human_blood_mass_cytometry' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Dataset: human_blood_mass_cytometry Number of results: 182 Expected number of results: 208 Percentage missing: 12%

  • warning Raw results Dataset 'human_cll_mass_cytometry' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Dataset: human_cll_mass_cytometry Number of results: 176 Expected number of results: 208 Percentage missing: 15%

  • warning Raw results Dataset 'human_cll_mass_cytometry' % failed

    Percentage of failed processes should be less than 10% Task: cyto_batch_integration Dataset: human_cll_mass_cytometry Succeeded processes: 22 Attempted processes: 26 Percentage failed: 15%

Method info 21

Batchadjust with all control sample.

CytofBatchadjust corrects batch effects across cytometry data by aligning signal intensity peaks for each channel across batches. The algorithm uses technical replicates included in each barcode set as reference points to anchor each batch. Adjustments factors are calibrated using anchor samples representing each barcode set, then applied to all samples composing a batch.

This implementation uses samples from each group as control samples.

Batchadjust with all control sample.

CytofBatchadjust corrects batch effects across cytometry data by aligning signal intensity peaks for each channel across batches. The algorithm uses technical replicates included in each barcode set as reference points to anchor each batch. Adjustments factors are calibrated using anchor samples representing each barcode set, then applied to all samples composing a batch.

This implementation uses samples only from one group as control samples.

ComBat batch correction for single-cell data, implemented in the scanpy package

Corrects for batch effects by fitting linear models, gains statistical power via an EB framework where information is borrowed across genes. This uses the implementation combat.py

cyCombine (all-controls, to-goal)code ↗docs ↗source ↗Pedersen et al., 2022

cyCombine run with all control samples, correcting to a goal batch

cyCombine perform batch integration using self-organizing maps and ComBat. It first uses self-organizing maps (SOM) to group similar cells into clusters, then applies a ComBat-based method to correct batch effects within each cluster. Here, we run cyCombine using all control samples (replicates, in cyCombine terms), with a square SOM grid and correct the batches to batch 1. The size of the SOM grid is varied linearly between value of 6-16, with default set to 8 following the default value in create_som function in cyCombine. Rlen, which is the number of times data is presented to the SOM network and can impact clustering quality. It is varied linearly between value of 10-20 with default set to 10.

cyCombine (all-controls, to-middle)code ↗docs ↗source ↗Pedersen et al., 2022

cyCombine run with all control samples, correcting to a midpoint

cyCombine perform batch integration using self-organizing maps and ComBat. It first uses self-organizing maps (SOM) to group similar cells into clusters, then applies a ComBat-based method to correct batch effects within each cluster. Here, we run cyCombine using all control samples (replicates, in cyCombine terms), with a square SOM grid and correct the batches to a midpoint derived from all batches. The size of the SOM grid is varied linearly between value of 6-16, with default set to 8 following the default value in create_som function in cyCombine. Rlen, which is the number of times data is presented to the SOM network and can impact clustering quality. It is varied linearly between value of 10-20 with default set to 10.

cyCombine (no-controls, to-goal)code ↗docs ↗source ↗Pedersen et al., 2022

cyCombine run without control samples, correcting to a goal batch

cyCombine perform batch integration using self-organizing maps and ComBat. It first uses self-organizing maps (SOM) to group similar cells into clusters, then applies a ComBat-based method to correct batch effects within each cluster. Here, we run cyCombine without any control samples (replicates, in cyCombine terms), with a square SOM grid and correct the batches to batch 1. The size of the SOM grid is varied linearly between value of 6-16, with default set to 8 following the default value in create_som function in cyCombine. Rlen, which is the number of times data is presented to the SOM network and can impact clustering quality. It is varied linearly between value of 10-20 with default set to 10.

cyCombine (no-controls, to-middle)code ↗docs ↗source ↗Pedersen et al., 2022

cyCombine run without control samples, correcting to a midpoint

cyCombine perform batch integration using self-organizing maps and ComBat. It first uses self-organizing maps (SOM) to group similar cells into clusters, then applies a ComBat-based method to correct batch effects within each cluster. Here, we run cyCombine without any control samples (replicates, in cyCombine terms), with a square SOM grid and correct the batches to a midpoint derived from all batches. The size of the SOM grid is varied linearly between value of 6-16, with default set to 8 following the default value in create_som function in cyCombine. Rlen, which is the number of times data is presented to the SOM network and can impact clustering quality. It is varied linearly between value of 10-20 with default set to 10.

cyCombine (one-control, to-goal)code ↗docs ↗source ↗Pedersen et al., 2022

cyCombine run with control sample from one group only, correcting to a goal batch

cyCombine perform batch integration using self-organizing maps and ComBat. It first uses self-organizing maps (SOM) to group similar cells into clusters, then applies a ComBat-based method to correct batch effects within each cluster. Here, we run cyCombine using control samples only from one biological group (replicates, in cyCombine terms), with a square SOM grid and correct the batches to batch 1. The size of the SOM grid is varied linearly between value of 6-16, with default set to 8 following the default value in create_som function in cyCombine. Rlen, which is the number of times data is presented to the SOM network and can impact clustering quality. It is varied linearly between value of 10-20 with default set to 10.

cyCombine (one-control, to-middle)code ↗docs ↗source ↗Pedersen et al., 2022

cyCombine run with control sample from one group only, correcting to a midpoint

cyCombine perform batch integration using self-organizing maps and ComBat. It first uses self-organizing maps (SOM) to group similar cells into clusters, then applies a ComBat-based method to correct batch effects within each cluster. Here, we run cyCombine using control samples only from one biological group (replicates, in cyCombine terms), with a square SOM grid and correct the batches to a midpoint derived from all batches. The size of the SOM grid is varied linearly between value of 6-16, with default set to 8 following the default value in create_som function in cyCombine. Rlen, which is the number of times data is presented to the SOM network and can impact clustering quality. It is varied linearly between value of 10-20 with default set to 10.

CytoNorm run with all control samples, correcting to a goal batch.

CytoNorm corrects batch effects by using reference control samples (aliquots of one sample, technical replicates) included with each batch. It clusters cells using FlowSOM, then trains a model on the control samples to learn how marker expression distributions differ across batches for each population. It then uses splines to align these distributions to a common reference (either a midpoint derived from all batches or to a single batch).

Here, we run CytoNorm using all control samples available, aligning the batches to batch 1.

The parameter nQ, which specifies the number of quantiles used when computing the splines is varied linearly between value of 80-120, with default set to 99 following the default value provided by CytoNorm. Clustering was performed by FlowSOM. The number of cells clustered by FlowSOM is set to be number of cells in the smallest control samples or 1,000,000, whichever is the smaller, multiplied by how many control samples there are in the data. The size of the SOM grid is varied linearly between value of 6-16, with default set to 15 following the default value provided by CytoNorm. The number of metaclusters is varied linearly between value of 8-20, with default set to 10 following the default value provided by CytoNorm.

CytoNorm (all-controls, to-middle)code ↗docs ↗source ↗Van Gassen et al., 2019

CytoNorm run with all control samples, correcting to a midpoint.

CytoNorm corrects batch effects by using reference control samples (aliquots of one sample, technical replicates) included with each batch. It clusters cells using FlowSOM, then trains a model on the control samples to learn how marker expression distributions differ across batches for each population. It then uses splines to align these distributions to a common reference (either a midpoint derived from all batches or to a single batch).

Here, we run CytoNorm using all control samples available, aligning to a midpoint derived from all batches.

The parameter nQ, which specifies the number of quantiles used when computing the splines is varied linearly between value of 80-120, with default set to 99 following the default value provided by CytoNorm. Clustering was performed by FlowSOM. The number of cells clustered by FlowSOM is set to be number of cells in the smallest control samples or 1,000,000, whichever is the smaller, multiplied by how many control samples there are in the data. The size of the SOM grid is varied linearly between value of 6-16, with default set to 15 following the default value provided by CytoNorm. The number of metaclusters is varied linearly between value of 8-20, with default set to 10 following the default value provided by CytoNorm.

CytoNorm run without control samples, correcting to a goal batch.

CytoNorm corrects batch effects by using reference control samples (aliquots of one sample, technical replicates) included with each batch. It clusters cells using FlowSOM, then trains a model on the control samples to learn how marker expression distributions differ across batches for each population. It then uses splines to align these distributions to a common reference (either a midpoint derived from all batches or to a single batch).

In this CytoNorm version, an aggregate of each batch is created and subsequently used as a proxy for the control samples, aligning the batches to batch 1.

The number of cells used to create an aggregate is set as the number of cells in the smallest sample or 1,000,000, whichever is the smaller, multiplied by how many samples there are in the batch. Clustering was performed by FlowSOM. The size of the SOM grid is varied linearly between value of 6-16, with default set to 15 following the default value provided by CytoNorm. The number of metaclusters is varied linearly between value of 8-20, with default set to 10 following the default value provided by CytoNorm. The number of cells clustered by FlowSOM is set as the number of cells in the smallest aggregate or 1,000,000, whichever is the smaller, multiplied by how many batches there are in the data (because 1 aggregate = 1 batch).

CytoNorm (no-controls, to-middle)code ↗docs ↗source ↗Van Gassen et al., 2019

CytoNorm run without control samples, correcting to a midpoint.

CytoNorm corrects batch effects by using reference control samples (aliquots of one sample, technical replicates) included with each batch. It clusters cells using FlowSOM, then trains a model on the control samples to learn how marker expression distributions differ across batches for each population. It then uses splines to align these distributions to a common reference (either a midpoint derived from all batches or to a single batch).

In this CytoNorm version, an aggregate of each batch is created and subsequently used as a proxy for the control samples, aligning to a midpoint derived from all batches.

The number of cells used to create an aggregate is set as the number of cells in the smallest sample or 1,000,000, whichever is the smaller, multiplied by how many samples there are in the batch. Clustering was performed by FlowSOM. The size of the SOM grid is varied linearly between value of 6-16, with default set to 15 following the default value provided by CytoNorm. The number of metaclusters is varied linearly between value of 8-20, with default set to 10 following the default value provided by CytoNorm. The number of cells clustered by FlowSOM is set as the number of cells in the smallest aggregate or 1,000,000, whichever is the smaller, multiplied by how many batches there are in the data (because 1 aggregate = 1 batch).

CytoNorm run with control sample from one group only, correcting to a goal batch

CytoNorm corrects batch effects by using reference control samples (aliquots of one sample, technical replicates) included with each batch. It clusters cells using FlowSOM, then trains a model on the control samples to learn how marker expression distributions differ across batches for each population. It then uses splines to align these distributions to a common reference (either a midpoint derived from all batches or to a single batch).

Here, we run CytoNorm using only control samples from one donor, aligning the batches to batch 1.

The parameter nQ, which specifies the number of quantiles used when computing the splines is varied linearly between value of 80-120, with default set to 99 following the default value provided by CytoNorm. Clustering was performed by FlowSOM. The number of cells clustered by FlowSOM is set to be number of cells in the smallest control samples or 1,000,000, whichever is the smaller, multiplied by how many control samples there are in the data. The size of the SOM grid is varied linearly between value of 6-16, with default set to 15 following the default value provided by CytoNorm. The number of metaclusters is varied linearly between value of 8-20, with default set to 10 following the default value provided by CytoNorm.

CytoNorm (one-control, to-middle)code ↗docs ↗source ↗Van Gassen et al., 2019

CytoNorm run with control sample from one group only, correcting to a midpoint

CytoNorm corrects batch effects by using reference control samples (aliquots of one sample, technical replicates) included with each batch. It clusters cells using FlowSOM, then trains a model on the control samples to learn how marker expression distributions differ across batches for each population. It then uses splines to align these distributions to a common reference (either a midpoint derived from all batches or to a single batch).

Here, we run CytoNorm using only control samples from one donor, aligning to a midpoint derived from all batches.

The parameter nQ, which specifies the number of quantiles used when computing the splines is varied linearly between value of 80-120, with default set to 99 following the default value provided by CytoNorm. Clustering was performed by FlowSOM. The number of cells clustered by FlowSOM is set to be number of cells in the smallest control samples or 1,000,000, whichever is the smaller, multiplied by how many control samples there are in the data. The size of the SOM grid is varied linearly between value of 6-16, with default set to 15 following the default value provided by CytoNorm. The number of metaclusters is varied linearly between value of 8-20, with default set to 10 following the default value provided by CytoNorm.

A deep generative model for correcting batch effects

CytoVI is a deep generative model that utilizes antibody-based single-cell profiles to learn a biologically meaningful latent representation of each cell. It is part of the scvi-tools framework and is built upon the variational autoencoder (VAE) architecture.

Batch effect correction using a per‐channel basis normalization method (gaussNorm)

This method batch-normalizes a set of cytometry data samples by identifying and aligning the high density regions (landmarks or peaks) for each channel. The data of each channel is shifted in such a way that the identified high density regions are moved to fixed locations called base landmarks. Normalization is achieved in three phases:

  1. identifying high-density regions (landmarks) for each flowFrame in the flowSet for a single channel
  2. computing the best matching between the landmarks and a set of fixed reference landmarks for each channel called base landmarks
  3. manipulating the data of each channel in such a way that each landmark is moved to its matching base landmark. Please note that this normalization is on a channel-by-channel basis

NOTE: The default implementation uses max.lms=2, although for some channels it is not possible to compute 2 landmarks, resulting in an error. In order to fully automate the batch normalization process, this implementation checks whether it is possible to compute 2 landmarks, and if not, it sets max.lms=1 for that channel.

Harmonypy is a port of the harmony R package

Harmony is a general-purpose R package with an efficient algorithm for integrating multiple data sets. It is especially useful for large single-cell datasets such as single-cell RNA-seq.

Uses a linear model and matrix decomposition to remove batch effects from a dataset

Limma removeBatchEffect is a method that uses a linear model and matrix decomposition to remove batch effects from a dataset. It first fits a linear model to the data, then decomposes the model matrix into a set of orthogonal components. The batch effect is then removed by subtracting the component corresponding to the batch effect from the data.

Batch integrate data to a goal batch using mutual nearest neighbors identified via Seurat reciprocal PCA.

Seurat RPCA performs batch integration by projecting each query dataset into the PCA space of a goal batch, and identifying anchors using reciprocal PCA (RPCA). RPCA identifies mutual nearest neighbors between the goal batch and the remaining batches in their shared low-dimensional space (PCA), which are used to align and integrate the batches into the goal batch.

We ran Seurat RPCA implemented in Seurat v4.4.0, available on their github. This is because subsequent version (>= v5) does not support getting corrected count matrix, only corrected PC space. See: https://github.com/satijalab/seurat/issues/8551.

We varied the number of PCs and nearest neighbours considered when running RPCA.

Batch integrate data to a midpoint using mutual nearest neighbors identified via Seurat reciprocal PCA.

Seurat RPCA performs batch integration by identifying mutual nearest neighbors (anchors) between all batches using reciprocal PCA (RPCA). Instead of merging batches into a single PCA space, each batch is projected into the PCA space of the others, and anchors are found where cells are mutual nearest neighbors across these projections. These anchors are then used to compute batch correction vectors, aligning the batches into a shared corrected space.

We ran Seurat RPCA implemented in Seurat v4.4.0, available on their github. This is because subsequent version (>= v5) does not support getting corrected count matrix, only corrected PC space. See: https://github.com/satijalab/seurat/issues/8551

We varied the number of PCs and nearest neighbours considered when running RPCA.

Control method info 5
No Integration

Control method returning the unintegrated data without performing batch correction.

The component works by reading and writing back the 'unintegrated' data without performing any operation.

Perfect Integration

Act as a positive control by returning samples from the same batch (batch 1) for both splits.

The method returns only samples from one batch (batch 1), for both the splits. This act as a positive control method for the metrics that compare technical replicates between runs ('horizontal' metrics), as the difference between a sample compared to itself should be zero. It also acts as a positive control for 'vertical' metrics, as the samples being returned are from the same batch.

Shuffle Integration

Randomly shuffle cells in the whole dataset.

This negative control randomly permutes cell-to-sample (hence batch) assignments while keeping each cell's measured markers unchanged. This destroys any biological and batch specific structure but preserves marker expression.

Purpose:

  • Provide a baseline to verify that integration methods outperform random assignment of cells to batches.

Example:

  • A cell from a KO sample in batch 1 may be reassigned to any sample in the whole data. It may be reassigned to another KO sample in batch 1 or 2, or to a WT sample in batch 1 or 2.
Shuffle Integration — within batches

Randomly reassign cells to any samples within the same batch.

This negative-control method randomly permutes cell-to-cell type assignments. Cells remain assigned to their original batch (batch effects preserved). Within each batch, cells are reassigned to random samples, destroying biological/sample-specific structure (e.g., KO vs WT differences).

Purpose:

  • Evaluate whether an integration method preserves differences between samples and biological groups while removing batch effects.

Example:

  • A cell from a KO sample in batch 1 may be reassigned to any sample in batch 1 (KO or WT), but it will never be moved to batch 2.
Shuffle Integration — within cell type

Randomly reassign cells to any cell types

This negative-control method randomly permutes cell-to-cell type assignments. Cells will be assigned to any cell types, regardless of their original cell type or sample of origin or batch of origin.

Purpose:

  • Evaluate whether an integration method preserves differences between cell types while removing batch effects.

Example:

  • A Neutrophil from a KO sample in batch 1 may be reassigned to any cell type (B cell, T cell, Monocyte, etc.) from any sample in any batch.
Metric info 8
Average Batch R-squared Cell Typelower is betterDraper & Smith, n.d.

Quantifies how strongly the batch covariate explains the variance in the data among technical replicates after correction (by taking into account cell type effect).

First, a simple linear model is fitted for each paired sample and marker to determine the fraction of variance ($R^{2}$) explained by the batch covariate $B$. The average batch R-squared is then computed as the average of the $R^{2}$ values across all paired samples, markers, cell types. As a result, $\overline{R^2_B}_{cell\ type}$ quantifies how much of the total variability in the data is driven by batch effects. Consequently, lower values are desirable.

$\overline{R^2_B}{cell\ type} = \frac{1}{NCM}\sum{\substack{(x_{\mathrm{split1}},,x_{\mathrm{split2}})\ \textit{paired samples}}}^{N} \sum_{j=1}^{C} \sum_{i=1}^{M},R^2!\bigl(\mathrm{marker}_i \mid B\bigr)$

Where:

  • $N$ is the number of paired samples, where $x_{\mathrm{split1}}$ and $x_{\mathrm{split2}}$ are the two technical replicates that have been batch-corrected. Technical replicates belong to different batches.
  • $M$ is the number of markers
  • $C$ is the number of cell types
  • $B$ is the batch covariate

A high value of $\overline{R^2_B}{global}$ or $\overline{R^2_B}{cell\ type}$ indicates that the batch variable explains a large portion of the variance in the data, which indicates a higher level of batch effects. A good performance on $\overline{R^2_B}{global}$ but not on $\overline{R^2_B}{cell\ type}$ might indicate that the batch effect correction is not addressing cell type specific batch effects.

Average Batch R-squared Globallower is betterDraper & Smith, n.d.

Quantifies how strongly the batch covariate explains the variance in the data among technical replicates after correction.

First, a simple linear model is fitted for each paired sample and marker to determine the fraction of variance ($R^{2}$) explained by the batch covariate $B$. The average batch R-squared is then computed as the average of the $R^{2}$ values across all paired samples, markers. As a result, $\overline{R^2_B}_{global}$ quantifies how much of the total variability in the data is driven by batch effects. Consequently, lower values are desirable.

$\overline{R^2_B}{global} = \frac{1}{N*M}\sum{\substack{(x_{\mathrm{split1}},,x_{\mathrm{split2}})\ \textit{paired samples}}}^{N} \sum_{i=1}^{M} ,R^2!\bigl(\mathrm{marker}_i \mid B\bigr)$

Where:

  • $N$ is the number of paired samples, where $x_{\mathrm{split1}}$ and $x_{\mathrm{split2}}$ are the two technical replicates that have been batch-corrected. Technical replicates belong to different batches.
  • $M$ is the number of markers
  • $B$ is the batch covariate

A higher value of $\overline{R^2_B}_{global}$ indicates that the batch variable explains more of the variance in the data, which indicates a higher level of batch effects.

Cell type Lisi score.

Compute the cell type local inverse simpson index (cLISI) for each cell, then return the median as the final score.

EMD Mean CT Horizontallower is betterRubner et al., 2000

Mean Earth Mover Distance calculated horizontally across donors for each cell type and marker.

Earth Mover Distance (EMD), also known as the Wasserstein metric, measures the difference between two probability distributions.

Here, EMD is used to compare marker expression distributions between paired samples from the same donor quantified across two different batches. For each paired sample, cell type, and marker, the marker expression values are first converted into probability distributions. This is done by binning the expression values into a range from -100 to 100 with a bin width of 0.1. The wasserstein_distance function from SciPy is then used to calculate the EMD between the two probability distributions belonging to the same cell type, marker, and a given paired samples. This is then repeated for every cell type, marker, and paired sample. Finally, the average of all these EMD values is computed and reported as the metric score.

A high score indicates large overall differences in the distributions of marker expressions between the paired samples, suggesting poor batch integration. A low score means the small differences in marker expression distributions between batches, indicating good batch integration.

EMD Mean CT Verticallower is betterRubner et al., 2000

Mean Earth Mover Distance across batch corrected samples, cell types, and markers.

Earth Mover Distance (EMD), also known as the Wasserstein metric, measures the difference between two probability distributions.

Here, EMD is used to compare marker expression distributions between all integrated samples from the same group. For each pair of samples, cell type, and marker, the marker expression values are first converted into probability distributions. This is done by binning the expression values into a range from -100 to 100 with a bin width of 0.1. The wasserstein_distance function from SciPy is then used to calculate the EMD between the two probability distributions belonging to the same cell type, marker, and a given paired samples. This is then repeated for every cell type, marker, and paired sample. Finally, the average of all these EMD values is computed and reported as the metric score.

A high score indicates overall, there is a large difference in distribution of marker expression after batch integration. A low score means that overall, the samples are well integrated.

FlowSOM Mean Mapping Similarityhigher is betterSofie Van Gassen, 2017Van Gassen et al., 2015

Assess the similarity between FlowSOM trees of integrated and validation samples.

The metric is based on the FlowSOM algorithm, a popular method which uses self-organizing maps for the viasualization/interpretation/clustering of cytometry data. The FlowSOM algorithm creates a tree structure that represents the relationships between different cell populations in the data.

For each paired sample of technical replicates (where 'split1' are integrated technical replicates from split 1 and 'split2' are integrated technical replicates from split 2)

  1. A FlowSOM tree is created using cells in the sample from split 1.
  2. Cells from the sample in split 2 are mapped onto the FlowSOM tree created in step 1.
  3. A similarity measure is computed by comparing cell type proportions of 'split2' and 'split1' in each flowsom cluster.

Ideally, the proportions of cell types in the flowsom clusters of the paired samples should be similar, as they are technical replicates.

The FlowSOM mapping similarity measure can be expressed as follows: $\text{FlowSOM mapping similarity} = 100 - \text{FlowSOM mapping dissimilarity}$

The $\text{FlowSOM mapping dissimilarity}$ is:

$\text{FlowSOM mapping dissimilarity} = \sum_{k=1}^{K}w_{k}\sum_{c=1}^{C}\abs{P^{split1}{k,c} - P^{split2}{k,c}}$

Where:

  • $K$ is the number of flowsom clusters
  • $C$ is the number of cell types
  • $w_{k}$ is the weight of flowsom cluster $k$. It refers to the number of cells in flowsom cluster $k$, divided by the total number of cells in the flowsom tree. Note: the flowsom tree contains all the cells of a sample pair (both integrated and validation).
  • $P^{split1}_{k,c}$ is the percentage of cell type $c$ in flowsom cluster $k$ of the split 1 sample
  • $P^{split2}_{k,c}$ is the percentage of cell type $c$ in flowsom cluster $k$ of the split 2 sample

The average FlowSOM mapping similarity among all paired samples is computed and reported as the final metric value. Unlabelled cells are excluded from the analysis.

Integration Lisi score.

Compute the integration local inverse simpson index (iLISI) for each cell, then return the median as the final score.

Ratio of inconsistent peakslower is betterVirtanen et al., 2020

Ratio of the number of cell‑type marker‑expression peaks between unintegrated and batch-integrated data.

The metric compares the number of cell type specific marker expression peaks between unintegrated and batch integrated data. The number of peaks is calculated using the scipy.signal.find_peaks function. The metric is calculated as the absolute difference between the number of peaks in the unintegrated and batch-normalized data. The (cell type) marker expression profiles are first smoothed using kernel density estimation (KDE) (scipy.stats.gaussian_kde), and then peaks are then identified using the scipy.signal.find_peaks function. For peak calling, the prominence parameter is set to 0.1 and the height parameter is set to 0.05*max_density. Ratio of inconsistent peaks is defined as number of cases where the number of peaks differ between the two splits in the batch normalized data divided by the total number of cases. Cases where there are different number of peaks between the two splits in the unintegrated data are ignored from the denominator. A lower score indicates better performance, means there are less cases with inconsistent peaks after batch integration.

Dataset info 3
Human Blood Mass Cytometry unlinked

Mass cytometry dataset of whole blood from 2 human donors. For each donor, 2 samples are included: one unstimulated and one stimulated. For each of these, aliquotes of the same original sample were divided into 2 batches to allow the creation of sample-paired replicates for benchmarking purposes.

CyTOF data (whole blood) from 2 human donors, each with unstimulated and stimulated samples. The stimulated samples were treated with Interferonα (IFNα) and Lipopolysaccharide (LPS). The mass cytometry cytometry panel includes 21 surface markers for cell phenotyping and 14 markers for functional analysis of the signaling responses. The data has been arcsinh transformed with cofactor 5 and pregated to only include events which fall either in the CD66-CD45+ or CD66+ gates.

Human CLL mass cytometry unlinked

Mass cytometry data of PBMC from CLL patients and healthy donors.

Mass cytometry data of PBMC from CLL patients and healthy donors. Data was originally used in the CytofRUV paper. Data has been preprocessed (bead normalised, arcsinh with cofactor 5, pregated on live single cells).

Mouse Spleen Flow Cytometry unlinked

Flow cytometry data of spleens of 8 mice. For each mouse, aliquotes of the same original sample were divided into 2 batches and measured with 2 different instrument settings to allow the creation of sample-paired replicates for benchmarking purposes.

Flow cytometry data of spleens from 4 WT (IKK2 fl/fl CD11c-cre +/+) and 4 KO (IKK2 fl/fl CD11c-cre Tg/+) B6 mice, measured with a 22-color panel and 2 different instrument settings. Data has been preprocessed (compensated with a batch-specific compensation matrix, logicle transformed, cleaned with PeacoQC and pregated on live single CD45+ cells).

References

  1. Daly, A. C., Prendergast, M. E., Hughes, A. J., & Burdick, J. A. (2021). Bioprinting for the Biologist. 10.1016/j.cell.2020.12.002 ↗
  2. Draper, N. R., & Smith, H. (n.d.). Applied regression analysis. John Wiley & Sons.
  3. Sofie Van Gassen, B. C. (2017). FlowSOM. 10.18129/B9.BIOC.FLOWSOM ↗
  4. Van Gassen, S., Callebaut, B., Van Helden, M. J., Lambrecht, B. N., Demeester, P., Dhaene, T., & Saeys, Y. (2015). FlowSOM: Using self‐organizing maps for visualization and interpretation of cytometry data. 10.1002/cyto.a.22625 ↗
  5. Van Gassen, S., Gaudilliere, B., Angst, M. S., Saeys, Y., & Aghaeepour, N. (2019). CytoNorm: A Normalization Algorithm for Cytometry Data. 10.1002/cyto.a.23904 ↗
  6. Van Gassen, S., Gaudilliere, B., Angst, M. S., Saeys, Y., & Aghaeepour, N. (2020). CytoNorm: a normalization algorithm for cytometry data. Cytometry Part A, 97(3), 268–278.
  7. Hahne, F., Khodabakhshi, A. H., Bashashati, A., Wong, C., Gascoyne, R. D., Weng, A. P., Seyfert‐Margolis, V., Bourcier, K., Asare, A., Lumley, T., Gentleman, R., & Brinkman, R. R. (2009). Per‐channel basis normalization methods for flow cytometry data. 10.1002/cyto.a.20823 ↗
  8. Hao, Y., Hao, S., Andersen-Nissen, E., Mauck, W. M., Zheng, S., Butler, A., Lee, M. J., Wilk, A. J., Darby, C., Zager, M., Hoffman, P., Stoeckius, M., Papalexi, E., Mimitou, E. P., Jain, J., Srivastava, A., Stuart, T., Fleming, L. M., Yeung, B., … Satija, R. (2021). Integrated analysis of multimodal single-cell data. Cell, 184(13), 3573-3587.e29. 10.1016/j.cell.2021.04.048 ↗
  9. Ingelfinger, F., Levy, N., Ergen, C., Bakulin, A., Becker, A., Boyeau, P., Kim, M., Ditz, D., Dirks, J., Maaskola, J., Wertheimer, T., Zeiser, R., Widmer, C. C., Amit, I., & Yosef, N. (2025). CytoVI: Deep generative modeling of antibody-based single cell technologies. 10.1101/2025.09.07.674699 ↗
  10. Johnson, W. E., Li, C., & Rabinovic, A. (2006). Adjusting batch effects in microarray expression data using empirical Bayes methods. Biostatistics, 8(1), 118–127. 10.1093/biostatistics/kxj037 ↗
  11. Korsunsky, I., Millard, N., Fan, J., Slowikowski, K., Zhang, F., Wei, K., Baglaenko, Y., Brenner, M., ru Po-Loh, & Raychaudhuri, S. (2019). Fast, sensitive and accurate integration of single-cell data with Harmony. Nature Methods, 16(12), 1289–1296. 10.1038/s41592-019-0619-0 ↗
  12. Leomazzi, L., Putri, G., Cannoodt, R., Quintelier, K., & Van Gassen, S. (2025). Benchmarking batch integration methods for mass cytometry data. bioRxiv – in Preparation.
  13. Pedersen, C. B., Dam, S. H., Barnkob, M. B., Leipold, M. D., Purroy, N., Rassenti, L. Z., Kipps, T. J., Nguyen, J., Lederer, J. A., Gohil, S. H., Wu, C. J., & Olsen, L. R. (2022). cyCombine allows for robust integration of single-cell cytometry datasets within and across technologies. 10.1038/s41467-022-29383-5 ↗
  14. Ritchie, M. E., Phipson, B., Wu, D., Hu, Y., Law, C. W., Shi, W., & Smyth, G. K. (2015). limma powers differential expression analyses for RNA-sequencing and microarray studies. 10.1093/nar/gkv007 ↗
  15. Rubner, Y., Tomasi, C., & Guibas, L. J. (2000). The Earth Mover’s Distance as a Metric for Image Retrieval. 10.1023/a:1026543900054 ↗
  16. Schuyler, R. P., Jackson, C., Garcia-Perez, J. E., Baxter, R. M., Ogolla, S., Rochford, R., Ghosh, D., Rudra, P., & Hsieh, E. W. (2019). Minimizing batch effects in mass cytometry data. Frontiers in Immunology, 10, 2367.
  17. Schuyler, R. P., Jackson, C., Garcia-Perez, J. E., Baxter, R. M., Ogolla, S., Rochford, R., Ghosh, D., Rudra, P., & Hsieh, E. W. Y. (2019). Minimizing Batch Effects in Mass Cytometry Data. 10.3389/fimmu.2019.02367 ↗
  18. Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., … Vázquez-Baeza, Y. (2020). SciPy 1.0: fundamental algorithms for scientific computing in Python. 10.1038/s41592-019-0686-2 ↗