All benchmarks

Batch Integration

Remove unwanted batch effects from scRNA-seq data while retaining biologically meaningful variation.

24 methods
7 control methods
6 datasets
16 metrics
5 releases
Task repository MIT v2.1.0-rc1

As single-cell technologies advance, single-cell datasets are growing both in size and complexity. Especially in consortia such as the Human Cell Atlas, individual studies combine data from multiple labs, each sequencing multiple individuals possibly with different technologies. This gives rise to complex batch effects in the data that must be computationally removed to perform a joint analysis. These batch integration methods must remove the batch effect while not removing relevant biological information. Currently, over 200 tools exist that aim to remove batch effects scRNA-seq datasets (Zappia et al., 2018). These methods balance the removal of batch effects with the conservation of nuanced biological information in different ways. This abundance of tools has complicated batch integration method choice, leading to several benchmarks on this topic (Luecken et al., 2021; Tran et al., 2020; Chazarra-Gil et al., 2021; Mereu et al., 2020). Yet, benchmarks use different metrics, method implementations and datasets. Here we build a living benchmarking task for batch integration methods with the vision of improving the consistency of method evaluation.

In this task we evaluate batch integration methods on their ability to remove batch effects in the data while conserving variation attributed to biological effects. As input, methods require either normalised or unnormalised data with multiple batches and consistent cell type labels. The batch integrated output can be a feature matrix, a low dimensional embedding and/or a neighbourhood graph. The respective batch-integrated representation is then evaluated using sets of metrics that capture how well batch effects are removed and whether biological variance is conserved. We have based this particular task on the latest, and most extensive benchmark of single-cell data integration methods.

Contributors

  • Michaela Mueller
    maintainerauthor
  • Malte Luecken
    author
  • Daniel Strobl
    author
  • Robrecht Cannoodt
    author
  • Luke Zappia
    author
  • Scott Gigante
    contributor
  • Kai Waldrant
    contributor
  • Martin Kim
    contributor
  • Sai Nirmayi Yasa
    contributor
  • Jeremie Kalfon
    contributor

Leaderboard

Methods ranked by scaled overall mean. Each cell encodes a score from 0 to 1 by size and intensity.

QC: Normalisation Visualisation 16 plots

Per metric: points placed by control-anchored scaled score (x); dashed lines mark scaled 0 and 1 (worst/best control); the lower axis shows the raw score. Points beyond [-0.2, 1.2] are clamped to the edge as triangles. Hover a dot or line to highlight it and read details.

methodcontrol
  • ARIhigher better
    Embed cell typesEmbed cell typesPerfect embedding by …Perfect embedding by celltype with jitterbatchelor mnnCorrectbatchelor mnnCorrectscANVIscANVIHarmonypyHarmonypyBBKNNBBKNNSCimilaritySCimilarityscVIscVIHarmonyHarmonyscGPT (MLflow model)scGPT (MLflow model)batchelor fastMNNbatchelor fastMNNscGPT (zero shot)scGPT (zero shot)TranscriptFormer (MLf…TranscriptFormer (MLflow model)scVI (MLflow model)scVI (MLflow model)UCEUCEscGPT (fine-tuned)scGPT (fine-tuned)LIGERLIGERUCE (MLflow model)UCE (MLflow model)CombatCombatpyligerpyligerShuffle integration b…Shuffle integration by cell typeNo integrationNo integrationSCALEXSCALEXscPRINTscPRINTNo integration by Bat…No integration by BatchmnnpymnnpyScanoramaScanoramaShuffle integration b…Shuffle integration by batchShuffle integrationShuffle integration-3.0e-40.250.50.75100.250.50.751rawscaled
  • ARI_batchhigher better
    Shuffle integrationShuffle integrationLIGERLIGERpyligerpyligerscPRINTscPRINTHarmonyHarmonyHarmonypyHarmonypyScanoramaScanoramaBBKNNBBKNNShuffle integration b…Shuffle integration by cell typebatchelor mnnCorrectbatchelor mnnCorrectEmbed cell typesEmbed cell typesPerfect embedding by …Perfect embedding by celltype with jitterSCALEXSCALEXscVIscVIscANVIscANVIscGPT (zero shot)scGPT (zero shot)scVI (MLflow model)scVI (MLflow model)batchelor fastMNNbatchelor fastMNNscGPT (MLflow model)scGPT (MLflow model)SCimilaritySCimilarityUCEUCEmnnpymnnpyCombatCombatTranscriptFormer (MLf…TranscriptFormer (MLflow model)No integrationNo integrationscGPT (fine-tuned)scGPT (fine-tuned)UCE (MLflow model)UCE (MLflow model)Shuffle integration b…Shuffle integration by batchNo integration by Bat…No integration by Batch0.7290.8020.8750.9481.02100.250.50.751rawscaled
  • ASW batchhigher better
    UCEUCEEmbed cell typesEmbed cell typesPerfect embedding by …Perfect embedding by celltype with jitterUCE (MLflow model)UCE (MLflow model)scGPT (MLflow model)scGPT (MLflow model)scGPT (zero shot)scGPT (zero shot)mnnpymnnpyscVIscVIShuffle integration b…Shuffle integration by cell typebatchelor fastMNNbatchelor fastMNNCombatCombatNo integrationNo integrationscPRINTscPRINTHarmonypyHarmonypyHarmonyHarmonyShuffle integrationShuffle integrationscANVIscANVIscVI (MLflow model)scVI (MLflow model)Shuffle integration b…Shuffle integration by batchScanoramaScanoramascGPT (fine-tuned)scGPT (fine-tuned)SCimilaritySCimilarityTranscriptFormer (MLf…TranscriptFormer (MLflow model)SCALEXSCALEXbatchelor mnnCorrectbatchelor mnnCorrectpyligerpyligerLIGERLIGERNo integration by Bat…No integration by Batch0.5330.6460.7590.8720.98400.250.50.751rawscaled
  • ASW Labelhigher better
    Embed cell typesEmbed cell typesPerfect embedding by …Perfect embedding by celltype with jitterbatchelor mnnCorrectbatchelor mnnCorrectscANVIscANVISCimilaritySCimilarityUCE (MLflow model)UCE (MLflow model)scGPT (fine-tuned)scGPT (fine-tuned)batchelor fastMNNbatchelor fastMNNHarmonypyHarmonypyNo integrationNo integrationShuffle integration b…Shuffle integration by cell typeHarmonyHarmonyscVI (MLflow model)scVI (MLflow model)CombatCombatscVIscVIUCEUCEscGPT (MLflow model)scGPT (MLflow model)scGPT (zero shot)scGPT (zero shot)pyligerpyligerSCALEXSCALEXLIGERLIGERscPRINTscPRINTmnnpymnnpyTranscriptFormer (MLf…TranscriptFormer (MLflow model)No integration by Bat…No integration by BatchShuffle integrationShuffle integrationShuffle integration b…Shuffle integration by batchScanoramaScanorama0.4180.5610.7040.8470.9900.250.50.751rawscaled
  • Cell Cycle Conservationhigher better
    batchelor mnnCorrectbatchelor mnnCorrectUCE (MLflow model)UCE (MLflow model)No integration by Bat…No integration by Batchbatchelor fastMNNbatchelor fastMNNscANVIscANVIscGPT (MLflow model)scGPT (MLflow model)scGPT (zero shot)scGPT (zero shot)No integrationNo integrationCombatCombatscGPT (fine-tuned)scGPT (fine-tuned)HarmonyHarmonyHarmonypyHarmonypyscVI (MLflow model)scVI (MLflow model)UCEUCEscPRINTscPRINTscVIscVITranscriptFormer (MLf…TranscriptFormer (MLflow model)SCimilaritySCimilaritypyligerpyligerEmbed cell typesEmbed cell typesShuffle integration b…Shuffle integration by cell typePerfect embedding by …Perfect embedding by celltype with jitterLIGERLIGERSCALEXSCALEXmnnpymnnpyScanoramaScanoramaShuffle integration b…Shuffle integration by batchShuffle integrationShuffle integration8.2e-30.2210.4350.6480.86100.250.50.751rawscaled
  • cLISIhigher better
    batchelor mnnCorrectbatchelor mnnCorrectEmbed cell typesEmbed cell typesPerfect embedding by …Perfect embedding by celltype with jitterscANVIscANVIscVIscVINo integrationNo integrationShuffle integration b…Shuffle integration by cell typeUCE (MLflow model)UCE (MLflow model)CombatCombatSCimilaritySCimilaritybatchelor fastMNNbatchelor fastMNNHarmonypyHarmonypyUCEUCEHarmonyHarmonyNo integration by Bat…No integration by BatchTranscriptFormer (MLf…TranscriptFormer (MLflow model)scGPT (fine-tuned)scGPT (fine-tuned)scVI (MLflow model)scVI (MLflow model)scGPT (zero shot)scGPT (zero shot)scGPT (MLflow model)scGPT (MLflow model)pyligerpyligerLIGERLIGERSCALEXSCALEXscPRINTscPRINTBBKNNBBKNNmnnpymnnpyScanoramaScanoramaShuffle integration b…Shuffle integration by batchShuffle integrationShuffle integration0.710.7830.8550.928100.250.50.751rawscaled
  • Graph Connectivityhigher better
    Perfect embedding by …Perfect embedding by celltype with jitterEmbed cell typesEmbed cell typesbatchelor mnnCorrectbatchelor mnnCorrectscANVIscANVIUCE (MLflow model)UCE (MLflow model)scVIscVIscGPT (fine-tuned)scGPT (fine-tuned)BBKNNBBKNNUCEUCEShuffle integration b…Shuffle integration by cell typeNo integrationNo integrationCombatCombatSCimilaritySCimilarityTranscriptFormer (MLf…TranscriptFormer (MLflow model)scGPT (MLflow model)scGPT (MLflow model)scGPT (zero shot)scGPT (zero shot)batchelor fastMNNbatchelor fastMNNHarmonyHarmonyscVI (MLflow model)scVI (MLflow model)HarmonypyHarmonypyscPRINTscPRINTSCALEXSCALEXLIGERLIGERpyligerpyligerNo integration by Bat…No integration by BatchmnnpymnnpyScanoramaScanoramaShuffle integration b…Shuffle integration by batchShuffle integrationShuffle integration0.0170.2620.5080.754100.250.50.751rawscaled
  • HVG overlaphigher better
    Shuffle integration b…Shuffle integration by batchShuffle integration b…Shuffle integration by cell typeCombatCombatShuffle integrationShuffle integrationbatchelor mnnCorrectbatchelor mnnCorrectmnnpymnnpySCALEXSCALEXScanoramaScanorama0.490.6180.7450.873100.250.50.751rawscaled
  • iLISIhigher better
    BBKNNBBKNNbatchelor mnnCorrectbatchelor mnnCorrectscGPT (fine-tuned)scGPT (fine-tuned)Shuffle integrationShuffle integrationUCE (MLflow model)UCE (MLflow model)pyligerpyligerLIGERLIGERmnnpymnnpyShuffle integration b…Shuffle integration by cell typeEmbed cell typesEmbed cell typesPerfect embedding by …Perfect embedding by celltype with jitterScanoramaScanoramaHarmonyHarmonyscPRINTscPRINTHarmonypyHarmonypyscGPT (zero shot)scGPT (zero shot)scGPT (MLflow model)scGPT (MLflow model)SCimilaritySCimilaritybatchelor fastMNNbatchelor fastMNNscANVIscANVISCALEXSCALEXscVIscVITranscriptFormer (MLf…TranscriptFormer (MLflow model)scVI (MLflow model)scVI (MLflow model)UCEUCECombatCombatShuffle integration b…Shuffle integration by batchNo integrationNo integrationNo integration by Bat…No integration by Batch00.1190.2380.3560.47500.250.50.751rawscaled
  • Isolated label ASWhigher better
    Perfect embedding by …Perfect embedding by celltype with jitterEmbed cell typesEmbed cell typesscANVIscANVISCimilaritySCimilarityscVIscVINo integrationNo integrationShuffle integration b…Shuffle integration by cell typeCombatCombatHarmonyHarmonyscVI (MLflow model)scVI (MLflow model)scGPT (zero shot)scGPT (zero shot)HarmonypyHarmonypyscGPT (MLflow model)scGPT (MLflow model)TranscriptFormer (MLf…TranscriptFormer (MLflow model)batchelor fastMNNbatchelor fastMNNNo integration by Bat…No integration by BatchSCALEXSCALEXpyligerpyligerLIGERLIGERscPRINTscPRINTUCEUCEShuffle integrationShuffle integrationScanoramaScanoramaShuffle integration b…Shuffle integration by batch0.3250.4910.6570.8240.9900.250.50.751rawscaled
  • Isolated label F1 scorehigher better
    Embed cell typesEmbed cell typesPerfect embedding by …Perfect embedding by celltype with jitterscANVIscANVIscVIscVIShuffle integration b…Shuffle integration by cell typeNo integrationNo integrationscVI (MLflow model)scVI (MLflow model)scPRINTscPRINTbatchelor fastMNNbatchelor fastMNNCombatCombatSCimilaritySCimilarityNo integration by Bat…No integration by BatchscGPT (MLflow model)scGPT (MLflow model)scGPT (zero shot)scGPT (zero shot)LIGERLIGERpyligerpyligerHarmonypyHarmonypyUCEUCEHarmonyHarmonyTranscriptFormer (MLf…TranscriptFormer (MLflow model)SCALEXSCALEXScanoramaScanoramaBBKNNBBKNNShuffle integration b…Shuffle integration by batchShuffle integrationShuffle integration9.0e-40.2510.50.75100.250.50.751rawscaled
  • kBET pegasushigher better
    Shuffle integrationShuffle integrationShuffle integration b…Shuffle integration by cell typeEmbed cell typesEmbed cell typesPerfect embedding by …Perfect embedding by celltype with jitterScanoramaScanoramapyligerpyligerLIGERLIGERscPRINTscPRINTbatchelor fastMNNbatchelor fastMNNHarmonyHarmonySCALEXSCALEXHarmonypyHarmonypyscVIscVIUCEUCENo integration by Bat…No integration by BatchCombatCombatscANVIscANVIscVI (MLflow model)scVI (MLflow model)No integrationNo integrationShuffle integration b…Shuffle integration by batchbatchelor mnnCorrectbatchelor mnnCorrectmnnpymnnpySCimilaritySCimilarityscGPT (MLflow model)scGPT (MLflow model)scGPT (zero shot)scGPT (zero shot)TranscriptFormer (MLf…TranscriptFormer (MLflow model)scGPT (fine-tuned)scGPT (fine-tuned)UCE (MLflow model)UCE (MLflow model)00.2380.4760.7130.95100.250.50.751rawscaled
  • kBET pegasus labelhigher better
    Perfect embedding by …Perfect embedding by celltype with jitterEmbed cell typesEmbed cell typesShuffle integration b…Shuffle integration by cell typeShuffle integrationShuffle integrationpyligerpyligerLIGERLIGERScanoramaScanoramascPRINTscPRINTbatchelor fastMNNbatchelor fastMNNHarmonypyHarmonypyHarmonyHarmonyscANVIscANVIUCEUCEShuffle integration b…Shuffle integration by batchscVIscVIscVI (MLflow model)scVI (MLflow model)SCALEXSCALEXNo integrationNo integrationCombatCombatscGPT (MLflow model)scGPT (MLflow model)SCimilaritySCimilarityscGPT (zero shot)scGPT (zero shot)TranscriptFormer (MLf…TranscriptFormer (MLflow model)No integration by Bat…No integration by BatchscGPT (fine-tuned)scGPT (fine-tuned)batchelor mnnCorrectbatchelor mnnCorrectUCE (MLflow model)UCE (MLflow model)mnnpymnnpy2.8e-30.240.4770.7140.9500.250.50.751rawscaled
  • NMIhigher better
    Embed cell typesEmbed cell typesPerfect embedding by …Perfect embedding by celltype with jitterbatchelor mnnCorrectbatchelor mnnCorrectscANVIscANVISCimilaritySCimilarityscVIscVIbatchelor fastMNNbatchelor fastMNNscGPT (MLflow model)scGPT (MLflow model)HarmonypyHarmonypyUCEUCEscGPT (zero shot)scGPT (zero shot)TranscriptFormer (MLf…TranscriptFormer (MLflow model)HarmonyHarmonyCombatCombatscVI (MLflow model)scVI (MLflow model)BBKNNBBKNNShuffle integration b…Shuffle integration by cell typeNo integrationNo integrationUCE (MLflow model)UCE (MLflow model)scGPT (fine-tuned)scGPT (fine-tuned)LIGERLIGERpyligerpyligerscPRINTscPRINTSCALEXSCALEXNo integration by Bat…No integration by BatchmnnpymnnpyScanoramaScanoramaShuffle integration b…Shuffle integration by batchShuffle integrationShuffle integration4.0e-40.250.50.75100.250.50.751rawscaled
  • NMI_batchhigher better
    Shuffle integrationShuffle integrationbatchelor mnnCorrectbatchelor mnnCorrectpyligerpyligerLIGERLIGERBBKNNBBKNNmnnpymnnpyHarmonyHarmonyHarmonypyHarmonypyShuffle integration b…Shuffle integration by cell typescPRINTscPRINTScanoramaScanoramascGPT (fine-tuned)scGPT (fine-tuned)SCALEXSCALEXbatchelor fastMNNbatchelor fastMNNscVI (MLflow model)scVI (MLflow model)scGPT (zero shot)scGPT (zero shot)SCimilaritySCimilarityEmbed cell typesEmbed cell typesPerfect embedding by …Perfect embedding by celltype with jitterUCE (MLflow model)UCE (MLflow model)scGPT (MLflow model)scGPT (MLflow model)scVIscVIscANVIscANVIUCEUCECombatCombatTranscriptFormer (MLf…TranscriptFormer (MLflow model)No integrationNo integrationShuffle integration b…Shuffle integration by batchNo integration by Bat…No integration by Batch0.3420.5060.6710.835100.250.50.751rawscaled
  • PCRhigher better
    No integration by Bat…No integration by BatchCombatCombatShuffle integrationShuffle integrationSCALEXSCALEXpyligerpyligerLIGERLIGERmnnpymnnpyHarmonyHarmonyHarmonypyHarmonypybatchelor mnnCorrectbatchelor mnnCorrectShuffle integration b…Shuffle integration by cell typescVIscVIscANVIscANVIScanoramaScanoramaEmbed cell typesEmbed cell typesPerfect embedding by …Perfect embedding by celltype with jitterbatchelor fastMNNbatchelor fastMNNscPRINTscPRINTTranscriptFormer (MLf…TranscriptFormer (MLflow model)SCimilaritySCimilarityscGPT (zero shot)scGPT (zero shot)scGPT (MLflow model)scGPT (MLflow model)scVI (MLflow model)scVI (MLflow model)UCEUCENo integrationNo integrationscGPT (fine-tuned)scGPT (fine-tuned)Shuffle integration b…Shuffle integration by batchUCE (MLflow model)UCE (MLflow model)00.250.50.75100.250.50.751rawscaled
QC: Indicator table 31 errors21 warnings2 silenced

Automated checks on the benchmark run and its results: missing values, score scaling, metric ranges and similar. Errors are high-severity issues that usually need a maintainer's attention; warnings are lower-severity signals. Findings that are expected for this task are listed separately as silenced.

31 high-severity issues need review. 997 of 1051 checks passed.

  • error Raw results Dataset 'cellxgene_census/hypomap' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Dataset: cellxgene_census/hypomap Number of results: 314 Expected number of results: 496 Percentage missing: 37%

  • error Raw results Dataset 'cellxgene_census/mouse_pancreas_atlas' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Dataset: cellxgene_census/mouse_pancreas_atlas Number of results: 314 Expected number of results: 496 Percentage missing: 37%

  • error Raw results Method 'batchelor_mnn_correct' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Method: batchelor_mnn_correct Number of results: 14 Expected number of results: 96 Percentage missing: 85%

  • error Raw results Method 'bbknn' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Method: bbknn Number of results: 45 Expected number of results: 96 Percentage missing: 53%

  • error Raw results Method 'geneformer' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Method: geneformer Number of results: 0 Expected number of results: 96 Percentage missing: 100%

  • error Raw results Method 'geneformer_mlflow' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Method: geneformer_mlflow Number of results: 0 Expected number of results: 96 Percentage missing: 100%

  • error Raw results Method 'mnnpy' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Method: mnnpy Number of results: 14 Expected number of results: 96 Percentage missing: 85%

  • error Raw results Method 'scgpt_finetuned' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Method: scgpt_finetuned Number of results: 13 Expected number of results: 96 Percentage missing: 86%

  • error Raw results Method 'scgpt_mlflow' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Method: scgpt_mlflow Number of results: 58 Expected number of results: 96 Percentage missing: 40%

  • error Raw results Method 'scgpt_zeroshot' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Method: scgpt_zeroshot Number of results: 58 Expected number of results: 96 Percentage missing: 40%

  • error Raw results Method 'scimilarity' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Method: scimilarity Number of results: 58 Expected number of results: 96 Percentage missing: 40%

  • error Raw results Method 'transcriptformer_mlflow' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Method: transcriptformer_mlflow Number of results: 57 Expected number of results: 96 Percentage missing: 41%

  • error Raw results Method 'uce_mlflow' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Method: uce_mlflow Number of results: 13 Expected number of results: 96 Percentage missing: 86%

  • error Raw results Metric 'isolated_label_asw' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: isolated_label_asw Number of results: 112 Expected number of results: 186 Percentage missing: 40%

  • error Raw results Metric 'isolated_label_f1' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: isolated_label_f1 Number of results: 117 Expected number of results: 186 Percentage missing: 37%

  • error Raw results Dataset 'cellxgene_census/hypomap' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Dataset: cellxgene_census/hypomap Succeeded processes: 21 Attempted processes: 31 Percentage failed: 32%

  • error Raw results Dataset 'cellxgene_census/mouse_pancreas_atlas' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Dataset: cellxgene_census/mouse_pancreas_atlas Succeeded processes: 21 Attempted processes: 31 Percentage failed: 32%

  • error Raw results Method 'batchelor_mnn_correct' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Method: batchelor_mnn_correct Succeeded processes: 1 Attempted processes: 6 Percentage failed: 83%

  • error Raw results Method 'geneformer' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Method: geneformer Succeeded processes: 0 Attempted processes: 6 Percentage failed: 100%

  • error Raw results Method 'geneformer_mlflow' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Method: geneformer_mlflow Succeeded processes: 0 Attempted processes: 6 Percentage failed: 100%

  • error Raw results Method 'mnnpy' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Method: mnnpy Succeeded processes: 1 Attempted processes: 6 Percentage failed: 83%

  • error Raw results Method 'scgpt_finetuned' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Method: scgpt_finetuned Succeeded processes: 1 Attempted processes: 6 Percentage failed: 83%

  • error Raw results Method 'scgpt_mlflow' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Method: scgpt_mlflow Succeeded processes: 4 Attempted processes: 6 Percentage failed: 33%

  • error Raw results Method 'scgpt_zeroshot' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Method: scgpt_zeroshot Succeeded processes: 4 Attempted processes: 6 Percentage failed: 33%

  • error Raw results Method 'scimilarity' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Method: scimilarity Succeeded processes: 4 Attempted processes: 6 Percentage failed: 33%

  • error Raw results Method 'transcriptformer_mlflow' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Method: transcriptformer_mlflow Succeeded processes: 4 Attempted processes: 6 Percentage failed: 33%

  • error Raw results Method 'uce_mlflow' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Method: uce_mlflow Succeeded processes: 1 Attempted processes: 6 Percentage failed: 83%

  • error Raw results Metric 'isolated_label_asw' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: batch_integration Metric: isolated_label_asw Control method scores: 35 Expected control method scores: 42 Percentage succeeded: 83%

  • error Raw results Metric 'isolated_label_f1' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: batch_integration Metric: isolated_label_f1 Control method scores: 35 Expected control method scores: 42 Percentage succeeded: 83%

  • error Scaling Metric 'hvg_overlap' % outside range

    Percentage of scaled scores outside control range should be less than 10% Task: batch_integration Metric: hvg_overlap Inside range: 0 Scaled scores: 38 Percentage outside: 100%

  • error Scaling Metric 'hvg_overlap' worst score % outside range

    The worst scaled score should be less than 10% outside the control range Task: batch_integration Metric: hvg_overlap Worst score: -0.62546283110225 Percentage outside range: 63%

Show 21 warnings
  • warning Raw results Task number of results

    Number of results should be equal to #datasets × #methods × #metrics Task: batch_integration Number of results: 2126 Number of datasets: 6 Number of methods: 31 Number of metrics: 16 Expected number of results: 2976

  • warning Raw results Dataset 'cellxgene_census/tabula_sapiens' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Dataset: cellxgene_census/tabula_sapiens Number of results: 371 Expected number of results: 496 Percentage missing: 25%

  • warning Raw results Dataset 'cellxgene_census/dkd' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Dataset: cellxgene_census/dkd Number of results: 379 Expected number of results: 496 Percentage missing: 24%

  • warning Raw results Dataset 'cellxgene_census/gtex_v9' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Dataset: cellxgene_census/gtex_v9 Number of results: 374 Expected number of results: 496 Percentage missing: 25%

  • warning Raw results Dataset 'cellxgene_census/immune_cell_atlas' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Dataset: cellxgene_census/immune_cell_atlas Number of results: 374 Expected number of results: 496 Percentage missing: 25%

  • warning Raw results Metric 'asw_batch' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: asw_batch Number of results: 140 Expected number of results: 186 Percentage missing: 25%

  • warning Raw results Metric 'asw_label' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: asw_label Number of results: 139 Expected number of results: 186 Percentage missing: 25%

  • warning Raw results Metric 'cell_cycle_conservation' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: cell_cycle_conservation Number of results: 140 Expected number of results: 186 Percentage missing: 25%

  • warning Raw results Metric 'ari' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: ari Number of results: 146 Expected number of results: 186 Percentage missing: 22%

  • warning Raw results Metric 'nmi' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: nmi Number of results: 146 Expected number of results: 186 Percentage missing: 22%

  • warning Raw results Metric 'ari_batch' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: ari_batch Number of results: 146 Expected number of results: 186 Percentage missing: 22%

  • warning Raw results Metric 'nmi_batch' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: nmi_batch Number of results: 146 Expected number of results: 186 Percentage missing: 22%

  • warning Raw results Metric 'graph_connectivity' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: graph_connectivity Number of results: 146 Expected number of results: 186 Percentage missing: 22%

  • warning Raw results Metric 'kbet_pg' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: kbet_pg Number of results: 140 Expected number of results: 186 Percentage missing: 25%

  • warning Raw results Metric 'kbet_pg_label' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: kbet_pg_label Number of results: 140 Expected number of results: 186 Percentage missing: 25%

  • warning Raw results Metric 'ilisi' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: ilisi Number of results: 145 Expected number of results: 186 Percentage missing: 22%

  • warning Raw results Metric 'clisi' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: clisi Number of results: 145 Expected number of results: 186 Percentage missing: 22%

  • warning Raw results Metric 'pcr' % missing

    Percentage of missing results should be less than 10% Task: batch_integration Metric: pcr Number of results: 140 Expected number of results: 186 Percentage missing: 25%

  • warning Raw results Dataset 'cellxgene_census/tabula_sapiens' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Dataset: cellxgene_census/tabula_sapiens Succeeded processes: 25 Attempted processes: 31 Percentage failed: 19%

  • warning Raw results Dataset 'cellxgene_census/gtex_v9' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Dataset: cellxgene_census/gtex_v9 Succeeded processes: 25 Attempted processes: 31 Percentage failed: 19%

  • warning Raw results Dataset 'cellxgene_census/immune_cell_atlas' % failed

    Percentage of failed processes should be less than 10% Task: batch_integration Dataset: cellxgene_census/immune_cell_atlas Succeeded processes: 25 Attempted processes: 31 Percentage failed: 19%

Show 2 silenced (expected) checks
  • silenced Raw results Metric 'hvg_overlap' % missing

    Expected: v3.0.0 bundles the embed, graph and feature sub-tasks. hvg_overlap only applies to feature-output methods, so it is missing for embed and graph methods by design.

    Percentage of missing results should be less than 10% Task: batch_integration Metric: hvg_overlap Number of results: 38 Expected number of results: 186 Percentage missing: 80%

  • silenced Raw results Metric 'hvg_overlap' number of control methods

    Expected: Only the feature-output control methods produce a corrected feature matrix, so hvg_overlap sees fewer controls than datasets x control methods.

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: batch_integration Metric: hvg_overlap Control method scores: 18 Expected control method scores: 42 Percentage succeeded: 43%

Method info 24

Fast mutual nearest neighbors correction

The fastMNN() approach is much simpler than the original mnnCorrect() algorithm, and proceeds in several steps.

  1. Perform a multi-sample PCA on the (cosine-)normalized expression values to reduce dimensionality.
  2. Identify MNN pairs in the low-dimensional space between a reference batch and a target batch.
  3. Remove variation along the average batch vector in both reference and target batches.
  4. Correct the cells in the target batch towards the reference, using locally weighted correction vectors.
  5. Merge the corrected target batch with the reference, and repeat with the next target batch.

Mutual nearest neighbors correction

Correct for batch effects in single-cell expression data using the mutual nearest neighbors method.

BBKNN creates k nearest neighbours graph by identifying neighbours within batches, then combining and processing them with UMAP for visualization.

BBKNN or batch balanced k nearest neighbours graph is built for each cell by identifying its k nearest neighbours within each defined batch separately, creating independent neighbour sets for each cell in each batch. These sets are then combined and processed with the UMAP algorithm for visualisation.

Adjusting batch effects in microarray expression data using empirical Bayes methods

An Empirical Bayes (EB) approach to correct for batch effects. It estimates batch-specific parameters by pooling information across genes in each batch and shrinks the estimates towards the overall mean of the batch effect estimates across all genes. These parameters are then used to adjust the data for batch effects, leading to more accurate and reproducible results.

Geneformer is a foundation transformer model pretrained on a large-scale corpus of single cell transcriptomes

Geneformer is a foundation transformer model pretrained on a large-scale corpus of single cell transcriptomes to enable context-aware predictions in network biology. For this task, Geneformer is used to create a batch-corrected cell embedding.

Geneformer is a foundation transformer model pretrained on a large-scale corpus of single cell transcriptomes

Geneformer is a foundation transformer model pretrained on a large-scale corpus of single cell transcriptomes to enable context-aware predictions in network biology. For this task, Geneformer is used to create a batch-corrected cell embedding.

Here, we use a version packaged as an MLflow model.

Fast, sensitive and accurate integration of single-cell data with Harmony

Harmony is a general-purpose R package with an efficient algorithm for integrating multiple data sets. It is especially useful for large single-cell datasets such as single-cell RNA-seq.

harmonypy is a port of the harmony R package by Ilya Korsunsky.

Harmony is a general-purpose R package with an efficient algorithm for integrating multiple data sets. It is especially useful for large single-cell datasets such as single-cell RNA-seq.

Linked Inference of Genomic Experimental Relationships

LIGER or linked inference of genomic experimental relationships uses iNMF deriving and implementing a novel coordinate descent algorithm to efficiently do the factorization. Joint clustering is performed and factor loadings are normalised.

Batch effect correction by matching mutual nearest neighbors, Python implementation.

An implementation of MNN correct in python featuring low memory usage, full multicore support and compatibility with the scanpy framework.

Batch effect correction by matching mutual nearest neighbors (Haghverdi et al, 2018) has been implemented as a function 'mnnCorrect' in the R package scran. Sadly it's extremely slow for big datasets and doesn't make full use of the parallel architecture of modern CPUs.

This project is a python implementation of the MNN correct algorithm which takes advantage of python's extendability and hackability. It seamlessly integrates with the scanpy framework and has multicore support in its bones.

Python implementation of LIGER (Linked Inference of Genomic Experimental Relationships

LIGER (installed as rliger) is a package for integrating and analyzing multiple single-cell datasets, developed by the Macosko lab and maintained/extended by the Welch lab. It relies on integrative non-negative matrix factorization to identify shared and dataset-specific factors.

Online single-cell data integration through projecting heterogeneous datasets into a common cell-embedding space

SCALEX is a method for integrating heterogeneous single-cell data online using a VAE framework. Its generalised encoder disentangles batch-related components from batch-invariant biological components, which are then projected into a common cell-embedding space.

Efficient integration of heterogeneous single-cell transcriptomes using Scanorama

Scanorama enables batch-correction and integration of heterogeneous scRNA-seq datasets. It is designed to be used in scRNA-seq pipelines downstream of noise-reduction methods, including those for imputation and highly-variable gene filtering. The results from Scanorama integration and batch correction can then be used as input to other tools for scRNA-seq clustering, visualization, and analysis.

scANVI is a deep learning method that considers cell type labels.

scANVI (single-cell ANnotation using Variational Inference; Python class SCANVI) is a semi-supervised model for single-cell transcriptomics data. In a sense, it can be seen as a scVI extension that can leverage the cell type knowledge for a subset of the cells present in the data sets to infer the states of the rest of the cells.

A foundation model for single-cell biology (fine-tuned)

scGPT is a foundation model for single-cell biology based on a generative pre-trained transformer and trained on a repository of over 33 million cells.

Here, we fine-tune the pre-trained model for the batch integration task.

A foundation model for single-cell biology

scGPT is a foundation model for single-cell biology based on a generative pre-trained transformer and trained on a repository of over 33 million cells.

Here, we use a version packaged as an MLflow model.

A foundation model for single-cell biology (zero shot)

scGPT is a foundation model for single-cell biology based on a generative pre-trained transformer and trained on a repository of over 33 million cells.

Here, we use zero-shot output from a pre-trained model to get an integrated embedding for the batch integration task.

SCimilarity provides unifying representation of single cell expression profiles

SCimilarity is a unifying representation of single cell expression profiles that quantifies similarity between expression states and generalizes to represent new studies without additional training

scPRINT is a large transformer model built for the inference of gene networks

scPRINT is a large transformer model built for the inference of gene networks (connections between genes explaining the cell's expression profile) from scRNAseq data.

It uses novel encoding and decoding of the cell expression profile and new pre-training methodologies to learn a cell model.

scPRINT can be used to perform the following analyses:

  • expression denoising: increase the resolution of your scRNAseq data
  • cell embedding: generate a low-dimensional representation of your dataset
  • label prediction: predict the cell type, disease, sequencer, sex, and ethnicity of your cells
  • gene network inference: generate a gene network from any cell or cell cluster in your scRNAseq dataset

scVI combines a variational autoencoder with a hierarchical Bayesian model.

scVI combines a variational autoencoder with a hierarchical Bayesian model. It uses the negative binomial distribution to describe gene expression of each cell, conditioned on unobserved factors and the batch variable. ScVI is run as implemented in Luecken et al.

scVI combines a variational autoencoder with a hierarchical Bayesian model (MLflow model)

scVI combines a variational autoencoder with a hierarchical Bayesian model. It uses the negative binomial distribution to describe gene expression of each cell, conditioned on unobserved factors and the batch variable.

This version uses a pre-trained MLflow model.

TranscriptFormer (MLflow model)code ↗docs ↗source ↗Pearce et al., 2025

Context-aware representations of single-cell transcriptomes by jointly modeling genes and transcripts

TranscriptFormer is designed to learn rich, context-aware representations of single-cell transcriptomes while jointly modeling genes and transcripts using a novel generative architecture.

It is a family of generative foundation models representing a cross-species generative cell atlas trained on up to 112 million cells spanning 1.53 billion years of evolution across 12 species.

Here, we use a version packaged as an MLflow model.

UCE offers a unified biological latent space that can represent any cell

Universal Cell Embedding (UCE) is a single-cell foundation model that offers a unified biological latent space that can represent any cell, regardless of tissue or species

UCE offers a unified biological latent space that can represent any cell

Universal Cell Embedding (UCE) is a single-cell foundation model that offers a unified biological latent space that can represent any cell, regardless of tissue or species

Here, we use a version packaged as an MLflow model.

Control method info 7
Embed cell types

Cells are embedded as a one-hot encoding of celltype labels

No integration

Original feature space is not modified

No integration by Batch

Cells are embedded by computing PCA independently on each batch

Perfect embedding by celltype with jitter

Cells are embedded as a one-hot encoding of celltype labels, with a small amount of random noise added to the embedding

Shuffle integration

Integrations are randomly permuted

Shuffle integration by batch

Integrations are randomly permuted within each batch

Shuffle integration by cell type

Integrations are randomly permuted within each cell type

Metric info 16

Adjusted Rand Index compares clustering overlap, correcting for random labels and considering correct overlaps and disagreements.

The Adjusted Rand Index (ARI) compares the overlap of two clusterings; it considers both correct clustering overlaps while also counting correct disagreements between two clusterings. We compared the cell-type labels with the NMI-optimized Louvain clustering computed on the integrated dataset. The adjustment of the Rand index corrects for randomly correct labels. An ARI of 0 or 1 corresponds to random labeling or a perfect match, respectively.

This version of Adjusted Rand Index compares clustering overlap, correcting for batches and considering correct overlaps and disagreements.

The Adjusted Rand Index (ARI) compares the overlap of two clusterings; it considers both correct clustering overlaps while also counting correct disagreements between two clusterings. We compared the batches with the NMI-optimized Louvain clustering computed on the integrated dataset. The adjustment of the Rand index corrects for randomly correct labels. An ARI_batch of 0 or 1 corresponds to no batch correction or well corrected batches, respectively.

ASW batchhigher is betterLuecken et al., 2021

Modified average silhouette width (ASW) of batch

We consider the absolute silhouette width, s(i), on batch labels per cell i. Here, 0 indicates that batches are well mixed, and any deviation from 0 indicates a batch effect: 𝑠batch(𝑖)=|𝑠(𝑖)|.

To ensure higher scores indicate better batch mixing, these scores are scaled by subtracting them from 1. As we expect batches to integrate within cell identity clusters, we compute the batchASWj score for each cell label j separately, using the equation: batchASW𝑗=1|𝐶𝑗|∑𝑖∈𝐶𝑗1−𝑠batch(𝑖),

where Cj is the set of cells with the cell label j and |Cj| denotes the number of cells in that set.

To obtain the final batchASW score, the label-specific batchASWj scores are averaged: batchASW=1|𝑀|∑𝑗∈𝑀batchASW𝑗.

Here, M is the set of unique cell labels.

ASW Labelhigher is betterLuecken et al., 2021

Average silhouette of cell identity labels (cell types)

For the bio-conservation score, the ASW was computed on cell identity labels and scaled to a value between 0 and 1 using the equation: celltypeASW=(ASW_C+1)/2,

where C denotes the set of all cell identity labels. For information about the batch silhouette score, check sil_batch.

Cell Cycle Conservationhigher is betterLuecken et al., 2021

Cell cycle conservation score based on principle component regression on cell cycle gene scores

The cell-cycle conservation score evaluates how well the cell-cycle effect can be captured before and after integration. We computed cell-cycle scores using Scanpy's score_cell_cycle function with a reference gene set from Tirosh et al for the respective cell-cycle phases. We used the same set of cell-cycle genes for mouse and human data (using capitalization to convert between the gene symbols). We then computed the variance contribution of the resulting S and G2/M phase scores using principal component regression (Principal component regression), which was performed for each batch separately. The differences in variance before, Varbefore, and after, Varafter, integration were aggregated into a final score between 0 and 1, using the equation: CCconservation=1−|Varafter−Varbefore|/Varbefore.

In this equation, values close to 0 indicate lower conservation and 1 indicates complete conservation of the variance explained by cell cycle. In other words, the variance remains unchanged within each batch for complete conservation, while any deviation from the preintegration variance contribution reduces the score.

cLISIhigher is betterLuecken et al., 2021

Local inverse Simpson's Index

Local Inverse Simpson's Index metrics adapted from Korsunsky et al. 2019 to run on all full feature, embedding and kNN integration outputs via shortest path-based distance computation on single-cell kNN graphs. The metric assesses whether clusters of cells in a single-cell RNA-seq dataset are well-mixed across a categorical cell type variable.

The original LISI score ranges from 0 to the number of categories, with the latter indicating good cell mixing. This is rescaled to a score between 0 and 1.

Graph Connectivityhigher is betterLuecken et al., 2021

Connectivity of the subgraph per cell type label

The graph connectivity metric assesses whether the kNN graph representation, G, of the integrated data directly connects all cells with the same cell identity label. For each cell identity label c, we created the subset kNN graph G(Nc;Ec) to contain only cells from a given label. Using these subset kNN graphs, we computed the graph connectivity score using the equation:

gc =1/|C| Σc∈C |LCC(G(Nc;Ec))|/|Nc|.

Here, C represents the set of cell identity labels, |LCC()| is the number of nodes in the largest connected component of the graph, and |Nc| is the number of nodes with cell identity c. The resultant score has a range of (0;1], where 1 indicates that all cells with the same cell identity are connected in the integrated kNN graph, and the lowest possible score indicates a graph where no cell is connected. As this score is computed on the kNN graph, it can be used to evaluate all integration outputs.

HVG overlaphigher is betterLuecken et al., 2021

Overlap of highly variable genes per batch before and after integration.

The HVG conservation score is a proxy for the preservation of the biological signal. If the data integration method returned a corrected data matrix, we computed the number of HVGs before and after correction for each batch via Scanpy's highly_variable_genes function (using the 'cell ranger' flavor). If available, we computed 500 HVGs per batch. If fewer than 500 genes were present in the integrated object for a batch, the number of HVGs was set to half the total genes in that batch. The overlap coefficient is as follows: overlap(𝑋,𝑌)=|𝑋∩𝑌|/min(|𝑋|,|𝑌|),

where X and Y denote the fraction of preserved informative genes. The overall HVG score is the mean of the per-batch HVG overlap coefficients.

iLISIhigher is betterLuecken et al., 2021

Local inverse Simpson's Index

Local Inverse Simpson's Index metrics adapted from Korsunsky et al. 2019 to run on all full feature, embedding and kNN integration outputs via shortest path-based distance computation on single-cell kNN graphs. The metric assesses whether clusters of cells in a single-cell RNA-seq dataset are well-mixed across a categorical batch variable.

The original LISI score ranges from 0 to the number of categories, with the latter indicating good cell mixing. This is rescaled to a score between 0 and 1.

Isolated label ASWhigher is betterLuecken et al., 2021

Evaluate how well isolated labels separate by average silhouette width

Isolated cell labels are defined as the labels present in the least number of batches in the integration task. The score evaluates how well these isolated labels separate from other cell identities.

The isolated label ASW score is obtained by computing the ASW of isolated versus non-isolated labels on the PCA embedding (ASW metric above) and scaling this score to be between 0 and 1. The final score for each metric version consists of the mean isolated score of all isolated labels.

Isolated label F1 scorehigher is betterLuecken et al., 2021

Evaluate how well isolated labels coincide with clusters

We developed two isolated label scores to evaluate how well the data integration methods dealt with cell identity labels shared by few batches. Specifically, we identified isolated cell labels as the labels present in the least number of batches in the integration task. The score evaluates how well these isolated labels separate from other cell identities. We implemented the isolated label metric in two versions: (1) the best clustering of the isolated label (F1 score) and (2) the global ASW of the isolated label. For the cluster-based score, we first optimize the cluster assignment of the isolated label using the F1 score˚ across louvain clustering resolutions ranging from 0.1 to 2 in resolution steps of 0.1. The optimal F1 score for the isolated label is then used as the metric score. The F1 score is a weighted mean of precision and recall given by the equation: 𝐹1=2×(precision×recall)/(precision+recall).

It returns a value between 0 and 1, where 1 shows that all of the isolated label cells and no others are captured in the cluster. For the isolated label ASW score, we compute the ASW of isolated versus nonisolated labels on the PCA embedding (ASW metric above) and scale this score to be between 0 and 1. The final score for each metric version consists of the mean isolated score of all isolated labels.

kBET pegasushigher is betterLuecken et al., 2021

kBET algorithm to determine how well batches are mixed within a cell type

The kBET algorithm (v.0.99.6, release 4c9dafa) determines whether the label composition of a k nearest neighborhood of a cell is similar to the expected (global) label composition (Buettner et al., Nat Meth 2019). The test is repeated for a random subset of cells, and the results are summarized as a rejection rate over all tested neighborhoods.

This implementation uses the pegasus.calc_kBET function.

In Open Problems we do not run kBET on graph outputs to avoid computation-intensive diffusion processes being run.

kBET pegasus labelhigher is betterLuecken et al., 2021

kBET algorithm to determine how well batches are mixed within a cell type

The kBET algorithm (v.0.99.6, release 4c9dafa) determines whether the label composition of a k nearest neighborhood of a cell is similar to the expected (global) label composition (Buettner et al., Nat Meth 2019). The test is repeated for a random subset of cells, and the results are summarized as a rejection rate over all tested neighborhoods.

This implementation uses the pegasus.calc_kBET function per cell type and the rejection rate is averaged over all cell types.

In Open Problems we do not run kBET on graph outputs to avoid computation-intensive diffusion processes being run.

NMI compares overlap by scaling using mean entropy terms and optimizing Louvain clustering to obtain the best match between clusters and labels.

Normalized Mutual Information (NMI) compares the overlap of two clusterings. We used NMI to compare the cell-type labels with Louvain clusters computed on the integrated dataset. The overlap was scaled using the mean of the entropy terms for cell-type and cluster labels. Thus, NMI scores of 0 or 1 correspond to uncorrelated clustering or a perfect match, respectively. We performed optimized Louvain clustering for this metric to obtain the best match between clusters and labels.

This version of NMI compares overlap by scaling using mean entropy terms and optimizing Louvain clustering to obtain the best outcome of batch correction.

Normalized Mutual Information (NMI) compares the overlap of two clusterings. We used NMI to compare the batches with Louvain clusters computed on the integrated dataset. The overlap was scaled using the mean of the entropy terms for cell-type and cluster labels, then subracted from 1. Thus, NMI_batch scores of 0 or 1 correspond to no batch correction or well corrected batches, respectively. We performed optimized Louvain clustering for this metric to obtain the best outcome of batch correction.

PCRhigher is betterLuecken et al., 2021

Compare explained variance by batch before and after integration

Principal component regression, derived from PCA, has previously been used to quantify batch removal. Briefly, the R2 was calculated from a linear regression of the covariate of interest (for example, the batch variable B) onto each principal component. The variance contribution of the batch effect per principal component was then calculated as the product of the variance explained by the ith principal component (PC) and the corresponding R2(PCi|B). The sum across all variance contributions by the batch effects in all principal components gives the total variance explained by the batch variable as follows: Var(𝐶|𝐵)=∑𝑖=1𝐺Var(𝐶|PC𝑖)×𝑅2(PC𝑖|𝐵),

where Var(C|PCi) is the variance of the data matrix C explained by the ith principal component.

Dataset info 6

Multimodal single cell sequencing implicates chromatin accessibility and genetic background in diabetic kidney disease progression

Multimodal single cell sequencing is a powerful tool for interrogating cell-specific changes in transcription and chromatin accessibility. We performed single nucleus RNA (snRNA-seq) and assay for transposase accessible chromatin sequencing (snATAC-seq) on human kidney cortex from donors with and without diabetic kidney disease (DKD) to identify altered signaling pathways and transcription factors associated with DKD. Both snRNA-seq and snATAC-seq had an increased proportion of VCAM1+ injured proximal tubule cells (PT_VCAM1) in DKD samples. PT_VCAM1 has a pro-inflammatory expression signature and transcription factor motif enrichment implicated NFkB signaling. We used stratified linkage disequilibrium score regression to partition heritability of kidney-function-related traits using publicly-available GWAS summary statistics. Cell-specific PT_VCAM1 peaks were enriched for heritability of chronic kidney disease (CKD), suggesting that genetic background may regulate chromatin accessibility and DKD progression. snATAC-seq found cell-specific differentially accessible regions (DAR) throughout the nephron that change accessibility in DKD and these regions were enriched for glucocorticoid receptor (GR) motifs. Changes in chromatin accessibility were associated with decreased expression of insulin receptor, increased gluconeogenesis, and decreased expression of the GR cytosolic chaperone, FKBP5, in the diabetic proximal tubule. Cleavage under targets and release using nuclease (CUT&RUN) profiling of GR binding in bulk kidney cortex and an in vitro model of the proximal tubule (RPTEC) showed that DAR co-localize with GR binding sites. CRISPRi silencing of GR response elements (GRE) in the FKBP5 gene body reduced FKBP5 expression in RPTEC, suggesting that reduced FKBP5 chromatin accessibility in DKD may alter cellular response to GR. We developed an open-source tool for single cell allele specific analysis (SALSA) to model the effect of genetic background on gene expression. Heterozygous germline single nucleotide variants (SNV) in proximal tubule ATAC peaks were associated with allele-specific chromatin accessibility and differential expression of target genes within cis-coaccessibility networks. Partitioned heritability of proximal tubule ATAC peaks with a predicted allele-specific effect was enriched for eGFR, suggesting that genetic background may modify DKD progression in a cell-specific manner.

Single-nucleus cross-tissue molecular reference maps to decipher disease gene function

Understanding the function of genes and their regulation in tissue homeostasis and disease requires knowing the cellular context in which genes are expressed in tissues across the body. Single cell genomics allows the generation of detailed cellular atlases in human tissues, but most efforts are focused on single tissue types. Here, we establish a framework for profiling multiple tissues across the human body at single-cell resolution using single nucleus RNA-Seq (snRNA-seq), and apply it to 8 diverse, archived, frozen tissue types (three donors per tissue). We apply four snRNA-seq methods to each of 25 samples from 16 donors, generating a cross-tissue atlas of 209,126 nuclei profiles, and benchmark them vs. scRNA-seq of comparable fresh tissues. We use a conditional variational autoencoder (cVAE) to integrate an atlas across tissues, donors, and laboratory methods. We highlight shared and tissue-specific features of tissue-resident immune cells, identifying tissue-restricted and non-restricted resident myeloid populations. These include a cross-tissue conserved dichotomy between LYVE1- and HLA class II-expressing macrophages, and the broad presence of LAM-like macrophages across healthy tissues that is also observed in disease. For rare, monogenic muscle diseases, we identify cell types that likely underlie the neuromuscular, metabolic, and immune components of these diseases, and biological processes involved in their pathology. For common complex diseases and traits analyzed by GWAS, we identify the cell types and gene modules that potentially underlie disease mechanisms. The experimental and analytical frameworks we describe will enable the generation of large-scale studies of how cellular and molecular processes vary across individuals and populations.

A unified single cell gene expression atlas of the murine hypothalamus

The hypothalamus plays a key role in coordinating fundamental body functions. Despite recent progress in single-cell technologies, a unified catalogue and molecular characterization of the heterogeneous cell types and, specifically, neuronal subtypes in this brain region are still lacking. Here we present an integrated reference atlas “HypoMap” of the murine hypothalamus consisting of 384,925 cells, with the ability to incorporate new additional experiments. We validate HypoMap by comparing data collected from SmartSeq2 and bulk RNA sequencing of selected neuronal cell types with different degrees of cellular heterogeneity.

Cross-tissue immune cell analysis reveals tissue-specific features in humans

Despite their crucial role in health and disease, our knowledge of immune cells within human tissues remains limited. We surveyed the immune compartment of 16 tissues from 12 adult donors by single-cell RNA sequencing and VDJ sequencing generating a dataset of ~360,000 cells. To systematically resolve immune cell heterogeneity across tissues, we developed CellTypist, a machine learning tool for rapid and precise cell type annotation. Using this approach, combined with detailed curation, we determined the tissue distribution of finely phenotyped immune cell types, revealing hitherto unappreciated tissue-specific features and clonal architecture of T and B cells. Our multitissue approach lays the foundation for identifying highly resolved immune cell types by leveraging a common reference dataset, tissue-integrated expression analysis, and antigen receptor sequencing.

Mouse pancreatic islet scRNA-seq atlas across sexes, ages, and stress conditions including diabetes

To better understand pancreatic β-cell heterogeneity we generated a mouse pancreatic islet atlas capturing a wide range of biological conditions. The atlas contains scRNA-seq datasets of over 300,000 mouse pancreatic islet cells, of which more than 100,000 are β-cells, from nine datasets with 56 samples, including two previously unpublished datasets. The samples vary in sex, age (ranging from embryonic to aged), chemical stress, and disease status (including T1D NOD model development and two T2D models, mSTZ and db/db) together with different diabetes treatments. Additional information about data fields is available in anndata uns field 'field_descriptions' and on https://github.com/theislab/mm_pancreas_atlas_rep/blob/main/resources/cellxgene.md.

A multiple-organ, single-cell transcriptomic atlas of humans

Tabula Sapiens is a benchmark, first-draft human cell atlas of nearly 500,000 cells from 24 organs of 15 normal human subjects. This work is the product of the Tabula Sapiens Consortium. Taking the organs from the same individual controls for genetic background, age, environment, and epigenetic effects and allows detailed analysis and comparison of cell types that are shared between tissues. Our work creates a detailed portrait of cell types as well as their distribution and variation in gene expression across tissues and within the endothelial, epithelial, stromal and immune compartments.

References

  1. Amelio, A., & Pizzuti, C. (2015). Is Normalized Mutual Information a Fair Measure for Comparing Community Detection Methods? Proceedings of the 2015 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining 2015. 10.1145/2808797.2809344 ↗
  2. Chazarra-Gil, R., van Dongen, S., Kiselev, V. Y., & Hemberg, M. (2021). Flexible comparison of batch correction methods for single-cell RNA-seq using BatchBench. Nucleic Acids Research, 49(7), e42–e42. 10.1093/nar/gkab004 ↗
  3. Chen, H., Venkatesh, M. S., Gomez Ortega, J., Mahesh, S. V., Nandi, T. N., Madduri, R. K., Pelka, K., & Theodoris, C. V. (2024). Quantized multi-task learning for context-specific representations of gene network dynamics. 10.1101/2024.08.16.608180 ↗
  4. Domínguez Conde, C., Xu, C., Jarvis, L. B., Rainbow, D. B., Wells, S. B., Gomes, T., Howlett, S. K., Suchanek, O., Polanski, K., King, H. W., Mamanova, L., Huang, N., Szabo, P. A., Richardson, L., Bolt, L., Fasouli, E. S., Mahbubani, K. T., Prete, M., Tuck, L., … Teichmann, S. A. (2022). Cross-tissue immune cell analysis reveals tissue-specific features in humans. Science, 376(6594). 10.1126/science.abl5197 ↗
  5. Cui, H., Wang, C., Maan, H., Pang, K., Luo, F., Duan, N., & Wang, B. (2024). scGPT: toward building a foundation model for single-cell multi-omics using generative AI. 10.1038/s41592-024-02201-0 ↗
  6. Eraslan, G., Drokhlyansky, E., Anand, S., Fiskin, E., Subramanian, A., Slyper, M., Wang, J., Van Wittenberghe, N., Rouhana, J. M., Waldman, J., Ashenberg, O., Lek, M., Dionne, D., Win, T. S., Cuoco, M. S., Kuksenko, O., Tsankov, A. M., Branton, P. A., Marshall, J. L., … Regev, A. (2022). Single-nucleus cross-tissue molecular reference maps toward understanding disease gene function. Science, 376(6594). 10.1126/science.abl4290 ↗
  7. Haghverdi, L., Lun, A. T. L., Morgan, M. D., & Marioni, J. C. (2018). Batch effects in single-cell RNA-sequencing data are corrected by matching mutual nearest neighbors. Nature Biotechnology, 36(5), 421–427. 10.1038/nbt.4091 ↗
  8. Heimberg, G., Kuo, T., DePianto, D., Heigl, T., Diamant, N., Salem, O., Scalia, G., Biancalani, T., Turley, S., Rock, J., Bravo, H. C., Kaminker, J., Heiden, J. A. V., & Regev, A. (2023). Scalable querying of human cell atlases via a foundational model reveals commonalities across fibrosis-associated macrophages. 10.1101/2023.07.18.549537 ↗
  9. Hie, B., Bryson, B., & Berger, B. (2019). Efficient integration of heterogeneous single-cell transcriptomes using Scanorama. Nature Biotechnology, 37(6), 685–691. 10.1038/s41587-019-0113-3 ↗
  10. Hrovatin, K., Bastidas-Ponce, A., Bakhti, M., Zappia, L., Büttner, M., Sallino, C., Sterr, M., Böttcher, A., Migliorini, A., Lickert, H., & Theis, F. J. (2023). Delineating mouse β-cell identity during lifetime and in diabetes with a single cell atlas. bioRxiv. 10.1101/2022.12.22.521557 ↗
  11. Hubert, L., & Arabie, P. (1985). Comparing partitions. Journal of Classification, 2(1), 193–218. 10.1007/bf01908075 ↗
  12. Johnson, W. E., Li, C., & Rabinovic, A. (2006). Adjusting batch effects in microarray expression data using empirical Bayes methods. Biostatistics, 8(1), 118–127. 10.1093/biostatistics/kxj037 ↗
  13. Jones, R. C., Karkanias, J., Krasnow, M. A., Pisco, A. O., Quake, S. R., Salzman, J., Yosef, N., Bulthaup, B., Brown, P., Harper, W., Hemenez, M., Ponnusamy, R., Salehi, A., Sanagavarapu, B. A., Spallino, E., Aaron, K. A., Concepcion, W., Gardner, J. M., Kelly, B., … Wyss-Coray, T. (2022). The Tabula Sapiens: A multiple-organ, single-cell transcriptomic atlas of humans. Science, 376(6594). 10.1126/science.abl4896 ↗
  14. Kalfon, J., Samaran, J., Peyré, G., & Cantini, L. (2024). scPRINT: pre-training on 50 million cells allows robust gene network predictions. 10.1101/2024.07.29.605556 ↗
  15. Kang, C. (n.d.). mnnpy.
  16. Korsunsky, I., Millard, N., Fan, J., Slowikowski, K., Zhang, F., Wei, K., Baglaenko, Y., Brenner, M., ru Po-Loh, & Raychaudhuri, S. (2019). Fast, sensitive and accurate integration of single-cell data with Harmony. Nature Methods, 16(12), 1289–1296. 10.1038/s41592-019-0619-0 ↗
  17. Lopez, R., Regier, J., Cole, M. B., Jordan, M. I., & Yosef, N. (2018). Deep generative modeling for single-cell transcriptomics. Nature Methods, 15(12), 1053–1058. 10.1038/s41592-018-0229-2 ↗
  18. Luecken, M. D., Büttner, M., Chaichoompu, K., Danese, A., Interlandi, M., Mueller, M. F., Strobl, D. C., Zappia, L., Dugas, M., Colomé-Tatché, M., & Theis, F. J. (2021). Benchmarking atlas-level data integration in single-cell genomics. Nature Methods, 19(1), 41–50. 10.1038/s41592-021-01336-8 ↗
  19. Mereu, E., Lafzi, A., Moutinho, C., Ziegenhain, C., McCarthy, D. J., Alvarez-Varela, A., Batlle, E., Sagar, Gruen, D., Lau, J. K., & others. (2020). Benchmarking single-cell RNA-sequencing protocols for cell atlas projects. Nature Biotechnology, 38(6), 747–755. 10.1038/s41587-020-0469-4 ↗
  20. Pearce, J. D., Simmonds, S. E., Mahmoudabadi, G., Krishnan, L., Palla, G., Istrate, A.-M., Tarashansky, A., Nelson, B., Valenzuela, O., Li, D., Quake, S. R., & Karaletsos, T. (2025). A Cross-Species Generative Cell Atlas Across 1.5 Billion Years of Evolution: The TranscriptFormer Single-cell Model. 10.1101/2025.04.25.650731 ↗
  21. Polański, K., Young, M. D., Miao, Z., Meyer, K. B., Teichmann, S. A., & Park, J.-E. (2019). BBKNN: fast batch alignment of single cell transcriptomes. Bioinformatics. 10.1093/bioinformatics/btz625 ↗
  22. Rosen, Y., Roohani, Y., Agrawal, A., Samotorcan, L., Consortium, T. S., Quake, S. R., & Leskovec, J. (2023). Universal Cell Embeddings: A Foundation Model for Cell Biology. 10.1101/2023.11.28.568918 ↗
  23. Steuernagel, L., Lam, B. Y. H., Klemm, P., Dowsett, G. K. C., Bauder, C. A., Tadross, J. A., Hitschfeld, T. S., del Rio Martin, A., Chen, W., de Solis, A. J., Fenselau, H., Davidsen, P., Cimino, I., Kohnke, S. N., Rimmington, D., Coll, A. P., Beyer, A., Yeo, G. S. H., & Brüning, J. C. (2022). HypoMap—a unified single-cell gene expression atlas of the murine hypothalamus. Nature Metabolism, 4(10), 1402–1419. 10.1038/s42255-022-00657-y ↗
  24. Theodoris, C. V., Xiao, L., Chopra, A., Chaffin, M. D., Al Sayed, Z. R., Hill, M. C., Mantineo, H., Brydon, E. M., Zeng, Z., Liu, X. S., & Ellinor, P. T. (2023). Transfer learning enables predictions in network biology. 10.1038/s41586-023-06139-9 ↗
  25. Tran, H. T. N., Ang, K. S., Chevrier, M., Zhang, X., Lee, N. Y. S., Goh, M., & Chen, J. (2020). A benchmark of batch-effect correction methods for single-cell RNA sequencing data. Genome Biology, 21(1). 10.1186/s13059-019-1850-9 ↗
  26. Welch, J. D., Kozareva, V., Ferreira, A., Vanderburg, C., Martin, C., & Macosko, E. Z. (2019). Single-Cell Multi-omic Integration Compares and Contrasts Features of Brain Cell Identity. Cell, 177(7), 1873-1887.e17. 10.1016/j.cell.2019.05.006 ↗
  27. Wilson, P. C., Muto, Y., Wu, H., Karihaloo, A., Waikar, S. S., & Humphreys, B. D. (2022). Multimodal single cell sequencing implicates chromatin accessibility and genetic background in diabetic kidney disease progression. Nature Communications, 13(1). 10.1038/s41467-022-32972-z ↗
  28. Xiong, L., Tian, K., Li, Y., Ning, W., Gao, X., & Zhang, Q. C. (2022). Online single-cell data integration through projecting heterogeneous datasets into a common cell-embedding space. Nature Communications, 13(1). 10.1038/s41467-022-33758-z ↗
  29. Zappia, L., Phipson, B., & Oshlack, A. (2018). Exploring the single-cell RNA-seq analysis landscape with the scRNA-tools database. PLOS Computational Biology, 14(6), e1006245. 10.1371/journal.pcbi.1006245 ↗