All datasets

Zebrafish

90k cells from zebrafish embryos throughout the first day of development, with and without a knockout of chordin, an important developmental gene. Dimensions: 26022 cells, 25258 genes. 24 cell types (avg. 1084±1156 cells per cell type).

danio rerio Organism
26,022 × 25,258 Dimensions
777.92 MiB Size
log_cp10k Normalization

Description

90k cells from zebrafish embryos throughout the first day of development, with and without a knockout of chordin, an important developmental gene. Dimensions: 26022 cells, 25258 genes. 24 cell types (avg. 1084±1156 cells per cell type).

Preview

An AnnData object with n_obs × n_vars = 26,022 × 25,258 with slots:

Data structure

Name Description Type Data type Size
obs
cell_type Classification of the cell type based on its characteristics and function within the tissue or organism. vector category 26022
batch A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc. vector category 26022
size_factors The size factors created by the normalisation method, if any. vector float32 26022
var
feature_name A human-readable name for the feature, usually a gene symbol. vector object 25258
hvg Whether or not the feature is considered to be a 'highly variable gene' vector bool 25258
hvg_score A ranking of the features by hvg. vector float64 25258
obsp
knn_connectivities K nearest neighbors connectivities matrix. sparsematrix float32 26022 × 26022
knn_distances K nearest neighbors distance matrix. sparsematrix float64 26022 × 26022
obsm
X_pca The resulting PCA embedding. densematrix float32 26022 × 50
varm
pca_loadings The PCA loadings matrix. densematrix float32 25258 × 50
layers
counts Raw counts sparsematrix float32 26022 × 25258
normalized Normalised expression values sparsematrix float32 26022 × 25258
uns
dataset_description Long description of the dataset. atomic str 1
dataset_id A unique identifier for the dataset. This is different from the `obs.dataset_id` field, which is the identifier for the dataset from which the cell data is derived. atomic str 1
dataset_name A human-readable name for the dataset. atomic str 1
dataset_organism The organism of the sample in the dataset. atomic str 1
dataset_reference Bibtex reference of the paper in which the dataset was published. atomic str 1
dataset_summary Short description of the dataset. atomic str 1
dataset_url Link to the original source of the dataset. atomic str 1
knn Supplementary K nearest neighbors data. dict 3
normalization_id Which normalization was used atomic str 1
pca_variance The PCA variance objects. dict 2

Download & explore

The processed dataset lives on the OpenProblems public S3 bucket, which is world-readable. Download the file directly, or copy its S3 URI to fetch it with your tool of choice.

dataset.h5ad 777.92 MiB
Download
s3://openproblems-data/resources/datasets/openproblems_v1/zebrafish/log_cp10k/dataset.h5ad

Used in

References

  1. Wagner, D. E., Weinreb, C., Collins, Z. M., Briggs, J. A., Megason, S. G., & Klein, A. M. (2018). Single-cell mapping of gene expression landscapes and lineage in the zebrafish embryo. Science, 360(6392), 981–987. 10.1126/science.aar4362 ↗