Zebrafish
90k cells from zebrafish embryos throughout the first day of development, with and without a knockout of chordin, an important developmental gene. Dimensions: 26022 cells, 25258 genes. 24 cell types (avg. 1084±1156 cells per cell type).
Description
90k cells from zebrafish embryos throughout the first day of development, with and without a knockout of chordin, an important developmental gene. Dimensions: 26022 cells, 25258 genes. 24 cell types (avg. 1084±1156 cells per cell type).
Preview
An AnnData object with
n_obs × n_vars = 26,022 × 25,258 with slots:
Data structure
| Name | Description | Type | Data type | Size |
|---|---|---|---|---|
| obs | ||||
cell_type | Classification of the cell type based on its characteristics and function within the tissue or organism. | vector | category | 26022 |
batch | A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc. | vector | category | 26022 |
size_factors | The size factors created by the normalisation method, if any. | vector | float32 | 26022 |
| var | ||||
feature_name | A human-readable name for the feature, usually a gene symbol. | vector | object | 25258 |
hvg | Whether or not the feature is considered to be a 'highly variable gene' | vector | bool | 25258 |
hvg_score | A ranking of the features by hvg. | vector | float64 | 25258 |
| obsp | ||||
knn_connectivities | K nearest neighbors connectivities matrix. | sparsematrix | float32 | 26022 × 26022 |
knn_distances | K nearest neighbors distance matrix. | sparsematrix | float64 | 26022 × 26022 |
| obsm | ||||
X_pca | The resulting PCA embedding. | densematrix | float32 | 26022 × 50 |
| varm | ||||
pca_loadings | The PCA loadings matrix. | densematrix | float32 | 25258 × 50 |
| layers | ||||
counts | Raw counts | sparsematrix | float32 | 26022 × 25258 |
normalized | Normalised expression values | sparsematrix | float32 | 26022 × 25258 |
| uns | ||||
dataset_description | Long description of the dataset. | atomic | str | 1 |
dataset_id | A unique identifier for the dataset. This is different from the `obs.dataset_id` field, which is the identifier for the dataset from which the cell data is derived. | atomic | str | 1 |
dataset_name | A human-readable name for the dataset. | atomic | str | 1 |
dataset_organism | The organism of the sample in the dataset. | atomic | str | 1 |
dataset_reference | Bibtex reference of the paper in which the dataset was published. | atomic | str | 1 |
dataset_summary | Short description of the dataset. | atomic | str | 1 |
dataset_url | Link to the original source of the dataset. | atomic | str | 1 |
knn | Supplementary K nearest neighbors data. | dict | 3 | |
normalization_id | Which normalization was used | atomic | str | 1 |
pca_variance | The PCA variance objects. | dict | 2 | |
Download & explore
The processed dataset lives on the OpenProblems public S3 bucket, which is world-readable. Download the file directly, or copy its S3 URI to fetch it with your tool of choice.
s3://openproblems-data/resources/datasets/openproblems_v1/zebrafish/log_cp10k/dataset.h5ad Used in
References
- Wagner, D. E., Weinreb, C., Collins, Z. M., Briggs, J. A., Megason, S. G., & Klein, A. M. (2018). Single-cell mapping of gene expression landscapes and lineage in the zebrafish embryo. Science, 360(6392), 981–987. 10.1126/science.aar4362 ↗