merfish Search Results


90
Johns Hopkins HealthCare merfish transcriptional data
Top panel shows an iterative optimization scheme to estimate geometric transformation ( φ ) and latent feature distribution ( π ) by minimizing the normed difference (error) between the geometric and feature-transformed CCFv3 section to target <t>MERFISH.</t> Middle panel illustrates the application of estimated geometric transformation ( φ ) to deform the CCFv3 atlas to MERFISH coordinates (left); the application of latent feature distribution ( π ) to generate gene distributions on initial CCFv3 geometry (middle); and the application of inverse geometric transformation ( φ −1 ) to deform MERFISH genes to CCFv3 coordinates (right). Gene with the highest probability of expression at each location is shown as a MERFISH feature. Bottom panel illustrates the results of mapping CCFv3 sections to corresponding MERFISH sections. a , f 10 μm atlas sections at Z = 385 and Z = 485 out of 1320 visually chosen to match MERFISH architecture ( e , j ) rendered as meshes at 100 μm. e , j MERFISH sections rendered as meshes at 50 μm, with mRNA density depicted as a feature. b , g Geometric mappings ( φ ) of CCFv3 sections to MERFISH coordinates with the approximate determinant of the Jacobian showing areas of contraction (blue) and expansion (red). c , h Estimated mRNA density per atlas region ( \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${w}_{i}^{{\prime} }$$\end{document} w i ′ in ), as given by π shown in CCFv3 coordinates. d , i Estimated mRNA density shown per atlas region following geometric deformation to MERFISH coordinates.
Merfish Transcriptional Data, supplied by Johns Hopkins HealthCare, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/merfish/merfish+transcriptional+data/pmc11045777-98-25-33
Average 90 stars, based on 1 article reviews
merfish transcriptional data - by Bioz Stars, 2026-10
90/100 stars
  Buy from Supplier

90
NextGen Sciences merfish (multiplexed error-robust fluorescence in situ hybridization)
Top panel shows an iterative optimization scheme to estimate geometric transformation ( φ ) and latent feature distribution ( π ) by minimizing the normed difference (error) between the geometric and feature-transformed CCFv3 section to target <t>MERFISH.</t> Middle panel illustrates the application of estimated geometric transformation ( φ ) to deform the CCFv3 atlas to MERFISH coordinates (left); the application of latent feature distribution ( π ) to generate gene distributions on initial CCFv3 geometry (middle); and the application of inverse geometric transformation ( φ −1 ) to deform MERFISH genes to CCFv3 coordinates (right). Gene with the highest probability of expression at each location is shown as a MERFISH feature. Bottom panel illustrates the results of mapping CCFv3 sections to corresponding MERFISH sections. a , f 10 μm atlas sections at Z = 385 and Z = 485 out of 1320 visually chosen to match MERFISH architecture ( e , j ) rendered as meshes at 100 μm. e , j MERFISH sections rendered as meshes at 50 μm, with mRNA density depicted as a feature. b , g Geometric mappings ( φ ) of CCFv3 sections to MERFISH coordinates with the approximate determinant of the Jacobian showing areas of contraction (blue) and expansion (red). c , h Estimated mRNA density per atlas region ( \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${w}_{i}^{{\prime} }$$\end{document} w i ′ in ), as given by π shown in CCFv3 coordinates. d , i Estimated mRNA density shown per atlas region following geometric deformation to MERFISH coordinates.
Merfish (Multiplexed Error Robust Fluorescence In Situ Hybridization), supplied by NextGen Sciences, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/merfish/merfish++multiplexed+error+robust+fluorescence+in+situ+hybridization+/pm33351222-138-7-2
Average 90 stars, based on 1 article reviews
merfish (multiplexed error-robust fluorescence in situ hybridization) - by Bioz Stars, 2026-10
90/100 stars
  Buy from Supplier

90
Allen Institute for Brain Science merfish dataset
a, CellMemory can characterize single-cell spatial omics from various sequencing platforms, including CosMx, <t>MERFISH,</t> Slide-seq, Stereo-seq, Xenium, and Slide-tags. b, The accuracy of spatial annotation tools is evaluated using datasets from mHypo (mouse hypothalamus measured by MERFISH), hNSCLC (human NSCLC measured by CosMx), and msSermato (mouse spermatogenesis measured by Slide-seq). Each dataset comprised three samples for replication. c, CellMemory was trained using single-cell data at the L2 cell type resolution, to generate CLS embeddings and annotations for Slide-tags cells. The right part is plotted by the L1 (original) and L2 (CellMemory) cell types in spatial coordinates. d, CellMemory model was built using single-cell data at the L3 resolution, to integrate single-cell and Slide-tags data. e, The cells highlighted in the co-embedding are Slide-tags cells, labeled with Slide-tags cell, L1 (original), L2 (CellMemory), and L3 (CellMemory) cell identity. f, The identification of Slide-tags cells at L3 resolution by CellMemory (L4 IT_2 and Micro-PVM_1) is displayed (the first line), along with the expression (the second line) and memory score (the third line) of TAGs ( VWC2L, F13A1 ). The left half represents the spatial coordinate, and the right half is the UMAP coordinate. g, UMAP representation of 4 million mouse whole brain cells from MERFISH, colored by subclasses. h, Annotation benchmark comparison of CellMemory with other state-of-art methods, including scGPT, Geneformer, and CellTypist. The reference dataset consists of mouse whole-brain 10x single-cell data, the query set comprises MERFISH cells derived from 59 coronal sections. i, Visualization of the 5 (total 59) mouse whole-brain sections, with cells colored by predictions of CellMemory. j, Spatial coordinates of mouse brain MERFISH section 36, colored by CellMemory predictions. k, Heatmap displaying the memory scores of TAGs for all subclasses of IT-ET Glut in section 36. l, Distribution of subclasses and their corresponding TAGs’ memory scores within the spatial coordinates of section 36.
Merfish Dataset, supplied by Allen Institute for Brain Science, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/merfish/merfish+dataset/bio_rxiv__2024__12__17__628533-130-22-19
Average 90 stars, based on 1 article reviews
merfish dataset - by Bioz Stars, 2026-10
90/100 stars
  Buy from Supplier

90
Epigenomics ag merfish
a, CellMemory can characterize single-cell spatial omics from various sequencing platforms, including CosMx, <t>MERFISH,</t> Slide-seq, Stereo-seq, Xenium, and Slide-tags. b, The accuracy of spatial annotation tools is evaluated using datasets from mHypo (mouse hypothalamus measured by MERFISH), hNSCLC (human NSCLC measured by CosMx), and msSermato (mouse spermatogenesis measured by Slide-seq). Each dataset comprised three samples for replication. c, CellMemory was trained using single-cell data at the L2 cell type resolution, to generate CLS embeddings and annotations for Slide-tags cells. The right part is plotted by the L1 (original) and L2 (CellMemory) cell types in spatial coordinates. d, CellMemory model was built using single-cell data at the L3 resolution, to integrate single-cell and Slide-tags data. e, The cells highlighted in the co-embedding are Slide-tags cells, labeled with Slide-tags cell, L1 (original), L2 (CellMemory), and L3 (CellMemory) cell identity. f, The identification of Slide-tags cells at L3 resolution by CellMemory (L4 IT_2 and Micro-PVM_1) is displayed (the first line), along with the expression (the second line) and memory score (the third line) of TAGs ( VWC2L, F13A1 ). The left half represents the spatial coordinate, and the right half is the UMAP coordinate. g, UMAP representation of 4 million mouse whole brain cells from MERFISH, colored by subclasses. h, Annotation benchmark comparison of CellMemory with other state-of-art methods, including scGPT, Geneformer, and CellTypist. The reference dataset consists of mouse whole-brain 10x single-cell data, the query set comprises MERFISH cells derived from 59 coronal sections. i, Visualization of the 5 (total 59) mouse whole-brain sections, with cells colored by predictions of CellMemory. j, Spatial coordinates of mouse brain MERFISH section 36, colored by CellMemory predictions. k, Heatmap displaying the memory scores of TAGs for all subclasses of IT-ET Glut in section 36. l, Distribution of subclasses and their corresponding TAGs’ memory scores within the spatial coordinates of section 36.
Merfish, supplied by Epigenomics ag, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/merfish/merfish/pm40140573-546-0-6
Average 90 stars, based on 1 article reviews
merfish - by Bioz Stars, 2026-10
90/100 stars
  Buy from Supplier

90
Allen Institute for Brain Science merfish probe counts
Overall training and architectural scheme for CellTransformer. ( a. ) During training, a single cell is drawn (we denote this the reference cell, boxed in red). We extract the reference cell’s spatial neighbors and partition the group into a masked reference cell and its observed spatial neighbors. ( b. ) Initially, the model encoder receives information about each cell and projects those features to d- dimensional latent variable space. Features interact across cells (tokens) through the self-attention mechanism, which is repeated n times. These per-cell representations and an extra token are then aggregated into a single vector representation, which we refer to as the neighborhood representation. This representation is concatenated to a mask token which is cell type specific and chosen to represent the type of the reference cell. A shallow transformer decoder (dotted lines) further refines these representations and then a linear projection is used to output parameters of a negative binomial distribution modeling of the <t>MERFISH</t> probe counts for the reference cell.
Merfish Probe Counts, supplied by Allen Institute for Brain Science, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/merfish/merfish+probe+counts/bio_rxiv__2024__05__05__592608-165-4-24
Average 90 stars, based on 1 article reviews
merfish probe counts - by Bioz Stars, 2026-10
90/100 stars
  Buy from Supplier

90
Allen Institute for Brain Science multiplexed error-robust fluorescence situ hybridization (merfish
Overall training and architectural scheme for CellTransformer. ( a. ) During training, a single cell is drawn (we denote this the reference cell, boxed in red). We extract the reference cell’s spatial neighbors and partition the group into a masked reference cell and its observed spatial neighbors. ( b. ) Initially, the model encoder receives information about each cell and projects those features to d- dimensional latent variable space. Features interact across cells (tokens) through the self-attention mechanism, which is repeated n times. These per-cell representations and an extra token are then aggregated into a single vector representation, which we refer to as the neighborhood representation. This representation is concatenated to a mask token which is cell type specific and chosen to represent the type of the reference cell. A shallow transformer decoder (dotted lines) further refines these representations and then a linear projection is used to output parameters of a negative binomial distribution modeling of the <t>MERFISH</t> probe counts for the reference cell.
Multiplexed Error Robust Fluorescence Situ Hybridization (Merfish, supplied by Allen Institute for Brain Science, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/merfish/multiplexed+error+robust+fluorescence+situ+hybridization++merfish/pm38092914-66-15-24
Average 90 stars, based on 1 article reviews
multiplexed error-robust fluorescence situ hybridization (merfish - by Bioz Stars, 2026-10
90/100 stars
  Buy from Supplier

90
Bio-Synthesis Inc merfish readout probes
Images of nuclei (DAPI), total polyA mRNA, and two <t>MERFISH</t> bits were obtained using epifluorescence (Epi) and spinning-disk confocal microscopy respectively. Both epifluorescence and confocal images were taken with 1 s exposure time. The MERFISH bit-1 and bit-2 images were high-pass filtered to remove cellular background.
Merfish Readout Probes, supplied by Bio-Synthesis Inc, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/merfish/merfish+readout+probes/pmc11677232-177-0-20
Average 90 stars, based on 1 article reviews
merfish readout probes - by Bioz Stars, 2026-10
90/100 stars
  Buy from Supplier

90
Aries Systems Corporation single cell resolution merfish
Images of nuclei (DAPI), total polyA mRNA, and two <t>MERFISH</t> bits were obtained using epifluorescence (Epi) and spinning-disk confocal microscopy respectively. Both epifluorescence and confocal images were taken with 1 s exposure time. The MERFISH bit-1 and bit-2 images were high-pass filtered to remove cellular background.
Single Cell Resolution Merfish, supplied by Aries Systems Corporation, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/merfish/single+cell+resolution+merfish/pmc11258913__giae042_giga___d___23___00259_revision_1-72-32-26
Average 90 stars, based on 1 article reviews
single cell resolution merfish - by Bioz Stars, 2026-10
90/100 stars
  Buy from Supplier

86
Vizgen Inc merfish
A) Nicheformer is pretrained on the SpatialCorpus-110M, a large data collection of over 110 million cells measured with dissociated and image-based spatial transcriptomics technologies. The SpatialCorpus-110M collection comprises single-cell data from Homo Sapiens and Mus Musculus across 17 distinct organs, 18 cell lines, and additional single-cell data from other anatomical systems and junctions. Shown is an exemplary UMAP visualization of a random 1% subset of the entire pretraining dataset (n=1,108,759 cells) of the non-integrated log1p-transformed normalized SpatialCorpus-110M colored by modality. B) Nicheformer includes a novel set of downstream tasks, ranging from spatial cell type, niche and region label prediction to neighborhood cell density and neighborhood composition prediction. We test our approach on large-scale, high-quality spatial transcriptomics data from the brain (mouse - <t>MERFISH),</t> liver (CosMx - human), lung (CosMx - human, Xenium - human), and colon (Xenium - human). Visualized are example slices of the respective datasets colored by niche labels (brain, liver, and lung) and cell density (lung and colon). C) The SpatialCorpus-110M is harmonized and mapped to orthologous gene names, as well as human and mouse-specific genes, to create the input for Nicheformer pretraining. We harmonized metadata information across all datasets, capturing species, modality, and assay. D) Each cell’s gene expression profile and metadata are fed into a gene rank tokenizer to obtain a tokenized representation for each cell. The tokenized cells serve as input for the Nicheformer transformer block to predict masked tokens. Finally, the Nicheformer embedding is generated by aggregating the gene tokens (Methods). E) The pretrained Nicheformer embedding is visualized as UMAP colored by modality. The UMAP shows a random 5% subsample of the entire Nicheformer embedding (n=4,903,086).
Merfish, supplied by Vizgen Inc, used in various techniques. Bioz Stars score: 86/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/merfish/merfish+probes/bio_rxiv__2024__04__15__589472-90-9-10
Average 86 stars, based on 1 article reviews
merfish - by Bioz Stars, 2026-10
86/100 stars
  Buy from Supplier

86
Vizgen Inc merfish imaging
a XY directional map of all cell types identified within the striatal section is shown. Medium spiny neurons, cortical neurons, astrocytes, oligodendrocytes, endothelial cells, and microglia are displayed in different colors - each dot represents a cell segmented by MERLIN. All cells within the XY coordinates of the striatum were subset, and here we display the UMAP of these subset striatal cell clusters split by major cell types and UMAP split by age (young = blue; aged = red). b Striatal astrocytes were subset from all striatal cells and clustered separately using the Louvian algorithm. The first UMAP shows striatal astrocyte subtypes from single cells, the second UMAP shows striatal astrocyte subtypes from <t>MERFISH,</t> and the third UMAP shows striatal astrocyte subtypes from integrated single-cell and MERFISH datasets. c – f Top 4 astrocyte subtypes (by abundance) are shown. The astrocyte subtype expression probability is quantified along the dorsal-ventral axis in 500 μm segments in young and aged mice. We divided the striatum into five 500 μm sections, starting at the base of the corpus callosum and moving ventrally. We quantified the density of each astrocyte subtype within each subregion. This astrocyte expression probability quantification was calculated by the number of astrocytes within a subcluster (A1–7), within each 500 μm subregion (0–5) ( X A1…A7 within Y 0…5 ) divided by the total number of astrocytes within that 500 μm section (Σ total ) normalized to the total number of astrocytes within each respective subcluster ( σ A1…A7 ) ([( X A1,..A7 within Y 0…5 /Σ total )/ σ A1…A7 ]). The regional change is quantified by subtracting the young astrocyte expression probability from the aged expression probability. These quantifications were statistically analyzed using a two-way repeated measures (for subregion) ANOVA. Asterisks (*) indicate significant differences ( p value < 0.05) across sub-regions, and hashtags (#) indicate significant differences across ages. Individual representative astrocyte maps for young and aged striatal sections are displayed to the right of the astrocyte density quantification, with astrocyte subtypes demarcated in their respective colors. In the graphs shown in ( c – f ), the corpus callosum is abbreviated as CC on the y -axis.
Merfish Imaging, supplied by Vizgen Inc, used in various techniques. Bioz Stars score: 86/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/merfish/imaging+merfish/pmc12475473-354-0-7
Average 86 stars, based on 1 article reviews
merfish imaging - by Bioz Stars, 2026-10
86/100 stars
  Buy from Supplier

86
Vizgen Inc vizgen merfish
Schematic overview of the Voyager framework. Voyager brings exploratory spatial data analysis (ESDA) methods initially developed for geospatial data to spatial -omics, with consistent user interface for different methods. Voyager is based on the SpatialFeatureExperiment (SFE) object. In R, SFE uses sf and terra to extend SingleCellExperiment (SCE) and SpatialExperiment (SPE). In Python, SFE extends AnnData with GeoPandas. Voyager implements plotting functions for gene expression, cell attributes, and spatial analysis results. The documentation website includes tutorials that demonstrate ESDA on data from multiple spatial -omics technologies, including Visium, Slide-seq, Xenium, CosMX, <t>MERFISH,</t> seqFISH, and CODEX. The website is built automatically with GitHub Actions and pkgdown for reproducibility, and Google Colab notebooks are automatically generated from the vignettes. Compatibility tests are used to make sure that the R and Python implementations return consistent results for core functionalities.
Vizgen Merfish, supplied by Vizgen Inc, used in various techniques. Bioz Stars score: 86/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/merfish/merfish/bio_rxiv__2023__07__20__549945-66-30-30
Average 86 stars, based on 1 article reviews
vizgen merfish - by Bioz Stars, 2026-10
86/100 stars
  Buy from Supplier

86
Spatial Transcriptomics Inc merfish
a , Nicheformer is pretrained on the SpatialCorpus-110M, a large data collection of over 110 million cells measured with dissociated and image-based spatial transcriptomics technologies. The SpatialCorpus-110M collection comprises single-cell data from Homo Sapiens and Mus Musculus across 17 distinct organs and 18 cell lines, and additional single-cell data from other anatomical systems and junctions. Shown is an exemplary uniform manifold approximation and projection (UMAP) visualization of a random 1% subset of the entire pretraining dataset ( n = 1,108,759 cells) of the non-integrated log1p-transformed normalized SpatialCorpus-110M colored by modality. b , Nicheformer includes a novel set of downstream tasks, ranging from spatial cell-type, niche and region label prediction to neighborhood cell density and neighborhood composition prediction. We test our approach on large-scale, high-quality spatial transcriptomics data from the brain (mouse, <t>MERFISH),</t> liver (CosMx, human), lung (CosMx, human; Xenium, human) and colon (Xenium, human). Visualized are example slices of the respective datasets colored by niche labels (brain, liver and lung) and cell density (lung and colon). c , The SpatialCorpus-110M is harmonized and mapped to orthologous gene names, as well as human and mouse-specific genes, to create the input for Nicheformer pretraining. We harmonized metadata information across all datasets, capturing species, modality and assay. d , Each cell’s gene expression profile and metadata are fed into a gene-rank tokenizer to obtain a tokenized representation for each cell. The tokenized cells serve as input for the Nicheformer transformer block to predict masked tokens. Finally, the Nicheformer embedding is generated by aggregating the gene tokens . e , The pretrained Nicheformer embedding is visualized as UMAP colored by modality. The UMAP shows a random 5% subsample of the entire Nicheformer embedding ( n = 4,903,086). NA, not applicable.
Merfish, supplied by Spatial Transcriptomics Inc, used in various techniques. Bioz Stars score: 86/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/merfish/merfish/pmc12695652-308-18-12
Average 86 stars, based on 1 article reviews
merfish - by Bioz Stars, 2026-10
86/100 stars
  Buy from Supplier

Image Search Results


Top panel shows an iterative optimization scheme to estimate geometric transformation ( φ ) and latent feature distribution ( π ) by minimizing the normed difference (error) between the geometric and feature-transformed CCFv3 section to target MERFISH. Middle panel illustrates the application of estimated geometric transformation ( φ ) to deform the CCFv3 atlas to MERFISH coordinates (left); the application of latent feature distribution ( π ) to generate gene distributions on initial CCFv3 geometry (middle); and the application of inverse geometric transformation ( φ −1 ) to deform MERFISH genes to CCFv3 coordinates (right). Gene with the highest probability of expression at each location is shown as a MERFISH feature. Bottom panel illustrates the results of mapping CCFv3 sections to corresponding MERFISH sections. a , f 10 μm atlas sections at Z = 385 and Z = 485 out of 1320 visually chosen to match MERFISH architecture ( e , j ) rendered as meshes at 100 μm. e , j MERFISH sections rendered as meshes at 50 μm, with mRNA density depicted as a feature. b , g Geometric mappings ( φ ) of CCFv3 sections to MERFISH coordinates with the approximate determinant of the Jacobian showing areas of contraction (blue) and expansion (red). c , h Estimated mRNA density per atlas region ( \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${w}_{i}^{{\prime} }$$\end{document} w i ′ in ), as given by π shown in CCFv3 coordinates. d , i Estimated mRNA density shown per atlas region following geometric deformation to MERFISH coordinates.

Journal: Nature Communications

Article Title: Cross-modality mapping using image varifolds to align tissue-scale atlases to molecular-scale measures with application to 2D brain sections

doi: 10.1038/s41467-024-47883-4

Figure Lengend Snippet: Top panel shows an iterative optimization scheme to estimate geometric transformation ( φ ) and latent feature distribution ( π ) by minimizing the normed difference (error) between the geometric and feature-transformed CCFv3 section to target MERFISH. Middle panel illustrates the application of estimated geometric transformation ( φ ) to deform the CCFv3 atlas to MERFISH coordinates (left); the application of latent feature distribution ( π ) to generate gene distributions on initial CCFv3 geometry (middle); and the application of inverse geometric transformation ( φ −1 ) to deform MERFISH genes to CCFv3 coordinates (right). Gene with the highest probability of expression at each location is shown as a MERFISH feature. Bottom panel illustrates the results of mapping CCFv3 sections to corresponding MERFISH sections. a , f 10 μm atlas sections at Z = 385 and Z = 485 out of 1320 visually chosen to match MERFISH architecture ( e , j ) rendered as meshes at 100 μm. e , j MERFISH sections rendered as meshes at 50 μm, with mRNA density depicted as a feature. b , g Geometric mappings ( φ ) of CCFv3 sections to MERFISH coordinates with the approximate determinant of the Jacobian showing areas of contraction (blue) and expansion (red). c , h Estimated mRNA density per atlas region ( \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${w}_{i}^{{\prime} }$$\end{document} w i ′ in ), as given by π shown in CCFv3 coordinates. d , i Estimated mRNA density shown per atlas region following geometric deformation to MERFISH coordinates.

Article Snippet: We demonstrate the efficacy of xIV-LDDMM for mapping tissue-scale atlases to cellular-scale data in mapping CCFv3 section Z = 675 to a section of cell-segmented MERFISH transcriptional data (courtesy of the JEFworks Lab, Johns Hopkins University) (Fig. ).

Techniques: Transformation Assay, Expressing

a – c Relative expression on target section of MERFISH transcriptomics data for three genes ( Gfap ( a ), Trp53i11 ( b ), Wipf3 ( c )) out of a set of twenty (shown at the right), with demonstrated spatial variability according to a computed mutual information score. e – g Predicted expression for the same three genes ( Gfap ( e ), Trp53i11 ( f ), Wipf3 ( g )) in each region of the CCFv3 section Z = 485 out of 1320, as part of the latent distribution over genes estimated in tandem with a geometric transformation to align the CCFv3 section to MERFISH section. d Gene with the highest probability in MERFISH target section. h Predicted gene with the highest probability in estimated latent distribution for the CCFv3 section.

Journal: Nature Communications

Article Title: Cross-modality mapping using image varifolds to align tissue-scale atlases to molecular-scale measures with application to 2D brain sections

doi: 10.1038/s41467-024-47883-4

Figure Lengend Snippet: a – c Relative expression on target section of MERFISH transcriptomics data for three genes ( Gfap ( a ), Trp53i11 ( b ), Wipf3 ( c )) out of a set of twenty (shown at the right), with demonstrated spatial variability according to a computed mutual information score. e – g Predicted expression for the same three genes ( Gfap ( e ), Trp53i11 ( f ), Wipf3 ( g )) in each region of the CCFv3 section Z = 485 out of 1320, as part of the latent distribution over genes estimated in tandem with a geometric transformation to align the CCFv3 section to MERFISH section. d Gene with the highest probability in MERFISH target section. h Predicted gene with the highest probability in estimated latent distribution for the CCFv3 section.

Article Snippet: We demonstrate the efficacy of xIV-LDDMM for mapping tissue-scale atlases to cellular-scale data in mapping CCFv3 section Z = 675 to a section of cell-segmented MERFISH transcriptional data (courtesy of the JEFworks Lab, Johns Hopkins University) (Fig. ).

Techniques: Expressing, Transformation Assay

a Top row depicts MERFISH target rendered as a mesh with cell density (left) and cell type with the highest probability (right) used to summarize cell type distributions. Bottom shows predicted cell density (left) and cell type with the highest probability (right) for the latent feature distribution estimated for each CCFv3 region in the native CCFv3 coordinates. b , c Top row depicts the probability of expression for each gene out of a subset of 6 selected from a total measured set of ~500 as those with high spatial variance. Bottom row depicts the estimated probability of expression for each gene for the latent feature distribution estimated for each CCFv3 region in the native CCFv3 coordinates.

Journal: Nature Communications

Article Title: Cross-modality mapping using image varifolds to align tissue-scale atlases to molecular-scale measures with application to 2D brain sections

doi: 10.1038/s41467-024-47883-4

Figure Lengend Snippet: a Top row depicts MERFISH target rendered as a mesh with cell density (left) and cell type with the highest probability (right) used to summarize cell type distributions. Bottom shows predicted cell density (left) and cell type with the highest probability (right) for the latent feature distribution estimated for each CCFv3 region in the native CCFv3 coordinates. b , c Top row depicts the probability of expression for each gene out of a subset of 6 selected from a total measured set of ~500 as those with high spatial variance. Bottom row depicts the estimated probability of expression for each gene for the latent feature distribution estimated for each CCFv3 region in the native CCFv3 coordinates.

Article Snippet: We demonstrate the efficacy of xIV-LDDMM for mapping tissue-scale atlases to cellular-scale data in mapping CCFv3 section Z = 675 to a section of cell-segmented MERFISH transcriptional data (courtesy of the JEFworks Lab, Johns Hopkins University) (Fig. ).

Techniques: Expressing

a Variance in estimated cell subtype probabilities per CCFv3 region across three replicates summed overall CCFv3 regions. b Spatial variance of cell subtype probabilities, estimated empirically from three pulled-back MERFISH sections, per CCFv3 region and summed across all regions. c – e Probability of excitatory neuron subtype 2 (star in a ) for each of the three mice in CCFv3 coordinates. Yellow arrow highlights the area of the dentate gyrus with differences in excitatory neuron subtype 2 probabilities. f – h Probability for astrocyte subtypes 1,2, and 3 (stars in b ) in empirical distribution computed from all three mice in CCFv3 coordinates (most likely cell type shown in Fig. h). Yellow arrow highlights area of CA1 with differences in astrocyte probability medially to laterally in subtypes 1 and 2 but not 3. Orange arrow highlights differences medially to laterally and left and right in areas of the pons in astrocyte probability for subtypes 1 and 2 but not 3. Black lines indicate boundaries between CCFv3 regions.

Journal: Nature Communications

Article Title: Cross-modality mapping using image varifolds to align tissue-scale atlases to molecular-scale measures with application to 2D brain sections

doi: 10.1038/s41467-024-47883-4

Figure Lengend Snippet: a Variance in estimated cell subtype probabilities per CCFv3 region across three replicates summed overall CCFv3 regions. b Spatial variance of cell subtype probabilities, estimated empirically from three pulled-back MERFISH sections, per CCFv3 region and summed across all regions. c – e Probability of excitatory neuron subtype 2 (star in a ) for each of the three mice in CCFv3 coordinates. Yellow arrow highlights the area of the dentate gyrus with differences in excitatory neuron subtype 2 probabilities. f – h Probability for astrocyte subtypes 1,2, and 3 (stars in b ) in empirical distribution computed from all three mice in CCFv3 coordinates (most likely cell type shown in Fig. h). Yellow arrow highlights area of CA1 with differences in astrocyte probability medially to laterally in subtypes 1 and 2 but not 3. Orange arrow highlights differences medially to laterally and left and right in areas of the pons in astrocyte probability for subtypes 1 and 2 but not 3. Black lines indicate boundaries between CCFv3 regions.

Article Snippet: We demonstrate the efficacy of xIV-LDDMM for mapping tissue-scale atlases to cellular-scale data in mapping CCFv3 section Z = 675 to a section of cell-segmented MERFISH transcriptional data (courtesy of the JEFworks Lab, Johns Hopkins University) (Fig. ).

Techniques:

a , d Original CCFv3 and DevCCF ontologies at location Z = 680 out of 1320. b CCFv3 geometry with predicted DevCCF atlas ontology. Delineations of original CCFv3 partitions are outlined in gray. e DevCCF atlas geometry with predicted CCFv3 ontology. Delineations of original DevCCF partitions are outlined in gray. c , f Entropy of predicted ontologies, with higher entropy values (light) indicating less 1:1 correspondence between ontologies. g Predicted cell type with highest probability per simplex in CCFv3 atlas following mapping to MERFISH target (shown in Fig. with xIV-LDDMM. j Predicted cell type with a highest probability per simplex in DevCCF atlas following mapping to same MERFISH target with xIV-LDDMM. h , k Estimated geometric transformation, φ 1 , in each setting applied to each atlas, with areas of expansion (red) and contraction (blue) as measured by the determinant of the Jacobian. White arrow highlights differences in ontologies in amygdala and striatum designation leading to different geometric transformations. i, l Entropy of estimated cell type distribution per simplex in atlas. Circled area of the hippocampus highlights differences in atlas ontologies leading to differences in the estimated entropy of cell type distributions.

Journal: Nature Communications

Article Title: Cross-modality mapping using image varifolds to align tissue-scale atlases to molecular-scale measures with application to 2D brain sections

doi: 10.1038/s41467-024-47883-4

Figure Lengend Snippet: a , d Original CCFv3 and DevCCF ontologies at location Z = 680 out of 1320. b CCFv3 geometry with predicted DevCCF atlas ontology. Delineations of original CCFv3 partitions are outlined in gray. e DevCCF atlas geometry with predicted CCFv3 ontology. Delineations of original DevCCF partitions are outlined in gray. c , f Entropy of predicted ontologies, with higher entropy values (light) indicating less 1:1 correspondence between ontologies. g Predicted cell type with highest probability per simplex in CCFv3 atlas following mapping to MERFISH target (shown in Fig. with xIV-LDDMM. j Predicted cell type with a highest probability per simplex in DevCCF atlas following mapping to same MERFISH target with xIV-LDDMM. h , k Estimated geometric transformation, φ 1 , in each setting applied to each atlas, with areas of expansion (red) and contraction (blue) as measured by the determinant of the Jacobian. White arrow highlights differences in ontologies in amygdala and striatum designation leading to different geometric transformations. i, l Entropy of estimated cell type distribution per simplex in atlas. Circled area of the hippocampus highlights differences in atlas ontologies leading to differences in the estimated entropy of cell type distributions.

Article Snippet: We demonstrate the efficacy of xIV-LDDMM for mapping tissue-scale atlases to cellular-scale data in mapping CCFv3 section Z = 675 to a section of cell-segmented MERFISH transcriptional data (courtesy of the JEFworks Lab, Johns Hopkins University) (Fig. ).

Techniques: Transformation Assay

a Original DAPI-stained image, digitized at 2.5 μm resolution for tissue section measured with MERFISH technology (Fig. . c Image-varifold particle representation (black points) of DAPI-stained image overlaying the corresponding CCFv3 section in their respective initial coordinate spaces. Thresholded foreground pixels from ( a ) converted to particle image-varifold representation over a feature space of ~30 binned grayscale values. d Alignment of CCFv3 section to DAPI particles following diffeomorphic transformation to the DAPI coordinate space. b Alignment of CCFv3 section and DAPI particles in image format, with the deformed CCFv3 section image generated by resampling the deformed CCFv3 particles onto a regular 2.5 μm grid. White arrows highlight areas of alignment in the area of the substantia inominata (SI) and layer 1 of the cortex whereas red arrows highlight areas of questionable alignment in the area of the olfactory tubercle (OT).

Journal: Nature Communications

Article Title: Cross-modality mapping using image varifolds to align tissue-scale atlases to molecular-scale measures with application to 2D brain sections

doi: 10.1038/s41467-024-47883-4

Figure Lengend Snippet: a Original DAPI-stained image, digitized at 2.5 μm resolution for tissue section measured with MERFISH technology (Fig. . c Image-varifold particle representation (black points) of DAPI-stained image overlaying the corresponding CCFv3 section in their respective initial coordinate spaces. Thresholded foreground pixels from ( a ) converted to particle image-varifold representation over a feature space of ~30 binned grayscale values. d Alignment of CCFv3 section to DAPI particles following diffeomorphic transformation to the DAPI coordinate space. b Alignment of CCFv3 section and DAPI particles in image format, with the deformed CCFv3 section image generated by resampling the deformed CCFv3 particles onto a regular 2.5 μm grid. White arrows highlight areas of alignment in the area of the substantia inominata (SI) and layer 1 of the cortex whereas red arrows highlight areas of questionable alignment in the area of the olfactory tubercle (OT).

Article Snippet: We demonstrate the efficacy of xIV-LDDMM for mapping tissue-scale atlases to cellular-scale data in mapping CCFv3 section Z = 675 to a section of cell-segmented MERFISH transcriptional data (courtesy of the JEFworks Lab, Johns Hopkins University) (Fig. ).

Techniques: Staining, Transformation Assay, Generated

a, CellMemory can characterize single-cell spatial omics from various sequencing platforms, including CosMx, MERFISH, Slide-seq, Stereo-seq, Xenium, and Slide-tags. b, The accuracy of spatial annotation tools is evaluated using datasets from mHypo (mouse hypothalamus measured by MERFISH), hNSCLC (human NSCLC measured by CosMx), and msSermato (mouse spermatogenesis measured by Slide-seq). Each dataset comprised three samples for replication. c, CellMemory was trained using single-cell data at the L2 cell type resolution, to generate CLS embeddings and annotations for Slide-tags cells. The right part is plotted by the L1 (original) and L2 (CellMemory) cell types in spatial coordinates. d, CellMemory model was built using single-cell data at the L3 resolution, to integrate single-cell and Slide-tags data. e, The cells highlighted in the co-embedding are Slide-tags cells, labeled with Slide-tags cell, L1 (original), L2 (CellMemory), and L3 (CellMemory) cell identity. f, The identification of Slide-tags cells at L3 resolution by CellMemory (L4 IT_2 and Micro-PVM_1) is displayed (the first line), along with the expression (the second line) and memory score (the third line) of TAGs ( VWC2L, F13A1 ). The left half represents the spatial coordinate, and the right half is the UMAP coordinate. g, UMAP representation of 4 million mouse whole brain cells from MERFISH, colored by subclasses. h, Annotation benchmark comparison of CellMemory with other state-of-art methods, including scGPT, Geneformer, and CellTypist. The reference dataset consists of mouse whole-brain 10x single-cell data, the query set comprises MERFISH cells derived from 59 coronal sections. i, Visualization of the 5 (total 59) mouse whole-brain sections, with cells colored by predictions of CellMemory. j, Spatial coordinates of mouse brain MERFISH section 36, colored by CellMemory predictions. k, Heatmap displaying the memory scores of TAGs for all subclasses of IT-ET Glut in section 36. l, Distribution of subclasses and their corresponding TAGs’ memory scores within the spatial coordinates of section 36.

Journal: bioRxiv

Article Title: Hierarchical Interpretation of Out-of-Distribution Cells Using Bottlenecked Transformer

doi: 10.1101/2024.12.17.628533

Figure Lengend Snippet: a, CellMemory can characterize single-cell spatial omics from various sequencing platforms, including CosMx, MERFISH, Slide-seq, Stereo-seq, Xenium, and Slide-tags. b, The accuracy of spatial annotation tools is evaluated using datasets from mHypo (mouse hypothalamus measured by MERFISH), hNSCLC (human NSCLC measured by CosMx), and msSermato (mouse spermatogenesis measured by Slide-seq). Each dataset comprised three samples for replication. c, CellMemory was trained using single-cell data at the L2 cell type resolution, to generate CLS embeddings and annotations for Slide-tags cells. The right part is plotted by the L1 (original) and L2 (CellMemory) cell types in spatial coordinates. d, CellMemory model was built using single-cell data at the L3 resolution, to integrate single-cell and Slide-tags data. e, The cells highlighted in the co-embedding are Slide-tags cells, labeled with Slide-tags cell, L1 (original), L2 (CellMemory), and L3 (CellMemory) cell identity. f, The identification of Slide-tags cells at L3 resolution by CellMemory (L4 IT_2 and Micro-PVM_1) is displayed (the first line), along with the expression (the second line) and memory score (the third line) of TAGs ( VWC2L, F13A1 ). The left half represents the spatial coordinate, and the right half is the UMAP coordinate. g, UMAP representation of 4 million mouse whole brain cells from MERFISH, colored by subclasses. h, Annotation benchmark comparison of CellMemory with other state-of-art methods, including scGPT, Geneformer, and CellTypist. The reference dataset consists of mouse whole-brain 10x single-cell data, the query set comprises MERFISH cells derived from 59 coronal sections. i, Visualization of the 5 (total 59) mouse whole-brain sections, with cells colored by predictions of CellMemory. j, Spatial coordinates of mouse brain MERFISH section 36, colored by CellMemory predictions. k, Heatmap displaying the memory scores of TAGs for all subclasses of IT-ET Glut in section 36. l, Distribution of subclasses and their corresponding TAGs’ memory scores within the spatial coordinates of section 36.

Article Snippet: CellMemory was trained using 781 scRNA-seq libraries, encompassing 4 million single-cell transcriptomes from the mouse brain, to analyze the Allen Institute for Brain Science (AIBS) MERFISH dataset .

Techniques: Sequencing, Labeling, Expressing, Comparison, Derivative Assay

Overall training and architectural scheme for CellTransformer. ( a. ) During training, a single cell is drawn (we denote this the reference cell, boxed in red). We extract the reference cell’s spatial neighbors and partition the group into a masked reference cell and its observed spatial neighbors. ( b. ) Initially, the model encoder receives information about each cell and projects those features to d- dimensional latent variable space. Features interact across cells (tokens) through the self-attention mechanism, which is repeated n times. These per-cell representations and an extra token are then aggregated into a single vector representation, which we refer to as the neighborhood representation. This representation is concatenated to a mask token which is cell type specific and chosen to represent the type of the reference cell. A shallow transformer decoder (dotted lines) further refines these representations and then a linear projection is used to output parameters of a negative binomial distribution modeling of the MERFISH probe counts for the reference cell.

Journal: bioRxiv

Article Title: Data-driven fine-grained region discovery in the mouse brain with transformers

doi: 10.1101/2024.05.05.592608

Figure Lengend Snippet: Overall training and architectural scheme for CellTransformer. ( a. ) During training, a single cell is drawn (we denote this the reference cell, boxed in red). We extract the reference cell’s spatial neighbors and partition the group into a masked reference cell and its observed spatial neighbors. ( b. ) Initially, the model encoder receives information about each cell and projects those features to d- dimensional latent variable space. Features interact across cells (tokens) through the self-attention mechanism, which is repeated n times. These per-cell representations and an extra token are then aggregated into a single vector representation, which we refer to as the neighborhood representation. This representation is concatenated to a mask token which is cell type specific and chosen to represent the type of the reference cell. A shallow transformer decoder (dotted lines) further refines these representations and then a linear projection is used to output parameters of a negative binomial distribution modeling of the MERFISH probe counts for the reference cell.

Article Snippet: We downloaded the log-transformed MERFISH probe counts and metadata from the Allen Institute public beta ( https://alleninstitute.github.io/abc_atlas_access/intro.html ) access for ABC-MWB for both the Allen Institute for Brain Science and Zhuang lab mice.

Techniques: Plasmid Preparation

Images of nuclei (DAPI), total polyA mRNA, and two MERFISH bits were obtained using epifluorescence (Epi) and spinning-disk confocal microscopy respectively. Both epifluorescence and confocal images were taken with 1 s exposure time. The MERFISH bit-1 and bit-2 images were high-pass filtered to remove cellular background.

Journal: eLife

Article Title: Three-dimensional single-cell transcriptome imaging of thick tissues

doi: 10.7554/eLife.90029

Figure Lengend Snippet: Images of nuclei (DAPI), total polyA mRNA, and two MERFISH bits were obtained using epifluorescence (Epi) and spinning-disk confocal microscopy respectively. Both epifluorescence and confocal images were taken with 1 s exposure time. The MERFISH bit-1 and bit-2 images were high-pass filtered to remove cellular background.

Article Snippet: MERFISH readout probes , conjugated to either Cy5, Cy3B, or Alexa488 dye molecules through a disulfide linkage, were purchased from Bio-Synthesis, Inc.

Techniques: Confocal Microscopy

( a ) A single-bit high-pass-filtered MERFISH confocal image of 242 genes in a brain tissue section taken with an exposure time of 0.1 s (left) and a magnified view of a single cell marked by the white box in the left image (right). ( b ) The correlation between the copy number of individual genes detected per field of view (FOV) using 0.1 s exposure time and those obtained using 1 s exposure time. The median ratio of the copy number and the Pearson correlation coefficient r are shown. The copy number per gene detected using 0.1 s exposure time is 24% of that detected using 1 s exposure time. ( c ) The same image as in ( a ) but after enhancement of signal-to-noise ratio (SNR) by a DL algorithm. ( d ) The same as ( b ) but after DL was used to enhance the SNR of the 0.1 s images. The copy number per gene detected with 0.1 s exposure time after DL-based enhancement is 89% of that detected using 1 s exposure time. Figure 1—source data 1. This source data file contains source data for .

Journal: eLife

Article Title: Three-dimensional single-cell transcriptome imaging of thick tissues

doi: 10.7554/eLife.90029

Figure Lengend Snippet: ( a ) A single-bit high-pass-filtered MERFISH confocal image of 242 genes in a brain tissue section taken with an exposure time of 0.1 s (left) and a magnified view of a single cell marked by the white box in the left image (right). ( b ) The correlation between the copy number of individual genes detected per field of view (FOV) using 0.1 s exposure time and those obtained using 1 s exposure time. The median ratio of the copy number and the Pearson correlation coefficient r are shown. The copy number per gene detected using 0.1 s exposure time is 24% of that detected using 1 s exposure time. ( c ) The same image as in ( a ) but after enhancement of signal-to-noise ratio (SNR) by a DL algorithm. ( d ) The same as ( b ) but after DL was used to enhance the SNR of the 0.1 s images. The copy number per gene detected with 0.1 s exposure time after DL-based enhancement is 89% of that detected using 1 s exposure time. Figure 1—source data 1. This source data file contains source data for .

Article Snippet: MERFISH readout probes , conjugated to either Cy5, Cy3B, or Alexa488 dye molecules through a disulfide linkage, were purchased from Bio-Synthesis, Inc.

Techniques:

( a ) Number of RNA molecules detected per field of view (FOV) at a single z-plane at the tissue depths of 10 µm and 90 µm in the first bit of the 242-gene MERFISH measurements in a 100-µm-thick section of the mouse cortex. ( b ) Logarithmic distribution of integrated photon counts of individual RNA molecules at the tissue depths of 10 µm and 90 µm identified in ( a ). In each boxplot, the midline represents the median value, the box represents the interquartile range (IQR), the lower whisker represents the smaller of the minimum data point or 1.5× IQR below the 25th percentile, and the upper whisker represents the greater of the maximum data point or 1.5× IQR above the 75th percentile. Molecule number (n, from left to right): 7565, 7914. ( c ) Number of RNA molecules detected per FOV at a single z-plane at tissue depths of 10 µm and 190 µm of in the first bit of the 156-gene MERFISH measurements in a 200-µm-thick section of the mouse hypothalamus. ( d ) Logarithmic distribution of integrated photon counts of individual RNA molecules at the tissue depths of 10 µm and 190 µm identified in ( c ). In each boxplot, the midline represents the median value, the box represents the IQR, the lower whisker represents the smaller of the minimum data point or 1.5× IQR below the 25th percentile, and the upper whisker represents the greater of the maximum data point or 1.5× IQR above the 75th percentile. Molecule number (n, from left to right): 9611, 7238. Figure 1—figure supplement 2—source data 1. This source data file contains source data for .

Journal: eLife

Article Title: Three-dimensional single-cell transcriptome imaging of thick tissues

doi: 10.7554/eLife.90029

Figure Lengend Snippet: ( a ) Number of RNA molecules detected per field of view (FOV) at a single z-plane at the tissue depths of 10 µm and 90 µm in the first bit of the 242-gene MERFISH measurements in a 100-µm-thick section of the mouse cortex. ( b ) Logarithmic distribution of integrated photon counts of individual RNA molecules at the tissue depths of 10 µm and 90 µm identified in ( a ). In each boxplot, the midline represents the median value, the box represents the interquartile range (IQR), the lower whisker represents the smaller of the minimum data point or 1.5× IQR below the 25th percentile, and the upper whisker represents the greater of the maximum data point or 1.5× IQR above the 75th percentile. Molecule number (n, from left to right): 7565, 7914. ( c ) Number of RNA molecules detected per FOV at a single z-plane at tissue depths of 10 µm and 190 µm of in the first bit of the 156-gene MERFISH measurements in a 200-µm-thick section of the mouse hypothalamus. ( d ) Logarithmic distribution of integrated photon counts of individual RNA molecules at the tissue depths of 10 µm and 190 µm identified in ( c ). In each boxplot, the midline represents the median value, the box represents the IQR, the lower whisker represents the smaller of the minimum data point or 1.5× IQR below the 25th percentile, and the upper whisker represents the greater of the maximum data point or 1.5× IQR above the 75th percentile. Molecule number (n, from left to right): 9611, 7238. Figure 1—figure supplement 2—source data 1. This source data file contains source data for .

Article Snippet: MERFISH readout probes , conjugated to either Cy5, Cy3B, or Alexa488 dye molecules through a disulfide linkage, were purchased from Bio-Synthesis, Inc.

Techniques: Whisker Assay

( a ) Example high-pass-filtered bit-1 images of a 242-gene MERFISH measurement in a 100-µm-thick section of mouse cortex stained with different concentrations of encoding probes. The concentration values refer to the concentration of each individual encoding probe. ( b ) Distribution of integrated photon counts of individual RNA molecules identified at different encoding probe concentrations. In each boxplot, the midline represents the median value, the box represents the interquartile range (IQR), the lower whisker represents the smaller of the minimum data point or 1.5× IQR below the 25th percentile, and the upper whisker represents the greater of the maximum data point or 1.5× IQR above the 75th percentile. Molecule number (n, from left to right): 31606, 29190, 83644. The signals from individual RNA molecules increased with the encoding probe concentration and reached saturation at ~1.0 nM per probe. We thus used 1 nM encoding probe concentrations for staining thick-tissue samples. ( c ) A 100-µm-thick mouse brain slice was stained with the 242-gene MERFISH encoding probes, followed by sequential hybridization with readout probes corresponding to the first, second, third, and fourth bit of the barcodes, each bit using a different readout probe concentration. High-pass-filtered bit-1, bit-2, bit-3, and bit-4 MERFISH images (each with a different concentration of readout probes) are shown. ( d ) Distribution of integrated photon counts of individual RNA molecules identified at different readout probe concentrations. In each boxplot, the midline represents the median value, the box represents the IQR, the lower whisker represents the smaller of the minimum data point or 1.5× IQR below the 25th percentile, and the upper whisker represents the greater of the maximum data point or 1.5× IQR above the 75th percentile. Molecule number (n, from left to right): 1939, 3815, 3869, 5248. The signal increased with readout probe concentration, but the background also increased when the probe concentration reached beyond 5 nM. We thus used 5 nM readout probe concentration for thick-tissue imaging. In addition to the probe concentrations, we also optimized readout probe incubation time. ( e ) The number of RNA molecules per field of view per z-plane and the normalized intensity of individual molecules at different tissue depths. The encoding probe concentration was 1 nM per encoding probe, the readout probe concentration was 5 nM, and the readout probe incubation time was 25 min for these measurements. Figure 2—figure supplement 1—source data 1. This source data file contains source data for .

Journal: eLife

Article Title: Three-dimensional single-cell transcriptome imaging of thick tissues

doi: 10.7554/eLife.90029

Figure Lengend Snippet: ( a ) Example high-pass-filtered bit-1 images of a 242-gene MERFISH measurement in a 100-µm-thick section of mouse cortex stained with different concentrations of encoding probes. The concentration values refer to the concentration of each individual encoding probe. ( b ) Distribution of integrated photon counts of individual RNA molecules identified at different encoding probe concentrations. In each boxplot, the midline represents the median value, the box represents the interquartile range (IQR), the lower whisker represents the smaller of the minimum data point or 1.5× IQR below the 25th percentile, and the upper whisker represents the greater of the maximum data point or 1.5× IQR above the 75th percentile. Molecule number (n, from left to right): 31606, 29190, 83644. The signals from individual RNA molecules increased with the encoding probe concentration and reached saturation at ~1.0 nM per probe. We thus used 1 nM encoding probe concentrations for staining thick-tissue samples. ( c ) A 100-µm-thick mouse brain slice was stained with the 242-gene MERFISH encoding probes, followed by sequential hybridization with readout probes corresponding to the first, second, third, and fourth bit of the barcodes, each bit using a different readout probe concentration. High-pass-filtered bit-1, bit-2, bit-3, and bit-4 MERFISH images (each with a different concentration of readout probes) are shown. ( d ) Distribution of integrated photon counts of individual RNA molecules identified at different readout probe concentrations. In each boxplot, the midline represents the median value, the box represents the IQR, the lower whisker represents the smaller of the minimum data point or 1.5× IQR below the 25th percentile, and the upper whisker represents the greater of the maximum data point or 1.5× IQR above the 75th percentile. Molecule number (n, from left to right): 1939, 3815, 3869, 5248. The signal increased with readout probe concentration, but the background also increased when the probe concentration reached beyond 5 nM. We thus used 5 nM readout probe concentration for thick-tissue imaging. In addition to the probe concentrations, we also optimized readout probe incubation time. ( e ) The number of RNA molecules per field of view per z-plane and the normalized intensity of individual molecules at different tissue depths. The encoding probe concentration was 1 nM per encoding probe, the readout probe concentration was 5 nM, and the readout probe incubation time was 25 min for these measurements. Figure 2—figure supplement 1—source data 1. This source data file contains source data for .

Article Snippet: MERFISH readout probes , conjugated to either Cy5, Cy3B, or Alexa488 dye molecules through a disulfide linkage, were purchased from Bio-Synthesis, Inc.

Techniques: Staining, Concentration Assay, Whisker Assay, Slice Preparation, Hybridization, Imaging, Incubation

( a ) Total copy number of decoded RNAs detected per field of view (FOV) per z-plane at different tissue depths in a 242-gene MERFISH measurement of the 100-µm-thick section. ( b ) Pearson correlation coefficients of RNA copy number of individual genes per FOV per z-plane detected at different tissue depths by MERFISH with the FPKM values measured by bulk RNA-seq. ( c ) Correlation of RNA copy number of individual genes per FOV per z-plane detected in the entire 100-µm-thick section by MERFISH with the FPKM values obtained by bulk RNA-seq. The Pearson correlation coefficient r is shown. ( d ) Example images of gel-embedded beads acquired in two rounds of imaging. Buffer exchanges were performed between imaging rounds, mimicking the MERFISH protocol. Because the gel expanded to a different extent in different rounds, the positions of beads changed from round to round in x, y, and z directions. Circles mark beads identified in both imaging rounds. Because the gel size changed, the x and y positions of the beads changed, and the brightness of these beads also changed due to the shift in their z positions. Arrows highlight beads detected in one of the imaging rounds, but not the other, due to the gel-size change, which move these beads out of focus. Figure 2—figure supplement 2—source data 1. This source data file contains source data for .

Journal: eLife

Article Title: Three-dimensional single-cell transcriptome imaging of thick tissues

doi: 10.7554/eLife.90029

Figure Lengend Snippet: ( a ) Total copy number of decoded RNAs detected per field of view (FOV) per z-plane at different tissue depths in a 242-gene MERFISH measurement of the 100-µm-thick section. ( b ) Pearson correlation coefficients of RNA copy number of individual genes per FOV per z-plane detected at different tissue depths by MERFISH with the FPKM values measured by bulk RNA-seq. ( c ) Correlation of RNA copy number of individual genes per FOV per z-plane detected in the entire 100-µm-thick section by MERFISH with the FPKM values obtained by bulk RNA-seq. The Pearson correlation coefficient r is shown. ( d ) Example images of gel-embedded beads acquired in two rounds of imaging. Buffer exchanges were performed between imaging rounds, mimicking the MERFISH protocol. Because the gel expanded to a different extent in different rounds, the positions of beads changed from round to round in x, y, and z directions. Circles mark beads identified in both imaging rounds. Because the gel size changed, the x and y positions of the beads changed, and the brightness of these beads also changed due to the shift in their z positions. Arrows highlight beads detected in one of the imaging rounds, but not the other, due to the gel-size change, which move these beads out of focus. Figure 2—figure supplement 2—source data 1. This source data file contains source data for .

Article Snippet: MERFISH readout probes , conjugated to either Cy5, Cy3B, or Alexa488 dye molecules through a disulfide linkage, were purchased from Bio-Synthesis, Inc.

Techniques: RNA Sequencing Assay, Imaging

( a ) Quantification of gel expansion factor in various buffers used in the MERFISH protocol. The initial gel size was the same as the coverslip, and the expansion factor after buffer exchange was determined as the ratio between the gel size after buffer exchange and the coverslip size. ( b ) In each round of MERFISH imaging, the sample is incubated with the readout probes in a wash buffer (containing either 10% ethylene carbonate [EC] or 10% formamide) for a duration of 15 min. Subsequently, the sample was rinsed with the wash buffer (without readout probes) to remove any excessive readout probes, followed by a treatment with imaging buffer containing glucose-oxidase-based or protocatechuate 3,4-dioxygenase rPCO-based oxygen scavenger system to reduce photobleaching. After the imaging process, the sample is treated with tris(2-carboxyethyl) phosphine buffer to cleave off the fluorescent dye linked to the readout probe through a disulfide bond, and finally washed with a solution of 2× saline-sodium citrate (SSC). All buffers, including wash, imaging, and cleavage buffers, contained 2× SSC. Gel-expansion factor in these buffers used in the MERFISH protocol was quantified and shown here. Reagents marked by * were selected for final use in the 3D thick-tissue MERFISH experiment. The dashed line highlights the expansion factor in the 2× SSC buffer alone. ( c ) XZ projection images of fiducial beads embedded in a gel undergoing buffer exchange for the indicated time period. Wash buffer containing 15% EC in 2× SSC causes noticeable gel distortion, which was recovered after 15 min in 2× SSC buffer without EC. Figure 2—figure supplement 3—source data 1. This source data file contains source data for .

Journal: eLife

Article Title: Three-dimensional single-cell transcriptome imaging of thick tissues

doi: 10.7554/eLife.90029

Figure Lengend Snippet: ( a ) Quantification of gel expansion factor in various buffers used in the MERFISH protocol. The initial gel size was the same as the coverslip, and the expansion factor after buffer exchange was determined as the ratio between the gel size after buffer exchange and the coverslip size. ( b ) In each round of MERFISH imaging, the sample is incubated with the readout probes in a wash buffer (containing either 10% ethylene carbonate [EC] or 10% formamide) for a duration of 15 min. Subsequently, the sample was rinsed with the wash buffer (without readout probes) to remove any excessive readout probes, followed by a treatment with imaging buffer containing glucose-oxidase-based or protocatechuate 3,4-dioxygenase rPCO-based oxygen scavenger system to reduce photobleaching. After the imaging process, the sample is treated with tris(2-carboxyethyl) phosphine buffer to cleave off the fluorescent dye linked to the readout probe through a disulfide bond, and finally washed with a solution of 2× saline-sodium citrate (SSC). All buffers, including wash, imaging, and cleavage buffers, contained 2× SSC. Gel-expansion factor in these buffers used in the MERFISH protocol was quantified and shown here. Reagents marked by * were selected for final use in the 3D thick-tissue MERFISH experiment. The dashed line highlights the expansion factor in the 2× SSC buffer alone. ( c ) XZ projection images of fiducial beads embedded in a gel undergoing buffer exchange for the indicated time period. Wash buffer containing 15% EC in 2× SSC causes noticeable gel distortion, which was recovered after 15 min in 2× SSC buffer without EC. Figure 2—figure supplement 3—source data 1. This source data file contains source data for .

Article Snippet: MERFISH readout probes , conjugated to either Cy5, Cy3B, or Alexa488 dye molecules through a disulfide linkage, were purchased from Bio-Synthesis, Inc.

Techniques: Buffer Exchange, Imaging, Incubation, Saline

( a ) 3D images of DAPI and total polyA mRNA from a single field of view (FOV) in a 100-µm-thick mouse brain tissue slice (top), alongside a single z-plane at tissue depth of 50 µm marked by the yellow box in the top image (bottom). ( b ) Maximum-projection images of 10 consecutive 1 µm z-planes of individual MERFISH bits of the region marked in yellow box in the bottom panel of ( a ). ( c ) RNA molecules identified in the same region as in ( b ) with RNA molecules color coded by their genetic identities. ( d ) The RNA copy number for individual genes per unit area (100 2 µm 2 ) per z-plane detected in the 100 µm MERFISH measurements of mouse cortex versus the FPKM from bulk RNA-seq. The Pearson correlation coefficient r is shown. ( e ) The Pearson correlation between RNA copy number for individual genes per z-plane at different tissue depths detected by MERFISH and the FPKM values of individual genes from bulk RNA-seq. ( f ) Number of detected RNA molecules per FOV at different tissue depths. ( g ) The RNA copy number for individual genes per unit area (100 2 µm 2 ) per z-plane detected in the 100 µm MERFISH measurements of mouse brain sections in this work versus that detected by 10-µm-thick-tissue MERFISH measurements using an epifluorescence setup . The Pearson correlation coefficient r is shown. ( h ) The RNA copy number of individual genes per cell detected in the 100-µm-thick-tissue section versus that in individual 10-µm-thick z-ranges of the same sample. The 100-µm-thick section was evenly divided into ten 10 µm z-ranges to determine the latter. Data in this figure were generated from a 100-μm-thick section of the cortical region collected from an adult mouse. Figure 2—source data 1. This source data file contains source data for .

Journal: eLife

Article Title: Three-dimensional single-cell transcriptome imaging of thick tissues

doi: 10.7554/eLife.90029

Figure Lengend Snippet: ( a ) 3D images of DAPI and total polyA mRNA from a single field of view (FOV) in a 100-µm-thick mouse brain tissue slice (top), alongside a single z-plane at tissue depth of 50 µm marked by the yellow box in the top image (bottom). ( b ) Maximum-projection images of 10 consecutive 1 µm z-planes of individual MERFISH bits of the region marked in yellow box in the bottom panel of ( a ). ( c ) RNA molecules identified in the same region as in ( b ) with RNA molecules color coded by their genetic identities. ( d ) The RNA copy number for individual genes per unit area (100 2 µm 2 ) per z-plane detected in the 100 µm MERFISH measurements of mouse cortex versus the FPKM from bulk RNA-seq. The Pearson correlation coefficient r is shown. ( e ) The Pearson correlation between RNA copy number for individual genes per z-plane at different tissue depths detected by MERFISH and the FPKM values of individual genes from bulk RNA-seq. ( f ) Number of detected RNA molecules per FOV at different tissue depths. ( g ) The RNA copy number for individual genes per unit area (100 2 µm 2 ) per z-plane detected in the 100 µm MERFISH measurements of mouse brain sections in this work versus that detected by 10-µm-thick-tissue MERFISH measurements using an epifluorescence setup . The Pearson correlation coefficient r is shown. ( h ) The RNA copy number of individual genes per cell detected in the 100-µm-thick-tissue section versus that in individual 10-µm-thick z-ranges of the same sample. The 100-µm-thick section was evenly divided into ten 10 µm z-ranges to determine the latter. Data in this figure were generated from a 100-μm-thick section of the cortical region collected from an adult mouse. Figure 2—source data 1. This source data file contains source data for .

Article Snippet: MERFISH readout probes , conjugated to either Cy5, Cy3B, or Alexa488 dye molecules through a disulfide linkage, were purchased from Bio-Synthesis, Inc.

Techniques: RNA Sequencing Assay, Generated

( a ) UMAP visualization of subclasses of cells identified in a 100-μm-thick section of the mouse cortex. Cells are color coded by subclass identities. IT: intratelencephalic projection neurons; ET: extratelencephalic projection neurons; CT: cortical-thalamic projection neurons; NP: near projection neurons; OPC: oligodendrocyte progenitor cells; Oligo: oligodendrocytes; Micro: microglia; Astro: astrocytes; VLMC: vascular leptomeningeal cells; Endo: endothelial cells; Peri: pericytes; SMC: smooth muscle cells; PVM: perivascular macrophages. ( b ) 3D spatial maps of the identified subclasses of excitatory neurons (left), inhibitory neurons (middle), and non-neuronal cells (right) within the 100 μm mouse cortex section. ( c ) UMAP visualization of major cell types identified in a 200-μm-thick section of the mouse anterior hypothalamus. Cells are color coded by cell type identities. ( d ) 3D spatial maps of the excitatory neuronal (left), inhibitory neuronal (middle), and non-neuronal (right) cell clusters identified in the 200-μm-thick mouse hypothalamus section. Cells are color coded by cell cluster identities in the top panels and two example excitatory neuronal (left), inhibitory neuronal (middle), and non-neuronal (right) cell clusters are shown in the bottom panels. ( e ) Boxplots showing the distributions of the nearest-neighbor distances from cells in individual inhibitory neuronal subclasses to cells in the same subclass (‘to self’) or other subclasses (‘to other’) in the mouse cortex obtained from the thick-tissue 3D MERFISH data. Cell numbers (n, from left to right): 44, 274, 125, 1161, 136, 797, 26, 330. *FDR <0.01 was determined with the Wilcoxon rank-sum one-sided test and adjusted to FDR by the Benjamini and Hochberg procedure. Only inhibitory neuronal clusters with at least 20 ‘self-self’ interacting pairs were examined and plotted. In each boxplot, the midline represents the median value, the box represents the interquartile range (IQR), the lower whisker represents the smaller of the minimum data point or 1.5× IQR below the 25th percentile, and the upper whisker represents the greater of the maximum data point or 1.5× IQR above the 75th percentile. ( f ) Distributions of the nearest-neighbor distances among all interneurons derived from the thick-tissue 3D MERFISH data of the mouse cortex. The distribution is fitted with a bimodal distribution (blue curve) with the two individual Gaussian peaks shown in red and green. ( g, h ) Same as ( e, f ) but for inhibitory neurons in the mouse anterior hypothalamus obtained from the thick-tissue 3D MERFISH data. Cell numbers in g (n, from left to right): 1404, 1222, 198, 567, 845, 1315, 866, 1263, 1149, 873, 455, 1517, 387, 1490, 409, 1321, 277, 1151, 773, 1871, 379, 830, 174, 1013, 598, 449, 811, 158, 194, 728, 103, 771, 119, 667, 217, 580, 98, 704, 1095, 1339, 186, 362, 82, 438, 41, 302. Data in this figure were generated from a 200-μm-thick section of the hypothalamic region collected from an adult mouse. Figure 3—source data 1. This source data file contains source data for .

Journal: eLife

Article Title: Three-dimensional single-cell transcriptome imaging of thick tissues

doi: 10.7554/eLife.90029

Figure Lengend Snippet: ( a ) UMAP visualization of subclasses of cells identified in a 100-μm-thick section of the mouse cortex. Cells are color coded by subclass identities. IT: intratelencephalic projection neurons; ET: extratelencephalic projection neurons; CT: cortical-thalamic projection neurons; NP: near projection neurons; OPC: oligodendrocyte progenitor cells; Oligo: oligodendrocytes; Micro: microglia; Astro: astrocytes; VLMC: vascular leptomeningeal cells; Endo: endothelial cells; Peri: pericytes; SMC: smooth muscle cells; PVM: perivascular macrophages. ( b ) 3D spatial maps of the identified subclasses of excitatory neurons (left), inhibitory neurons (middle), and non-neuronal cells (right) within the 100 μm mouse cortex section. ( c ) UMAP visualization of major cell types identified in a 200-μm-thick section of the mouse anterior hypothalamus. Cells are color coded by cell type identities. ( d ) 3D spatial maps of the excitatory neuronal (left), inhibitory neuronal (middle), and non-neuronal (right) cell clusters identified in the 200-μm-thick mouse hypothalamus section. Cells are color coded by cell cluster identities in the top panels and two example excitatory neuronal (left), inhibitory neuronal (middle), and non-neuronal (right) cell clusters are shown in the bottom panels. ( e ) Boxplots showing the distributions of the nearest-neighbor distances from cells in individual inhibitory neuronal subclasses to cells in the same subclass (‘to self’) or other subclasses (‘to other’) in the mouse cortex obtained from the thick-tissue 3D MERFISH data. Cell numbers (n, from left to right): 44, 274, 125, 1161, 136, 797, 26, 330. *FDR <0.01 was determined with the Wilcoxon rank-sum one-sided test and adjusted to FDR by the Benjamini and Hochberg procedure. Only inhibitory neuronal clusters with at least 20 ‘self-self’ interacting pairs were examined and plotted. In each boxplot, the midline represents the median value, the box represents the interquartile range (IQR), the lower whisker represents the smaller of the minimum data point or 1.5× IQR below the 25th percentile, and the upper whisker represents the greater of the maximum data point or 1.5× IQR above the 75th percentile. ( f ) Distributions of the nearest-neighbor distances among all interneurons derived from the thick-tissue 3D MERFISH data of the mouse cortex. The distribution is fitted with a bimodal distribution (blue curve) with the two individual Gaussian peaks shown in red and green. ( g, h ) Same as ( e, f ) but for inhibitory neurons in the mouse anterior hypothalamus obtained from the thick-tissue 3D MERFISH data. Cell numbers in g (n, from left to right): 1404, 1222, 198, 567, 845, 1315, 866, 1263, 1149, 873, 455, 1517, 387, 1490, 409, 1321, 277, 1151, 773, 1871, 379, 830, 174, 1013, 598, 449, 811, 158, 194, 728, 103, 771, 119, 667, 217, 580, 98, 704, 1095, 1339, 186, 362, 82, 438, 41, 302. Data in this figure were generated from a 200-μm-thick section of the hypothalamic region collected from an adult mouse. Figure 3—source data 1. This source data file contains source data for .

Article Snippet: MERFISH readout probes , conjugated to either Cy5, Cy3B, or Alexa488 dye molecules through a disulfide linkage, were purchased from Bio-Synthesis, Inc.

Techniques: Whisker Assay, Derivative Assay, Generated

A) Nicheformer is pretrained on the SpatialCorpus-110M, a large data collection of over 110 million cells measured with dissociated and image-based spatial transcriptomics technologies. The SpatialCorpus-110M collection comprises single-cell data from Homo Sapiens and Mus Musculus across 17 distinct organs, 18 cell lines, and additional single-cell data from other anatomical systems and junctions. Shown is an exemplary UMAP visualization of a random 1% subset of the entire pretraining dataset (n=1,108,759 cells) of the non-integrated log1p-transformed normalized SpatialCorpus-110M colored by modality. B) Nicheformer includes a novel set of downstream tasks, ranging from spatial cell type, niche and region label prediction to neighborhood cell density and neighborhood composition prediction. We test our approach on large-scale, high-quality spatial transcriptomics data from the brain (mouse - MERFISH), liver (CosMx - human), lung (CosMx - human, Xenium - human), and colon (Xenium - human). Visualized are example slices of the respective datasets colored by niche labels (brain, liver, and lung) and cell density (lung and colon). C) The SpatialCorpus-110M is harmonized and mapped to orthologous gene names, as well as human and mouse-specific genes, to create the input for Nicheformer pretraining. We harmonized metadata information across all datasets, capturing species, modality, and assay. D) Each cell’s gene expression profile and metadata are fed into a gene rank tokenizer to obtain a tokenized representation for each cell. The tokenized cells serve as input for the Nicheformer transformer block to predict masked tokens. Finally, the Nicheformer embedding is generated by aggregating the gene tokens (Methods). E) The pretrained Nicheformer embedding is visualized as UMAP colored by modality. The UMAP shows a random 5% subsample of the entire Nicheformer embedding (n=4,903,086).

Journal: bioRxiv

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1101/2024.04.15.589472

Figure Lengend Snippet: A) Nicheformer is pretrained on the SpatialCorpus-110M, a large data collection of over 110 million cells measured with dissociated and image-based spatial transcriptomics technologies. The SpatialCorpus-110M collection comprises single-cell data from Homo Sapiens and Mus Musculus across 17 distinct organs, 18 cell lines, and additional single-cell data from other anatomical systems and junctions. Shown is an exemplary UMAP visualization of a random 1% subset of the entire pretraining dataset (n=1,108,759 cells) of the non-integrated log1p-transformed normalized SpatialCorpus-110M colored by modality. B) Nicheformer includes a novel set of downstream tasks, ranging from spatial cell type, niche and region label prediction to neighborhood cell density and neighborhood composition prediction. We test our approach on large-scale, high-quality spatial transcriptomics data from the brain (mouse - MERFISH), liver (CosMx - human), lung (CosMx - human, Xenium - human), and colon (Xenium - human). Visualized are example slices of the respective datasets colored by niche labels (brain, liver, and lung) and cell density (lung and colon). C) The SpatialCorpus-110M is harmonized and mapped to orthologous gene names, as well as human and mouse-specific genes, to create the input for Nicheformer pretraining. We harmonized metadata information across all datasets, capturing species, modality, and assay. D) Each cell’s gene expression profile and metadata are fed into a gene rank tokenizer to obtain a tokenized representation for each cell. The tokenized cells serve as input for the Nicheformer transformer block to predict masked tokens. Finally, the Nicheformer embedding is generated by aggregating the gene tokens (Methods). E) The pretrained Nicheformer embedding is visualized as UMAP colored by modality. The UMAP shows a random 5% subsample of the entire Nicheformer embedding (n=4,903,086).

Article Snippet: For spatial transcriptomics, we curated image-based spatial datasets, specifically MERFISH (Vizgen MERSCOPE), 10x Genomics Xenium, Nanostring CosMx , and In Situ Sequencing (ISS) data ( , Supp.

Techniques: Transformation Assay, Gene Expression, Blocking Assay, Generated

A) Single-cells resolved in space on an example slice (n=114,396 cells) of the MERFISH mouse brain dataset with niche label superimposed. B) Test-set F1-macro of niche and brain region label prediction of the fine-tuned Nicheformer model, the linear probing model, and a linear probing baseline computed based on embeddings generated with scVI and PCA, respectively. C) UMAP of dissociated scRNAseq dataset with original author cell type label superimposed. D) Nicheformer can transfer spatial niche and region labels onto dissociated single-cell data. E) Nicheformer accurately classifies cells from the dissociated motor cortex to relevant cell types (n=9 out of 33 distinct ones in the classifier) trained on the whole mouse brain MERFISH dataset. F-G) Nicheformer correctly projects dissociated single-cells to niche (F) and region (G) labels to provide spatially dependent labels. F) Nicheformer misclassified parts of L2/3 IT neurons as residing in the subpallium GABAergic niche (highlighted in the red box). Additionally, the deep cortical excitatory neurons L6b, L6 CT, L6 IT, and L6 IT Car3 (highlighted in the red box) should be classified as pallium glutamatergic niche instead of subpallium GABAergic by Nicheformer. G) Most of the non-neuronal cells (84.7 % of all non-neuronal cells, n=133) were misclassified as not belonging to the isocortex or the adjacent brain regions (highlighted in the red box). H) Cell type abundances in the scRNA-seq dataset measuring the primary motor cortex in the mouse. I-K) Classification uncertainty of label transfer of the dissociated scRNA-seq dataset to the MERFISH mouse brain data for cell type label (I), niche label (J), and region label (K) with a value of 0 representing a high uncertainty and 1 being a lower uncertainty, i.e., high certainty. K) Observed high uncertainty for parts of the Glut and GABA neurons for the region prediction of the isocortex, CTXsp and OLF, which are neighboring brain regions.

Journal: bioRxiv

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1101/2024.04.15.589472

Figure Lengend Snippet: A) Single-cells resolved in space on an example slice (n=114,396 cells) of the MERFISH mouse brain dataset with niche label superimposed. B) Test-set F1-macro of niche and brain region label prediction of the fine-tuned Nicheformer model, the linear probing model, and a linear probing baseline computed based on embeddings generated with scVI and PCA, respectively. C) UMAP of dissociated scRNAseq dataset with original author cell type label superimposed. D) Nicheformer can transfer spatial niche and region labels onto dissociated single-cell data. E) Nicheformer accurately classifies cells from the dissociated motor cortex to relevant cell types (n=9 out of 33 distinct ones in the classifier) trained on the whole mouse brain MERFISH dataset. F-G) Nicheformer correctly projects dissociated single-cells to niche (F) and region (G) labels to provide spatially dependent labels. F) Nicheformer misclassified parts of L2/3 IT neurons as residing in the subpallium GABAergic niche (highlighted in the red box). Additionally, the deep cortical excitatory neurons L6b, L6 CT, L6 IT, and L6 IT Car3 (highlighted in the red box) should be classified as pallium glutamatergic niche instead of subpallium GABAergic by Nicheformer. G) Most of the non-neuronal cells (84.7 % of all non-neuronal cells, n=133) were misclassified as not belonging to the isocortex or the adjacent brain regions (highlighted in the red box). H) Cell type abundances in the scRNA-seq dataset measuring the primary motor cortex in the mouse. I-K) Classification uncertainty of label transfer of the dissociated scRNA-seq dataset to the MERFISH mouse brain data for cell type label (I), niche label (J), and region label (K) with a value of 0 representing a high uncertainty and 1 being a lower uncertainty, i.e., high certainty. K) Observed high uncertainty for parts of the Glut and GABA neurons for the region prediction of the isocortex, CTXsp and OLF, which are neighboring brain regions.

Article Snippet: For spatial transcriptomics, we curated image-based spatial datasets, specifically MERFISH (Vizgen MERSCOPE), 10x Genomics Xenium, Nanostring CosMx , and In Situ Sequencing (ISS) data ( , Supp.

Techniques: Generated

A) We define the neighborhood of a cell as its local neighborhood given a radius and an index cell. The neighborhood cell density is then defined by the number of cells in the neighborhood, and the neighborhood compositions are the proportions of neighboring cell types. B) Neighborhoods are computed at multiple resolutions resulting in different neighborhood size distributions. Each barplot shows the distribution of the number of neighbors across the brain, liver, and lung datasets. We extract neighborhoods with the mean number of neighbors 10, 20, 50, and 100 for each dataset. C) The fine-tuned and linear probing Nicheformer models outperform zero-shot models trained on scVI and PCA embeddings in terms of mean absolute error across all neighborhood sizes and all three organs, the brain, liver, and lung. D) Left: Fine-tuned Nicheformer performance on the MERFISH mouse brain data grouped by index cell type. Shown are the absolute error values between predicted and observed neighborhood composition vectors for held-out test cells. For each box in (D), the centerline defines the median, the height of the box is given by the interquartile range (IQR), the whiskers are given by 1.5 × IQR, and outliers are given as points beyond the minimum or maximum whisker. Center: Index cell type abundances in the entire MERFISH mouse brain dataset. Right: UMAPs of MERFISH mouse brain Nicheformer embedding with the selected index cell type as color superimposed. E) UMAP of the Nicheformer embedding of all immune cells in the MERFISH mouse brain dataset with region label as color superimposed.

Journal: bioRxiv

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1101/2024.04.15.589472

Figure Lengend Snippet: A) We define the neighborhood of a cell as its local neighborhood given a radius and an index cell. The neighborhood cell density is then defined by the number of cells in the neighborhood, and the neighborhood compositions are the proportions of neighboring cell types. B) Neighborhoods are computed at multiple resolutions resulting in different neighborhood size distributions. Each barplot shows the distribution of the number of neighbors across the brain, liver, and lung datasets. We extract neighborhoods with the mean number of neighbors 10, 20, 50, and 100 for each dataset. C) The fine-tuned and linear probing Nicheformer models outperform zero-shot models trained on scVI and PCA embeddings in terms of mean absolute error across all neighborhood sizes and all three organs, the brain, liver, and lung. D) Left: Fine-tuned Nicheformer performance on the MERFISH mouse brain data grouped by index cell type. Shown are the absolute error values between predicted and observed neighborhood composition vectors for held-out test cells. For each box in (D), the centerline defines the median, the height of the box is given by the interquartile range (IQR), the whiskers are given by 1.5 × IQR, and outliers are given as points beyond the minimum or maximum whisker. Center: Index cell type abundances in the entire MERFISH mouse brain dataset. Right: UMAPs of MERFISH mouse brain Nicheformer embedding with the selected index cell type as color superimposed. E) UMAP of the Nicheformer embedding of all immune cells in the MERFISH mouse brain dataset with region label as color superimposed.

Article Snippet: For spatial transcriptomics, we curated image-based spatial datasets, specifically MERFISH (Vizgen MERSCOPE), 10x Genomics Xenium, Nanostring CosMx , and In Situ Sequencing (ISS) data ( , Supp.

Techniques: Whisker Assay

a XY directional map of all cell types identified within the striatal section is shown. Medium spiny neurons, cortical neurons, astrocytes, oligodendrocytes, endothelial cells, and microglia are displayed in different colors - each dot represents a cell segmented by MERLIN. All cells within the XY coordinates of the striatum were subset, and here we display the UMAP of these subset striatal cell clusters split by major cell types and UMAP split by age (young = blue; aged = red). b Striatal astrocytes were subset from all striatal cells and clustered separately using the Louvian algorithm. The first UMAP shows striatal astrocyte subtypes from single cells, the second UMAP shows striatal astrocyte subtypes from MERFISH, and the third UMAP shows striatal astrocyte subtypes from integrated single-cell and MERFISH datasets. c – f Top 4 astrocyte subtypes (by abundance) are shown. The astrocyte subtype expression probability is quantified along the dorsal-ventral axis in 500 μm segments in young and aged mice. We divided the striatum into five 500 μm sections, starting at the base of the corpus callosum and moving ventrally. We quantified the density of each astrocyte subtype within each subregion. This astrocyte expression probability quantification was calculated by the number of astrocytes within a subcluster (A1–7), within each 500 μm subregion (0–5) ( X A1…A7 within Y 0…5 ) divided by the total number of astrocytes within that 500 μm section (Σ total ) normalized to the total number of astrocytes within each respective subcluster ( σ A1…A7 ) ([( X A1,..A7 within Y 0…5 /Σ total )/ σ A1…A7 ]). The regional change is quantified by subtracting the young astrocyte expression probability from the aged expression probability. These quantifications were statistically analyzed using a two-way repeated measures (for subregion) ANOVA. Asterisks (*) indicate significant differences ( p value < 0.05) across sub-regions, and hashtags (#) indicate significant differences across ages. Individual representative astrocyte maps for young and aged striatal sections are displayed to the right of the astrocyte density quantification, with astrocyte subtypes demarcated in their respective colors. In the graphs shown in ( c – f ), the corpus callosum is abbreviated as CC on the y -axis.

Journal: Nature Communications

Article Title: Aging in mice alters regionally enriched striatal astrocytes

doi: 10.1038/s41467-025-63429-8

Figure Lengend Snippet: a XY directional map of all cell types identified within the striatal section is shown. Medium spiny neurons, cortical neurons, astrocytes, oligodendrocytes, endothelial cells, and microglia are displayed in different colors - each dot represents a cell segmented by MERLIN. All cells within the XY coordinates of the striatum were subset, and here we display the UMAP of these subset striatal cell clusters split by major cell types and UMAP split by age (young = blue; aged = red). b Striatal astrocytes were subset from all striatal cells and clustered separately using the Louvian algorithm. The first UMAP shows striatal astrocyte subtypes from single cells, the second UMAP shows striatal astrocyte subtypes from MERFISH, and the third UMAP shows striatal astrocyte subtypes from integrated single-cell and MERFISH datasets. c – f Top 4 astrocyte subtypes (by abundance) are shown. The astrocyte subtype expression probability is quantified along the dorsal-ventral axis in 500 μm segments in young and aged mice. We divided the striatum into five 500 μm sections, starting at the base of the corpus callosum and moving ventrally. We quantified the density of each astrocyte subtype within each subregion. This astrocyte expression probability quantification was calculated by the number of astrocytes within a subcluster (A1–7), within each 500 μm subregion (0–5) ( X A1…A7 within Y 0…5 ) divided by the total number of astrocytes within that 500 μm section (Σ total ) normalized to the total number of astrocytes within each respective subcluster ( σ A1…A7 ) ([( X A1,..A7 within Y 0…5 /Σ total )/ σ A1…A7 ]). The regional change is quantified by subtracting the young astrocyte expression probability from the aged expression probability. These quantifications were statistically analyzed using a two-way repeated measures (for subregion) ANOVA. Asterisks (*) indicate significant differences ( p value < 0.05) across sub-regions, and hashtags (#) indicate significant differences across ages. Individual representative astrocyte maps for young and aged striatal sections are displayed to the right of the astrocyte density quantification, with astrocyte subtypes demarcated in their respective colors. In the graphs shown in ( c – f ), the corpus callosum is abbreviated as CC on the y -axis.

Article Snippet: MERFISH imaging was performed on an automated Vizgen Alpha Instrument using imaging buffers, hybridization buffers, and parameter files provided by Vizgen.

Techniques: Expressing

a Top aging astrocyte markers were assessed using MAST differential expression, and the top 25 up and 25 downregulated transcripts are shown in the heatmap (see also Supplementary Data ). Dorsal enriched genes are marked with an asterisk (*) and ventral enriched genes are marked with a hashtag (#). b Representative image of MERFISH RNA counts shows age increases in Gfap in the dorsal striatum. c Representative images of the dorsal striatum are shown for young and aged mice. GFAP coverage per 500 μm 2 in the dorsal striatum; across young and aged mice ( n = 4–5 mice; 2-way RM ANOVA with Bonferroni post hoc (* p = 0.00032 (aged dorsal compared ventral) and * p = 0.00029 (aged dorsal compared to young dorsal), data are presented as mean values ± SEM). d Representative image of MERFISH RNA counts shows S100b expression in the dorsal striatum. e Quantification of S100β + cells per 500 μm 2 , in the dorsal, medial, and ventral striatum; across young and aged mice ( n = 4–5 mice; 2-way RM ANOVA with Bonferroni post hoc * p = 0.036, data are presented as mean values ± SEM). f IPA analysis was performed on age-induced DEGs in striatal astrocytes, and each black bar indicates the number of genes per pathway, circle size indicates the −log( p value) (right-tailed Fisher’s Exact Test), and circle color indicates the activation score number. g Aging gene score is calculated by the average expression levels of the top 20 differentially expressed genes in age on a single-cell level, subtracted by the aggregated expression of control feature sets. Each dot represents an astrocyte, and the color of the dot is the relative change in the aging score. h Heatmaps of the top 10 shared genes between our aging and A1 astrocyte subtype genes (Log2FC > 0.01) with mouse aging (Log2FC > 0.01), human striatal astrocyte aging, human striatal astrocyte Huntington’s Disease, and human Parkinson’s Disease using the DEG (see data availability and Supplementary Data ). Venn diagrams visualize the total overlap across each gene set. i UpSet plot of the overlap of murine aging striatal data and the past studies.

Journal: Nature Communications

Article Title: Aging in mice alters regionally enriched striatal astrocytes

doi: 10.1038/s41467-025-63429-8

Figure Lengend Snippet: a Top aging astrocyte markers were assessed using MAST differential expression, and the top 25 up and 25 downregulated transcripts are shown in the heatmap (see also Supplementary Data ). Dorsal enriched genes are marked with an asterisk (*) and ventral enriched genes are marked with a hashtag (#). b Representative image of MERFISH RNA counts shows age increases in Gfap in the dorsal striatum. c Representative images of the dorsal striatum are shown for young and aged mice. GFAP coverage per 500 μm 2 in the dorsal striatum; across young and aged mice ( n = 4–5 mice; 2-way RM ANOVA with Bonferroni post hoc (* p = 0.00032 (aged dorsal compared ventral) and * p = 0.00029 (aged dorsal compared to young dorsal), data are presented as mean values ± SEM). d Representative image of MERFISH RNA counts shows S100b expression in the dorsal striatum. e Quantification of S100β + cells per 500 μm 2 , in the dorsal, medial, and ventral striatum; across young and aged mice ( n = 4–5 mice; 2-way RM ANOVA with Bonferroni post hoc * p = 0.036, data are presented as mean values ± SEM). f IPA analysis was performed on age-induced DEGs in striatal astrocytes, and each black bar indicates the number of genes per pathway, circle size indicates the −log( p value) (right-tailed Fisher’s Exact Test), and circle color indicates the activation score number. g Aging gene score is calculated by the average expression levels of the top 20 differentially expressed genes in age on a single-cell level, subtracted by the aggregated expression of control feature sets. Each dot represents an astrocyte, and the color of the dot is the relative change in the aging score. h Heatmaps of the top 10 shared genes between our aging and A1 astrocyte subtype genes (Log2FC > 0.01) with mouse aging (Log2FC > 0.01), human striatal astrocyte aging, human striatal astrocyte Huntington’s Disease, and human Parkinson’s Disease using the DEG (see data availability and Supplementary Data ). Venn diagrams visualize the total overlap across each gene set. i UpSet plot of the overlap of murine aging striatal data and the past studies.

Article Snippet: MERFISH imaging was performed on an automated Vizgen Alpha Instrument using imaging buffers, hybridization buffers, and parameter files provided by Vizgen.

Techniques: Quantitative Proteomics, Expressing, Activation Assay, Control

Schematic overview of the Voyager framework. Voyager brings exploratory spatial data analysis (ESDA) methods initially developed for geospatial data to spatial -omics, with consistent user interface for different methods. Voyager is based on the SpatialFeatureExperiment (SFE) object. In R, SFE uses sf and terra to extend SingleCellExperiment (SCE) and SpatialExperiment (SPE). In Python, SFE extends AnnData with GeoPandas. Voyager implements plotting functions for gene expression, cell attributes, and spatial analysis results. The documentation website includes tutorials that demonstrate ESDA on data from multiple spatial -omics technologies, including Visium, Slide-seq, Xenium, CosMX, MERFISH, seqFISH, and CODEX. The website is built automatically with GitHub Actions and pkgdown for reproducibility, and Google Colab notebooks are automatically generated from the vignettes. Compatibility tests are used to make sure that the R and Python implementations return consistent results for core functionalities.

Journal: bioRxiv

Article Title: Voyager: exploratory single-cell genomics data analysis with geospatial statistics

doi: 10.1101/2023.07.20.549945

Figure Lengend Snippet: Schematic overview of the Voyager framework. Voyager brings exploratory spatial data analysis (ESDA) methods initially developed for geospatial data to spatial -omics, with consistent user interface for different methods. Voyager is based on the SpatialFeatureExperiment (SFE) object. In R, SFE uses sf and terra to extend SingleCellExperiment (SCE) and SpatialExperiment (SPE). In Python, SFE extends AnnData with GeoPandas. Voyager implements plotting functions for gene expression, cell attributes, and spatial analysis results. The documentation website includes tutorials that demonstrate ESDA on data from multiple spatial -omics technologies, including Visium, Slide-seq, Xenium, CosMX, MERFISH, seqFISH, and CODEX. The website is built automatically with GitHub Actions and pkgdown for reproducibility, and Google Colab notebooks are automatically generated from the vignettes. Compatibility tests are used to make sure that the R and Python implementations return consistent results for core functionalities.

Article Snippet: Voyager has a comprehensive documentation website that features tutorials on applying EDA and ESDA to datasets from multiple spatial -omics technologies, including 10X Visium and Xenium , Nanostring CosMX , Vizgen MERFISH , Slide-seq , seqFISH , and CODEX ( ).

Techniques: Gene Expression, Generated

Applications of Voyager on spatial transcriptomics datasets. A) In a mouse skeletal muscle dataset, the total UMI counts, or library size per spot (nCounts), are plotted in space as blue open circles and myofibers are colored in red according to their cross section areas. Only spots that intersect tissue are plotted. The H&E image is plotted on the side as a reference. B) Scatter plot of the number of genes detected per spot (nGenes) vs. nCounts, colored by mean area of myofibers that intersect each spot. C) Simulated (density plot) and observed (vertical line) difference between Moran’s I in nCounts of spots that intersect tissue (in) and that of spots that don’t (out). D) The 20 most positive and 20 most negative eigenvalues from MULTISPATI PCA of a mouse liver MERFISH dataset. As other eigenvalues were not computed, there is a break after PC20 in this plot. E) The most positive and negative gene loadings for PCs 1, 2, and 40. F) A subset of the MERFISH data showing a portal triad (near top right) and two central veins (left and bottom right), with cell polygons colored by their projections into 2 PCs with the most positive eigenvalues and the PC with the most negative eigenvalue (“PC40”). The first 2 PCs show zonation.

Journal: bioRxiv

Article Title: Voyager: exploratory single-cell genomics data analysis with geospatial statistics

doi: 10.1101/2023.07.20.549945

Figure Lengend Snippet: Applications of Voyager on spatial transcriptomics datasets. A) In a mouse skeletal muscle dataset, the total UMI counts, or library size per spot (nCounts), are plotted in space as blue open circles and myofibers are colored in red according to their cross section areas. Only spots that intersect tissue are plotted. The H&E image is plotted on the side as a reference. B) Scatter plot of the number of genes detected per spot (nGenes) vs. nCounts, colored by mean area of myofibers that intersect each spot. C) Simulated (density plot) and observed (vertical line) difference between Moran’s I in nCounts of spots that intersect tissue (in) and that of spots that don’t (out). D) The 20 most positive and 20 most negative eigenvalues from MULTISPATI PCA of a mouse liver MERFISH dataset. As other eigenvalues were not computed, there is a break after PC20 in this plot. E) The most positive and negative gene loadings for PCs 1, 2, and 40. F) A subset of the MERFISH data showing a portal triad (near top right) and two central veins (left and bottom right), with cell polygons colored by their projections into 2 PCs with the most positive eigenvalues and the PC with the most negative eigenvalue (“PC40”). The first 2 PCs show zonation.

Article Snippet: Voyager has a comprehensive documentation website that features tutorials on applying EDA and ESDA to datasets from multiple spatial -omics technologies, including 10X Visium and Xenium , Nanostring CosMX , Vizgen MERFISH , Slide-seq , seqFISH , and CODEX ( ).

Techniques:

a , Nicheformer is pretrained on the SpatialCorpus-110M, a large data collection of over 110 million cells measured with dissociated and image-based spatial transcriptomics technologies. The SpatialCorpus-110M collection comprises single-cell data from Homo Sapiens and Mus Musculus across 17 distinct organs and 18 cell lines, and additional single-cell data from other anatomical systems and junctions. Shown is an exemplary uniform manifold approximation and projection (UMAP) visualization of a random 1% subset of the entire pretraining dataset ( n = 1,108,759 cells) of the non-integrated log1p-transformed normalized SpatialCorpus-110M colored by modality. b , Nicheformer includes a novel set of downstream tasks, ranging from spatial cell-type, niche and region label prediction to neighborhood cell density and neighborhood composition prediction. We test our approach on large-scale, high-quality spatial transcriptomics data from the brain (mouse, MERFISH), liver (CosMx, human), lung (CosMx, human; Xenium, human) and colon (Xenium, human). Visualized are example slices of the respective datasets colored by niche labels (brain, liver and lung) and cell density (lung and colon). c , The SpatialCorpus-110M is harmonized and mapped to orthologous gene names, as well as human and mouse-specific genes, to create the input for Nicheformer pretraining. We harmonized metadata information across all datasets, capturing species, modality and assay. d , Each cell’s gene expression profile and metadata are fed into a gene-rank tokenizer to obtain a tokenized representation for each cell. The tokenized cells serve as input for the Nicheformer transformer block to predict masked tokens. Finally, the Nicheformer embedding is generated by aggregating the gene tokens . e , The pretrained Nicheformer embedding is visualized as UMAP colored by modality. The UMAP shows a random 5% subsample of the entire Nicheformer embedding ( n = 4,903,086). NA, not applicable.

Journal: Nature Methods

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1038/s41592-025-02814-z

Figure Lengend Snippet: a , Nicheformer is pretrained on the SpatialCorpus-110M, a large data collection of over 110 million cells measured with dissociated and image-based spatial transcriptomics technologies. The SpatialCorpus-110M collection comprises single-cell data from Homo Sapiens and Mus Musculus across 17 distinct organs and 18 cell lines, and additional single-cell data from other anatomical systems and junctions. Shown is an exemplary uniform manifold approximation and projection (UMAP) visualization of a random 1% subset of the entire pretraining dataset ( n = 1,108,759 cells) of the non-integrated log1p-transformed normalized SpatialCorpus-110M colored by modality. b , Nicheformer includes a novel set of downstream tasks, ranging from spatial cell-type, niche and region label prediction to neighborhood cell density and neighborhood composition prediction. We test our approach on large-scale, high-quality spatial transcriptomics data from the brain (mouse, MERFISH), liver (CosMx, human), lung (CosMx, human; Xenium, human) and colon (Xenium, human). Visualized are example slices of the respective datasets colored by niche labels (brain, liver and lung) and cell density (lung and colon). c , The SpatialCorpus-110M is harmonized and mapped to orthologous gene names, as well as human and mouse-specific genes, to create the input for Nicheformer pretraining. We harmonized metadata information across all datasets, capturing species, modality and assay. d , Each cell’s gene expression profile and metadata are fed into a gene-rank tokenizer to obtain a tokenized representation for each cell. The tokenized cells serve as input for the Nicheformer transformer block to predict masked tokens. Finally, the Nicheformer embedding is generated by aggregating the gene tokens . e , The pretrained Nicheformer embedding is visualized as UMAP colored by modality. The UMAP shows a random 5% subsample of the entire Nicheformer embedding ( n = 4,903,086). NA, not applicable.

Article Snippet: The spatial part of the SpatialCorpus-110M consists of datasets measured with image-based spatial transcriptomics technologies, namely CosMx, ISS, MERFISH and 10x Xenium.

Techniques: Transformation Assay, Gene Expression, Blocking Assay, Generated

A) Shown are the F1 scores for niche classification in the CosMx human liver (top left) and lung (top right) datasets, cell type classification in MERFISH mouse brain (bottom right) and the MSE for niche regression in MERFISH mouse brain (bottom left) obtained by different models trained on different data subsets. The results demonstrate a clear advantage of training on spatial data compared to dissociated data. A model trained on just 1% of spatial data significantly outperforms models trained on the same or even three times the amount of dissociated data, reinforcing the fundamental difference between these modalities. This suggests that no amount of dissociated data can fully compensate for the spatial context when evaluated on spatial tasks. Additionally, computational efficiency plays a crucial role: the model trained on a smaller dissociated subset (1%) performs better than one trained on a larger subset (3%) because both were trained for the same duration, leading to more updates per sample in the smaller dataset. Furthermore, stratified training offers advantages only in specific cases, such as the liver, which can be explained by the distribution of tissue types in the random subset - since they are overly present in SpatialCorpus-110M. For example, brain cells are more abundant in the random subset than in the stratified one, potentially influencing performance. The results are found statistically significant even after adjusting for FDR. B) Shown are the F1 score curves of two different models trained on different modalities: spatial and dissociated respectively. Both models have the same number of parameters and have been training for the same amount of time. The task is performed by linear probing. The model trained on MERFISH data notably outperforms the model trained on RNA-seq, highlighting a significant distribution shift between technologies. C) Shown are the F1 scores for niche classification in the CosMx human liver (top left) and lung (top right) datasets, cell type classification in MERFISH mouse brain (bottom right) and the MSE for niche regression in MERFISH mouse brain (bottom right) obtained by different models trained on different data subsets. As in the previous data split test, a broad coverage train distribution is necessary to achieve good performance across a variety of scenarios. In this case, models trained uniquely in mouse data underperform in downstream tasks based on human data (top row); and models trained on only human data underperform in downstream tasks based on mouse data (bottom row). A model trained on a combination of mouse and human data performs on pair in both cases. Results were found statistically significant even after FDR correction.

Journal: Nature Methods

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1038/s41592-025-02814-z

Figure Lengend Snippet: A) Shown are the F1 scores for niche classification in the CosMx human liver (top left) and lung (top right) datasets, cell type classification in MERFISH mouse brain (bottom right) and the MSE for niche regression in MERFISH mouse brain (bottom left) obtained by different models trained on different data subsets. The results demonstrate a clear advantage of training on spatial data compared to dissociated data. A model trained on just 1% of spatial data significantly outperforms models trained on the same or even three times the amount of dissociated data, reinforcing the fundamental difference between these modalities. This suggests that no amount of dissociated data can fully compensate for the spatial context when evaluated on spatial tasks. Additionally, computational efficiency plays a crucial role: the model trained on a smaller dissociated subset (1%) performs better than one trained on a larger subset (3%) because both were trained for the same duration, leading to more updates per sample in the smaller dataset. Furthermore, stratified training offers advantages only in specific cases, such as the liver, which can be explained by the distribution of tissue types in the random subset - since they are overly present in SpatialCorpus-110M. For example, brain cells are more abundant in the random subset than in the stratified one, potentially influencing performance. The results are found statistically significant even after adjusting for FDR. B) Shown are the F1 score curves of two different models trained on different modalities: spatial and dissociated respectively. Both models have the same number of parameters and have been training for the same amount of time. The task is performed by linear probing. The model trained on MERFISH data notably outperforms the model trained on RNA-seq, highlighting a significant distribution shift between technologies. C) Shown are the F1 scores for niche classification in the CosMx human liver (top left) and lung (top right) datasets, cell type classification in MERFISH mouse brain (bottom right) and the MSE for niche regression in MERFISH mouse brain (bottom right) obtained by different models trained on different data subsets. As in the previous data split test, a broad coverage train distribution is necessary to achieve good performance across a variety of scenarios. In this case, models trained uniquely in mouse data underperform in downstream tasks based on human data (top row); and models trained on only human data underperform in downstream tasks based on mouse data (bottom row). A model trained on a combination of mouse and human data performs on pair in both cases. Results were found statistically significant even after FDR correction.

Article Snippet: The spatial part of the SpatialCorpus-110M consists of datasets measured with image-based spatial transcriptomics technologies, namely CosMx, ISS, MERFISH and 10x Xenium.

Techniques: RNA Sequencing

a , Analysis of layer-wise attention patterns reveals that Nicheformer’s later layers consistently pay more attention to contextual tokens across all tissues and modalities, demonstrating a clear and robust hierarchical processing pattern. b , Maximum layer-wise attention paid to gene tokens. For all tissues and modalities, Nicheformer’s middle layers pay the most attention to the gene tokens. c , Single cells resolved in space on an example slice ( n = 2,292 cells) of the MERFISH female mouse brain dataset with cell-type label superimposed. d , e , Single cells resolved in space on an example slice ( n = 2,269 cells) of the MERFISH male mouse brain dataset with the cell-type label ( b ) and CCF acronym label ( c ) superimposed. ADP, anterodorsal preoptic nucleus; AVP, anteroventral preoptic nucleus; HY, hypothalamus; MB, midbrain; MEPO, median preoptic nucleus; MPO, medial preoptic nucleus; NA, nucleus accumbens; IIn, second cranial nerve; OV, organum vasculosum laminae terminalis. f , g , Absolute difference of layer-wise attention scores between male and female MERFISH mouse brain sections show per transformer block of the SDGs considering just the HY GABA cells ( d ) and the entire AVPV section ( e ). h , Maximum layer-wise attention difference across layers between male and female HY GABA cells. The attention paid to random genes and SDGs is equal across all layers except in layers 9 and 10, where there is an increment in the maximum attention paid to the SDGs in comparison to the attention paid to the random set of genes. i , Volcano plot showing the differentially expressed genes (DEGs) highlighting the genes with highest attention difference between sexes (red), and highlighting the SDGs (blue). The genes found with highest attention differences are not among the most differentially expressed. P values were obtained from two-sided Wald tests and adjusted for multiple comparisons using the Benjamini–Hochberg procedure.

Journal: Nature Methods

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1038/s41592-025-02814-z

Figure Lengend Snippet: a , Analysis of layer-wise attention patterns reveals that Nicheformer’s later layers consistently pay more attention to contextual tokens across all tissues and modalities, demonstrating a clear and robust hierarchical processing pattern. b , Maximum layer-wise attention paid to gene tokens. For all tissues and modalities, Nicheformer’s middle layers pay the most attention to the gene tokens. c , Single cells resolved in space on an example slice ( n = 2,292 cells) of the MERFISH female mouse brain dataset with cell-type label superimposed. d , e , Single cells resolved in space on an example slice ( n = 2,269 cells) of the MERFISH male mouse brain dataset with the cell-type label ( b ) and CCF acronym label ( c ) superimposed. ADP, anterodorsal preoptic nucleus; AVP, anteroventral preoptic nucleus; HY, hypothalamus; MB, midbrain; MEPO, median preoptic nucleus; MPO, medial preoptic nucleus; NA, nucleus accumbens; IIn, second cranial nerve; OV, organum vasculosum laminae terminalis. f , g , Absolute difference of layer-wise attention scores between male and female MERFISH mouse brain sections show per transformer block of the SDGs considering just the HY GABA cells ( d ) and the entire AVPV section ( e ). h , Maximum layer-wise attention difference across layers between male and female HY GABA cells. The attention paid to random genes and SDGs is equal across all layers except in layers 9 and 10, where there is an increment in the maximum attention paid to the SDGs in comparison to the attention paid to the random set of genes. i , Volcano plot showing the differentially expressed genes (DEGs) highlighting the genes with highest attention difference between sexes (red), and highlighting the SDGs (blue). The genes found with highest attention differences are not among the most differentially expressed. P values were obtained from two-sided Wald tests and adjusted for multiple comparisons using the Benjamini–Hochberg procedure.

Article Snippet: The spatial part of the SpatialCorpus-110M consists of datasets measured with image-based spatial transcriptomics technologies, namely CosMx, ISS, MERFISH and 10x Xenium.

Techniques: Blocking Assay, Comparison

A-C) Region ( A ), niche ( B ), and cell type ( C ) label distribution across all tissue sections in the MERFISH mouse brain data with the test set highlighted. D) Spatial allocation of cells in the five test tissue sections of the MERFISH mouse brain E) UMAP visualization of the Nicheformer embedding of the MERFISH mouse brain dataset colored by region label. F) Exemplary brain slice of the MERFISH mouse brain dataset colored by region label.

Journal: Nature Methods

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1038/s41592-025-02814-z

Figure Lengend Snippet: A-C) Region ( A ), niche ( B ), and cell type ( C ) label distribution across all tissue sections in the MERFISH mouse brain data with the test set highlighted. D) Spatial allocation of cells in the five test tissue sections of the MERFISH mouse brain E) UMAP visualization of the Nicheformer embedding of the MERFISH mouse brain dataset colored by region label. F) Exemplary brain slice of the MERFISH mouse brain dataset colored by region label.

Article Snippet: The spatial part of the SpatialCorpus-110M consists of datasets measured with image-based spatial transcriptomics technologies, namely CosMx, ISS, MERFISH and 10x Xenium.

Techniques: Slice Preparation

a , Single cells resolved in space on an example slice ( n = 114,396 cells) of the MERFISH mouse brain dataset with niche label superimposed. b , Test-set F1 macro of niche and brain region label prediction of the fine-tuned Nicheformer model, the linear-probing model and a linear-probing baseline computed based on embeddings generated with Geneformer, scGPT, scVI and PCA, respectively. For scVI and PCA, both embeddings generated from a random 1% subset of the SpatialCorpus as well as embeddings generated from the training set of the original dataset are evaluated. c , UMAP of dissociated scRNA-seq dataset with original author cell-type label superimposed. ET, extratelencephalic neurons; IT, intratelencephalic neurons; CT, corticothalamic neurons; NP, near-projecting neurons; OPC, oligodendrocyte precursor cells. d , Nicheformer can transfer spatial niche and region labels onto dissociated single-cell data. e , Nicheformer accurately classifies cells from the dissociated motor cortex to relevant cell types ( n = 9 of 33 distinct ones in the classifier) trained on the whole mouse brain MERFISH dataset. f , g , Nicheformer correctly projects dissociated single cells to niche ( f ) and region ( g ) labels to provide spatially dependent labels. STRd, dorsal striatum; STRv, ventral striatum; RHP, retrohippocampal region; HIP, hippocampal formation; TH, thalamus. f , Nicheformer misclassified parts of layer 2/3 (L2/3) IT neurons as residing in the subpallium GABAergic niche (highlighted in the red box). Additionally, the deep cortical excitatory neurons L6b, L6 CT, L6 IT, and L6 IT Car3 (highlighted in the red box) should be classified as pallium glutamatergic niche instead of subpallium GABAergic by Nicheformer. g , Most of the non-neuronal cells (84.7% of all non-neuronal cells, n = 133) were misclassified as not belonging to the isocortex or the adjacent brain regions (highlighted in the red box). h , Cell-type abundances in the scRNA-seq dataset measuring the primary motor cortex in the mouse. i – k , Classification uncertainty of label transfer of the dissociated scRNA-seq dataset to the MERFISH mouse brain data for cell-type label ( i ), niche label ( j ) and region label ( k ) with a value of 0 representing a high uncertainty and 1 being a lower uncertainty, that is, high certainty. k , Observed high uncertainty for parts of the Glut and GABA neurons for the region prediction of the isocortex, CTXsp and OLF, which are neighboring brain regions.

Journal: Nature Methods

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1038/s41592-025-02814-z

Figure Lengend Snippet: a , Single cells resolved in space on an example slice ( n = 114,396 cells) of the MERFISH mouse brain dataset with niche label superimposed. b , Test-set F1 macro of niche and brain region label prediction of the fine-tuned Nicheformer model, the linear-probing model and a linear-probing baseline computed based on embeddings generated with Geneformer, scGPT, scVI and PCA, respectively. For scVI and PCA, both embeddings generated from a random 1% subset of the SpatialCorpus as well as embeddings generated from the training set of the original dataset are evaluated. c , UMAP of dissociated scRNA-seq dataset with original author cell-type label superimposed. ET, extratelencephalic neurons; IT, intratelencephalic neurons; CT, corticothalamic neurons; NP, near-projecting neurons; OPC, oligodendrocyte precursor cells. d , Nicheformer can transfer spatial niche and region labels onto dissociated single-cell data. e , Nicheformer accurately classifies cells from the dissociated motor cortex to relevant cell types ( n = 9 of 33 distinct ones in the classifier) trained on the whole mouse brain MERFISH dataset. f , g , Nicheformer correctly projects dissociated single cells to niche ( f ) and region ( g ) labels to provide spatially dependent labels. STRd, dorsal striatum; STRv, ventral striatum; RHP, retrohippocampal region; HIP, hippocampal formation; TH, thalamus. f , Nicheformer misclassified parts of layer 2/3 (L2/3) IT neurons as residing in the subpallium GABAergic niche (highlighted in the red box). Additionally, the deep cortical excitatory neurons L6b, L6 CT, L6 IT, and L6 IT Car3 (highlighted in the red box) should be classified as pallium glutamatergic niche instead of subpallium GABAergic by Nicheformer. g , Most of the non-neuronal cells (84.7% of all non-neuronal cells, n = 133) were misclassified as not belonging to the isocortex or the adjacent brain regions (highlighted in the red box). h , Cell-type abundances in the scRNA-seq dataset measuring the primary motor cortex in the mouse. i – k , Classification uncertainty of label transfer of the dissociated scRNA-seq dataset to the MERFISH mouse brain data for cell-type label ( i ), niche label ( j ) and region label ( k ) with a value of 0 representing a high uncertainty and 1 being a lower uncertainty, that is, high certainty. k , Observed high uncertainty for parts of the Glut and GABA neurons for the region prediction of the isocortex, CTXsp and OLF, which are neighboring brain regions.

Article Snippet: The spatial part of the SpatialCorpus-110M consists of datasets measured with image-based spatial transcriptomics technologies, namely CosMx, ISS, MERFISH and 10x Xenium.

Techniques: Generated

A) Downstream task metrics (MSE) for models trained in the MERFISH mouse brain dataset using linear probing on Nicheformer, UCE and CellPLM embeddings. The downstream tasks evaluated are niche regression for 4 different radius sizes. In all cases, Nicheformer outperforms both CellPLM and UCE, being the differences statistically significant. B) F1 Score for region and niche prediction in the MERFISH mouse brain dataset. Likewise, Nicheformer outperforms CellPLM and UCE and the differences are statistically significant. The arrows indicate which direction is the optimal one. For F1 Score, the higher the better; for MSE, the lower the better. C) Downstream task metrics (MSE) for models trained in the CosMX human liver dataset using linear probing on Nicheformer, UCE and CellPLM embeddings. The downstream tasks evaluated are niche regression for 4 different radius sizes. In all cases, Nicheformer outperforms both CellPLM and UCE, being the differences statistically significant. D) Downstream task metrics (MSE) for models trained in the CosMX human liver dataset using linear probing on Nicheformer, UCE and CellPLM embeddings. The downstream tasks evaluated are niche regression for 4 different radius sizes. In all cases, Nicheformer outperforms both CellPLM and UCE, being the differences statistically significant.

Journal: Nature Methods

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1038/s41592-025-02814-z

Figure Lengend Snippet: A) Downstream task metrics (MSE) for models trained in the MERFISH mouse brain dataset using linear probing on Nicheformer, UCE and CellPLM embeddings. The downstream tasks evaluated are niche regression for 4 different radius sizes. In all cases, Nicheformer outperforms both CellPLM and UCE, being the differences statistically significant. B) F1 Score for region and niche prediction in the MERFISH mouse brain dataset. Likewise, Nicheformer outperforms CellPLM and UCE and the differences are statistically significant. The arrows indicate which direction is the optimal one. For F1 Score, the higher the better; for MSE, the lower the better. C) Downstream task metrics (MSE) for models trained in the CosMX human liver dataset using linear probing on Nicheformer, UCE and CellPLM embeddings. The downstream tasks evaluated are niche regression for 4 different radius sizes. In all cases, Nicheformer outperforms both CellPLM and UCE, being the differences statistically significant. D) Downstream task metrics (MSE) for models trained in the CosMX human liver dataset using linear probing on Nicheformer, UCE and CellPLM embeddings. The downstream tasks evaluated are niche regression for 4 different radius sizes. In all cases, Nicheformer outperforms both CellPLM and UCE, being the differences statistically significant.

Article Snippet: The spatial part of the SpatialCorpus-110M consists of datasets measured with image-based spatial transcriptomics technologies, namely CosMx, ISS, MERFISH and 10x Xenium.

Techniques:

A) Downstream task metrics (MSE) for models trained in the MERFISH mouse brain using linear probing on Nicheformer and PCA embeddings with increasingly more principal components. The downstream tasks evaluated are niche regression for 4 different radius sizes. In all cases, Nicheformer outperforms PCA, even though the PCA substantially improves with the more principal components employed. Differences are found statistically significant between the best PCA performing model and Nicheformer. B) F1 Score for region and niche prediction. Interestingly, PCA ends up outperforming Nicheformer in the case of linear probing for the region classification and performing as good as Nicheformer for the niche classification. However, fine tuning Nicheformer is still better. C) Downstream task metrics (MSE) for models trained in the CosMX human liver dataset using linear probing on Nicheformer and PCA embeddings with increasingly more principal components. The downstream tasks evaluated are niche regression for 4 different radius sizes. In all cases, Nicheformer outperforms PCA, even though the PCA substantially improves with the more principal components employed. Differences are found statistically significant between the best PCA performing model and Nicheformer. D) Downstream task metrics (MSE) for models trained in the CosMX human lung dataset using linear probing on Nicheformer and PCA embeddings with increasingly more principal components. The downstream tasks evaluated are niche regression for 4 different radius sizes. In all cases, Nicheformer outperforms PCA, even though the PCA substantially improves with the more principal components employed. Differences are found statistically significant between the best PCA performing model and Nicheformer.

Journal: Nature Methods

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1038/s41592-025-02814-z

Figure Lengend Snippet: A) Downstream task metrics (MSE) for models trained in the MERFISH mouse brain using linear probing on Nicheformer and PCA embeddings with increasingly more principal components. The downstream tasks evaluated are niche regression for 4 different radius sizes. In all cases, Nicheformer outperforms PCA, even though the PCA substantially improves with the more principal components employed. Differences are found statistically significant between the best PCA performing model and Nicheformer. B) F1 Score for region and niche prediction. Interestingly, PCA ends up outperforming Nicheformer in the case of linear probing for the region classification and performing as good as Nicheformer for the niche classification. However, fine tuning Nicheformer is still better. C) Downstream task metrics (MSE) for models trained in the CosMX human liver dataset using linear probing on Nicheformer and PCA embeddings with increasingly more principal components. The downstream tasks evaluated are niche regression for 4 different radius sizes. In all cases, Nicheformer outperforms PCA, even though the PCA substantially improves with the more principal components employed. Differences are found statistically significant between the best PCA performing model and Nicheformer. D) Downstream task metrics (MSE) for models trained in the CosMX human lung dataset using linear probing on Nicheformer and PCA embeddings with increasingly more principal components. The downstream tasks evaluated are niche regression for 4 different radius sizes. In all cases, Nicheformer outperforms PCA, even though the PCA substantially improves with the more principal components employed. Differences are found statistically significant between the best PCA performing model and Nicheformer.

Article Snippet: The spatial part of the SpatialCorpus-110M consists of datasets measured with image-based spatial transcriptomics technologies, namely CosMx, ISS, MERFISH and 10x Xenium.

Techniques:

A-B) Spatial allocation of cells in the healthy CosMx liver section colored by training and test split used for training Nicheformer ( A ) and niche label ( B ). C) Niche label distribution in the training and test set for the healthy CosMx liver dataset. D) Spatial allocation of cells in the cancer CosMx liver section colored by training and test split used for training Nicheformer in the cancer CosMx liver section. E) Distribution of cell type labels in the healthy and cancer CosMx liver data in both training and test set. F) Test-set F1-macro of niche label prediction of the fine-tuned Nicheformer model, the linear probing model, the linear probing model evaluated on a Nicheformer model longer trained in the liver training-set, and a linear probing baseline computed based on embeddings generated with scVI and PCA, respectively. G) The fine-tuned, a multi-task MLP on top of the Nicheformer embedding and the linear probing Nicheformer models outperform zero-shot models trained on scVI and PCA embeddings in terms of mean absolute error across all neighborhood sizes and all three organs, the brain, liver, and lung. H) Left: Fine-tuned Nicheformer performance on the CosMx human liver data grouped by index cell type. Shown are the absolute error values between predicted and observed niche composition vectors for held-out test cells. For each box in (H), the centerline defines the median, the height of the box is given by the interquartile range (IQR), the whiskers are given by 1.5 × IQR and outliers are given as points beyond the minimum or maximum whisker. Right: Index cell type abundances in the entire CosMx human liver dataset. I-M) Nicheformer label transfer classification uncertainty from spatial to dissociated assays in the MERFISH mouse brain dataset. I-K) Cell type ( I ), niche ( J ), and region ( K ) predicted label uncertainty across all cell types in the scRNA-seq mouse brain data. Nicheformer assigns lower uncertainty to plausible labels given the nature of the dataset and high uncertainty to labels not present in the primary motor cortex. The highlighted boxes show cell types, niches and regions one would not expect to find in the primary motor cortex. Nicheformer correctly shows a high uncertainty in those. L-M) Spatial allocation of cells in an exemplary section of the MERFISH mouse brain dataset colored by the pallium glutamatergic niche label ( L ) and the subpallium GABAergic niche label ( M ), respectively.

Journal: Nature Methods

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1038/s41592-025-02814-z

Figure Lengend Snippet: A-B) Spatial allocation of cells in the healthy CosMx liver section colored by training and test split used for training Nicheformer ( A ) and niche label ( B ). C) Niche label distribution in the training and test set for the healthy CosMx liver dataset. D) Spatial allocation of cells in the cancer CosMx liver section colored by training and test split used for training Nicheformer in the cancer CosMx liver section. E) Distribution of cell type labels in the healthy and cancer CosMx liver data in both training and test set. F) Test-set F1-macro of niche label prediction of the fine-tuned Nicheformer model, the linear probing model, the linear probing model evaluated on a Nicheformer model longer trained in the liver training-set, and a linear probing baseline computed based on embeddings generated with scVI and PCA, respectively. G) The fine-tuned, a multi-task MLP on top of the Nicheformer embedding and the linear probing Nicheformer models outperform zero-shot models trained on scVI and PCA embeddings in terms of mean absolute error across all neighborhood sizes and all three organs, the brain, liver, and lung. H) Left: Fine-tuned Nicheformer performance on the CosMx human liver data grouped by index cell type. Shown are the absolute error values between predicted and observed niche composition vectors for held-out test cells. For each box in (H), the centerline defines the median, the height of the box is given by the interquartile range (IQR), the whiskers are given by 1.5 × IQR and outliers are given as points beyond the minimum or maximum whisker. Right: Index cell type abundances in the entire CosMx human liver dataset. I-M) Nicheformer label transfer classification uncertainty from spatial to dissociated assays in the MERFISH mouse brain dataset. I-K) Cell type ( I ), niche ( J ), and region ( K ) predicted label uncertainty across all cell types in the scRNA-seq mouse brain data. Nicheformer assigns lower uncertainty to plausible labels given the nature of the dataset and high uncertainty to labels not present in the primary motor cortex. The highlighted boxes show cell types, niches and regions one would not expect to find in the primary motor cortex. Nicheformer correctly shows a high uncertainty in those. L-M) Spatial allocation of cells in an exemplary section of the MERFISH mouse brain dataset colored by the pallium glutamatergic niche label ( L ) and the subpallium GABAergic niche label ( M ), respectively.

Article Snippet: The spatial part of the SpatialCorpus-110M consists of datasets measured with image-based spatial transcriptomics technologies, namely CosMx, ISS, MERFISH and 10x Xenium.

Techniques: Generated, Whisker Assay

a , We define the neighborhood of a cell as its local neighborhood given a radius and an index cell. The neighborhood cell density is then defined by the number of cells in the neighborhood, and the neighborhood compositions are the proportions of neighboring cell types. b , Neighborhoods are computed at multiple resolutions resulting in different neighborhood size distributions. Each barplot shows the distribution of the number of neighbors across the brain, liver and lung datasets. We extract neighborhoods with the mean number of neighbors 10, 20, 50 and 100 for each dataset. Neighborh., neighborhood. c , The fine-tuned and linear-probing Nicheformer models outperform for brain and lung linear-probing models trained on Geneformer, scGPT, scVI and PCA embeddings in terms of mean absolute error across all neighborhood sizes. Still, it struggles to outperform all benchmarks in liver, where scVI models are very competitive. This is an issue related to the previous liver performance reported in the previous section (Extended Data Figs. and ). d , Left, Fine-tuned Nicheformer performance on the MERFISH mouse brain data grouped by index cell type. Shown are the absolute error values between predicted and observed neighborhood composition vectors for held-out test cells. For each box in d , the centerline defines the median, the height of the box is given by the interquartile range (IQR), the whiskers are given by 1.5 times the IQR, and outliers are given as points beyond the minimum or maximum whisker. Center, Index cell-type abundances in the entire MERFISH mouse brain dataset. Right, UMAPs of MERFISH mouse brain Nicheformer embedding with the selected index cell type as color superimposed. e , UMAP of the Nicheformer embedding of all immune cells in the MERFISH mouse brain dataset with region label as color superimposed.

Journal: Nature Methods

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1038/s41592-025-02814-z

Figure Lengend Snippet: a , We define the neighborhood of a cell as its local neighborhood given a radius and an index cell. The neighborhood cell density is then defined by the number of cells in the neighborhood, and the neighborhood compositions are the proportions of neighboring cell types. b , Neighborhoods are computed at multiple resolutions resulting in different neighborhood size distributions. Each barplot shows the distribution of the number of neighbors across the brain, liver and lung datasets. We extract neighborhoods with the mean number of neighbors 10, 20, 50 and 100 for each dataset. Neighborh., neighborhood. c , The fine-tuned and linear-probing Nicheformer models outperform for brain and lung linear-probing models trained on Geneformer, scGPT, scVI and PCA embeddings in terms of mean absolute error across all neighborhood sizes. Still, it struggles to outperform all benchmarks in liver, where scVI models are very competitive. This is an issue related to the previous liver performance reported in the previous section (Extended Data Figs. and ). d , Left, Fine-tuned Nicheformer performance on the MERFISH mouse brain data grouped by index cell type. Shown are the absolute error values between predicted and observed neighborhood composition vectors for held-out test cells. For each box in d , the centerline defines the median, the height of the box is given by the interquartile range (IQR), the whiskers are given by 1.5 times the IQR, and outliers are given as points beyond the minimum or maximum whisker. Center, Index cell-type abundances in the entire MERFISH mouse brain dataset. Right, UMAPs of MERFISH mouse brain Nicheformer embedding with the selected index cell type as color superimposed. e , UMAP of the Nicheformer embedding of all immune cells in the MERFISH mouse brain dataset with region label as color superimposed.

Article Snippet: The spatial part of the SpatialCorpus-110M consists of datasets measured with image-based spatial transcriptomics technologies, namely CosMx, ISS, MERFISH and 10x Xenium.

Techniques: Whisker Assay

A-B ) Spatial allocation of cells in the training set ( A ) and test set ( B ) tissue sections colored by cell type. C ) Distribution of cell type labels in the training and test set in the CosMx human lung dataset. D-C) Histogram of output token L2 norms for CosMx human lung and liver cells. D-C) The histograms display the distribution of the average L2 norm of output tokens for lung ( D ) and liver ( E ) cells. The modality token, marked by an arrow, exhibits a notably higher norm compared to other tokens. These norms reflect the representation magnitudes in the model’s output space. Including contextual tokens in cell representation aggregation led to poor label transfer performance. This is because aggregation is performed via mean pooling, where tokens with higher norms disproportionately influence the result. Additionally, contextual tokens appear in all cells, whereas the other tokens shown here are present only in specific subsets. As a result, while contextual tokens contribute to all cells, non-contextual tokens contribute only to the cells in which they appear. F-H) Orthologs versus non orthologs comparison. F) Venn diagram showing the number of genes of the non orthologs-trained model (9026) and the orthologs-trained model (7407). The 1619 genes of difference are genes that have a corresponding ortholog but we choose not to use the mapping. G) Niche regression in the MERFISH mouse brain dataset is the only downstream task - among the tested ones - in which there is a statistical significant difference (t-test) between both models. No statistical significance was found in the case of niche prediction for the CosMX human datasets. H) Boxplots showing the distribution of similarities between tokens measured as cosine similarity. We use the official Ensembl releases to map ortholog genes and assess if they are more similar between them than to random genes and we find that they are actually less similar.

Journal: Nature Methods

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1038/s41592-025-02814-z

Figure Lengend Snippet: A-B ) Spatial allocation of cells in the training set ( A ) and test set ( B ) tissue sections colored by cell type. C ) Distribution of cell type labels in the training and test set in the CosMx human lung dataset. D-C) Histogram of output token L2 norms for CosMx human lung and liver cells. D-C) The histograms display the distribution of the average L2 norm of output tokens for lung ( D ) and liver ( E ) cells. The modality token, marked by an arrow, exhibits a notably higher norm compared to other tokens. These norms reflect the representation magnitudes in the model’s output space. Including contextual tokens in cell representation aggregation led to poor label transfer performance. This is because aggregation is performed via mean pooling, where tokens with higher norms disproportionately influence the result. Additionally, contextual tokens appear in all cells, whereas the other tokens shown here are present only in specific subsets. As a result, while contextual tokens contribute to all cells, non-contextual tokens contribute only to the cells in which they appear. F-H) Orthologs versus non orthologs comparison. F) Venn diagram showing the number of genes of the non orthologs-trained model (9026) and the orthologs-trained model (7407). The 1619 genes of difference are genes that have a corresponding ortholog but we choose not to use the mapping. G) Niche regression in the MERFISH mouse brain dataset is the only downstream task - among the tested ones - in which there is a statistical significant difference (t-test) between both models. No statistical significance was found in the case of niche prediction for the CosMX human datasets. H) Boxplots showing the distribution of similarities between tokens measured as cosine similarity. We use the official Ensembl releases to map ortholog genes and assess if they are more similar between them than to random genes and we find that they are actually less similar.

Article Snippet: The spatial part of the SpatialCorpus-110M consists of datasets measured with image-based spatial transcriptomics technologies, namely CosMx, ISS, MERFISH and 10x Xenium.

Techniques: Comparison

Shown are the cumulative explained variance ratios obtained after performing PCA. for the MERFISH brain mouse (top), CosMx human liver (middle) and CosMx human lung (bottom) datasets. Notice that this accounts for the explained variance in the train set, not in the test set (the PCA is computed in the train set and the test data transformer using the principal components obtained). The red line indicates the 90% of explained variance.

Journal: Nature Methods

Article Title: Nicheformer: a foundation model for single-cell and spatial omics

doi: 10.1038/s41592-025-02814-z

Figure Lengend Snippet: Shown are the cumulative explained variance ratios obtained after performing PCA. for the MERFISH brain mouse (top), CosMx human liver (middle) and CosMx human lung (bottom) datasets. Notice that this accounts for the explained variance in the train set, not in the test set (the PCA is computed in the train set and the test data transformer using the principal components obtained). The red line indicates the 90% of explained variance.

Article Snippet: The spatial part of the SpatialCorpus-110M consists of datasets measured with image-based spatial transcriptomics technologies, namely CosMx, ISS, MERFISH and 10x Xenium.

Techniques: