seq Search Results


86
10X Genomics cite seq
Cite Seq, supplied by 10X Genomics, used in various techniques. Bioz Stars score: 86/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/seq/cite+seq/pm39322753-416-124-90
Average 86 stars, based on 1 article reviews
cite seq - by Bioz Stars, 2026-09
86/100 stars
  Buy from Supplier

86
10X Genomics pbmc 4k
a Description of the sequencing budget allocation problem. Consider estimating the underlying gene distribution (top) from the noisy read counts obtained via sequencing (bottom). With a fixed number of reads to be sequenced, deep sequencing of a few cells accurately estimates each individual cell but lacks coverage of the entire distribution (left), whereas a shallow sequencing of many cells covers the entire population but introduces a lot of noise (right). b Optimal tradeoff. The memory T-cell marker gene S100A4 has 41.7k reads in the <t>pbmc_4k</t> dataset. For estimating the underlying gamma distribution \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${X}_{g} \sim {\rm{Gamma}}({r}_{g},{\theta }_{g})$$\end{document} X g ~ Gamma ( r g , θ g ) , the relative error is plotted as a function of the sequencing depth, where the optimal error is obtained at a depth of one read per cell (orange star) and is two times smaller than that at the current depth of pbmc_4k (red triangle). c Experimental design. To determine the sequencing depth for an experiment, first the relative gene expression level can be obtained via pilot experiments or previous studies (top left). Then the researcher can select a set of genes of interest (i.e., some marker genes highlighted as black dots), of which the smallest relative expression level \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${p}^{* }$$\end{document} p * ( MS4A1 ) defines the reliable detection limit. Finally, the optimal sequencing depth is determined as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${n}_{{\rm{reads}}}^{* }=1/{p}^{* }$$\end{document} n reads * = 1 ∕ p * (top right). The errors under different tradeoffs are visualized as a function of the genes ordered from the most expressed to the least (bottom). The optimal sequencing budget allocation (orange) minimizes the worst-case error over all the genes of interest (left of the red dashed line), whereas both the deeper sequencing (green) and the shallower sequencing (blue) yield worse results.
Pbmc 4k, supplied by 10X Genomics, used in various techniques. Bioz Stars score: 86/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/seq/data+pbmc+scrna+seq/pmc07005864-337-36-15
Average 86 stars, based on 1 article reviews
pbmc 4k - by Bioz Stars, 2026-09
86/100 stars
  Buy from Supplier

86
GenScript corporation seq id
a Description of the sequencing budget allocation problem. Consider estimating the underlying gene distribution (top) from the noisy read counts obtained via sequencing (bottom). With a fixed number of reads to be sequenced, deep sequencing of a few cells accurately estimates each individual cell but lacks coverage of the entire distribution (left), whereas a shallow sequencing of many cells covers the entire population but introduces a lot of noise (right). b Optimal tradeoff. The memory T-cell marker gene S100A4 has 41.7k reads in the <t>pbmc_4k</t> dataset. For estimating the underlying gamma distribution \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${X}_{g} \sim {\rm{Gamma}}({r}_{g},{\theta }_{g})$$\end{document} X g ~ Gamma ( r g , θ g ) , the relative error is plotted as a function of the sequencing depth, where the optimal error is obtained at a depth of one read per cell (orange star) and is two times smaller than that at the current depth of pbmc_4k (red triangle). c Experimental design. To determine the sequencing depth for an experiment, first the relative gene expression level can be obtained via pilot experiments or previous studies (top left). Then the researcher can select a set of genes of interest (i.e., some marker genes highlighted as black dots), of which the smallest relative expression level \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${p}^{* }$$\end{document} p * ( MS4A1 ) defines the reliable detection limit. Finally, the optimal sequencing depth is determined as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${n}_{{\rm{reads}}}^{* }=1/{p}^{* }$$\end{document} n reads * = 1 ∕ p * (top right). The errors under different tradeoffs are visualized as a function of the genes ordered from the most expressed to the least (bottom). The optimal sequencing budget allocation (orange) minimizes the worst-case error over all the genes of interest (left of the red dashed line), whereas both the deeper sequencing (green) and the shallower sequencing (blue) yield worse results.
Seq Id, supplied by GenScript corporation, used in various techniques. Bioz Stars score: 86/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/seq/seq+id/us10435668-683-43-49
Average 86 stars, based on 1 article reviews
seq id - by Bioz Stars, 2026-09
86/100 stars
  Buy from Supplier

86
10X Genomics 10x genomics scrna seq libraries
(a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate scRNA-seq data across four platforms <t>(10X</t> Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.
10x Genomics Scrna Seq Libraries, supplied by 10X Genomics, used in various techniques. Bioz Stars score: 86/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/seq/libraries+scrna+seq/pmc11245320-481-36-36
Average 86 stars, based on 1 article reviews
10x genomics scrna seq libraries - by Bioz Stars, 2026-09
86/100 stars
  Buy from Supplier

86
10X Genomics 10x genomics scrna seq
(a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate scRNA-seq data across four platforms <t>(10X</t> Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.
10x Genomics Scrna Seq, supplied by 10X Genomics, used in various techniques. Bioz Stars score: 86/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/seq/scrna+seq/pmc11245320-633-21-21
Average 86 stars, based on 1 article reviews
10x genomics scrna seq - by Bioz Stars, 2026-09
86/100 stars
  Buy from Supplier

86
10X Genomics 10x genomics scrna datasets
(a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate <t>scRNA-seq</t> data across four platforms <t>(10X</t> Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.
10x Genomics Scrna Datasets, supplied by 10X Genomics, used in various techniques. Bioz Stars score: 86/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/seq/datasets+scrna+seq/pmc11245320-971-1-1
Average 86 stars, based on 1 article reviews
10x genomics scrna datasets - by Bioz Stars, 2026-09
86/100 stars
  Buy from Supplier

86
Biokey American Instrument Inc scrna seq datasets
(a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate <t>scRNA-seq</t> data across four platforms <t>(10X</t> Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.
Scrna Seq Datasets, supplied by Biokey American Instrument Inc, used in various techniques. Bioz Stars score: 86/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/seq/datasets+scrna+seq/pmc13138057-430-1-20
Average 86 stars, based on 1 article reviews
scrna seq datasets - by Bioz Stars, 2026-09
86/100 stars
  Buy from Supplier

86
Medicago seq id
(a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate <t>scRNA-seq</t> data across four platforms <t>(10X</t> Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.
Seq Id, supplied by Medicago, used in various techniques. Bioz Stars score: 86/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/seq/id+seq/us12516336-417-51-131
Average 86 stars, based on 1 article reviews
seq id - by Bioz Stars, 2026-09
86/100 stars
  Buy from Supplier

95
Addgene inc hstarr seq ori vector
(a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate <t>scRNA-seq</t> data across four platforms <t>(10X</t> Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.
Hstarr Seq Ori Vector, supplied by Addgene inc, used in various techniques. Bioz Stars score: 95/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/seq/hSTARR-seq_ORI+vector+(Plasmid+%2399296)/pmc09522591-41-12-14
Average 95 stars, based on 1 article reviews
hstarr seq ori vector - by Bioz Stars, 2026-09
95/100 stars
  Buy from Supplier

93
fluidigm c1tm single cell mrna seq high throughput ht
(a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate <t>scRNA-seq</t> data across four platforms <t>(10X</t> Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.
C1tm Single Cell Mrna Seq High Throughput Ht, supplied by fluidigm, used in various techniques. Bioz Stars score: 93/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/seq/C1+Single-Cell+mRNA+Seq+HT/bio_rxiv__2020__10__14__339259-202-5-22
Average 93 stars, based on 1 article reviews
c1tm single cell mrna seq high throughput ht - by Bioz Stars, 2026-09
93/100 stars
  Buy from Supplier

94
fluidigm mrna seq
(a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate <t>scRNA-seq</t> data across four platforms <t>(10X</t> Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.
Mrna Seq, supplied by fluidigm, used in various techniques. Bioz Stars score: 94/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/seq/C1+Single-Cell+mRNA+Seq+IFC/pmc06631348-200-13-17
Average 94 stars, based on 1 article reviews
mrna seq - by Bioz Stars, 2026-09
94/100 stars
  Buy from Supplier

92
fluidigm c1 single cell auto prep ifc
(a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate <t>scRNA-seq</t> data across four platforms <t>(10X</t> Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.
C1 Single Cell Auto Prep Ifc, supplied by fluidigm, used in various techniques. Bioz Stars score: 92/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
https://www.bioz.com/product/seq/C1+Single-Cell+mRNA+Seq+IFC/pmc06303105-319-22-30
Average 92 stars, based on 1 article reviews
c1 single cell auto prep ifc - by Bioz Stars, 2026-09
92/100 stars
  Buy from Supplier

Image Search Results


a Description of the sequencing budget allocation problem. Consider estimating the underlying gene distribution (top) from the noisy read counts obtained via sequencing (bottom). With a fixed number of reads to be sequenced, deep sequencing of a few cells accurately estimates each individual cell but lacks coverage of the entire distribution (left), whereas a shallow sequencing of many cells covers the entire population but introduces a lot of noise (right). b Optimal tradeoff. The memory T-cell marker gene S100A4 has 41.7k reads in the pbmc_4k dataset. For estimating the underlying gamma distribution \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${X}_{g} \sim {\rm{Gamma}}({r}_{g},{\theta }_{g})$$\end{document} X g ~ Gamma ( r g , θ g ) , the relative error is plotted as a function of the sequencing depth, where the optimal error is obtained at a depth of one read per cell (orange star) and is two times smaller than that at the current depth of pbmc_4k (red triangle). c Experimental design. To determine the sequencing depth for an experiment, first the relative gene expression level can be obtained via pilot experiments or previous studies (top left). Then the researcher can select a set of genes of interest (i.e., some marker genes highlighted as black dots), of which the smallest relative expression level \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${p}^{* }$$\end{document} p * ( MS4A1 ) defines the reliable detection limit. Finally, the optimal sequencing depth is determined as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${n}_{{\rm{reads}}}^{* }=1/{p}^{* }$$\end{document} n reads * = 1 ∕ p * (top right). The errors under different tradeoffs are visualized as a function of the genes ordered from the most expressed to the least (bottom). The optimal sequencing budget allocation (orange) minimizes the worst-case error over all the genes of interest (left of the red dashed line), whereas both the deeper sequencing (green) and the shallower sequencing (blue) yield worse results.

Journal: Nature Communications

Article Title: Determining sequencing depth in a single-cell RNA-seq experiment

doi: 10.1038/s41467-020-14482-y

Figure Lengend Snippet: a Description of the sequencing budget allocation problem. Consider estimating the underlying gene distribution (top) from the noisy read counts obtained via sequencing (bottom). With a fixed number of reads to be sequenced, deep sequencing of a few cells accurately estimates each individual cell but lacks coverage of the entire distribution (left), whereas a shallow sequencing of many cells covers the entire population but introduces a lot of noise (right). b Optimal tradeoff. The memory T-cell marker gene S100A4 has 41.7k reads in the pbmc_4k dataset. For estimating the underlying gamma distribution \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${X}_{g} \sim {\rm{Gamma}}({r}_{g},{\theta }_{g})$$\end{document} X g ~ Gamma ( r g , θ g ) , the relative error is plotted as a function of the sequencing depth, where the optimal error is obtained at a depth of one read per cell (orange star) and is two times smaller than that at the current depth of pbmc_4k (red triangle). c Experimental design. To determine the sequencing depth for an experiment, first the relative gene expression level can be obtained via pilot experiments or previous studies (top left). Then the researcher can select a set of genes of interest (i.e., some marker genes highlighted as black dots), of which the smallest relative expression level \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${p}^{* }$$\end{document} p * ( MS4A1 ) defines the reliable detection limit. Finally, the optimal sequencing depth is determined as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${n}_{{\rm{reads}}}^{* }=1/{p}^{* }$$\end{document} n reads * = 1 ∕ p * (top right). The errors under different tradeoffs are visualized as a function of the genes ordered from the most expressed to the least (bottom). The optimal sequencing budget allocation (orange) minimizes the worst-case error over all the genes of interest (left of the red dashed line), whereas both the deeper sequencing (green) and the shallower sequencing (blue) yield worse results.

Article Snippet: They are publicly available and can be downloaded via the following links: pbmc_4k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/pbmc4k pbmc_8k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/pbmc8k brain_1k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/neurons_900 brain_2k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/neurons_2000 brain_9k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/neuron_9k brain_1.3m: https://support.10xgenomics.com/single-cell-gene-expression/datasets/1.3.0/1M_neurons 293T_1k, 3T3_1k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/hgmm_1k 293T_6k, 3T3_6k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/hgmm_6k 293T_12k, 3T3_12k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/hgmm_12k We note that pbmc_4k and pbmc_8k are from the same donor; brain_1k and brain_9k are also from the same donor.

Techniques: Sequencing, Marker, Expressing

a Top: for estimating the coefficient of variation (CV), the plug-in estimates become more inflated as the sequencing depth becomes shallower (from right to left along the x axis), whereas the EB estimates are consistent. 3-std confidence intervals are provided for this panel. Middle: both brain_1k and brain_1.3m are from the mouse brain, and hence each gene should have a similar CV value between the two datasets. This is indeed the case for the EB estimator (right), which is adaptive to different sequencing depths. However, as brain_1k is twice deeper than brain_1.3m, the plug-in estimates are biased that most points are above the 45-degree line (red). Bottom: distribution recovery for the gene GZMA from a dataset that is subsampled to be five times shallower (left). The EB estimator provides a reasonable estimation for both the zero proportion and the tail shape, resulting in a small total variation error (right). b Feature selection and PCA. The task is to first select features (genes) based on CV, and then perform PCA on the selected features. The results on the full data (pbmc_4k) and a subsampled (three times shallower) are compared. EB estimates are more consistent between the full data and the subsampled data for both the CV ranks (top) and the PCA plots (bottom).

Journal: Nature Communications

Article Title: Determining sequencing depth in a single-cell RNA-seq experiment

doi: 10.1038/s41467-020-14482-y

Figure Lengend Snippet: a Top: for estimating the coefficient of variation (CV), the plug-in estimates become more inflated as the sequencing depth becomes shallower (from right to left along the x axis), whereas the EB estimates are consistent. 3-std confidence intervals are provided for this panel. Middle: both brain_1k and brain_1.3m are from the mouse brain, and hence each gene should have a similar CV value between the two datasets. This is indeed the case for the EB estimator (right), which is adaptive to different sequencing depths. However, as brain_1k is twice deeper than brain_1.3m, the plug-in estimates are biased that most points are above the 45-degree line (red). Bottom: distribution recovery for the gene GZMA from a dataset that is subsampled to be five times shallower (left). The EB estimator provides a reasonable estimation for both the zero proportion and the tail shape, resulting in a small total variation error (right). b Feature selection and PCA. The task is to first select features (genes) based on CV, and then perform PCA on the selected features. The results on the full data (pbmc_4k) and a subsampled (three times shallower) are compared. EB estimates are more consistent between the full data and the subsampled data for both the CV ranks (top) and the PCA plots (bottom).

Article Snippet: They are publicly available and can be downloaded via the following links: pbmc_4k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/pbmc4k pbmc_8k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/pbmc8k brain_1k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/neurons_900 brain_2k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/neurons_2000 brain_9k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/neuron_9k brain_1.3m: https://support.10xgenomics.com/single-cell-gene-expression/datasets/1.3.0/1M_neurons 293T_1k, 3T3_1k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/hgmm_1k 293T_6k, 3T3_6k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/hgmm_6k 293T_12k, 3T3_12k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/hgmm_12k We note that pbmc_4k and pbmc_8k are from the same donor; brain_1k and brain_9k are also from the same donor.

Techniques: Sequencing, Selection

a Top: the EB-estimated Pearson correlation for some marker genes in pbmc_4k are visualized, ordered by different cell populations (top). The clear block-diagonal structure implies that the EB estimator is capable of capturing the gene functional groups. As a comparison, the plug-in estimator also recovers those modules but with a weaker contrast (bottom left panel, plug-in with 100%). Bottom: a subsample experiment further shows that the EB estimator can recover the module with 5% of the data. For the plug-in estimator, the first block (T cells) is blurred with 25% of the data, and the entire structure vanishes with 10% of the data. b Gene network based on the EB-estimated Pearson correlation using the pbmc_4k dataset. Most gene modules correspond to important cell types or functions, including T cells, B cells, NK-cells, myeloid-derived cells, megakaryocytes/platelets, ribosomal protein genes, and mitochondrially encoded protein-coding genes. c Left: the estimated Pearson correlations between all genes and LCK (1st panel) and CD3D (2nd panel), two known T-cell markers. There are three modes for the EB-estimated values, where the positive mode, the zero mode, and the negative mode correspond to genes in the same module, different modules, and irrelevant genes, respectively. The plug-in estimated values are nonetheless much closer to zero even for the truly correlated ones, indicating an artificial shrinkage of the estimated values. Right: two instances where the EB estimates are significantly different from the plug-in estimates. The axes represent read counts, and the color codes the number of cells. Both gene pairs are biologically validated (see Gene network analysis in Methods). See also Supplementary Figs. – for more examples.

Journal: Nature Communications

Article Title: Determining sequencing depth in a single-cell RNA-seq experiment

doi: 10.1038/s41467-020-14482-y

Figure Lengend Snippet: a Top: the EB-estimated Pearson correlation for some marker genes in pbmc_4k are visualized, ordered by different cell populations (top). The clear block-diagonal structure implies that the EB estimator is capable of capturing the gene functional groups. As a comparison, the plug-in estimator also recovers those modules but with a weaker contrast (bottom left panel, plug-in with 100%). Bottom: a subsample experiment further shows that the EB estimator can recover the module with 5% of the data. For the plug-in estimator, the first block (T cells) is blurred with 25% of the data, and the entire structure vanishes with 10% of the data. b Gene network based on the EB-estimated Pearson correlation using the pbmc_4k dataset. Most gene modules correspond to important cell types or functions, including T cells, B cells, NK-cells, myeloid-derived cells, megakaryocytes/platelets, ribosomal protein genes, and mitochondrially encoded protein-coding genes. c Left: the estimated Pearson correlations between all genes and LCK (1st panel) and CD3D (2nd panel), two known T-cell markers. There are three modes for the EB-estimated values, where the positive mode, the zero mode, and the negative mode correspond to genes in the same module, different modules, and irrelevant genes, respectively. The plug-in estimated values are nonetheless much closer to zero even for the truly correlated ones, indicating an artificial shrinkage of the estimated values. Right: two instances where the EB estimates are significantly different from the plug-in estimates. The axes represent read counts, and the color codes the number of cells. Both gene pairs are biologically validated (see Gene network analysis in Methods). See also Supplementary Figs. – for more examples.

Article Snippet: They are publicly available and can be downloaded via the following links: pbmc_4k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/pbmc4k pbmc_8k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/pbmc8k brain_1k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/neurons_900 brain_2k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/neurons_2000 brain_9k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/neuron_9k brain_1.3m: https://support.10xgenomics.com/single-cell-gene-expression/datasets/1.3.0/1M_neurons 293T_1k, 3T3_1k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/hgmm_1k 293T_6k, 3T3_6k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/hgmm_6k 293T_12k, 3T3_12k: https://support.10xgenomics.com/single-cell-gene-expression/datasets/2.1.0/hgmm_12k We note that pbmc_4k and pbmc_8k are from the same donor; brain_1k and brain_9k are also from the same donor.

Techniques: Marker, Blocking Assay, Functional Assay, Derivative Assay

(a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate scRNA-seq data across four platforms (10X Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate scRNA-seq data across four platforms (10X Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques: Modification, Sequencing, RNA Sequencing

The violin plot shows the number of genes detected in each cell across 20 scRNA-seq datasets. The plot was generated using Seurat (version 3.1). Each dot represents a single cell. The violin shapes summarize the data distributions, which are colored in the background to signify each of the 20 different scRNA seq datasets. Each scRNA-seq dataset is plotted on the X-axis; the Y-axis shows the corresponding number of genes detected in a cell (nGene) for that dataset. The average number of genes detected in each cell was about 4000 and most of the cells had 2500–7500 genes, except for samples C1_LLU_A and C1_LLU_B. The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: The violin plot shows the number of genes detected in each cell across 20 scRNA-seq datasets. The plot was generated using Seurat (version 3.1). Each dot represents a single cell. The violin shapes summarize the data distributions, which are colored in the background to signify each of the 20 different scRNA seq datasets. Each scRNA-seq dataset is plotted on the X-axis; the Y-axis shows the corresponding number of genes detected in a cell (nGene) for that dataset. The average number of genes detected in each cell was about 4000 and most of the cells had 2500–7500 genes, except for samples C1_LLU_A and C1_LLU_B. The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques: Generated

(a–c) Evaluation of the UMI-based (10X) data with Cell Ranger, UMI-Tools, or zUMIs. (d–e) Evaluation of data from non-UMI based technologies C1 full-length transcript, C1 HT, and ICELL8 full-length transcript using FeatureCounts, Kallisto, or RSEM. (a) Bar plot showing the number of cells captured with UMI-based technology; (b) and (d) Box plot showing the number of genes detected per cell in UMI-based and non-UMI based technologies, respectively; (c) and (e) Violin plots showing the gene expression correlation and consensus genes [represented by IoU (Intersection over Union)] per cell between any two pipelines in UMI-based and non-UMI based technologies, respectively. The sample sizes (n) used to derive statistics in (b) and (d) were: (b) 10X_LLU_A, n= 3045 cells; 10X_NCI_A, n=6425 cells; 10X_NCI_M_A, n=6483 cells; 10X_LLU_B, n=1439 cells; 10X_NCI_B, n=3296 cells; 10X_NCI_M_B, n=3273 cells; (d) C1_LLU_A, n=80 cells; C1_FDA_HT_A, n=203 cells; ICELL8_SE_A, n=600 cells; ICELL8_PE_A, n=598 cells; C1_LLU_B, n=66 cells; C1_FDA_B, n=241 cells; ICELL8_SE_B, n=600 cells; ICELL8_PE_B, n=596 cells. For detailed statistics regarding minima, maxima, centre, bounds of box and whiskers and percentile related to the figure, please refer to Supplementary Table 5.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a–c) Evaluation of the UMI-based (10X) data with Cell Ranger, UMI-Tools, or zUMIs. (d–e) Evaluation of data from non-UMI based technologies C1 full-length transcript, C1 HT, and ICELL8 full-length transcript using FeatureCounts, Kallisto, or RSEM. (a) Bar plot showing the number of cells captured with UMI-based technology; (b) and (d) Box plot showing the number of genes detected per cell in UMI-based and non-UMI based technologies, respectively; (c) and (e) Violin plots showing the gene expression correlation and consensus genes [represented by IoU (Intersection over Union)] per cell between any two pipelines in UMI-based and non-UMI based technologies, respectively. The sample sizes (n) used to derive statistics in (b) and (d) were: (b) 10X_LLU_A, n= 3045 cells; 10X_NCI_A, n=6425 cells; 10X_NCI_M_A, n=6483 cells; 10X_LLU_B, n=1439 cells; 10X_NCI_B, n=3296 cells; 10X_NCI_M_B, n=3273 cells; (d) C1_LLU_A, n=80 cells; C1_FDA_HT_A, n=203 cells; ICELL8_SE_A, n=600 cells; ICELL8_PE_A, n=598 cells; C1_LLU_B, n=66 cells; C1_FDA_B, n=241 cells; ICELL8_SE_B, n=600 cells; ICELL8_PE_B, n=596 cells. For detailed statistics regarding minima, maxima, centre, bounds of box and whiskers and percentile related to the figure, please refer to Supplementary Table 5.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques: Gene Expression

(a) Batch-effect correction in Scenario #1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395; and Sample B, B-lymphocyte line HCC1395BL). Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Batch-effect correction in Scenario #2, where five scRNA-seq datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells were generated separately at the four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Batch-effect correction in Scenario #3, where five scRNA-seq datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from the B lymphocytes were generated separately at the four centers on the same four platforms; (d) Batch-effect correction in Scenario #4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, 10X_NCI_M_Mix5_F2) were generated from 5% or 10% breast cancer cells spiked into B lymphocytes, and analyzed with the 10X Genomics platform at two centers in four different batches. Each dataset is indicated by a unique color in panels (a) to (d). Idealized projection of cells for the four different scenarios is presented on the left. *Note for BBKNN, only UMAP is available and shown. Silhouette width score quantifying the clusterability for (e) Scenario #1 or (f) Scenario #4, corresponding to panels (a) and (d), respectively. (g) kBET acceptance score quantifying the mixability, calculated using the cross-platform/center scRNA-seq data acquired either from breast cancer cells only or from B-lymphocytes only for all four scenarios (a-d, also labeled as Scenarios #1–#4).

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a) Batch-effect correction in Scenario #1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395; and Sample B, B-lymphocyte line HCC1395BL). Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Batch-effect correction in Scenario #2, where five scRNA-seq datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells were generated separately at the four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Batch-effect correction in Scenario #3, where five scRNA-seq datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from the B lymphocytes were generated separately at the four centers on the same four platforms; (d) Batch-effect correction in Scenario #4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, 10X_NCI_M_Mix5_F2) were generated from 5% or 10% breast cancer cells spiked into B lymphocytes, and analyzed with the 10X Genomics platform at two centers in four different batches. Each dataset is indicated by a unique color in panels (a) to (d). Idealized projection of cells for the four different scenarios is presented on the left. *Note for BBKNN, only UMAP is available and shown. Silhouette width score quantifying the clusterability for (e) Scenario #1 or (f) Scenario #4, corresponding to panels (a) and (d), respectively. (g) kBET acceptance score quantifying the mixability, calculated using the cross-platform/center scRNA-seq data acquired either from breast cancer cells only or from B-lymphocytes only for all four scenarios (a-d, also labeled as Scenarios #1–#4).

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques: Generated, Labeling

Batch-effect corrections were performed for the following four scenarios: (a) Scenario 1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395 and Sample B, B-lymphocyte line HCC1395BL); Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Scenario 2, where five datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells (Sample A, HCC1395) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Scenario 3, where five datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from B-lymphocytes (Sample B, HCC1395BL) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); and (d) Scenario 4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, 10X_NCI_M_Mix5_F2) were generated from 5% or 10% of breast cancer cells (Sample A, HCC1395) spiked into the B-lymphocytes (Sample B, HCC1395BL) and analyzed with the 10X Genomics platform at two centers (LLU and NCI) in four different batches. *For BBKNN, only UMAPs were available and shown in (a-d). The HCC1395 breast cancer cells (Sample A) were labeled in red and the HCC1395BL B lymphocytes (Sample B) were labeled in blue. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). The top 2000 HVGs were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Batch-effect corrections were performed for the following four scenarios: (a) Scenario 1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395 and Sample B, B-lymphocyte line HCC1395BL); Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Scenario 2, where five datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells (Sample A, HCC1395) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Scenario 3, where five datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from B-lymphocytes (Sample B, HCC1395BL) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); and (d) Scenario 4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, 10X_NCI_M_Mix5_F2) were generated from 5% or 10% of breast cancer cells (Sample A, HCC1395) spiked into the B-lymphocytes (Sample B, HCC1395BL) and analyzed with the 10X Genomics platform at two centers (LLU and NCI) in four different batches. *For BBKNN, only UMAPs were available and shown in (a-d). The HCC1395 breast cancer cells (Sample A) were labeled in red and the HCC1395BL B lymphocytes (Sample B) were labeled in blue. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). The top 2000 HVGs were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques: Generated, Labeling

Boxplot of silhouette values stratified by eight normalization methods across 14 datasets, including (a) 10X_LLU, (b) 10X_NCI, (c) 10X_NCI_M, (d) C1_FDA_HT, (e) C1_LLU, (f) ICELL8_PE, and (g) ICELL8_SE in breast cancer cells (HCC1395; Sample A) and B lymphocytes (HCC1395BL; Sample B). Eight normalization methods included SCTransform, Scran Deconvolution, CPM, LogCPM, TMM, DESeq, Quantile, and Linnorm. For each dataset, reads of each cell were down-sampled to two different read depths (10K and 100K per cell) before calculating the silhouette width values. LogCPM normalization performed fairly well and was used as the default normalization for our subsequent batch-effect correction analyses. Two normalization methods developed for bulk cell RNA-seq (TMM and Quantile) had the lowest scores. The sample sizes (n) used to derive statistics were: 10X_LLU_A, n= 3560 cells, 10X_LLU_B, n=1770 cells; 10X_NCI_A, n=4284 cells, 10X_NCI_B, n=4136 cells; 10X_NCI_M_A, n=1372 cells, 10X_NCI_M_B, n=2082 cells; C1_LLU_A, n=160 cells, C1_LLU_B, n=132 cells; C1_FDA_HT_A, n=318 cells, C1_FDA_HT_B, n=374 cells; ICELL8_SE_A, n=1134 cells, ICELL8_SE_B, n=1078 cells; ICELL8_PE_A, n=980 cells, ICELL8_PE_B, n=954 cells). For detailed statistics regarding minima, maxima, centre, bounds of box and whiskers and percentile related to the figure, please refer to Supplementary Table 6.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Boxplot of silhouette values stratified by eight normalization methods across 14 datasets, including (a) 10X_LLU, (b) 10X_NCI, (c) 10X_NCI_M, (d) C1_FDA_HT, (e) C1_LLU, (f) ICELL8_PE, and (g) ICELL8_SE in breast cancer cells (HCC1395; Sample A) and B lymphocytes (HCC1395BL; Sample B). Eight normalization methods included SCTransform, Scran Deconvolution, CPM, LogCPM, TMM, DESeq, Quantile, and Linnorm. For each dataset, reads of each cell were down-sampled to two different read depths (10K and 100K per cell) before calculating the silhouette width values. LogCPM normalization performed fairly well and was used as the default normalization for our subsequent batch-effect correction analyses. Two normalization methods developed for bulk cell RNA-seq (TMM and Quantile) had the lowest scores. The sample sizes (n) used to derive statistics were: 10X_LLU_A, n= 3560 cells, 10X_LLU_B, n=1770 cells; 10X_NCI_A, n=4284 cells, 10X_NCI_B, n=4136 cells; 10X_NCI_M_A, n=1372 cells, 10X_NCI_M_B, n=2082 cells; C1_LLU_A, n=160 cells, C1_LLU_B, n=132 cells; C1_FDA_HT_A, n=318 cells, C1_FDA_HT_B, n=374 cells; ICELL8_SE_A, n=1134 cells, ICELL8_SE_B, n=1078 cells; ICELL8_PE_A, n=980 cells, ICELL8_PE_B, n=954 cells). For detailed statistics regarding minima, maxima, centre, bounds of box and whiskers and percentile related to the figure, please refer to Supplementary Table 6.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques: RNA Sequencing

Five different batches of scRNA-seq data (10X_LLU_A, 10X_LLU_B, 10X_NCI_A, 10X_NCI_B, and 10X_NCI_Mix5) generated at two sites (LLU and NCI) are shown either as t-SNE plots (panels a-d) or as UMAPs (panels e-h). (a) LogNormalized, scaled data with no regression; (b) LogNormalized, scaled data filtered with mitochondrial (Mito) gene regression >5% and UMI normalization by Seurat v3; (c) ScTransform with no regression; (d) SCTransform with mitochondrial gene regression and UMI normalization; (e) LogNormalized, scaled data with no regression; (f) scaled data with mitochondrial gene regression and UMI normalization; (g) SCTransform with no regression; and (h) SCTransform with mitochondrial gene regression and UMI normalization.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Five different batches of scRNA-seq data (10X_LLU_A, 10X_LLU_B, 10X_NCI_A, 10X_NCI_B, and 10X_NCI_Mix5) generated at two sites (LLU and NCI) are shown either as t-SNE plots (panels a-d) or as UMAPs (panels e-h). (a) LogNormalized, scaled data with no regression; (b) LogNormalized, scaled data filtered with mitochondrial (Mito) gene regression >5% and UMI normalization by Seurat v3; (c) ScTransform with no regression; (d) SCTransform with mitochondrial gene regression and UMI normalization; (e) LogNormalized, scaled data with no regression; (f) scaled data with mitochondrial gene regression and UMI normalization; (g) SCTransform with no regression; and (h) SCTransform with mitochondrial gene regression and UMI normalization.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques: Generated

Batch-effect corrections were performed for the following four scenarios: (a) Scenario 1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395 and Sample B, B-lymphocyte line HCC1395BL); Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Scenario 2, where five datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells (Sample A, HCC1395) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Scenario 3, where five datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from B-lymphocytes (Sample B, HCC1395BL) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); and (d) Scenario 4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were generated from 5% or 10% of breast cancer cells (Sample A, HCC1395), spiked into the B-lymphocytes (Sample B, HCC1395BL), and analyzed with the 10X Genomics platform at two centers (LLU and NCI) in four different batches. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). The top 2000 highly variable genes (HVGs) of these datasets were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Batch-effect corrections were performed for the following four scenarios: (a) Scenario 1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395 and Sample B, B-lymphocyte line HCC1395BL); Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Scenario 2, where five datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells (Sample A, HCC1395) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Scenario 3, where five datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from B-lymphocytes (Sample B, HCC1395BL) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); and (d) Scenario 4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were generated from 5% or 10% of breast cancer cells (Sample A, HCC1395), spiked into the B-lymphocytes (Sample B, HCC1395BL), and analyzed with the 10X Genomics platform at two centers (LLU and NCI) in four different batches. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). The top 2000 highly variable genes (HVGs) of these datasets were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques: Generated

t-SNE plots and UMAPs showing the batch-effect corrections performed by seven methods using 20 scRNA-seq datasets across different platforms. Datasets from 10X were down-sampled to 1200 cells per dataset. *Note, for BBKNN, only UMAP was available and shown. The scRNA-seq datasets are colored to identify the four different platforms: 10X 3´ scRNA-seq platform (red), C1 3´ HT scRNA-seq platform (yellow), C1 full-length scRNA-seq platform (light blue), and ICELL8 full-length scRNA-seq platform (dark blue). Batch correction methods included: Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). Scanorama failed to separate two cell types into discrete clusters when non-10X platforms were included in the analysis. The top 2000 HVGs across all datasets were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: t-SNE plots and UMAPs showing the batch-effect corrections performed by seven methods using 20 scRNA-seq datasets across different platforms. Datasets from 10X were down-sampled to 1200 cells per dataset. *Note, for BBKNN, only UMAP was available and shown. The scRNA-seq datasets are colored to identify the four different platforms: 10X 3´ scRNA-seq platform (red), C1 3´ HT scRNA-seq platform (yellow), C1 full-length scRNA-seq platform (light blue), and ICELL8 full-length scRNA-seq platform (dark blue). Batch correction methods included: Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). Scanorama failed to separate two cell types into discrete clusters when non-10X platforms were included in the analysis. The top 2000 HVGs across all datasets were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques:

(a) t-SNE plot and (b) UMAP showing batch-effect corrections using twelve 10X Genomics scRNA-seq datasets consisting of both mixed and non-mixed samples from two sites (LLU and NCI) in different batches after Scanorama (version 1.4.) batch correction. (c) t-SNE plot and (d) UMAP showing projections of batch-effect corrections using six 10X scRNA-seq datasets consisting of only non-mixed samples from two sites (LLU and NCI) in different batches after Scanorama (version 1.4.) batch correction. Different colors represent different datasets. All the datasets were down-sampled to 1200 cells per dataset. After the batch correction, cells from the same cell line type clustered together and mixed adequately within the same cell types. All the data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a) t-SNE plot and (b) UMAP showing batch-effect corrections using twelve 10X Genomics scRNA-seq datasets consisting of both mixed and non-mixed samples from two sites (LLU and NCI) in different batches after Scanorama (version 1.4.) batch correction. (c) t-SNE plot and (d) UMAP showing projections of batch-effect corrections using six 10X scRNA-seq datasets consisting of only non-mixed samples from two sites (LLU and NCI) in different batches after Scanorama (version 1.4.) batch correction. Different colors represent different datasets. All the datasets were down-sampled to 1200 cells per dataset. After the batch correction, cells from the same cell line type clustered together and mixed adequately within the same cell types. All the data were preprocessed using CellRanger 3.1.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques:

t-SNE plots and UMAPs showing batch-effect corrections performed by seven methods using 14 non-mixture scRNA-seq datasets across different platforms and sites. Six spiked-in mixture scRNA-seq datasets (10X_LLU_Mix10, 10X_NCI_Mix5, 10X_NCI_Mix5_F, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were removed from the 20 datasets in Scenario 1 for batch-effect correction evaluation. The fourteen non-mixture scRNA-seq datasets are from both breast cancer cells (10X_LLU_A, 10X_NCI_A, 10X_NCI_M_A, C1_FDA_HT_A, C1_LLU_A, ICELL8_SE_A, and ICELL8_PE_A) and B-lymphocytes (10X_LLU_B, 10X_NCI_B, 10X_NCI_M_B, C1_FDA_HT_B, C1_LLU_B, ICELL8_SE_B, and ICELL8_PE_B). Datasets from 10X were down-sampled to 1200 cells per dataset. *Note, for BBKNN, only UMAP was available and shown. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). All the 10X data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: t-SNE plots and UMAPs showing batch-effect corrections performed by seven methods using 14 non-mixture scRNA-seq datasets across different platforms and sites. Six spiked-in mixture scRNA-seq datasets (10X_LLU_Mix10, 10X_NCI_Mix5, 10X_NCI_Mix5_F, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were removed from the 20 datasets in Scenario 1 for batch-effect correction evaluation. The fourteen non-mixture scRNA-seq datasets are from both breast cancer cells (10X_LLU_A, 10X_NCI_A, 10X_NCI_M_A, C1_FDA_HT_A, C1_LLU_A, ICELL8_SE_A, and ICELL8_PE_A) and B-lymphocytes (10X_LLU_B, 10X_NCI_B, 10X_NCI_M_B, C1_FDA_HT_B, C1_LLU_B, ICELL8_SE_B, and ICELL8_PE_B). Datasets from 10X were down-sampled to 1200 cells per dataset. *Note, for BBKNN, only UMAP was available and shown. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). All the 10X data were preprocessed using CellRanger 3.1.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques:

Panels (a-c) show results obtained using fastMNN when the spiked-in (mixed) datasets (i.e., 10X_LLU_Mix10, 10X_NCI_Mix5, 10X_NCI_Mix5_F, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were imported into the pipeline before other non-mixed scRNA-seq datasets from the 20 scRNA-seq datasets of Scenario 1. (a) t-SNE vs. UMAP with color-coding by dataset; (b) tSNE vs. UMAP, colored by cell types (HCC1395, red; HCC1395BL, blue); and (c) A silhouette score = 0.52 showing that fastMNN correctly separated the two cell types into two clusters representing breast cancer cells and B lymphocytes. Panels (d-f) show results obtained using fastMNN when the non-mixed datasets were imported into the pipeline before the mixture datasets. (d) tSNE vs. UMAP with color-coding by datasets or (e) tSNE vs. UMAP colored by cell types; and (f) A low silhouette score of 0.22 showing that fastMNN had difficulty correctly separating the two cell types in this case. Batch-effect corrections were performed using fastMNN (SeuratWrappers v0.1.0) and silhouette width scores were calculated using the silhouette function from the R package cluster (v.2.0.8). Datasets from 10X were down-sampled to 1200 cells per dataset. The order of dataset input is shown on the top of the Figures (a, b, c or d, e, f).

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Panels (a-c) show results obtained using fastMNN when the spiked-in (mixed) datasets (i.e., 10X_LLU_Mix10, 10X_NCI_Mix5, 10X_NCI_Mix5_F, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were imported into the pipeline before other non-mixed scRNA-seq datasets from the 20 scRNA-seq datasets of Scenario 1. (a) t-SNE vs. UMAP with color-coding by dataset; (b) tSNE vs. UMAP, colored by cell types (HCC1395, red; HCC1395BL, blue); and (c) A silhouette score = 0.52 showing that fastMNN correctly separated the two cell types into two clusters representing breast cancer cells and B lymphocytes. Panels (d-f) show results obtained using fastMNN when the non-mixed datasets were imported into the pipeline before the mixture datasets. (d) tSNE vs. UMAP with color-coding by datasets or (e) tSNE vs. UMAP colored by cell types; and (f) A low silhouette score of 0.22 showing that fastMNN had difficulty correctly separating the two cell types in this case. Batch-effect corrections were performed using fastMNN (SeuratWrappers v0.1.0) and silhouette width scores were calculated using the silhouette function from the R package cluster (v.2.0.8). Datasets from 10X were down-sampled to 1200 cells per dataset. The order of dataset input is shown on the top of the Figures (a, b, c or d, e, f).

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques:

Scatter plots displaying the gene expression profile correlations between each of seven scRNA-seq datasets (10X_LLU, 10X_NCI, 10X_NCI_M, C1_FDA, C1_LLU, ICELL8_SE, and ICELL8_PE) vs. their corresponding bulk RNA-seq dataset (BK_RNA-seq) for either (a) breast cancer cells or (b) B lymphocytes. The commonly detected transcripts [(log(CPM +1) normalized] across all datasets were used (15,553 genes for breast cancer cells and 15,201 genes for B lymphocytes) to generate the scatter plots. Each dot represents each gene as a point in each scatterplot; x,y values represent the gene expression variation in a pair of compared datasets. The middle diagonal bar charts display the distribution of the most abundant or rare genes in each dataset and also provide the labels for the respective datasets. The Pearson correlation coefficient R between each of the datasets compared is shown to display the consistency of the different RNA-seq datasets.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Scatter plots displaying the gene expression profile correlations between each of seven scRNA-seq datasets (10X_LLU, 10X_NCI, 10X_NCI_M, C1_FDA, C1_LLU, ICELL8_SE, and ICELL8_PE) vs. their corresponding bulk RNA-seq dataset (BK_RNA-seq) for either (a) breast cancer cells or (b) B lymphocytes. The commonly detected transcripts [(log(CPM +1) normalized] across all datasets were used (15,553 genes for breast cancer cells and 15,201 genes for B lymphocytes) to generate the scatter plots. Each dot represents each gene as a point in each scatterplot; x,y values represent the gene expression variation in a pair of compared datasets. The middle diagonal bar charts display the distribution of the most abundant or rare genes in each dataset and also provide the labels for the respective datasets. The Pearson correlation coefficient R between each of the datasets compared is shown to display the consistency of the different RNA-seq datasets.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques: Gene Expression, RNA Sequencing

Feature plots generated across 20 scRNA-seq datasets using the top 10 DEGs specific for (a) breast cancer cells before batch-effect correction; (b) breast cancer cells after fastMNN batch-effect correction; (c) B lymphocytes before batch correction; and (d) B lymphocytes after fastMNN batch-effect correction. Datasets from 10X were down-sampled to 1200 cells per dataset. In feature plots, genes with relatively high expression in each cell are highlighted in brick red (corresponding to breast cancer cells; Sample A) or blue (corresponding to B cells; Sample B).

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Feature plots generated across 20 scRNA-seq datasets using the top 10 DEGs specific for (a) breast cancer cells before batch-effect correction; (b) breast cancer cells after fastMNN batch-effect correction; (c) B lymphocytes before batch correction; and (d) B lymphocytes after fastMNN batch-effect correction. Datasets from 10X were down-sampled to 1200 cells per dataset. In feature plots, genes with relatively high expression in each cell are highlighted in brick red (corresponding to breast cancer cells; Sample A) or blue (corresponding to B cells; Sample B).

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques: Generated, Expressing

(a) Gene detection sensitivity measured separately for each of the three classes of scRNA-seq protocol: 10X-, non-10X-based 3´ tagging, and full-length. (b) Normalization methods ranked by their clusterability as measured by Z-scores (either the median or the variance of the silhouette width across the 14 datasets). (c) Batch-correction methods ranked by their clusterability as measured by Z-score from the harmonic mean of the silhouette scores (Scenarios #1 and #4). (d) Batch-correction methods ranked by their mixability as measured by Z-score from the harmonic mean of kBET acceptance scores (Scenarios #1–#4). Z-scores are plotted as circles with their size and color shade scaled to the Z-score value from large to small, and dark blue to light blue. Note that larger Z-score values imply better performance, except for clusterability variance, where a smaller value is preferred: *Larger is better; **Smaller is better. (e) Best practice recommendations for single-cell RNA-seq analysis. #The current version of Scanorama did not correct batch effects for data from multiple platforms; however, it worked well when only 10X Genomics data were analyzed. ##Seurat v.3 was suitable for biologically similar samples, but over-corrected batch effects and misclassified cell types if large fractions of distinct cell types were present in different batches.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a) Gene detection sensitivity measured separately for each of the three classes of scRNA-seq protocol: 10X-, non-10X-based 3´ tagging, and full-length. (b) Normalization methods ranked by their clusterability as measured by Z-scores (either the median or the variance of the silhouette width across the 14 datasets). (c) Batch-correction methods ranked by their clusterability as measured by Z-score from the harmonic mean of the silhouette scores (Scenarios #1 and #4). (d) Batch-correction methods ranked by their mixability as measured by Z-score from the harmonic mean of kBET acceptance scores (Scenarios #1–#4). Z-scores are plotted as circles with their size and color shade scaled to the Z-score value from large to small, and dark blue to light blue. Note that larger Z-score values imply better performance, except for clusterability variance, where a smaller value is preferred: *Larger is better; **Smaller is better. (e) Best practice recommendations for single-cell RNA-seq analysis. #The current version of Scanorama did not correct batch effects for data from multiple platforms; however, it worked well when only 10X Genomics data were analyzed. ##Seurat v.3 was suitable for biologically similar samples, but over-corrected batch effects and misclassified cell types if large fractions of distinct cell types were present in different batches.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques: RNA Sequencing

(a, un-corrected) UMAP of 10 datasets (10X: PBMCs 68K, PBMCs 3K, CD19+ B cells, CD14+ monocytes, CD4+ helper T cells, CD56+ NK cells, CD8+ cytotoxic T cells, CD4+CD45RO+ memory T cells, CD4+CD25+ regulatory T cells; Drop-seq: PBMCs) out of 26 datasets from Hie et al.8 before batch correction by Scanorama. (b, corrected-based on dataset) UMAP of 10 different datasets shown in (a) from Hie et al. after batch correction by Scanorama, colored to identify the datasets. (c, corrected-based on platform) UMAP of 10 different datasets shown in (a) from Hie et al. colored to identify the two different platforms used (10X Genomics and Drop-seq); note poor results using Drop-seq. (d, un-corrected) UMAP of 8 datasets (breast cancer cells: C1_FDA_HT_A, C1_LLU_A, ICELL8_SE_A, and ICELL8_PE_A; and B lymphocytes: C1_FDA_HT_B, C1_LLU_B, ICELL8_SE_B, and ICELL8_PE_B) out of 20 datasets in our study analyzed using three different non-10X sequencing platforms before batch correction by Scanorama. (e, corrected-based on dataset) UMAP of 8 datasets shown in (d) after batch correction by Scanorama, colored to identify the datasets. Note lack of discrimination between different cell types. (f, corrected-based on platform) UMAP of 8 datasets shown in (d) after batch correction by Scanorama, colored to identify the platforms (C1_FDA_HT, blue; C1, purple; ICELL8, pink). The PBMC datasets were downloaded from http://scanorama.csail.mit.edu/data_light.tar.gz. Our eight datasets were preprocessed using the featureCounts pipeline and batch-effect correction was performed using Scanorama V1.4.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a, un-corrected) UMAP of 10 datasets (10X: PBMCs 68K, PBMCs 3K, CD19+ B cells, CD14+ monocytes, CD4+ helper T cells, CD56+ NK cells, CD8+ cytotoxic T cells, CD4+CD45RO+ memory T cells, CD4+CD25+ regulatory T cells; Drop-seq: PBMCs) out of 26 datasets from Hie et al.8 before batch correction by Scanorama. (b, corrected-based on dataset) UMAP of 10 different datasets shown in (a) from Hie et al. after batch correction by Scanorama, colored to identify the datasets. (c, corrected-based on platform) UMAP of 10 different datasets shown in (a) from Hie et al. colored to identify the two different platforms used (10X Genomics and Drop-seq); note poor results using Drop-seq. (d, un-corrected) UMAP of 8 datasets (breast cancer cells: C1_FDA_HT_A, C1_LLU_A, ICELL8_SE_A, and ICELL8_PE_A; and B lymphocytes: C1_FDA_HT_B, C1_LLU_B, ICELL8_SE_B, and ICELL8_PE_B) out of 20 datasets in our study analyzed using three different non-10X sequencing platforms before batch correction by Scanorama. (e, corrected-based on dataset) UMAP of 8 datasets shown in (d) after batch correction by Scanorama, colored to identify the datasets. Note lack of discrimination between different cell types. (f, corrected-based on platform) UMAP of 8 datasets shown in (d) after batch correction by Scanorama, colored to identify the platforms (C1_FDA_HT, blue; C1, purple; ICELL8, pink). The PBMC datasets were downloaded from http://scanorama.csail.mit.edu/data_light.tar.gz. Our eight datasets were preprocessed using the featureCounts pipeline and batch-effect correction was performed using Scanorama V1.4.

Article Snippet: Abbreviations and notations for : 10X_LLU , single cells were captured using a 10X Genomics Chromium controller; scRNA-seq was done at the LLU Center for Genomics using the standard 10X Genomics protocol (26+98 bp); 10X_NCI_M , 10X Genomics scRNA-seq libraries were prepared and sequenced at the NCI sequencing facility using a modified 10X sequencing protocol (26+56 bp); 10X_NCI , the same 10X Genomics scRNA-seq libraries prepared at the NCI sequencing facility were also sequenced at LLU using the standard 10X sequencing protocol (26+98 bp); C1_FDA_HT , single cells were captured using a Fluidigm C1 HT IFC and the scRNA-seq libraries were sequenced at the FDA sequencing facility (75×2 bp, PE); C1_LLU , single cells were captured using a Fluidigm C1 IFC chip and the scRNA-seq libraries were sequenced at the LLU Center for Genomics (150×2 bp, PE, ~4–4.77M reads/cell); ICELL8_PE , single cells were captured using an ICELL8 chip at TBU and scRNA-seq libraries were paired end sequenced (75×2 bp, PE) at TBU; ICELL8_SE , the same scRNA-seq libraries generated at TBU site were also sequenced, single-end (SE), at LLU (150×1 bp, SE, ~1M reads/cell).

Techniques: Sequencing

(a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate scRNA-seq data across four platforms (10X Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate scRNA-seq data across four platforms (10X Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques: Modification, Sequencing, RNA Sequencing

The violin plot shows the number of genes detected in each cell across 20 scRNA-seq datasets. The plot was generated using Seurat (version 3.1). Each dot represents a single cell. The violin shapes summarize the data distributions, which are colored in the background to signify each of the 20 different scRNA seq datasets. Each scRNA-seq dataset is plotted on the X-axis; the Y-axis shows the corresponding number of genes detected in a cell (nGene) for that dataset. The average number of genes detected in each cell was about 4000 and most of the cells had 2500–7500 genes, except for samples C1_LLU_A and C1_LLU_B. The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: The violin plot shows the number of genes detected in each cell across 20 scRNA-seq datasets. The plot was generated using Seurat (version 3.1). Each dot represents a single cell. The violin shapes summarize the data distributions, which are colored in the background to signify each of the 20 different scRNA seq datasets. Each scRNA-seq dataset is plotted on the X-axis; the Y-axis shows the corresponding number of genes detected in a cell (nGene) for that dataset. The average number of genes detected in each cell was about 4000 and most of the cells had 2500–7500 genes, except for samples C1_LLU_A and C1_LLU_B. The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques: Generated

(a–c) Evaluation of the UMI-based (10X) data with Cell Ranger, UMI-Tools, or zUMIs. (d–e) Evaluation of data from non-UMI based technologies C1 full-length transcript, C1 HT, and ICELL8 full-length transcript using FeatureCounts, Kallisto, or RSEM. (a) Bar plot showing the number of cells captured with UMI-based technology; (b) and (d) Box plot showing the number of genes detected per cell in UMI-based and non-UMI based technologies, respectively; (c) and (e) Violin plots showing the gene expression correlation and consensus genes [represented by IoU (Intersection over Union)] per cell between any two pipelines in UMI-based and non-UMI based technologies, respectively. The sample sizes (n) used to derive statistics in (b) and (d) were: (b) 10X_LLU_A, n= 3045 cells; 10X_NCI_A, n=6425 cells; 10X_NCI_M_A, n=6483 cells; 10X_LLU_B, n=1439 cells; 10X_NCI_B, n=3296 cells; 10X_NCI_M_B, n=3273 cells; (d) C1_LLU_A, n=80 cells; C1_FDA_HT_A, n=203 cells; ICELL8_SE_A, n=600 cells; ICELL8_PE_A, n=598 cells; C1_LLU_B, n=66 cells; C1_FDA_B, n=241 cells; ICELL8_SE_B, n=600 cells; ICELL8_PE_B, n=596 cells. For detailed statistics regarding minima, maxima, centre, bounds of box and whiskers and percentile related to the figure, please refer to Supplementary Table 5.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a–c) Evaluation of the UMI-based (10X) data with Cell Ranger, UMI-Tools, or zUMIs. (d–e) Evaluation of data from non-UMI based technologies C1 full-length transcript, C1 HT, and ICELL8 full-length transcript using FeatureCounts, Kallisto, or RSEM. (a) Bar plot showing the number of cells captured with UMI-based technology; (b) and (d) Box plot showing the number of genes detected per cell in UMI-based and non-UMI based technologies, respectively; (c) and (e) Violin plots showing the gene expression correlation and consensus genes [represented by IoU (Intersection over Union)] per cell between any two pipelines in UMI-based and non-UMI based technologies, respectively. The sample sizes (n) used to derive statistics in (b) and (d) were: (b) 10X_LLU_A, n= 3045 cells; 10X_NCI_A, n=6425 cells; 10X_NCI_M_A, n=6483 cells; 10X_LLU_B, n=1439 cells; 10X_NCI_B, n=3296 cells; 10X_NCI_M_B, n=3273 cells; (d) C1_LLU_A, n=80 cells; C1_FDA_HT_A, n=203 cells; ICELL8_SE_A, n=600 cells; ICELL8_PE_A, n=598 cells; C1_LLU_B, n=66 cells; C1_FDA_B, n=241 cells; ICELL8_SE_B, n=600 cells; ICELL8_PE_B, n=596 cells. For detailed statistics regarding minima, maxima, centre, bounds of box and whiskers and percentile related to the figure, please refer to Supplementary Table 5.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques: Gene Expression

(a) Batch-effect correction in Scenario #1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395; and Sample B, B-lymphocyte line HCC1395BL). Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Batch-effect correction in Scenario #2, where five scRNA-seq datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells were generated separately at the four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Batch-effect correction in Scenario #3, where five scRNA-seq datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from the B lymphocytes were generated separately at the four centers on the same four platforms; (d) Batch-effect correction in Scenario #4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, 10X_NCI_M_Mix5_F2) were generated from 5% or 10% breast cancer cells spiked into B lymphocytes, and analyzed with the 10X Genomics platform at two centers in four different batches. Each dataset is indicated by a unique color in panels (a) to (d). Idealized projection of cells for the four different scenarios is presented on the left. *Note for BBKNN, only UMAP is available and shown. Silhouette width score quantifying the clusterability for (e) Scenario #1 or (f) Scenario #4, corresponding to panels (a) and (d), respectively. (g) kBET acceptance score quantifying the mixability, calculated using the cross-platform/center scRNA-seq data acquired either from breast cancer cells only or from B-lymphocytes only for all four scenarios (a-d, also labeled as Scenarios #1–#4).

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a) Batch-effect correction in Scenario #1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395; and Sample B, B-lymphocyte line HCC1395BL). Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Batch-effect correction in Scenario #2, where five scRNA-seq datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells were generated separately at the four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Batch-effect correction in Scenario #3, where five scRNA-seq datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from the B lymphocytes were generated separately at the four centers on the same four platforms; (d) Batch-effect correction in Scenario #4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, 10X_NCI_M_Mix5_F2) were generated from 5% or 10% breast cancer cells spiked into B lymphocytes, and analyzed with the 10X Genomics platform at two centers in four different batches. Each dataset is indicated by a unique color in panels (a) to (d). Idealized projection of cells for the four different scenarios is presented on the left. *Note for BBKNN, only UMAP is available and shown. Silhouette width score quantifying the clusterability for (e) Scenario #1 or (f) Scenario #4, corresponding to panels (a) and (d), respectively. (g) kBET acceptance score quantifying the mixability, calculated using the cross-platform/center scRNA-seq data acquired either from breast cancer cells only or from B-lymphocytes only for all four scenarios (a-d, also labeled as Scenarios #1–#4).

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques: Generated, Labeling

Batch-effect corrections were performed for the following four scenarios: (a) Scenario 1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395 and Sample B, B-lymphocyte line HCC1395BL); Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Scenario 2, where five datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells (Sample A, HCC1395) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Scenario 3, where five datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from B-lymphocytes (Sample B, HCC1395BL) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); and (d) Scenario 4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, 10X_NCI_M_Mix5_F2) were generated from 5% or 10% of breast cancer cells (Sample A, HCC1395) spiked into the B-lymphocytes (Sample B, HCC1395BL) and analyzed with the 10X Genomics platform at two centers (LLU and NCI) in four different batches. *For BBKNN, only UMAPs were available and shown in (a-d). The HCC1395 breast cancer cells (Sample A) were labeled in red and the HCC1395BL B lymphocytes (Sample B) were labeled in blue. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). The top 2000 HVGs were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Batch-effect corrections were performed for the following four scenarios: (a) Scenario 1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395 and Sample B, B-lymphocyte line HCC1395BL); Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Scenario 2, where five datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells (Sample A, HCC1395) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Scenario 3, where five datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from B-lymphocytes (Sample B, HCC1395BL) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); and (d) Scenario 4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, 10X_NCI_M_Mix5_F2) were generated from 5% or 10% of breast cancer cells (Sample A, HCC1395) spiked into the B-lymphocytes (Sample B, HCC1395BL) and analyzed with the 10X Genomics platform at two centers (LLU and NCI) in four different batches. *For BBKNN, only UMAPs were available and shown in (a-d). The HCC1395 breast cancer cells (Sample A) were labeled in red and the HCC1395BL B lymphocytes (Sample B) were labeled in blue. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). The top 2000 HVGs were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques: Generated, Labeling

Boxplot of silhouette values stratified by eight normalization methods across 14 datasets, including (a) 10X_LLU, (b) 10X_NCI, (c) 10X_NCI_M, (d) C1_FDA_HT, (e) C1_LLU, (f) ICELL8_PE, and (g) ICELL8_SE in breast cancer cells (HCC1395; Sample A) and B lymphocytes (HCC1395BL; Sample B). Eight normalization methods included SCTransform, Scran Deconvolution, CPM, LogCPM, TMM, DESeq, Quantile, and Linnorm. For each dataset, reads of each cell were down-sampled to two different read depths (10K and 100K per cell) before calculating the silhouette width values. LogCPM normalization performed fairly well and was used as the default normalization for our subsequent batch-effect correction analyses. Two normalization methods developed for bulk cell RNA-seq (TMM and Quantile) had the lowest scores. The sample sizes (n) used to derive statistics were: 10X_LLU_A, n= 3560 cells, 10X_LLU_B, n=1770 cells; 10X_NCI_A, n=4284 cells, 10X_NCI_B, n=4136 cells; 10X_NCI_M_A, n=1372 cells, 10X_NCI_M_B, n=2082 cells; C1_LLU_A, n=160 cells, C1_LLU_B, n=132 cells; C1_FDA_HT_A, n=318 cells, C1_FDA_HT_B, n=374 cells; ICELL8_SE_A, n=1134 cells, ICELL8_SE_B, n=1078 cells; ICELL8_PE_A, n=980 cells, ICELL8_PE_B, n=954 cells). For detailed statistics regarding minima, maxima, centre, bounds of box and whiskers and percentile related to the figure, please refer to Supplementary Table 6.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Boxplot of silhouette values stratified by eight normalization methods across 14 datasets, including (a) 10X_LLU, (b) 10X_NCI, (c) 10X_NCI_M, (d) C1_FDA_HT, (e) C1_LLU, (f) ICELL8_PE, and (g) ICELL8_SE in breast cancer cells (HCC1395; Sample A) and B lymphocytes (HCC1395BL; Sample B). Eight normalization methods included SCTransform, Scran Deconvolution, CPM, LogCPM, TMM, DESeq, Quantile, and Linnorm. For each dataset, reads of each cell were down-sampled to two different read depths (10K and 100K per cell) before calculating the silhouette width values. LogCPM normalization performed fairly well and was used as the default normalization for our subsequent batch-effect correction analyses. Two normalization methods developed for bulk cell RNA-seq (TMM and Quantile) had the lowest scores. The sample sizes (n) used to derive statistics were: 10X_LLU_A, n= 3560 cells, 10X_LLU_B, n=1770 cells; 10X_NCI_A, n=4284 cells, 10X_NCI_B, n=4136 cells; 10X_NCI_M_A, n=1372 cells, 10X_NCI_M_B, n=2082 cells; C1_LLU_A, n=160 cells, C1_LLU_B, n=132 cells; C1_FDA_HT_A, n=318 cells, C1_FDA_HT_B, n=374 cells; ICELL8_SE_A, n=1134 cells, ICELL8_SE_B, n=1078 cells; ICELL8_PE_A, n=980 cells, ICELL8_PE_B, n=954 cells). For detailed statistics regarding minima, maxima, centre, bounds of box and whiskers and percentile related to the figure, please refer to Supplementary Table 6.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques: RNA Sequencing

Five different batches of scRNA-seq data (10X_LLU_A, 10X_LLU_B, 10X_NCI_A, 10X_NCI_B, and 10X_NCI_Mix5) generated at two sites (LLU and NCI) are shown either as t-SNE plots (panels a-d) or as UMAPs (panels e-h). (a) LogNormalized, scaled data with no regression; (b) LogNormalized, scaled data filtered with mitochondrial (Mito) gene regression >5% and UMI normalization by Seurat v3; (c) ScTransform with no regression; (d) SCTransform with mitochondrial gene regression and UMI normalization; (e) LogNormalized, scaled data with no regression; (f) scaled data with mitochondrial gene regression and UMI normalization; (g) SCTransform with no regression; and (h) SCTransform with mitochondrial gene regression and UMI normalization.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Five different batches of scRNA-seq data (10X_LLU_A, 10X_LLU_B, 10X_NCI_A, 10X_NCI_B, and 10X_NCI_Mix5) generated at two sites (LLU and NCI) are shown either as t-SNE plots (panels a-d) or as UMAPs (panels e-h). (a) LogNormalized, scaled data with no regression; (b) LogNormalized, scaled data filtered with mitochondrial (Mito) gene regression >5% and UMI normalization by Seurat v3; (c) ScTransform with no regression; (d) SCTransform with mitochondrial gene regression and UMI normalization; (e) LogNormalized, scaled data with no regression; (f) scaled data with mitochondrial gene regression and UMI normalization; (g) SCTransform with no regression; and (h) SCTransform with mitochondrial gene regression and UMI normalization.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques: Generated

Batch-effect corrections were performed for the following four scenarios: (a) Scenario 1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395 and Sample B, B-lymphocyte line HCC1395BL); Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Scenario 2, where five datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells (Sample A, HCC1395) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Scenario 3, where five datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from B-lymphocytes (Sample B, HCC1395BL) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); and (d) Scenario 4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were generated from 5% or 10% of breast cancer cells (Sample A, HCC1395), spiked into the B-lymphocytes (Sample B, HCC1395BL), and analyzed with the 10X Genomics platform at two centers (LLU and NCI) in four different batches. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). The top 2000 highly variable genes (HVGs) of these datasets were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Batch-effect corrections were performed for the following four scenarios: (a) Scenario 1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395 and Sample B, B-lymphocyte line HCC1395BL); Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Scenario 2, where five datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells (Sample A, HCC1395) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Scenario 3, where five datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from B-lymphocytes (Sample B, HCC1395BL) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); and (d) Scenario 4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were generated from 5% or 10% of breast cancer cells (Sample A, HCC1395), spiked into the B-lymphocytes (Sample B, HCC1395BL), and analyzed with the 10X Genomics platform at two centers (LLU and NCI) in four different batches. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). The top 2000 highly variable genes (HVGs) of these datasets were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques: Generated

t-SNE plots and UMAPs showing the batch-effect corrections performed by seven methods using 20 scRNA-seq datasets across different platforms. Datasets from 10X were down-sampled to 1200 cells per dataset. *Note, for BBKNN, only UMAP was available and shown. The scRNA-seq datasets are colored to identify the four different platforms: 10X 3´ scRNA-seq platform (red), C1 3´ HT scRNA-seq platform (yellow), C1 full-length scRNA-seq platform (light blue), and ICELL8 full-length scRNA-seq platform (dark blue). Batch correction methods included: Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). Scanorama failed to separate two cell types into discrete clusters when non-10X platforms were included in the analysis. The top 2000 HVGs across all datasets were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: t-SNE plots and UMAPs showing the batch-effect corrections performed by seven methods using 20 scRNA-seq datasets across different platforms. Datasets from 10X were down-sampled to 1200 cells per dataset. *Note, for BBKNN, only UMAP was available and shown. The scRNA-seq datasets are colored to identify the four different platforms: 10X 3´ scRNA-seq platform (red), C1 3´ HT scRNA-seq platform (yellow), C1 full-length scRNA-seq platform (light blue), and ICELL8 full-length scRNA-seq platform (dark blue). Batch correction methods included: Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). Scanorama failed to separate two cell types into discrete clusters when non-10X platforms were included in the analysis. The top 2000 HVGs across all datasets were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques:

(a) t-SNE plot and (b) UMAP showing batch-effect corrections using twelve 10X Genomics scRNA-seq datasets consisting of both mixed and non-mixed samples from two sites (LLU and NCI) in different batches after Scanorama (version 1.4.) batch correction. (c) t-SNE plot and (d) UMAP showing projections of batch-effect corrections using six 10X scRNA-seq datasets consisting of only non-mixed samples from two sites (LLU and NCI) in different batches after Scanorama (version 1.4.) batch correction. Different colors represent different datasets. All the datasets were down-sampled to 1200 cells per dataset. After the batch correction, cells from the same cell line type clustered together and mixed adequately within the same cell types. All the data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a) t-SNE plot and (b) UMAP showing batch-effect corrections using twelve 10X Genomics scRNA-seq datasets consisting of both mixed and non-mixed samples from two sites (LLU and NCI) in different batches after Scanorama (version 1.4.) batch correction. (c) t-SNE plot and (d) UMAP showing projections of batch-effect corrections using six 10X scRNA-seq datasets consisting of only non-mixed samples from two sites (LLU and NCI) in different batches after Scanorama (version 1.4.) batch correction. Different colors represent different datasets. All the datasets were down-sampled to 1200 cells per dataset. After the batch correction, cells from the same cell line type clustered together and mixed adequately within the same cell types. All the data were preprocessed using CellRanger 3.1.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques:

t-SNE plots and UMAPs showing batch-effect corrections performed by seven methods using 14 non-mixture scRNA-seq datasets across different platforms and sites. Six spiked-in mixture scRNA-seq datasets (10X_LLU_Mix10, 10X_NCI_Mix5, 10X_NCI_Mix5_F, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were removed from the 20 datasets in Scenario 1 for batch-effect correction evaluation. The fourteen non-mixture scRNA-seq datasets are from both breast cancer cells (10X_LLU_A, 10X_NCI_A, 10X_NCI_M_A, C1_FDA_HT_A, C1_LLU_A, ICELL8_SE_A, and ICELL8_PE_A) and B-lymphocytes (10X_LLU_B, 10X_NCI_B, 10X_NCI_M_B, C1_FDA_HT_B, C1_LLU_B, ICELL8_SE_B, and ICELL8_PE_B). Datasets from 10X were down-sampled to 1200 cells per dataset. *Note, for BBKNN, only UMAP was available and shown. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). All the 10X data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: t-SNE plots and UMAPs showing batch-effect corrections performed by seven methods using 14 non-mixture scRNA-seq datasets across different platforms and sites. Six spiked-in mixture scRNA-seq datasets (10X_LLU_Mix10, 10X_NCI_Mix5, 10X_NCI_Mix5_F, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were removed from the 20 datasets in Scenario 1 for batch-effect correction evaluation. The fourteen non-mixture scRNA-seq datasets are from both breast cancer cells (10X_LLU_A, 10X_NCI_A, 10X_NCI_M_A, C1_FDA_HT_A, C1_LLU_A, ICELL8_SE_A, and ICELL8_PE_A) and B-lymphocytes (10X_LLU_B, 10X_NCI_B, 10X_NCI_M_B, C1_FDA_HT_B, C1_LLU_B, ICELL8_SE_B, and ICELL8_PE_B). Datasets from 10X were down-sampled to 1200 cells per dataset. *Note, for BBKNN, only UMAP was available and shown. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). All the 10X data were preprocessed using CellRanger 3.1.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques:

Panels (a-c) show results obtained using fastMNN when the spiked-in (mixed) datasets (i.e., 10X_LLU_Mix10, 10X_NCI_Mix5, 10X_NCI_Mix5_F, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were imported into the pipeline before other non-mixed scRNA-seq datasets from the 20 scRNA-seq datasets of Scenario 1. (a) t-SNE vs. UMAP with color-coding by dataset; (b) tSNE vs. UMAP, colored by cell types (HCC1395, red; HCC1395BL, blue); and (c) A silhouette score = 0.52 showing that fastMNN correctly separated the two cell types into two clusters representing breast cancer cells and B lymphocytes. Panels (d-f) show results obtained using fastMNN when the non-mixed datasets were imported into the pipeline before the mixture datasets. (d) tSNE vs. UMAP with color-coding by datasets or (e) tSNE vs. UMAP colored by cell types; and (f) A low silhouette score of 0.22 showing that fastMNN had difficulty correctly separating the two cell types in this case. Batch-effect corrections were performed using fastMNN (SeuratWrappers v0.1.0) and silhouette width scores were calculated using the silhouette function from the R package cluster (v.2.0.8). Datasets from 10X were down-sampled to 1200 cells per dataset. The order of dataset input is shown on the top of the Figures (a, b, c or d, e, f).

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Panels (a-c) show results obtained using fastMNN when the spiked-in (mixed) datasets (i.e., 10X_LLU_Mix10, 10X_NCI_Mix5, 10X_NCI_Mix5_F, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were imported into the pipeline before other non-mixed scRNA-seq datasets from the 20 scRNA-seq datasets of Scenario 1. (a) t-SNE vs. UMAP with color-coding by dataset; (b) tSNE vs. UMAP, colored by cell types (HCC1395, red; HCC1395BL, blue); and (c) A silhouette score = 0.52 showing that fastMNN correctly separated the two cell types into two clusters representing breast cancer cells and B lymphocytes. Panels (d-f) show results obtained using fastMNN when the non-mixed datasets were imported into the pipeline before the mixture datasets. (d) tSNE vs. UMAP with color-coding by datasets or (e) tSNE vs. UMAP colored by cell types; and (f) A low silhouette score of 0.22 showing that fastMNN had difficulty correctly separating the two cell types in this case. Batch-effect corrections were performed using fastMNN (SeuratWrappers v0.1.0) and silhouette width scores were calculated using the silhouette function from the R package cluster (v.2.0.8). Datasets from 10X were down-sampled to 1200 cells per dataset. The order of dataset input is shown on the top of the Figures (a, b, c or d, e, f).

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques:

Scatter plots displaying the gene expression profile correlations between each of seven scRNA-seq datasets (10X_LLU, 10X_NCI, 10X_NCI_M, C1_FDA, C1_LLU, ICELL8_SE, and ICELL8_PE) vs. their corresponding bulk RNA-seq dataset (BK_RNA-seq) for either (a) breast cancer cells or (b) B lymphocytes. The commonly detected transcripts [(log(CPM +1) normalized] across all datasets were used (15,553 genes for breast cancer cells and 15,201 genes for B lymphocytes) to generate the scatter plots. Each dot represents each gene as a point in each scatterplot; x,y values represent the gene expression variation in a pair of compared datasets. The middle diagonal bar charts display the distribution of the most abundant or rare genes in each dataset and also provide the labels for the respective datasets. The Pearson correlation coefficient R between each of the datasets compared is shown to display the consistency of the different RNA-seq datasets.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Scatter plots displaying the gene expression profile correlations between each of seven scRNA-seq datasets (10X_LLU, 10X_NCI, 10X_NCI_M, C1_FDA, C1_LLU, ICELL8_SE, and ICELL8_PE) vs. their corresponding bulk RNA-seq dataset (BK_RNA-seq) for either (a) breast cancer cells or (b) B lymphocytes. The commonly detected transcripts [(log(CPM +1) normalized] across all datasets were used (15,553 genes for breast cancer cells and 15,201 genes for B lymphocytes) to generate the scatter plots. Each dot represents each gene as a point in each scatterplot; x,y values represent the gene expression variation in a pair of compared datasets. The middle diagonal bar charts display the distribution of the most abundant or rare genes in each dataset and also provide the labels for the respective datasets. The Pearson correlation coefficient R between each of the datasets compared is shown to display the consistency of the different RNA-seq datasets.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques: Gene Expression, RNA Sequencing

Feature plots generated across 20 scRNA-seq datasets using the top 10 DEGs specific for (a) breast cancer cells before batch-effect correction; (b) breast cancer cells after fastMNN batch-effect correction; (c) B lymphocytes before batch correction; and (d) B lymphocytes after fastMNN batch-effect correction. Datasets from 10X were down-sampled to 1200 cells per dataset. In feature plots, genes with relatively high expression in each cell are highlighted in brick red (corresponding to breast cancer cells; Sample A) or blue (corresponding to B cells; Sample B).

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Feature plots generated across 20 scRNA-seq datasets using the top 10 DEGs specific for (a) breast cancer cells before batch-effect correction; (b) breast cancer cells after fastMNN batch-effect correction; (c) B lymphocytes before batch correction; and (d) B lymphocytes after fastMNN batch-effect correction. Datasets from 10X were down-sampled to 1200 cells per dataset. In feature plots, genes with relatively high expression in each cell are highlighted in brick red (corresponding to breast cancer cells; Sample A) or blue (corresponding to B cells; Sample B).

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques: Generated, Expressing

(a) Gene detection sensitivity measured separately for each of the three classes of scRNA-seq protocol: 10X-, non-10X-based 3´ tagging, and full-length. (b) Normalization methods ranked by their clusterability as measured by Z-scores (either the median or the variance of the silhouette width across the 14 datasets). (c) Batch-correction methods ranked by their clusterability as measured by Z-score from the harmonic mean of the silhouette scores (Scenarios #1 and #4). (d) Batch-correction methods ranked by their mixability as measured by Z-score from the harmonic mean of kBET acceptance scores (Scenarios #1–#4). Z-scores are plotted as circles with their size and color shade scaled to the Z-score value from large to small, and dark blue to light blue. Note that larger Z-score values imply better performance, except for clusterability variance, where a smaller value is preferred: *Larger is better; **Smaller is better. (e) Best practice recommendations for single-cell RNA-seq analysis. #The current version of Scanorama did not correct batch effects for data from multiple platforms; however, it worked well when only 10X Genomics data were analyzed. ##Seurat v.3 was suitable for biologically similar samples, but over-corrected batch effects and misclassified cell types if large fractions of distinct cell types were present in different batches.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a) Gene detection sensitivity measured separately for each of the three classes of scRNA-seq protocol: 10X-, non-10X-based 3´ tagging, and full-length. (b) Normalization methods ranked by their clusterability as measured by Z-scores (either the median or the variance of the silhouette width across the 14 datasets). (c) Batch-correction methods ranked by their clusterability as measured by Z-score from the harmonic mean of the silhouette scores (Scenarios #1 and #4). (d) Batch-correction methods ranked by their mixability as measured by Z-score from the harmonic mean of kBET acceptance scores (Scenarios #1–#4). Z-scores are plotted as circles with their size and color shade scaled to the Z-score value from large to small, and dark blue to light blue. Note that larger Z-score values imply better performance, except for clusterability variance, where a smaller value is preferred: *Larger is better; **Smaller is better. (e) Best practice recommendations for single-cell RNA-seq analysis. #The current version of Scanorama did not correct batch effects for data from multiple platforms; however, it worked well when only 10X Genomics data were analyzed. ##Seurat v.3 was suitable for biologically similar samples, but over-corrected batch effects and misclassified cell types if large fractions of distinct cell types were present in different batches.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques: RNA Sequencing

(a, un-corrected) UMAP of 10 datasets (10X: PBMCs 68K, PBMCs 3K, CD19+ B cells, CD14+ monocytes, CD4+ helper T cells, CD56+ NK cells, CD8+ cytotoxic T cells, CD4+CD45RO+ memory T cells, CD4+CD25+ regulatory T cells; Drop-seq: PBMCs) out of 26 datasets from Hie et al.8 before batch correction by Scanorama. (b, corrected-based on dataset) UMAP of 10 different datasets shown in (a) from Hie et al. after batch correction by Scanorama, colored to identify the datasets. (c, corrected-based on platform) UMAP of 10 different datasets shown in (a) from Hie et al. colored to identify the two different platforms used (10X Genomics and Drop-seq); note poor results using Drop-seq. (d, un-corrected) UMAP of 8 datasets (breast cancer cells: C1_FDA_HT_A, C1_LLU_A, ICELL8_SE_A, and ICELL8_PE_A; and B lymphocytes: C1_FDA_HT_B, C1_LLU_B, ICELL8_SE_B, and ICELL8_PE_B) out of 20 datasets in our study analyzed using three different non-10X sequencing platforms before batch correction by Scanorama. (e, corrected-based on dataset) UMAP of 8 datasets shown in (d) after batch correction by Scanorama, colored to identify the datasets. Note lack of discrimination between different cell types. (f, corrected-based on platform) UMAP of 8 datasets shown in (d) after batch correction by Scanorama, colored to identify the platforms (C1_FDA_HT, blue; C1, purple; ICELL8, pink). The PBMC datasets were downloaded from http://scanorama.csail.mit.edu/data_light.tar.gz. Our eight datasets were preprocessed using the featureCounts pipeline and batch-effect correction was performed using Scanorama V1.4.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a, un-corrected) UMAP of 10 datasets (10X: PBMCs 68K, PBMCs 3K, CD19+ B cells, CD14+ monocytes, CD4+ helper T cells, CD56+ NK cells, CD8+ cytotoxic T cells, CD4+CD45RO+ memory T cells, CD4+CD25+ regulatory T cells; Drop-seq: PBMCs) out of 26 datasets from Hie et al.8 before batch correction by Scanorama. (b, corrected-based on dataset) UMAP of 10 different datasets shown in (a) from Hie et al. after batch correction by Scanorama, colored to identify the datasets. (c, corrected-based on platform) UMAP of 10 different datasets shown in (a) from Hie et al. colored to identify the two different platforms used (10X Genomics and Drop-seq); note poor results using Drop-seq. (d, un-corrected) UMAP of 8 datasets (breast cancer cells: C1_FDA_HT_A, C1_LLU_A, ICELL8_SE_A, and ICELL8_PE_A; and B lymphocytes: C1_FDA_HT_B, C1_LLU_B, ICELL8_SE_B, and ICELL8_PE_B) out of 20 datasets in our study analyzed using three different non-10X sequencing platforms before batch correction by Scanorama. (e, corrected-based on dataset) UMAP of 8 datasets shown in (d) after batch correction by Scanorama, colored to identify the datasets. Note lack of discrimination between different cell types. (f, corrected-based on platform) UMAP of 8 datasets shown in (d) after batch correction by Scanorama, colored to identify the platforms (C1_FDA_HT, blue; C1, purple; ICELL8, pink). The PBMC datasets were downloaded from http://scanorama.csail.mit.edu/data_light.tar.gz. Our eight datasets were preprocessed using the featureCounts pipeline and batch-effect correction was performed using Scanorama V1.4.

Article Snippet: Nevertheless, many overlapping genes were detected (96.6–97.3%) with a high correlation (R=0.997–0.998) between the standard and modified sequencing protocol for the 10X Genomics scRNA-seq ( Supplementary Fig. 1 ).

Techniques: Sequencing

(a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate scRNA-seq data across four platforms (10X Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a) Schematic overview of the study design (see detailed descriptions and notations in the Methods). Two reference cell lines (Sample A, HCC1395; and Sample B, HCC1395BL) were used to generate scRNA-seq data across four platforms (10X Genomics, Fluidigm C1, Fluidigm C1 HT, and Takara Bio ICELL8), four testing sites (LLU, NCI, FDA, and TBU). At the LLU and NCI sites (10X), mixed single-cell captures and library constructions were also prepared with either 10% or 5% cancer cells spiked into the B lymphocytes. At the NCI site, single-cell captures and library constructions were also performed with methanol-fixed cell mixtures (5% cancer cells spiked into B lymphocytes, Fixed 1 & 2). One set of 10X scRNA libraries from NCI was also sequenced using a shorter modified sequencing method. Bulk cell RNA-seq was also obtained from these cell lines, each in triplicate. See Methods for details about study design. (b) For both the breast cancer cell line (Sample A) and the B lymphocyte line (Sample B) across 14 pair-wise datasets, percentage of reads mapped to the exonic region (blue), non-exonic region (orange), or not mapped to the human genome (gray). For unique molecular identifier (UMI) methods (10X), dark blue indicates the exonic reads with UMIs. (c) Median number of genes detected per cell at different sequencing read depths.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques: Modification, Sequencing, RNA Sequencing

The violin plot shows the number of genes detected in each cell across 20 scRNA-seq datasets. The plot was generated using Seurat (version 3.1). Each dot represents a single cell. The violin shapes summarize the data distributions, which are colored in the background to signify each of the 20 different scRNA seq datasets. Each scRNA-seq dataset is plotted on the X-axis; the Y-axis shows the corresponding number of genes detected in a cell (nGene) for that dataset. The average number of genes detected in each cell was about 4000 and most of the cells had 2500–7500 genes, except for samples C1_LLU_A and C1_LLU_B. The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: The violin plot shows the number of genes detected in each cell across 20 scRNA-seq datasets. The plot was generated using Seurat (version 3.1). Each dot represents a single cell. The violin shapes summarize the data distributions, which are colored in the background to signify each of the 20 different scRNA seq datasets. Each scRNA-seq dataset is plotted on the X-axis; the Y-axis shows the corresponding number of genes detected in a cell (nGene) for that dataset. The average number of genes detected in each cell was about 4000 and most of the cells had 2500–7500 genes, except for samples C1_LLU_A and C1_LLU_B. The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques: Generated

(a–c) Evaluation of the UMI-based (10X) data with Cell Ranger, UMI-Tools, or zUMIs. (d–e) Evaluation of data from non-UMI based technologies C1 full-length transcript, C1 HT, and ICELL8 full-length transcript using FeatureCounts, Kallisto, or RSEM. (a) Bar plot showing the number of cells captured with UMI-based technology; (b) and (d) Box plot showing the number of genes detected per cell in UMI-based and non-UMI based technologies, respectively; (c) and (e) Violin plots showing the gene expression correlation and consensus genes [represented by IoU (Intersection over Union)] per cell between any two pipelines in UMI-based and non-UMI based technologies, respectively. The sample sizes (n) used to derive statistics in (b) and (d) were: (b) 10X_LLU_A, n= 3045 cells; 10X_NCI_A, n=6425 cells; 10X_NCI_M_A, n=6483 cells; 10X_LLU_B, n=1439 cells; 10X_NCI_B, n=3296 cells; 10X_NCI_M_B, n=3273 cells; (d) C1_LLU_A, n=80 cells; C1_FDA_HT_A, n=203 cells; ICELL8_SE_A, n=600 cells; ICELL8_PE_A, n=598 cells; C1_LLU_B, n=66 cells; C1_FDA_B, n=241 cells; ICELL8_SE_B, n=600 cells; ICELL8_PE_B, n=596 cells. For detailed statistics regarding minima, maxima, centre, bounds of box and whiskers and percentile related to the figure, please refer to Supplementary Table 5.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a–c) Evaluation of the UMI-based (10X) data with Cell Ranger, UMI-Tools, or zUMIs. (d–e) Evaluation of data from non-UMI based technologies C1 full-length transcript, C1 HT, and ICELL8 full-length transcript using FeatureCounts, Kallisto, or RSEM. (a) Bar plot showing the number of cells captured with UMI-based technology; (b) and (d) Box plot showing the number of genes detected per cell in UMI-based and non-UMI based technologies, respectively; (c) and (e) Violin plots showing the gene expression correlation and consensus genes [represented by IoU (Intersection over Union)] per cell between any two pipelines in UMI-based and non-UMI based technologies, respectively. The sample sizes (n) used to derive statistics in (b) and (d) were: (b) 10X_LLU_A, n= 3045 cells; 10X_NCI_A, n=6425 cells; 10X_NCI_M_A, n=6483 cells; 10X_LLU_B, n=1439 cells; 10X_NCI_B, n=3296 cells; 10X_NCI_M_B, n=3273 cells; (d) C1_LLU_A, n=80 cells; C1_FDA_HT_A, n=203 cells; ICELL8_SE_A, n=600 cells; ICELL8_PE_A, n=598 cells; C1_LLU_B, n=66 cells; C1_FDA_B, n=241 cells; ICELL8_SE_B, n=600 cells; ICELL8_PE_B, n=596 cells. For detailed statistics regarding minima, maxima, centre, bounds of box and whiskers and percentile related to the figure, please refer to Supplementary Table 5.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques: Gene Expression

(a) Batch-effect correction in Scenario #1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395; and Sample B, B-lymphocyte line HCC1395BL). Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Batch-effect correction in Scenario #2, where five scRNA-seq datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells were generated separately at the four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Batch-effect correction in Scenario #3, where five scRNA-seq datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from the B lymphocytes were generated separately at the four centers on the same four platforms; (d) Batch-effect correction in Scenario #4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, 10X_NCI_M_Mix5_F2) were generated from 5% or 10% breast cancer cells spiked into B lymphocytes, and analyzed with the 10X Genomics platform at two centers in four different batches. Each dataset is indicated by a unique color in panels (a) to (d). Idealized projection of cells for the four different scenarios is presented on the left. *Note for BBKNN, only UMAP is available and shown. Silhouette width score quantifying the clusterability for (e) Scenario #1 or (f) Scenario #4, corresponding to panels (a) and (d), respectively. (g) kBET acceptance score quantifying the mixability, calculated using the cross-platform/center scRNA-seq data acquired either from breast cancer cells only or from B-lymphocytes only for all four scenarios (a-d, also labeled as Scenarios #1–#4).

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a) Batch-effect correction in Scenario #1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395; and Sample B, B-lymphocyte line HCC1395BL). Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Batch-effect correction in Scenario #2, where five scRNA-seq datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells were generated separately at the four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Batch-effect correction in Scenario #3, where five scRNA-seq datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from the B lymphocytes were generated separately at the four centers on the same four platforms; (d) Batch-effect correction in Scenario #4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, 10X_NCI_M_Mix5_F2) were generated from 5% or 10% breast cancer cells spiked into B lymphocytes, and analyzed with the 10X Genomics platform at two centers in four different batches. Each dataset is indicated by a unique color in panels (a) to (d). Idealized projection of cells for the four different scenarios is presented on the left. *Note for BBKNN, only UMAP is available and shown. Silhouette width score quantifying the clusterability for (e) Scenario #1 or (f) Scenario #4, corresponding to panels (a) and (d), respectively. (g) kBET acceptance score quantifying the mixability, calculated using the cross-platform/center scRNA-seq data acquired either from breast cancer cells only or from B-lymphocytes only for all four scenarios (a-d, also labeled as Scenarios #1–#4).

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques: Generated, Labeling

Batch-effect corrections were performed for the following four scenarios: (a) Scenario 1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395 and Sample B, B-lymphocyte line HCC1395BL); Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Scenario 2, where five datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells (Sample A, HCC1395) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Scenario 3, where five datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from B-lymphocytes (Sample B, HCC1395BL) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); and (d) Scenario 4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, 10X_NCI_M_Mix5_F2) were generated from 5% or 10% of breast cancer cells (Sample A, HCC1395) spiked into the B-lymphocytes (Sample B, HCC1395BL) and analyzed with the 10X Genomics platform at two centers (LLU and NCI) in four different batches. *For BBKNN, only UMAPs were available and shown in (a-d). The HCC1395 breast cancer cells (Sample A) were labeled in red and the HCC1395BL B lymphocytes (Sample B) were labeled in blue. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). The top 2000 HVGs were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Batch-effect corrections were performed for the following four scenarios: (a) Scenario 1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395 and Sample B, B-lymphocyte line HCC1395BL); Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Scenario 2, where five datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells (Sample A, HCC1395) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Scenario 3, where five datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from B-lymphocytes (Sample B, HCC1395BL) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); and (d) Scenario 4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, 10X_NCI_M_Mix5_F2) were generated from 5% or 10% of breast cancer cells (Sample A, HCC1395) spiked into the B-lymphocytes (Sample B, HCC1395BL) and analyzed with the 10X Genomics platform at two centers (LLU and NCI) in four different batches. *For BBKNN, only UMAPs were available and shown in (a-d). The HCC1395 breast cancer cells (Sample A) were labeled in red and the HCC1395BL B lymphocytes (Sample B) were labeled in blue. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). The top 2000 HVGs were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques: Generated, Labeling

Boxplot of silhouette values stratified by eight normalization methods across 14 datasets, including (a) 10X_LLU, (b) 10X_NCI, (c) 10X_NCI_M, (d) C1_FDA_HT, (e) C1_LLU, (f) ICELL8_PE, and (g) ICELL8_SE in breast cancer cells (HCC1395; Sample A) and B lymphocytes (HCC1395BL; Sample B). Eight normalization methods included SCTransform, Scran Deconvolution, CPM, LogCPM, TMM, DESeq, Quantile, and Linnorm. For each dataset, reads of each cell were down-sampled to two different read depths (10K and 100K per cell) before calculating the silhouette width values. LogCPM normalization performed fairly well and was used as the default normalization for our subsequent batch-effect correction analyses. Two normalization methods developed for bulk cell RNA-seq (TMM and Quantile) had the lowest scores. The sample sizes (n) used to derive statistics were: 10X_LLU_A, n= 3560 cells, 10X_LLU_B, n=1770 cells; 10X_NCI_A, n=4284 cells, 10X_NCI_B, n=4136 cells; 10X_NCI_M_A, n=1372 cells, 10X_NCI_M_B, n=2082 cells; C1_LLU_A, n=160 cells, C1_LLU_B, n=132 cells; C1_FDA_HT_A, n=318 cells, C1_FDA_HT_B, n=374 cells; ICELL8_SE_A, n=1134 cells, ICELL8_SE_B, n=1078 cells; ICELL8_PE_A, n=980 cells, ICELL8_PE_B, n=954 cells). For detailed statistics regarding minima, maxima, centre, bounds of box and whiskers and percentile related to the figure, please refer to Supplementary Table 6.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Boxplot of silhouette values stratified by eight normalization methods across 14 datasets, including (a) 10X_LLU, (b) 10X_NCI, (c) 10X_NCI_M, (d) C1_FDA_HT, (e) C1_LLU, (f) ICELL8_PE, and (g) ICELL8_SE in breast cancer cells (HCC1395; Sample A) and B lymphocytes (HCC1395BL; Sample B). Eight normalization methods included SCTransform, Scran Deconvolution, CPM, LogCPM, TMM, DESeq, Quantile, and Linnorm. For each dataset, reads of each cell were down-sampled to two different read depths (10K and 100K per cell) before calculating the silhouette width values. LogCPM normalization performed fairly well and was used as the default normalization for our subsequent batch-effect correction analyses. Two normalization methods developed for bulk cell RNA-seq (TMM and Quantile) had the lowest scores. The sample sizes (n) used to derive statistics were: 10X_LLU_A, n= 3560 cells, 10X_LLU_B, n=1770 cells; 10X_NCI_A, n=4284 cells, 10X_NCI_B, n=4136 cells; 10X_NCI_M_A, n=1372 cells, 10X_NCI_M_B, n=2082 cells; C1_LLU_A, n=160 cells, C1_LLU_B, n=132 cells; C1_FDA_HT_A, n=318 cells, C1_FDA_HT_B, n=374 cells; ICELL8_SE_A, n=1134 cells, ICELL8_SE_B, n=1078 cells; ICELL8_PE_A, n=980 cells, ICELL8_PE_B, n=954 cells). For detailed statistics regarding minima, maxima, centre, bounds of box and whiskers and percentile related to the figure, please refer to Supplementary Table 6.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques: RNA Sequencing

Five different batches of scRNA-seq data (10X_LLU_A, 10X_LLU_B, 10X_NCI_A, 10X_NCI_B, and 10X_NCI_Mix5) generated at two sites (LLU and NCI) are shown either as t-SNE plots (panels a-d) or as UMAPs (panels e-h). (a) LogNormalized, scaled data with no regression; (b) LogNormalized, scaled data filtered with mitochondrial (Mito) gene regression >5% and UMI normalization by Seurat v3; (c) ScTransform with no regression; (d) SCTransform with mitochondrial gene regression and UMI normalization; (e) LogNormalized, scaled data with no regression; (f) scaled data with mitochondrial gene regression and UMI normalization; (g) SCTransform with no regression; and (h) SCTransform with mitochondrial gene regression and UMI normalization.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Five different batches of scRNA-seq data (10X_LLU_A, 10X_LLU_B, 10X_NCI_A, 10X_NCI_B, and 10X_NCI_Mix5) generated at two sites (LLU and NCI) are shown either as t-SNE plots (panels a-d) or as UMAPs (panels e-h). (a) LogNormalized, scaled data with no regression; (b) LogNormalized, scaled data filtered with mitochondrial (Mito) gene regression >5% and UMI normalization by Seurat v3; (c) ScTransform with no regression; (d) SCTransform with mitochondrial gene regression and UMI normalization; (e) LogNormalized, scaled data with no regression; (f) scaled data with mitochondrial gene regression and UMI normalization; (g) SCTransform with no regression; and (h) SCTransform with mitochondrial gene regression and UMI normalization.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques: Generated

Batch-effect corrections were performed for the following four scenarios: (a) Scenario 1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395 and Sample B, B-lymphocyte line HCC1395BL); Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Scenario 2, where five datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells (Sample A, HCC1395) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Scenario 3, where five datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from B-lymphocytes (Sample B, HCC1395BL) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); and (d) Scenario 4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were generated from 5% or 10% of breast cancer cells (Sample A, HCC1395), spiked into the B-lymphocytes (Sample B, HCC1395BL), and analyzed with the 10X Genomics platform at two centers (LLU and NCI) in four different batches. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). The top 2000 highly variable genes (HVGs) of these datasets were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Batch-effect corrections were performed for the following four scenarios: (a) Scenario 1, where all 20 scRNA-seq datasets were combined, including mixed and non-mixed, with large proportions of two dissimilar types of cells (Sample A, breast cancer cell line HCC1395 and Sample B, B-lymphocyte line HCC1395BL); Datasets from 10X were down-sampled to 1200 cells per dataset. (b) Scenario 2, where five datasets (10X_LLU_A, 10X_NCI_A, C1_FDA_HT_A, C1_LLU_A, and ICELL8_SE_A) from the breast cancer cells (Sample A, HCC1395) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); (c) Scenario 3, where five datasets (10X_LLU_B, 10X_NCI_B, C1_FDA_HT_B, C1_LLU_B, and ICELL8_SE_B) from B-lymphocytes (Sample B, HCC1395BL) were generated separately at four centers (LLU, NCI, FDA, and TBU) on four platforms (10X, Fluidigm C1, Fluidigm C1_HT, and TBU ICELL8); and (d) Scenario 4, where four datasets (10X_LLU_Mix10, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were generated from 5% or 10% of breast cancer cells (Sample A, HCC1395), spiked into the B-lymphocytes (Sample B, HCC1395BL), and analyzed with the 10X Genomics platform at two centers (LLU and NCI) in four different batches. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). The top 2000 highly variable genes (HVGs) of these datasets were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques: Generated

t-SNE plots and UMAPs showing the batch-effect corrections performed by seven methods using 20 scRNA-seq datasets across different platforms. Datasets from 10X were down-sampled to 1200 cells per dataset. *Note, for BBKNN, only UMAP was available and shown. The scRNA-seq datasets are colored to identify the four different platforms: 10X 3´ scRNA-seq platform (red), C1 3´ HT scRNA-seq platform (yellow), C1 full-length scRNA-seq platform (light blue), and ICELL8 full-length scRNA-seq platform (dark blue). Batch correction methods included: Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). Scanorama failed to separate two cell types into discrete clusters when non-10X platforms were included in the analysis. The top 2000 HVGs across all datasets were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: t-SNE plots and UMAPs showing the batch-effect corrections performed by seven methods using 20 scRNA-seq datasets across different platforms. Datasets from 10X were down-sampled to 1200 cells per dataset. *Note, for BBKNN, only UMAP was available and shown. The scRNA-seq datasets are colored to identify the four different platforms: 10X 3´ scRNA-seq platform (red), C1 3´ HT scRNA-seq platform (yellow), C1 full-length scRNA-seq platform (light blue), and ICELL8 full-length scRNA-seq platform (dark blue). Batch correction methods included: Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). Scanorama failed to separate two cell types into discrete clusters when non-10X platforms were included in the analysis. The top 2000 HVGs across all datasets were used as the gene set for batch correction. All the 10X data were preprocessed using CellRanger 3.1.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques:

(a) t-SNE plot and (b) UMAP showing batch-effect corrections using twelve 10X Genomics scRNA-seq datasets consisting of both mixed and non-mixed samples from two sites (LLU and NCI) in different batches after Scanorama (version 1.4.) batch correction. (c) t-SNE plot and (d) UMAP showing projections of batch-effect corrections using six 10X scRNA-seq datasets consisting of only non-mixed samples from two sites (LLU and NCI) in different batches after Scanorama (version 1.4.) batch correction. Different colors represent different datasets. All the datasets were down-sampled to 1200 cells per dataset. After the batch correction, cells from the same cell line type clustered together and mixed adequately within the same cell types. All the data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a) t-SNE plot and (b) UMAP showing batch-effect corrections using twelve 10X Genomics scRNA-seq datasets consisting of both mixed and non-mixed samples from two sites (LLU and NCI) in different batches after Scanorama (version 1.4.) batch correction. (c) t-SNE plot and (d) UMAP showing projections of batch-effect corrections using six 10X scRNA-seq datasets consisting of only non-mixed samples from two sites (LLU and NCI) in different batches after Scanorama (version 1.4.) batch correction. Different colors represent different datasets. All the datasets were down-sampled to 1200 cells per dataset. After the batch correction, cells from the same cell line type clustered together and mixed adequately within the same cell types. All the data were preprocessed using CellRanger 3.1.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques:

t-SNE plots and UMAPs showing batch-effect corrections performed by seven methods using 14 non-mixture scRNA-seq datasets across different platforms and sites. Six spiked-in mixture scRNA-seq datasets (10X_LLU_Mix10, 10X_NCI_Mix5, 10X_NCI_Mix5_F, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were removed from the 20 datasets in Scenario 1 for batch-effect correction evaluation. The fourteen non-mixture scRNA-seq datasets are from both breast cancer cells (10X_LLU_A, 10X_NCI_A, 10X_NCI_M_A, C1_FDA_HT_A, C1_LLU_A, ICELL8_SE_A, and ICELL8_PE_A) and B-lymphocytes (10X_LLU_B, 10X_NCI_B, 10X_NCI_M_B, C1_FDA_HT_B, C1_LLU_B, ICELL8_SE_B, and ICELL8_PE_B). Datasets from 10X were down-sampled to 1200 cells per dataset. *Note, for BBKNN, only UMAP was available and shown. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). All the 10X data were preprocessed using CellRanger 3.1.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: t-SNE plots and UMAPs showing batch-effect corrections performed by seven methods using 14 non-mixture scRNA-seq datasets across different platforms and sites. Six spiked-in mixture scRNA-seq datasets (10X_LLU_Mix10, 10X_NCI_Mix5, 10X_NCI_Mix5_F, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were removed from the 20 datasets in Scenario 1 for batch-effect correction evaluation. The fourteen non-mixture scRNA-seq datasets are from both breast cancer cells (10X_LLU_A, 10X_NCI_A, 10X_NCI_M_A, C1_FDA_HT_A, C1_LLU_A, ICELL8_SE_A, and ICELL8_PE_A) and B-lymphocytes (10X_LLU_B, 10X_NCI_B, 10X_NCI_M_B, C1_FDA_HT_B, C1_LLU_B, ICELL8_SE_B, and ICELL8_PE_B). Datasets from 10X were down-sampled to 1200 cells per dataset. *Note, for BBKNN, only UMAP was available and shown. Batch correction methods included Seurat v3.1, fastMNN (SeuratWrappers v0.1.0), Scanorama V1.4, BBKNN V1.3.5, Harmony V0.99.9, limma V3.40.4, and Combat (sva V3.32.1). All the 10X data were preprocessed using CellRanger 3.1.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques:

Panels (a-c) show results obtained using fastMNN when the spiked-in (mixed) datasets (i.e., 10X_LLU_Mix10, 10X_NCI_Mix5, 10X_NCI_Mix5_F, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were imported into the pipeline before other non-mixed scRNA-seq datasets from the 20 scRNA-seq datasets of Scenario 1. (a) t-SNE vs. UMAP with color-coding by dataset; (b) tSNE vs. UMAP, colored by cell types (HCC1395, red; HCC1395BL, blue); and (c) A silhouette score = 0.52 showing that fastMNN correctly separated the two cell types into two clusters representing breast cancer cells and B lymphocytes. Panels (d-f) show results obtained using fastMNN when the non-mixed datasets were imported into the pipeline before the mixture datasets. (d) tSNE vs. UMAP with color-coding by datasets or (e) tSNE vs. UMAP colored by cell types; and (f) A low silhouette score of 0.22 showing that fastMNN had difficulty correctly separating the two cell types in this case. Batch-effect corrections were performed using fastMNN (SeuratWrappers v0.1.0) and silhouette width scores were calculated using the silhouette function from the R package cluster (v.2.0.8). Datasets from 10X were down-sampled to 1200 cells per dataset. The order of dataset input is shown on the top of the Figures (a, b, c or d, e, f).

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Panels (a-c) show results obtained using fastMNN when the spiked-in (mixed) datasets (i.e., 10X_LLU_Mix10, 10X_NCI_Mix5, 10X_NCI_Mix5_F, 10X_NCI_M_Mix5, 10X_NCI_M_Mix5_F, and 10X_NCI_M_Mix5_F2) were imported into the pipeline before other non-mixed scRNA-seq datasets from the 20 scRNA-seq datasets of Scenario 1. (a) t-SNE vs. UMAP with color-coding by dataset; (b) tSNE vs. UMAP, colored by cell types (HCC1395, red; HCC1395BL, blue); and (c) A silhouette score = 0.52 showing that fastMNN correctly separated the two cell types into two clusters representing breast cancer cells and B lymphocytes. Panels (d-f) show results obtained using fastMNN when the non-mixed datasets were imported into the pipeline before the mixture datasets. (d) tSNE vs. UMAP with color-coding by datasets or (e) tSNE vs. UMAP colored by cell types; and (f) A low silhouette score of 0.22 showing that fastMNN had difficulty correctly separating the two cell types in this case. Batch-effect corrections were performed using fastMNN (SeuratWrappers v0.1.0) and silhouette width scores were calculated using the silhouette function from the R package cluster (v.2.0.8). Datasets from 10X were down-sampled to 1200 cells per dataset. The order of dataset input is shown on the top of the Figures (a, b, c or d, e, f).

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques:

Scatter plots displaying the gene expression profile correlations between each of seven scRNA-seq datasets (10X_LLU, 10X_NCI, 10X_NCI_M, C1_FDA, C1_LLU, ICELL8_SE, and ICELL8_PE) vs. their corresponding bulk RNA-seq dataset (BK_RNA-seq) for either (a) breast cancer cells or (b) B lymphocytes. The commonly detected transcripts [(log(CPM +1) normalized] across all datasets were used (15,553 genes for breast cancer cells and 15,201 genes for B lymphocytes) to generate the scatter plots. Each dot represents each gene as a point in each scatterplot; x,y values represent the gene expression variation in a pair of compared datasets. The middle diagonal bar charts display the distribution of the most abundant or rare genes in each dataset and also provide the labels for the respective datasets. The Pearson correlation coefficient R between each of the datasets compared is shown to display the consistency of the different RNA-seq datasets.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Scatter plots displaying the gene expression profile correlations between each of seven scRNA-seq datasets (10X_LLU, 10X_NCI, 10X_NCI_M, C1_FDA, C1_LLU, ICELL8_SE, and ICELL8_PE) vs. their corresponding bulk RNA-seq dataset (BK_RNA-seq) for either (a) breast cancer cells or (b) B lymphocytes. The commonly detected transcripts [(log(CPM +1) normalized] across all datasets were used (15,553 genes for breast cancer cells and 15,201 genes for B lymphocytes) to generate the scatter plots. Each dot represents each gene as a point in each scatterplot; x,y values represent the gene expression variation in a pair of compared datasets. The middle diagonal bar charts display the distribution of the most abundant or rare genes in each dataset and also provide the labels for the respective datasets. The Pearson correlation coefficient R between each of the datasets compared is shown to display the consistency of the different RNA-seq datasets.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques: Gene Expression, RNA Sequencing

Feature plots generated across 20 scRNA-seq datasets using the top 10 DEGs specific for (a) breast cancer cells before batch-effect correction; (b) breast cancer cells after fastMNN batch-effect correction; (c) B lymphocytes before batch correction; and (d) B lymphocytes after fastMNN batch-effect correction. Datasets from 10X were down-sampled to 1200 cells per dataset. In feature plots, genes with relatively high expression in each cell are highlighted in brick red (corresponding to breast cancer cells; Sample A) or blue (corresponding to B cells; Sample B).

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: Feature plots generated across 20 scRNA-seq datasets using the top 10 DEGs specific for (a) breast cancer cells before batch-effect correction; (b) breast cancer cells after fastMNN batch-effect correction; (c) B lymphocytes before batch correction; and (d) B lymphocytes after fastMNN batch-effect correction. Datasets from 10X were down-sampled to 1200 cells per dataset. In feature plots, genes with relatively high expression in each cell are highlighted in brick red (corresponding to breast cancer cells; Sample A) or blue (corresponding to B cells; Sample B).

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques: Generated, Expressing

(a) Gene detection sensitivity measured separately for each of the three classes of scRNA-seq protocol: 10X-, non-10X-based 3´ tagging, and full-length. (b) Normalization methods ranked by their clusterability as measured by Z-scores (either the median or the variance of the silhouette width across the 14 datasets). (c) Batch-correction methods ranked by their clusterability as measured by Z-score from the harmonic mean of the silhouette scores (Scenarios #1 and #4). (d) Batch-correction methods ranked by their mixability as measured by Z-score from the harmonic mean of kBET acceptance scores (Scenarios #1–#4). Z-scores are plotted as circles with their size and color shade scaled to the Z-score value from large to small, and dark blue to light blue. Note that larger Z-score values imply better performance, except for clusterability variance, where a smaller value is preferred: *Larger is better; **Smaller is better. (e) Best practice recommendations for single-cell RNA-seq analysis. #The current version of Scanorama did not correct batch effects for data from multiple platforms; however, it worked well when only 10X Genomics data were analyzed. ##Seurat v.3 was suitable for biologically similar samples, but over-corrected batch effects and misclassified cell types if large fractions of distinct cell types were present in different batches.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a) Gene detection sensitivity measured separately for each of the three classes of scRNA-seq protocol: 10X-, non-10X-based 3´ tagging, and full-length. (b) Normalization methods ranked by their clusterability as measured by Z-scores (either the median or the variance of the silhouette width across the 14 datasets). (c) Batch-correction methods ranked by their clusterability as measured by Z-score from the harmonic mean of the silhouette scores (Scenarios #1 and #4). (d) Batch-correction methods ranked by their mixability as measured by Z-score from the harmonic mean of kBET acceptance scores (Scenarios #1–#4). Z-scores are plotted as circles with their size and color shade scaled to the Z-score value from large to small, and dark blue to light blue. Note that larger Z-score values imply better performance, except for clusterability variance, where a smaller value is preferred: *Larger is better; **Smaller is better. (e) Best practice recommendations for single-cell RNA-seq analysis. #The current version of Scanorama did not correct batch effects for data from multiple platforms; however, it worked well when only 10X Genomics data were analyzed. ##Seurat v.3 was suitable for biologically similar samples, but over-corrected batch effects and misclassified cell types if large fractions of distinct cell types were present in different batches.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques: RNA Sequencing

(a, un-corrected) UMAP of 10 datasets (10X: PBMCs 68K, PBMCs 3K, CD19+ B cells, CD14+ monocytes, CD4+ helper T cells, CD56+ NK cells, CD8+ cytotoxic T cells, CD4+CD45RO+ memory T cells, CD4+CD25+ regulatory T cells; Drop-seq: PBMCs) out of 26 datasets from Hie et al.8 before batch correction by Scanorama. (b, corrected-based on dataset) UMAP of 10 different datasets shown in (a) from Hie et al. after batch correction by Scanorama, colored to identify the datasets. (c, corrected-based on platform) UMAP of 10 different datasets shown in (a) from Hie et al. colored to identify the two different platforms used (10X Genomics and Drop-seq); note poor results using Drop-seq. (d, un-corrected) UMAP of 8 datasets (breast cancer cells: C1_FDA_HT_A, C1_LLU_A, ICELL8_SE_A, and ICELL8_PE_A; and B lymphocytes: C1_FDA_HT_B, C1_LLU_B, ICELL8_SE_B, and ICELL8_PE_B) out of 20 datasets in our study analyzed using three different non-10X sequencing platforms before batch correction by Scanorama. (e, corrected-based on dataset) UMAP of 8 datasets shown in (d) after batch correction by Scanorama, colored to identify the datasets. Note lack of discrimination between different cell types. (f, corrected-based on platform) UMAP of 8 datasets shown in (d) after batch correction by Scanorama, colored to identify the platforms (C1_FDA_HT, blue; C1, purple; ICELL8, pink). The PBMC datasets were downloaded from http://scanorama.csail.mit.edu/data_light.tar.gz. Our eight datasets were preprocessed using the featureCounts pipeline and batch-effect correction was performed using Scanorama V1.4.

Journal: Nature biotechnology

Article Title: A multi-center study benchmarking single-cell RNA sequencing technologies using reference samples

doi: 10.1038/s41587-020-00748-9

Figure Lengend Snippet: (a, un-corrected) UMAP of 10 datasets (10X: PBMCs 68K, PBMCs 3K, CD19+ B cells, CD14+ monocytes, CD4+ helper T cells, CD56+ NK cells, CD8+ cytotoxic T cells, CD4+CD45RO+ memory T cells, CD4+CD25+ regulatory T cells; Drop-seq: PBMCs) out of 26 datasets from Hie et al.8 before batch correction by Scanorama. (b, corrected-based on dataset) UMAP of 10 different datasets shown in (a) from Hie et al. after batch correction by Scanorama, colored to identify the datasets. (c, corrected-based on platform) UMAP of 10 different datasets shown in (a) from Hie et al. colored to identify the two different platforms used (10X Genomics and Drop-seq); note poor results using Drop-seq. (d, un-corrected) UMAP of 8 datasets (breast cancer cells: C1_FDA_HT_A, C1_LLU_A, ICELL8_SE_A, and ICELL8_PE_A; and B lymphocytes: C1_FDA_HT_B, C1_LLU_B, ICELL8_SE_B, and ICELL8_PE_B) out of 20 datasets in our study analyzed using three different non-10X sequencing platforms before batch correction by Scanorama. (e, corrected-based on dataset) UMAP of 8 datasets shown in (d) after batch correction by Scanorama, colored to identify the datasets. Note lack of discrimination between different cell types. (f, corrected-based on platform) UMAP of 8 datasets shown in (d) after batch correction by Scanorama, colored to identify the platforms (C1_FDA_HT, blue; C1, purple; ICELL8, pink). The PBMC datasets were downloaded from http://scanorama.csail.mit.edu/data_light.tar.gz. Our eight datasets were preprocessed using the featureCounts pipeline and batch-effect correction was performed using Scanorama V1.4.

Article Snippet: The 10X Genomics scRNA datasets were preprocessed using CellRanger 3.1.

Techniques: Sequencing