RNA sequencing, usually abbreviated RNA-seq, is a family of methods for determining the sequences and relative abundances of RNA molecules in a biological sample. It is used to study the transcriptome—the collection of RNA transcripts present under particular conditions—and to investigate gene expression, transcript structure, and differences between samples. Most methods convert RNA into complementary DNA before sequencing, although direct RNA sequencing reads native RNA molecules. RNA-seq can measure expression and reveal features such as splice junctions that are not readily captured by conventional expression microarrays. (pubmed.ncbi.nlm.nih.gov)
Biological basis and measurement
RNA-seq samples the products of transcription and subsequent RNA processing. Depending on the protocol, it can examine messenger RNA, noncoding transcripts, or selected classes of small RNA. The resulting sequences provide evidence about which genes are expressed and which transcript forms are present. Reads crossing exon boundaries can identify RNA splicing events, including alternative splicing. (pubmed.ncbi.nlm.nih.gov)
An RNA-seq read is a sequence observation, not automatically a count of one original RNA molecule. Fragmentation can produce several sequenceable fragments from one transcript, while amplification can produce multiple copies of a fragment. Abundance estimation therefore depends on library preparation, transcript length, sampling depth, and computational treatment of ambiguous reads. Technical biases associated with sequence composition and fragment position can also affect the measurement. (pmc.ncbi.nlm.nih.gov)
Experimental workflow
A typical complementary-DNA-based experiment has several stages:
- RNA extraction and quality assessment. RNA is isolated from the sample, and its quantity and integrity are assessed. Degraded RNA and very small amounts of starting material require methods adapted to those limitations.
- Selection or depletion. Poly(A) enrichment concentrates many messenger RNAs by capturing their polyadenylated tails. Alternatively, depletion of ribosomal RNA leaves a broader mixture of transcripts. These approaches differ in the RNA populations they retain.
- Library construction. RNA is converted into complementary DNA. Depending on the method, fragmentation occurs before or after this conversion. Sequencing adapters and sample identifiers are added, and many workflows use polymerase chain reaction amplification.
- Sequencing. The library is read using a DNA sequencing platform.
- Computational analysis. Reads are evaluated for quality, assigned to genomic or transcript sequences, and used for abundance estimation or transcript discovery. (nature.com)
Library design determines which questions the data can answer. Strand-specific methods retain information about the strand from which a transcript originated, helping distinguish overlapping transcripts. Methods that sequence transcript ends emphasize gene-level quantification, whereas broader transcript coverage provides more evidence about internal structure and splice patterns. “Total RNA” sequencing does not imply equally effective recovery of every RNA class; depletion, size selection, and other preparation steps still shape the observed population. (nature.com)
Major approaches
Bulk RNA-seq measures RNA pooled from many cells. Its expression profile combines signals from the cell types in the sample. A difference between tissues can consequently reflect altered expression within cells, a change in cell-type proportions, or both. (nature.com)
Single-cell RNA sequencing assigns expression measurements to individual cells. Barcodes identify cells, and many methods also use unique molecular identifiers to help distinguish captured molecules from amplification copies. These methods reveal cellular heterogeneity, but measurements are sparse: failure to detect a transcript in one cell does not necessarily establish its biological absence. (nature.com)
Long-read RNA-seq sequences longer complementary-DNA or RNA molecules, providing more direct evidence about combinations of splice junctions within individual transcripts. Direct RNA sequencing, a distinct but overlapping approach, passes native RNA through nanopores and infers sequence from changes in electrical signal. It avoids sequencing an amplified DNA representation of the RNA and can retain information relevant to RNA modifications, although interpreting such signals requires appropriate models and validation. (nature.com)
Spatial transcriptomics includes sequencing-based methods that attach positional barcodes to captured transcripts. These methods associate expression measurements with locations in tissue sections. Spatial resolution depends on the capture system: a measured location may encompass multiple cells rather than correspond to a single cell. Other spatial-transcriptomic technologies use imaging rather than sequencing. (pubmed.ncbi.nlm.nih.gov)
Computational analysis
Reads can be aligned to a reference genome or transcriptome. Genome-based analysis uses splice-aware alignment to accommodate reads spanning exon junctions. Transcriptome-based methods estimate which annotated transcripts could have generated the observed fragments. When several transcripts share sequence, assignment is uncertain; probabilistic methods distribute evidence among compatible transcripts. Transcript assembly can also reconstruct candidate transcripts from aligned reads. (nature.com)
The output may be a gene-level count table, transcript-level abundance estimates, splice-junction measurements, or reconstructed transcript structures. Gene-level aggregation is generally less demanding than distinguishing similar transcript isoforms, because reads shared by several isoforms may still identify their common gene. Benchmarks show that performance depends on the analytical task and that transcript-level quantification presents particular difficulties. (pmc.ncbi.nlm.nih.gov)
Normalization addresses differences in sequencing depth and sample composition. Transcripts per million (TPM) expresses abundance after accounting for transcript length and scaling the sample total. Count-based comparisons between conditions commonly use sample-specific scaling factors instead. These measures serve different purposes: normalized expression values should not be treated as interchangeable with the count inputs required by a particular statistical model. (pmc.ncbi.nlm.nih.gov)
Differential expression and experimental design
Differential-expression analysis asks whether RNA abundance differs systematically between conditions. Methods such as DESeq2 use generalized linear models with negative-binomial count distributions to account for variability exceeding simple Poisson sampling. They estimate effect sizes and perform statistical hypothesis tests, often sharing information across genes to stabilize estimates when replicate numbers are small. Because many genes are tested, analyses commonly control the false discovery rate. Statistical significance and the magnitude of an expression change are distinct quantities. (pmc.ncbi.nlm.nih.gov)
Biological replication is central to experimental design. Additional sequencing reads improve sampling but do not replace independent biological samples. In single-cell comparisons, cells from the same individual are not equivalent to independent individuals. One approach, pseudobulk analysis, aggregates counts for a cell type within each biological replicate and then compares replicates. Methods that ignore between-replicate variation can produce false discoveries. (nature.com)
Technical differences between laboratories, preparation protocols, and sequencing runs can influence results. Reference materials, consistent processing, and documented analysis procedures support reproducibility, but normalization alone does not eliminate every technical bias. (nature.com)
History and applications
Several studies published in 2008 established high-throughput sequencing as a practical approach to transcriptome measurement. Mortazavi and colleagues demonstrated quantification of mammalian transcripts and direct detection of splice-crossing reads. In 2009, Tang and colleagues reported whole-transcriptome messenger-RNA sequencing from individual cells. Subsequent developments included sequencing-based spatial transcriptomics in 2016 and highly parallel nanopore direct RNA sequencing in 2018. (pubmed.ncbi.nlm.nih.gov)
Applications include comparing expression between experimental conditions, identifying previously unannotated transcripts, characterizing alternative splicing, and detecting candidate fusion transcripts or expressed sequence variants. Single-cell methods distinguish expression programs within heterogeneous populations, while spatial methods relate those programs to tissue organization. Different applications impose different requirements on sequencing depth, read length, library preparation, and analysis. (nature.com)
Limitations and interpretation
RNA-seq does not recover all transcripts equally. RNA degradation, amplification, sequence composition, transcript length, and preparation chemistry can distort coverage or abundance estimates. Low-abundance transcripts are harder to detect, and similar sequences can prevent unambiguous assignment. More sequencing can reduce sampling uncertainty, but it cannot by itself remove systematic preparation biases or resolve every shared sequence. (nature.com)
Expression differences also require biological interpretation. In bulk samples, changes in cellular composition can resemble changes in regulation. In single-cell studies, sparse sampling and dependence among cells affect statistical inference. Transcript discovery and expressed-variant detection have different error profiles from gene-level quantification, so a pipeline validated for one task is not automatically reliable for another. (nature.com)
References
- Mapping and quantifying mammalian transcriptomes by RNA-Seqpubmed.ncbi.nlm.nih.gov
- Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2pmc.ncbi.nlm.nih.gov
- mRNA-Seq whole-transcriptome analysis of a single cellnature.com
- Comparative analysis of RNA sequencing methods for degraded or low-input samplesnature.com
- Multi-platform assessment of transcriptome profiling using RNA-seq in the ABRF next-generation sequencing studynature.com
- Sequencing of first-strand cDNA library reveals full-length transcriptomesnature.com
- Salmon: fast and bias-aware quantification of transcript expression using dual-phase inferencepmc.ncbi.nlm.nih.gov
- Gaining comprehensive biological insight into the transcriptome by performing a broad-spectrum RNA-seq analysisnature.com
- QuantSeq. 3′ Sequencing combined with Salmon provides a fast, reliable approach for high throughput RNA expression analysisnature.com
- Massively parallel digital transcriptional profiling of single cellsnature.com
- Highly parallel direct RNA sequencing on an array of nanoporesnature.com
- Visualization and analysis of gene expression in tissue sections by spatial transcriptomicspubmed.ncbi.nlm.nih.gov