The Secret Life of read junctions: how much information can we get from an RNA-seq experiment?
Estefania Mancini
2 de octubre de 2026
2 min readEver since I started working on alternative splicing, I’ve become meticulous about one specific type of short-read alignment: junction reads.
Junction reads can only be aligned by introducing a gap, that is, a jump between two regions of the genome. They don’t align in a complete, continuous stretch, and that is precisely why they hold so much biological information. From them we can extract:
- Well known genes and isoforms quantification
- Known and novel alternative splicing events
- Gene fusions
- circRNAs
- Variants in transcripts
- rRNA percentage and bacterial/viral abundance, from reads not mapped to the genome of interest
The key: three decisions
Reference genome: version and annotation consistent with the research question.
Reference transcriptome: defines which junctions are annotated and which are novel.
Aligner parameters: each analysis requires a different configuration.
An example depending on the question
Alternative splicing. To quantify known events, a good annotation is essential. If we are looking for undescribed events, for example when studying the effect of an aberrant protein, we need to allow unannotated donor/acceptor sites and filter by minimum coverage, because junction support differs between the two cases.
Fusions. We are interested in junctions that span more than one gene, so the aligner must retain chimeric alignments. Coverage is essential to separate true fusions from artifacts.
circRNAs. Back-splice junctions appear as alignments in reverse order, which a standard aligner discards. You need to enable chimeric detection or use dedicated tools.
Unaligned reads. They can be realigned against bacterial, viral genomes or ribosomal databases
An RNA-seq experiment is much more than estimating differential gene expression. Choosing the references well and tuning the aligner before aligning determines how much information you recover. And you, which of these layers are you looking at in your data?