Aug 2026· Bioinform.· Vol 42· 0 citations· 23 references
Computer ScienceMedicine
TL;DR
NextLongIso is presented, a scalable and reproducible Nextflow pipeline that enables coordinated analysis of multiple layers of transcript regulation and facilitates the transition from transcript identification to functional interpretation of transcriptomic variation.
Abstract
Abstract Summary Long-read RNA sequencing technologies, including Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT), enable direct characterization of full-length transcripts and transcriptome complexity. However, analysis of long-read RNA-seq data remains fragmented across multiple tools, limiting the ability to obtain a unified view of transcript structure, expression, and regulatory variation in long-read transcriptomes. We present NextLongIso, a scalable and reproducible Nextflow pipeline that enables coordinated analysis of multiple layers of transcript regulation. Rather than focusing solely on transcript reconstruction, NextLongIso integrates transcript discovery with downstream regulatory analyses to jointly characterize alternative splicing, isoform switching, transcript boundary dynamics (including alternative promoters and polyadenylation), and transposable element-associated transcription from both PacBio and ONT datasets. By eliminating complex cross-tool data harmonization, this unified framework facilitates the transition from transcript identification to functional interpretation of transcriptomic variation. Availability and Implementation NextLongIso is implemented in Nextflow and is freely available at github: https://github.com/YidanSunResearchLab/nf-LongIso.git and Zenodo: https://doi.org/10.5281/zenodo.21049837.
This review examines the experimental and computational foundations of scLR-seq, including platform selection, library design, cell barcode and unique molecular identifier recovery, transcript discovery, and isoform quantification, and summarize emerging insights into isoform usage, alternative splicing, transcription start and end site selection, allele-specific expression, fusion transcripts, transposable element-derived transcripts, and RNA modifications.
Short-read sequencing-based single-cell transcriptomics represents the current gold standard for studying cellular transcriptomes but remains limited in its ability to resolve full-length transcript isoforms and splicing patterns. Long-read single-cell and single-nucleus RNA sequencing (LR sc/snRNA-seq) enables the transcriptome-wide characterization of full-length isoforms at cellular resolution, yet the relative performance of commercially available workflows remains insufficiently explored. Here, using nuclei extracted from a standardized multi-species benchmark sample and Oxford Nanopore Technologies long-read sequencing, we systematically benchmarked four LR snRNA-seq strategies: 10x Genomics 3’, 10x Genomics 5’, ArgenTag, and Parse Biosciences. Comparing transcriptome features qualitatively and quantitatively, as well as the concordance with matched short-read data and the ability to resolve cellular heterogeneity, we identified substantial method-specific differences in read length and yield, transcript coverage, isoform detection, and recovery of sample-specific biological information, with the 10x Genomics 3’ and 5’ assays emerging as the most balanced approaches for comprehensive isoform-resolved single-nucleus transcriptomics. Altogether, our study provides a systematic assessment of four commercially available workflows for performing LR snRNA-seq and highlights key methodological trade-offs related to distinct library preparation strategies, thus providing practical guidance for future isoform-resolved transcriptome studies at the single-nucleus level.
RNA-Seq, analyses of RNA abundance by next-generation sequencing, has become a near-universal tool in modern biology. Availability of streamlined protocols and kits, straightforward ability to multiplex hundreds of samples, low cost of short-read sequencing, and well-established analytical pipelines make RNA-Seq a method of choice when even a few genes need to be analyzed in parallel. While many tools have been developed for quality control, mapping, and visualization of RNA-Seq data, managing all these individually still requires substantial familiarity with shell scripting and R, and remains a bottleneck for laboratories with limited computational background. We assembled FetchR, an intuitive pipeline with built-in, clear explanations of features and outputs, for local analyses of RNA-Seq data from either own .fastq files or data imported from Sequence Read Archive via the ENA Portal API. The pipeline operates in Windows Subsystem for Linux (WSL) and is installed via a single script that handles all individual tools, as well as their dependencies and updates, including the reference genome annotation(s), and system requirements. The outputs include standard quality control checks, data visualization, read summation, differential gene expression analyses, visualization, and exploratory analyses using Gene Ontology and Gene Set Enrichment Analyses, as well as detailed logs of every step for subsequent reproducible reporting.
Dustin R. Fetch, Alexey A. Soshnev· bioRxiv· 0 citations
Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.
C. Kouam, Jackson Mingle, Pilar Álvarez Jerez et al.· bioRxiv· 0 citations
Motivation Single-cell RNA sequencing (scRNA-seq) allows for the detailed analysis of dynamic cellular processes. In particular, this has been enabled by the estimation of RNA velocity, the derivative of gene expression, from separate count matrices for different splice states, which provides information about a cell’s immediate future even in snapshot data. Useful velocity estimates strongly depend on accurate counts for spliced and unspliced transcripts. Velocyto remains the standard tool for spliced and unspliced mRNA molecule quantification. However, despite considerable advances in scRNA-seq protocols, velocyto has not been updated to account for peculiarities of new protocols, such as popular approaches based on 5’ chemistry. Results To address this shortcoming, we present tidesurf, a command line tool for the quantification of spliced and unspliced transcript molecules from scRNA-seq libraries. Employing it on four different publicly available 10x Genomics Chromium datasets, we show the accuracy on various datasets generated with either 3’ or 5’ chemistry, whereas velocyto’s results are highly erroneous for the latter. Considering broader applicability, our results highlight tidesurf as a potential replacement for velocyto. Availability and implementation A Python implementation of tidesurf is available from PyPI and at github.com/janschleicher/tidesurf. Code for reproducing the analyses is available at github.com/janschleicher/tidesurf projects.
Jan T. Schleicher, M. Claassen· bioRxiv· 1 citation
Abstract Transcription—the process by which genomic DNA is converted into RNA—is a highly dynamic and tightly controlled process across all domains of life. Although bacteria were once regarded as relatively simple organisms, their transcriptomes are now recognized to be remarkably complex, heterogeneous, and subject to multilayered regulation. Despite the availability of an abundance of sequenced bacterial genomes, a comprehensive understanding of how bacteria tune their transcriptional output to adapt to changing environments remains lacking. To this end, SEnd-seq (simultaneous 5′ and 3′ end sequencing) was developed as a high-throughput approach uniquely capable of simultaneously capturing both 5′ and 3′ ends of individual RNA molecules, enabling the reconstruction of full-length transcripts. By capturing each RNA molecule as a distinct molecular entity with single-nucleotide resolution, SEnd-seq has uncovered previously unrecognized transcriptional features across diverse bacterial species, including even the well-studied Escherichia coli. This method performs robustly across a wide range of RNA species and organisms, including hard-to-lyse pathogens such as Mycobacterium tuberculosis. Moreover, SEnd-seq exhibits high sensitivity for detecting low-abundance RNA and is compatible with various target RNA enrichment strategies, as well as genetic, chemical, and functional perturbations, enabling context-specific transcriptomic analyses. As a versatile and broadly adaptable technology, SEnd-seq provides comprehensive insights into transcriptional regulation, RNA processing, and the coordination between RNA-based processes, thereby uncovering potential targets for antibiotic development. In the present review, we summarize the features and applications of SEnd-seq and discuss its future methodological development and expansion into broader biological and biomedical research contexts.