Active Human Transposable Elements: Long-Read Sequencing Technologies, Computational Analysis, and Implications for Human Disease
Abstract
Transposable elements (TEs) account for nearly half of the human genome and shape chromatin organization, gene regulation, and genome evolution. However, their contributions to human physiology and disease remain incompletely understood. The most active elements in humans, LINE-1 (L1), Alu, and SVA, retain some copies with the ability to evade epigenetic repression and mobilize via target-primed reverse transcription (TPRT), whereas copies become inactive through various fragmentations and mutations. TE activity contributes to genomic instability and has been implicated in aging, cancer, neurological disorders, chromatin organization, and epigenetic regulation. Studying TE is challenging due to their repetitive and polymorphic nature. Recent advances in sequencing technologies and short- and long-read sequencing platforms, combined with specialized bioinformatic pipelines, currently enable more comprehensive characterization of TE insertions, deletions, expression, and epigenetic status. Computational approaches vary in sensitivity, specificity, and resource requirements, and their performance is influenced by sequencing modality, coverage, and the reference genome used. Assembly-based and read-based methods, as well as integrating methylation data or single-cell data, provide complementary insights into TE biology. This review summarizes the biology of active human TE, surveys state-of-the-art short- and long-read pipelines for TE analysis, and highlights their applications in studies of aging, cancer, and other complex diseases. We also provide practical guidance for selecting appropriate sequencing strategies and tools for TE-focused projects, and discuss emerging approaches and open questions in the field.