Skip to content
Open access

HPRC2: A human pangenome reference with near-complete coverage of common genetic variation

Julian K. Lucas Prajna Hebbar Wen-Wei Liao Juan F. Macias-Velasco Adam M. Novak Mobin Asri Jennifer R. Balacco A. Blair Davide Bolognini Jana Ebler J. Gardner M. Geleta Cristian Groza Andrea Guarracino Peter Heringer Glenn Hickey S. Koren Shuangjia Lu Maximillian G. Marin Christopher Markovic Mira Mastoras Capucine Mayoud B. McNulty Julian Menendez A. Minkina S. Mohanty J. Monlong K. Munson Keisuke K. Oshima D. Porubsky T. Ranallo-Benavidez A. Raveane W. Seligmann R. Shemirani Yoshihiko Suzuki Jack A. S. Tierney I. Violich DongAhn Yoo Xiaoyu Zhuo Derek Albracht I. Alexandrov Jamie Allen Alawi Alsheikh-Ali Casey T. Andrews D. Antipov L. Antonacci-Fulton A. Arguello Marcelo Ayllon Edward A. Belter H. Bender K. Bonini S. Buonaiuto Shuo Cao Ann M. Mc Cartney Pi-Chuan Chang Xian H. Chang Jitender Cheema Claudio Ciofi Hiram Clawson Sarah Cody Vincenza Colonna Holland C. Conwell M. Diekhans M. Diroma Z. Dong Danilo Dubocanin Jordan M. Eizenga Parsa Eskandar Eddie Ferro Sarah M. Ford Willard W. Ford Adam Frankish Mallory A Freeberg Qichen Fu Shenghan Gao Yan Gao Gage H. Garcia O. Garcia John E. Garza Mohammadmersad Ghorbani Tina A. Graves-Lindsay Bida Gu Leanne Haggerty Nancy F. Hansen Yue Hao T. Hillaker S. N. Hossain Neng Huang Sarah E. Hunt T. Hunt N. Jafarzadeh Nivesh Jain Maryam Jehangir Juan Jiang Juhyun Kim Bonhwang Koo Milinn Kremitzki Daofeng Li Ronghan Li Jiadong Lin Tianjie Liu Ryan Lorig-Roach Hailey Loucks J. Loveland Jianguo Lu Walfred Ma Franco Marsico Jack A. Medico Younes Mokrab Shabir Moosa Avelina Moreno-Ochando Shinichi Morishita Jonathan M. Mudge N. Mwaniki Nasna Nassir C. Natali Shloka Negi L. Ni Faith Okamoto C. Owa S. Paez C. Peano Brandon D. Pickett Laura Pignata Timofey Prodanov A. Radhakrishnan B. Raney A. Rechtsteiner Luyao Ren F. Ryabov S. Sacco F. Salehi Aarushi Sehgal Mahsa Shabani Shadi Shahatit V. Shivakumar Swati Sinha Linnéa Smeds Steven J. Solar Marco Sollitto Nicole Soranzo M. Suner Arda Söylev Chad Tomlinson F. F. Tricomi M. T. Ungaro Rahul Varki B. Walenz Charles Wang Lisa Wang Aaron M. Wenger Conor V. Whelan Zilan Xin Zheng Xu Wenjin Zhang Ying Zhou Giulia Zunino Nicolas Altemose Floris P. Barthel Christina Boucher Guillaume Bourque Andrew Carroll Monika Cechova Mark J. P. Chaisson Haoyu Cheng Robert Cook-Deegan Daniel Doerr R. Durbin Anna-Sophie Fiston-Lavier G. Formenti Stephanie M. Fullerton Robert S. Fulton Shilpa Garg Nanibaa’ A. Garrison Richard E. Green C. Greider Melissa Gymrek Maximilian Haeussler Mohammad Amiruddin Hashmi David Haussler Alexander G. Ioannidis Charles H. Langley Ben Langmead Heather A. Lawson Glennis A. Logsdon Kateryna D. Makova Fergal J. Martin Matthew W. Mitchell P. Ossorio Nadia Pisanti P. Prins Mikko Rautiainen A. Rhie M. Schatz Laura B. Scheinfeldt Kishwar Shafin Jouni Sirén Andrew B. Stergachis Ahmad N. Abou Tayoun Mohammed Uddin F. Villani Mitchell R. Vollger Kai Ye E. Eichler Erik Garrison Ira M. Hall Erich D. Jarvis E. Kenny Heng Li J. Lotempio Tobias Marschall K. Miga A. Phillippy Ting Wang Benedict Paten
Jul 2026 · bioRxiv · 1 citation
Medicine Biology

Abstract

A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium’s (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.