Following the consortium's successful completion of the CHM13 genome, we are branching out and bringing T2T assemblies to additional human genomes and important model organisms.
T2T-CHM13 was the first truly complete, gap-free sequence of a human genome, released by the Telomere-to-Telomere (T2T) Consortium in 2022. Assembled from the nearly homozygous CHM13 hydatidiform mole cell line, it added over 250 million base pairs of sequence missing from the previous reference (GRCh38) and resolved long-inaccessible regions, including all centromeric satellite arrays, segmental duplications, and the short arms of the acrocentric chromosomes. The v2.0 release included a complete Y chromosome sequenced from the HG002 sample.
HG002 extends the T2T approach from a single set of chromosomes to the complete diploid genome of an actual person. While the CHM13 project used a specialized cell line that did not contain genetic information from both parents, HG002 is the genome of a real individual—specifically, the Genome in a Bottle (GIAB) benchmark sample that is widely used to measure the accuracy of DNA sequencing technologies. Because a normal human genome carries two copies of every chromosome, one inherited from each parent, this project developed new methods to resolve both of those copies separately and completely. Read why complete diploid genomes are important to the future of personalized medicine.
The WashU pedigree is a set of complete T2T genome assemblies spanning a multi-generational family. Four individuals were sequenced across three generations (grandmother, grandfather, mother, and granddaughter) from a family of admixed ancestry recruited in St. Louis. By assembling complete, parent-of-origin–resolved genomes for related individuals, the project makes it possible to trace Mendelian inheritance, meiotic recombination, and de novo mutation directly through the most repetitive parts of the genome at base-pair resolution. It also broadens the diversity of complete human reference genomes, with matched cell line resources available for further study.
Consortium partners are finishing a rapidly growing list of near-perfect T2T genomes for individuals of diverse ancestries as well as various cell lines important for biomedical research.
Selected data:
YAO GitHub (East Asian individual)
I002C GitHub (South Asian individual)
KOLF2.1J GitHub (iPSC)
H9 GitHub (ESC)
HG008 Data (pancreatic cancer genome)
We working with the Human Pangenome Reference Consortium to generate complete, T2T genomes for hundreds of humans from around the world, to serve as a new reference for human genomic diversity.
Completing a genome from telomere to telomere is no longer just a human project. The same T2T long-read sequencing and assembly methods are now being applied across the tree of life, producing the first truly complete references for the animals that biologists rely on for research into evolution, biodiversity, agriculture, and health. A complete reference makes each animal a far better research model, improving everything from read mapping and variant calling to gene annotation in studies of disease, evolution, and development.
Much of this work is led by the T2T Consortium together with partners such as the Vertebrate Genomes Project and the Ruminant T2T Consortium.
Selected data: