Researchers at Tecnológico de Monterrey sequenced high-coverage whole genomes from 1,481 volunteers in Mexico, identifying nearly 3 million short genetic variants absent from major reference databases, according to research published in Nature Communications.
The dataset revealed more than 47.2 million single nucleotide variants and 8.1 million short insertions and deletions across the oriGen Project cohort. Nearly 3 million non-singleton short variants were missing from both dbSNP and the Mexico City Prospective Study. Admixture analysis demonstrated that researchers require a Mexican training dataset to accurately estimate ancestry compositions through genetic similarity.
Genetic tests showed copy number variations tied to Mexican Indigenous American ancestry, designated MX-AMR, at several specific locations including the LCE1D and RHD genes. About 3.1 percent of participants carried homozygous deletions in RHD, the gene determining the Rh blood group. Among participants with high MX-AMR ancestry, that frequency dropped to 0.6 percent. The deletion is uncommon in East Asians, while the Rh-negative trait is rare in Indigenous American populations. The authors concluded the data supports the hypothesis that Rh-negative blood expanded during the Spanish conquest rather than by genetic drift.
Medical findings extended to drug processing, where 10 percent of volunteers carried a heterozygous 22-42128945-C-T loss-of-function variant in CYP2D6. This enzyme metabolizes painkillers and tamoxifen.
Funding for the research came from Tecnológico de Monterrey and FEMSA, with supplementary Azure credits awarded by Microsoft’s AI for Good Research Lab. Grant IJXT070-23DG01001 supported postdoctoral research before Nature Communications published the peer-reviewed paper on September 5, 2026.
