A genetic variant (rs3211938) linked to the CD36 gene had an allele frequency of about 7.7 percent among individuals with over 50 percent African ancestry but only 0.2 percent among those with less than 10 percent African ancestry. Because this genetic variant is rare outside African-ancestry populations, it went undetected by the widely used Genotype-Tissue Expression (GTEx) reference, which is built predominantly on European-ancestry participants. This means that prominent associations in African-ancestry populations could be completely missed in further research. This rs3211938 case illustrates a broader pattern in genome-wide association studies (GWAS), which compare genetic variants across large groups of people to discover patterns that are associated with diseases and traits. Melinda Mills and Charles Rahal’s 2019 study reviewed thousands of GWAS and found that in 2017, 88 percent of participants were of European ancestry. Data were also concentrated geographically; among studies that recruited from one country, 71.8 percent of participants came from the United States, the United Kingdom, or Iceland. To bridge this underrepresentation gap, countries such as India created GenomeIndia, which has sequenced 10,000 people drawn from 83 populations. Several Middle Eastern countries also launched their own genomic projects. As genomic data sovereignty increases, two distinct powers emerge. Bargaining power helps countries control their data, while productive power allows countries to turn their genomic data into biomedical innovation. A state can strengthen the former without securing the latter.
What Representation Changes
A reasonable scientific objection is that much of human biology is shared across populations, so genomic representation may matter more for equity than for scientific discovery. This is partly true: for body mass index and type 2 diabetes, scientists estimate that 70 to 80 percent of the decline in accuracy across populations comes from differences in common variants and how they are inherited, rather than completely different disease-causing variants. However, those differences still affect the accuracy of genetic tools. One important tool is polygenic risk scores, which are predictive models that can estimate an individual’s susceptibility to diseases. They are built from ancestrally and geographically concentrated data, so their predictions reflect that skew. In the UK Biobank, European-derived polygenic risk scores for hemoglobin A1c explained 1.6 percent of variation among European-ancestry participants, but only 0.07 percent among African-ancestry participants. The same scores captured far more information for populations included in the data than for populations that were not included.
This is not an isolated example; across 28 traits, models built from more diverse cohorts identified 380 significant gene-trait associations compared to 268 with the original GTEx models, an almost 42 percent increase. These associations help identify which genes and biological pathways are worth pursuing. When countries are underrepresented, many of their associations go undetected. Therefore, they may have less evidence to support research into targets that are relevant to them, giving scientists fewer reasons to prioritize those pathways for innovation. For those countries, building their own genomic bank does more than close an equity gap. It also allows them to discover connections missed by earlier references and take a greater role in biomedical innovation.
From Data to Discovery
However, genomic banks are just the first step toward scientific output. The next step is productive power: turning data into discovery. This requires linked patient records and computational capacity to analyze them, backed by strong institutions and networks. For some countries, the barrier starts even earlier. Across the World Health Organization’s African Region, only 10 percent of deaths are registered, limiting researchers’ ability to link genetic variants to health outcomes. Countries are learning that building a dataset is much faster than building the necessary ecosystem to support it.
Human Heredity and Health in Africa (H3Africa) is a consortium trying to harness productive power. H3Africa built a network of over 500 researchers across 30 African countries, alongside biorepositories and bioinformatics infrastructure. However, its individual genome-wide studies often had fewer than 10,000 participants, limiting the detection of rarer genetic variants. More participants allow researchers to distinguish noise from genetic patterns and better identify genetic associations. To address this, H3Africa used meta-analysis with other African and international datasets to increase statistical power. Through these strong connections and networks, it was able to scale beyond what its original datasets could provide. Despite these networks, its cohorts still lagged behind. Established datasets hold an advantage that compounds over time. Mills and Rahal found that the most frequently used cohorts tracked participants for long periods of time and had extensive information about participants beyond their DNA. Once these groups of participants exist, researchers reuse them for further discoveries across diseases, creating a snowball effect in which previous investments make future breakthroughs easier. Without this data, newer programs like H3Africa struggle. While H3Africa is still making significant progress through its intricate networks and research capacity, its cohorts lack the scale, depth of clinical information, and years of follow-up available in many established datasets.
The Limits of Control
Many historically exploited countries fear the parachute researcher, who comes in from outside, gathers data, and then exports the resulting intellectual property and economic value. Eva Hilberg argues that genomic sovereignty can prevent uncompensated transfers of biological data and give countries more control over who accesses their data and under what conditions. Mexico implemented this idea by increasing government control over genomic resources and the export of genetic material, thereby strengthening its bargaining power. It did so not only to potentially limit foreign extraction but also to promote local innovation. In 2012, Ernesto Schwartz-Marín and Alberto Arellano-Méndez argued that rather than becoming widely shared resources, genomic data were monopolized among a few elite researchers. This made access to data more limited and created tension among researchers over who could use it, which in practice reduced opportunities for broader collaboration. This showed that greater control does not guarantee broader scientific capacity. In 2015, Augusto Rojas-Martínez similarly argued that policies hindered large-scale collaborations and therefore potentially weakened Mexico’s own genomic development instead. By contrast, H3Africa built African-led research, biorepositories, and networks while still focusing on international research. Despite partly relying on funding from outside organizations, it developed successful research capacity and collaborations across Africa. Therefore, while limiting genomic data can increase bargaining power, it does not necessarily produce increased scientific discovery and capacity.
Linking Control and Capacity
Two forms of power are at play here: bargaining power and productive power. While they reinforce each other, one does not automatically come with the other. India’s GenomeIndia dataset gives it control over genomic evidence, but that control alone does not provide the infrastructure, expertise, and networks needed to turn data into innovation. This is why more countries can create genomic datasets for their populations without necessarily closing the research gap. The states with the greatest genomic power are not those that can only set the terms for others’ access to genomic data, but those that connect their researchers to networks where data can be combined at scale and turned into scientific advances.


