Ninety-seven percent of genomes assembled in East Africa are analyzed outside the continent. Africa represents between 15 and 17% of the world’s population, but less than 2% of genomic data circulating in international research. This asymmetry does not stem from a lack of biological material: the continent is one of the richest in genetic diversity on the planet. It stems from absent infrastructure, and this absence has a cost that Africans pay without perceiving its dividends.

The Essential Points

  • African genetic diversity, the most extensive in the world, is treated as raw material for export: sequences are collected locally, then analyzed and monetized elsewhere.
  • Genomic data of African origin represents 1.82% of the global total, according to a study published in Clinical Microbiology Reviews in 2026, while 97% of genomes assembled in East Africa are processed outside the continent.
  • The structural obstacle lies in infrastructure: Africa holds less than 1% of global data center capacity, and 96.4% of Tanzanian researchers work on personal computers due to lack of local computing resources.
  • India and China broke with this pattern by investing massively in domestic computing infrastructure before developing their analytical capacities; Africa is beginning to chart a comparable path, and investment decisions over the coming years will be decisive.
  • The stakes go beyond bioinformatics: whoever analyzes African genomic data holds the keys to future medicines, diagnostics, and public health policies calibrated to these populations.

African Genetic Wealth, the Most Under-Represented in Science

Sub-Saharan Africa is the cradle of humanity. It is also the region of the world where genetic diversity is greatest, precisely because Homo sapiens evolved there for hundreds of thousands of years before migrating. This evolutionary depth means that genetic variants present in African populations cover a spectrum that European or Asian genomes cannot reproduce. For precision medicine, for research on resistance to infectious diseases, for pharmacogenomics, this diversity is a first-rate scientific resource.

Yet an analysis published in Clinical Microbiology Reviews in 2026 establishes that African genomic data represents 1.82% of the global total. This figure says something precise: African genetic wealth is known, recognized, actively solicited by international researchers, and massively absent from the databases that structure global research. UK Biobank and many European cohorts are overwhelmingly European; All of Us is more diverse, but does not replace proportional representation of populations from the African continent. Diagnostic algorithms trained on this data perform worse on African patients. Polygenic scores and certain precision medicine tools derived from predominantly European cohorts are often less well calibrated or less performant in underrepresented populations, particularly those of African ancestry.

This imbalance has a history. The first large genome-wide association studies (GWAS) were conducted in the 2000s with the most accessible populations for Western researchers: populations of European ancestry. The institutional momentum, reference databases, bioinformatics tools: everything was built on these foundations. Correcting this trajectory requires more than political will: it requires data, local analytical capacities, and institutions capable of producing them sustainably.

Why 97% of East African Genomes Are Analyzed Abroad

The local insufficiency of sequencing, storage, and computing capacities is an important factor in transfers outside Africa, but it is not the sole determinant. Computing power is still very limited and unevenly distributed in Africa compared to major global centers, but it is not absent.

The Brookings report Foresight Africa 2026 and work published in Frontiers in Bioinformatics paint a coherent picture. Africa represents less than 1% of global data center capacity. In a survey of 84 respondents with bioinformatics knowledge in Tanzania, 96.4% reported using a personal computer for this work; access to advanced infrastructure was limited. Assembling a single complete human genome requires several hundred gigabytes of storage and tens of hours of intensive computing. On a laptop, this operation takes days, when it does not simply run up against insuperable material limits.

The resulting protocol is well documented: biological samples, or raw sequencing data, are sent to servers in Europe, the United States, or China. Analysis is conducted there. Results come back, sometimes. Some African data is hosted outside the continent after processing, but this is neither universal nor automatically permanent. Legally, often.

Contractually, always.

The situation amounts to an architecture of incentives, not deliberate coordination. African researchers ask legitimate scientific questions and have access to populations and samples that their foreign collaborators do not. These foreign collaborators, for their part, have the infrastructure, funding, and analytical pipelines.

Collaborations are real and publications are co-signed, but capacities, infrastructure, and data repositories remain partially externalized; the actual distribution of control, intellectual property, and analytical value varies according to projects, contracts, and governance arrangements.

This capture mechanism recalls a broader pattern: assembling without designing is a known economic trap in manufacturing. Genomics reproduces its structure.

The Public Health Stakes That This Dependence Compromises

African underrepresentation in global genomic databases has measurable clinical consequences, and they affect African patients first and foremost.

Genetic tests for predisposition to breast cancer, Alzheimer’s disease, or cardiovascular disease have been validated on cohorts with European majorities. Applied to patients of African ancestry, their diagnostic accuracy drops measurably. Variants of uncertain significance, mutations about which we do not yet know whether they are benign or pathogenic, are overrepresented in African patients precisely because these populations were absent from the studies that allowed them to be classified. An African patient who takes a genetic test more often receives an inconclusive result than a European patient. The inequality is inscribed in the training data, not in the biology.

Beyond diagnostic tests, pharmacogenomics is affected. The way an organism metabolizes a drug depends in part on its genetic profile. The liver enzymes that break down antivirals, antidepressants, or anticoagulants vary across populations. These variations are better characterized in European populations, which means that standard dosing recommendations are better calibrated for them. For African populations, data is lacking, and dosing becomes an approximation.

Research on infectious diseases is also affected. Malaria, tuberculosis, Lassa fever, trypanosomiasis: these diseases kill principally in Africa. Understanding the genetic bases of resistance or susceptibility to these infections in African populations requires African genomic data analyzed by teams that understand these epidemiological contexts. The absence of local infrastructure delays this understanding, and this delay translates into lives.

The Choices of India and China, and Lessons for Africa

India and China went through a similar phase in the 1990s and 2000s. Their health data and genomic sequences went to the West. Their researchers worked on insufficient machines. Their scientific contributions remained marginal in major international journals.

The two countries made an explicit strategic choice: invest in computing infrastructure before investing in science itself. China launched national bioinformatics programs starting in the 2000s, built data centers dedicated to genomics, and produced institutes like the BGI (Beijing Genomics Institute), today one of the world’s largest genome sequencers. India developed public high-performance computing centers, supported national genomic consortia, and built a sufficiently dense base of bioinformatics researchers to treat significant data volumes locally. Today, both countries are exporters of analytical capacity, no longer merely exporters of raw data.

The African path is taking shape, and certain actors deserve to be named. The African BioGenome Project aims to sequence and analyze the genomes of 100,000 African species, with an explicit requirement for local data processing. H3Africa (Human Heredity and Health in Africa), supported by the American NIH and the British Wellcome Trust, has funded genomics laboratories on the continent and is beginning to produce locally processed data. South Africa has substantial computing and genomics capacities, but it is not the only African country with Tier III data centers, and the Tier III label is not sufficient to establish large-scale genomic capacity; several countries—Kenya, Ethiopia, Nigeria—have launched national digital infrastructure strategies that explicitly include health.

These initiatives are real. They remain modest against the scale of the backlog. The parallel with shared research infrastructures in Europe is instructive: even in a context of far superior funding, governance and the distribution of access to shared infrastructure remain open problems. Africa will build its own in a more constrained environment, which makes the quality of choices regarding architecture—federal, national, or continental—all the more decisive.

The Infrastructure That Could Change the Trajectory by 2040

The question is concrete: who will build African data centers, according to what rules of access and with what safeguards on data sovereignty? The answers given in the coming five years will condition the continent’s capacity to retain the value of its own biology over the following two decades.

Three trajectories are open. The first is that of laissez-faire: hyperscalers—Google, Microsoft, Amazon—deploy their African infrastructure according to commercial logic. They are already doing so, with data centers in Johannesburg, Lagos, and Nairobi. This infrastructure is real and alleviates immediate constraints. But data entrusted to a cloud provider subject to American jurisdiction can be targeted by American legal obligations, including when stored abroad, without excluding the application of other jurisdictions, and rights over a model and its training data depend on contracts, applicable law, the status of creators, and intellectual property rules; they do not automatically belong to their creators.

Analytical capacity remains externalized, even if physically localized.

The second trajectory is that of national or continental public infrastructure, financed by states or multilateral institutions, with explicit rules of data sovereignty. The African Union adopted the Malabo Convention on cybersecurity and personal data protection in 2022, which establishes a legal framework. The African Open Science Platform, championed notably by the Academy of Sciences of South Africa, works toward data-sharing standards that maintain African control of scientific data. These frameworks exist. Their implementation depends on funding that few African states can mobilize alone.

The third trajectory is hybrid: partnerships between African public institutions and international technology actors, with contractual clauses on data localization and the sharing of analytical capacity. Agreements of this type are being negotiated, notably within collaborations between African universities and European institutions. Their legal robustness and capacity to resist commercial pressure remain to be proven over time. Experience shows that technological dependence builds quickly and breaks down slowly: initial incentive architectures have considerable inertia.

The signals to watch are concrete. The number of Tier II or Tier III certified data centers operational outside South Africa will say whether infrastructure is truly distributing. National genomic data policies adopted by countries like Kenya, Nigeria, or Ethiopia will say whether African states intend to legally protect their biological heritage. And scientific publications conducted by African teams on African data analyzed locally will constitute the most direct measure of change: that is where scientific value begins to remain on the continent.

The Governance of Genomic Data, an Open Political Challenge

Computing infrastructure is the necessary condition. Data governance is the sufficient condition. And on this second terrain, the challenges are of a different nature.

Genomic data are personal data of particular sensitivity: they identify individuals, reveal medical predispositions, and concern biological families beyond the sequenced person. In contexts where data protection systems are recent and regulatory capacities limited, the risk of ceding this data to commercial actors without informed consent from populations is real. Documented cases in sub-Saharan Africa have highlighted collections of biological samples conducted without rigorous consent protocols in the 1990s and 2000s.

Institutional response is being built. Several countries have adopted specific biobank legislation. National ethics committees are professionalizing. African researchers in bioethics, notably within the H3Africa consortium, work on consent models adapted to African community contexts, where the notion of individual consent must be articulated with collective decision-making structures. This work is less visible than infrastructure announcements, but it is as decisive for African genomic sovereignty as the servers themselves.

The trajectory of French genomic data illustrates that even with considerable means, the governance of biological data is an unresolved problem in well-resourced countries. Africa is building its responses in a context of more severe constraint, but with the possibility of not repeating the architectural errors that have led elsewhere to unusable silos.

The coming decisions depend on determining the right to model African biology, the conditions framing this right, and the designated beneficiaries. The answer lies as much in political choices as in technical choices. It will be given by African actors or, failing that, by those who already hold the infrastructure.


Sources

  1. Frontiers in Bioinformatics (2026), study on bioinformatics capacities in East Africa, including Tanzania
  2. Clinical Microbiology Reviews (2026), analysis of African genomic data representation in global databases: https://doi.org/10.1128/cmr.00387-25
  3. Brookings Institution, Foresight Africa 2026, chapter on digital infrastructure and health data
  4. African BioGenome Project, https://africanbiogenome.org
  5. H3Africa Consortium, Human Heredity and Health in Africa, publications and reports available at https://h3africa.org
  6. African Open Science Platform, reports on the governance of African scientific data
  7. African Union, Malabo Convention on cybersecurity and personal data protection (2022)