The Chan Zuckerberg Biohub published, at the end of May 2026, an atlas of more than one billion protein structures generated by the ESMFold2 model. Free. Downloadable. Without usage restrictions. It is the largest structural biology database ever made freely available, and it arrives at the precise moment when the race to control data about living organisms is accelerating between academic laboratories and technology giants.

The stakes go far beyond biology. The fundamental question is simple: who will control the data infrastructures that will feed medicine in the coming decades?

The Essentials

  • At the end of May 2026, the Chan Zuckerberg Biohub publishes an atlas of more than one billion protein structures in open access, generated by the ESMFold2 model developed by the Biohub.
  • The atlas covers 6.8 billion proteins from across the tree of life, notably including metagenomic sequences (soil, ocean, etc.), and not specifically the human proteome — coverage that remains the specialty of DeepMind/EMBL-EBI’s AlphaFold database.
  • The publication occurs in a context of growing tension between open science and the proprietarization of biological data, particularly since Alphabet transformed AlphaFold into a brick in a commercial service.
  • The real test will be adoption by academic and pharmaceutical teams worldwide: an open database has value only if it is used and enriched.

What a Protein Structure Predicts — and Why It Took a Century

The protein is the molecule that acts. It transports oxygen, destroys viruses, transmits nerve signals, regulates gene expression. Its function depends almost entirely on its three-dimensional shape: two proteins composed of the same amino acids in a different order will have different shapes, and therefore radically different functions.

Predicting this shape from the genetic sequence — the “protein folding” problem — has resisted researchers for more than fifty years. The reason is mathematical: a medium-sized protein can theoretically take an astronomically large number of configurations. The experimental approach, through X-ray crystallography or cryo-electron microscopy, produces results of remarkable precision, but it takes months and costs tens of thousands of euros per structure.

In 2021, DeepMind changed the game. AlphaFold2, the model from Alphabet’s subsidiary, predicted the structure of nearly 200 million proteins with precision comparable to experimental methods. The corresponding database, developed with EMBL-EBI, was published in open access and covers the nearly complete human proteome. Immediate result: a measurable acceleration of research into neglected diseases, rare therapeutic targets, enzymes of industrial interest. Within months, AlphaFold2 became one of the most cited scientific articles in recent history.

But DeepMind did not remain in this posture. In 2024, Alphabet launched AlphaFold3, whose model parameters were partially restricted, with a separate commercial use license. Open science had won a battle in 2021; free access to the infrastructure was not guaranteed nonetheless.

ESMFold2 and the Biohub’s Billion Structures

The ESMFold2 model is developed by the Chan Zuckerberg Biohub. It follows in the lineage of ESMFold (v1), which had been developed by Meta AI Research (FAIR), but the teams behind ESMFold2 — notably those from EvolutionaryScale around Alex Rives — have since been recruited by the Biohub. Unlike AlphaFold, which uses multiple sequence alignment steps to predict structures, ESMFold relies on a large protein language model — an approach that makes it much faster, at the cost of slightly lower precision on certain categories of complex proteins.

This difference in speed is precisely what makes the Biohub’s atlas possible. Generating a billion structures using current experimental methods would take thousands of years and unimaginable resources. With ESMFold2 deployed at scale on the Biohub’s computing infrastructure, it takes a few months.

The Chan Zuckerberg Biohub is a philanthropic organization founded in 2016 by Mark Zuckerberg and Priscilla Chan, endowed with an initial funding of 600 million dollars. Its mandate is explicitly oriented toward open science and interdisciplinary collaboration. The atlas published in May 2026 follows this trajectory: structures are downloadable, metadata is documented, commercial use is not restricted.

This last point deserves attention. An “open” structural biology atlas that prohibited commercial use would have real academic value but limited scope in drug development, where value chains inevitably pass through industrial actors. The absence of restriction is a deliberate choice.

The atlas covers 6.8 billion proteins from across the tree of life, with a large metagenomic component — sequences from soil, oceans, and other environments. Its purpose is therefore not to specifically map the human proteome, a role filled by the AlphaFold database, but to offer unprecedented coverage of the diversity of life as a whole.

The Tension Between Open Access and Capture — a Story We Know From Elsewhere

Molecular biology is not the first field to experience this tension. In AI itself, the open source movement — Meta’s LLaMA, Mistral in Europe, Falcon from the United Arab Emirates — clashes with the growing centralization of the most powerful models at OpenAI, Anthropic, and Google. The parallel is not merely rhetorical: in both cases, the question is whether fundamental cognitive infrastructures will remain common goods or become resources controlled by a few actors. The article published here on the absorption of the best AI labs by Microsoft and Amazon illustrates exactly this capture mechanism — without formal acquisition, through the sheer weight of computational resources.

In genomics, the tension is even older. The Human Genome Project (1990-2003) was partly a race between an international public consortium and Celera Genomics, Craig Venter’s company that wanted to patent the genome. The consortium won, in part, by publishing its data as it went to render patents moot. But the lesson did not definitively settle the debate: since then, companies like 23andMe, Ancestry.com and their successors have built proprietary genomic databases of tens of millions of profiles, without the question of their long-term governance being resolved.

The protein atlas follows in this heritage. Proteins are the level at which drugs act. Antibodies, small molecules, cell therapies — all target specific proteins in specific conformations. Whoever possesses the complete map of the human proteome in three dimensions possesses a discovery infrastructure that no one else can easily reproduce.

What an Open Database Concretely Changes for Pharmaceutical Research

Drug discovery follows a well-documented funnel. Of ten thousand candidate molecules identified in the initial phase, approximately ten reach human clinical trials, and only one obtains marketing authorization. The average cost of an approved drug exceeds estimates of 1 to 2 billion dollars, but these figures are widely debated depending on the scope retained.

The phase where a protein structure database most changes things is the virtual screening phase. Before physically synthesizing and testing thousands of molecules, research teams simulate computationally which are likely to bind to a target protein in a given conformation. The quality of this simulation depends directly on the precision of the available protein structure.

For major known targets — proteins involved in common cancers, cardiovascular diseases, diabetes — experimental structures already exist in databases like the Protein Data Bank (PDB), which today contains more than 220,000 structures. But for less-studied proteins, emerging pathogens, variants related to rare diseases, structures often exist only in predictive form or not at all.

This is where the Biohub’s atlas changes the equation for teams without industry resources. A university in Kenya, a Brazilian laboratory working on tropical diseases, a Vietnamese team on local pathogens: none of them can fund a large-scale crystallography campaign. With an open atlas, they have access to the same starting map as teams at Novartis or Pfizer.

This reasoning has a limit that must be named. The precision of ESMFold2 is not uniform. On intrinsically disordered proteins, multiprotein complexes, membrane proteins — categories that are particularly interesting for pharmacology — protein language models produce less reliable predictions than the hybrid approaches used by AlphaFold3 or experimental methods. The atlas opens access; it does not solve all precision problems.

The Long Battle for Data on Living Things

The horizon of this atlas goes beyond structural biology. What is at stake, over twenty to thirty years, is the control architecture of precision medicine.

Precision medicine rests on a principle: diseases are not uniform. Two lung cancers can have entirely different molecular profiles, and therefore respond to different treatments. To map these profiles and identify appropriate treatments, you must cross genomic, proteomic, metabolomic, and clinical data at very large scale. Those who possess this crossed data possess predictive capacity.

Today, this capacity is concentrating. Biobanks from the British NHS and major American academic centers (NIH All of Us, UK Biobank, FinnGen in Finland) remain in controlled but open access. But proprietary data is also accumulating: Tempus AI, founded by Eric Lefkofsky, built one of the largest private clinical oncology databases in the United States, valued at several billion dollars in 2024. Flatiron Health (acquired by Roche) aggregates oncology data from electronic medical records of hundreds of hospitals. The question of the governance of this data — who accesses it, at what price, for what uses — is subject to no stable international framework.

The Biohub’s publication does not solve this question. But it creates a valuable precedent: a fundamental data infrastructure as large as this can be maintained in open access, at acceptable cost, with philanthropic funding. This precedent will carry weight in upcoming regulatory debates, particularly in Europe where the health data regulation (EHDS, European Health Data Space) has been attempting since 2022 to organize the sharing of medical data at continental scale.

What makes the tension even more acute is the convergence between structural data and genomic data. Companies like Isomorphic Laboratories (DeepMind subsidiary, founded in 2021) or Recursion Pharmaceuticals work on models that integrate protein structures, cellular data, and clinical data to predict not only the shape of a protein, but directly the effect of a drug. If these integrated models remain closed, the Biohub’s open atlas will be valuable but insufficient: it will be a free piece in a puzzle whose keys are proprietary. This is a structural fragility of the openness strategy, and it deserves to be named clearly.

Conversely, if the atlas is indeed used and enriched by the global academic community — if teams in Africa, Southeast Asia, Latin America build upon it — it creates a critical mass that complicates capture. No company can easily replace an infrastructure on which thousands of teams have founded their work. This is exactly what happened with Linux, with Apache software, with Wikipedia: the size of the user community becomes a rampart against proprietarization.

This dynamic has a generational component. Researchers beginning their careers today with free access to an atlas of a billion structures will have a different intuition about what is “normal” in terms of data access. They will build practices, tools, and expectations that will make it politically difficult, twenty years later, to close access to comparable infrastructures. The technical precedent is also a cultural precedent.

What Adoption Will Say in Eighteen Months

The atlas is published. The next question is empirical: how many teams use it, for what, with what results?

The metrics to watch are known. The number of downloads and citations is the first signal, but it is imperfect: a database downloaded without published results demonstrates nothing. More solid markers will be publications that cite the atlas as a source of primary structures, deposits of new structures validated experimentally that confirm or correct ESMFold2 predictions, and clinical trials of drugs whose initial discovery used data from the atlas.

The AlphaFold2 precedent is encouraging. According to EMBL-EBI data, the AlphaFold database had been downloaded more than 1.5 million times and cited in more than 14,000 scientific publications within two years of its publication. Teams working on malaria, leishmaniasis, Chagas disease — all neglected tropical diseases — cited AlphaFold structures as the starting point for their screening programs. The Medicines for Malaria Venture identified new therapeutic targets from AlphaFold structures; several candidate molecules are today in preclinical trials.

The Biohub’s atlas starts from a broader base, with a different model and potentially more extensive coverage. Its precision limitations are documented and known. Transparency about these limitations is in itself good practice: a database that overstates its precision is dangerous; a database honest about its error margins is reliably usable.

In eighteen months, if teams in Southeast Asia or sub-Saharan Africa publish work using the atlas as infrastructure, the Biohub will have demonstrated something that philanthropic institutes are rarely able to demonstrate: that open science can be an effective strategy, not just an ideal. If the atlas remains principally used by teams that would have had access to alternative resources anyway, the publication will have had value, but it will not have changed the geography of research.

The tipping point is not guaranteed. But for the first time in years in this sector, a major actor has deliberately chosen, and at large scale, not to capture what it could have captured.


Sources

  1. Futura Sciences / Chan Zuckerberg Biohub — Atlas of one billion proteins, May 2026
  2. EMBL-EBI — AlphaFold Database (usage data and download statistics): alphafold.ebi.ac.uk
  3. Protein Data Bank — rcsb.org (statistics of deposited structures)
  4. Nature — Original AlphaFold2 publication (Jumper et al., 2021), link not guaranteed
  5. DeepMind / Isomorphic Laboratories — announcement of creation, 2021 (institutional sources, no guaranteed URL)
  6. NIH All of Us Research Program — allofus.nih.gov
  7. UK Biobank — ukbiobank.ac.uk
  8. Official Biohub statement on ESMFold2 / ESM Atlas (27 May 2026) — biohub.org/news/world-model-of-protein-biology/
  9. Google DeepMind — AlphaFold (official page): deepmind.google/science/alphafold/
  10. AlphaFold3 — Model parameters usage conditions (GitHub DeepMind): github.com/google-deepmind/alphafold3
  11. EMBL-EBI — AlphaFold Database: alphafold.ebi.ac.uk
  12. Wikipedia — Chan Zuckerberg Biohub: en.wikipedia.org/wiki/Chan_Zuckerberg_Biohub
  13. Science (2012) — The Protein-Folding Problem, 50 Years On: science.org/doi/10.1126/science.1219021
  14. Nature — 23andMe plans to sell its huge genetic database: nature.com/articles/d41586-025-01004-3
  15. STBIO — Cost of a protein structure: stbio.de/kosten_en.html