European cultural heritage exists in abundance—museums, libraries, archives—but its presence in the digital corpora that train AI remains fragmentary. Europeana, the European platform for digital heritage, brings together 60 million objects from more than 4,000 institutions. Google Books has digitized approximately 40 million books, from a massively Anglophone corpus, in a comparable timeframe. When large language models train on this data, they learn to describe, classify, and interpret global cultural heritage based on what was digitized first, and by whom.

The Essentials

  • The corpora on which AIs train determine the narratives they will produce: digitizing European cultural heritage is also a matter of cultural sovereignty.
  • Europeana brings together 60 million objects from more than 4,000 European institutions, but a 2024 European Commission report indicates that digitization targets set for 2030 are at risk.
  • Google Books has digitized approximately 40 million books from a corpus dominated by English, offering AI models a massively Anglo-American cultural reference point.
  • If European resources remain underrepresented in training data, AIs will reconstruct European cultural heritage from external sources, with their own angles and gaps.
  • Public and academic initiatives seek to correct this imbalance, but their pace and funding remain the decisive variables.

The Corpus Is the Message

There is a simple principle behind large language models: they learn what they are shown. A model trained on millions of English texts about Italian Renaissance painting will know much about that subject, but viewed from London or New York. A model fed by Italian sources, catalogues from Florentine museums, correspondence from Roman art historians, would know other things, and would state them differently.

Leonardo da Vinci’s birth dates do not change depending on the language of the corpus. However, interpretation, framing, associations of ideas, and implicit hierarchies between works, artists, and periods depend on who wrote, in what language, and whether those writings were digitized. Researchers in legal and cultural semiotics have documented this dynamic: a biased corpus produces biased inferences, even on apparently neutral subjects like cultural heritage.

The issue is therefore less technical than political. The digitization of cultural heritage involves as much the definition of what an AI will judge plausible, typical, or central, as opposed to what it will treat as marginal or exotic, as it does conservation and access.

60 Million Objects, but 2030 Targets at Risk

Europeana is a substantial achievement. Launched in 2008, the platform today aggregates 60 million digitized objects—paintings, manuscripts, photographs, maps, musical scores—from more than 4,000 institutions across Europe. It is one of the largest cultural corpora in the world, and its existence testifies to genuine political will to make European cultural heritage accessible online.

But a European Commission report published in 2024 tempers this optimism. The targets set by the European digital strategy for 2030, particularly regarding the pace of digitization, are judged to be at risk. The causes are multiple: insufficient funding in several member states, unequal technical capacities between large national institutions and small regional museums, and lack of sufficient coordination between countries. The mass of 60 million objects conceals a more fragmented reality: some collections are digitized at 80%, others at less than 5%.

This delay has a direct consequence for AIs. An incomplete corpus constitutes a gap in available training data: what is not digitized does not exist for a language model. This void is then filled by other sources, generally Anglophone ones.

An article on the disappearance of French cultural heritage already posed this diagnosis for physical transmission. Digitization is the digital version of the same problem: without organized effort, it is abandonment by default.

Google Books as Global Reference

Google Books began in 2004. Over two decades, the project has digitized approximately 40 million books, largely from American and British university libraries—Harvard, Stanford, Oxford, Michigan. The corpus is considerable, and its influence on language models is difficult to overstate: GPT, Gemini, and their competitors have all been trained, directly or indirectly, on texts that derive in part from this collection.

The linguistic bias is documented. English overwhelmingly dominates the books digitized by Google, followed distantly by Spanish, French, and German. The languages of Central and Eastern Europe—Polish, Romanian, Hungarian, Czech—are structurally underrepresented. Languages with limited international reach, such as Maltese, Irish Gaelic, or Luxembourgish, are virtually absent.

This imbalance is not a moral fault of Google. It reflects the reality of the libraries that participated in the project and the languages in which the most easily accessible books under copyright were published. But its effects are real. An AI model that has read ten times more texts in English about Baroque art than in Polish or Hungarian will have associations, references, and hierarchies different from those of a historian from Kraków or Budapest.

The practical question concerns what Europe will do about this gap.

Corrections Attempted by European Institutions

Several public and academic actors have taken the measure of the problem and are working on concrete responses.

Europeana itself has evolved. The platform no longer merely aggregates metadata: it works on data interoperability so that the corpora it brings together can be used as training data for European AI models. The Europeana Data Space program, launched as part of the European cloud for cultural data, aims precisely to make these 60 million objects exploitable by researchers and AI developers, not merely browsable by internet users.

The Una Europa university alliance, which brings together eleven European universities, has undertaken work on the representation of cultural heritage in digital corpora. The central idea is that the linguistic and cultural diversity of the continent must translate into diversity of data; otherwise, AI tools produced in Europe will reason, despite themselves, with imported categories.

The European Commission, for its part, has integrated cultural heritage digitization into its digital strategy, and the AI Act adopted in 2024 introduces transparency requirements on the training data of high-risk models. These provisions could, over time, create regulatory pressure on AI developers to document and diversify their corpora. It is an indirect lever, but a real one.

This point connects with broader reflection already published here: AI integrates itself into organizations as a system rather than as a tool. The same logic applies to cultural institutions: an AI integrated without reflection on its training data reproduces inherited hierarchies without reexamining them.

Cultural Soft Power Is Played Out in Corpora

The history of cultural soft power has always played out on infrastructures: the libraries of the Roman empire, the French academies of the seventeenth century, Hollywood in the twentieth. Each era has produced centers of cultural production and distribution whose influence transcended political borders.

The training of AIs is the new infrastructure. A language model massively nourished by Anglophone sources will describe the world, including European cultural heritage, with the categories, references, and hierarchies of those sources. This does not mean that current AIs are false in their descriptions of European cultural heritage. But they have a filtered reading of it, and this reading will spread on a large scale: in educational tools, cultural mediation applications, research assistants, automatic translations of museum catalogues.

By 2030, if European digitization targets remain at risk, the dependence of AIs on Anglophone corpora to interpret European cultural heritage could establish itself as a stable state rather than a transitional phase. The window for building competitive corpora is not unlimited: next-generation models will train on data available in the next five to ten years.

The Decisive Variables: Funding, Coordination, Data Openness

Three factors determine whether Europe can close the gap.

Funding first. Large-scale digitization is expensive: physical digitization, metadata, rights, hosting, maintenance. Several member states have underfunded their national programs, leaving regional institutions without resources. The Commission’s 2024 report identifies this funding gap as among the principal obstacles to 2030 targets. A targeted budgetary boost, notably through structural funds or the Creative Europe program, would constitute the most direct lever.

Coordination next. The 4,000 Europeana institutions work with very different standards, formats, and levels of maturity. Data interoperability, a necessary condition for these corpora to serve in training AIs, requires harmonization that neither the market nor national institutions will spontaneously produce. This is a classic role for a European coordination body, provided it has sufficient mandate and resources.

Data openness finally. Part of the digitized collections remains under complex rights regimes that prevent their use for model training. Works in the public domain are theoretically free, but the digitizations themselves are sometimes claimed as protected by neighboring rights by the institutions that funded the work. Simplifying these regimes, notably for noncommercial research and AI training uses, would unlock part of the existing corpus without waiting for new digitizations.

These three undertakings are known. They have appeared in European reports for several years. What is missing, according to the 2024 report, is execution speed. The urgency comes from the AI calendar, not from that of cultural heritage.

Algorithms and Archives

The question posed by the development of AIs for European cultural heritage is not new in its structure. It resembles the one that the JdP has documented for investigative journalism facing algorithms: when the infrastructures of distribution and classification are controlled by actors whose corpora and interests diverge from those of content producers, those producers progressively lose control of their own visibility and their own interpretation.

For cultural heritage, the dynamic is analogous, but with an added temporal asymmetry. A news article can be republished, indexed differently, recovered in a new corpus tomorrow. A training corpus for a large language model is constituted once, at a given moment, and its biases structure the model for years. The window for action is therefore narrower, and the cost of inaction more lasting.

Europe possesses a mass advantage: 60 million digitized objects already form a considerable corpus, produced by institutions endowed with a legitimacy and historical depth that no private project can reproduce quickly. The condition is that this corpus becomes exploitable, technically, legally, and financially, before the next generations of models have completed their training phase.

Who digitizes first writes the history that AI will tell. Europe has already digitized much. It must now decide whether it also wants to define how this data feeds the tools that will reinterpret its own heritage.


Sources

  1. European Commission, Digital Strategy: 2024 Report on the Digitization of Cultural Heritage
  2. Europeana, aggregated data 2023 (europeana.eu)
  3. Springer International Journal of Semiotics and Law, 2023, digital corpora and cultural biases in language models
  4. Una Europa, work on the representation of cultural heritage in AI training corpora
  5. European Commission, AI Act, 2024, provisions on training data transparency (eur-lex.europa.eu)