Every time a large AI model trains on the outputs of its predecessor, it erases a little more of what was already marginal in the original data. This isn’t a fixable bug: it’s a statistical property of iterative training. And what the machine forgets first are rare languages, narrative traditions that never went digital, literary canons that never carried much weight in English-language corpora. Global culture is narrowing. The phenomenon is discrete, progressive, and the corpus choices made today will inscribe it in the infrastructure for a generation.
The essentials
- Indiscriminate recursive training on synthetic outputs can introduce irreversible defects under certain conditions.
- A significant portion of European screen time is devoted to American content; non-Western creators are underrepresented in training data.
- Platforms and their recommendation systems influence visibility and revenue in a market marked by concentration of revenue among major labels and popular artists.
- The mechanism is self-reinforcing: exclusion of peripheral creators impoverishes future data, which produces even narrower models.
- Several paths exist: diversity quotas in corpora, regulation of recommendation algorithms, collective rights for creators. Their deployment remains limited.
Model collapse, or how a loop consumes diversity
Language models learn from large collections of textual data, historically mostly human but potentially including synthetic content; images and sounds fall under multimodal models. When real data is replaced at each generation by synthetic data, information from the initial distribution can be lost at each iteration; this result is not universal. Not by chance: what is lost first is what was already marginal in the original data.
Shumailov et al. (2024) studied multiple generations of models trained recursively on generated data. The study shows degradation of generative models in scenarios of recursive training on synthetic data, without evaluating a universal cultural threshold. Languages with low digital representation tend to be even less represented in subsequent generations. Non-Anglophone literary references fade away.
Narrative structures unique to oral cultures or non-digital epic traditions risk not being faithfully reproduced by successive models.
The technical term is “model collapse,” the collapse of the model toward a reduced set of probable responses. But behind the phrase lies something very concrete: a model that no longer knows how to write a Fulani tale, that no longer recognizes the structures of contemporary Japanese novels, that systematically associates “literature” with a handful of Western canons.
This mechanism results from the statistical properties of iterative training on synthetic data and cannot be corrected by debugging alone. The distribution of outputs converges toward its dominant mode, and the dominant mode of current digital corpora is, massively, American and Anglophone.
61% of European screen time on American screens
The dominance of American content on global platforms predates generative AI. Large generative models risk reinforcing this imbalance at levels where they influence content production.
According to the European Audiovisual Observatory, American works represented 61.2% of SVOD time in nine EU countries studied between September 2022 and September 2023. Netflix, YouTube, and Spotify algorithms optimize for engagement, and average engagement rewards the familiar, the smooth, the already-seen, what cognitive science researchers call “processing fluency”—the ease with which the brain assimilates content it recognizes.
Non-Western traditions are underrepresented or misrepresented in training data, while algorithmic curation can disadvantage less well-known creators. Daryani et al. examine the risks of cultural homogenization and convergence of narrative structures by large language models.
The corpora used to train large models include limited representation of content from oral traditions, non-digitized archives, or minority languages. The 7,000 languages recorded by UNESCO are represented asymmetrically: about a hundred languages concentrate most of the available data, the others are very weakly represented. This imbalance risks reflecting in the models unless explicit rebalancing and evaluation measures are taken.
Creator compensation has collapsed in ten years
The attention economy has always been unequal. But the transition to generative AI has caused a rupture in degree that looks like a change in nature.
Creators can be compensated through royalties and licenses. Digital revenues take an increasingly large share of creator income, but precarity, income instability, and intellectual property risks are intensifying. The gap is considerable. Some musicians whose works are used to train generative models receive no explicit agreement or identified compensation, with often insufficient transparency about how their data is being used.
The question touches on the viability of the system as much as on distributive justice. The relationship between cultural value and creator compensation had already fragmented well before AI. Generative AI amplifies this movement by adding a third party: the model that produces content “inspired” by the entire corpus, with no one able to clearly identify what contribution weighed how much in the result.
Some creators reduce their activity or adapt their production to better match the expectations of recommendation algorithms. Cultural diversity risks being impoverished by this, affecting future training data. The loop closes.
UNESCO’s attempts and their limits
The institutional framework is not silent. Later UNESCO instruments, notably digital guidance and the 2021 Recommendation on AI Ethics, address algorithmic issues for cultural diversity. This is a step. But the convention is a soft law instrument; it sets orientations, not binding obligations.
The European Union went further. The Directive on Audiovisual Media Services requires major platforms to maintain a minimum 30% quota of European content in their catalogs. But a catalog quota is not a recommendation quota: a European film present in Netflix’s library but never promoted by the recommendation algorithm changes nothing about the ratio of 61% of viewing time devoted to American content.
Platform regulation has generally focused on content supply, while real influence is exerted through recommendation. Imposing catalog diversity is not enough if recommendation algorithms continue to direct users toward content that dominates average engagement. The question of algorithmic transparency—obligation to publish recommendation criteria, independent audit of cultural bias—remains largely open in most jurisdictions.
This problem of algorithmic legitimacy extends beyond the cultural domain: when an automated system decides what is visible, the question of its decision-making criteria becomes a first-order political question.
The 2026-2030 window, before degradation is inscribed in infrastructure
The coming years are decisive: the corpus and weighting choices for large models currently in design will influence digital cultural production in the next cycle.
Several technical paths have been the subject of serious publications. Deliberate corpus construction, through active selection and weighting to represent underrepresented languages and traditions, is the most promising in the short term. It requires that large model training teams consider cultural diversity as a quality objective, on par with factual accuracy or logical consistency. Digitizing oral archives, translating minor texts, and compensating multilingual annotators represents a cost that private labs have no reason to bear alone: the challenge is organizational and economic as much as technical.
A second lever involves regulation of training data. Requiring model developers to publish the composition of their corpora and respect minimum thresholds for linguistic and cultural diversity is technically feasible; the European Union has already included transparency requirements in the AI Act, but without quantitative specifications on cultural diversity. Completing this framework with precise and auditable metrics is a political decision that can be made within the 2026-2028 window.
A third lever is economic: creating a collective remuneration mechanism for creators whose works feed training corpora. Several models exist—collective management in music, neighboring rights for the press—but none has yet been adapted to the complexity of large language models. UNESCO 2025 recommends exploring this path; member states have not yet translated this recommendation into legislation.
These three levers presume that we must collectively decide that cultural diversity is a common good deserving of public investment and regulatory constraint. This assumption is debated. Some actors argue that overly heavy intervention risks fragmenting the internet into closed zones, each region imposing its own cultural quotas, a scenario where surface diversity masks the balkanization of exchanges. The argument deserves to be taken seriously. But it rests on the hypothesis that the market, left free, would produce more diversity.
Without intervention, there is a risk of content concentration generated by a limited number of labs toward dominant cultural registers. Shumailov et al. (2024) show degradation of generative models in scenarios of recursive training on synthetic data, without evaluating a universal cultural threshold.
Several generations of large models have already integrated synthetic data into their training.
Initiatives already underway in the field
Several initiatives show that pressure is not absent from the ground.
Common Voice, Mozilla’s project to collect voice data in underrepresented languages, has reached more than 30,000 hours of recordings in about a hundred languages—a fraction of what is needed, but proof that community collection is possible. Masakhane, a collective of African researchers specializing in natural language processing, produces linguistic resources for sub-Saharan African languages and publishes its work in open access. The AI4Bharat initiative does the same for Indian languages.
These projects operate outside major commercial platforms and have limited resources. Their impact remains limited for lack of obligation for developers to integrate them. Regulation could intervene by funding the production of missing data and establishing representation criteria, without dictating the aesthetic choices of models.
South Korea, which has been funding since 2021 a national program to digitize its intangible cultural heritage in formats usable by AI models, offers an example of what targeted public policy can do. France engaged similar reflection through the plan for the French language and multilingualism announced in 2024. These initiatives remain isolated; their alignment at the European or international scale has not yet occurred.
The governance of the corpora that will form the cultural memory of models remains undetermined. Industrial labs play a dominant role in developing state-of-the-art models, but corpus choices are not exclusively private; national regulators could claim them, with the fragmentation risk that implies; there is no multilateral body specialized exclusively in AI corpus governance, while UNESCO already possesses competent bodies on digital cultural diversity. The decisions made over the coming years will determine whether cultural diversity is taken into account in the design of the next generation of models.
Sources
- Frontiers Communication (2026), Degradation of cultural diversity by generative models: https://www.frontiersin.org/journals/communication/articles/10.3389/fcomm.2026.1828344/full
- Zolynski et al. (2026), Model collapse and cultural granularity (cited in Frontiers Communication 2026)
- Daryani et al. (2026), Recommendation algorithms and pseudo-cultural diversity, SAGE Publications
- ACM CHI (2026), Cultural Heritage and representation in training corpora
- UNESCO (2025), Working Group on the Convention on the Diversity of Cultural Expressions



