The planet has approximately 7,000 living languages. The large language models that power today’s voice assistants, translation tools, educational platforms, and search engines handle only a handful of them properly. This is not a technical accident. It is the result of a precise, documented mechanism that remains largely uncontested: models learn from what already exists online, and what exists online reflects the power imbalances, infrastructure gaps, and investment disparities that have structured the digital world for thirty years.

The Pacific offers a striking illustration. It is home to a significant number of languages — Papua New Guinea alone has approximately 840, accounting for more than 10% of the global total. But Pacific island languages are vastly underrepresented in the training data of large language models. This imbalance goes unnoticed in major technology hubs. Yet it produces a concrete effect: where the digital realm takes over from oral transmission, these languages are absent from accessible tools and services.

This case is not isolated. It illustrates a structural phenomenon affecting African languages, indigenous languages of the Americas, dialects of Southeast Asia, and many others. Wherever a language lacks digital presence, models handle it poorly. And wherever models handle it poorly, its digital presence shrinks further.

Key Points

  • The Pacific alone concentrates more than 1,200 living languages, but very few general-purpose large language models cover them; there exists at least one open model optimized for Samoan, but the total absence of LLMs for all languages mentioned is not established (UNESCO, 2025).
  • Large language models are trained on massively English-language data: English represents between 45% and 65% of tokens according to available estimates, based on corpus analyses like Common Crawl, Wikipedia, and international academic literature.
  • LLMs are less factually accurate in languages less represented in training data — an empirically established finding from several recent studies on multilingual model performance.
  • Possible trajectories through 2035 include silent digital extinction, heritage preservation without active use, or revitalization through open corpora and community partnerships.
  • Initiatives exist — Mozilla’s Common Voice, MzansiLM in South Africa, Orange’s projects with Meta and OpenAI for Wolof and Pulaar — but they remain underfunded and fragmented relative to the scale of the challenge.

A Mechanical Imbalance, Not Inevitability

Understanding why so many languages are absent from training corpora requires examining how these corpora are built. Large language models learn from texts available on the web, in digital libraries, in news archives. A language with little digital textual presence — few websites, few digitized books, few online media outlets — is mechanically underrepresented in these corpora.

This is not simply a volume bias. It is a structural bias that reinforces itself. A language underrepresented in training data produces models that handle it poorly, which can discourage text production in that language, and worsen the underrepresentation. The cycle closes.

Researchers have a name for this phenomenon: representation bias. Some groups, some cultures, some languages are overrepresented in corpora; others are nearly absent. Models reproduce and amplify these imbalances because they optimize for dominant statistical patterns. What research increasingly establishes is that LLMs are less factually accurate in languages with fewer speakers and less presence in training data. This is not a defect correctable at the margins: it is a direct consequence of the architecture of these systems themselves.

There is also a form of cross-interference: even when a multilingual model generates content in a minority language, traces of English — dominant in its training data — intrude into the linguistic form, syntax, and cultural associations. The language produced is technically present but culturally distorted.

Pacific Languages: A Case Study of Universal Mechanics

The Pacific is one of the world’s regions richest in linguistic diversity. Papua New Guinea alone is home to approximately 840 distinct languages. Vanuatu, with its 300,000 inhabitants, has more than 130. These languages carry systems of navigation, botany, customary law, and cosmology contained in no English dictionary.

For centuries, their transmission functioned: oral, embodied, intergenerational. That channel is not dead. But a second channel has been added alongside it, and these languages do not occupy that one.

For several Pacific island languages, the cycle of digital underrepresentation can begin and prove difficult to reverse without intervention. Constraints are multiple: limited connectivity, restricted digital skills, insufficient funding for corpus-capture projects — recording, transcription, annotation, structuring — and the need to adapt interoperable standards between often fragmented initiatives.

Other obstacles beyond technical complexity limit model development for these languages. Models have been built for low-resource languages in Africa, Southeast Asia, and the Americas. In the Pacific, investment in digital linguistic infrastructure has been limited.

Africa and the Americas: The Same Mechanism, Different Faces

The Pacific is just one case among many. In Africa, where an estimated one-third of the world’s languages are spoken, the situation is structurally similar. African languages remain largely underrepresented in the global digital universe, despite tens of millions of speakers for some of them. Factors such as multilingual complexity, dominant policies favoring colonial languages, weak institutional support, and lack of digital infrastructure continue to place these languages in a weakened position within AI systems.

The effects are measurable. Where GPT-4o exceeds 90% accuracy in English, its performance drops below 40% for several African languages — a gap documented by recent benchmarks. Languages such as Hausa, Yoruba, and Igbo, spoken by tens of millions of people, have historically lacked high-quality datasets for integration into advanced AI systems.

Initiatives are emerging to correct course. In South Africa, researchers at the University of Cape Town developed MzansiLM, a model capable of handling the country’s eleven official languages, supported by a specific corpus called MzansiText. Orange is collaborating with OpenAI and Meta to fine-tune models for Wolof and Pulaar, two West African languages spoken by tens of millions of people. These initiatives illustrate an emerging dynamic, but they remain exceptions in a landscape dominated by a few dozen languages.

The phenomenon extends beyond Africa. In the Americas, indigenous languages — from Quechua to Nahuatl to First Nations languages in Canada — are in a comparable situation. In Southeast Asia, hundreds of regional dialects remain absent from general-purpose models. In all these cases, the structure is identical: preexisting weak digital presence, training corpora that ignore it, a model that marginalizes it further.

Transmission Under Digital Pressure

Oral transmission has its own conditions for survival. It requires copresence, community continuity, spaces where the language is useful in daily life. These conditions are eroding in many regions worldwide, through urbanization, economic migration, and instruction in colonial or dominant languages.

But the digital shift changes the nature of the question. When a twelve-year-old Tongan child seeks information about his own culture, he passes through a search engine that answers in English. When a Yoruba student in Nigeria uses a voice assistant, it responds to him in a language that is not his. When a community wants to produce digital content in its language, it runs into gaps in available tools: adapted keyboards, spell-checkers, speech synthesis.

The digital realm creates pressure toward dominant languages in spaces that the oral realm never occupied. Young generations quickly learn that their mother tongue functions in the family sphere and remains invisible in the digital sphere.

According to some analyses, the shift toward large language models as infrastructure for cultural transmission risks creating a new filter of legitimacy. Underrepresentation in models can reduce a culture’s digital visibility and access. Without support, communities risk inadequate representation, but they can also participate in data governance and create or adapt their own tools.

This mechanism aligns with what we have documented on AI’s impact on cultural diversity: AI tools amplify what is already dominant and marginalize what is already fragile.

Three Paths to 2035

The UNESCO 2025 roadmap calls for strengthening digital presence of languages underprovided in digital infrastructure. The question is what they will be in the digital space, and what that presence or absence will do to their long-term vitality.

Three trajectories stand out depending on the scope and nature of interventions undertaken by 2035.

The first trajectory is silent digital extinction. Lack of targeted investment limits representation of underprovided languages in training corpora. Digital tools can function in dominant languages but remain lacunary or even nonexistent for others. Over time, generations learning increasingly on screens might develop the idea that their mother tongue belongs to the private sphere while dominant languages govern useful digital uses. In this scenario, oral transmission could be affected by the contraction of usage contexts.

The second trajectory concerns heritage preservation. Corpora are captured and archives created, but without the structure necessary to use them as living tools. The language survives as an object of documentation, accessible to researchers and presented in digital archives, but remains absent from daily use and regular instruction. This is the fate of many languages that have benefited from documentation efforts without benefiting from revitalization efforts. Archive replaces practice.

The third trajectory requires a change of regime. Coordinated public investments, bringing together international institutions, states, regional universities, and communities, could facilitate the creation of linguistic corpora for open models. Standards for community ownership of linguistic data can help communities control uses of their data. Integration of these models into educational tools, digital platforms, and local media could strengthen language presence. The language could have digital space presence that reinforces rather than supplants its oral or family transmission.

This is the most resource- and coordination-intensive trajectory.

What Already Exists and What Is Still Missing

Several tools and projects already exist in this domain. Mozilla Common Voice is one of the main open platforms for participatory collection of multilingual voice data. It functions on a participatory model: speakers record phrases, others validate them. Dozens of languages have been added in recent years, including several African and American indigenous languages. Few Pacific island languages have significant corpora there.

Researchers are working on corpus-capture projects for languages at risk of digital extinction. But these projects remain, in their vast majority, in a framing or underfunding phase. Community data governance protocols are not yet finalized.

Universities in the Pacific region — the University of the South Pacific, the National University of Papua New Guinea — have linguistic documentation programs. Their researchers produce grammars, lexicons, audio archives. But the distance between an academic linguistic archive and a corpus usable for training a language model is considerable. This conversion work requires specialized skills whose availability in the region is limited.

The question of data rights is equally sensitive — and it arises with equal urgency in Africa, the Americas, all regions where community data has been captured and exploited without benefit to the communities concerned. Experience with ethnographic and genetic research has created lasting distrust. Any credible project will need to propose community intellectual property frameworks guaranteeing that linguistic data belongs to speakers, not to platforms exploiting it.

This challenge meets a broader tension we have documented elsewhere: adopting digital tools without training locals amounts to outsourcing one’s own trajectory. The issue is symmetrical: without data and community participation, tools risk misrepresenting local languages and needs.

Signals That Will Show the Way

Three indicators will allow, in coming years, to position underprovided languages between these trajectories.

The first is the growth in the number of languages with an open corpus on platforms like Hugging Face or Common Voice. Some resources exist for Samoan, Fijian, Wolof, or Swahili, but remain limited for the vast majority of languages concerned. Significant growth by 2028 would indicate the dynamic has changed.

The second is the progress of discussions initiated by UNESCO on standards for digital preservation of endangered languages. This consultation concerns follow-up to the 2015 UNESCO Recommendation on documentary heritage, including digital. Without a coordinated framework, initiatives remain fragmented.

The third is the investment decision of states themselves — whether governments of Pacific islands, African states, or countries hosting indigenous communities. Some have pursued policies promoting their national languages in educational systems. Extension of these policies to the digital realm, through integration of linguistic corpora into school curricula, funding of linguist-engineers, and partnerships with regional universities, would constitute a solid indicator of commitment toward the third trajectory.

AI is not condemned to be a linguistic steamroller. It can also be a revitalization tool, if the data exists and if community rights are respected. Experiments in this direction have been conducted for Māori in New Zealand, with encouraging results: corpora built with speakers, models trained, progressive integration into educational tools. The path is known. What is lacking, in most cases, is not the method. It is political will and funding.


Sources

  1. Artsjournal – The Great Renegotiation: Five Ideas About Where Culture Is Going in 2026
  2. UNESCO Endangered Languages Project 2025, Report on Languages at Risk of Digital Extinction (unesco.org/endangered-languages)
  3. Georgetown Institute Asia-Oceania, Pacific Digital Divide: Brief on Digital Colonialism and Linguistic Representation, 2026
  4. Mozilla Common Voice, Inventory of Available Languages (commonvoice.mozilla.org)
  5. Hugging Face Datasets, Inventory of Low-Resource Language Corpora (huggingface.co/datasets)
  6. Alain Grandjean, Large Language Models and Paradigmatic Biases: An Analytical Framework by Discipline, 2026 (alaingrandjean.fr)
  7. Centre for International Governance Innovation, Key Points on African Languages and AI, Ife Adebara, Memo No. 216, November 2025 (cigionline.org)
  8. AfricaNLP Workshop 2025, Multilingual and Multicultural-aware LLMs (sites.google.com/view/africanlp2025)
  9. Orange Newsroom, Orange Integrates African Regional Languages into Open-Source AI Models, 2025 (newsroom.orange.com)
  10. Osiris.sn, AI: A Model Adapted to 11 Official Languages Emerges in South Africa, 2025 (osiris.sn)