In 2018, France ranked third in the European assessment of open public data maturity, at 83%, behind Ireland and Spain, yet its researchers navigate between silos that are not uniformly connected. The country has established rules for free reuse of public research data and targeted opening obligations for certain project-based funding calls, invested in several distinct infrastructures, and French policy has identified the lack of links between infrastructures and planned to build a coordinated ecosystem. Opening obligations and policies preceded certain national support mechanisms, without the transfer of costs to researchers being demonstrated by available sources.

The essentials

  • In 2018, France ranked third in Europe for open data (83%), but since July 8, 2022, Research Data Gouv has constituted a national unifying framework, without however unifying all services and disciplinary practices between Research Data Gouv, HAL, NAKALA and Data Terra, although some have interoperability mechanisms within their own scope.
  • In 2018-2020, several French research communities lacked infrastructure dedicated to data, according to a CNRS census.
  • In 2026, the Research Data Gouv repository contained more than 120,000 files at the time of consultation: a real database. Research Data Gouv combines a multidisciplinary repository, institutional spaces, and a catalog of harvested resources.
  • Data management and curation require significant work; the existence or extent of a transfer of this work to researchers depends on how funding and support mechanisms are organized.
  • The challenge for 2028-2032 is to build a common coordination layer without erasing the diversity of disciplinary practices.

83% maturity, but the seams still need stitching

The 83% score from the 2018 Open Data Maturity study placed France third in Europe for open data. This figure measures the existence of declared policies, legal mandates, and infrastructures. It does not measure whether these infrastructures communicate with each other.

Research Data Gouv aggregates datasets from multiple institutions. HAL centralizes open access publications. NAKALA inherits cultural and humanities and social sciences data. Data Terra federates environmental and Earth data. Each responds to a sectoral or disciplinary logic.

Some infrastructures have interoperable services, but the question of frictionless movement between all spaces remains only partially resolved.

France’s situation is not atypical. Most countries that advanced rapidly on opening mandates built the repositories before the connections. The United Kingdom with UKRI, the Netherlands with DANS, and Germany with the NFDI (Nationale Forschungsdateninfrastruktur) all followed the same sequence: first the repositories, then coordination. Some funded coordination as an entire undertaking, while France initially relied on coordination flowing from mandates.

80% of communities without support: the starting point was low

The figure that illuminates everything comes from the CNRS census covering the 2018-2020 period: the survey indicated that several research communities did not have the appropriate support or infrastructure for their data.

This starting point explains both the real progress and current limitations. Over 120,000 files deposited on Research Data Gouv at the time of consultation in 2026 represent real construction from a near-zero base. Entire communities, in social sciences, biology, physics, have learned to describe their data, to choose licenses, to document their methodologies. This is collective learning that takes time and produces concrete results.

But progress remains uneven. Disciplines that already had a culture of data sharing—genomics, high-energy physics, quantitative economics—migrated faster to the platforms. Humanities and social sciences, arts, clinical disciplines advance more slowly, often because the nature of their data poses specific questions: anonymization, intellectual property rights, ethical sensitivity. NAKALA responds partly to these needs, but articulations with Research Data Gouv depend on local initiatives.

The researcher as default integrator

The absence of a shared standard between multiple platforms can shift alignment work to the least equipped actor to handle it. A researcher facing this friction has three options, all costly. They can document their data two or three times in distinct formats, mobilizing time that no project funding covers. They can choose a single platform and forgo the visibility offered by others, reducing the effective scope of openness. Or they can delegate the task to a team member whose core competence is not metadata management.

In each of these cases, fragmentation creates daily trade-offs between research time and administrative tasks.

Here is the concrete mechanism that transforms a political ambition into a work burden: when two platforms do not share a common metadata schema, someone has to do the translation.

Depositing a dataset on Research Data Gouv requires filling in metadata in a specific format. If that same dataset must be referenced on HAL, because the associated publication is deposited there, a publication deposited in HAL can reference the DOI of the dataset retained in its original repository. If the data also fall within a domain covered by Data Terra, additional articulation may be required.

This cost is invisible in open science maturity statistics. It appears in no research budget as a distinct line item. It dissolves into the unfunded time of researchers, into hours of doctoral students mobilized on curation tasks rather than research, into stops in deposits when the burden becomes too heavy.

This phenomenon resembles, in its logic, other public policies where compliance costs are externalized to the actor least well positioned to bear them. France pays twice for imposed part-time work, an article published in these pages, documented a comparable mechanism: a rule creates a hidden cost that official measurement never records.

The cost of coordination: the example of the German NFDI

Germany launched the NFDI, Nationale Forschungsdateninfrastruktur, in 2020 with dedicated federal and Länder funding for building interconnected disciplinary consortia. The central idea: explicitly fund the coordination layer, not just the repositories. Each NFDI consortium negotiates common standards with others, under the supervision of an umbrella organization.

The assessment after five years is mixed but instructive. Coordination between consortia remains difficult, disciplinary interests carry weight, and harmonization of metadata takes much longer than expected. But the model at least named the problem as a governance problem to be funded, not as an emergent property of researcher goodwill.

France has actors capable of playing this role. The Committee for Open Science, attached to the Ministry of Higher Education, has produced recommendations on interoperability. France Université Numérique and the well-structured network of URFIST, Regional Units for Scientific and Technical Information Training, train researchers in data practices. But none of these actors currently has a mandate and budget to impose common standards between Research Data Gouv, HAL, NAKALA and Data Terra.

Three concrete paths already being pursued by stakeholders

Several initiatives are attempting to bridge the gap without waiting for a centralized reform.

The OAI-PMH protocol, Open Archives Initiative Protocol for Metadata Harvesting, is already used by HAL to expose its metadata to external aggregators. Extending its use to Research Data Gouv and NAKALA with a harmonized metadata profile would reduce friction without imposing a merger of platforms. Working groups within the research library community have been pushing in this direction for several years.

The persistent identifier is a second path. Each dataset deposited receives a DOI, Digital Object Identifier. Each publication on HAL receives a HAL identifier. Systematically linking these identifiers at the time of deposit, via a common registry, would allow crossing silos without merging their architectures. DataCite, the organization that manages DOIs for research data, already provides the infrastructure; the effort lies in systematic adoption of the practice by institutions.

The third path is institutional. Several universities have created research data engineer positions, sometimes called data stewards, whose explicit function is to bridge between platforms. Strasbourg University, Paris-Saclay and a few others have invested in these profiles. The problem: these positions remain too rare, often funded on projects, and their sustainability depends on fluctuating budgets. These are human bridges where technical bridges are also needed.

The structural issue behind the platforms

The difficulty lies partly in how coordination is perceived in research governance: as a service rendered to existing platforms rather than as an autonomous infrastructure. When the success of a device is measured by the volume of deposits rather than by their discoverability and reusability across other spaces, the incentive to invest in interoperability remains limited. Each platform optimizes for its own indicators, which collectively produces fragmentation. Requalifying coordination as an infrastructure in its own right, with its own performance metrics, is therefore a prerequisite for any sustainable progress, regardless of the technical choice made to implement it.

The debate on platforms raises a broader issue: decentralizing research production, with autonomous laboratories, disciplines equipped with their own norms, and institutions attached to their prerogatives, while centralizing the coordination layer enough for data to circulate.

Experience from public data policies in other sectors suggests that the answer lies in a clear distinction between two levels: minimum common standards, which must be imposed and funded centrally, and specific disciplinary practices, which can remain decentralized. This is the model that data.gouv.fr uses for administrative data, or that INSPIRE attempts for European geographic data.

Applied to research, this would mean defining a minimal metadata profile that any platform benefiting from public funding should respect, creating a common identifier registry, and explicitly funding maintenance of this coordination layer. Existing platforms would retain their identity and specific functions. But a query crossing the entire device would become possible.

The Committee for Open Science has recommended similar orientations in its 2021-2024 national plan. The question for 2025-2028 is how to move to execution: who receives the mandate, who receives the budget, and how do you measure success other than with an aggregated maturity score that doesn’t see the underlying fragmentation.

The over 120,000 files of Research Data Gouv constitute a serious starting point. The communities that have learned to describe and share their data represent real collective capital. The next step is to explicitly fund what connects the platforms, each retaining its own utility, rather than rebuilding them.


Sources

  1. CNRS, Sharing research data more effectively: https://www.cnrs.fr/en/update/sharing-research-data-more-effectively
  2. Open Science Monitor, European Commission, open science dashboard (open-science.ec.europa.eu)
  3. Research Data Gouv, national portal for research data (recherche.data.gouv.fr)
  4. HAL, multidisciplinary open archive (hal.science)
  5. NAKALA, research data platform in SSH, Huma-Num (nakala.fr)
  6. Data Terra, infrastructure for Earth and environmental data (data-terra.org)
  7. Nationale Forschungsdateninfrastruktur (NFDI), 2023 activity report (nfdi.de)
  8. Committee for Open Science, National Plan for Open Science 2021-2024, Ministry of Higher Education and Research (ouvrirlascience.fr)
  9. DataCite, international organization for persistent research data identifiers (datacite.org)