Skip to main content

Guides

Guides for Researchers

How to comply with Horizon Europe mandate

for Research Data Management

  • What are the requirements?

    Proper Research Data Management (RDM) is a mandatory Open Science practice in Horizon Europe projects that generate or reuse digital research data. Beneficiaries must manage these data responsibly, in line with the FAIR principles (Findable, Accessible, Interoperable, Reusable).

    At a minimum, beneficiaries must:

    • Establish by month 6 a Data Management Plan (DMP), and keep it updated throughout the project, covering the data generated and/or reused in the project.
    • Deposit the data in a ‘trusted’ and ‘metadata-ready’ repository as soon as possible or within the deadlines set out in the DMP. When required by the call conditions, use a repository federated in the European Open Science Cloud (EOSC).
    • Ensure Open Access to the deposited data via the repository as soon as possible or within the timeline set in the DMP, under a CC0 or CC BY licence (or equivalent), following the principle “as open as possible, as closed as necessary”. This means that the data access can be restricted to safeguard legitimate interests and constraints. If open access is not provided to some or all data, the reasons must be explained and justified in the DMP.
    • Provide open metadata for deposited datasets (CC0 or equivalent,) and provide information via the repository about any related research outputs and any other tools or instruments needed to reuse or validate the data.
    • Provide information (via the same repository) about any research output or any other tools and instruments needed to re-use or validate the data.

    Keep in mind that ‘research data’ is a very broad concept and certainly not limited to numerical and tabular data.

  • The FAIR principles

    • Horizon Europe emphasises managing digital research data, and where relevant other research outputs, in line with the FAIR principles, which means making them Findable, Accessible, Interoperable and Reusable. Through appropriate metadata, identifiers, standards and clear access/reuse conditions, they can be reused by both humans and machines.
    • Depositing data in a FAIR-enabling, trusted and metadata-ready repository is a key step towards FAIRness, as it supports persistent identifiers, rich metadata, standard access mechanisms and preservation. FAIR does not necessarily mean open: access may be restricted where justified, provided that metadata remain openly accessible and conditions are clearly described within the DMP.
    • There is no one-size-fits-all approach to FAIR. What is appropriate and feasible largely depends on the research domain, data types involved, as well as on the specificities of the project (e.g. sensitivities, constraints, project context), and should be documented and reviewed in the DMP.

    FAIRdataprinciples foster

    Image: https://book.fosteropenscience.eu/

    • Developing a Data Management Plan (DMP)

      Developing a Data Management Plan (DMP) is a core step in complying with Research Data Management (RDM) requirements in Horizon Europe.

      A DMP describes how digital research data (and, where relevant, other research outputs) will be managed both during and after a research project. It identifies key actions and strategies to ensure that research data are of a high-quality, properly documented, secure, preserved, sustainable, and where possible, accessible and reusable. It also outlines how the FAIR principles align with any applicable legal, ethical, or commercial constraints.

      DMP timeline

      • Proposal stage: a formal DMP is not required but a short (typically one-page) DMP-like document is required as part of the proposal. Note that in cases of a public emergency, and where required by the Work Programme, a full DMP may be required at proposal submission or by grant signature.
      • Project implementation: a full initial version of the DMP must be submitted as a project deliverable, normally by month 6.
      • During the project: the DMP is a living document and must be regularly updated to reflect changes in data generation, reuse, access conditions, repositories, or decisions taken. For projects lasting longer than 12 months, at least one updated version must be submitted as a deliverable.
      • End of the project: a final DMP must be delivered, describing how the data have been managed, preserved, and shared.

      Whenever possible, beneficiaries are encouraged to make DMPs publicly available as non-restricted deliverables, under open access and a CC BY licence, to support reuse. This may be done by uploading the DMP on a trusted repository (e.g. Zenodo), which will make it available on the CORDIS website.

      What to include in a DMP?

      Horizon Europe provides a recommended DMP template, which beneficiaries are encouraged, but not required, to use.

      To support both proposal preparation and DMP development, beneficiaries may use ARGOS, OpenAIRE’s free and open-source online tool for creating and publishing DMPs. ARGOS offers templates aligned with Horizon Europe requirements, together with contextual guidance for each section.

    • Providing access to research data in trusted repositories

      Trusted repositories

      Trusted repositories are infrastructures that provide reliable and long-term access to digital resources such as data, publications, etc. Usually, these repositories go through assessment or certification processes to guarantee that certain quality criteria are met.

      In Horizon Europe, the following are considered trusted repositories:

      • Certified repositories, for example those with CoreTrustSeal, Nestor Seal DIN31644, or ISO16363 certifications.  
      • Domain-specific repositories that are internationally recognised, commonly used and endorsed by the research communities relevant to your project.
      • General-purpose repositories or institutional repositories that, without official certification, present the essential characteristics of trusted repositories (e.g. security provisions, services for the creation of machine actionable metadata, long-term preservation of data, etc.).


      You can find more information on how to find a trustworthy repository for your data on our dedicated page.

      When should data be deposited?

      In Horizon Europe, data should be deposited as soon as possible after its generation and, at the latest, by the end of the project.

      There are, however, some additional requirements:

      • Data underpinning a scientific publication should be deposited at the latest at the time of publication, and in line with standard community practices.  
      • In cases of public emergency, if requested by the granting authority, immediate open access should be provided. If exceptions to open access to data apply, data should be made available at least to the legal entities that need the data to address the public emergency.
      • In exceptional cases, data can be deposited after the project has finished.

      'As open as possible, as closed as necessary'

      In Horizon Europe, research data should be made open access by default and licensed under the latest version of CC BY (attribution required) or CC0 (public domain), or equivalent.

      However, it is recognized that data should be ‘as open as possible, as closed as necessary’, and exceptions can be made when providing open access to data:

      • Is against the beneficiary’s legitimate interests, including regarding commercial exploitation;
      • Is contrary to any other constraints, such as data protection rules, privacy, confidentiality, trade secrets, Union competitive interests, security rules, intellectual property rights or;  
      • Would be against other obligations under the Grant Agreement. 

      In such cases, data can be kept restricted, closed or under embargo, but beneficiaries must explain in the DMP the legitimate exception(s) under which they choose to restrict access to (some of the) research data. 

      Find more information on How do I know if my research data is protected? 

      Metadata compliance

      Horizon Europe requires that, when you deposit data in a trusted repository, this should be described with rich metadata in line with the FAIR principles. In practice, this means that you should choose a repository that has a number of mandatory and recommended metadata fields. 

      Metadata should:

      • At least include the following fields: author(s), dataset description or abstract, date of dataset deposit or publication date, dataset deposit venue, dataset license (CC0 or CC BY by default) and dataset embargo period (if any).
      • Also include information about Horizon Europe or Euratom funding: grant project name, acronym and number. Ideally, the repository will have dedicated fields for this information. If not, you can include them in other appropriate fields, such as the abstract.
      • Be open access under a CC0 licence or equivalent. This is also recommended in cases where data must be closed or restricted but there are no compelling reasons for metadata not to be findable and accessible.

      Most (trusted) repositories will require that you fill in a metadata form about the data or files that you will publish, which should cover these requirements.

      For additional information, you can check the page on metadata.


      IMPORTANT!

      The European Research Council (ERC) requested in 2023 an analysis of the ‘metadata readiness’ of repositories (Lazzeri 2024). Only five repositories (<intR>²Dok, DANS, HAL, and Zenodo), met the “Essential” readiness level requirements. Many other repositories were at a “Close-to-Essential” level, and were making modifications to add the missing metadata fields. 

      In practical terms, you need to check directly with your repository whether they have all the mandatory metadata fields required by the European Commission. If not, one work around is to deposit your data in Zenodo (for HE compliance) and a subject-specific or your institutional repository (to improve visibility and discoverability, and a better community engagement and networking).

    • Validation (and re-use) requirements

      Horizon Europe requires that, at the time of depositing research data in a trusted repository, you must also provide access (via the repository) to information about any research output or any other tools or instruments needed to re-use or validate the research data.

      • Research outputs, tools and instruments” may include data, software, algorithms, protocols, models, workflows, electronic notebooks and others.
      • What type of information should be provided? A detailed description of the research output/tool/instrument, how to access it, any dependencies on commercial products (e.g. software), potential version/type, potential parameters, etc.
      • Besides information about these outputs, beneficiaries are encouraged to provide open (digital) access to the research outputs, tools and instruments themselves unless legitimate interests or constraints apply.

      Typically, the repository will allow you to provide links to relevant information (e.g. links to related publications, to a code repository, to related datasets, etc.).

    • Costs of Research Data Management

      Research Data Management activities need to be costed into research, in terms of the time and resources needed. By planning early, costs can be significantly reduced. Under Horizon Europe, eligible costs are governed by the Grant Agreement rules (Article 6.2 and general eligibility principles); they must be actual, incurred by the beneficiary, necessary for the action, recorded in accounts, and incurred during the action (with limited exceptions e.g. final reporting). They must be foreseen in the project budget/Description of Action and incurred during the duration of the project.

      To estimate costs for RDM, you can check the online RDM-costing tool and the infographic on ‘What will it cost t o manage and share my data?’.

      • View our webinars recordings

      • How can OpenAIRE help?

        The following additional support materials can help you with the RDM requirements in Horizon Europe projects:

         

        argos logo newOpenAIRE also offers tools for research data management: ARGOS is an OpenAIRE service that simplifies the management, validation, monitoring and maintenance of DMPs. [tool]

        Links and further information

        The following sources were used and contain more extensive information on how to address open science in Horizon Europe proposals:

         

        Links and further information

        The following sources were used and contain more extensive information about Research Data Management, and how to address Open Science in Horizon Europe proposals:

      • Glossary

      Last Edit: April 22, 2026

      Horizon Europe related guides

      Still have questions?

      Contact us via our Helpdesk.
      We try to respond within 48 hours.

      Continue reading

      This Glossary provides the definitions of and practical advice regarding the key terms mentioned in the context of RDM requirements in Horizon Europe. Other terms relating to RDM can be found in the From Science Europe Data Glossary.


       

       

      Anonymisation

      Anonymisation is the process of removing personally identifiable information (information that directly or indirectly relates to an identified or identifiable person) from datasets containing sensitive data. As a result, data subject is no longer identifiable. As opposed to pseudonymisation, anonymisation is not reversible, which means that the re-identification of the data subject is not possible.

      Practical advice:

      • OpenAIRE has developed a tool that can be used for anonymisation Amnesia.

       

       

      Backup

      Data backup is a process of creating a copy of data in a digital format and storing it on another device to ensure that data are saved and to prevent data loss.
      Backups can be full (all files are backed up whenever a backup is made) or partial (only a part of the files, e.g. new files, are backed up).

      Practical advice:

      • One backup should be at a physically separate location.
      • Backups should be made from the master copy.
      • The backup location should be as secure as the master copy location.

       

       

      Controlled vocabularies and ontologies

      Controlled vocabulary is an organised and standardised arrangement of predefined terms (words and phrases) that are used to index content in an information system with the aim of facilitating information retrieval. Controlled vocabularies connect variant terms and synonyms for concepts, link concepts in a logical order and organise them into categories, so as to provide a consistent way to describe data. They can be general and discipline-specific, and can take the form of subject heading lists, thesauri, authority files, taxonomies and alphanumeric classification schemes.

      Ontologies are not controlled vocabularies, but they use controlled vocabularies to establish a formal specification of a conceptual model in which concepts and categories of concepts, properties, relationships among concepts and categories, functions, constraints, and axioms are defined.

      Practical advice:

      • Use standard domain specific controlled vocabularies and/or ontologies wherever possible to better align your outputs with similar data in your field of research. 
      • Use well-documented vocabularies in which terms are assigned persistent identifiers
      • Use repositories that enable you to add terms from controlled vocabularies.
      • A useful domain agnostic resource for finding controlled vocabularies and ontologies can be found here.
      • Check also other resources.

       

       

      Data Management Plan

      Data Management Plan (DMP) is a formal document that outlines how data will be handled throughout the research data lifecycle – from planning, through collecting, analysing, publishing, preserving, to sharing and reusing.

      In Horizon Europe, DMPs are mandatory and a template to guide the preparation of DMPs is provided. The list of points to be addressed includes data types and formats, compliance with the FAIR principles, (metadata, repositories, controlled vocabularies, licences, etc.),  legal requirements (intellectual property rights, GDPR), costs of preservation, data security and ethics, retention periods.

      Online tools that can facilitate the preparation of DMPs are available:

      Practical advice:

      • A DMP is a living document, which should be updated as the project develops. Any deviations from the original proposal can be documented and explained.
      • If possible, make the DMP publicly available.

       

       

      Documentation

      Data documentation includes various types of information that can help find, assess, understand/interpret, and (re)use research data – e.g. information about methods, protocols, datasets to be used and data files, preliminary findings, etc. Documentation helps understand the context in which data were created, as well as the structure and the content of data. Data should be documented through all stages of the research data lifecycle. Detailed and rich documentation ensures reproducibility and upholds research integrity. Documentation also includes metadata.

      Practical advice:

      • Various tools, such as e-lab notebooks, are available to support you in the process of  creating documentation

       

       

      File format

      File format is a standard way of encoding information so that it can be stored in a computer file. Digital research data may be stored in a wide variety of file formats, depending on the devices and tools used in data collection and processing.

      File formats may be proprietary (the encoding-scheme is designed and owned by a company or organisation, and is not published, due to which files can be opened only by those who have particular software or hardware tools) and/or prone to obsolescence (legacy formats, bit rot).

      To ensure that users can access and understand data and that data can be preserved in the long term, use open (defined by an openly published specification that anyone can use) and lossless formats (ensuring that no data or quality loss will occur during file manipulation).

      Practical advice:

      Check also this OpenAIRE Guide.


       

       

      File naming convention

      File naming convention is a framework for generating file names that have a consistent structure, while describing the content of files and their relations to other files.

      Practical advice:

      • Use standardised syntax for file naming to aid better searchability and the ability to perform batch processes. These can take the form of YYMMDD_filename, and it is recommended to use version suffixes wherever possible when creating versions of a file from a master copy. These could, for example, be generated after cleaning or analysis steps.
      • Define the file naming convention in an early stage of your research and apply it consistently throughout the research data lifecycle.

       

       

      Licence

      Licence is a written agreement by means of which the copyright holder defines the rights granted to the users. In a digital environment, standardised licences based on a set of predefined reuse conditions, such as Creative Commons, are used. Licences for digital objects are machine readable. 

      In Horizon Europe, CC BY or CC0 (or equivalent) open licence is required for data in open access, while metadata deposited data must be open under the CC0 or equivalent licence.

      Practical advice:

      • When depositing your data in a repository, consult the repository’s licence policy to check whether it is compliant with the Horizon Europe requirements.
      • If you are combining already published data into a new dataset, check the compatibility of data licences.

       

       

      Metadata

      Metadata are data that provide information about other data, e.g. a description of the content of the data, the date when the data were produced or collected, tools and devices used to obtain data, file formats and sizes, the names of the people who created or collected data, relevant persistent identifiers, etc. Metadata should be created and provided in accordance with commonly used metadata standards, which may be general or discipline specific. This ensures that metadata can be understood by humans and processed and exchanged by machines. 

      Practical advice:

      • When preparing a DMP, you will be required to mention the metadata standards you will “follow to make your data interoperable”. A useful resource for finding available metadata standards can be found here. Choose a standard that is commonly used in your discipline. Do not invent your own!
      • Some metadata are automatically generated (by devices used to create and capture data) and embedded in data files, e.g. in digital audio and video recordings,while some have to be produced manually. 
      • When depositing your data in a repository, you will be guided by the user interface through the process of providing metadata. In the input form, some metadata fields are mandatory, which means that the procedure cannot be completed if this information is not provided. It is highly recommended to provide as detailed information as possible even in non-mandatory fields.

       

       

      Persistent identifier (PID)

      Persistent identifier (PID) is a long-lasting reference to a resource that provides the information required to reliably identify, verify and locate the resource. In a digital environment, PIDs have the form of URLs. When pasted in a browser, they take users to the resource.

      Apart from digital resources, PIDs can also relate to researchers (e.g. ORCID, ISNI), institutions (e.g. ROR), grants, instruments and devices, etc. In this case, a PID leads to the record describing a researcher, an institution, etc. in the relevant registry.

      Examples of PIDs include DOIs, ORCIDs, ISBN, Handles, etc. 

      Practical advice

      • Apply PIDs to your research outputs. PIDs are typically automatically assigned by trustworthy repositories once your data are deposited there.  
      • It is also recommended to register yourself with ORCID to get a PID for you as an individual.

       

       

      Pseudonymisation

      Pseudonymisation (pseudo anonymisation) is the processing of personal data in such a way that the data can no longer be related to the data subject without the use of additional information. However, the additional information must be kept separately and subject to technical and organisational measures to ensure that data subjects remain unidentifiable. As opposed to anonymisation, pseudonymisation is a reversible process, which means that data subjects can be re-identified if access to the additional information is enabled.

      Practical advice

      • OpenAIRE has developed a tool that can be used for pseudonymisation - Amnesia.

       

       

      Repository

      Repository is a digital platform that ingests, stores, manages, preserves, and provides access to digital content. A repository should support a commonly accepted metadata standard and have a protocol enabling metadata exchange. 

      Repositories are usually classified into: 

      • subject/disciplinary, 
      • institutional and 
      • generalist repositories. 

      In Horizon Europe research data should be deposited in a trusted repository, i.e. in a repository that operates in accordance with relevant standards and best practice, provides long-term access and preservation, and ensures compliance with the FAIR principles. Trusted repositories include certified or community-recognised repositories, stable institutional repositories and generalist repositories (such as Zenodo).

      Practical advice:

      • Use the re3data repository registry to identify appropriate repositories.
      • The use of domain specific repositories is the most desirable. Such repositories will employ domain specific metadata standards and controlled vocabularies and/or ontologies which will enhance the ability to do analyses across similar datasets and even across domains.
      • Institutional and generalist repositories should only be considered if no domain specific repository can be found or if there is a mandate to deposit in a specific repository. If it is necessary to deposit data in multiple repositories, this should be done so as to maintain the same persistent identifier across those multiple copies.

       

       

      Sensitive data

      Sensitive data is information that should be protected against unauthorised disclosure because unauthorised access may negatively affect the privacy of an individual, trade and business secrets or even security. In the context of research, sensitive data usually include personally identifiable information (names, date and place of birth, place of living, employment information, etc.), health information,  and other private or confidential data.

      Practical advice:

      • Even in case of sensitive data, it is still possible to adhere to FAIR principles and open science by making the metadata freely available, while not enabling public access to the underlying data. These data can be managed through access control mechanisms and anonymisation and pseudonymisation.

       

       

      Storage

      Data storage is a computing technology that enables saving data in a digital format on computer components and recording media, including cloud services.

      In the context of Research Data Management, it is necessary to ensure that data are stored securely until the end of the project and throughout the minimum retention period. Storage options may include:

      • Portable devices,
      • University network drives,
      • Cloud services.

      Practical advice

      • Storage on local storage spaces on hard drives, pen drives, etc. is discouraged because these devices are vulnerable and data loss may occur. If portable devices are used, it should be ensured that there are copies on networked drives and backup storage.
      • Using in-house or institutionally approved spaces is recommended, especially if regular backups are enabled. 
      • Cloud services are suitable for collaboration with partners from partner institutions. However, it should be checked whether the selected cloud service makes regular backups and whether it falls under European jurisdiction.

      The 2019 PSI – Open Data Directive

      A checklist for IP, Research Data and Open Science

      The 2019 PSI – Open Data Directive represents the latest major upgrade of the PSI legislation in the EU. It amends previous PSI directives which first introduced a regulatory framework for Public Sector Information establishing the principle of reuse by default. With the 2019 revision Public Sector Bodies covered by the Directive can now benefit from a more comprehensive set of rules that regulate the reuse of data and information that they hold. The PSI – Open Data Directive poses an even stronger requirement on reuse by default, expands the type of PSB covered and, very importantly for Open Science, establishes the principle that research data resulting from publicly funded research must be Open Access by default.

      What is the PSI?

      PSI Stands for Public Sector Information and has been a directive since 2003 in order to regulate use of public sector information and to encourage as much public sector to be openly available. In 2017 it changed to the Open Data Directive which provides a common legal framework for government held data.


      • What is the PSI?

        PSI Stands for Public Sector Information and has been a directive since 2003 in order to regulate use of public sector information and to encourage as much public sector to be openly available. In 2017 it changed to the Open Data Directive which provides a common legal framework for government held data.

        The 2019 PSI – Open Data Directive represents the latest major upgrade of the PSI legislation in the EU. It amends previous PSI directives which first introduced a regulatory framework for Public Sector Information establishing the principle of reuse by default. With the 2019 revision Public Sector Bodies covered by the Directive can now benefit from a more comprehensive set of rules that regulate the reuse of data and information that they hold. The PSI – Open Data Directive poses an even stronger requirement on reuse by default, expands the type of PSB covered and, very importantly for Open Science, establishes the principle that research data resulting from publicly funded research must be Open Access by default.

      • I. Overview

        The PSI – Open Data Directive (Directive EU 2019/1024 of the European Parliament and of the Council of 20 June 2019 on open data and the re-use of public sector information), which entered into force on 16 July 2019 and gives Member States 24 months to transpose it into domestic legislation, constitutes the latest major upgrade of the PSI legislation in the EU. It amends Directives 2003/98/EC and 2013/37/EU which represented respectively the first substantial intervention by the EU into the field of Public Sector Information and its significant 10-year update which further reinforced the principle of reuse by default. With the 2019 revision, which is based on principles such as transparency and fair competition in the internal market, Public Sector Bodies (PSB) covered by the Directive can now benefit from a more comprehensive legal framework that regulates the reuse of data and information that they hold. The PSI – Open Data Directive poses an even stronger requirement of reuse by default, expands the type of PSB covered and, very importantly for Open Science, establishes the principle that research data resulting from publicly funded research must be open access by default.

        There are a number of specific provisions that may impact, in certain cases significantly, Open Science practices in the EU once transposed into domestic law by Member States (MS). The following is a list of the most relevant.

        1. The scope of the directive covers documents held by Public Sector Bodies (PSB) in Member States at national, regional and local levels

        This includes national governments, ministries, state agencies and municipalities, as well as organisations funded mostly by or under the control of public authorities (e.g. meteorological institutes). Since the 2013 update, museums, libraries, including university libraries and archives are likewise included, although special rules apply. With the 2019 reform the scope has been further expanded to certain public undertakings under specific rules.

        2. The Directive regulates the reuse of documents held by PSB

        Documents are defined as any representation of acts, facts or information — and any compilation of such acts, facts or information — whatever its medium (paper, or electronic form or as a sound, visual or audiovisual recording). The definition of ‘document’ is not intended to cover computer programmes, however, Member States (MS) may extend the application of this Directive to computer programmes

        3. PSI legislation regulates the reuse of documents held by PSB and establishes the rule that such documents shall be reusable for commercial and non commercial purposes (Art. 1)

        It should be noted that access to information, i.e. the principal way in which the covered type of information is made available, is not directly regulated by PSI, which nevertheless identifies The Charter of Fundamental Rights of the European Union (see Recital 5 PSI-OD) as the legal basis for national access to information rules and also establishes that MS shall encourage public sector bodies and public undertakings to produce and make available documents falling within the scope of the Directive in accordance with the principle of ‘open by design and by default’.

        Access to information legislation is regulated at the MS level. EU institutions have their own procedures to access their documents. Interestingly, Art. 1(6) 2019 PSI – Open Data Directive establishes that the Sui Generis Database Right (SGDR) of Directive 96/9/EC shall not be exercised by public sector bodies in order to prevent the re-use of documents or to restrict re-use beyond the limits set by the 2019 Directive.

        4. Research data resulting from public funding

        Member States will be asked to develop policies for open access to publicly funded research data. This is perhaps the most important innovation brought by the 2019 PSI-OD amendment for Open Science. It implies a mandatory open access status for all research data produced with public funding (Recital 27).

        5. Provision on “high value datasets”

        Another important aspect is the provision on “high value datasets” defined as documents the re-use of which is associated with important benefits for the society and economy, which will be governed by a separate set of rules. Thematically, these datasets identify geospatial, earth observation and environment, meteorological, statistics, companies and company ownership, and mobility as high value areas. The EC will identify the relevant dataset in 2021.

        6. Transparency requirements

        Stronger transparency requirements for public–private agreements involving public sector information, avoiding exclusive arrangements.

      • II. Focus on Research Data and Open Science

        In particular, regarding point 4) above, it should be noted how the 2019 Directive clarifies that under the national open access policies, publicly funded research data should be made open as the default option. However, this “open by default” rule should apply only to research data that have already been made publicly available by researchers, research performing organisations or research funding organisations through an institutional or subject-based repository. This should not impose extra costs for the retrieval of the datasets or require additional curation of data.

        The Directive also clarifies that MS may extend the application of the Directive to research data made publicly available through other data infrastructures than repositories, through open access publications, as an attached file to an article, a data paper or a paper in a data journal. Documents other than research data (e.g. scholarly articles, publications, etc) should continue to be exempt from the scope of this Directive.

        Additionally, concerns in relation to privacy, protection of personal data, confidentiality, national security, legitimate commercial interests, such as trade secrets, and to intellectual property rights of third parties should be duly taken into account, according to the principle ‘as open as possible, as closed as necessary’. Moreover, research data which are excluded from access on grounds of national security, defence or public security should not be covered by this Directive (see Recital 28).

      • III. The importance of national implementations and the development of national open access policies

        These are very important provisions in order to achieve one of the most important goals of Open Science, namely the accessibility, reusability and verifiability of scientific results. Access to the data that brought to certain result is a key element for the achievement of Open Science. The Directive is quite clear in terms of the basic rules, however, these rules will need to be implemented by MS in what the Directive calls national open access policies.

        It is fundamental that these national open access policies follow a common and coordinated approach in order to avoid potential fragmentation at the MS level. Space for intervention at the MS level could be found in areas such as:

        • The inclusion of software in the definition of documents of PSBs covered by the Directive;
        • The extension to other data infrastructures such as open access publications, as an attached file to an article, a data paper or a paper in a data journal as the type of first publication that will trigger the open by default rule;
        • National OA policies should contain a clear identification of a licence or licence type in order to avoid potential issues connected with licence identification, compatibility and maintenance;
        • In this sense EC Decision of 22/02/2019 “adopting Creative Commons as an open licence under the European Commission’s reuse policy” C(2019) 1655 final, which establishes CC BY 4.0 as the European Commission default licence for reuse policy and CC0 for raw data, metadata or other documents of comparable nature, while not binding for MS, indicates a clear and strong orientation for the convergence towards a common standard licence as the default option. MS should be encouraged to follow this common pattern for the generality of cases and only diverge when special circumstances justify a different approach.

      Factsheet

      DOI

      Still have questions?

      Contact us via our Helpdesk.
      We try to respond within 48 hours.

      Continue reading

      Guides for Researchers

      Making Sense of the FAIR Principles

      Understanding the ideas behind Findable, Accessible, Interoperable and Reusable

      BACK TO GUIDES

      • What is FAIR

        Digital research outputs are most valuable when they can be used beyond the project in which they were created. To make this possible, researchers need to organise and manage their data so they remain easy to find, understand and reuse, both during a project and in the future.

        The FAIR Principles provide a framework for achieving this. They are a set of guiding principles that help researchers manage research data so they remain useful over time for themselves, collaborators, other researchers and computer systems.

        The FAIR Principles were introduced in 2016 by a group of researchers and data experts to improve the management and sharing of digital research outputs in an increasingly data-driven research environment (Wilkinson et al., 2016). As the amount of research data continues to grow, it is no longer realistic for people alone to find and work with all available data. The FAIR Principles therefore encourage researchers to organise and describe their data so that both people and computer systems can find, understand and use them effectively. This ability for computer systems to automatically find, access and process data is known as machine-actionability. It is becoming increasingly important as technologies such as artificial intelligence rely on data that can be processed automatically and consistently.

        FAIR is an acronym for four principles:

        • Findable: Data and their descriptions should be easy to find.
        • Accessible: Once found, it should be clear how the data can be accessed.
        • Interoperable: Data should be organised so they can be combined with other data and used across different software and systems.
        • Reusable: Data should include enough information to allow them to be understood and reused appropriately.

        FAIR is not the same as Open Data. Data can be FAIR without being openly available, for example when access may remain under embargo or be shared through controlled access for legal, ethical, commercial or security reasons. Likewise, making data openly available does not automatically make them FAIR if they lack sufficient metadata, documentation or other information needed for reuse. The goal is to make research data as useful as possible, regardless of whether they can be shared openly.

        Finally, FAIR is best understood as a set of guiding principles rather than a technical standard or certification. It is also a continuum rather than a binary concept. Researchers do not either achieve FAIR completely or not at all. Instead, small improvements, such as organising files consistently, documenting how data were collected or choosing an appropriate repository, can all make research data more FAIR.

      • Why FAIR matters

        Applying the FAIR Principles helps maximise the value of research by making data easier to discover, understand and reuse. FAIR data management supports transparency and reproducibility, enables collaboration and data reuse, and helps maximise the scientific, societal and economic value of research. The following sections highlight some of the main benefits of FAIR research data.

        Help researchers discover existing data

        Research data can only be reused if other researchers know they exist. Making data easier to find helps researchers discover relevant datasets, build on existing work and avoid collecting data that already exist. This saves time, reduces duplication of effort and increases the value of research over time.

        Example: A study analysing more than 2,000 links to shared research data found that data stored in recognised repositories were much more likely to remain available over time than data shared through ordinary web links. The authors concluded that depositing data in repositories is a more reliable way to ensure research data remain discoverable in the future (Briney, 2024).

        Enable new research through data reuse

        Collecting research data often requires considerable time, funding and participant effort. When research data are shared in ways that enable reuse, they can continue to generate new knowledge long after the original project has ended. Researchers can investigate new questions, test different hypotheses and combine data from multiple studies, extending the value of the original research (Piwowar & Vision, 2013).

        Example: UK Biobank was established to support future health research using data from around 500,000 participants. During the COVID-19 pandemic, researchers rapidly reused the resource by linking participants' health records with SARS-CoV-2 testing data and follow-up information. This allowed them to investigate the risk factors and long-term effects of COVID-19 without recruiting a new cohort of participants, saving considerable time and resources (Feng et al., 2023).

        Improve transparency and reproducibility

        Research findings are easier to understand and verify when researchers clearly document how data were collected, processed and analysed. Sharing data together with the information needed to interpret and reproduce the analyses helps others evaluate the research, identify potential errors and build confidently on previous work (Munafò et al., 2017).

        Example: After the journal Management Science introduced a policy requiring authors to share their data and code, the proportion of papers with replication materials increased substantially. For papers where reviewers could access the required data and software, more than 95% of the published results could be fully or largely reproduced. The main barriers to reproducibility were inaccessible data, missing code and insufficient documentation (Fišar et al., 2023).

        Maximise the impact of research

        Research data create the greatest value when they can be reused beyond the original project. FAIR data support innovation, inform public policy, enable evidence-based decision making and accelerate scientific and technological advances. By making research data easier to discover, access and combine, FAIR practices increase the likelihood that research will generate benefits for science, industry, government and society.

        Example: The Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology, has accumulated decades of experimentally determined protein structures through contributions from researchers worldwide. This shared collection of high-quality data enabled the development of AI-based protein structure prediction systems such as AlphaFold, demonstrating how openly shared research data can support major scientific breakthroughs. The resulting AlphaFold Protein Structure Database is now widely used across structural biology, medicine and drug discovery, extending the impact of these shared data even further (Berman & Burley, 2025; Váradi & Velankar, 2022).

      • What does each principle mean

        The FAIR Principles provide a framework for making research data more useful to both people and computers. Rather than being a checklist, they describe four complementary characteristics of well-managed data. The table below summarises each principle before looking at them in more detail.

        Principle In short... How to put it into practice
        Findable Make your data easy to discover. Describe your data clearly, give it a permanent identifier, and deposit it in a trusted repository.
        Accessible Make your data available under clear conditions. Store your data in a repository, explain how it can be accessed, and keep its description available even if the data cannot be shared.
        Interoperable Make your data work with other data and software. Use common file formats, agreed standards and consistent terminology.
        Reusable Make your data easy to understand and reuse. Provide enough documentation, include a licence, and ensure the data are well described and of good quality.

        Findable

        What it means

        Data cannot be cited or reused if nobody knows it exists. Making data findable means that both people and computers can discover your dataset, identify it correctly and understand what it contains without opening every file.

        Even if access to the data is restricted, other researchers should still be able to discover that the dataset exists.

        How to make your data Findable

        Make it easy for other people to discover that your data exist.

        • Store your data in a trusted repository where others can search for them.
        • Give your dataset a permanent reference so it can always be found.
        • Provide a clear description of what the dataset contains.
        • Use a descriptive title and relevant keywords.
        • Connect the dataset to related publications, software or other research outputs.

        Example

        Researchers from the NIAID Systems Biology Consortium found that many infectious disease datasets were difficult to discover because they were scattered across multiple repositories and could not easily be found through services such as Google Dataset Search. By improving how nearly 400 datasets and software tools from 15 research centres were described, the consortium made them much easier to discover while allowing them to remain in their original repositories. This illustrates that making data findable is about improving their discoverability, not necessarily moving them into a single database or making them openly available (Tsueng et al., 2022).

        Accessible

        What it means

        Making data accessible means ensuring that, once a dataset has been found, it is clear how it can be retrieved. Accessible does not necessarily mean openly available. Access may be immediate, restricted or not possible, depending on legal, ethical, commercial or security considerations. In these cases, the conditions for accessing the data should be clearly explained.

        Even if the data itself cannot be accessed, information describing the dataset should be available so that the existence of the research is not lost. This allows others to understand what data were produced, understand how access can be requested, and why they may no longer be available.

        How to make your data Accessible

        Make it clear how other people can obtain your data.

        • Store your data in a trusted place that provides long-term access.
        • Explain whether the data are openly available or whether access is restricted.
        • Describe how people can request access when restrictions apply.
        • Keep a public description of the dataset available, even if the data themselves cannot be shared.

        Example

        Many biodiversity databases make observations of endangered species discoverable while protecting their precise locations. A review of 39 biodiversity databases found that sensitive records are often shared under restricted access or with reduced location precision to reduce the risk of poaching, illegal collection and habitat disturbance. Researchers can therefore discover that the data exist and, where appropriate, request access without exposing information that could threaten vulnerable species (Kaehrle & Eschenfelder, 2025).

        Interoperable

        What it means

        Making data interoperable means ensuring that data can be combined with other datasets and used by different software, tools and researchers. Data should not only make sense to the researchers who created them, but also be understandable and usable by others working in the same field or in different disciplines.

        Interoperability also helps computer systems exchange and interpret data automatically, making large-scale analyses and data integration possible.

        How to make your data Interoperable

        Organise your data so they can be combined with other data and used by different software.

        • Use file formats that are widely supported.
        • Organise your data consistently.
        • Use clear and consistent names for variables, categories and measurements.
        • Avoid abbreviations or codes that only your research group understands.
        • Link related datasets and research outputs together.

        Example

        Psychology and Neuroscience: Researchers studying the brain often collect different types of information, such as brain scans, cognitive tests and clinical assessments. However, these data have traditionally been organised differently by individual research groups, making it difficult to combine results from different studies or analyse them using the same software. To address this, researchers developed the **Brain Imaging Data Structure (BIDS)**, a common framework for organising brain imaging datasets. By adopting a consistent structure, research teams from different institutions can more easily analyse datasets using the same software, compare findings across studies and combine data for larger collaborative projects (Gorgolewski et al., 2016).

        Digital Humanities: Museums, libraries and archives often hold related historical objects, documents and artworks, but each institution has traditionally managed its collections independently. This makes it difficult for researchers to search across collections or study cultural heritage on a larger scale. Initiatives such as Europeana, and research projects such as the Sampo series (e.g. WarSampo), use a common approach that allows collections from many different organisations to be searched and explored together, even though they remain managed by the contributing institutions (Hyvönen, 2022; O'Neill & Stapleton, 2022).

        Reusable

        What it means

        Making data reusable means providing enough information for others to understand, evaluate and use the data correctly, both now and in the future. Someone who was not involved in the original study should be able to interpret what the data represent, how they were collected and any limitations that should be considered.

        Reusable data remain valuable long after the original project has finished because future researchers can build on them with confidence. But, without sufficient documentation, even openly available data may be difficult to interpret correctly.

        How to make your data Reusable

        Provide enough information for other researchers to understand and use your data correctly.

        • Explain how the data were collected, processed and analysed.
        • Describe what the dataset contains.
        • State clearly how other people may use the data.
        • Follow established practices in your research field.
        • Check that the data are complete, accurate and well organised before sharing.

        Example

        A researcher combined data from five potato field experiments conducted over 11 years in four different countries to investigate how different potato varieties respond to environmental conditions (Hurtado-Lopez et al., 2015). Although the datasets had been collected independently and in different formats, a substantial proportion could be successfully reused to produce new findings. However, the process required considerable effort because important information was scattered across publications, theses and personal communications, and some data could not be reused because essential details were missing. Papoutsoglou et al. (2023) later revisited this case to show how better documentation and community standards could have made the data much easier to understand and reuse.

      • FAIR and Open Data

        FAIR and Open Data are closely related but are not the same thing. FAIR focuses on ensuring that research data are well described and reusable, whereas Open Data concerns making data openly available for reuse. Data can be FAIR without being open, and openly available data are not necessarily FAIR.

        Read the guide on Open Research Data for a more in-depth description.

      • Putting FAIR into Practice

        The FAIR Principles are not a separate task to complete at the end of a research project. They are achieved through a series of good research data management practices that help make data easier to find, access, combine and reuse. The checklist below provides a practical starting point. Each topic is explored in more detail in later chapters of this guide.

        FAIR Checklist

        Plan for data sharing from the start

        It is much easier to make data FAIR if you consider how they will be organised, documented and shared before data collection begins. Planning ahead reduces the amount of work needed later and helps identify any legal, ethical or technical issues early in the project.

        Data Management Plans

        Organise your files consistently

        Use a logical folder structure and meaningful file names that make sense to both you and other researchers. Keeping related files together and following a consistent naming approach makes datasets easier to navigate and reduces confusion over time.

        Organising Research Data

        Document your data

        Imagine that someone unfamiliar with your project needs to understand your dataset several years from now. Provide enough information to explain what the data contain, how they were collected, how variables or columns should be interpreted, and whether any processing has already been carried out.

        Describing Your Data 

        Choose suitable file formats

        Whenever possible, use file formats that are widely supported and likely to remain usable in the future. Open or commonly used formats reduce the risk that data become difficult to access because specialised software is no longer available.

        File Formats

        Deposit your data in a trusted repository

        Rather than storing data only on a personal computer, institutional website or as supplementary files to a publication, deposit them in a trusted data repository that is designed for long-term preservation and discovery. Many repositories also help make datasets easier to cite and access.

        Choosing a Repository

        Describe how your data may be reused

        Clearly explain any conditions for reuse. If your data can be shared openly, make this explicit. If access must be restricted, explain why and describe how other researchers can request access if appropriate.

        Licensing Research Data 

        Link related research outputs

        Where possible, connect your dataset with related publications, software, protocols, preprints and other research outputs. These links provide important context and help others understand how the data were generated and used.

        Persistent Identifiers and Linking Research Outputs

        Follow practices used in your research community

        Many disciplines have established ways of organising, describing and sharing data. Following these practices makes it easier for other researchers in your field to understand and reuse your work.

        Standards and Best Practices

        Think beyond publication

        Making data FAIR is not simply about publishing a dataset. Consider whether someone who discovers your data in several years' time would still be able to understand what they contain, how they were created and how they may be reused.

         

        Review your dataset before sharing

        Before depositing your data, check that files are complete, organised and understandable. Remove unnecessary duplicates, ensure documentation is up to date and confirm that any confidential or sensitive information has been handled appropriately.

        Preparing Data for Sharing

         

      • Common Misconceptions

        "A dataset must satisfy all four FAIR Principles before it has any value."
        FALSE: FAIR is not an all-or-nothing concept. Many datasets already satisfy some FAIR Principles and can become more FAIR over time. Even small improvements, such as providing better documentation or depositing data in a repository, can significantly increase their value and reusability.


        "Making data FAIR requires specialised technical knowledge."
        FALSE: Many FAIR practices involve straightforward research data management, such as organising files, documenting datasets and depositing them in an appropriate repository. While some disciplines use specialised standards and technologies, researchers can make substantial progress towards FAIR without becoming technical experts.


        "The FAIR Principles are a checklist that every dataset must follow in exactly the same way."
        FALSE: The FAIR Principles provide general guidance rather than prescriptive rules. The most appropriate way to implement FAIR depends on the type of data, the discipline and any legal, ethical or technical constraints.


        "Interoperable means the data must be machine-readable or use artificial intelligence."
        FALSE: Interoperability simply means that data can be combined with other datasets and used by different software. In many cases, adopting common file formats, consistent terminology and established disciplinary practices is enough to make data much easier to integrate and reuse.


        "Making data FAIR only benefits researchers who reuse the data."
        FALSE: FAIR also helps the original research team. Well-organised, well-documented data are easier to understand, share with collaborators, revisit months or years later, and build upon in future projects.


        "Only completed datasets should be managed according to the FAIR Principles."
        FALSE: FAIR is most effective when applied throughout a project rather than only at the point of publication. Organising and documenting data as they are created reduces effort later and improves day-to-day research workflows.

      • Key Takeaways

        • The FAIR Principles provide a framework for making research data easier to find, access, combine and reuse.
        • Plan for FAIR from the start of your project rather than waiting until publication.
        • The four FAIR Principles are complementary. Together they make data more useful for both people and computers.
        • Organise and document your data as you create them.
        • Use trusted repositories, appropriate file formats and clear conditions for reuse.
        • Remember that FAIR and Open are complementary but different concepts. Even restricted data can be FAIR when they are well described and accompanied by clear access conditions.
        • Small improvements can make data significantly more FAIR. Perfection is not required.
        • Making data FAIR benefits both the original research team and the wider research community by improving data organisation, collaboration, transparency and opportunities for reuse.
      • References

        Foundational Documents

        Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J., Da Silva Santos, L. O. B., Bourne, P., Bouwman, J., Brookes, A., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C., Finkers, R., . . . Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3. https://doi.org/10.1038/sdata.2016.18. Licence: CC BY 4.0

        Research Evidence

        Briney, K. A. (2024). Measuring data rot: An analysis of the continued availability of shared data from a Single University. PLOS ONE, 19, e0304781 - e0304781. https://doi.org/10.1371/journal.pone.0304781. Licence: CC BY 4.0

        Berman, H. M., & Burley, S. (2025). Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society. Quarterly Reviews of Biophysics, 58. https://doi.org/10.1017/s0033583525000034. Licence: CC BY 4.0

        Feng, Q., Lacey, B., Bešević, J., Omiyale, W., Conroy, M. C., Starkey, F., Calvin, C. M., Callen, H., Bramley, L., Welsh, S., Young, A., Effingham, M., Young, A., Collins, R., Holliday, J., & Allen, N. (2023). UK biobank: Enhanced assessment of the epidemiology and long-term impact of coronavirus disease-2019. Cambridge Prisms: Precision Medicine, 1. https://doi.org/10.1017/pcm.2023.18. Licence: CC BY 4.0

        Fišar, M., Greiner, B., Huber, C., Katok, E., & Ozkes, A. (2023). Reproducibility in Management Science. Manag. Sci., 70, 1343-1356. https://doi.org/10.1287/mnsc.2023.03556. Licence: © 2023 INFORMS

        Gorgolewski, K. J., Auer, T., Calhoun, V., Craddock, R., Das, S., Duff, E., Flandin, G., Ghosh, S. S., Glatard, T., Halchenko, Y., Handwerker, D., Hanke, M., Keator, D., Li, X., Michael, Z., Maumet, C., Nichols, B. N., Nichols, T. E., Pellman, J., . . . Poldrack, R. (2016). The brain imaging data structure, a format for organizing and describing outputs of neuroimaging experiments. Scientific Data, 3. https://doi.org/10.1038/sdata.2016.44. Licence: CC BY 4.0

        Hurtado-Lopez, P. X., Tessema, B. B., Schnabel, S. K., Maliepaard, C., Van der Linden, C. G., Eilers, P. H. C., ... & Visser, R. G. F. (2015). Understanding the genetic basis of potato development using a multi-trait QTL analysis. Euphytica204(1), 229-241. https://doi.org/10.1007/s10681-015-1431-2. Licence: CC BY 4.0

        Hyvönen, E. (2022). Digital humanities on the Semantic Web: Sampo model and portal series. Semantic Web, 14, 729-744. https://doi.org/10.3233/sw-223034. Licence: CC BY 4.0

        Kaehrle, M., & Eschenfelder, K. (2025). Justifying Biodiversity Data Access Restrictions: A Global Comparison of Data Policies. Proceedings of the Association for Information Science and Technology, 62. https://doi.org/10.1002/pra2.1259

        Munafo, M., Nosek, B. A., Bishop, D., Button, K., Chambers, C. D., Sert, N. P., Simonsohn, U., Wagenmakers, E., Ware, J., & Ioannidis, J. (2017). A manifesto for reproducible science. Nature human behaviour, 1. https://doi.org/10.1038/s41562-016-0021. Licence: CC BY 4.0

        O’Neill, B., & Stapleton, L. (2022). Digital cultural heritage standards: from silo to semantic web. Ai & Society, 37, 891 - 903. https://doi.org/10.1007/s00146-021-01371-1. Licence: CC BY 4.0

        Papoutsoglou, E. A., Athanasiadis, I., Visser, R., & Finkers, R. (2023). The benefits and struggles of FAIR data: the case of reusing plant phenotyping data. Scientific Data, 10. https://doi.org/10.1038/s41597-023-02364-z. Licence: CC BY 4.0

        Piwowar, H. A., & Vision, T. (2013). Data reuse and the open data citation advantage. PeerJ, 1. https://doi.org/10.7717/peerj.175. Licence: CC BY 4.0

        Tsueng, G., Alvarado Cano, M. A., Bento, J., Czech, C., Kang, M., Pache, L., Rasmussen, L., Savidge, T., Starren, J., Wu, Q., Xin, J., Yeaman, M., Zhou, X., Su, A. I., Wu, C., Brown, L., Shabman, R. S., & Hughes, L. D. (2022). Developing a standardized but extendable framework to increase the findability of infectious disease datasets. Scientific Data, 10. https://doi.org/10.1038/s41597-023-01968-9. Licence: CC BY 4.0

        Váradi, M., & Velankar, S. (2022). The impact of AlphaFold Protein Structure Database on the fields of life sciences. PROTEOMICS, 23. https://doi.org/10.1002/pmic.202200128. Licence: CC BY 4.0

        Research Infrastructures and Resources

        AlphaFold. https://alphafold.com/

        Brain Imaging Data Structure (BIDS). https://bids.neuroimaging.io/

        Europeana. https://www.europeana.eu/en

        RCSB Protein Data Bank. https://www.rcsb.org/

        UK BioBank. https://www.ukbiobank.ac.uk/

        WarSampo. https://www.sotasampo.fi/en/

      DOI

      Author: Jonathan England

      Created: August 2026

      Guides for Researchers on RDM

      • Data formats for preservation

      • How to deal with non-digital data

      • How to deal with sensitive data

      • How to find a trustworthy repository for your data

      • How to make your data FAIR

      • Raw data, backup and versioning

      Still have questions?

      Contact us via our Helpdesk.
      We try to respond within 48 hours.

      Continue reading

      Guides for Researchers

      What Do I Need to Think About Before, During and After My Research?

      The journey of research data, from planning to long-term preservation.

      BACK TO GUIDES

      • Three stages

        Research Data Management (RDM) refers to the organisation, documentation, storage, protection, sharing and preservation of research data throughout a research project.

        The research literature consistently associates good RDM with improved data quality and integrity, reproducibility, collaboration, project management and risk reduction. Although much of the evidence comes from reviews and case studies rather than controlled intervention studies, these benefits are reported consistently across disciplines (Perrier et al., 2017).

        Rather than being a single task completed at the end of a project, RDM begins before any data are collected and continues throughout the research process. Decisions made early in a project, such as how research data will be organised, documented, stored and protected, influence every subsequent stage of the research.

        For simplicity, this guide groups research data management into three broad stages:

        Stage When? Typical activities
        Plan Before the project begins Plan how research data will be collected, organised, documented and stored, and consider roles and responsibilities, ethical and legal requirements, sharing and long-term preservation.
        Manage Throughout the project Organise files and folders, document research data, store and back up data securely, protect sensitive information, and support collaborative working.
        Preserve During and after the project ends Select valuable research data for long-term preservation and, where appropriate, share them through trusted repositories to support verification and future reuse.

        Note: You may have seen other guides present research data management as a lifecycle with more stages. This guide adopts a simpler three-stage model: Plan, Manage and Preserve, to emphasise when key decisions are made throughout a project, while covering the same core RDM activities (Töwe & Barrilari, 2020; Yu et al., 2017).

      • When Things Go Wrong

        Research Data Management isn't just about following good practices. It's also about avoiding problems that can compromise months or even years of research. The following examples are based on published research and real-world incidents. They illustrate how seemingly small decisions about planning, managing and preserving research data can have significant consequences later in a project.

        A Data Transfer That Triggered a GDPR Complaint

        Researchers in Sweden were conducting a survey on ageing and housing in collaboration with a housing company. To recruit participants, the company transferred personal data for around 23,000 people to the university research team so invitations could be sent. The study had ethical approval, formal data-sharing agreements, and the data were stored securely. However, the housing company had not obtained individuals' consent to disclose their personal data to the university for research purposes.

        Within days, a formal GDPR complaint was submitted. The incident had to be reported to the Swedish Authority for Privacy Protection, all affected individuals were informed, and the research team spent four months responding to around 600 emails, many from people requesting that their personal data be deleted. Recruitment was delayed while a new recruitment process was designed (Zingmark & Iwarsson, 2025).

        What could have helped?

        • Planning how personal data would be transferred between project partners before recruitment.
        • Clearly defining responsibilities for handling personal data.
        • Ensuring participant information and consent covered all intended data transfers and uses.

        RDM stage: Plan

        Which File Is the Right One?

        A research team is preparing the final analyses for publication. Several versions of the dataset exist, with names such as Final, Final_v2, and Final_FINAL. Nobody is completely sure which version contains the latest corrections. After days of analysis, the team discovers they have been working on the wrong file and must repeat part of the work.

        This may sound familiar. In a survey of 488 psychology researchers, version-control errors were among the most commonly reported data management mistakes. Researchers also frequently reported ambiguous file naming and analysing the wrong version of a dataset. While many mistakes resulted only in lost time and frustration, around one in five respondents said their most serious data-management mistake had major or extreme consequences, including project failure or erroneous conclusions (Kovács et al., 2021).

        What could have helped?

        • Agree on a shared file naming convention.
        • Use version control instead of creating multiple "final" files.
        • Store files in one agreed location rather than personal folders.
        • Keep the original raw data separate from processed data.

        RDM stage: Manage

        The Data Are Available Upon Request

        A researcher wants to verify the findings of a published study. The paper states that the underlying data are "available from the corresponding author upon request." Years after the research was published, they send an email asking for the dataset. The email address no longer works. After tracking down one of the authors, they are told that they cannot remember where the data have been stored.

        This is not an isolated case. A group of researchers contacted the authors of 516 published studies to request the original datasets. Overall, they managed to find a working email for 74% of papers, but were able to recover only 19% of the datasets, and for papers published before 2000, this fell to 11%. When authors explained why the data were unavailable, the most common reasons were that the files had been lost or were stored on inaccessible media. The study found that the likelihood of a dataset still existing declined by 17% for every year after publication, leading the authors to conclude that long-term preservation cannot rely solely on individual researchers (Vines et al., 2013).

        What could have helped?

        • Depositing research data in a trusted repository instead of relying on "available upon request".
        • Preserving research data independently of individual researchers or personal storage.
        • Planning for long-term preservation before the project ended.
        • Including sufficient documentation so the data remain understandable over time.

        RDM stage: Preserve

        Rescuing a 20-Year Dataset

        Researchers working with historic Australian vegetation data wanted to preserve decades of ecological observations for future research. Although much of the data still existed, recovering them proved to be a major challenge. Some documentation had been created using obsolete software and stored on floppy disks that could no longer be read using modern computers. Other information had to be reconstructed from paper records, obsolete file formats, and outdated spatial references. Recovering and documenting the datasets required years of work and an estimated AU$200,000 in funding before the data could finally be deposited in modern repositories (Specht et al., 2018).

        What could have helped?

        • Preserving research data in sustainable, widely supported file formats.
        • Depositing data in a trusted repository before storage technologies became obsolete.
        • Maintaining complete documentation alongside the data.
        • Planning for long-term preservation throughout the project rather than years afterwards.

        RDM stage: Preserve

      • Plan Before You Start

        As a researcher, you would probably never write a project or grant proposal without first thinking carefully about your research questions, methods and planned analyses. Putting those ideas into writing often reveals assumptions, raises new questions and helps turn an initial idea into a coherent research project.

        Research data should be approached in the same way.

        Based on habits or past experience, you usually already have an idea of how you want to organise your files, where to access your data from, or share responsibilities within the team(s). However, ideas that seem obvious often become less straightforward when they need to be discussed with others or written down. Planning forces these decisions to be made explicitly rather than relying on assumptions.

        For example, it may seem obvious that everyone will use the same folder structure, store data in the same location or document their work in the same way. However, unless these decisions are discussed and recorded, different team members often develop their own approaches. These differences may appear minor at first, but they can gradually lead to confusion, duplicated work or disagreements as the project progresses.

        Planning is therefore less about predicting every aspect of the research than about making practical decisions while there is still time to think about them. Once data collection and analysis are underway, the focus is on the research itself, and it is rarely the ideal moment to redesign folder structures, decide where files should be stored, or clarify who is responsible for documenting the data.

        Writing these decisions down, usually in the form of a Data Management Plan (DMP), also provides a shared reference for everyone involved. It records not only what was decided, but also why. As projects evolve, decisions can always be revisited and updated, but documenting them early helps keep everyone aligned and avoids having to rethink the same questions later.

      • Manage Throughout the Project

        Once data collection begins, research data management becomes part of the research process itself. Rather than being a separate administrative task, it consists of a series of small decisions made throughout the project that help keep research organised, understandable and secure.

        Many of these decisions only take a few seconds at the time: following the file naming convention established during the planning phase, recording how a dataset was processed, saving files in the agreed location, or updating documentation after making changes to the data. Individually they may seem insignificant, but together they build a record of the research that makes it easier to continue working, collaborate with others, and understand what was done weeks, months or even years later. The best time to document a decision is usually when you make it, not weeks later when you are trying to remember why you made it.

        Research projects rarely remain unchanged. New people may join, methods may evolve, additional data may be collected, or unexpected challenges may require changes to the original plan. Research data management should therefore be viewed as an ongoing process rather than a fixed set of rules established at the beginning of the project. As the research evolves, documentation and planning should evolve with it.

      • Preserve and Share What Matters

        As a project comes to an end, the focus shifts from supporting day-to-day research to deciding what should be kept for the future. Not every file created during a project needs to be preserved. Temporary files, intermediate analyses and duplicate copies can often be discarded once they are no longer needed. Instead, the aim is to identify the research data and documentation that have long-term value.

        Preparing research data for long-term preservation is more than simply storing a copy somewhere safe. The data should remain easy to find, accessible to the right people, understandable without relying on memory, and reusable in the future. Future users may include your future self, your collaborators, reviewers, other researchers, or anyone who needs to verify or build upon your work. Preserving research data therefore means preserving not only the files themselves, but also the information needed to understand and interpret them. Achieving this often involves organising the final version of the data, completing the documentation, choosing appropriate file formats, selecting a trusted repository, and deciding how the data will be accessed.

        Many of these ideas are captured by the FAIR Principles, which provide a framework for making research data easier to find, access, understand and reuse over time. Another guide explains the FAIR principles in more detail. For now, it is enough to remember that preserving research data means ensuring they remain useful long after the project has finished.

        Where appropriate, research data may also be shared with others. Sharing does not necessarily mean making data openly available to everyone. Depending on ethical, legal, contractual or commercial considerations, data may be shared openly, under controlled access, or only with specific collaborators. The important point is to choose an approach that is appropriate for both the research and the data involved.

      • Typical Considerations

        Managing research data throughout a project typically involves:

        Planning how research data will be managed, shared and preserved during and after the project

        Consideration Example Why it matters
        Research data Identify the types of research data that will be collected, created or reused during the project. Helps anticipate storage, documentation, ethical and preservation requirements.
        Roles and responsibilities Decide who is responsible for data management tasks such as documentation, backups, quality checks and maintaining the Data Management Plan. Prevents important tasks from being overlooked or assumed to be someone else's responsibility.
        Resources and budget Consider whether the project requires additional storage, software licences, repository fees, specialist staff (e.g. data stewards), training, or other RDM-related resources. Check whether these costs are eligible and should be included in the project budget. Ensures the project has the resources needed to manage, preserve and, where appropriate, share research data throughout its lifetime.
        Storage strategy Depending on the project, research data may be stored or accessed in multiple locations, such as an institutional server, secure cloud storage, a high-performance computing (HPC) system, or a local computer used for analysis. Prevents confusion about where the latest version of the data is stored, identifies where copies exist, and helps ensure the data remain secure and accessible throughout the project.
        Access to data Decide who can view, edit or delete files, and whether access differs between team members. Protects sensitive information and reduces the risk of accidental changes or data loss.
        Personal and sensitive data Consider whether personal or sensitive data will be collected, how they will be protected, stored and shared, or by when they will be deleted. Ensures meeting ethical, legal and institutional requirements before data collection begins.
        Collaboration and exchanging How will collaborators exchange files or work on shared data? For example, using institutional platforms or secure cloud storage instead of email or personal cloud accounts. Helps collaborators work efficiently while reducing security and compliance risks.
        Data Management Plan (DMP) Document the project's data management decisions in a document and update it as the project evolves. Provides a shared reference for the research team and records both the decisions made and the reasoning behind them.

        Managing research data during the project

        Consideration Example Why it matters
        File naming Use a consistent way of naming files, e.g. Interview_Group-A_001_2026-08-06.xlsx instead of Interview group A patient 1.xlsx. Makes files easier to identify, find and share within the research team.
        Folder structure Organise files using the agreed folder structure and keep it consistent throughout the project. Helps find information quickly and avoids duplicate or misplaced files.
        Version control Use a simple versioning system (e.g. v1, v2) or more advanced version control tools such as Git. Prevents accidental overwriting and makes it possible to recover or compare previous versions.
        Backups Make use of the automatic institutional backups where available, or establish a regular backup routine. Protects against hardware failure, accidental deletion and cyberattacks.
        Documentation Record how data were collected, processed and analysed using README files, codebooks or protocols. Makes the data understandable to everyone and to your future self.
        Quality assurance Based on the schedule set in the DMP, regularly check for missing values, inconsistencies or entry errors during the project. Helps identify problems early, before they affect analyses or become difficult to correct.
        Data Management Plan (DMP) Update the DMP when significant changes occur during the project. Ensures the documented plan continues to reflect how research data are actually being managed.

        Preserving and sharing research data during and after the project

        Consideration Example Why it matters
        Selecting what to preserve Decide which datasets, documentation, code and other research outputs have long-term value. Avoids preserving unnecessary material while ensuring valuable research outputs are retained.
        Retention requirements Consider your field's best practices, and institutional, funder or legal requirements for how long data should be retained. Helps ensure compliance while planning for long-term stewardship.
        Preparing data for preservation Remove duplicate files, organise the final version of the data and complete the documentation. Makes the preserved dataset easier to understand and reuse in the future.
        File formats Choose sustainable, widely supported formats where appropriate (e.g. CSV, TXT, PDF/A). Reduces the risk that files become difficult to open over time.
        Selecting a repository Deposit data in an appropriate trusted repository rather than relying on personal websites or "available upon request". Improves long-term preservation, discoverability and access.
        Selecting a Licence Choose an appropriate licence describing how others may access, use, modify and share your research data (e.g. CC0). Clarifies what others are allowed to do with your research data and encourages appropriate reuse.
        Repository information (metadata) When depositing research data, complete as much of the requested information as possible (e.g. title, authors, description, keywords, funding information and licences). Helps others discover, understand and correctly interpret the data.
        Access conditions Decide whether data will be openly available, shared under controlled access, or restricted. Balances openness with ethical, legal and contractual obligations.
        Persistent identifiers Obtain a DOI or another persistent identifier for the dataset where appropriate. Allows datasets to be cited reliably and found in the future.
      • Common Misconceptions

        A Data Management Plan is just an administrative requirement.
        FALSE: A Data Management Plan helps transform assumptions into agreed decisions, records why choices were made, and keeps collaborators aligned as the project evolves.


        I'll organise my files when I have time.
        REALITY: Small habits established throughout the project are much easier than trying to reorganise everything at the end.


        I'll remember how I processed the data.
        REALITY: Documentation written while you work is far more reliable than relying on memory weeks or months later.


        Everything should be preserved forever.
        FALSE: Not all research data need to be kept. The goal is to preserve the research outputs that have long-term value, together with enough documentation to understand and reuse them.


        'Available upon request' is a sufficient preservation strategy.
        FALSE: Research has shown that data kept only by individual researchers often become unavailable over time. Trusted repositories provide much more reliable long-term preservation.

      • Key Takeaways

        • Research data management is a continuous process that spans planning, managing, preserving and sharing research data.
        • Many RDM decisions only take a few minutes, but they can prevent significant problems later in a project.
        • Writing decisions down helps transform assumptions into shared ways of working.
        • Good documentation supports both collaborators and your future self.
        • Preserving research data means keeping them understandable and usable, not simply storing a copy somewhere.
      • References

        Kovács, M., Hoekstra, R., & Aczél, B. (2021). The Role of Human Fallibility in Psychological Research: A Survey of Mistakes in Data Management. Advances in Methods and Practices in Psychological Science, 4. https://doi.org/10.1177/25152459211045930. Licence: CC BY-NC 4.0

        Perrier, L., Blondal, E., Ayala, A. P., Dearborn, D., Kenny, T., Lightfoot, D., Reka, R., Thuna, M., Trimble, L., & MacDonald, H. (2017). Research data management in academic institutions: A scoping review. PLoS ONE, 12. https://doi.org/10.1371/journal.pone.0178261. Licence: CC BY 4.0

        Specht, A., Bolton, M., Kingsford, B., Specht, R., & Belbin, L. (2018). A story of data won, data lost and data re-found: the realities of ecological data preservation. Biodiversity Data Journal. https://doi.org/10.3897/bdj.6.e28073. Licence: CC BY 4.0

        Töwe, M., & Barillari, C. (2020). Who does what?–Research data management at ETH Zurich. Data Science Journal19(1), 36-36. https://doi.org/10.5334/dsj-2020-036. Licence: CC BY 4.0

        Vines, T., Albert, A. Y., Andrew, R., Debarre, F., Bock, D. G., Franklin, M. T., Gilbert, K., Moore, J.-S., Renaut, S., & Rennison, D. (2013). The availability of research data declines rapidly with article age. Current biology : CB, 24 1, 94-97 . https://doi.org/10.1016/j.cub.2013.11.014. Licence: © 2014 Elsevier Ltd. All rights reserved

        Yu, F., Deuble, R., & Morgan, H. (2017). Designing research data management services based on the research lifecycle–a consultative leadership approach. Journal of the Australian Library and Information Association66(3), 287-298. https://doi.org/10.1080/24750158.2017.1364835. Licence: © 2017 Australian Library & Information Association

        Zingmark, M., & Iwarsson, S. (2025). Case study on challenges in research with public partners: A personal data incident during recruitment for a survey study on ageing and housing. BMC Research Notes, 18. https://doi.org/10.1186/s13104-025-07246-8. Licence: CC BY 4.0

      DOI

      Author: Jonathan England

      Created: August 2026

      Guides for Researchers on RDM

      • Data formats for preservation

      • How to deal with non-digital data

      • How to deal with sensitive data

      • How to find a trustworthy repository for your data

      • How to make your data FAIR

      • Raw data, backup and versioning

      Still have questions?

      Contact us via our Helpdesk.
      We try to respond within 48 hours.

      Continue reading