Skip to main content

Guides

Guides for Researchers

How to comply with Horizon Europe mandate for research data

BACK TO GUIDES

  • What are the requirements?

    Proper Research Data Management (RDM) is a mandatory Open Science practice in Horizon Europe projects that generate or reuse digital research data. Beneficiaries must manage these data responsibly, in line with the FAIR principles (Findable, Accessible, Interoperable, Reusable).

    At a minimum, beneficiaries must:

    • Establish by month 6 a Data Management Plan (DMP), and keep it updated throughout the project, covering the data generated and/or reused in the project.
    • Deposit the data in a ‘trusted’ and ‘metadata-ready’ repository as soon as possible or within the deadlines set out in the DMP. When required by the call conditions, use a repository federated in the European Open Science Cloud (EOSC).
    • Ensure Open Access to the deposited data via the repository as soon as possible or within the timeline set in the DMP, under a CC0 or CC BY licence (or equivalent), following the principle “as open as possible, as closed as necessary”. This means that the data access can be restricted to safeguard legitimate interests and constraints. If open access is not provided to some or all data, the reasons must be explained and justified in the DMP.
    • Provide open metadata for deposited datasets (CC0 or equivalent,) and provide information via the repository about any related research outputs and any other tools or instruments needed to reuse or validate the data.
    • Provide information (via the same repository) about any research output or any other tools and instruments needed to re-use or validate the data.

    Keep in mind that ‘research data’ is a very broad concept and certainly not limited to numerical and tabular data.

  • The FAIR principles

    • Horizon Europe emphasises managing digital research data, and where relevant other research outputs, in line with the FAIR principles, which means making them Findable, Accessible, Interoperable and Reusable. Through appropriate metadata, identifiers, standards and clear access/reuse conditions, they can be reused by both humans and machines.
    • Depositing data in a FAIR-enabling, trusted and metadata-ready repository is a key step towards FAIRness, as it supports persistent identifiers, rich metadata, standard access mechanisms and preservation. FAIR does not necessarily mean open: access may be restricted where justified, provided that metadata remain openly accessible and conditions are clearly described within the DMP.
    • There is no one-size-fits-all approach to FAIR. What is appropriate and feasible largely depends on the research domain, data types involved, as well as on the specificities of the project (e.g. sensitivities, constraints, project context), and should be documented and reviewed in the DMP.

    FAIRdataprinciples foster

    Image: https://book.fosteropenscience.eu/

    • Developing a Data Management Plan (DMP)

      Developing a Data Management Plan (DMP) is a core step in complying with Research Data Management (RDM) requirements in Horizon Europe.

      A DMP describes how digital research data (and, where relevant, other research outputs) will be managed both during and after a research project. It identifies key actions and strategies to ensure that research data are of a high-quality, properly documented, secure, preserved, sustainable, and where possible, accessible and reusable. It also outlines how the FAIR principles align with any applicable legal, ethical, or commercial constraints.

      DMP timeline

      • Proposal stage: a formal DMP is not required but a short (typically one-page) DMP-like document is required as part of the proposal. Note that in cases of a public emergency, and where required by the Work Programme, a full DMP may be required at proposal submission or by grant signature.
      • Project implementation: a full initial version of the DMP must be submitted as a project deliverable, normally by month 6.
      • During the project: the DMP is a living document and must be regularly updated to reflect changes in data generation, reuse, access conditions, repositories, or decisions taken. For projects lasting longer than 12 months, at least one updated version must be submitted as a deliverable.
      • End of the project: a final DMP must be delivered, describing how the data have been managed, preserved, and shared.

      Whenever possible, beneficiaries are encouraged to make DMPs publicly available as non-restricted deliverables, under open access and a CC BY licence, to support reuse. This may be done by uploading the DMP on a trusted repository (e.g. Zenodo), which will make it available on the CORDIS website.

      What to include in a DMP?

      Horizon Europe provides a recommended DMP template, which beneficiaries are encouraged, but not required, to use.

      To support both proposal preparation and DMP development, beneficiaries may use ARGOS, OpenAIRE’s free and open-source online tool for creating and publishing DMPs. ARGOS offers templates aligned with Horizon Europe requirements, together with contextual guidance for each section.

    • Providing access to research data in trusted repositories

      Trusted repositories

      Trusted repositories are infrastructures that provide reliable and long-term access to digital resources such as data, publications, etc. Usually, these repositories go through assessment or certification processes to guarantee that certain quality criteria are met.

      In Horizon Europe, the following are considered trusted repositories:

      • Certified repositories, for example those with CoreTrustSeal, Nestor Seal DIN31644, or ISO16363 certifications.  
      • Domain-specific repositories that are internationally recognised, commonly used and endorsed by the research communities relevant to your project.
      • General-purpose repositories or institutional repositories that, without official certification, present the essential characteristics of trusted repositories (e.g. security provisions, services for the creation of machine actionable metadata, long-term preservation of data, etc.).


      You can find more information on how to find a trustworthy repository for your data on our dedicated page.

      When should data be deposited?

      In Horizon Europe, data should be deposited as soon as possible after its generation and, at the latest, by the end of the project.

      There are, however, some additional requirements:

      • Data underpinning a scientific publication should be deposited at the latest at the time of publication, and in line with standard community practices.  
      • In cases of public emergency, if requested by the granting authority, immediate open access should be provided. If exceptions to open access to data apply, data should be made available at least to the legal entities that need the data to address the public emergency.
      • In exceptional cases, data can be deposited after the project has finished.

      'As open as possible, as closed as necessary'

      In Horizon Europe, research data should be made open access by default and licensed under the latest version of CC BY (attribution required) or CC0 (public domain), or equivalent.

      However, it is recognized that data should be ‘as open as possible, as closed as necessary’, and exceptions can be made when providing open access to data:

      • Is against the beneficiary’s legitimate interests, including regarding commercial exploitation;
      • Is contrary to any other constraints, such as data protection rules, privacy, confidentiality, trade secrets, Union competitive interests, security rules, intellectual property rights or;  
      • Would be against other obligations under the Grant Agreement. 

      In such cases, data can be kept restricted, closed or under embargo, but beneficiaries must explain in the DMP the legitimate exception(s) under which they choose to restrict access to (some of the) research data. 

      Find more information on How do I know if my research data is protected? 

      Metadata compliance

      Horizon Europe requires that, when you deposit data in a trusted repository, this should be described with rich metadata in line with the FAIR principles. In practice, this means that you should choose a repository that has a number of mandatory and recommended metadata fields. 

      Metadata should:

      • At least include the following fields: author(s), dataset description or abstract, date of dataset deposit or publication date, dataset deposit venue, dataset license (CC0 or CC BY by default) and dataset embargo period (if any).
      • Also include information about Horizon Europe or Euratom funding: grant project name, acronym and number. Ideally, the repository will have dedicated fields for this information. If not, you can include them in other appropriate fields, such as the abstract.
      • Be open access under a CC0 licence or equivalent. This is also recommended in cases where data must be closed or restricted but there are no compelling reasons for metadata not to be findable and accessible.

      Most (trusted) repositories will require that you fill in a metadata form about the data or files that you will publish, which should cover these requirements.

      For additional information, you can check the page on metadata.


      IMPORTANT!

      The European Research Council (ERC) requested in 2023 an analysis of the ‘metadata readiness’ of repositories (Lazzeri 2024). Only five repositories (<intR>²Dok, DANS, HAL, and Zenodo), met the “Essential” readiness level requirements. Many other repositories were at a “Close-to-Essential” level, and were making modifications to add the missing metadata fields. 

      In practical terms, you need to check directly with your repository whether they have all the mandatory metadata fields required by the European Commission. If not, one work around is to deposit your data in Zenodo (for HE compliance) and a subject-specific or your institutional repository (to improve visibility and discoverability, and a better community engagement and networking).

    • Validation (and re-use) requirements

      Horizon Europe requires that, at the time of depositing research data in a trusted repository, you must also provide access (via the repository) to information about any research output or any other tools or instruments needed to re-use or validate the research data.

      • Research outputs, tools and instruments” may include data, software, algorithms, protocols, models, workflows, electronic notebooks and others.
      • What type of information should be provided? A detailed description of the research output/tool/instrument, how to access it, any dependencies on commercial products (e.g. software), potential version/type, potential parameters, etc.
      • Besides information about these outputs, beneficiaries are encouraged to provide open (digital) access to the research outputs, tools and instruments themselves unless legitimate interests or constraints apply.

      Typically, the repository will allow you to provide links to relevant information (e.g. links to related publications, to a code repository, to related datasets, etc.).

    • Costs of Research Data Management

      Research Data Management activities need to be costed into research, in terms of the time and resources needed. By planning early, costs can be significantly reduced. Under Horizon Europe, eligible costs are governed by the Grant Agreement rules (Article 6.2 and general eligibility principles); they must be actual, incurred by the beneficiary, necessary for the action, recorded in accounts, and incurred during the action (with limited exceptions e.g. final reporting). They must be foreseen in the project budget/Description of Action and incurred during the duration of the project.

      To estimate costs for RDM, you can check the online RDM-costing tool and the infographic on ‘What will it cost t o manage and share my data?’.

      • Webinar recording

        We organise this webinar every 3-4 months. We update every time with the latest recommendations and information.

        Watch video
      • How can OpenAIRE help?

        The following additional support materials can help you with the RDM requirements in Horizon Europe projects:

         

        argos logo newOpenAIRE also offers tools for research data management: ARGOS is an OpenAIRE service that simplifies the management, validation, monitoring and maintenance of DMPs. [tool]

        Links and further information

        The following sources were used and contain more extensive information on how to address open science in Horizon Europe proposals:

         

        Links and further information

        The following sources were used and contain more extensive information about Research Data Management, and how to address Open Science in Horizon Europe proposals:

      • Glossary

      Created: June 2022

      Updated: August 2026

      View our other guides
      [question]
      Still have questions?

      Contact Us

      Continue reading

      This Glossary provides the definitions of and practical advice regarding the key terms mentioned in the context of RDM requirements in Horizon Europe. Other terms relating to RDM can be found in the From Science Europe Data Glossary.


       

       

      Anonymisation

      Anonymisation is the process of removing personally identifiable information (information that directly or indirectly relates to an identified or identifiable person) from datasets containing sensitive data. As a result, data subject is no longer identifiable. As opposed to pseudonymisation, anonymisation is not reversible, which means that the re-identification of the data subject is not possible.

      Practical advice:

      • OpenAIRE has developed a tool that can be used for anonymisation Amnesia.

       

       

      Backup

      Data backup is a process of creating a copy of data in a digital format and storing it on another device to ensure that data are saved and to prevent data loss.
      Backups can be full (all files are backed up whenever a backup is made) or partial (only a part of the files, e.g. new files, are backed up).

      Practical advice:

      • One backup should be at a physically separate location.
      • Backups should be made from the master copy.
      • The backup location should be as secure as the master copy location.

       

       

      Controlled vocabularies and ontologies

      Controlled vocabulary is an organised and standardised arrangement of predefined terms (words and phrases) that are used to index content in an information system with the aim of facilitating information retrieval. Controlled vocabularies connect variant terms and synonyms for concepts, link concepts in a logical order and organise them into categories, so as to provide a consistent way to describe data. They can be general and discipline-specific, and can take the form of subject heading lists, thesauri, authority files, taxonomies and alphanumeric classification schemes.

      Ontologies are not controlled vocabularies, but they use controlled vocabularies to establish a formal specification of a conceptual model in which concepts and categories of concepts, properties, relationships among concepts and categories, functions, constraints, and axioms are defined.

      Practical advice:

      • Use standard domain specific controlled vocabularies and/or ontologies wherever possible to better align your outputs with similar data in your field of research. 
      • Use well-documented vocabularies in which terms are assigned persistent identifiers
      • Use repositories that enable you to add terms from controlled vocabularies.
      • A useful domain agnostic resource for finding controlled vocabularies and ontologies can be found here.
      • Check also other resources.

       

       

      Data Management Plan

      Data Management Plan (DMP) is a formal document that outlines how data will be handled throughout the research data lifecycle – from planning, through collecting, analysing, publishing, preserving, to sharing and reusing.

      In Horizon Europe, DMPs are mandatory and a template to guide the preparation of DMPs is provided. The list of points to be addressed includes data types and formats, compliance with the FAIR principles, (metadata, repositories, controlled vocabularies, licences, etc.),  legal requirements (intellectual property rights, GDPR), costs of preservation, data security and ethics, retention periods.

      Online tools that can facilitate the preparation of DMPs are available:

      Practical advice:

      • A DMP is a living document, which should be updated as the project develops. Any deviations from the original proposal can be documented and explained.
      • If possible, make the DMP publicly available.

       

       

      Documentation

      Data documentation includes various types of information that can help find, assess, understand/interpret, and (re)use research data – e.g. information about methods, protocols, datasets to be used and data files, preliminary findings, etc. Documentation helps understand the context in which data were created, as well as the structure and the content of data. Data should be documented through all stages of the research data lifecycle. Detailed and rich documentation ensures reproducibility and upholds research integrity. Documentation also includes metadata.

      Practical advice:

      • Various tools, such as e-lab notebooks, are available to support you in the process of  creating documentation

       

       

      File format

      File format is a standard way of encoding information so that it can be stored in a computer file. Digital research data may be stored in a wide variety of file formats, depending on the devices and tools used in data collection and processing.

      File formats may be proprietary (the encoding-scheme is designed and owned by a company or organisation, and is not published, due to which files can be opened only by those who have particular software or hardware tools) and/or prone to obsolescence (legacy formats, bit rot).

      To ensure that users can access and understand data and that data can be preserved in the long term, use open (defined by an openly published specification that anyone can use) and lossless formats (ensuring that no data or quality loss will occur during file manipulation).

      Practical advice:

      Check also this OpenAIRE Guide.


       

       

      File naming convention

      File naming convention is a framework for generating file names that have a consistent structure, while describing the content of files and their relations to other files.

      Practical advice:

      • Use standardised syntax for file naming to aid better searchability and the ability to perform batch processes. These can take the form of YYMMDD_filename, and it is recommended to use version suffixes wherever possible when creating versions of a file from a master copy. These could, for example, be generated after cleaning or analysis steps.
      • Define the file naming convention in an early stage of your research and apply it consistently throughout the research data lifecycle.

       

       

      Licence

      Licence is a written agreement by means of which the copyright holder defines the rights granted to the users. In a digital environment, standardised licences based on a set of predefined reuse conditions, such as Creative Commons, are used. Licences for digital objects are machine readable. 

      In Horizon Europe, CC BY or CC0 (or equivalent) open licence is required for data in open access, while metadata deposited data must be open under the CC0 or equivalent licence.

      Practical advice:

      • When depositing your data in a repository, consult the repository’s licence policy to check whether it is compliant with the Horizon Europe requirements.
      • If you are combining already published data into a new dataset, check the compatibility of data licences.

       

       

      Metadata

      Metadata are data that provide information about other data, e.g. a description of the content of the data, the date when the data were produced or collected, tools and devices used to obtain data, file formats and sizes, the names of the people who created or collected data, relevant persistent identifiers, etc. Metadata should be created and provided in accordance with commonly used metadata standards, which may be general or discipline specific. This ensures that metadata can be understood by humans and processed and exchanged by machines. 

      Practical advice:

      • When preparing a DMP, you will be required to mention the metadata standards you will “follow to make your data interoperable”. A useful resource for finding available metadata standards can be found here. Choose a standard that is commonly used in your discipline. Do not invent your own!
      • Some metadata are automatically generated (by devices used to create and capture data) and embedded in data files, e.g. in digital audio and video recordings,while some have to be produced manually. 
      • When depositing your data in a repository, you will be guided by the user interface through the process of providing metadata. In the input form, some metadata fields are mandatory, which means that the procedure cannot be completed if this information is not provided. It is highly recommended to provide as detailed information as possible even in non-mandatory fields.

       

       

      Persistent identifier (PID)

      Persistent identifier (PID) is a long-lasting reference to a resource that provides the information required to reliably identify, verify and locate the resource. In a digital environment, PIDs have the form of URLs. When pasted in a browser, they take users to the resource.

      Apart from digital resources, PIDs can also relate to researchers (e.g. ORCID, ISNI), institutions (e.g. ROR), grants, instruments and devices, etc. In this case, a PID leads to the record describing a researcher, an institution, etc. in the relevant registry.

      Examples of PIDs include DOIs, ORCIDs, ISBN, Handles, etc. 

      Practical advice

      • Apply PIDs to your research outputs. PIDs are typically automatically assigned by trustworthy repositories once your data are deposited there.  
      • It is also recommended to register yourself with ORCID to get a PID for you as an individual.

       

       

      Pseudonymisation

      Pseudonymisation (pseudo anonymisation) is the processing of personal data in such a way that the data can no longer be related to the data subject without the use of additional information. However, the additional information must be kept separately and subject to technical and organisational measures to ensure that data subjects remain unidentifiable. As opposed to anonymisation, pseudonymisation is a reversible process, which means that data subjects can be re-identified if access to the additional information is enabled.

      Practical advice

      • OpenAIRE has developed a tool that can be used for pseudonymisation - Amnesia.

       

       

      Repository

      Repository is a digital platform that ingests, stores, manages, preserves, and provides access to digital content. A repository should support a commonly accepted metadata standard and have a protocol enabling metadata exchange. 

      Repositories are usually classified into: 

      • subject/disciplinary, 
      • institutional and 
      • generalist repositories. 

      In Horizon Europe research data should be deposited in a trusted repository, i.e. in a repository that operates in accordance with relevant standards and best practice, provides long-term access and preservation, and ensures compliance with the FAIR principles. Trusted repositories include certified or community-recognised repositories, stable institutional repositories and generalist repositories (such as Zenodo).

      Practical advice:

      • Use the re3data repository registry to identify appropriate repositories.
      • The use of domain specific repositories is the most desirable. Such repositories will employ domain specific metadata standards and controlled vocabularies and/or ontologies which will enhance the ability to do analyses across similar datasets and even across domains.
      • Institutional and generalist repositories should only be considered if no domain specific repository can be found or if there is a mandate to deposit in a specific repository. If it is necessary to deposit data in multiple repositories, this should be done so as to maintain the same persistent identifier across those multiple copies.

       

       

      Sensitive data

      Sensitive data is information that should be protected against unauthorised disclosure because unauthorised access may negatively affect the privacy of an individual, trade and business secrets or even security. In the context of research, sensitive data usually include personally identifiable information (names, date and place of birth, place of living, employment information, etc.), health information,  and other private or confidential data.

      Practical advice:

      • Even in case of sensitive data, it is still possible to adhere to FAIR principles and open science by making the metadata freely available, while not enabling public access to the underlying data. These data can be managed through access control mechanisms and anonymisation and pseudonymisation.

       

       

      Storage

      Data storage is a computing technology that enables saving data in a digital format on computer components and recording media, including cloud services.

      In the context of Research Data Management, it is necessary to ensure that data are stored securely until the end of the project and throughout the minimum retention period. Storage options may include:

      • Portable devices,
      • University network drives,
      • Cloud services.

      Practical advice

      • Storage on local storage spaces on hard drives, pen drives, etc. is discouraged because these devices are vulnerable and data loss may occur. If portable devices are used, it should be ensured that there are copies on networked drives and backup storage.
      • Using in-house or institutionally approved spaces is recommended, especially if regular backups are enabled. 
      • Cloud services are suitable for collaboration with partners from partner institutions. However, it should be checked whether the selected cloud service makes regular backups and whether it falls under European jurisdiction.

      The 2019 PSI – Open Data Directive

      A checklist for IP, Research Data and Open Science

      The 2019 PSI – Open Data Directive represents the latest major upgrade of the PSI legislation in the EU. It amends previous PSI directives which first introduced a regulatory framework for Public Sector Information establishing the principle of reuse by default. With the 2019 revision Public Sector Bodies covered by the Directive can now benefit from a more comprehensive set of rules that regulate the reuse of data and information that they hold. The PSI – Open Data Directive poses an even stronger requirement on reuse by default, expands the type of PSB covered and, very importantly for Open Science, establishes the principle that research data resulting from publicly funded research must be Open Access by default.

      What is the PSI?

      PSI Stands for Public Sector Information and has been a directive since 2003 in order to regulate use of public sector information and to encourage as much public sector to be openly available. In 2017 it changed to the Open Data Directive which provides a common legal framework for government held data.


      • What is the PSI?

        PSI Stands for Public Sector Information and has been a directive since 2003 in order to regulate use of public sector information and to encourage as much public sector to be openly available. In 2017 it changed to the Open Data Directive which provides a common legal framework for government held data.

        The 2019 PSI – Open Data Directive represents the latest major upgrade of the PSI legislation in the EU. It amends previous PSI directives which first introduced a regulatory framework for Public Sector Information establishing the principle of reuse by default. With the 2019 revision Public Sector Bodies covered by the Directive can now benefit from a more comprehensive set of rules that regulate the reuse of data and information that they hold. The PSI – Open Data Directive poses an even stronger requirement on reuse by default, expands the type of PSB covered and, very importantly for Open Science, establishes the principle that research data resulting from publicly funded research must be Open Access by default.

      • I. Overview

        The PSI – Open Data Directive (Directive EU 2019/1024 of the European Parliament and of the Council of 20 June 2019 on open data and the re-use of public sector information), which entered into force on 16 July 2019 and gives Member States 24 months to transpose it into domestic legislation, constitutes the latest major upgrade of the PSI legislation in the EU. It amends Directives 2003/98/EC and 2013/37/EU which represented respectively the first substantial intervention by the EU into the field of Public Sector Information and its significant 10-year update which further reinforced the principle of reuse by default. With the 2019 revision, which is based on principles such as transparency and fair competition in the internal market, Public Sector Bodies (PSB) covered by the Directive can now benefit from a more comprehensive legal framework that regulates the reuse of data and information that they hold. The PSI – Open Data Directive poses an even stronger requirement of reuse by default, expands the type of PSB covered and, very importantly for Open Science, establishes the principle that research data resulting from publicly funded research must be open access by default.

        There are a number of specific provisions that may impact, in certain cases significantly, Open Science practices in the EU once transposed into domestic law by Member States (MS). The following is a list of the most relevant.

        1. The scope of the directive covers documents held by Public Sector Bodies (PSB) in Member States at national, regional and local levels

        This includes national governments, ministries, state agencies and municipalities, as well as organisations funded mostly by or under the control of public authorities (e.g. meteorological institutes). Since the 2013 update, museums, libraries, including university libraries and archives are likewise included, although special rules apply. With the 2019 reform the scope has been further expanded to certain public undertakings under specific rules.

        2. The Directive regulates the reuse of documents held by PSB

        Documents are defined as any representation of acts, facts or information — and any compilation of such acts, facts or information — whatever its medium (paper, or electronic form or as a sound, visual or audiovisual recording). The definition of ‘document’ is not intended to cover computer programmes, however, Member States (MS) may extend the application of this Directive to computer programmes

        3. PSI legislation regulates the reuse of documents held by PSB and establishes the rule that such documents shall be reusable for commercial and non commercial purposes (Art. 1)

        It should be noted that access to information, i.e. the principal way in which the covered type of information is made available, is not directly regulated by PSI, which nevertheless identifies The Charter of Fundamental Rights of the European Union (see Recital 5 PSI-OD) as the legal basis for national access to information rules and also establishes that MS shall encourage public sector bodies and public undertakings to produce and make available documents falling within the scope of the Directive in accordance with the principle of ‘open by design and by default’.

        Access to information legislation is regulated at the MS level. EU institutions have their own procedures to access their documents. Interestingly, Art. 1(6) 2019 PSI – Open Data Directive establishes that the Sui Generis Database Right (SGDR) of Directive 96/9/EC shall not be exercised by public sector bodies in order to prevent the re-use of documents or to restrict re-use beyond the limits set by the 2019 Directive.

        4. Research data resulting from public funding

        Member States will be asked to develop policies for open access to publicly funded research data. This is perhaps the most important innovation brought by the 2019 PSI-OD amendment for Open Science. It implies a mandatory open access status for all research data produced with public funding (Recital 27).

        5. Provision on “high value datasets”

        Another important aspect is the provision on “high value datasets” defined as documents the re-use of which is associated with important benefits for the society and economy, which will be governed by a separate set of rules. Thematically, these datasets identify geospatial, earth observation and environment, meteorological, statistics, companies and company ownership, and mobility as high value areas. The EC will identify the relevant dataset in 2021.

        6. Transparency requirements

        Stronger transparency requirements for public–private agreements involving public sector information, avoiding exclusive arrangements.

      • II. Focus on Research Data and Open Science

        In particular, regarding point 4) above, it should be noted how the 2019 Directive clarifies that under the national open access policies, publicly funded research data should be made open as the default option. However, this “open by default” rule should apply only to research data that have already been made publicly available by researchers, research performing organisations or research funding organisations through an institutional or subject-based repository. This should not impose extra costs for the retrieval of the datasets or require additional curation of data.

        The Directive also clarifies that MS may extend the application of the Directive to research data made publicly available through other data infrastructures than repositories, through open access publications, as an attached file to an article, a data paper or a paper in a data journal. Documents other than research data (e.g. scholarly articles, publications, etc) should continue to be exempt from the scope of this Directive.

        Additionally, concerns in relation to privacy, protection of personal data, confidentiality, national security, legitimate commercial interests, such as trade secrets, and to intellectual property rights of third parties should be duly taken into account, according to the principle ‘as open as possible, as closed as necessary’. Moreover, research data which are excluded from access on grounds of national security, defence or public security should not be covered by this Directive (see Recital 28).

      • III. The importance of national implementations and the development of national open access policies

        These are very important provisions in order to achieve one of the most important goals of Open Science, namely the accessibility, reusability and verifiability of scientific results. Access to the data that brought to certain result is a key element for the achievement of Open Science. The Directive is quite clear in terms of the basic rules, however, these rules will need to be implemented by MS in what the Directive calls national open access policies.

        It is fundamental that these national open access policies follow a common and coordinated approach in order to avoid potential fragmentation at the MS level. Space for intervention at the MS level could be found in areas such as:

        • The inclusion of software in the definition of documents of PSBs covered by the Directive;
        • The extension to other data infrastructures such as open access publications, as an attached file to an article, a data paper or a paper in a data journal as the type of first publication that will trigger the open by default rule;
        • National OA policies should contain a clear identification of a licence or licence type in order to avoid potential issues connected with licence identification, compatibility and maintenance;
        • In this sense EC Decision of 22/02/2019 “adopting Creative Commons as an open licence under the European Commission’s reuse policy” C(2019) 1655 final, which establishes CC BY 4.0 as the European Commission default licence for reuse policy and CC0 for raw data, metadata or other documents of comparable nature, while not binding for MS, indicates a clear and strong orientation for the convergence towards a common standard licence as the default option. MS should be encouraged to follow this common pattern for the generality of cases and only diverge when special circumstances justify a different approach.

      Factsheet

      DOI

      Still have questions?

      Contact us via our Helpdesk.
      We try to respond within 48 hours.

      Continue reading

      Guides for Researchers

      How can identifiers improve the dissemination of your research outputs?

      Connect all your research products with your person identifier

      Guides for Researchers on RDM

      • How to deal with non-digital data

      • How to deal with sensitive data

      • How to find a trustworthy repository for your data

      • How to make your data FAIR

      • Raw data, backup and versioning

      Still have questions?

      Contact us via our Helpdesk.
      We try to respond within 48 hours.

      Continue reading

      Guides for Researchers

      What Changes When My Research Involves Sensitive Data?

      Balancing Openness and Protection

      BACK TO GUIDES

      • Why Sensitive Data Require Additional Care

        Some research projects may involve information that requires additional care. This may include personal data about participants, such as interview recordings, survey responses, medical records or photographs. However, sensitivity is not limited to personal data. Research data may also be considered sensitive because their disclosure could harm communities, reveal the location of endangered species, compromise cultural heritage sites, expose commercially valuable information, or create security risks.

        However, the presence of sensitive information does not prevent good research data management. Sensitive data still need to be organised, documented, stored, preserved, and shared appropriately. The difference is that additional safeguards may be needed to protect the interests of the individuals, populations, organisations or resources involved.

        Examples:

        • Clinical trial datasets containing patient health information are often shared through controlled-access repositories, allowing access only to authorised researchers rather than the general public.
        • Research data may be kept confidential temporarily to protect patent opportunities or meet obligations to commercial partners before being shared more widely.
        • The precise locations of endangered species are often generalised or withheld to reduce the risk of harm, while the remaining data remain available for research and conservation purposes.
        • Research data relating to Indigenous communities may be shared under community-defined conditions rather than made openly available, ensuring that decisions about access and reuse remain aligned with Indigenous rights, interests and cultural protocols.

        We present in more details these examples in the section "Not Everything Must Be Open" of our guide about Open Data.

        This chapter approaches sensitive data through the same three stages introduced in previous guides: planning before data collection, managing data during the project, and preparing data for preservation and sharing. Rather than focusing on legal terminology, the emphasis is on the practical decisions researchers may need to make throughout the research process.

        A final point is worth emphasising from the outset: sensitive data do not have to be openly available to be valuable. Even when access must be restricted, research data can still be organised, documented and preserved in ways that make them Findable, Accessible, Interoperable and Reusable (FAIR). This distinction between FAIR and Open is discussed further in the chapters Making Sense of the FAIR Principles and Should My Research Data Be Open?.

      • Planning Before Collecting Data

        Many challenges associated with sensitive data can be avoided through early planning. Before data collection begins, it is often easier to decide what information will be collected, who will have access to it, how it will be protected, and whether it may be shared in the future. Once data collection is underway, changing these decisions can become more difficult.

        Planning is particularly important because decisions made at the beginning of a project often determine what options remain available later. For example, if human participants are not informed that their data may be shared for future research, it may become impossible to share or reuse the data beyond the original project. Similarly, if identifiable information is collected unnecessarily, additional obligations and risks may be introduced that could otherwise have been avoided.

        Example: A Data Transfer That Triggered a GDPR Complaint

        Researchers in Sweden received personal data for around 23,000 residents from a housing company to recruit participants for a survey on ageing and housing. Although the study had ethical approval and formal agreements were in place, individuals had not consented to this transfer of their personal data.

        A GDPR complaint was submitted within days. The incident had to be reported to the Swedish privacy authority, all affected individuals were informed, and the research team spent four months responding to around 600 emails. Recruitment was delayed while a new process was developed (Zingmark & Iwarsson, 2025).

        Many of these considerations overlap with the broader planning activities discussed in What Do I Need to Think About Before, During and After My Research?. However, projects involving personal or sensitive information often require additional attention to privacy, confidentiality and future access arrangements.

        Before collecting data, it can be useful to consider the following questions:

        Question Why it matters
        Do I need to collect identifiable information at all? Collecting only the information necessary for the research reduces risks and simplifies data management.
        What information should participants receive about the research? Participants should understand how their data will be collected, used, stored, preserved and potentially shared.
        Will personal or sensitive information need to be shared with project partners? Data-sharing arrangements may need to be agreed before data collection begins.
        Who will have access to identifiable information? Access should be limited to people who genuinely need it for the project.
        Can identifying information be separated from the research data? Separating identifiers from research data can reduce risks during storage, analysis and sharing.
        How long will the data need to be kept? Retention requirements may come from funders, institutions, legislation, ethics approvals or disciplinary practices.
        What are the plans for preservation and sharing after the project? Decisions made now may determine whether the data can later be shared openly, shared under controlled access, or preserved for future research.

        The 3Rs

        Depending on your research field, you might already be familiar with the 3Rs in animal research ethics: Replacement, Reduction and Refinement (Russell & Burch, 1959). These principles encourage researchers to replace animals where possible, reduce the number used, and refine procedures to minimise harm.

        A similar mindset can be applied to working with sensitive data. Before collecting data, it is worth considering whether all the planned information is genuinely necessary, whether identifiable information can be avoided or reduced, and whether data collection procedures can be designed to minimise risks for participants.

        For example, do you need participants' names, or would a participant ID be sufficient? Do you need exact dates of birth, or would age ranges answer the research question equally well? Could identifiable information be collected separately from the research data? Asking these questions before data collection begins can reduce risks while still meeting the objectives of the research.

        A common misconception is that considerations about data sharing only become relevant at the end of a project. In practice, many opportunities and constraints are established before the first data are collected. Thinking about future preservation and sharing from the outset does not mean that all data will eventually be made openly available. Rather, it helps ensure that appropriate options remain available when the project reaches its conclusion.

      • Considerations During the Project

        Once data collection begins, the focus shifts from planning to implementation. Decisions about access, storage, documentation and security now become part of the day-to-day management of the research. Many of the same practices discussed in 'Where Should My Research Data Be Stored During The Project?' apply here, but the presence of personal or sensitive information often requires additional safeguards.

        In practice, managing sensitive data is rarely about a single security measure. Rather, it involves a combination of decisions that reduce the likelihood of accidental disclosure, unauthorised access, data loss or misuse. Small actions, such as storing files in the agreed location, restricting access to authorised team members, or separating identifiable information from research data, can significantly reduce risks over the lifetime of a project.

        Projects also evolve. New collaborators may join, responsibilities may change, and additional data may be collected. It is therefore useful to review the arrangements established during the planning phase periodically and confirm that they still reflect how the project is operating in practice.

        The following considerations commonly arise when managing personal or sensitive data during a project:

        Consideration Example Why it matters
        Storage location Store identifiable information on approved institutional infrastructure rather than personal devices or personal cloud services. Reduces the risk of unauthorised access and helps comply with institutional requirements.
        Access control Limit access to identifiable information to researchers who genuinely need it. Reduces the risk of accidental disclosure or inappropriate use.
        Separation of identifiers Store participant names and contact details separately from research data whenever possible. Limits the consequences of a potential data breach and facilitates future sharing.
        Data transfer Use secure institutional transfer tools rather than email attachments or personal file-sharing services. Protects sensitive information during exchange with collaborators.
        Backups Ensure sensitive data are included in institutional or project backup procedures. Protects against accidental deletion, hardware failure and cyberattacks.
        Documentation Record how data were collected, processed and modified throughout the project. Helps maintain transparency and supports future interpretation of the data.
        Retention and deletion Periodically review whether identifiable information still needs to be retained. Reduces unnecessary risks and supports compliance with agreed retention policies.

        When working with personal data, it is also good practice to collect and retain only the information that remains necessary for the project. Over time, some identifiable information may no longer be required, while other data may be suitable for anonymisation or preparation for future sharing. Considering these possibilities throughout the project often makes the final preservation and sharing stage much easier.

      • Preparing Data for Sharing

        As a project comes to an end, researchers often face an important question: what, if anything, can be shared? When personal or sensitive information is involved, the answer is not always straightforward. However, sensitive data should not automatically be excluded from preservation or sharing.

        In many cases, it is possible to reduce risks while still making research outputs available for verification, reuse or future research. The appropriate approach depends on the nature of the data, the commitments made to participants, ethical and legal requirements, and the potential risks associated with disclosure.

        Anonymisation

        One option is to anonymise the data before sharing them.

        Anonymised data have been processed so that individuals can no longer be identified, directly or indirectly, using reasonably available means. When data are truly anonymised, they are no longer considered personal data under the GDPR.

        Depending on the nature of the research, anonymisation may involve removing direct identifiers such as names and email addresses, aggregating variables, generalising locations, or reducing the level of detail available in the dataset.

        However, you should be aware that anonymisation can be difficult to achieve in practice, particularly for small datasets, rare populations, or datasets that could be linked with other information sources.

        To support this process, OpenAIRE provides access to Amnesia, a free anonymisation tool that helps researchers assess and reduce disclosure risks before sharing datasets. A dedicated guide is available to help you getting started.

        Pseudonymisation

        In some cases, anonymisation is not possible or desirable.

        Pseudonymisation replaces direct identifiers with codes while retaining a mechanism that could allow individuals to be re-identified if necessary. For example, participant names may be replaced with participant IDs, while a separate file links those IDs to the original identities.

        Unlike anonymised data, pseudonymised data remain personal data under the GDPR because re-identification remains possible. As a result, additional safeguards and access restrictions are usually required.

        Anonymised data Pseudonymised data
        Individuals cannot reasonably be identified. Re-identification remains possible.
        No longer considered personal data under the GDPR. Still considered personal data under the GDPR.
        May often be shared openly. Usually requires controlled or restricted access.

        Controlled and Restricted Access

        Not all datasets can be shared openly, even after anonymisation efforts.

        Examples may include health data, sensitive interview data, commercially valuable information, Indigenous knowledge, or the locations of endangered species. In these situations, access may need to be restricted to authorised users under specific conditions.

        Rather than relying on statements such as "available upon request", it is generally preferable to deposit data in a trusted repository that can manage controlled access requests. This helps ensure that data remain discoverable and accessible through a transparent process, even after the original researchers have moved on.

        Retention and Deletion

        Preparing data for sharing also involves deciding what should be preserved, for how long, and what should be deleted.

        In some projects, identifiable information may no longer be needed once data collection has ended and can be deleted according to the agreed retention policy. In others, legal, ethical or contractual obligations may require data to be retained for a specific period.

        It is also possible for identifiable information to be deleted while anonymised research data are preserved and shared for future use. These decisions are often easier when retention requirements have been considered during the planning phase.

        Deleting Sensitive Data Securely

        When highly sensitive data need to be deleted, moving files to the recycle bin or trash folder is often insufficient because the data may still be recoverable. Depending on the sensitivity of the information, institutional requirements and the risks associated with the data, researchers may need to apply data sanitisation techniques to make the information inaccessible.

        It can involve different approaches, ranging from software-based methods to the physical destruction of storage media. The appropriate method will depend on the type of storage, the sensitivity of the information and the level of effort that a potential attacker might invest in attempting to recover the data.

        Technique Example
        Cryptographic erasure Permanently deleting the encryption keys protecting an encrypted drive or dataset, making the data inaccessible.
        Overwriting (data wiping) Replacing the original data with new data to reduce the possibility of recovery.
        Physical destruction Shredding, crushing or otherwise destroying storage devices containing highly sensitive information.

        Guidance on secure deletion and data sanitisation techniques is provided in the NIST Guidelines for Media Sanitization (Chandramouli & Hibbard, 2025)

      • Sensitive Data Can Still Be FAIR

        A common misconception is that sensitive data cannot be FAIR because they cannot be openly shared.

        In reality, the FAIR principles and Open Data are not the same thing. Research data can remain Findable, Accessible under clearly defined conditions, Interoperable and Reusable even when access is restricted. For example, a repository record may remain publicly visible while access to the underlying data is granted only to authorised users.

        The goal is therefore not necessarily to make sensitive data open, but to make them as FAIR as possible while protecting the individuals, communities, organisations or resources involved.

      • Key Takeaways

        • Sensitive data are not limited to personal data. Research data may also be considered sensitive because they could harm individuals, communities, endangered species, cultural heritage, commercial interests or other protected resources if disclosed inappropriately.
        • Many challenges associated with sensitive data can be avoided through early planning. Decisions about consent, access, retention, preservation and future sharing are often easier to make before data collection begins.
        • Managing sensitive data during a project involves practical measures such as controlling access, using appropriate storage solutions, documenting decisions, protecting confidential information and regularly reviewing data-management arrangements.
        • Anonymised and pseudonymised data are not the same. Anonymised data no longer allow individuals to be identified and are no longer considered personal data under the GDPR, whereas pseudonymised data remain personal data and require appropriate safeguards.
        • Not all sensitive data can be shared openly. Depending on the nature of the data and the associated risks, access may need to be restricted, controlled or limited to specific users.
        • FAIR and Open are not the same thing. Sensitive data can still be Findable, Accessible under appropriate conditions, Interoperable and Reusable even when access to the data is restricted.
        • The goal is not necessarily to make sensitive data open, but to manage, preserve and share them responsibly while protecting the people, organisations and resources involved.
      • References

        Foundational Documents

        Chandramouli, R. and Hibbard, E.A. (2025). Guidelines for Media Sanitization, National Institute of Standards and Technology, Gaithersburg, MD, , NIST Special Publication (SP) NIST SP 800-88r2. https://doi.org/10.6028/NIST.SP.800-88r2

        Russell, W.M.S. and Burch, R.L., (1959). The Principles of Humane Experimental Technique, Methuen, London. ISBN 0-900767-78-2. 
        A digital version of the Principles may be accessed for free on the website of Johns Hopkins University's Center for Alternatives to Animal Testing (CAAT).

        Research Evidence

        Zingmark, M., & Iwarsson, S. (2025). Case study on challenges in research with public partners: A personal data incident during recruitment for a survey study on ageing and housing. BMC Research Notes, 18https://doi.org/10.1186/s13104-025-07246-8. Licence: CC BY 4.0

        Research Infrastructures and Resources

        Amnesia. OpenAIRE. https://amnesia.openaire.eu/

        Amnesia eBook. OpenAIRE. https://catalogue.openaire.eu/ebooks/Amnesia_Ebook_OpenAIRE.pdf

      DOI

      Author: Jonathan England

      Created: August 2026

      View our other guides
      [question]
      Still have questions?

      Contact Us

      Continue reading