Skip to main content
Guides for Researchers

What Changes When My Research Involves Sensitive Data?

Balancing Openness and Protection

Why Sensitive Data Require Additional Care

Some research projects may involve information that requires additional care. This may include personal data about participants, such as interview recordings, survey responses, medical records or photographs. However, sensitivity is not limited to personal data. Research data may also be considered sensitive because their disclosure could harm communities, reveal the location of endangered species, compromise cultural heritage sites, expose commercially valuable information, or create security risks.

However, the presence of sensitive information does not prevent good research data management. Sensitive data still need to be organised, documented, stored, preserved, and shared appropriately. The difference is that additional safeguards may be needed to protect the interests of the individuals, populations, organisations or resources involved.

Examples:

  • Clinical trial datasets containing patient health information are often shared through controlled-access repositories, allowing access only to authorised researchers rather than the general public.
  • Research data may be kept confidential temporarily to protect patent opportunities or meet obligations to commercial partners before being shared more widely.
  • The precise locations of endangered species are often generalised or withheld to reduce the risk of harm, while the remaining data remain available for research and conservation purposes.
  • Research data relating to Indigenous communities may be shared under community-defined conditions rather than made openly available, ensuring that decisions about access and reuse remain aligned with Indigenous rights, interests and cultural protocols.

We present in more details these examples in the section "Not Everything Must Be Open" of our guide about Open Data.

This chapter approaches sensitive data through the same three stages introduced in previous guides: planning before data collection, managing data during the project, and preparing data for preservation and sharing. Rather than focusing on legal terminology, the emphasis is on the practical decisions researchers may need to make throughout the research process.

A final point is worth emphasising from the outset: sensitive data do not have to be openly available to be valuable. Even when access must be restricted, research data can still be organised, documented and preserved in ways that make them Findable, Accessible, Interoperable and Reusable (FAIR). This distinction between FAIR and Open is discussed further in the chapters Making Sense of the FAIR Principles and Should My Research Data Be Open?.

Planning Before Collecting Data

Many challenges associated with sensitive data can be avoided through early planning. Before data collection begins, it is often easier to decide what information will be collected, who will have access to it, how it will be protected, and whether it may be shared in the future. Once data collection is underway, changing these decisions can become more difficult.

Planning is particularly important because decisions made at the beginning of a project often determine what options remain available later. For example, if human participants are not informed that their data may be shared for future research, it may become impossible to share or reuse the data beyond the original project. Similarly, if identifiable information is collected unnecessarily, additional obligations and risks may be introduced that could otherwise have been avoided.

Example: A Data Transfer That Triggered a GDPR Complaint

Researchers in Sweden received personal data for around 23,000 residents from a housing company to recruit participants for a survey on ageing and housing. Although the study had ethical approval and formal agreements were in place, individuals had not consented to this transfer of their personal data.

A GDPR complaint was submitted within days. The incident had to be reported to the Swedish privacy authority, all affected individuals were informed, and the research team spent four months responding to around 600 emails. Recruitment was delayed while a new process was developed (Zingmark & Iwarsson, 2025).

Many of these considerations overlap with the broader planning activities discussed in What Do I Need to Think About Before, During and After My Research?. However, projects involving personal or sensitive information often require additional attention to privacy, confidentiality and future access arrangements.

Before collecting data, it can be useful to consider the following questions:

Question Why it matters
Do I need to collect identifiable information at all? Collecting only the information necessary for the research reduces risks and simplifies data management.
What information should participants receive about the research? Participants should understand how their data will be collected, used, stored, preserved and potentially shared.
Will personal or sensitive information need to be shared with project partners? Data-sharing arrangements may need to be agreed before data collection begins.
Who will have access to identifiable information? Access should be limited to people who genuinely need it for the project.
Can identifying information be separated from the research data? Separating identifiers from research data can reduce risks during storage, analysis and sharing.
How long will the data need to be kept? Retention requirements may come from funders, institutions, legislation, ethics approvals or disciplinary practices.
What are the plans for preservation and sharing after the project? Decisions made now may determine whether the data can later be shared openly, shared under controlled access, or preserved for future research.

The 3Rs

Depending on your research field, you might already be familiar with the 3Rs in animal research ethics: Replacement, Reduction and Refinement (Russell & Burch, 1959). These principles encourage researchers to replace animals where possible, reduce the number used, and refine procedures to minimise harm.

A similar mindset can be applied to working with sensitive data. Before collecting data, it is worth considering whether all the planned information is genuinely necessary, whether identifiable information can be avoided or reduced, and whether data collection procedures can be designed to minimise risks for participants.

For example, do you need participants' names, or would a participant ID be sufficient? Do you need exact dates of birth, or would age ranges answer the research question equally well? Could identifiable information be collected separately from the research data? Asking these questions before data collection begins can reduce risks while still meeting the objectives of the research.

A common misconception is that considerations about data sharing only become relevant at the end of a project. In practice, many opportunities and constraints are established before the first data are collected. Thinking about future preservation and sharing from the outset does not mean that all data will eventually be made openly available. Rather, it helps ensure that appropriate options remain available when the project reaches its conclusion.

Considerations During the Project

Once data collection begins, the focus shifts from planning to implementation. Decisions about access, storage, documentation and security now become part of the day-to-day management of the research. Many of the same practices discussed in 'Where Should My Research Data Be Stored During The Project?' apply here, but the presence of personal or sensitive information often requires additional safeguards.

In practice, managing sensitive data is rarely about a single security measure. Rather, it involves a combination of decisions that reduce the likelihood of accidental disclosure, unauthorised access, data loss or misuse. Small actions, such as storing files in the agreed location, restricting access to authorised team members, or separating identifiable information from research data, can significantly reduce risks over the lifetime of a project.

Projects also evolve. New collaborators may join, responsibilities may change, and additional data may be collected. It is therefore useful to review the arrangements established during the planning phase periodically and confirm that they still reflect how the project is operating in practice.

The following considerations commonly arise when managing personal or sensitive data during a project:

Consideration Example Why it matters
Storage location Store identifiable information on approved institutional infrastructure rather than personal devices or personal cloud services. Reduces the risk of unauthorised access and helps comply with institutional requirements.
Access control Limit access to identifiable information to researchers who genuinely need it. Reduces the risk of accidental disclosure or inappropriate use.
Separation of identifiers Store participant names and contact details separately from research data whenever possible. Limits the consequences of a potential data breach and facilitates future sharing.
Data transfer Use secure institutional transfer tools rather than email attachments or personal file-sharing services. Protects sensitive information during exchange with collaborators.
Backups Ensure sensitive data are included in institutional or project backup procedures. Protects against accidental deletion, hardware failure and cyberattacks.
Documentation Record how data were collected, processed and modified throughout the project. Helps maintain transparency and supports future interpretation of the data.
Retention and deletion Periodically review whether identifiable information still needs to be retained. Reduces unnecessary risks and supports compliance with agreed retention policies.

When working with personal data, it is also good practice to collect and retain only the information that remains necessary for the project. Over time, some identifiable information may no longer be required, while other data may be suitable for anonymisation or preparation for future sharing. Considering these possibilities throughout the project often makes the final preservation and sharing stage much easier.

Preparing Data for Sharing

As a project comes to an end, researchers often face an important question: what, if anything, can be shared? When personal or sensitive information is involved, the answer is not always straightforward. However, sensitive data should not automatically be excluded from preservation or sharing.

In many cases, it is possible to reduce risks while still making research outputs available for verification, reuse or future research. The appropriate approach depends on the nature of the data, the commitments made to participants, ethical and legal requirements, and the potential risks associated with disclosure.

Anonymisation

One option is to anonymise the data before sharing them.

Anonymised data have been processed so that individuals can no longer be identified, directly or indirectly, using reasonably available means. When data are truly anonymised, they are no longer considered personal data under the GDPR.

Depending on the nature of the research, anonymisation may involve removing direct identifiers such as names and email addresses, aggregating variables, generalising locations, or reducing the level of detail available in the dataset.

However, you should be aware that anonymisation can be difficult to achieve in practice, particularly for small datasets, rare populations, or datasets that could be linked with other information sources.

To support this process, OpenAIRE provides access to Amnesia, a free anonymisation tool that helps researchers assess and reduce disclosure risks before sharing datasets. A dedicated guide is available to help you getting started.

Pseudonymisation

In some cases, anonymisation is not possible or desirable.

Pseudonymisation replaces direct identifiers with codes while retaining a mechanism that could allow individuals to be re-identified if necessary. For example, participant names may be replaced with participant IDs, while a separate file links those IDs to the original identities.

Unlike anonymised data, pseudonymised data remain personal data under the GDPR because re-identification remains possible. As a result, additional safeguards and access restrictions are usually required.

Anonymised data Pseudonymised data
Individuals cannot reasonably be identified. Re-identification remains possible.
No longer considered personal data under the GDPR. Still considered personal data under the GDPR.
May often be shared openly. Usually requires controlled or restricted access.

Controlled and Restricted Access

Not all datasets can be shared openly, even after anonymisation efforts.

Examples may include health data, sensitive interview data, commercially valuable information, Indigenous knowledge, or the locations of endangered species. In these situations, access may need to be restricted to authorised users under specific conditions.

Rather than relying on statements such as "available upon request", it is generally preferable to deposit data in a trusted repository that can manage controlled access requests. This helps ensure that data remain discoverable and accessible through a transparent process, even after the original researchers have moved on.

Retention and Deletion

Preparing data for sharing also involves deciding what should be preserved, for how long, and what should be deleted.

In some projects, identifiable information may no longer be needed once data collection has ended and can be deleted according to the agreed retention policy. In others, legal, ethical or contractual obligations may require data to be retained for a specific period.

It is also possible for identifiable information to be deleted while anonymised research data are preserved and shared for future use. These decisions are often easier when retention requirements have been considered during the planning phase.

Deleting Sensitive Data Securely

When highly sensitive data need to be deleted, moving files to the recycle bin or trash folder is often insufficient because the data may still be recoverable. Depending on the sensitivity of the information, institutional requirements and the risks associated with the data, researchers may need to apply data sanitisation techniques to make the information inaccessible.

It can involve different approaches, ranging from software-based methods to the physical destruction of storage media. The appropriate method will depend on the type of storage, the sensitivity of the information and the level of effort that a potential attacker might invest in attempting to recover the data.

Technique Example
Cryptographic erasure Permanently deleting the encryption keys protecting an encrypted drive or dataset, making the data inaccessible.
Overwriting (data wiping) Replacing the original data with new data to reduce the possibility of recovery.
Physical destruction Shredding, crushing or otherwise destroying storage devices containing highly sensitive information.

Guidance on secure deletion and data sanitisation techniques is provided in the NIST Guidelines for Media Sanitization (Chandramouli & Hibbard, 2025)

Sensitive Data Can Still Be FAIR

A common misconception is that sensitive data cannot be FAIR because they cannot be openly shared.

In reality, the FAIR principles and Open Data are not the same thing. Research data can remain Findable, Accessible under clearly defined conditions, Interoperable and Reusable even when access is restricted. For example, a repository record may remain publicly visible while access to the underlying data is granted only to authorised users.

The goal is therefore not necessarily to make sensitive data open, but to make them as FAIR as possible while protecting the individuals, communities, organisations or resources involved.

Key Takeaways

  • Sensitive data are not limited to personal data. Research data may also be considered sensitive because they could harm individuals, communities, endangered species, cultural heritage, commercial interests or other protected resources if disclosed inappropriately.
  • Many challenges associated with sensitive data can be avoided through early planning. Decisions about consent, access, retention, preservation and future sharing are often easier to make before data collection begins.
  • Managing sensitive data during a project involves practical measures such as controlling access, using appropriate storage solutions, documenting decisions, protecting confidential information and regularly reviewing data-management arrangements.
  • Anonymised and pseudonymised data are not the same. Anonymised data no longer allow individuals to be identified and are no longer considered personal data under the GDPR, whereas pseudonymised data remain personal data and require appropriate safeguards.
  • Not all sensitive data can be shared openly. Depending on the nature of the data and the associated risks, access may need to be restricted, controlled or limited to specific users.
  • FAIR and Open are not the same thing. Sensitive data can still be Findable, Accessible under appropriate conditions, Interoperable and Reusable even when access to the data is restricted.
  • The goal is not necessarily to make sensitive data open, but to manage, preserve and share them responsibly while protecting the people, organisations and resources involved.

References

Foundational Documents

Chandramouli, R. and Hibbard, E.A. (2025). Guidelines for Media Sanitization, National Institute of Standards and Technology, Gaithersburg, MD, , NIST Special Publication (SP) NIST SP 800-88r2. https://doi.org/10.6028/NIST.SP.800-88r2

Russell, W.M.S. and Burch, R.L., (1959). The Principles of Humane Experimental Technique, Methuen, London. ISBN 0-900767-78-2. 
A digital version of the Principles may be accessed for free on the website of Johns Hopkins University's Center for Alternatives to Animal Testing (CAAT).

Research Evidence

Zingmark, M., & Iwarsson, S. (2025). Case study on challenges in research with public partners: A personal data incident during recruitment for a survey study on ageing and housing. BMC Research Notes, 18https://doi.org/10.1186/s13104-025-07246-8. Licence: CC BY 4.0

Research Infrastructures and Resources

Amnesia. OpenAIRE. https://amnesia.openaire.eu/

Amnesia eBook. OpenAIRE. https://catalogue.openaire.eu/ebooks/Amnesia_Ebook_OpenAIRE.pdf

DOI

Author: Jonathan England

Created: August 2026