Skip to main content
Guides for Researchers

How Do I Plan My Research Data Management?

Planning now to avoid problems later

What kind of data?

Many of the decisions you make during a research project depend on the type of data you are working with. Before deciding where data will be stored, how they will be documented, who should have access to them, or whether they can be shared, it is important to understand what data the project is likely to produce.

Identifying the Data in Your Project

As discussed in a previous guide, all research fields produce research data, although this may not always be immediately obvious. Identifying the types of data your project will generate is an important first step because it influences many of the decisions that follow. For example, a study based on interview recordings may require different storage arrangements and access controls than a project analysing publicly available datasets. A project generating large numbers of images or videos may require significantly more storage capacity than one based on spreadsheets or text files. Some types of data may need to be retained for many years, while others may have limited long-term value.

Thinking Beyond Raw Data

It can also be useful to think about the different stages of your data. The files collected directly from instruments, participants or observations are often not the same as the files used for analysis or included in publications. For example, a project may begin with video recordings of participant behaviour, from which observations, measurements or coded responses are extracted and analysed. Similarly, a sequencing instrument may produce raw output files that are later transformed into processed datasets used for downstream analyses. A project may therefore produce raw data, processed data, analysis scripts, visualisations and documentation, each of which may need to be managed differently.

Formats Matter Too

The formats of your data can also have practical implications. Some file formats require specialised software to open, process or analyse them. If collaborators do not have access to that software, sharing and collaborative work may become more difficult. Thinking about formats early can help identify potential barriers and, where appropriate, allow the use of more accessible or interoperable alternatives.


Questions to Consider

When planning your project, consider questions such as:

  • What data will be collected, created or reused?
  • What formats will the data be in?
  • Approximately how much data will be generated?
  • Will the data contain personal, sensitive or confidential information?
  • Will multiple versions of the data be created during processing and analysis?
  • Are there existing datasets, software or other materials that will be reused?

You do not need precise answers to all of these questions before the project begins. However, having a general understanding of the data you will be working with makes it easier to make informed decisions about storage, documentation, collaboration, sharing and long-term preservation later on.

Who manages what?

Why Responsibilities Matter

Many research projects involve more than one person. Even in small projects, tasks may be shared between researchers, students, technicians, data stewards, software developers or external collaborators. When responsibilities are not clearly discussed, it is easy to assume that someone else is taking care of an important task.

For example, a research team may assume that backups are being performed automatically, only to discover after a hardware failure that no backup system was in place. Similarly, documentation may be postponed because everyone assumes another team member is recording key decisions, data processing steps or methodological changes.

Identifying responsibilities at the start of a project can help avoid confusion later. This does not necessarily require assigning a different person to every task. In some projects, a single researcher may be responsible for most aspects of data management. In larger collaborations, responsibilities may be distributed across several team members.

Shared Responsibility versus Assigned Responsibilities

Research data are often produced collectively, and many team members may contribute to their collection, processing and analysis. In that sense, responsibility for the data may be shared across the project. However, shared ownership does not automatically mean that individual tasks have clear owners.

A research team may collectively value good documentation, regular backups and organised files, but unless specific responsibilities are assigned, these tasks can easily be overlooked. Clarifying who is responsible for particular activities does not reduce the collaborative nature of the project; it helps ensure that important work is completed consistently.

In practice, responsibility for the data may be shared across the project while specific tasks are assigned to individual team members. These may include maintaining documentation, managing access permissions, performing quality checks, updating the Data Management Plan, or making decisions about data sharing and preservation.

What Responsibilities Should Be Discussed?

The way responsibilities are assigned will depend on the size of the project, the composition of the team and local practices. The example below illustrates how responsibilities might be distributed within a research project. Here is an example of responsibilities.

Responsibility Example responsible person
Organising files and folders Anyone generating the data
Managing storage and backups Research technician, IT support team or designated researcher(s)
Recording documentation and metadata Anyone collecting or processing the data
Monitoring data quality Researcher(s) responsible for data collection or analysis
Managing access permissions Principal Investigator(s), project coordinator(s) or data owner(s)
Making decisions about data access, sharing and preservation Principal Investigator(s) or project lead(s)
Updating the Data Management Plan Project coordinator(s), principal Investigator(s) or designated team member(s)
Preparing data for sharing or preservation Researcher(s) responsible for the dataset, often with support from a data steward
Responding to requests for access to data after the project ends Principal Investigator(s), repository contact person or institution responsible for the data

What Happens When People Join or Leave?

Responsibilities should also be reviewed when team members join or leave the project. A departing researcher may have important knowledge about how data were collected, processed or organised. Without a plan for transferring this knowledge, valuable information can be lost even when the data themselves remain available.

This is particularly important for projects involving students, PhD and postdoctoral researchers or fixed-term staff. Important knowledge about the data may otherwise leave the project with the individual who generated or managed them.

In collaborative projects, it can be useful to document responsibilities explicitly. This does not need to be complicated. A short list of tasks and responsible individuals may be sufficient to ensure that everyone understands their role and knows who to contact when questions arise.


Questions to Consider

When discussing responsibilities, consider questions such as:

  • Who is responsible for organising and documenting the data?
  • Who manages storage and backups?
  • Who can grant or remove access to data?
  • Who updates the Data Management Plan?
  • What happens if a team member leaves the project?
  • Who will prepare the data for sharing or long-term preservation?

Taking time to clarify responsibilities at the beginning of a project can help prevent misunderstandings, reduce the risk of important tasks being overlooked and ensure that data remain usable throughout the research lifecycle.

Where will the data be stored?

One Project, Multiple Storage Locations

It is common for research data to exist in several locations at the same time. For example, raw data may be stored on an institutional server, working copies may be analysed on researchers' computers or using specialised infrastructure such as high-performance computing systems, while shared files may be accessed through a cloud platform.

This is not necessarily a problem. However, when multiple copies exist, it is important to understand their purpose and to know which version should be considered the primary version. Without clear arrangements, different team members may work on different copies of the same file, creating confusion and increasing the risk of errors.

When planning storage, it can be useful to identify:

  • Where data will be collected or created
  • Where they will be analysed
  • Where shared working copies will be stored
  • Whether multiple copies will exist
  • Which location will contain the authoritative version

Choosing Suitable Storage

The most appropriate storage solution will depend on the nature of the project and the data involved.

Factors that may influence storage decisions include:

  • The volume of data being generated
  • Whether the data contain personal or sensitive information
  • The number of people who require access
  • Whether collaborators are based at different institutions
  • Whether specialised computing infrastructure is required
  • Institutional policies and available services
  • How long the data need to remain accessible

For example, a project generating several terabytes of imaging data may require different storage arrangements than a project based on interview transcripts or survey responses.

Access and Security

Storage is not only about where data are kept but also about who can access them.

Some projects may require all team members to have access to the same files, while others may need different levels of access depending on roles and responsibilities. Projects involving personal, confidential or commercially sensitive information may require additional protections and restrictions.

Considering access requirements at the start of a project can help prevent situations where researchers cannot access the data they need or where sensitive information is made available more broadly than intended.

Planning for Backups

No storage solution is immune to accidental deletion, hardware failure, cyberattacks or human error. Understanding how data will be backed up is therefore just as important as deciding where they will be stored.

Researchers should not assume that backups are being performed automatically. Different storage services provide different levels of protection, and responsibilities for backups may vary between institutions and platforms.

Even if backup procedures are discussed in more detail later, it is useful to establish from the beginning:

  • Whether backups exist
  • How frequently they are created
  • Who is responsible for them
  • How data can be recovered if something goes wrong

Questions to Consider

When planning where data will be stored, consider questions such as:

  • Where will data be collected, stored and analysed?
  • Will the project use multiple storage locations?
  • Which copy should be considered the authoritative version?
  • Who needs access to the data?
  • Will the data contain personal, sensitive or confidential information?
  • Are institutional storage services available?
  • How will the data be backed up?
  • Who is responsible for managing storage and backups?
  • What would happen if the primary copy of the data became unavailable tomorrow?

Taking time to plan storage arrangements before data collection begins can help ensure that data remain accessible, secure and usable throughout the project.

How will we collaborate?

Before data collection begins, it is worth agreeing not only where data will be stored, but also how they will move between researchers, systems and institutions throughout the project. Many data management problems arise not because researchers lack the right tools, but because collaborators work in different ways. Files may be stored in different locations, multiple versions of the same document may be created, or team members may use platforms that others cannot access.

Agreeing on how data will be shared and managed before the project begins can help avoid these problems. Investing time early to establish a workflow can make collaboration more efficient and reduce the risk of confusion, duplicated work or accidental data loss.

Choosing Collaboration Tools

One of the first decisions is how collaborators will access and share data.

Researchers often use familiar tools such as Google Drive, Dropbox, OneDrive or WeTransfer because they are convenient and easy to use. However, the suitability of these tools depends on the nature of the project and the data involved.

For example, personal accounts may not meet institutional requirements or data protection obligations (e.g. GDPR), particularly when personal, sensitive or confidential information is involved. In contrast, many institutions provide approved services such as institutional Google Workspace, Microsoft 365, Nextcloud or other collaboration platforms that may offer stronger security, access management and contractual protections.

When selecting collaboration tools, consider:

  • Whether the tool is approved or supported by your institution
  • Whether it is suitable for the type of data being handled
  • Whether collaborators from other institutions can access it
  • How access permissions are managed
  • What happens to files if a collaborator leaves the project

The goal is not necessarily to avoid commercial tools, but to make an informed decision about which solution best fits the project's needs.


In addition to institutional services, researchers in Europe may also have access to services provided through the European Open Science Cloud (EOSC). The EOSC EU Node provides collaboration services such as file synchronisation and sharing, which can help research teams work together across institutions while maintaining a shared environment for project files.

EOSC EU Node service Typical use
File Sync & Share Sharing and synchronising project files between collaborators
Large File Transfer Secure transfer of large files that may be impractical to exchange through email or standard cloud services
Bulk Data Transfer Transfer of very large datasets between storage systems or institutions

Sharing Files or Working Together?

Not all collaborative activities require the same type of workflow.

In some situations, researchers simply need access to the latest version of a file. In others, several people may need to work on the same document or dataset at the same time.

These requirements may seem similar, but they often require different approaches. For example, sharing a report with collaborators for review is very different from having multiple researchers simultaneously entering observations into a spreadsheet.

Before selecting tools or workflows, consider:

  • Do collaborators only need to access the latest version of a file?
  • Will multiple people edit the same file?
  • Will several people enter data at the same time?
  • How will changes be tracked and reviewed?

Answering these questions early can help avoid version conflicts and reduce the risk of data being accidentally overwritten.

Designing a Data Collection Workflow

In some cases, direct access to a dataset may not be the most effective way to collaborate.

For example, allowing multiple researchers to edit the same spreadsheet can increase the risk of accidental deletions, inconsistent data entry or unintended modifications. As projects grow, these risks often become more difficult to manage.

Instead, it may be useful to separate data entry from data storage. Researchers might enter information through an online form, survey platform or database interface, while the underlying dataset remains protected from accidental changes.

For larger projects involving structured data and many contributors, a database may provide a more robust alternative to shared spreadsheets. Managed database services, such as Database as a Service (DBaaS) available through EOSC EU Node, allow researchers to use databases without managing the underlying infrastructure.

The most appropriate workflow will depend on the project, but it is worth considering how data will move between collection, storage and analysis before data collection begins.


Example: Collecting Data from Multiple Researchers

A research project involves ten researchers collecting observations from different locations. The initial plan is for everyone to enter data directly into a shared spreadsheet.

However, the team identifies several potential problems:

  • Rows may be accidentally deleted or modified.
  • Different researchers may use different formats for dates or measurements.
  • It may be difficult to identify who entered or modified a particular record.
  • Simultaneous editing may create conflicts.

Instead, the team decides to use an online form. Researchers enter observations through the form, which automatically populates a central dataset. The project coordinator reviews new entries regularly, while the original data remain protected from accidental modification.

Although this requires more planning at the start of the project, it reduces errors and provides a clearer record of how the data were collected.

Sharing Large Datasets

For projects generating large volumes of data, standard file-sharing approaches may not be sufficient. Uploading and downloading large datasets repeatedly can be slow and inefficient, particularly when collaborators are distributed across different institutions or countries.

Depending on the workflow, it may be preferable to provide collaborators with direct access to shared storage or to use dedicated transfer services. For example, the EOSC EU Node provides services specifically designed for transferring large files and high-volume datasets between researchers and institutions.

When planning collaboration workflows, consider:

  • How much data will need to be exchanged?
  • How frequently will data be transferred?
  • Do collaborators need local copies of the data?
  • Is remote access to a shared storage system possible?
  • Are there storage or transfer limits imposed by the chosen platform?

Planning for Change

Research teams rarely remain unchanged throughout a project. New collaborators may join, students may graduate and researchers may move to different institutions.

When selecting collaboration tools and workflows, consider how easily new team members can be added and how access can be removed when someone leaves. Important files, documentation and project knowledge should not depend entirely on a single individual's personal account or computer.

Planning for these changes from the start can help ensure that data and documentation remain accessible throughout the lifetime of the project.


Questions to Consider

When planning how collaborators will work together, consider questions such as:

  • Which platform will be used to share project files?
  • Is the chosen platform appropriate for the type of data being handled?
  • Can all collaborators access the necessary tools and services?
  • Do researchers need access to files, or do they need to work on the same files simultaneously?
  • How will changes and new versions be managed?
  • How will data be collected and entered into the project?
  • Who manages access permissions?
  • What happens to files and accounts when a team member leaves the project?

Taking time to establish clear collaboration workflows before the project begins can help reduce confusion, improve efficiency and support consistent management of research data throughout the project.

What resources?

Many aspects of research data management require resources. These may include storage space, software, specialised infrastructure, training, staff time or expert support. While some of these resources are available through institutions or research infrastructures, others may need to be planned and budgeted for before a project begins.

It is often easier to identify resource needs at the start of a project than to find additional funding, storage capacity or technical support once data collection is already underway.

Storage and Computing Resources

The type and volume of data produced by a project can have a significant impact on resource requirements.

Projects generating large volumes of images, video, sequencing data or simulation outputs may require substantial storage capacity and specialised computing infrastructure. Other projects may have relatively modest storage requirements but require secure environments for handling sensitive information.

Questions to consider include:

  • How much data is expected to be generated?
  • Will storage requirements increase over time?
  • Is specialised computing infrastructure required?
  • Are secure environments needed for sensitive data?
  • Are institutional or disciplinary services available?

Software and Tools

Research projects often depend on software for data collection, processing, analysis and collaboration.

Some tools may be freely available, while others require licences or subscriptions. In collaborative projects, it is also important to consider whether all team members have access to the software needed to work with the data.

For example, collaborators may encounter difficulties if a dataset can only be opened or analysed using specialised software that is unavailable to part of the research team.

Expertise and Support

Not every project requires specialist support, but some may benefit from expertise that is not available within the immediate research team.

Depending on the project, support may be available from:

  • Data stewards
  • Librarians
  • Research software engineers
  • IT services
  • Legal or data protection specialists
  • Ethics committees
  • Research infrastructures

Planning for Preservation and Sharing

Resource needs do not necessarily end when data collection or analysis is complete.

Preparing data for long-term preservation or sharing may require additional effort, including organising files, improving documentation, selecting repositories and reviewing data for ethical or legal considerations.

Some repositories may also charge fees for storage or preservation services.


Questions to Consider

When planning project resources, consider questions such as:

  • What storage and computing resources will be required?
  • Does the project require specialised software or licences?
  • Do all collaborators have access to the necessary tools?
  • Is additional training needed?
  • Will specialist support be required?
  • Are there costs associated with preservation or sharing?
  • Have sufficient time and effort been allocated to managing the data throughout the project?

Thinking about resource requirements before the project begins can help ensure that the people, infrastructure and support needed to manage the data are available when they are needed.

What constraints?

Researchers often begin a project with ideas about how data will be collected, analysed, shared or preserved. However, these decisions are not always entirely within the research team's control.

Legal requirements, contractual obligations, institutional policies and technical limitations can all influence how data are managed throughout a project. Identifying these constraints before data collection begins can help avoid situations where researchers need to redesign workflows, restrict access to data or abandon plans for sharing data later in the project.

Personal and Sensitive Information

One of the most common constraints concerns data relating to people.

Projects involving personal information, health data, interviews, survey responses or other potentially sensitive information may require additional safeguards. These may affect how data are collected, stored, accessed, shared and preserved.

For example, researchers may need to:

  • Limit access to authorised individuals during the project
  • Store data in approved environments
  • Remove identifying information before sharing
  • Define retention periods
  • Obtain clear and appropriate consent from participants.
  • Consider whether the consent process and participant information allow for future uses of the data, such as reuse in follow-up studies, data sharing or long-term preservation.

Contractual and Funder Requirements

Research projects may also be subject to requirements imposed by funders, institutions, partner organisations or contracts.

For example, agreements may specify:

  • How long data must be retained
  • Whether data should be shared
  • Whether access restrictions apply
  • Which storage services may be used
  • Which repository should be used
  • Who owns particular datasets or outputs

Existing Data and Third-Party Materials

Many projects reuse existing datasets, software, images, documents or other research materials.

However, access to these materials does not necessarily imply permission to redistribute, modify or share them. Licences, terms of use and intellectual property rights may influence what researchers are allowed to do.

This can be particularly important if the project plans to combine existing datasets with newly collected data or make derivative datasets available to others.

Technical and Practical Constraints

Not all constraints are legal or ethical. Projects may also face practical limitations related to infrastructure, software, storage capacity, computing resources or available expertise.

For example:

  • A dataset may be too large to transfer easily.
  • Collaborators may not have access to the same software.
  • Institutional systems may impose storage limits.
  • Uploading on a repository may be limited to a certain size.
  • Long-term preservation may only be possible for selected outputs.
  • Specialised software may only be available to part of the research team.

Constraints Can Change

Constraints identified at the start of a project may change over time.

New collaborators may join the project, institutional policies may evolve, additional datasets may be incorporated, or new legal and ethical considerations may emerge.

For this reason, it can be useful to review assumptions periodically rather than treating them as fixed for the duration of the project.


Questions to Consider

When planning a project, consider questions such as:

  • Will personal or sensitive information be collected?
  • Are there restrictions on who can access the data?
  • Do funders, institutions or partners impose specific requirements?
  • Are existing datasets or third-party materials being reused?
  • Are there licensing, consent or intellectual property considerations?
  • Could the data be reused in future studies or projects?
  • Do technical limitations affect how the data can be managed?
  • Could any of these constraints affect plans for sharing or preserving the data?

Identifying constraints early does not necessarily mean a project becomes more complicated. In many cases, it simply helps researchers design workflows that are realistic, compliant and aligned with the needs of the research.

How and why record decisions?

Why Record Decisions?

Many of the topics discussed in this chapter involve making decisions before data collection begins. For example, the research team may agree where data will be stored, who is responsible for backups, which collaboration tools will be used, or how sensitive information will be protected.

These decisions are often discussed during meetings, shared in email threads or agreed informally between collaborators. However, if they are not recorded in a central place, they can easily be forgotten, misunderstood, unknown to some, or applied inconsistently as the project progresses.

Recording decisions provides a shared reference for the research team and helps ensure that everyone is working according to the same understanding. It can also make it easier for new team members to understand how data are being managed and why particular choices were made.

What Should Be Recorded?

The goal is not to document every detail of the project. Instead, it can be useful to record decisions that are likely to affect how data are managed during the project.

Examples include:

  • The types of data that will be collected, created or reused
  • Roles and responsibilities
  • Storage locations and backup arrangements
  • Collaboration tools and workflows
  • Requirements relating to personal or sensitive data
  • Institutional or funder requirements
  • Assumptions, constraints and planned approaches

The level of detail will depend on the size and complexity of the project.

A Data Management Plan

Many research projects record these decisions in a Data Management Plan (DMP).

A DMP is a document that describes how research data will be managed during and after a project. Depending on the funder, institution or discipline, a DMP may be required or recommended.

A DMP should not be viewed (only) as an administrative exercise completed at the beginning of a project and then forgotten. Its main value is that it provides a central place to record decisions, assumptions and responsibilities, and can be updated as the project evolves.

Keeping Information Up to Date

Research projects rarely proceed exactly as planned. New collaborators may join, storage locations may change, additional datasets may be incorporated, or new requirements may emerge.

For this reason, recorded decisions should be reviewed periodically and updated when significant changes occur.

The objective is not to create perfect documentation but to ensure that important information remains available when it is needed.


Questions to Consider

  • Where will project decisions be recorded?
  • Who is responsible for maintaining this information?
  • Will a Data Management Plan be created?
  • Who in the institution can help me with my DMP?
  • How will changes be documented?
  • How often should decisions be reviewed?
  • How will new team members access this information?

Recording decisions does not eliminate the need for discussion and review. However, it provides a shared reference point that can help the project remain organised and consistent as it develops.

Key Takeaways

  • Understand the data before collecting them.
  • Agree who is responsible for what.
  • Decide where data will be stored and how collaborators will work together.
  • Identify the resources and constraints that may affect the project.
  • Record important decisions so that they are not lost or misunderstood.
  • Review and update these decisions as the project evolves.

DOI

Author: Jonathan England

Created: August 2026