What Do I Need to Think About Before, During and After My Research?
Three stages
Research Data Management (RDM) refers to the organisation, documentation, storage, protection, sharing and preservation of research data throughout a research project.
The research literature consistently associates good RDM with improved data quality and integrity, reproducibility, collaboration, project management and risk reduction. Although much of the evidence comes from reviews and case studies rather than controlled intervention studies, these benefits are reported consistently across disciplines (Perrier et al., 2017).
Rather than being a single task completed at the end of a project, RDM begins before any data are collected and continues throughout the research process. Decisions made early in a project, such as how research data will be organised, documented, stored and protected, influence every subsequent stage of the research.
For simplicity, this guide groups research data management into three broad stages:
| Stage | When? | Typical activities |
|---|---|---|
| Plan | Before the project begins | Plan how research data will be collected, organised, documented and stored, and consider roles and responsibilities, ethical and legal requirements, sharing and long-term preservation. |
| Manage | Throughout the project | Organise files and folders, document research data, store and back up data securely, protect sensitive information, and support collaborative working. |
| Preserve | During and after the project ends | Select valuable research data for long-term preservation and, where appropriate, share them through trusted repositories to support verification and future reuse. |
Note: You may have seen other guides present research data management as a lifecycle with more stages. This guide adopts a simpler three-stage model: Plan, Manage and Preserve, to emphasise when key decisions are made throughout a project, while covering the same core RDM activities (Töwe & Barrilari, 2020; Yu et al., 2017).
When Things Go Wrong
Research Data Management isn't just about following good practices. It's also about avoiding problems that can compromise months or even years of research. The following examples are based on published research and real-world incidents. They illustrate how seemingly small decisions about planning, managing and preserving research data can have significant consequences later in a project.
A Data Transfer That Triggered a GDPR Complaint
Researchers in Sweden were conducting a survey on ageing and housing in collaboration with a housing company. To recruit participants, the company transferred personal data for around 23,000 people to the university research team so invitations could be sent. The study had ethical approval, formal data-sharing agreements, and the data were stored securely. However, the housing company had not obtained individuals' consent to disclose their personal data to the university for research purposes.
Within days, a formal GDPR complaint was submitted. The incident had to be reported to the Swedish Authority for Privacy Protection, all affected individuals were informed, and the research team spent four months responding to around 600 emails, many from people requesting that their personal data be deleted. Recruitment was delayed while a new recruitment process was designed (Zingmark & Iwarsson, 2025).
What could have helped?
- Planning how personal data would be transferred between project partners before recruitment.
- Clearly defining responsibilities for handling personal data.
- Ensuring participant information and consent covered all intended data transfers and uses.
RDM stage: Plan
Which File Is the Right One?
A research team is preparing the final analyses for publication. Several versions of the dataset exist, with names such as Final, Final_v2, and Final_FINAL. Nobody is completely sure which version contains the latest corrections. After days of analysis, the team discovers they have been working on the wrong file and must repeat part of the work.
This may sound familiar. In a survey of 488 psychology researchers, version-control errors were among the most commonly reported data management mistakes. Researchers also frequently reported ambiguous file naming and analysing the wrong version of a dataset. While many mistakes resulted only in lost time and frustration, around one in five respondents said their most serious data-management mistake had major or extreme consequences, including project failure or erroneous conclusions (Kovács et al., 2021).
What could have helped?
- Agree on a shared file naming convention.
- Use version control instead of creating multiple "final" files.
- Store files in one agreed location rather than personal folders.
- Keep the original raw data separate from processed data.
RDM stage: Manage
The Data Are Available Upon Request
A researcher wants to verify the findings of a published study. The paper states that the underlying data are "available from the corresponding author upon request." Years after the research was published, they send an email asking for the dataset. The email address no longer works. After tracking down one of the authors, they are told that they cannot remember where the data have been stored.
This is not an isolated case. A group of researchers contacted the authors of 516 published studies to request the original datasets. Overall, they managed to find a working email for 74% of papers, but were able to recover only 19% of the datasets, and for papers published before 2000, this fell to 11%. When authors explained why the data were unavailable, the most common reasons were that the files had been lost or were stored on inaccessible media. The study found that the likelihood of a dataset still existing declined by 17% for every year after publication, leading the authors to conclude that long-term preservation cannot rely solely on individual researchers (Vines et al., 2013).
What could have helped?
- Depositing research data in a trusted repository instead of relying on "available upon request".
- Preserving research data independently of individual researchers or personal storage.
- Planning for long-term preservation before the project ended.
- Including sufficient documentation so the data remain understandable over time.
RDM stage: Preserve
Rescuing a 20-Year Dataset
Researchers working with historic Australian vegetation data wanted to preserve decades of ecological observations for future research. Although much of the data still existed, recovering them proved to be a major challenge. Some documentation had been created using obsolete software and stored on floppy disks that could no longer be read using modern computers. Other information had to be reconstructed from paper records, obsolete file formats, and outdated spatial references. Recovering and documenting the datasets required years of work and an estimated AU$200,000 in funding before the data could finally be deposited in modern repositories (Specht et al., 2018).
What could have helped?
- Preserving research data in sustainable, widely supported file formats.
- Depositing data in a trusted repository before storage technologies became obsolete.
- Maintaining complete documentation alongside the data.
- Planning for long-term preservation throughout the project rather than years afterwards.
RDM stage: Preserve
Plan Before You Start
As a researcher, you would probably never write a project or grant proposal without first thinking carefully about your research questions, methods and planned analyses. Putting those ideas into writing often reveals assumptions, raises new questions and helps turn an initial idea into a coherent research project.
Research data should be approached in the same way.
Based on habits or past experience, you usually already have an idea of how you want to organise your files, where to access your data from, or share responsibilities within the team(s). However, ideas that seem obvious often become less straightforward when they need to be discussed with others or written down. Planning forces these decisions to be made explicitly rather than relying on assumptions.
For example, it may seem obvious that everyone will use the same folder structure, store data in the same location or document their work in the same way. However, unless these decisions are discussed and recorded, different team members often develop their own approaches. These differences may appear minor at first, but they can gradually lead to confusion, duplicated work or disagreements as the project progresses.
Planning is therefore less about predicting every aspect of the research than about making practical decisions while there is still time to think about them. Once data collection and analysis are underway, the focus is on the research itself, and it is rarely the ideal moment to redesign folder structures, decide where files should be stored, or clarify who is responsible for documenting the data.
Writing these decisions down, usually in the form of a Data Management Plan (DMP), also provides a shared reference for everyone involved. It records not only what was decided, but also why. As projects evolve, decisions can always be revisited and updated, but documenting them early helps keep everyone aligned and avoids having to rethink the same questions later.
Manage Throughout the Project
Once data collection begins, research data management becomes part of the research process itself. Rather than being a separate administrative task, it consists of a series of small decisions made throughout the project that help keep research organised, understandable and secure.
Many of these decisions only take a few seconds at the time: following the file naming convention established during the planning phase, recording how a dataset was processed, saving files in the agreed location, or updating documentation after making changes to the data. Individually they may seem insignificant, but together they build a record of the research that makes it easier to continue working, collaborate with others, and understand what was done weeks, months or even years later. The best time to document a decision is usually when you make it, not weeks later when you are trying to remember why you made it.
Research projects rarely remain unchanged. New people may join, methods may evolve, additional data may be collected, or unexpected challenges may require changes to the original plan. Research data management should therefore be viewed as an ongoing process rather than a fixed set of rules established at the beginning of the project. As the research evolves, documentation and planning should evolve with it.
Preserve and Share What Matters
As a project comes to an end, the focus shifts from supporting day-to-day research to deciding what should be kept for the future. Not every file created during a project needs to be preserved. Temporary files, intermediate analyses and duplicate copies can often be discarded once they are no longer needed. Instead, the aim is to identify the research data and documentation that have long-term value.
Preparing research data for long-term preservation is more than simply storing a copy somewhere safe. The data should remain easy to find, accessible to the right people, understandable without relying on memory, and reusable in the future. Future users may include your future self, your collaborators, reviewers, other researchers, or anyone who needs to verify or build upon your work. Preserving research data therefore means preserving not only the files themselves, but also the information needed to understand and interpret them. Achieving this often involves organising the final version of the data, completing the documentation, choosing appropriate file formats, selecting a trusted repository, and deciding how the data will be accessed.
Many of these ideas are captured by the FAIR Principles, which provide a framework for making research data easier to find, access, understand and reuse over time. Another guide explains the FAIR principles in more detail. For now, it is enough to remember that preserving research data means ensuring they remain useful long after the project has finished.
Where appropriate, research data may also be shared with others. Sharing does not necessarily mean making data openly available to everyone. Depending on ethical, legal, contractual or commercial considerations, data may be shared openly, under controlled access, or only with specific collaborators. The important point is to choose an approach that is appropriate for both the research and the data involved.
Typical Considerations
Managing research data throughout a project typically involves:
Planning how research data will be managed, shared and preserved during and after the project
| Consideration | Example | Why it matters |
|---|---|---|
| Research data | Identify the types of research data that will be collected, created or reused during the project. | Helps anticipate storage, documentation, ethical and preservation requirements. |
| Roles and responsibilities | Decide who is responsible for data management tasks such as documentation, backups, quality checks and maintaining the Data Management Plan. | Prevents important tasks from being overlooked or assumed to be someone else's responsibility. |
| Resources and budget | Consider whether the project requires additional storage, software licences, repository fees, specialist staff (e.g. data stewards), training, or other RDM-related resources. Check whether these costs are eligible and should be included in the project budget. | Ensures the project has the resources needed to manage, preserve and, where appropriate, share research data throughout its lifetime. |
| Storage strategy | Depending on the project, research data may be stored or accessed in multiple locations, such as an institutional server, secure cloud storage, a high-performance computing (HPC) system, or a local computer used for analysis. | Prevents confusion about where the latest version of the data is stored, identifies where copies exist, and helps ensure the data remain secure and accessible throughout the project. |
| Access to data | Decide who can view, edit or delete files, and whether access differs between team members. | Protects sensitive information and reduces the risk of accidental changes or data loss. |
| Personal and sensitive data | Consider whether personal or sensitive data will be collected, how they will be protected, stored and shared, or by when they will be deleted. | Ensures meeting ethical, legal and institutional requirements before data collection begins. |
| Collaboration and exchanging | How will collaborators exchange files or work on shared data? For example, using institutional platforms or secure cloud storage instead of email or personal cloud accounts. | Helps collaborators work efficiently while reducing security and compliance risks. |
| Data Management Plan (DMP) | Document the project's data management decisions in a document and update it as the project evolves. | Provides a shared reference for the research team and records both the decisions made and the reasoning behind them. |
Managing research data during the project
| Consideration | Example | Why it matters |
|---|---|---|
| File naming | Use a consistent way of naming files, e.g. Interview_Group-A_001_2026-08-06.xlsx instead of Interview group A patient 1.xlsx. |
Makes files easier to identify, find and share within the research team. |
| Folder structure | Organise files using the agreed folder structure and keep it consistent throughout the project. | Helps find information quickly and avoids duplicate or misplaced files. |
| Version control | Use a simple versioning system (e.g. v1, v2) or more advanced version control tools such as Git. | Prevents accidental overwriting and makes it possible to recover or compare previous versions. |
| Backups | Make use of the automatic institutional backups where available, or establish a regular backup routine. | Protects against hardware failure, accidental deletion and cyberattacks. |
| Documentation | Record how data were collected, processed and analysed using README files, codebooks or protocols. | Makes the data understandable to everyone and to your future self. |
| Quality assurance | Based on the schedule set in the DMP, regularly check for missing values, inconsistencies or entry errors during the project. | Helps identify problems early, before they affect analyses or become difficult to correct. |
| Data Management Plan (DMP) | Update the DMP when significant changes occur during the project. | Ensures the documented plan continues to reflect how research data are actually being managed. |
Preserving and sharing research data during and after the project
| Consideration | Example | Why it matters |
|---|---|---|
| Selecting what to preserve | Decide which datasets, documentation, code and other research outputs have long-term value. | Avoids preserving unnecessary material while ensuring valuable research outputs are retained. |
| Retention requirements | Consider your field's best practices, and institutional, funder or legal requirements for how long data should be retained. | Helps ensure compliance while planning for long-term stewardship. |
| Preparing data for preservation | Remove duplicate files, organise the final version of the data and complete the documentation. | Makes the preserved dataset easier to understand and reuse in the future. |
| File formats | Choose sustainable, widely supported formats where appropriate (e.g. CSV, TXT, PDF/A). | Reduces the risk that files become difficult to open over time. |
| Selecting a repository | Deposit data in an appropriate trusted repository rather than relying on personal websites or "available upon request". | Improves long-term preservation, discoverability and access. |
| Selecting a Licence | Choose an appropriate licence describing how others may access, use, modify and share your research data (e.g. CC0). | Clarifies what others are allowed to do with your research data and encourages appropriate reuse. |
| Repository information (metadata) | When depositing research data, complete as much of the requested information as possible (e.g. title, authors, description, keywords, funding information and licences). | Helps others discover, understand and correctly interpret the data. |
| Access conditions | Decide whether data will be openly available, shared under controlled access, or restricted. | Balances openness with ethical, legal and contractual obligations. |
| Persistent identifiers | Obtain a DOI or another persistent identifier for the dataset where appropriate. | Allows datasets to be cited reliably and found in the future. |
Common Misconceptions
A Data Management Plan is just an administrative requirement.
FALSE: A Data Management Plan helps transform assumptions into agreed decisions, records why choices were made, and keeps collaborators aligned as the project evolves.
I'll organise my files when I have time.
REALITY: Small habits established throughout the project are much easier than trying to reorganise everything at the end.
I'll remember how I processed the data.
REALITY: Documentation written while you work is far more reliable than relying on memory weeks or months later.
Everything should be preserved forever.
FALSE: Not all research data need to be kept. The goal is to preserve the research outputs that have long-term value, together with enough documentation to understand and reuse them.
'Available upon request' is a sufficient preservation strategy.
FALSE: Research has shown that data kept only by individual researchers often become unavailable over time. Trusted repositories provide much more reliable long-term preservation.
Key Takeaways
- Research data management is a continuous process that spans planning, managing, preserving and sharing research data.
- Many RDM decisions only take a few minutes, but they can prevent significant problems later in a project.
- Writing decisions down helps transform assumptions into shared ways of working.
- Good documentation supports both collaborators and your future self.
- Preserving research data means keeping them understandable and usable, not simply storing a copy somewhere.
References
Kovács, M., Hoekstra, R., & Aczél, B. (2021). The Role of Human Fallibility in Psychological Research: A Survey of Mistakes in Data Management. Advances in Methods and Practices in Psychological Science, 4. https://doi.org/10.1177/25152459211045930. Licence: CC BY-NC 4.0
Perrier, L., Blondal, E., Ayala, A. P., Dearborn, D., Kenny, T., Lightfoot, D., Reka, R., Thuna, M., Trimble, L., & MacDonald, H. (2017). Research data management in academic institutions: A scoping review. PLoS ONE, 12. https://doi.org/10.1371/journal.pone.0178261. Licence: CC BY 4.0
Specht, A., Bolton, M., Kingsford, B., Specht, R., & Belbin, L. (2018). A story of data won, data lost and data re-found: the realities of ecological data preservation. Biodiversity Data Journal. https://doi.org/10.3897/bdj.6.e28073. Licence: CC BY 4.0
Töwe, M., & Barillari, C. (2020). Who does what?–Research data management at ETH Zurich. Data Science Journal, 19(1), 36-36. https://doi.org/10.5334/dsj-2020-036. Licence: CC BY 4.0
Vines, T., Albert, A. Y., Andrew, R., Debarre, F., Bock, D. G., Franklin, M. T., Gilbert, K., Moore, J.-S., Renaut, S., & Rennison, D. (2013). The availability of research data declines rapidly with article age. Current biology : CB, 24 1, 94-97 . https://doi.org/10.1016/j.cub.2013.11.014. Licence: © 2014 Elsevier Ltd. All rights reserved
Yu, F., Deuble, R., & Morgan, H. (2017). Designing research data management services based on the research lifecycle–a consultative leadership approach. Journal of the Australian Library and Information Association, 66(3), 287-298. https://doi.org/10.1080/24750158.2017.1364835. Licence: © 2017 Australian Library & Information Association
Zingmark, M., & Iwarsson, S. (2025). Case study on challenges in research with public partners: A personal data incident during recruitment for a survey study on ageing and housing. BMC Research Notes, 18. https://doi.org/10.1186/s13104-025-07246-8. Licence: CC BY 4.0
Author: Jonathan England
Created: August 2026