Skip to main content
Guides for Researchers

Where Should I Share My Research Data?

Choosing a Repository for Long-Term Preservation and Reuse

Why Use a Repository

Researchers have many ways to make files available online. Data may be uploaded as supplementary materials alongside a publication, placed on a project website, stored on an institutional web page, or shared through cloud storage services. While these approaches can make files accessible in the short term, they do not necessarily provide the features needed for long-term preservation, discovery and reuse.

Research data repositories are specifically designed to support the long-term management of research outputs. They typically provide proper technical infrastructure, preservation policies, access controls and citation mechanisms that help ensure data remain discoverable and accessible over time. They can therefore play an important role in supporting the FAIR principles and facilitating future reuse.

Sharing option Good for Potential limitations
Supplementary files attached to publications Directly findable and accessible alongside an article. Not optimised for discovery as standalone datasets, may be difficult to find independently from the article, and may not provide dedicated preservation or access-management features.
Project websites Can help with communication and dissemination of project outputs and results. May disappear when funding ends, staff leave, or the website is redesigned. Usually lack long-term preservation infrastructure and repository services.
Personal or laboratory websites Informal dissemination and showcasing research outputs. Depend on individual maintenance and may disappear when the researcher moves institution or retires.
Email on request Cannot easily support large datasets, is not discoverable, and can create barriers or delays for potential users.
Cloud storage services (e.g. Google Drive, Dropbox, OneDrive) Providing convenient access to files, spreadsheets and other resources. Generally not designed for long-term preservation, citation, or discovery through research data catalogues and search services.
Institutional shared drives Providing access to files within an institution or to specific collaborators. Typically not intended for public dissemination, discovery or long-term preservation.
Research data repositories Long-term preservation, sharing, citation and reuse of research data. Requires preparation of data, documentation and metadata before deposit.

Making files available online is therefore not always the same as depositing them in a trusted repository.

Repositories often provide features that are difficult to maintain through ordinary websites, including:

  • Persistent identifiers (such as DOIs) that provide stable citations.
  • Structured metadata that improve discoverability.
  • Long-term preservation policies and infrastructure.
  • Support for licences and access conditions.
  • Controlled-access mechanisms for sensitive data.
  • Citation information to support attribution and scholarly credit.
  • Links between datasets, publications, software and other research outputs.

For these reasons, many funders, institutions and journals recommend, and in some cases require, that research data be deposited in an appropriate repository rather than being shared only through supplementary materials, project websites or personal web pages.

Not All Repositories Are the Same

Research data repositories vary considerably in their scope, purpose and audience. Some are designed for specific disciplines, while others accept data from any field. Some are managed by research institutions, while others are operated by national initiatives, international organisations or research infrastructures.

In many cases, the most appropriate repository is not simply the first one that accepts your files, but the one most likely to make the data discoverable and reusable by the intended research community.

Types of Repositories

Several broad categories of repositories exist:

Repository type Description Examples
Disciplinary repositories Serve specific research communities and often support disciplinary standards and metadata.

PANGAEA (earth and environmental sciences), European Genome-phenome Archive, ARIADNE (archaeology)

Institutional repositories Managed by universities or research organisations to preserve and share outputs produced by their researchers. University repositories
National or regional repositories Support researchers within a country or region and may align with national Open Science policies and infrastructures. AUSSDA (Austria), DataFirst (African countries)
General-purpose repositories Accept data from a wide range of disciplines and research areas. Zenodo

Good to know: Some repositories are part of larger research infrastructures, such as CLARIN, GBIF, CESSDA and ELIXIR, which support specific research communities through shared standards, services and data platforms.

Finding a Suitable Repository

If you do not already know which repository is commonly used in your field, several services can help identify suitable options.

Resource What it is What we like about it
re3data Global registry of research data repositories. Covers thousands of repositories worldwide. It has a very practical visual tool to browse by subject or by country.
FAIRsharing Registry of standards, repositories and data policies. Goes beyond repository discovery by connecting repositories with the standards and policies that recommend them, helping researchers make more informed decisions about where and how to share data.
Research infrastructures Domain-specific services that support particular research communities. Many infrastructures maintain lists of recommended repositories and community standards. Examples include ELIXIR (life sciences), CLARIN (language resources and digital humanities), and CESSDA (social sciences).
Institutional or funder guidance Recommendations provided by universities, research organisations or libraries. Can help identify repositories that satisfy disciplinary, institutional or policy requirements.

In many cases, the easiest starting point is to search re3data, consult FAIRsharing, or review guidance provided by your institution, funder or disciplinary community. If a recognised disciplinary repository exists, it is often a good candidate because it is more likely to be used and searched by researchers in the relevant field.

How Do I Choose a Repository

When several repositories are available, selecting the most appropriate one can sometimes be challenging. In general, it is often preferable to choose a repository that is recognised and used by the intended research community, as this increases the likelihood that the data will be discovered, understood and reused.

The best repository is not necessarily the largest or most widely known. Instead, researchers should consider whether the repository is appropriate for the type of data being shared, supports any relevant disciplinary standards, and provides the features needed to preserve and manage access to the data over time.

The following questions may help when evaluating potential repositories:

Question Why it matters
Does my discipline have a recognised repository? Community repositories are often the most visible and widely used.
Does the repository assign persistent identifiers (e.g. DOIs)? Supports citation, linking and long-term referencing.
Does it support the required access conditions? Some datasets require restricted or controlled access rather than open publication.
Does it provide adequate metadata and documentation options? Improves discoverability and reuse.
Does it accept my file types and data volumes? Some repositories have limitations.
Does the repository support licences and access controls that match my needs? Some repositories offer more flexibility than others for managing reuse and access conditions.
Are there eligibility requirements, storage limits or fees? Practical considerations may influence repository selection.
Is the repository trusted by my community, funder or journal? Can improve visibility and compliance.
Does it have a clear preservation policy? Long-term preservation helps ensure continued access to the data.

💡 Some repositories hold certifications such as CoreTrustSeal, which indicate that they meet recognised standards for trustworthy data stewardship and preservation.

No single repository will be suitable for every type of research data. The most appropriate choice will depend on the discipline, the nature of the data, and any legal, ethical, institutional or funder requirements that apply to the project.

What Information Will a Repository Ask For

Depositing data in a repository usually involves more than uploading a set of files. To help others discover, understand and reuse the data, repositories typically ask researchers to provide descriptive information about the dataset. This information is commonly known as metadata, or "data about data".

Metadata help potential users understand what the dataset contains, who created it, when it was produced, and how it relates to other research outputs. They also allow repositories, search engines and discovery services to index and retrieve datasets more effectively.

Although requirements vary between repositories, many ask for similar types of information.

Information Why it matters
Title Provides a clear and concise description of the dataset.
Creator(s) Identifies the researchers responsible for the data and supports attribution. Make sure you follow the repository's recommended naming format (e.g. Patel, Anika Priya versus Anika Priya Patel versus Patel, A.P.) to ensure names are displayed and indexed consistently.
ORCID iD A unique identifier for researchers. It helps distinguish people with similar names, improves attribution, and supports connections between datasets, publications and other research outputs.
Description or abstract Explains what the dataset contains, how it was created and what it may be useful for.
Keywords Improves discoverability through repository searches and search engines.
Dates Records when the data were collected, created or published.
Funding information Links the dataset to the projects or funders that supported the research.
Licence Clarifies how others may access, use and reuse the data.
Access conditions Explains whether the data are openly available, restricted, under embargo, or subject to controlled access.
Related outputs Connects the dataset to publications, software, protocols or other research outputs.

Providing good metadata is not simply an administrative requirement. In many cases, users will encounter the metadata long before they access the dataset itself. Clear titles, descriptions and keywords can therefore have a significant influence on whether data are discovered and reused.

Most repositories require only a minimum set of metadata before a dataset can be deposited. However, taking the time to provide richer descriptions, additional keywords, links to related outputs and other contextual information can substantially improve discoverability, understanding and reuse. Well-prepared metadata often have a greater influence on whether a dataset is found and reused than the data files themselves.

Repositories may also request additional information depending on the discipline. For example, disciplinary repositories often support specialised metadata standards that describe particular types of data in greater detail. As discussed earlier in this guide, using community standards can improve interoperability and make data easier for others in the field to understand and reuse.

Tip: Imagine that someone discovers your dataset through a search engine several years from now. Would the title, description and metadata provide enough information for them to understand what the data contain, how they were generated, and whether the dataset is relevant to their work, without having to open it?

Linking Related Research Outputs

Research projects often produce more than a single dataset. Publications, software, protocols, reports, preprints and other outputs may all contribute to the research process. Linking these outputs together helps provide context, improves discoverability and allows others to better understand how the research was conducted.

Many repositories allow researchers to establish connections between datasets and related outputs during the deposition process. When possible, provide identifiers for related outputs rather than simply mentioning them in a description.

Research output Why link it? What to do
Publications Allows readers to access the research outputs supported by the data and understand how the data were used. Add the DOI assigned by the publisher.
Preprints Connects the dataset to earlier versions of research outputs. Add the DOI, Handle or other identifier assigned by the preprint server.
Software and code Enables others to reproduce analyses, workflows and results. Link to the software repository and provide a DOI if one has been assigned.
Protocols and methods Provides details about how the data were collected, processed or analysed. Add a DOI or repository link if the protocol has been published or deposited.
Other datasets Helps document relationships between datasets generated by the same project or research programme. Link related datasets using their DOI or repository identifier.
Training materials Supports understanding and reuse of the dataset. Add links to relevant tutorials, documentation or educational resources.
Projects or grants Provides funding context and connects outputs from the same project. Include grant or project identifiers when requested by the repository.

Repositories increasingly support these connections through persistent identifiers (PIDs). A PID is a long-lasting identifier that uniquely identifies a person, organisation or research output, even if its location changes over time.

Common examples include:

Identifier Used for Example
DOI (Digital Object Identifier) Publications, datasets, software and other research outputs 10.xxxx/xxxxx
ORCID iD Researchers and contributors 0000-0000-0000-0000
ROR ID Research organisations and institutions https://ror.org/...
Grant identifiers Research projects and funding awards Funder-specific identifiers

Additional persistent identifiers exist for specific types of research outputs, including software (SWHIDs), research projects (RAiDs), physical samples (IGSNs) and research resources (RRIDs).

By linking outputs through persistent identifiers, repositories and discovery services can automatically connect datasets, publications, software, researchers, institutions and funding information. This helps create a more complete and interoperable scholarly record.

Tip: When depositing data, take a few minutes to link any related publications, software, protocols or project information. These connections can significantly improve the visibility, discoverability and reusability of your research outputs.

Use PIDs to Share and Cite Research Outputs

Persistent identifiers are most useful when they are actively used. Whenever possible, use the persistent identifier of a research output rather than a repository-specific or temporary web address.

For example, when sharing a dataset deposited in Zenodo, prefer:

https://doi.org/10.5281/zenodo.20631491

rather than:

https://zenodo.org/records/20631491

The DOI will continue to resolve to the dataset even if the repository changes its internal structure or web addresses. Repository URLs may change over time, while persistent identifiers are designed to provide stable, long-term access.

The same principle applies when sharing research outputs in:

  • publications;
  • reference lists;
  • data availability statements;
  • project reports;
  • websites;
  • presentations;
  • social media posts.

Using persistent identifiers helps ensure that readers can reliably locate the resource and allows repositories, discovery services and citation systems to track connections between research outputs.

A Final Check Before Depositing Data

Before depositing a dataset in a repository, it can be helpful to perform one final review. This does not need to be a lengthy process, but it can help identify issues that may reduce the discoverability, accessibility or reusability of the data.

The following questions can serve as a practical checklist:

Question Why it matters
Have I selected the data that should be shared? Ensure that the dataset contains the materials needed to understand, verify or reuse the research.
Have unnecessary files been removed? Temporary files, duplicates and obsolete versions can create confusion.
Are the data organised and clearly named? Consistent file names and folder structures make datasets easier to navigate.
Is sufficient documentation available? README files, codebooks, protocols and data dictionaries help users understand the data.
Have community standards been considered? Disciplinary conventions can improve interoperability and reuse.
Are the file formats appropriate for sharing and preservation? Open or widely used formats generally improve long-term accessibility.
Have access conditions been defined? Users should understand whether the data are openly available, restricted or subject to controlled access.
Has an appropriate licence been selected? Reuse conditions should be clearly communicated.
Have personal or sensitive data been addressed appropriately? Anonymisation, pseudonymisation or access controls may be necessary.
Have metadata been completed? Good metadata improve discoverability and understanding.
Have related outputs been linked? Publications, software, protocols and other outputs provide valuable context.
Are persistent identifiers available and being used? PIDs support citation, linking and long-term discovery.

No dataset will ever be perfectly documented or completely future-proof. The goal is not perfection, but to ensure that others can find, understand, access and reuse the data as effectively as possible.

Key Takeaways

  • Preparing data for sharing involves more than uploading files to a repository.
  • Not every file created during a project needs to be preserved or shared.
  • Documentation, metadata and community standards are often as important as the data themselves.
  • Choosing appropriate file formats, licences and access conditions can improve long-term reuse.
  • Trusted repositories provide services that support preservation, discovery, citation and access management.
  • Different types of repositories serve different communities and needs.
  • Rich metadata make datasets easier to discover and understand.
  • Linking datasets to publications, software, protocols and other outputs provides important context.
  • Persistent identifiers help connect research outputs and support citation and discoverability.
  • Investing a little extra effort before deposition can significantly increase the visibility and reusability of research data.

DOI

Author: Jonathan England

Created: August 2026