Skip to main content

Guides

Findable

Make your data (and metadata) findable by ensuring it:

  • Has a globally unique, persistent identifier (PID) (e.g., a DOI) 
  • Has rich, machine-readable discovery (citation/descriptive) metadata (e.g., title, creators, abstract, keywords, dates, methods, license) 
  • Is registered/indexed in a searchable resource (typically a repository/catalog) so it is discoverable by people and machines 

Persistent identifiers (PIDs) are important because they unambiguously identify your data and support reliable citation and linking across systems. A common PID for datasets is a Digital Object Identifier (DOI)Choose a repository that mints/registers PIDs and exposes the record through a stable landing page (e.g., Zenodo registers a DOI for uploads via DataCite). 

The metadata describing your data supports findability, citation, and reuse, and, critically supportsmachine-actionability (so services can index, link, and reuse your records at scale). Follow community metadata standards where available, or widely used cross-domain standards (e.g., Dublin CoreDCC/DataCite Metadata Schema), and prefer standards and vocabularies that are maintained and widely implemented by repositories. 

To identify appropriate standards and vocabularies, consult:

DCC metadata guidance (disciplinary and general metadata resources)

Accessible

Make your data accessible by ensuring it:

  • Is retrievable by its identifier using standardized, open protocols (e.g., HTTPS; APIs where relevant) 
  • Supports authentication/authorization where necessary (for sensitive, embargoed, or restricted data) 
  • Keeps metadata publicly accessible even if the data are restricted or no longer available 

Remember: not all data must be open to be FAIR. Data can be restricted and still be FAIR, as long asthe metadata are openly accessible and the access conditions are described clearly

This aligns with the European Commission’s principle: “as open as possible, as closed as necessary.” 

 
As Open as Possible, As Closed as Necessary
 

Where can I keep my data (for the long term)?

Not necessarily “open to everyone,” but safe, preservable, and persistently accessible. Look for a repository that:

  • Preserves data for the long term (including format-risk and preservation planning) 
  • Makes (meta)data findable (searchable record pages; indexing/harvesting) 
  • Supports rich, standardized, machine-readable metadata 
  • Captures access conditions and reuse terms (license) in the record metadata 

You can deposit data to a general repository (e.g., ZenodoDataverse installations) or a subject-specific repository (e.g., Dryad, domain repositories). Prefer discipline repositories when they exist, because they often enforce domain standards and community metadata. 

To identify suitable repositories for your discipline, search:

  • re3data (registry of research data repositories; searchable by discipline and features) 
  • FAIRsharing (also indexes repositories and which standards/policies they align with) 

Interoperable

Make your data interoperable by using:

  • Open, documented file formats (avoid proprietary-only formats where feasible) 
  • Community-agreed schemas/standards and machine-readable structures (so tools can parse and integrate your data) 
  • Controlled vocabularies, thesauri, and ontologies for consistent meaning (semantics) across systems 
  • Qualified links to related entities via PIDs (e.g., link dataset DOI ↔ article DOI ↔ ORCID iD; include funder/organization IDs where applicable) 

Interoperable data can be integrated with other data, applications, and workflows. Think about not creating data that can only be read in proprietary software, and when you must use proprietary tools export/share an interoperable version (e.g., CSV/TSV plus a data dictionary; open exchange formats) so others can reuse it without specialized software.

Reusable

Make your data reusable by ensuring it:

  • Is well-documented (so others—and your future self—can interpret it correctly) 
  • Has clear reuse terms via a license, recorded in the metadata 
  • Includes provenance and context (how data were generated/processed; versions; assumptions) 
  • Meets domain-relevant community standards where they exist 

Create documentation, e.g., a README (plain text or Markdown preferred for machine-use; PDF only if formatting is essential) that includes:

  • For each filename: what it contains and how it relates to figures/tables/publications 
  • For tabular data: definitions of columns/rows, codes (incl. missing values), and units 
  • Processing/cleaning steps that affect interpretation 
  • Links to related datasets stored elsewhere (with PIDs where possible) 
  • Contact point (and ideally ORCID iD) for questions 

Source (README guidance): Dryad’s “Creating a README.” 

Data should have a clear license to govern the terms of reuse. If reuse is intended, avoid “no license” (which creates legal ambiguity and often blocks reuse). Practical guidance from the DCC can help you choose an appropriate data license and understand trade-offs. You can also check this practical guide on licensing research data.

Where possible, to maximize reuse, consider widely recognized licenses such as CC BY 4.0 (attribution) or CC0 (public domain dedication/waiver), and document any restrictions in both the metadata and DMP. 

Check out: EUDAT’s License Selector wizard to support consistent license choice, especially when balancing derivatives, share-alike, and commercial reuse questions.

Find out more on HE online manual  

Use this checklist to reflect on how well your data meet the FAIR principles: Findable, Accessible, Interoperable, and Reusable.

Findable

  • Does the dataset have a clear title and description?
  • Is rich metadata provided?
  • Are keywords, subject terms, and disciplinary categories included?
  • Has the dataset been assigned a persistent identifier, such as a DOI?
  • Is the dataset deposited in a searchable repository or catalogue?

Accessible

  • Can users access the data or metadata through a clear landing page?
  • Are access conditions clearly stated?
  • Is the licence or reuse condition visible?
  • If the data cannot be open, is the restriction explained?
  • Will metadata remain available even if access to the data is restricted?

Interoperable

  • Are open, standard, and machine-readable formats used where possible?
  • Are community standards, vocabularies, or ontologies used?
  • Are variables, units, methods, and relationships clearly documented?
  • Can the data be combined with other datasets or systems?
  • Are links provided to related publications, software, protocols, or datasets?

Reusable

  • Is the dataset accompanied by enough documentation for others to understand and reuse it?
  • Are methods, provenance, processing steps, and quality checks described?
  • Are versioning and file naming clear?
  • Is a clear licence applied?
  • Are ethical, legal, or confidentiality restrictions explained?

Jones, S. & Grootveld, M. (2017, November). How FAIR are your data? Zenodo. http://doi.org/10.5281/zenodo.1065991

 

key message symbol icon sq

Key Message

FAIR does not always mean fully open. Data should be as open as possible and as closed as necessary, but the metadata, documentation, access conditions, and preservation plan should make the research output understandable and reusable over time.

 

 

FAIR HOW TO

Source: U.S. National Library of Medicine (NLM), Common Data Elements (CDE) Tutorial.

 

 

In deciding where to store your data, you may have a number of choices about who will look after it. The choice may be straightforward if you have an established data management facility in your domain or institution, or even within your research group or department. When data preservation standards or norms exist in your discipline, these should be followed. Your research funder may recommend a data centre or self-deposit archive. In order of preference:

  1. Use an external data archive or repository already established for your research domain to preserve the data according to recognised standards in your discipline.
  2. If available, use an institutional research data repository, or your research group’s established data management facilities.
  3. Use a cost-free data repository such as Zenodo.
  4. Search for other data repositories here: re3data.org. On top of specific research disciplines you can filter on access categories, data usage licenses, trustworthy data repositories (with a certificate or explicitly adhering to archival standards) and whether a repository gives the data a persistent identifier.

res3data categories

Remember – you don’t need to keep everything! Work with your library to help you determine which data you need to retain for validation and/or reuse.

When choosing a repository it is important to consider factors such as whether the repository:

  • Gives your submitted dataset a persistent and unique identifier. This is essential for sustainable citations – both for data and publications – and to make sure that research outputs in disparate repositories can be linked back to particular researchers and grants.
  • Provides a landing page for each dataset, with metadata that helps others find it, tell what it is, relate it to publications, and cite it. This makes your research more visible and stimulates reuse of the data.
  • Helps you to track how the data has been used by providing access and download statistics.
    Responds to community needs and is preferably certified as a ‘trustworthy data repository’, with an explicit ambition to keep the data available in the long term.
  • Matches your particular data needs (e.g. formats accepted; access, back-up and recovery, and sustainability of the service). Most of this information should be contained within the data repository’s policy pages.
  • Offers clear terms and conditions that meet legal requirements (e.g. for data protection) and allow reuse without unnecessary licensing conditions.
  • Provides guidance on how to cite the data that has been deposited.
  • Charges for its services.

Your institution may offer Research Data Management support to help you deal with these issues and get the most out of the investment put into your research. This could involve:

  • Registering datasets with the institution’s Data Catalogue to help make the research more visible. National registry services are also being established to harvest institutional data catalogue records to make the data visible at a national level.
  • Depositing the dataset with an institutional repository to maintain a long-term record of its safekeeping and, if it is publicly available, the access and download statistics.
  • Providing advice on rich metadata and other documentation, to help make the data both discoverable and understandable.
    Selecting an appropriate licence for using your data – also when you make them available Open Access.