Skip to main content
Guides for Researchers

Making Sense of the FAIR Principles

Understanding the ideas behind Findable, Accessible, Interoperable and Reusable

What is FAIR

Digital research outputs are most valuable when they can be used beyond the project in which they were created. To make this possible, researchers need to organise and manage their data so they remain easy to find, understand and reuse, both during a project and in the future.

The FAIR Principles provide a framework for achieving this. They are a set of guiding principles that help researchers manage research data so they remain useful over time for themselves, collaborators, other researchers and computer systems.

The FAIR Principles were introduced in 2016 by a group of researchers and data experts to improve the management and sharing of digital research outputs in an increasingly data-driven research environment (Wilkinson et al., 2016). As the amount of research data continues to grow, it is no longer realistic for people alone to find and work with all available data. The FAIR Principles therefore encourage researchers to organise and describe their data so that both people and computer systems can find, understand and use them effectively. This ability for computer systems to automatically find, access and process data is known as machine-actionability. It is becoming increasingly important as technologies such as artificial intelligence rely on data that can be processed automatically and consistently.

FAIR is an acronym for four principles:

  • Findable: Data and their descriptions should be easy to find.
  • Accessible: Once found, it should be clear how the data can be accessed.
  • Interoperable: Data should be organised so they can be combined with other data and used across different software and systems.
  • Reusable: Data should include enough information to allow them to be understood and reused appropriately.

FAIR is not the same as Open Data. Data can be FAIR without being openly available, for example when access may remain under embargo or be shared through controlled access for legal, ethical, commercial or security reasons. Likewise, making data openly available does not automatically make them FAIR if they lack sufficient metadata, documentation or other information needed for reuse. The goal is to make research data as useful as possible, regardless of whether they can be shared openly.

Finally, FAIR is best understood as a set of guiding principles rather than a technical standard or certification. It is also a continuum rather than a binary concept. Researchers do not either achieve FAIR completely or not at all. Instead, small improvements, such as organising files consistently, documenting how data were collected or choosing an appropriate repository, can all make research data more FAIR.

Why FAIR matters

Applying the FAIR Principles helps maximise the value of research by making data easier to discover, understand and reuse. FAIR data management supports transparency and reproducibility, enables collaboration and data reuse, and helps maximise the scientific, societal and economic value of research. The following sections highlight some of the main benefits of FAIR research data.

Help researchers discover existing data

Research data can only be reused if other researchers know they exist. Making data easier to find helps researchers discover relevant datasets, build on existing work and avoid collecting data that already exist. This saves time, reduces duplication of effort and increases the value of research over time.

Example: A study analysing more than 2,000 links to shared research data found that data stored in recognised repositories were much more likely to remain available over time than data shared through ordinary web links. The authors concluded that depositing data in repositories is a more reliable way to ensure research data remain discoverable in the future (Briney, 2024).

Enable new research through data reuse

Collecting research data often requires considerable time, funding and participant effort. When research data are shared in ways that enable reuse, they can continue to generate new knowledge long after the original project has ended. Researchers can investigate new questions, test different hypotheses and combine data from multiple studies, extending the value of the original research (Piwowar & Vision, 2013).

Example: UK Biobank was established to support future health research using data from around 500,000 participants. During the COVID-19 pandemic, researchers rapidly reused the resource by linking participants' health records with SARS-CoV-2 testing data and follow-up information. This allowed them to investigate the risk factors and long-term effects of COVID-19 without recruiting a new cohort of participants, saving considerable time and resources (Feng et al., 2023).

Improve transparency and reproducibility

Research findings are easier to understand and verify when researchers clearly document how data were collected, processed and analysed. Sharing data together with the information needed to interpret and reproduce the analyses helps others evaluate the research, identify potential errors and build confidently on previous work (Munafò et al., 2017).

Example: After the journal Management Science introduced a policy requiring authors to share their data and code, the proportion of papers with replication materials increased substantially. For papers where reviewers could access the required data and software, more than 95% of the published results could be fully or largely reproduced. The main barriers to reproducibility were inaccessible data, missing code and insufficient documentation (Fišar et al., 2023).

Maximise the impact of research

Research data create the greatest value when they can be reused beyond the original project. FAIR data support innovation, inform public policy, enable evidence-based decision making and accelerate scientific and technological advances. By making research data easier to discover, access and combine, FAIR practices increase the likelihood that research will generate benefits for science, industry, government and society.

Example: The Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology, has accumulated decades of experimentally determined protein structures through contributions from researchers worldwide. This shared collection of high-quality data enabled the development of AI-based protein structure prediction systems such as AlphaFold, demonstrating how openly shared research data can support major scientific breakthroughs. The resulting AlphaFold Protein Structure Database is now widely used across structural biology, medicine and drug discovery, extending the impact of these shared data even further (Berman & Burley, 2025; Váradi & Velankar, 2022).

What does each principle mean

The FAIR Principles provide a framework for making research data more useful to both people and computers. Rather than being a checklist, they describe four complementary characteristics of well-managed data. The table below summarises each principle before looking at them in more detail.

Principle In short... How to put it into practice
Findable Make your data easy to discover. Describe your data clearly, give it a permanent identifier, and deposit it in a trusted repository.
Accessible Make your data available under clear conditions. Store your data in a repository, explain how it can be accessed, and keep its description available even if the data cannot be shared.
Interoperable Make your data work with other data and software. Use common file formats, agreed standards and consistent terminology.
Reusable Make your data easy to understand and reuse. Provide enough documentation, include a licence, and ensure the data are well described and of good quality.

Findable

What it means

Data cannot be cited or reused if nobody knows it exists. Making data findable means that both people and computers can discover your dataset, identify it correctly and understand what it contains without opening every file.

Even if access to the data is restricted, other researchers should still be able to discover that the dataset exists.

How to make your data Findable

Make it easy for other people to discover that your data exist.

  • Store your data in a trusted repository where others can search for them.
  • Give your dataset a permanent reference so it can always be found.
  • Provide a clear description of what the dataset contains.
  • Use a descriptive title and relevant keywords.
  • Connect the dataset to related publications, software or other research outputs.

Example

Researchers from the NIAID Systems Biology Consortium found that many infectious disease datasets were difficult to discover because they were scattered across multiple repositories and could not easily be found through services such as Google Dataset Search. By improving how nearly 400 datasets and software tools from 15 research centres were described, the consortium made them much easier to discover while allowing them to remain in their original repositories. This illustrates that making data findable is about improving their discoverability, not necessarily moving them into a single database or making them openly available (Tsueng et al., 2022).

Accessible

What it means

Making data accessible means ensuring that, once a dataset has been found, it is clear how it can be retrieved. Accessible does not necessarily mean openly available. Access may be immediate, restricted or not possible, depending on legal, ethical, commercial or security considerations. In these cases, the conditions for accessing the data should be clearly explained.

Even if the data itself cannot be accessed, information describing the dataset should be available so that the existence of the research is not lost. This allows others to understand what data were produced, understand how access can be requested, and why they may no longer be available.

How to make your data Accessible

Make it clear how other people can obtain your data.

  • Store your data in a trusted place that provides long-term access.
  • Explain whether the data are openly available or whether access is restricted.
  • Describe how people can request access when restrictions apply.
  • Keep a public description of the dataset available, even if the data themselves cannot be shared.

Example

Many biodiversity databases make observations of endangered species discoverable while protecting their precise locations. A review of 39 biodiversity databases found that sensitive records are often shared under restricted access or with reduced location precision to reduce the risk of poaching, illegal collection and habitat disturbance. Researchers can therefore discover that the data exist and, where appropriate, request access without exposing information that could threaten vulnerable species (Kaehrle & Eschenfelder, 2025).

Interoperable

What it means

Making data interoperable means ensuring that data can be combined with other datasets and used by different software, tools and researchers. Data should not only make sense to the researchers who created them, but also be understandable and usable by others working in the same field or in different disciplines.

Interoperability also helps computer systems exchange and interpret data automatically, making large-scale analyses and data integration possible.

How to make your data Interoperable

Organise your data so they can be combined with other data and used by different software.

  • Use file formats that are widely supported.
  • Organise your data consistently.
  • Use clear and consistent names for variables, categories and measurements.
  • Avoid abbreviations or codes that only your research group understands.
  • Link related datasets and research outputs together.

Example

Psychology and Neuroscience: Researchers studying the brain often collect different types of information, such as brain scans, cognitive tests and clinical assessments. However, these data have traditionally been organised differently by individual research groups, making it difficult to combine results from different studies or analyse them using the same software. To address this, researchers developed the **Brain Imaging Data Structure (BIDS)**, a common framework for organising brain imaging datasets. By adopting a consistent structure, research teams from different institutions can more easily analyse datasets using the same software, compare findings across studies and combine data for larger collaborative projects (Gorgolewski et al., 2016).

Digital Humanities: Museums, libraries and archives often hold related historical objects, documents and artworks, but each institution has traditionally managed its collections independently. This makes it difficult for researchers to search across collections or study cultural heritage on a larger scale. Initiatives such as Europeana, and research projects such as the Sampo series (e.g. WarSampo), use a common approach that allows collections from many different organisations to be searched and explored together, even though they remain managed by the contributing institutions (Hyvönen, 2022; O'Neill & Stapleton, 2022).

Reusable

What it means

Making data reusable means providing enough information for others to understand, evaluate and use the data correctly, both now and in the future. Someone who was not involved in the original study should be able to interpret what the data represent, how they were collected and any limitations that should be considered.

Reusable data remain valuable long after the original project has finished because future researchers can build on them with confidence. But, without sufficient documentation, even openly available data may be difficult to interpret correctly.

How to make your data Reusable

Provide enough information for other researchers to understand and use your data correctly.

  • Explain how the data were collected, processed and analysed.
  • Describe what the dataset contains.
  • State clearly how other people may use the data.
  • Follow established practices in your research field.
  • Check that the data are complete, accurate and well organised before sharing.

Example

A researcher combined data from five potato field experiments conducted over 11 years in four different countries to investigate how different potato varieties respond to environmental conditions (Hurtado-Lopez et al., 2015). Although the datasets had been collected independently and in different formats, a substantial proportion could be successfully reused to produce new findings. However, the process required considerable effort because important information was scattered across publications, theses and personal communications, and some data could not be reused because essential details were missing. Papoutsoglou et al. (2023) later revisited this case to show how better documentation and community standards could have made the data much easier to understand and reuse.

FAIR and Open Data

FAIR and Open Data are closely related but are not the same thing. FAIR focuses on ensuring that research data are well described and reusable, whereas Open Data concerns making data openly available for reuse. Data can be FAIR without being open, and openly available data are not necessarily FAIR.

Read the guide on Open Research Data for a more in-depth description.

Putting FAIR into Practice

The FAIR Principles are not a separate task to complete at the end of a research project. They are achieved through a series of good research data management practices that help make data easier to find, access, combine and reuse. The checklist below provides a practical starting point. Each topic is explored in more detail in later chapters of this guide.

FAIR Checklist

Plan for data sharing from the start

It is much easier to make data FAIR if you consider how they will be organised, documented and shared before data collection begins. Planning ahead reduces the amount of work needed later and helps identify any legal, ethical or technical issues early in the project.

Data Management Plans

Organise your files consistently

Use a logical folder structure and meaningful file names that make sense to both you and other researchers. Keeping related files together and following a consistent naming approach makes datasets easier to navigate and reduces confusion over time.

Organising Research Data

Document your data

Imagine that someone unfamiliar with your project needs to understand your dataset several years from now. Provide enough information to explain what the data contain, how they were collected, how variables or columns should be interpreted, and whether any processing has already been carried out.

Describing Your Data 

Choose suitable file formats

Whenever possible, use file formats that are widely supported and likely to remain usable in the future. Open or commonly used formats reduce the risk that data become difficult to access because specialised software is no longer available.

File Formats

Deposit your data in a trusted repository

Rather than storing data only on a personal computer, institutional website or as supplementary files to a publication, deposit them in a trusted data repository that is designed for long-term preservation and discovery. Many repositories also help make datasets easier to cite and access.

Choosing a Repository

Describe how your data may be reused

Clearly explain any conditions for reuse. If your data can be shared openly, make this explicit. If access must be restricted, explain why and describe how other researchers can request access if appropriate.

Licensing Research Data 

Link related research outputs

Where possible, connect your dataset with related publications, software, protocols, preprints and other research outputs. These links provide important context and help others understand how the data were generated and used.

Persistent Identifiers and Linking Research Outputs

Follow practices used in your research community

Many disciplines have established ways of organising, describing and sharing data. Following these practices makes it easier for other researchers in your field to understand and reuse your work.

Standards and Best Practices

Think beyond publication

Making data FAIR is not simply about publishing a dataset. Consider whether someone who discovers your data in several years' time would still be able to understand what they contain, how they were created and how they may be reused.

 

Review your dataset before sharing

Before depositing your data, check that files are complete, organised and understandable. Remove unnecessary duplicates, ensure documentation is up to date and confirm that any confidential or sensitive information has been handled appropriately.

Preparing Data for Sharing

 

Common Misconceptions

"A dataset must satisfy all four FAIR Principles before it has any value."
FALSE: FAIR is not an all-or-nothing concept. Many datasets already satisfy some FAIR Principles and can become more FAIR over time. Even small improvements, such as providing better documentation or depositing data in a repository, can significantly increase their value and reusability.


"Making data FAIR requires specialised technical knowledge."
FALSE: Many FAIR practices involve straightforward research data management, such as organising files, documenting datasets and depositing them in an appropriate repository. While some disciplines use specialised standards and technologies, researchers can make substantial progress towards FAIR without becoming technical experts.


"The FAIR Principles are a checklist that every dataset must follow in exactly the same way."
FALSE: The FAIR Principles provide general guidance rather than prescriptive rules. The most appropriate way to implement FAIR depends on the type of data, the discipline and any legal, ethical or technical constraints.


"Interoperable means the data must be machine-readable or use artificial intelligence."
FALSE: Interoperability simply means that data can be combined with other datasets and used by different software. In many cases, adopting common file formats, consistent terminology and established disciplinary practices is enough to make data much easier to integrate and reuse.


"Making data FAIR only benefits researchers who reuse the data."
FALSE: FAIR also helps the original research team. Well-organised, well-documented data are easier to understand, share with collaborators, revisit months or years later, and build upon in future projects.


"Only completed datasets should be managed according to the FAIR Principles."
FALSE: FAIR is most effective when applied throughout a project rather than only at the point of publication. Organising and documenting data as they are created reduces effort later and improves day-to-day research workflows.

Key Takeaways

  • The FAIR Principles provide a framework for making research data easier to find, access, combine and reuse.
  • Plan for FAIR from the start of your project rather than waiting until publication.
  • The four FAIR Principles are complementary. Together they make data more useful for both people and computers.
  • Organise and document your data as you create them.
  • Use trusted repositories, appropriate file formats and clear conditions for reuse.
  • Remember that FAIR and Open are complementary but different concepts. Even restricted data can be FAIR when they are well described and accompanied by clear access conditions.
  • Small improvements can make data significantly more FAIR. Perfection is not required.
  • Making data FAIR benefits both the original research team and the wider research community by improving data organisation, collaboration, transparency and opportunities for reuse.

References

Foundational Documents

Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J., Da Silva Santos, L. O. B., Bourne, P., Bouwman, J., Brookes, A., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C., Finkers, R., . . . Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3. https://doi.org/10.1038/sdata.2016.18. Licence: CC BY 4.0

Research Evidence

Briney, K. A. (2024). Measuring data rot: An analysis of the continued availability of shared data from a Single University. PLOS ONE, 19, e0304781 - e0304781. https://doi.org/10.1371/journal.pone.0304781. Licence: CC BY 4.0

Berman, H. M., & Burley, S. (2025). Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society. Quarterly Reviews of Biophysics, 58. https://doi.org/10.1017/s0033583525000034. Licence: CC BY 4.0

Feng, Q., Lacey, B., Bešević, J., Omiyale, W., Conroy, M. C., Starkey, F., Calvin, C. M., Callen, H., Bramley, L., Welsh, S., Young, A., Effingham, M., Young, A., Collins, R., Holliday, J., & Allen, N. (2023). UK biobank: Enhanced assessment of the epidemiology and long-term impact of coronavirus disease-2019. Cambridge Prisms: Precision Medicine, 1. https://doi.org/10.1017/pcm.2023.18. Licence: CC BY 4.0

Fišar, M., Greiner, B., Huber, C., Katok, E., & Ozkes, A. (2023). Reproducibility in Management Science. Manag. Sci., 70, 1343-1356. https://doi.org/10.1287/mnsc.2023.03556. Licence: © 2023 INFORMS

Gorgolewski, K. J., Auer, T., Calhoun, V., Craddock, R., Das, S., Duff, E., Flandin, G., Ghosh, S. S., Glatard, T., Halchenko, Y., Handwerker, D., Hanke, M., Keator, D., Li, X., Michael, Z., Maumet, C., Nichols, B. N., Nichols, T. E., Pellman, J., . . . Poldrack, R. (2016). The brain imaging data structure, a format for organizing and describing outputs of neuroimaging experiments. Scientific Data, 3. https://doi.org/10.1038/sdata.2016.44. Licence: CC BY 4.0

Hurtado-Lopez, P. X., Tessema, B. B., Schnabel, S. K., Maliepaard, C., Van der Linden, C. G., Eilers, P. H. C., ... & Visser, R. G. F. (2015). Understanding the genetic basis of potato development using a multi-trait QTL analysis. Euphytica204(1), 229-241. https://doi.org/10.1007/s10681-015-1431-2. Licence: CC BY 4.0

Hyvönen, E. (2022). Digital humanities on the Semantic Web: Sampo model and portal series. Semantic Web, 14, 729-744. https://doi.org/10.3233/sw-223034. Licence: CC BY 4.0

Kaehrle, M., & Eschenfelder, K. (2025). Justifying Biodiversity Data Access Restrictions: A Global Comparison of Data Policies. Proceedings of the Association for Information Science and Technology, 62. https://doi.org/10.1002/pra2.1259

Munafo, M., Nosek, B. A., Bishop, D., Button, K., Chambers, C. D., Sert, N. P., Simonsohn, U., Wagenmakers, E., Ware, J., & Ioannidis, J. (2017). A manifesto for reproducible science. Nature human behaviour, 1. https://doi.org/10.1038/s41562-016-0021. Licence: CC BY 4.0

O’Neill, B., & Stapleton, L. (2022). Digital cultural heritage standards: from silo to semantic web. Ai & Society, 37, 891 - 903. https://doi.org/10.1007/s00146-021-01371-1. Licence: CC BY 4.0

Papoutsoglou, E. A., Athanasiadis, I., Visser, R., & Finkers, R. (2023). The benefits and struggles of FAIR data: the case of reusing plant phenotyping data. Scientific Data, 10. https://doi.org/10.1038/s41597-023-02364-z. Licence: CC BY 4.0

Piwowar, H. A., & Vision, T. (2013). Data reuse and the open data citation advantage. PeerJ, 1. https://doi.org/10.7717/peerj.175. Licence: CC BY 4.0

Tsueng, G., Alvarado Cano, M. A., Bento, J., Czech, C., Kang, M., Pache, L., Rasmussen, L., Savidge, T., Starren, J., Wu, Q., Xin, J., Yeaman, M., Zhou, X., Su, A. I., Wu, C., Brown, L., Shabman, R. S., & Hughes, L. D. (2022). Developing a standardized but extendable framework to increase the findability of infectious disease datasets. Scientific Data, 10. https://doi.org/10.1038/s41597-023-01968-9. Licence: CC BY 4.0

Váradi, M., & Velankar, S. (2022). The impact of AlphaFold Protein Structure Database on the fields of life sciences. PROTEOMICS, 23. https://doi.org/10.1002/pmic.202200128. Licence: CC BY 4.0

Research Infrastructures and Resources

AlphaFold. https://alphafold.com/

Brain Imaging Data Structure (BIDS). https://bids.neuroimaging.io/

Europeana. https://www.europeana.eu/en

RCSB Protein Data Bank. https://www.rcsb.org/

UK BioBank. https://www.ukbiobank.ac.uk/

WarSampo. https://www.sotasampo.fi/en/

DOI

Author: Jonathan England

Created: August 2026

Still have questions?

Contact us via our Helpdesk.
We try to respond within 48 hours.