Should My Research Data Be Open?
What is Open Data
Open data are data that anyone can freely access, use, modify, and share for any purpose (Open Knowledge Foundation, Open Definition 2.1). The goal of open data is to make information available so that it can be reused, combined with other data, and built upon to generate new knowledge, products, or services. Open data benefit a wide range of users, including researchers, businesses, governments, organisations, and the public.
However, simply making data publicly available does not necessarily make them open. To be considered open, data must satisfy both some legal and technical requirements.
From a legal perspective, the data should be released under an open licence (e.g. CC BY 4.0) or dedicated to the public domain (e.g. CC0 1.0). An open licence is a legal agreement that grants anyone permission to access, reuse, modify, and redistribute the data, including for commercial purposes, while imposing only limited conditions. These conditions typically include acknowledging the original creator (attribution), indicating when changes have been made, or requiring modified versions to be shared under the same licence (share-alike). Open licences must not restrict who can use the data or how they are used. Any use of the data must still comply with applicable laws and regulations (e.g. those relating to privacy, data protection, intellectual property, and defamation), respect the rights of others, and not falsely imply endorsement by the creator or rights holder.
Openness also depends on how the data are made available. From a technical perspective, the data should be provided in a machine-readable format, meaning that computers can automatically process the data without requiring manual extraction. For example, a CSV spreadsheet is machine-readable, whereas a scanned table saved as a PDF is generally not. Whenever possible, data should also be shared using open formats, whose specifications are publicly available and can be accessed using freely available software. This helps ensure that the data remain accessible over time and can be reused by a wide range of people and applications, without relying on proprietary software.
Open data are often associated with information produced by governments and public authorities. In Europe, open government data generally refers to public-sector information that is made available for anyone to reuse (European Union, Open Data Directive). However, not all publicly funded or publicly held data can be made open. Some datasets must remain restricted because they contain personal data, commercially confidential information, intellectual property, or information whose disclosure could pose security or ethical risks. Consequently, open-data policies seek to maximise openness while respecting legitimate legal, ethical, and security constraints. In the context of research data, the European Commission summarises this principle as making data "as open as possible, as closed as necessary" (European Commission, Horizon Europe Programme Guide (PDF, 1.1 MB)).
Why Open Data Matters
Why should researchers share their data? The following sections highlight some of the main benefits of doing so. These are not simply theoretical advantages. A growing body of empirical research has investigated the effects of making research data openly available. For example, a recent scoping review identified 67 studies on the academic impacts of Open/FAIR Data, with most evidence relating to data reuse, citation impact, reproducibility and research efficiency (Klebel et al., 2025).
It improves transparency and trust in research
Openly sharing research data allows other researchers to verify analyses, reproduce findings and identify potential errors. This greater transparency increases accountability, making it more difficult for errors and questionable research practices to go unnoticed. Although sharing data alone does not guarantee reproducibility, it is one of the foundations of trustworthy research.
Example: After Management Science introduced a mandatory data and code disclosure policy in 2019, only 12% of articles published under the previous voluntary approach provided replication materials. Under the new policy, when data were accessible, more than 95% of studies could be successfully reproduced. Across all articles, including those where data could not be accessed, the reproducibility rate was 68% (Fiŝar et al., 2024). The study highlights that mandatory sharing policies can substantially improve reproducibility, but also that accessible, well-documented data are essential.
It enables new discoveries through data reuse
Research data often remain valuable long after the original study has finished. When datasets are openly available, they can be re-analysed to answer new research questions, combined with other datasets for larger studies, or reused in systematic reviews and meta-analyses. This allows researchers to generate new knowledge without repeating data collection, accelerating scientific progress while making better use of existing resources.
Example: Piwowar and Vision (2013) analysed 10,555 gene expression studies and found that openly archived datasets were frequently reused by researchers who had not contributed to the original work. For every 100 datasets deposited, they estimated that around 40 third-party data reuse papers had been published within 2 years, 100 within 4 years, and more than 150 within 5 years. This demonstrates that a single shared dataset can support numerous independent studies long after the original research has been completed.
It saves time and resources
Collecting research data is often expensive, time-consuming and may involve considerable effort from participants. Reusing existing datasets reduces unnecessary duplication, allowing researchers to focus their time and resources on answering new questions rather than repeating work that has already been completed. This increases the overall efficiency of the research system.
Example: In an empirical study of ecological research, collecting a reusable dataset required a median of 30 person-days, whereas preparing the dataset for sharing required 5 person-days. Researchers reusing an existing dataset needed a median of only 3 person-days to discover, assess and integrate it into their work (Asghar & Daniel, 2024). Although these estimates are specific to ecology, they illustrate how reusing existing datasets can substantially reduce the effort needed to generate new research.
It increases the visibility of research
Making research data openly available increases the likelihood that other researchers will discover, reuse and cite both the dataset and the associated publication. Across multiple disciplines, observational studies consistently report a citation advantage for publications that make their underlying data openly available. Although these studies cannot demonstrate a direct causal relationship, the association has been observed across large and diverse samples.
Example: A study of approximately 122,000 publications found that articles with datasets shared in an online repository received an average 4.3% citation advantage compared with similar publications (Colavizza et al., 2024). Earlier studies reported even larger associations: publications linking to repository data received up to 25.4% more citations in an analysis of over 530,000 journal articles (Colavizza et al., 2019), while a controlled study of 10,555 gene expression papers found a 9% citation advantage after accounting for factors such as journal impact, open access status and author publication history (Piwowar & Vision, 2013). This suggests that making research data openly available can increase the visibility and reach of research outputs.
It creates value beyond academia
Open research data can benefit governments, businesses, educators, non-governmental organisations and the public. By making data openly available, researchers enable others to develop new services, inform public policy, protect the environment and address societal challenges. In many cases, the same dataset continues to support new applications and services long after the original research project has ended.
Example: Freely available Earth observation data from the European Union's Copernicus programme are used to monitor mangrove ecosystems in West Africa. Satellite observations allow environmental agencies and researchers to track changes in mangrove cover over time, assess restoration efforts and support conservation decisions across large areas that would be difficult and costly to survey using field observations alone. This example illustrates how openly shared research data can support environmental management and decision-making well beyond the research community.
Not Everything Must Be Open
Open data is an important objective of Open Science, but not all research data can or should be made publicly available. Researchers are responsible for balancing openness with other legitimate interests, including the protection of individuals, communities, organisations and society.
The European Commission summarises this principle as making research data "as open as possible, as closed as necessary" (Horizon Europe Programme Guide (PDF, 1.1 MB)). Openness should therefore be the default, but access may be restricted when there are valid legal, ethical, commercial or security reasons for doing so.
When data cannot be shared openly, this does not mean that they should be hidden completely. Researchers should make as much information available as possible, for example by:
- publishing metadata so others can discover the dataset;
- explaining why access is restricted;
- delaying open access where appropriate, for example while protecting intellectual property through a patent application or until associated research has been published;
- providing information on how access can be requested, where appropriate;
- sharing anonymised, aggregated or synthetic data where possible.
Common reasons for restricting access
Personal and sensitive data
Research involving identifiable individuals is subject to data protection legislation such as the General Data Protection Regulation (GDPR) in the European Union. Even after direct identifiers have been removed, datasets may still contain information that could allow individuals to be re-identified when combined with other sources. Depending on the level of risk, data may need to be anonymised, pseudonymised or made available only through controlled-access procedures.
Example: Clinical trial datasets containing patient health information are often shared through controlled-access repositories such as the European Genome-phenome Archive, where researchers must apply for access rather than downloading the data openly.
Commercial confidentiality and intellectual property
Research carried out with industry partners or involving commercially sensitive information may be subject to confidentiality agreements or intellectual property considerations. In some cases, researchers may also delay making data openly available while protecting inventions through patents or other forms of commercial exploitation. Once these obligations have been met, the data may be released openly or under specific access conditions.
Example:
Studies of synthetic biology researchers have shown that they balance the principles of Open Science with the need to preserve patentability and fulfil obligations to commercial partners (McLennan & Maslen, 2025). This is because, under Article 54 of the European Patent Convention, public disclosure before filing a patent application may destroy the novelty required for patent protection.
Security and public safety
Some datasets could present risks if released without restriction. Examples include information relating to critical infrastructure, cyber-security, dual-use research, or the precise locations of endangered species or culturally sensitive sites. Restricting access helps reduce the risk of misuse while still allowing legitimate research where appropriate.
Example:
Biodiversity databases such as the Global Biodiversity Information Facility (GBIF) protect endangered species by generalising or withholding precise location data while still making the remaining dataset and metadata openly available (Chapman, 2020).
Ethical responsibilities
Even where data sharing is legally permitted, ethical responsibilities may justify restricting access. Participants may have consented only to specific uses of their data, or sharing could increase the risk of harm, stigma or discrimination for individuals or communities. Researchers should ensure that data sharing remains consistent with ethical approvals and participant expectations.
Example:
In a study involving survivors of sexual assault, researchers found that participants wanted to retain control over how their interview data were shared, citing concerns about privacy, confidentiality and personal safety. The study highlights that protecting participants from potential harm may justify restricting access to qualitative research data, even within an Open Science framework (Campbell et al., 2023; Campbell et al., 2022).
Indigenous and traditional knowledge
Research involving Indigenous Peoples or local communities may include traditional knowledge, cultural heritage or community-held information that should not be openly shared without permission. Decisions about access should respect community governance, cultural protocols and any agreements established with knowledge holders.
Example:
The CARE Principles for Indigenous Data Governance recognise that Indigenous data should be governed according to the rights and interests of Indigenous communities. In practice, this may mean restricting or conditioning access to traditional knowledge and cultural data rather than making them openly available, balancing openness with Indigenous data sovereignty (Carroll et al., 2020; Jennings et al., 2023).
Levels of access
Not all datasets fall into a simple "open" or "closed" category. Different levels of access can be used depending on the sensitivity of the data.
| Access level | Description | Example |
|---|---|---|
| Open | Anyone can access the data immediately. | Environmental sensor measurements published in a repository. |
| Embargoed | Data become openly available after a defined period. | Experimental data released after publication of the associated article or after a patent application has been filed. |
| Restricted | Access is granted only under specified conditions. | Clinical research data available only to approved researchers. |
| Closed | The data are not shared, but metadata remain publicly available. | Highly sensitive data that cannot legally or ethically be disclosed. |
Key takeaway
Making data "as open as possible, as closed as necessary" does not mean choosing between openness and protection. It means making informed decisions that maximise the reuse of research data while respecting legal, ethical, commercial and security obligations.
Open and FAIR Are Different
Open Data and FAIR Data are often mentioned together, but they describe different concepts and should not be used interchangeably. Understanding this distinction is essential because many research policies, funders and repositories encourage researchers to make their data both Open and FAIR.
While both aim to maximise the value of research data, they address different questions. Open Data concerns the conditions under which data can be accessed and reused, whereas the FAIR Principles describe how data should be organised and described so they can be found, accessed, combined with other data and reused by both humans and machines (Wilkinson et al., 2016).
FAIR data are not always open
One of the most common misconceptions is that FAIR data must also be open. In reality, the FAIR Principles do not require unrestricted access to data. Instead, they recognise that access may need to be controlled for legal, ethical, commercial or security reasons. What matters is that the conditions for accessing the data are clearly described and that information describing the dataset (metadata) remains publicly available so that others can discover it and understand how access can be obtained (Wilkinson et al., 2016).
Example: The European Genome-phenome Archive (EGA) stores and shares genetic, phenotypic and clinical research data. Because these datasets often contain personal data, including special categories such as genetic and health information, they cannot be made publicly available. Instead, researchers must request access through a Data Access Committee. However, the datasets remain discoverable through publicly available metadata and clear access procedures, illustrating that research data can be FAIR without being open.
Open data are not necessarily FAIR
However, not all open data are FAIR. Simply making a dataset publicly available does not guarantee that others will be able to find, understand or reuse it. Without sufficient metadata, documentation, licensing information and provenance, openly available datasets can be difficult to discover, interpret, integrate into new studies or reproduce.
Example: A practical reanalysis study by Vadadokhau et al. (2026) illustrates this challenge. Six teams independently reanalysed six publicly available proteomics datasets using a common workflow. Missing sample relationships, processing parameters, reference files and inconsistent filenames produced substantial discrepancies in the reproduced results, including 13,068 versus 4,923 identified proteins and 108 versus 11 differentially expressed proteins in some analyses. The authors argued that such datasets can become "data tombs" (publicly accessible but not practically reusable) when essential metadata and documentation are missing.
FAIR and Open are complementary
Ideally, research data should be both Open and FAIR. Making data openly available increases opportunities for verification, reuse and collaboration, while applying the FAIR Principles ensures that the data are organised, documented and described in ways that maximise their long-term value. Open data that are not FAIR may be difficult to discover, interpret, integrate into new studies or reuse, whereas FAIR data that cannot be made open can still be discovered and accessed through appropriate procedures.
For this reason, many research funders and organisations promote the principle that research data should be "as open as possible, as closed as necessary" (European Commission, Horizon Europe Programme Guide (PDF, 1.1 MB)). This recognises that openness should be the default wherever possible, while acknowledging that some datasets require restrictions to protect privacy, confidentiality, security or other legitimate interests.
Regardless of whether data can be shared openly, applying the FAIR Principles helps ensure that they remain discoverable, accessible under appropriate conditions and reusable over time.
Historical note:
In 2010, Tim Berners-Lee introduced the Five-Star Open Data Model as a simple framework for encouraging organisations to publish data in increasingly reusable ways. Building on his earlier work on Linked Data (2006), the model describes a progression from simply making data available on the web under an open licence to publishing linked, machine-readable data using open standards.
The Five-Star Model remains influential, particularly in government open data initiatives, where it has helped promote the publication of interoperable data on the web. However, research data management today generally places greater emphasis on the FAIR Principles, which provide a broader framework for managing and describing research data so that they can be discovered, accessed and reused, regardless of whether they can be shared openly.
It is also important to recognise that the four- and five-star levels are not an expectation for most researchers. These higher levels rely on Linked Data technologies and semantic interoperability, where datasets are connected through shared vocabularies, ontologies and persistent identifiers. Such capabilities are typically provided by research infrastructures and knowledge graphs, including the European Open Science Cloud (EOSC), the OpenAIRE Graph, Wikidata, and discipline-specific infrastructures, rather than by individual researchers. For most researchers, publishing well-documented datasets in a trusted repository and applying the FAIR Principles represent good research data management practice.
Rating What improves? Research data example ★ The data are openly available. A dataset is published in a repository under a CC BY or CC0 licence, even if it is only available as a PDF. ★★ The data are structured. Experimental results are shared as a spreadsheet instead of a scanned table in a PDF. ★★★ The data use an open format. Tabular data are provided as CSV instead of XLSX. ★★★★ The data use standard identifiers. Researchers are identified using ORCID iDs, species using identifiers from a recognised taxonomy. ★★★★★ The data are linked to other data. A biodiversity dataset links species to recognised taxonomic databases, sampling locations to geographic identifiers, and related publications through persistent identifiers.
Common Misconceptions
"Open data means all research data must be made publicly available."
FALSE: Not all research data can or should be openly shared. Legal, ethical, commercial and security considerations may require restricting access. The aim is to make data "as open as possible, as closed as necessary".
"Publicly available means the data are open."
FALSE: Simply making a dataset available online does not necessarily make it open. To be considered open, the data should be accompanied by an open licence that clearly permits reuse and redistribution. Without explicit permission, others may be able to access the data but not legally reuse them.
"If my data are FAIR, they are automatically open."
FALSE: FAIR and Open are complementary but different concepts. FAIR data may remain under embargo or controlled access where justified.
"If my data cannot be open, FAIR does not apply."
FALSE: FAIR principles still apply to restricted datasets. Metadata, documentation, persistent identifiers and clear access conditions help ensure that restricted data remain Findable, Accessible under well-defined conditions, Interoperable and Reusable.
"Open data means giving up ownership or copyright."
FALSE: Sharing data openly does not mean giving up ownership. Researchers or their institutions generally retain ownership and intellectual property rights while granting others permission to reuse the data through an appropriate licence. Even when attribution is not legally required (for example, under CC0), citing the original dataset and acknowledging its creators remain established scholarly norms.
"Open data are only useful for reproducing studies."
FALSE: Open research data can be reused in many ways beyond reproducibility, including meta-analyses, teaching, benchmarking, method development, validation studies, machine learning, policy development and answering new research questions beyond the original purpose of the study.
"Sharing open data always means losing my competitive advantage."
FALSE: Researchers may legitimately delay data sharing through embargoes or controlled access while completing planned analyses or protecting sensitive information. Once shared, datasets can increase research visibility, foster new collaborations and receive formal citations, helping ensure that the original creators receive appropriate credit.
"Sensitive or personal data can never be shared."
FALSE: While some data cannot be made openly available, many sensitive datasets can still be shared responsibly through anonymisation, pseudonymisation, controlled-access repositories or data use agreements. Even when the data themselves cannot be shared, the metadata can often remain publicly available so others can discover the dataset and understand how access may be requested.
Key Takeaways
- Open data means making research data available for others to access, reuse and redistribute under an open licence.
- Research data should be as open as possible, as closed as necessary, balancing openness with legal, ethical, commercial and security considerations.
- FAIR and Open are complementary but distinct concepts. Data can be FAIR without being open, and openly available data are not necessarily FAIR.
- Openly sharing data does not mean giving up ownership or intellectual property rights. Appropriate licences define how others may reuse the data.
- Open research data support transparency, reproducibility, collaboration and innovation by enabling reuse in many different contexts.
- Even when data cannot be shared openly, publishing metadata and explaining access conditions helps make research more discoverable and reusable.
References
Foundational Documents
Berners-Lee, T. (2010). 5-star deployment scheme for Open Data. https://5stardata.info/. Licence: CC0
Berners-Lee, T. (2006). Linked Data. https://www.w3.org/DesignIssues/LinkedData.html
Creative Commons. CC0 1.0 Universal Deed. https://creativecommons.org/publicdomain/zero/1.0/. Licence: CC BY 4.0
Carroll, S. R., Garba, I., Figueroa-Rodríguez, O. L., Holbrook, J., Lovett, R., Materechera, S., … Hudson, M. (2020). The CARE Principles for Indigenous Data Governance. Data Science Journal, 19(1), 43. https://doi.org/10.5334/dsj-2020-043. Licence: CC BY 4.0
Creative Commons. CC BY 4.0 Attribution International Deed. https://creativecommons.org/licenses/by/4.0/. Licence: CC BY 4.0
European Commission. Horizon Europe Progamme Guide. https://ec.europa.eu/info/funding-tenders/opportunities/docs/2021-2027/horizon/guidance/programme-guide_horizon_en.pdf. Licence: CC BY 4.0
European Commission. Legal framework of EU data protection. https://commission.europa.eu/law/law-topic/data-protection/legal-framework-eu-data-protection_en. Licence: CC BY 4.0
European Patent Organisation. _Convention on the Grant of European Patents (European Patent Convention), Article 54 – Novelty. https://www.epo.org/en/legal/epc/2020/a54.html
European Union. Copernicus. https://eu-space.europa.eu/programmes/earth-observation-copernicus/copernicus. Licence: CC BY 4.0
European Union. (2019). Directive (EU) 2019/1024 of the European Parliament and of the Council of 20 June 2019 on open data and the re-use of public sector information (recast). http://data.europa.eu/eli/dir/2019/1024/oj. Licence: Reuse permitted under Commission Decision 2011/833/EU on the reuse of Commission documents.
Open Knowledge Foundation. Open Definition 2.1. https://opendefinition.org/od/2.1/. Licence: CC BY 4.0
Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J., Da Silva Santos, L. O. B., Bourne, P., Bouwman, J., Brookes, A., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C., Finkers, R., . . . Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3. https://doi.org/10.1038/sdata.2016.18. Licence: CC BY 4.0
Research Evidence
Asghar, A., & Daniel, F. (2024). Is Data Sharing Time-Efficient in Ecology? An Empirical Test and Extension of the Break-Even Reuse Model. Bull. Comput. Data Sci., 5, 1-11. https://doi.org/10.71448/bcds2454-1. Licence: CC BY 4.0
Campbell, R., Javorka, M., Engleton, J., Fishwick, K., Gregory, K., & Goodman-Williams, R. (2023). Open-Science Guidance for Qualitative Research: An Empirically Validated Approach for De-Identifying Sensitive Narrative Data. Advances in Methods and Practices in Psychological Science, 6. https://doi.org/10.1177/25152459231205832. Licence: CC BY-NC 4.0
Campbell, R., Goodman-Williams, R., Engleton, J., Javorka, M., & Gregory, K. (2022). Open science and data sharing in trauma research: Developing a trauma-informed protocol for archiving sensitive qualitative data.. Psychological trauma : theory, research, practice and policy. https://doi.org/10.1037/tra0001358. Licence: © 2022 American Psychological Association
Chapman, A.D. (2020). Current Best Practices for Generalizing Sensitive Species Occurrence Data. https://doi.org/10.15468/doc-5jp4-5g10. Licence: CC BY-SA 4.0
Colavizza, G., Cadwallader, L., LaFlamme, M., Dozot, G., Lecorney, S., Rappo, D., & Hrynaszkiewicz, I. (2024). An analysis of the effects of sharing research data, code, and preprints on citations. PLOS ONE, 19. https://doi.org/10.1371/journal.pone.0311493. Licence: CC BY 4.0
Colavizza, G., Hrynaszkiewicz, I., Staden, I., Whitaker, K., & McGillivray, B. (2019). The citation advantage of linking publications to research data. PLoS ONE, 15. https://doi.org/10.1371/journal.pone.0230416. Licence: CC BY 4.0
European Union. (2026). OBSERVER: How Earth Observation supports mangrove monitoring in West Africa. https://eu-space.europa.eu/news/observer-how-earth-observation-supports-mangrove-monitoring-west-africa. Licence: CC BY 4.0
Fišar, M., Greiner, B., Huber, C., Katok, E., & Ozkes, A. (2023). Reproducibility in Management Science. Manag. Sci., 70, 1343-1356. https://doi.org/10.1287/mnsc.2023.03556. Licence: © 2023 INFORMS
Jennings, L., Anderson, T., Martinez, A., Sterling, R., Chavez, D. D., Garba, I., ... & Carroll, S. R. (2023). Applying the ‘CARE Principles for Indigenous Data Governance’to ecology and biodiversity research. Nature ecology & evolution, 7(10), 1547-1551. https://doi.org/10.1038/s41559-023-02161-2. Licence: © 2023, Springer Nature Limited
Klebel, T., Traag, V., Grypari, I., Stoy, L., & Ross-Hellauer, T. (2025). The academic impact of Open Science: a scoping review. Royal Society Open Science, 12. https://doi.org/10.1098/rsos.241248. Licence: CC BY 4.0
McLennan, A., & Maslen, S. (2025). “Open science” meets commercial realities: a qualitative study of factors influencing sharing in synthetic biology research in Australia. Frontiers in Bioengineering and Biotechnology, 13. https://doi.org/10.3389/fbioe.2025.1604509. Licence: CC BY 4.0
Piwowar, H. A., & Vision, T. (2013). Data reuse and the open data citation advantage. PeerJ, 1. https://doi.org/10.7717/peerj.175. Licence: CC BY 4.0
Vadadokhau, U., Soliman, M., Castillon, L., Pastor Muñoz, P., Id, L., Gayathri, S. N., Srivastava, A., Runeberg, T., González-Armijos, T., Šapovalovaitė, K., Sakalauskaite, M., Adhikari, S., Abe, O., Tohmola, T., Li, H., Sundaresan, S., Vesikukka, H., Roininen, J., Zangene, E., . . . Jafari, M. (2026). Preventing Proteomics Data Tombs Through Collective Responsibility and Community Engagement. Scientific Data, 13. https://doi.org/10.1038/s41597-026-06614-8. Licence: CC BY-NC-ND 4.0
Research Infrastructures and Resources
European Genome-phenome Archive (EGA). https://ega-archive.org/
European Open Science Cloud (EOSC EU Node). https://open-science-cloud.ec.europa.eu/
Global Biodiversity Information Facility. https://www.gbif.org/
OpenAIRE Graph. https://graph.openaire.eu/
Wikidata. https://www.wikidata.org/
Author: Jonathan England
Created: August 2026