A growing push to share health and research data is running into a problem that cannot be solved by technology alone: people may stop trusting the institutions collecting and distributing their information. In a new analysis published in Nature Communications, researchers Kristoffer Hoeyer and Lene Skovgaard examine what happens when data-sharing practices lose public legitimacy—and what institutions can do before that loss of trust becomes irreversible. Their central message is direct: the question is not simply whether data can be shared, but whether sharing is considered justified by the people whose lives are represented in the datasets.
The distinction matters because modern science increasingly depends on combining information from hospitals, laboratories, government registers, wearable devices, genomic databases and online platforms. Linking these sources can reveal patterns that are invisible in isolated datasets, from disease risk and treatment outcomes to population-level responses during an epidemic. Technically, data integration can improve statistical power and enable researchers to identify correlations across millions of records. Yet the same process can make information travel far beyond the context in which it was originally collected. A blood sample donated for one study may later contribute to another project; administrative records gathered for public services may become valuable to commercial or international research networks. Each additional use can create scientific value, but also raise questions about authority, consent and control.
Hoeyer and Skovgaard focus on the gap between formal permission and social acceptance. A data-sharing system may comply with privacy law, obtain approval from an ethics committee and use encryption, while still appearing illegitimate to the public. Legitimacy is broader than legal compliance: it concerns whether people believe decisions are made fairly, transparently and for defensible purposes. This distinction becomes especially important when consent is broad, indirect or difficult to interpret. Traditional informed consent assumes that a person can understand the purpose of a project before deciding whether to participate. In large-scale data ecosystems, however, future users, research questions and commercial partners may not yet be known when data are collected.
That uncertainty has helped drive the development of technical safeguards such as de-identification, pseudonymization, access controls and secure data enclaves. De-identification removes direct identifiers, while pseudonymization replaces names with codes that can be linked back under controlled conditions. Secure enclaves allow approved researchers to analyze data without downloading raw files. These measures reduce certain risks, but they do not eliminate them. Data can sometimes be re-identified by combining supposedly anonymous records with other information, particularly when datasets contain rare diseases, precise locations, dates of birth or detailed genetic signatures. More fundamentally, privacy protection does not answer the question of whether a use is socially acceptable. A dataset may be technically secure and still be used in a way that communities regard as exploitative or unfair.
The researchers’ argument places data governance at the center of the debate. Governance includes the rules, institutions and decision-making processes that determine who may access data, for what purposes and under which conditions. It also includes mechanisms for auditing use, responding to complaints and imposing sanctions when agreements are breached. Effective governance must therefore be more than a one-time approval process. It needs to operate throughout the life of a dataset, from collection and storage to linkage, analysis, publication and eventual deletion or archiving. This lifecycle approach recognizes that risks can change as data are combined with new sources or applied to new policy and commercial questions.
One of the most difficult issues is the unequal distribution of power. Individuals often provide data to institutions they cannot realistically negotiate with, such as national health services, schools, employers or digital platforms. The benefits of data sharing may be distributed across society, while the risks—including discrimination, surveillance or loss of autonomy—can fall disproportionately on particular groups. Communities that have historically experienced medical abuse or state monitoring may reasonably demand stronger protections than a generic consent form provides. For these groups, promises of public benefit may not be persuasive unless they are accompanied by meaningful participation, independent oversight and evidence that benefits will be shared fairly.
The analysis also challenges the idea that public trust can be repaired mainly through better communication. Clear explanations are essential, but transparency alone cannot compensate for decisions that people have no opportunity to influence. Institutions may publish privacy policies and technical documents while leaving the public unable to question the underlying purposes of data use. A more durable approach could involve public deliberation, patient and community representatives on governance bodies, accessible information about who is using data, and channels through which affected people can contest decisions. These mechanisms do not make disagreement disappear, but they can make data systems more accountable and help transform passive subjects into participants in governance.
For scientists, the debate has practical consequences. Public legitimacy is an operational condition for research, not merely an ethical ideal. If people fear that their information will be sold, shared with authorities or used without meaningful oversight, they may refuse to participate, withdraw consent where possible or provide incomplete information. Such behavior can introduce selection bias, reducing the representativeness and scientific value of datasets. Researchers may also face delays, legal disputes and restrictions that make valuable studies harder to conduct. By contrast, trustworthy systems can support sustained participation and improve the quality of the evidence used in medicine and public health. The challenge is to design access rules that protect people without making data so difficult to use that socially valuable research becomes impossible.
The authors’ perspective arrives as artificial intelligence intensifies pressure to make more data available. Machine-learning systems can identify complex patterns in medical images, electronic health records and genetic data, but their performance often depends on large and diverse training datasets. This creates incentives to pool information across institutions and borders, sometimes involving private companies and opaque computational models. Yet an algorithm can magnify existing inequalities if its training data underrepresent certain populations or if its predictions are used in high-stakes decisions without adequate review. Technical sophistication does not remove the need to justify why data are collected, who benefits from the resulting system and how errors will be addressed.
The broader lesson is that responsible data sharing requires institutions to treat legitimacy as something that must be continuously earned. That means matching scientific ambition with proportionality, limiting uses that exceed the expectations attached to collection, protecting vulnerable communities and creating real routes for public oversight. It also means acknowledging that some data uses may be legally permissible but politically or ethically unacceptable. As health systems and research organizations build increasingly interconnected data infrastructures, the future of discovery will depend not only on faster computation or larger databases, but on whether people believe these systems respect their rights. The success of data-driven science may ultimately hinge on a deceptively simple question: who gets to decide what happens to information about human lives?
Subject of Research: Public legitimacy, governance and responsible sharing of health and research data
Article Title: What is the problem and what can be done when data sharing challenges public legitimacy?
Article References: Hoeyer, K., Skovgaard, L. “What is the problem and what can be done when data sharing challenges public legitimacy?” Nature Communications (2026). https://doi.org/10.1038/s41467-026-77006-0
Image Credits: AI Generated
DOI: 10.1038/s41467-026-77006-0
Keywords: Data sharing, public legitimacy, data governance, privacy, informed consent, health data, research ethics, public trust, artificial intelligence, accountability

