When Replication Packages Reveal Too Much: Personally Identifiable Information in the Social Sciences
September 5, 2026 · Abel Brodeur and David Valenta

Open science has transformed research. Across the social sciences, journals increasingly require authors to share data and code so that published findings can be verified, reproduced, and extended. Replication packages are now a central part of transparent research practice.
But transparency comes with responsibilities. One of the most important is making sure that the materials researchers share do not inadvertently expose the identities of research participants.
In a new I4R Discussion Paper (https://www.econstor.eu/handle/10419/343373/), On the Prevalence of Personally Identifiable Information (PII) in the Social Sciences, we study how often personally identifiable information appears in publicly available replication packages. What we found is striking: 21.1% of the replication packages in our sample of articles relying on author-collected data contained some form of PII, with 14.4% containing direct PII.

Why We Studied This
This project grew directly out of I4R’s replication work. Through Replication Games and other I4R activities, we repeatedly encountered replication packages that appeared to contain personally identifiable information. In some cases, the problem was relatively straightforward: names, email addresses, or phone numbers had not been removed from a dataset that did not contain otherwise particularly sensitive information. In other cases, the disclosures were more serious, involving highly sensitive information that could plausibly identify participants and potentially expose them to harm.
In one instance, a replication package associated with a study conducted in a low-income country contained the names, aliases, and self-reported criminal activities of research participants and names of their accomplices, information that had been disclosed to researchers in confidence and never reported to authorities. Public disclosure of such information could potentially expose participants to legal jeopardy, retaliation, or physical danger.
These experiences raised an important question: were these isolated mistakes, or were they part of a broader pattern in how replication packages are prepared and shared?
This paper is our attempt to answer that question systematically.
What We Did
We assembled a dataset of 327 publicly available replication packages associated with articles published between 2020 and 2024 in 11 leading journals across economics, political science, and psychology. We focused on studies using author-collected individual-level data, including field experiments, laboratory experiments, online experiments, and surveys.
The replication packages came from repositories such as Dataverse, OpenICPSR, OSF, and Zenodo, as well as supplementary materials posted on journal websites.
To detect PII at scale, we developed a semi-automated screening procedure and combined it with manual review. The automated tool flagged variables that might contain identifying information, and each flagged case was then manually checked before being classified as containing PII.
We distinguish between:
- Direct PII, such as names, email addresses, phone numbers, precise locations, or Prolific/MTurk IDs;
- Indirect PII, where combinations of variables could allow re-identification;
- IP addresses, which we report separately.
Our Main Finding
The headline result is simple but important: PII is much more common in replication packages than many researchers likely realize.
Across the 327 packages we examined:
- 21.1% contained some form of PII;
- 14.4% contained direct PII;
- 7.0% contained indirect PII;
- 7.0% contained IP addresses.
These are likely lower-bound estimates, especially for indirect PII.
What Kind of PII Did We Find?
One of the most striking patterns in the paper is the range of identifying information found in publicly archived materials. Among the 46 packages containing direct identifiers, names were the most common and appeared in 24 cases. Other commonly present PII were platform (Prolific/MTurk) IDs and phone numbers which appeared in 14 and 12 cases.
In many cases, names and contact details appeared in places researchers may not have expected—especially in open-ended survey responses. Participants sometimes entered their own names, email addresses, or other personal details into free-text fields, even when those fields were not designed to collect identifying information.

We also found that IP addresses and location information were quite common and were likely captured automatically by survey platforms such as Qualtrics or SurveyMonkey and then remained in the publicly shared data.
Where Is the Likelihood of PII?
Not all research settings carry the same level of risk.
We find especially high PII prevalence in:
- field experiments (29.2%),
- online experiments/surveys (26.6%),
- compared with laboratory experiments (8.3%).
We also find large differences by country income group. Replication packages involving participants only from low- or lower-middle-income countries were substantially more likely to contain PII than packages involving participants only from upper-middle- or high-income countries.
Field experiments in lower income settings contained PII in 35.6% of cases and laboratory experiments in 25.0% of cases while in higher income settings they contained PII only in 18.2% and 7.1% of cases.

These patterns matter for both methodological and ethical reasons. Studies in lower-income settings often rely on more granular participant-level data and more complex fieldwork. At the same time, participants in those settings may face greater risks and weaker protections if identifying information is disclosed.

Differences Across Fields
Across disciplines, economics has the highest overall prevalence of PII in our sample:
- Economics: 25.9%
- Political science: 17.8%
- Psychology: 16.7%
But the story is more nuanced than a simple field ranking. Much of the variation appears to be tied to the types of data collected and the contexts in which data are collected. For example, field experiments and studies conducted in developing-country settings are more common in some journals than in others, and those types of studies also appear to face greater disclosure risks.

So while the raw differences across fields are informative, the broader lesson is that research design and data environment matter a great deal.
Also, these figures should therefore be interpreted as the prevalence of PII among the subset of articles meeting our eligibility criteria (e.g., having a replication package and author-collected data), rather than as an estimate for all articles or replication packages published by a specific journal. The proportion of such articles by journal varies wildly, ranging from about 7.5% in JDE to over 93% in JEPS.
What Happened After We Contacted Authors?
For each replication package in which we identified PII, we followed a structured notification process: we contacted the authors and copied the relevant journal editors-in chief and data editors when applicable.
The good news is that the response was largely constructive.
Within 20 days of notification, about 70% of flagged packages had either been removed so that the identified PII was no longer publicly accessible.
Of note, once a package has been downloaded, removing it from a repository does not eliminate all copies already in circulation. Some repositories also make it easier than others to correct problems or notify prior downloaders.
What Should Change?
One of the most encouraging aspects of this project is that many of the problems we identify are fixable.
For researchers, a few simple practices could go a long way:
- conduct a column-by-column audit before public deposition;
- separate contact information from analytical data;
- review open-ended text fields carefully;
- disable automatic IP and location capture unless truly necessary;
- apply stronger de-identification practices for granular geographic data;
- review not only final analytic datasets, but also raw, intermediate, and auxiliary files.
While we recognize the primary responsibility for protecting research participants rests with the authors who need to ensure the data they publish do not contain PII, there is potentially a role journals and repositories could play in mitigating the issue.
For journals, there is a clear opportunity to integrate PII screening into existing reproducibility workflows. Automated checks followed by targeted manual review could catch many of the most obvious problems before materials become public.
For repositories, better tools for automated screening, version control, withdrawal, and user notification could make both prevention and remediation more effective.
Our PII Detection Tool Is Publicly Available
We are currently developing a website to help with detecting direct (and to some extent indirect) PII. Have a look: PII Checker.
This website is designed for data editors, repositories and authors obtaining IRB consent to use the tool. Our tool is for replication packages that are already public or just about to be made public. A few data editors in our sample have already started using it. We are also working on a version of our tool that researchers could run locally.
Response from Journals, IRBs and Data Repositories
The response from journals, institutional review boards, and data repositories has already been encouraging. We shared the paper with editors-in-chief and data editors a few weeks ago, and a number of journals have expressed interest in incorporating our PII detection tool into their data-review processes. We have also shared our work with several IRBs, some of which are interested in encouraging researchers to use our tool—or similar tools—to check their data before making replication materials public. Major data repositories have responded as well. One repository founder told us that our findings were distressing and that they intend to take additional steps to reduce the risk of sensitive information being made publicly available.
There are also promising developments at journals where our results suggest that improvements are particularly important. The Journal of Development Economics (JDE), for example, has two new editors-in-chief who take both reproducibility and participant privacy seriously. We are meeting with one of them to discuss whether our tool could become part of the journal's data-review process. JDE is also introducing a data editor, creating an opportunity to integrate privacy checks directly into its reproducibility procedures. These early conversations make us optimistic that responsible data sharing can quickly become a standard part of the publication and reproducibility workflow.
Open Science and Privacy Should Go Together
The broader message of the paper is not that data sharing is a mistake. Quite the opposite.
Replication packages are essential for credible, reproducible research. But if open science is going to work well, it also has to protect the people whose information makes that science possible.
Our findings suggest that the challenge is not a fundamental conflict between transparency and privacy. It is a problem of implementation. The kinds of PII exposures we document often stem from recurring, predictable, and preventable failures.
That is why we think PII governance should become a routine part of the reproducibility process, not an afterthought.
As journals, repositories, and researchers continue to expand data-sharing practices, the goal should be clear: replication packages should be both reproducible and privacy-preserving.