Reproducibility and Robustness in Development Economics (R2E-Dev) – A Large-Scale Meta-Replication
October 7, 2026 · Jörg Ankel-Peters, Abel Brodeur, and Julian Rose
The recent I4R Nature paper reported on a meta-replication of 110 articles from leading economics and political science journals and demonstrated that robustness can be scrutinized at scale – mostly through replications conducted during and after I4R Replication Games. We are now about to launch results from a new I4R project, R2E-Dev. It builds on this approach but takes a deliberately more focused and more coordinated route. We concentrate on one field—development economics—and place greater emphasis on standardizing how studies are selected, how robustness checks are identified and implemented, how their quality is assessed, and how replication reports are reviewed. Our guiding doctrine is to restrict the analysis to conservative robustness checks. We implement this doctrine by requiring robustness checks not only to be theoretically and statistically defensible, but also to remain close to the research question of the original study. The objective, clearly, is to come up with a higher bound in terms of robustness replicability in economics. We think this is helpful for the broader replicability debate in the social sciences because otherwise attention is immediately diverted to discussions about whether analytical choices are defensible or not.
R2E-Dev has conducted robustness replications of 66 empirical studies in development economics, randomly drawn from two volumes of the American Economic Review, the American Economic Journal: Applied Economics, and the Journal of Development Economics. Another innovation is that for each study, an independent replication fellow was hired. Yes, replicators were contracted and paid (2,500 EUR), and in exchange were expected to follow a tightly specified protocol (but all are very decent scholars and have internalized the project’s spirit – so we hardly had to enforce anything). The replication fellows examined the computational reproducibility, audited the original code, and implemented robustness stress tests along the lines described above. All robustness checks were pre-specified prior to the analysis. We, the R2E-Dev core team, reviewed each replication report and provided detailed internal referee reports to ensure the conservativeness doctrine. All R2E-Dev reports follow a standardized reporting structure, including a new visualization tool called the Robustness Dashboard.
This sounds like an epic endeavor, and some of you will ask: isn’t this something AI can do? Yes, indeed, AI will take over much of this work in the near future, and I4R is working with full force on the required tools. In fact, our project, R2E Dev, because of its heavily controlled nature will be an essential ground-truthing sample for this. More on this soon.
The R2E-Dev reports will be released gradually, with the meta-paper planned for early 2027. Along the way, we will publish a series of blog posts on the project, including a closer look at the availability and completeness of reproduction packages and the role of raw data, spotlights on selected individual replications, and a discussion of what makes a robustness check “conservative.” The series will culminate in the results of the meta-paper. To follow the project as it develops, subscribe to the I4R newsletter and follow I4R [X and Bluesky] and Jörg (X and Bluesky) on social media.
In the following, we will provide more details on the methodology.

66 papers, 66 replication projects
We first constructed a sampling frame of 162 eligible papers. Our main sampling years were 2020 and 2021. Papers had to belong to micro-development economics, use experimental or observational data, and estimate a causal effect. Where the 2020/21 pool did not provide enough eligible studies for the planned sample, we drew replacement papers from 2018 and 2019.
The final sample of replicated papers consists of 66 studies: 22 from each of the three journals, between studies using experimental and observational data.
Rather than having a small core team conduct all replications, we recruited external researchers as replication fellows. From around 170 applicants, we selected researchers ranging from PhD students and postdocs to professors. Each fellow is responsible for one replication project, receives the same €2,500 honorarium, and follows the same replication protocol. The protocol and its rationale are described in our Protocol for structured robustness reproductions and replicability assessments, published in Q Open.
This structure combines independent scrutiny with common incentives and procedures. Importantly, the incentive structure is deliberately designed to decouple researchers’ rewards from the outcome of the replication: neither the honorarium nor authorship depend on whether the replication uncovers problems or confirms the original findings. Replication fellows become first authors of their individual I4R replication reports and co-authors of the eventual meta-paper.
The protocol
A central challenge in robustness replication is that the replicator also has researcher degrees of freedom. If one searches long enough across specifications, it will often be possible to find an alternative specification under which a published result changes.
Our protocol is designed to restrict that flexibility.
Each replication follows the same broad sequence: identifying the study's main outcomes, computationally reproducing the original results, auditing the underlying code, and conducting a robustness replication. The robustness component asks which analytically relevant decisions made by the original authors have plausible alternatives.
These alternatives can concern, for example, control variables, outcome or treatment definitions, functional form and fixed effects, inference, or decisions made during data preparation. Depending on the original study, combining these choices can generate anything from a handful of alternative specifications to a much larger multiverse of analytical paths.
Importantly, the goal is explicitly not to find specifications that make a result disappear. Analytical choices should represent defensible alternatives that a researcher could reasonably have chosen when conducting the original analysis. The checks are pre-specified by the replication fellow before their results are evaluated, limiting the scope for reverse specification searching or “null hacking.”
The protocol therefore serves two purposes. Practically, it gives 66 independent researchers a common workflow and standardizes the structure of the replication process. Methodologically, it makes analytical choices more explicit and comparable across otherwise very different studies.
Visualizing robustness
Standardization also extends to how the results are presented. Across all replication reports, we use the Robustness Dashboard developed by Bensch et al. (2025). It provides a compact summary of the multiverse of analytical paths and complements specification curves, which allow readers to inspect the individual estimates and identify which analytical choices drive non-confirmatory results.
Using both consistently across the 66 replications serves two purposes. For an individual report, readers can quickly see whether sensitivity is concentrated in a few analytical paths or is more pervasive, and then use the specification curve to understand why. Across the project, the common visualization makes otherwise heterogeneous replications easier to read and compare.
This distinction is important for our approach. While all included robustness checks are intended to be reasonable and conservative, some are inevitably more conservative, and therefore less contestable, than others. The combination of the Robustness Dashboard and the specification curves helps readers quickly identify which analytical choices drive non-confirmatory results and assess how far these choices depart from the original specification.
What is a “reasonable” robustness check?
A protocol alone cannot resolve every judgment call. Whether an alternative specification is informative ultimately depends on the substantive question, the data, and the identification strategy of the individual study.
For this reason, the review stage is a central part of the project design.
Each draft replication report is reviewed by two members of the core team and, additionally, peer-reviewed by another replication fellow. The purpose is not simply to improve exposition. Reviewers scrutinize the proposed robustness checks and ask whether they are theoretically and methodologically justified.
We applied several deliberately conservative principles. Alternative specifications should essentially test the same theoretical claim as the original specification. Replicators needed to justify why each alternative is sensible. We excluded exercises that are primarily tests of treatment-effect heterogeneity or that altered the tested theory in other ways. We also avoided including several variations of the same analytical decision (e.g. trimming at 1%, 2% and 5%) in order to not artificially inflate the specification multiverse.
The result is inevitably more restrictive than allowing every replicator to decide independently what constitutes a useful reanalysis. That restriction is intentional. For the subsequent meta-analysis, we want sensitivity to reflect a multiverse of reasonable analytical alternatives, rather than differences in replicator objectives, effort, or preferences.
From individual replications to meta-evidence
Each replication is intended to stand on its own.
After review and revision, reports are shared with the original authors before public release and subsequently published as I4R Discussion Papers. These reports document the computational reproduction, relevant coding issues, the robustness choices implemented, and the resulting evidence for the individual study.
The common protocol, review process, and reporting structure allow us to combine evidence across all 66 studies. The individual reports therefore form the building blocks of a meta-paper examining the reproducibility and robustness of empirical research in development economics.
This two-level structure is central to the project. Individual replication reports provide detailed scrutiny of particular papers; the meta-replication asks what we can learn when the same type of scrutiny is applied systematically across a broader sample of published research.