When we talk about conducting replications, we often talk about tips and tricks how to do them better. Research groups around the world are building AI-powered replication pipelines to make computational reproducibility checks faster. As journals embrace open science and discussions increasingly turn to scaling replication efforts, the future of transparent science looks promising. However, there is one reminder I’d like to post: to self, to any new replicator coming across this page looking for wisdom, and to the community at large. Replications need one key safeguard: A replicator who, until proven otherwise, assumes that they are the one who might be wrong (rather than the original authors).
Don’t get me wrong. Everybody who has ever replicated a paper knows the rush of adrenaline when finding a discrepancy. The numbers do not match. The code is broken. The result is gone! AAA!
But before we call it a discrepancy (or coding error), we have a duty to ask a much harder question: What if the mistake is ours?
Notice the inherent asymmetry in replication: we are scrutinizing somebody else’s work. If they made a mistake, we correct the scientific record. If we made a mistake, we risk unfairly damaging somebody else’s reputation by questioning their competence, integrity, or both.
The reputational harm could be substantial. For this reason, I think replicators should hold themselves to an unusually high standard before concluding that someone else's work is wrong.
Of course, high confidence in a conclusion requires putting in the work. Thinking deeply. Considering alternative explanations. Yet, even replicators face the temptations available to other researchers:
1) delegation of (unpleasant) work to others, especially less experienced research assistants, without checking what was done and
2) relying on AI without verifying its output, especially the more complex bits like new code or interpretation of results.
Remember: It is always the replicator – not the RA, not the LLM – who is responsible.
Before saying that a claim fails replication, I think it is worth asking "What evidence would convince me that the mistake is mine?" Then actively look for it. Check my assumptions. Re-run my code. Contact the original authors where appropriate. It is time-consuming, and this effort usually goes unseen, but it is essential.
In the end, replication is an exercise in skepticism – but the first target of that skepticism should be ourselves.
