The Bermuda Triangle of Open Science: Where Data and Code Disappear
July 20, 2026 · Ghina Abdul Baki

Somewhere between the author's hard drive, the journal's editorial pipeline, and the repository link in the availability statement lies a kind of Bermuda Triangle: a region into which promised data and code enter and are never seen again. Computational reproducibility runs on a scarce resource: researcher time. Every hour a reproducer spends searching this triangle for a dataset that was promised but never delivered is an hour not spent checking whether a published result holds. In my own experience, much of the friction in reproduction comes from something surprisingly mundane: vague, broken, or misleading claims about what materials are available and where. What follows is a reflection drawn from that experience, and a plea for a simple norm: say what you are actually providing, and mean it.
"Available upon request" is a promise rarely kept
The phrase "data available upon request" sounds cooperative. In practice, it is often where data goes to disappear. The routine is familiar to anyone who has done this work: you email the corresponding author, you wait, you send a polite reminder, you wait again, and more often than not the exchange simply trails off. It seems the promise was made to satisfy a submission checkbox.

If the data can be shared, share it. Post it in a repository at the moment of publication. This way you save the researcher time, and also your time searching for it in your dusty folders. And if the data genuinely cannot be shared — proprietary, confidential, restricted-access, or maybe you have just spent tons of hours to build it and do not want to share it — then say exactly that, plainly, in the availability statement. An honest "no" costs the reader thirty seconds. A hollow "upon request" costs weeks.
The blame relay: passing the request down the line
Sometimes the request does get a reply, only to enter a different kind of loop. The paper designates a corresponding author; the corresponding author responds, helpfully enough, that the data and analysis were handled by a coauthor, and that I should contact them instead. The coauthor never replies. The chain ends there, with each party having technically done their part: one answered, one was merely busy, and the availability statement remains formally true. Yet the practical effect is identical to silence, with the added cost of a second email and a second wait.

"Corresponding author" is supposed to mean the person who takes responsibility for the paper's claims, including the claim that the data is available. Passing the request down the authorship line turns a commitment into a relay race where the baton is always in someone else's hand.
Links that lead nowhere, and nobody checks
A second, quieter failure mode: the availability statement contains a link, and the link is a dead end. Sometimes the repository page is simply empty. Sometimes the link has rotted since publication. And in the strangest cases I have encountered, the "data availability" link redirects the reader back to the landing page of the paper itself, a perfect closed loop of performative compliance.

These failures are not all equivalent. Link rot is at least partly forgivable: personal websites move, institutional pages get reorganized. The fix is well known: deposit materials in archival repositories that issue persistent identifiers rather than linking to a homepage that will not survive the author's next job move. But a link that never pointed anywhere useful, or a page that was empty from day one, is different. Someone wrote that link into the manuscript, and someone at the journal approved it, and apparently nobody clicked it.
This raises the question that keeps recurring throughout this reflection: who audits this? Availability statements are, at most journals, self-reported and unverified. The researcher only discovers the truth after clicking, refreshing, checking their own connection, and slowly realizing the problem is on the other side.
Data without code, code without data
A related form of half-compliance is posting only one of the two things reproduction requires. It is worth being precise about what each half buys us, because the asymmetry is interesting.
Data without code makes reproduction difficult and sometimes impossible, especially when the empirical analysis is not fully described in the manuscript. But data alone is not worthless: it at least makes the detection of data problems, including fabrication, feasible, without a single line of the author's code.
Code without data is the mirror image. Nothing can be reproduced, because there is nothing to run the code on (short of simulating your own data just to watch the pipeline work). But it, too, has some residual value: careful readers can trace the code against the manuscript and detect coding errors, precisely where code and manuscript mismatch.

So each half permits one kind of audit. Only both together enable reproduction. If you can only provide one half, the least you can do is state clearly which half it is, rather than writing "replication materials available" or "data available" when it is only the code (and vice versa) and letting researchers discover, hours later, which crucial ingredient is missing.
When the author complied and the journal did not
The failures above sit with authors. But there is a failure mode that flips the usual story. I have come across cases where authors state, credibly, that they submitted their (links to) data and code to the journal at initial submission, and the journal simply never posted the materials. The files vanished somewhere in the editorial pipeline.
Who do we blame here? Not the author. There is a real irony in this: journals increasingly present their open science policies as a mark of rigor, yet the purely administrative task of posting a file can fall through the cracks.

What troubles me most is that this failure damages the wrong party. A researcher who finds no materials will naturally suspect the author and not the journal. A simple remedy suggests itself, with a role for each side. Journals already display "received / revised / accepted" dates on every article; they could just as easily display a timestamped record of when replication materials were received and posted. If the manuscript can be tracked, so can the deposit. Authors, for their part, could treat the data availability statement as part of the proofs, checking not only that the wording is right, but, once the article is live, that the link actually leads somewhere.
A closing thought
Looking back across these experiences, a common thread emerges. Every failure described here shares one feature: the cost primarily falls on the person trying to verify the work.

I do not think most of this stems from bad faith. Much of it is probably inattention. The original Bermuda Triangle owes its legend to mystery: ships and planes vanishing without explanation, beyond anyone's control. The triangle of open science enjoys no such excuse. Nothing here is swallowed by unknowable forces; materials disappear through untested links, unanswered emails, and unposted files, each one traceable, each one preventable. We do not need to solve a mystery for this triangle to disappear from the map entirely, and with it, the hours so many researchers lose sailing through it.