I4R.
← Back to BlogResearch

Not Wrong: The Replication Crisis is a Metatheory Crisis?

August 27, 2026 · Nate Breznau

What happens when you fill a conference room in Barcelona with social and behavioral scientists, a few policymakers and science administrators, and I lead a discussion about metatheoretical multiverse analysis with a 90-second timer? That’s what we found out at a recent Institute for Replication (I4R) workshop where I tried to flip the script from analytical robustness, coding discrepancies and replication packages, the usual I4R stuff, to something theoretically driving the reproducibility crisis that we often ignore.

Wanna know what that is? It’s theory.

We talk about reproducibility as a methodological problem. We debate p-hacking, publication bias, and analytical forking paths. But the elephant just chillin in the corner is that our theories are weak. There are too many other plausible explanations for what we observe and measure. Sure, the things we measure are complex, but without more clarity, even Commisioner Odenthal isn’t gonna find us a clue to what we are testing. Meaning that we don’t really know what to make of our findings. Other than hoping they are publishable and packaging them neatly to increase those chances. The findings are not right or wrong, they are not really interpretable. At least not with weak theory.

Part of the problem is that unlike the ‘harder’ sciences, consciousness is in the equation. We are measuring phenomena far more complex than quantum physics. Another part is simply that we do not spend much time on theory. In one of my favorite esoteric and less-mainstream open science writings, Anne Scheel pointed out “Why Most Psychological Research Findings Are Not Even Wrong”. Anne, I see your discipline specific arguments and raise you all social and behavioral science disciplines. We often don't know what the estimand is. We don’t know what we are looking at. If our theories don't explain the phenomenon in the first place, debating whether an empirical finding is statistically "right" or "wrong" is moot.

To confront this, I pitched a method I’m working on with my former doctoral student turned postdoc Hung H.V. Nguyen. It’s called Metatheoretical Multiverse Analysis (MMA) and I wanna kick some ‘theoric’ with it (cool huh? It rhymes with lyric and could mean theory in action…). What follows is a summary of the ensuing debate, full of interdisciplinary friction, economic curmudgeonism, and philosophy of science moments.

The Pitch: Metatheoretical Multiverse Analysis (MMA)

To understand why metatheoretical multiverse analysis is necessary, we have to look at how we currently treat plausible theoretical arguments about the data-generating model… Say what? I mean how we use and apply logic.

Imagine you are testing the effect of X1 on Y. But there is an unobserved confounder, X2, which causes both X1 and Y. If X2 is not in your test, your results are uninterpretable. You don’t know if X2 is confounding the results. And, don’t even get me started about X3. This might be a collider. But the theories are not clear in your subfield. So when we try to compare them we end up with several conflicts or unknown paths. These are all alternatively plausible theories and from them we have a multiverse of theory.

Now, imagine I have five equally plausible alternative theories explaining the same phenomenon. This means there’s a 20% chance the theory I am using is the correct one. But I cannot imagine any social or behavioral science where there are only five plausible theories. There are likely thousands. We only test five because it takes entire careers to write semantic theory. And since it takes 20-40 years for most of our discipline to forget those long-winded theories of the past maybe this is a sinking ship…. I digress. Let me re-gress instead (fancied opposite of digress which happens to be a statistical procedure too).

This is where Metatheoretical Multiverse Analysis (MMA) comes punching (or wrestling or kicking) in. If a standard multiverse analysis runs all reasonable empirical specifications for a given dataset, a metatheoretical multiverse analysis would logically compare all these theories. Somehow…. That’s where the idea gets a little sticky. Metatheory is mostly something people write about, rather than try to formalize and analyze with math. But if they could, our method would then show them where they need to invest their theory-building efforts.

I opened the floor. The timer started.

The Economist’s Dilemma: HARKing and the Illusion of Theory

The pushback was immediate, insightful, and brutally honest. The first counterargument highlighted just how difficult theoretical work is. As one applied economist noted, reading James Heckman makes you realize how brilliant deep theory can be, but spending your career trying to come up with sufficient conditions to definitively disentangle two competing theories is a great way to find yourself out of academia before you get tenure. Remember the long-winded argument?

But a more damning critique came from the reality of how theory is actually utilized in modern economics. To commit a “statistical sin” and assign causality, one researcher pointed out that the lack of robustness in our fields might stem from the fact that HARKing (Hypothesizing After Results are Known) is practically the norm.

In economics, the workflow rarely starts with a pristine a priori theory. Instead, a researcher finds an intriguing empirical pattern in the data, and then builds a formal mathematical model to justify the findings. We only take the time to do the exhausting math if we already have a paper we want to publish. Often, the formal model is something requested by Reviewer 2 at a top journal ex-post. Because the theory is engineered to fit the data, it doesn't improve our prior hypothesis in a meaningful way, nor does it guarantee robustness when exposed to new data.

Interestingly, someone from the I4R team chimed in with data to back this up. After checking roughly 15,000 robustness checks in the I4R database, they ran an AI classification to score papers from 0 to 100 on ‘how economic’ they were (i.e., true economic theory vs. a paper on TV habits published in an econ journal). The finding? There was absolutely no relationship between the "econ-ness" of the paper (the presence of formal modeling) and its robustness.

So, the means that even if we were to invest in theory development, we wouldn’t get anywhere because… basically, we suck at it.

The AI Revolution

These days there’s always gotta be something about AI. If not everything. The egomaniac AI that I asked to help me write this just couldn’t wait to point out how many people were talking about it.

Funnily we came to an AI discussion through a strong defense of the structural approach in economics (after the initial dust had settled and we could again breathe the cool Barcelona air-conditioned summer air). Economics, one participant noted, used to be purely theoretical because, prior to the 1970s, we simply didn't have the capacity to handle large data. The empirical revolution brought an era of data-mining into play. P-hacking came into full swing. The academic industrial complex had taken over thanks to secondary data spewing forth from society.

But the tide now is actually flowing back toward structural modeling. Thanks to… AI? The massive amounts of data we have today, combined with AI's unparalleled ability to explore patterns, is killing the human comparative advantage in pure empirical data mining. AI will always be better at finding patterns (at least an AI or data scientist tells us this, Gemini still can’t perform better regressions than I can IMO). The only place human researchers will retain a comparative advantage is in thinking, being creative, and structuring the data generating process. We don't need one perfect theory to describe reality anymore (not that that ever worked for us) good approximations are good enough now – and good approximations means…. Drum roll and someone on the mic saying “Ya’ll ready for this?!”. PREDICTION. The better we can predict things, the less we will need to explain them, they will just become some sort of facts in our life worlds. Won’t they?

Metatheoretical analysis might just be the structural framework we need to survive the AI transition. I mean, I invented it, so I’d like to think so.

Bias, DAGs, and Kung-Fu Nancy Cartwright

As the 90-second buzzer kept interrupting and resetting, the conversation evolved into the philosophy and sociology of science itself.

From a legal and equality perspective, an important question was raised about those “five theories” we tend to rely on (you know that when equally plausible reduce the chances that any one is correct down to 20%). Where do they come from? Historically, they have been generated by a very specific demographic—predominantly white men from the Global North. By restricting our empirical tests to a handful of established theories, we inadvertently perpetuate biases and ignore alternative paradigms that might emerge from the Global South. A metatheoretical multiverse approach, by automatically generating and considering thousands of models, might offer a mechanical antidote to this historical bias by forcing us to acknowledge the vast space of un-theorized realities.

But are DAGs really capable of saving us? The room had its doubts. Hey Mister Jack… I’m talking to you. The world, as one researcher passionately argued, is not a clean DAG with X1, X2, and Y. It has X15, Y7562, and a zillion unobserved mediators. Furthermore, DAGs can’t handle cyclical relationships and feedback loops the most famous relationship in economics, price and quantity, is entirely cyclical, endogenous.

This brought us to Nancy Cartwright and the philosophy of science. If we are mapping out thousands of theories, we must remember the Popperian ideal: a model must be falsifiable. If a theoretical model cannot be thrown out under certain conditions, it ceases to be a model and becomes a religion. Metatheoretical multiverse analysis is only useful if we have the empirical tools to actually falsify the branches of the multiverse we generate.

The crow cheers with Nancy’s MMA kick to the face.

The Preregistration Battleground

You cannot talk about theory and open science without stumbling into the debate on preregistration. I posed the question: ‘Does preregistration inappropriately constrain our theories?’ because I want to seem smart and provocative. If there are hundreds of theories, forcing a researcher to write down a specific one beforehand essentially chokes the multiverse before it can breathe. Choking is forbidden in MMA by the way.

The responses were polarized:

Some argued that preregistration forces a hypothetico-deductive model onto fields that generate knowledge inductively. In economic history or sociology, research is exploratory. You learn from the data. Preregistration actively harms inductive discovery.

Others pointed out that preregistration is incredibly difficult if not overrated for secondary data analysis. Datasets are messy, collected for non-research purposes, and require deep exploration just to understand how missing variables are coded. Recently someone pointed out on LinkedIn that my praise of Neumark is overrated. An MMA sweep kick, totally permissible in the sport - touché.

Conversely, a third group (there’s always a third group otherwise it feels incomplete) advocated that preregistration is simply a record of where you started. It prevents the ex-post invention of stories and increases transparency. If ‘Mostly Harmless Econometrics’ had a chapter telling students to pre-specify their hypotheses before opening the dataset, it would have transformed the culture of economics entirely. To late, the historical institutionalists won.

Where Do We Go From Here?

As we wrapped up the session (and prepared to face the blistering heat outside in the name of finding Fideuà), the consensus was clear: our methodological tools have far outpaced our theoretical foundations. We have built incredibly sophisticated empirical engines, but we are putting them in theoretical chassis that are fundamentally flawed. I liked this shift. Oh wait I kinda led the discussion there.

Metatheoretical multiverse analysis is not a magic bullet. It will not solve the fact that social science involves the unpredictable chaos of human consciousness, nor will it easily map the cyclical, non-DAG-friendly feedback loops of the global economy. But it is a start. It is a way to stop pretending that our opportunistic, post-hoc theories are the only valid models of reality. It might trigger some epistemological soul searching if nothing else. By mapping the vast space of plausible theories and systematically testing where they align and where they conflict, we can begin to rebuild the credibility of our disciplines from the ground up – in theory (which is probably weak, so take it with salt).

Furthermore, as the discussion highlighted, we need to completely overhaul how we incentivize work. Open science has an image problem we often make it look boring, framing it as an adversarial compliance checklist rather than a thrilling pursuit of truth. We leave it up to early-career researchers to awkwardly teach their supervisors about open code and data. And shoulder them with the onus of spending hours making all their work open and reproducible, hours that their forefathers (yeah, they were 95% men) didn’t need to do. Heck these forefathers could publish like 2 papers per year and get tenure.

But since its my blog post, I get the last word: If we want to fix the replication crisis, we cannot just mandate better code and policies. We have to foster better theory. We need to acknowledge the sheer size of the theoretical multiverse, embrace the complexity, and start theorizing before regressionizing. In this case, sadly, the answer is not as simple as 42.