I4R.
← Back to BlogPublishing

The Correlated Referee

August 24, 2026 · Ghina Abdul Baki

The Correlated Referee

If every journal (or any committee that makes a decision) screens work with the same AI, the reviewers stop being independent, and independence was the point.

Fig. 1"Three referees. One of them."

This month, the American Economic Association and the Econometric Society began using Refine, an AI tool that checks manuscripts for errors in reasoning, calculation, and references, as part of how they publish. The framing has been deliberately narrow: a technical proofreader, not a referee; applied late in the process; never the thing that decides a paper’s fate. Authors in the pilot were overwhelmingly in favor.

Taken at face value, that’s welcome. Catching a miscalculated standard error before it reaches print helps everyone. But peer review is not only a filter for errors; it is the mechanism by which a field decides what counts as knowledge. And once an AI sits anywhere in that mechanism, it is worth asking a question: what happens when everyone starts using the same one?

Why review works when it works

Fig. 2"Disagreement: the last renewable resource."

Peer review is resilient for an unflattering reason. It runs on many fallible opinions that do not agree with each other. Three referees read a paper three ways. It is rejected at one journal and welcomed at another. The strange result that one reviewer waves off is the one another finds important enough to fight for. That disagreement is not a defect to be engineered away; it is what lets an unusual, unfashionable, or genuinely new paper survive somewhere long enough to prove itself. The reviewer who champions the odd paper against the consensus is the system’s escape valve.

Independence is doing the quiet work here. The value of a second opinion comes entirely from its being a second opinion, uncorrelated with the first. Remove the independence and you do not have three referees anymore. You have one, consulted three times.

What a shared tool actually does

Fig. 3"Five doors. One arm. Same verdict, everywhere at once."

Kleinberg and Raghavan (2021) formalized the cost of this and proved something sharper than the obvious worry. The obvious worry is that if everyone relies on one system, everyone is exposed when it fails. Their 2021 result is more unsettling: even with no failure at all, under entirely normal operation, the overall quality of decisions can fall when everyone adopts the same algorithm, even if that algorithm is more accurate than any alternative each decision-maker could have used on their own. They call it algorithmic monoculture. A better tool, universally adopted, can make the collective outcome worse, because good options are no longer rejected by one judge here and rescued by another there. They are rejected by the same judgment everywhere at once.

The mapping onto peer review is almost exact. Let the papers be the candidates and the journals the decision-makers. Once they lean on a shared AI to gauge whether a paper is sound, the referees become correlated reads of a single system rather than independent minds. A paper that system happens to underrate (an uncommon method, an out-of-fashion specification, a finding that runs against the priors baked into the model) is not turned down at one journal and picked up at the next. It is turned down at all of them, for the same reason, simultaneously. The escape valve closes.

This concern is not confined to journal publication. It applies anywhere a decision is meant to rest on independent reviewers, such as grant panels, tenure and promotion cases, conference selection, and thesis committees among them. Wherever we build a shared AI to help those readers judge, we risk turning several independent opinions into correlated reads of one system, and in the less formal settings there is no pilot, no announcement, and no policy to slow it down.

Extend that a few years and the literature starts to look strangely uniform: papers drift toward the specifications the tool prefers and the results it finds plausible. The cause is not laziness or capture. It is that everyone behaved sensibly and used the best instrument available. This is the part worth holding onto, the homogenization does not require the tool to be bad. It follows precisely from the tool being good and everyone, reasonably, trusting it.

Correctness vs. taste

Fig. 4"Left of the line, arithmetic. Right of it, opinion in a lab coat."

The argument only holds if we are careful about where it applies, so let me draw the line myself. Some judgments have a right answer. Arithmetic does. A proof holds or it does not. A citation points where it claims to or it does not. On these, universal convergence is not monoculture; it’s just correctness, and I want it. If every journal runs the same check and catches the same real error, good.

The damage is on the other side, in the soft calls that wear technical clothing: Is this robustness check persuasive? Is this the right specification? Is this result interesting enough to publish? These are exactly the questions where competent experts are supposed to diverge, and where correlated judgment quietly narrows the field. So the concern was never the proofreader. It is the drift: the natural, well-meaning slide from “the tool caught a mistake” to “the tool wasn’t convinced,” from verifying arithmetic to arbitrating taste. Every tool in this space is one design decision from that line.

What to protect

Fig. 5"Keep one mind in the room that hasn’t been consulted."

I’m not against AI in review. It catches things we miss, and I would rather it caught them. What I want to defend, deliberately and before it quietly erodes, is the one property that made review trustworthy in the first place: independence. The dissenting read that lands somewhere instead of being filtered out everywhere. That probably means resisting the convenience of routing every judgment through a single oracle, treating “the tool approved it” as the beginning of a decision rather than its end, and keeping more than one genuinely independent mind in the room.

Fig. 6"One thought, beautifully distributed."

A field that reviews itself with one mind will, in time, only be able to think one thing.

References

Jon Kleinberg and Manish Raghavan, “Algorithmic Monoculture and Social Welfare,” Proceedings of the National Academy of Sciences 118, no. 22 (2021): e2018340118, https://doi.org/10.1073/pnas.2018340118.