The Rhetoric Bias: Why AI Reviewers Fall for Style Over Substance

Imagine submitting the same scientific paper twice. Same hypotheses, same methods, same data, same conclusions. The first version is rejected. The second is accepted. The only difference? In the second, you framed the evidence a little more confidently and called your contribution "novel" one more time.
If a human reviewer did this, we would call it sloppy. If an AI reviewer did this, we would call it a bug. But a new study shows it is exactly what happens when large language models review research. And the implications reach far beyond academic publishing.
AI is being handed the job of judging arguments. It reviews papers, scores essays, ranks proposals, and—on platforms like DebateAI—argues both sides of contested questions. We assume that because AI is fast, consistent, and seemingly objective, it can separate good reasoning from bad. The new paper suggests we should not assume that at all.
The study in brief
In "How Can Rhetoric Reward-Hack AI Reviewers?" researchers built a controlled experiment to test whether rhetorical changes alone could change an AI reviewer's mind. They started with 120 real submissions to ICLR 2026, a top machine-learning conference. Then they used LLMs to rewrite those papers along six rhetorical dimensions—things like evidence framing, novelty stance, scope framing, and confidence tone—without changing the underlying results.
Then they had five different LLM reviewers score the original and rewritten versions. Across 4,200 manuscripts, the verdicts shifted. Not randomly. Systematically.
The biggest levers were evidence framing and novelty stance. Papers that presented their evidence more assertively, and that claimed their contribution was more groundbreaking, got better overall assessments. Scope framing mattered too, though less. Other dimensions had smaller or less stable effects.
The effect was not just academic. Rewriting changed accept-reject decisions. When the same paper was dressed up in stronger rhetoric, low-scoring papers tended to rise. High-scoring papers sometimes fell—perhaps because the heightened tone triggered more skepticism in already-critical reviews. The clearest swings happened in the middle, where the decision boundary lives.
Why this should unsettle us
Peer review is supposed to be about substance. Does the method hold up? Are the claims justified? Is the contribution real? We outsource those questions to AI because we hope it will be more consistent and less tired than humans. We imagine it reading the paper the way a careful scientist would: checking the proof, comparing baselines, weighing limitations.
Instead, the AI reviewer responded like a human—but worse in one important way. Humans can be swayed by a confident abstract or a punchy narrative, but they can also notice when style is substituting for content. They might read the methods. They might ask a skeptical follow-up. The LLM reviewers in this study had no such escape valve. They scored the paper in one pass, and their scores moved with the rhetoric.
The paper also tested whether stricter prompting would help. It did not consistently. A strict review protocol lowered average scores by more than a point, but it did not reliably reduce rhetorical sensitivity. Telling the AI to be more careful made it harsher; it did not make it better at telling style from substance.
The same problem hides everywhere
This is not just about conference submissions. The same AI review architecture is being sold into education, hiring, grant evaluation, legal research, and content moderation. A student submits an essay. A grant panel runs proposals through an LLM summary. A hiring tool screens cover letters. In each case, the system is asked to judge the quality of an argument. And in each case, the system is vulnerable to the same rhetorical hacks.
A confident but wrong claim can outscore a cautious but right one. A proposal that promises revolution can beat one that promises incremental progress, even when the incremental work is more likely to succeed. An essay with the right emotional arc can receive a higher mark than one with stronger reasoning but plainer prose.
This is the opposite of the dream. We wanted AI to strip away charisma, polish, and status and evaluate the argument underneath. Instead, it appears to reward the polish even more predictably than humans do, because it cannot access the deeper signals humans use—collegial reputation, methodological intuition, the smell test—to compensate.
Why AI reviewers are especially vulnerable
A tired human reviewer might skim the abstract, be charmed by a clean narrative, and miss a flaw in the methods. But that same reviewer can also be provoked by overreach. A paper that claims too much can trigger skepticism. A familiar author with a history of shaky claims can set off alarms. A reviewer might forward the paper to a colleague who knows the literature and say, "Does this seem right to you?"
An AI reviewer has none of these defenses. It has no memory of being wrong before. It has no colleague in the next office. It cannot run the experiment again or check whether the code matches the prose. It sees only the text in front of it, and it responds to the patterns in that text.
Those patterns include the statistical fingerprints of success. Papers that get accepted tend to use certain verbs, frame their evidence strongly, and emphasize novelty. The AI learns those fingerprints and applies them. The result is a reviewer that rewards the costume of good science even when the body inside is different.
The study's strict-prompt test is revealing here. When the researchers told the AI reviewer to be more rigorous, the average score dropped by 1.36 points. But the rhetorical sensitivity did not reliably disappear. The AI became harder to please across the board; it did not become better at seeing through rhetorical tricks. Stricter instructions changed the threshold, not the judgment.
That matters for anyone building AI evaluation tools. You cannot fix this problem by telling the model to be more careful. Carefulness is not the same as discernment.
The hierarchy of rhetorical leverage
The study's most useful finding is not that rhetoric matters; it is which rhetoric matters most.
Evidence framing came first. How you describe your support—"we demonstrate," "we prove," "we establish" versus "we suggest," "we observe," "we find that"—moved overall assessments more than anything else. The AI reviewer rewarded certainty even when the data were identical.
Novelty stance came next. Calling your work a "fundamental shift" rather than an "incremental improvement" changed scores. This is especially worrying in science, where most real progress is incremental and most "revolutionary" claims are overblown.
Scope framing followed. Papers that broadened or narrowed their claimed contribution moved the needle, though less reliably. Confidence tone, clarity emphasis, and limitation framing had smaller or more inconsistent effects.
This hierarchy is a map of what AI reviewers cannot currently see. They cannot tell whether "we demonstrate" is earned. They cannot tell whether "novel" is true. They can only tally the words and respond to their statistical associations with high-quality papers in the training data.
The regression-to-the-mean effect
One of the strangest findings: the direction of rhetorical influence depended on where the AI reviewer started. Low-scoring papers tended to rise. High-scoring papers tended to fall. The biggest contrasts appeared in the middle of the score range.
This looks like a kind of automated regression to the mean. The AI reviewer, unsure what to do with a paper, lets the rhetoric pull it toward a safer central verdict. A weak paper that sounds confident gets a bump. A strong paper that sounds too confident gets pushed down. The reviewer is not evaluating the argument so much as adjusting its presentation toward a default.
For researchers, this creates a perverse incentive. If your work is genuinely weak, you should write it boldly. If your work is genuinely strong, you should write it modestly. The AI reviewer becomes not a judge of quality but a corrector of rhetorical calibration.
What debate training gives you
If AI reviewers are this easy to hack, the natural response is to train people to hack them. Write the abstract the algorithm wants. Front-load the novelty claim. Use the confident verbs. That will work, for a while, but it is a losing arms race. The algorithms will change, and the people who spent years optimizing for them will have optimized for the wrong thing.
The better response is to become the kind of reader—and writer—who can see through the hack.
That is what debate training builds. A competitive debater learns to separate the rhetorical frame from the actual argument. They can listen to a case that sounds devastating and immediately ask: what is the causal mechanism? What evidence is missing? What would the other side say? They can also take a poorly framed strong argument and reframe it so its strength becomes visible.
The skill is not prompt engineering. It is argument literacy. It is knowing that "we demonstrate" is not the same as demonstrating, that "novel" is a claim to be evaluated, not a fact to be accepted, and that confidence is a presentation choice independent of truth.
You can practice it on almost any AI-generated summary. Read a paragraph and ask: if the strongest critic rewrote this in the least favorable light, what would change? Which adjectives would disappear? Which hedges would appear? Then ask whether the underlying claim still stands. If it does, the argument is solid. If it collapses, you have found a rhetoric-shaped void.
This is the mental motion the AI reviewer cannot perform. It requires holding two versions of the same argument in your head at once: the confident presentation and the stripped-down claim. It requires asking not "does this sound right?" but "would it still be right if it were said badly?" That is the standard good judges use, and it is the standard we need to teach.
The honest standard
None of this means AI has no place in evaluation. Used well, it can summarize, check consistency, flag missing citations, and surface patterns across thousands of submissions. But the new research draws a clear line: AI should not be the final judge of an argument's quality until it can reliably distinguish rhetorical performance from substantive merit.
Until then, the honest standard is human-in-the-loop review. Let AI do the first pass. Let it sort, summarize, and highlight. But keep the final verdict human—and train those humans to spot the rhetorical tricks that fool both them and the machine.
This is also why watching AI argue matters. On DebateAI, you do not just see one AI answer. You see two AIs argue against each other, expose weak points, and defend their positions under pressure. That structure does not eliminate rhetorical bias, but it surfaces it. A claim that sounds persuasive in isolation looks different when another model is actively looking for the gap between confidence and evidence.
Conclusion
We are rushing to put AI in the judge's chair: of papers, of essays, of arguments, of ideas. We are doing it because we hope AI will be fairer, faster, and more consistent than we are. The new research is a warning that fairness is harder than speed.
The AI reviewer is not malicious. It is not even particularly unusual. It is a pattern-matching engine that has learned, from human writing, that confident claims and novelty claims tend to appear in good papers. So it rewards them. And because it cannot step outside the text to ask whether those claims are justified, it rewards them blindly.
That is not judgment. That is mimicry. Real judgment requires the ability to hold a claim up against the world and see if it holds. AI cannot do that yet. The best we can do, while we wait, is to keep humans in the loop and to teach the humans to be better judges than the machines they are supervising.
Style should never be enough. For AI reviewers, it still is.
Related Posts

The Rhetoric Hack: How AI Judges Fall for Style Over Substance
New research shows AI peer reviewers can be reward-hacked by rhetorical style alone—no facts changed. What that means for the dream of using AI to evaluate arguments.

The Peer Pressure Machine: How AI Falls for Bad Arguments Under Pressure
New research shows GPT-4o can be talked into abandoning correct answers after just three turns of misleading persuasion. What AI's peer-pressure problem reveals about the skill of thinking for yourself.

The Explanation Trap: How AI Rationales Make Us Stop Thinking for Ourselves
New research shows that AI-generated rationales can degrade human judgment and cause cognitive atrophy. Why explanations that feel like reasoning may be the most dangerous AI output of all.
You just read the argument. Can you make one?
The AI takes the other side, every time. Three rounds, one scored verdict.
Argue today's Daily