The Rhetoric Hack: How AI Judges Fall for Style Over Substance

We spend a lot of time worrying that AI will become too good at arguing. The more urgent problem may be what happens when we put it in the judge's chair.
A new paper on arXiv tests exactly that. Researchers took 120 real machine-learning submissions, used AI to rewrite them along six rhetorical dimensions while keeping the science identical, and then had AI reviewers score them. The result: the reviewers were systematically swayed by style. Same data. Same methods. Different words. Different scores.
This is not a quirk of academic peer review. It is a warning about any system—legal, financial, educational, political—where AI is asked to evaluate arguments.
The experiment: same paper, different wrapping
The setup from Li et al. is admirably clean. They started with 120 anonymized submissions to ICLR 2026 and generated 4,200 full-paper variants. Two large language models acted as "rewriters," shifting the prose along six rhetorical dimensions—including evidence framing, novelty stance, and scope framing—in opposite directions. Then five different large language models acted as "reviewers," scoring the manuscripts under both standard and strict review protocols.
The underlying research never changed. The numbers stayed the same. The methodology stayed the same. The conclusions stayed the same. Only the presentation moved.
And it mattered. Evidence framing and novelty stance produced the largest positive-negative contrasts in overall assessment. Scope framing formed a weaker but still meaningful second tier. The remaining dimensions had smaller or less stable effects. AI reviewers were not just noisy; they were noisy in predictable, exploitable ways.
The paper also tested more elaborate attacks: joint rewriting, recursive rewriting, and reviewer-guided rewriting. None reliably produced bigger gains. Joint rewriting depended heavily on which rewriter you used. Reviewer guidance did not consistently beat an unguided second pass. Repeated rewriting yielded diminishing, configuration-dependent returns. Across conditions, the rewriter mostly determined the gap between opposing variants, while the reviewer determined the size and direction of the score effect.
In other words, the hack is not infinite. But it does not need to be infinite to matter.
What "rhetorical sensitivity" really means
The authors call this "rhetorical sensitivity": the degree to which a reviewer's judgment shifts based on how a claim is dressed, independent of what the claim actually says.
For humans, this is ancient news. Aristotle divided persuasion into ethos, pathos, and logos precisely because we do not evaluate claims in isolation. We are moved by confidence, novelty, scope, stakes, solidarity, and status. A paper framed as "resolving a long-standing puzzle" lands differently than the same paper framed as "a minor refinement." A result presented as "surprising" feels more important than the same result presented as "expected."
We expect machines to be different. The whole pitch of AI peer review, AI essay grading, and AI debate judging is that software might strip away the prose and weigh the evidence. Less ego, less fatigue, less favoritism. More objectivity.
The experiment says: not yet. AI reviewers learned enough about scientific rhetoric to be moved by it, but not enough to see through it. They are sensitive to the same signals human reviewers are sensitive to, without the human safeguards.
Why this flips the usual AI debate worry
Most recent coverage of AI and argument focuses on one danger: AI is becoming too persuasive. It can change minds, mimic expertise, gish-gallop, flatter, overwhelm, and lie with confidence. We have covered some of that here.
That is real. But it is only half the story.
The other half is that AI is also too easy to persuade. Put an LLM in the evaluator role and it becomes the target, not the attacker. A skilled rhetorician—or another LLM trained to optimize rhetoric—can move the judge's score without changing a single fact.
This is the rhetoric hack. It is especially dangerous because it hides inside plausible-sounding evaluation. The output does not look wrong. It looks like a slightly better score for a slightly better-presented argument. Multiply that across thousands of papers, motions, briefs, or debates, and the cumulative effect is a system that rewards presentation over proof.
The courtroom echo
The warning is not theoretical. The same week the arXiv paper appeared, Josh Morrow, a partner at Lehotsky Cohn LLP, published a survey of AI-generated text in federal appellate opinions.
Morrow took roughly 2,250 published opinions from the regional U.S. courts of appeals in 2026 and ran them through Pangram, an AI-detection tool. More than 50 showed signs of AI-generated prose. For a baseline, he tested opinions from January 2022. Pangram flagged zero.
Morrow is careful. He does not name judges. He does not claim the opinions are defective. He even says AI drafting can sharpen thinking and prose. But he points to the central risk: writing an opinion is not just recording a conclusion. It is part of the reasoning. Organizing premises, confronting gaps, reconciling competing considerations, articulating why one argument prevails over another—those activities happen in the act of writing. If substantial passages arrive essentially finished and survive largely unedited, some of that rigor may leak out.
Then there is the deeper question. If AI can draft the opinion, what about the judgment itself? For now, the text cannot tell us. But the direction is clear: we are moving from "AI helps humans reason" to "AI reasons, and humans sign off."
That transition only works if AI reasoning is robust. The peer-review experiment suggests it is not.
The banker in the room
Goldman Sachs partner Chris Churchman made the point even more bluntly in a firm podcast this month. He runs Marquee, the bank's digital platform for institutional clients, and is watching AI consume the routine analytical work that used to train junior bankers.
"There's a huge danger here that in the era of AI, we outsource our reasoning to these models, and we have cognitive atrophy that stops us being able to reason from first principles ourselves," he said.
Then he landed on a line that should be printed on every AI output: "Look, in the end, I'm better at sounding thorough than being thorough."
That is the crux. AI is a style engine. It can make a weak argument sound balanced, a thin analysis sound comprehensive, a conventional idea sound novel. And when another AI evaluates that output, the evaluation itself becomes vulnerable to the same stylistic hacks.
The result is a closed loop: style produces style, and nobody checks the load-bearing premises. Churchman's worry about cognitive atrophy applies to organizations, not just individuals. If firms delegate first-principles reasoning to systems that are themselves fooled by rhetoric, they do not merely lose a skill. They lose the ability to notice that they have lost it.
Why humans are still part of the defense
None of this means AI evaluation is useless. Done well, it can surface inconsistencies, check coverage, compare against standards, and operate at a scale no human team can match.
But the experiment shows that AI reviewers need human oversight the way human reviewers need sleep: not as a luxury, but as a prerequisite for functioning.
The reason is adversarial imagination. A good human reviewer asks: "How could someone game this system?" They look for the missing control, the framing trick, the result that sounds too clean. They know that confidence is cheap and that novelty claims are often old ideas in new hats. They treat polish as a variable, not a virtue.
LLM reviewers do not have this reflex. They read what is there, not what is missing. They respond to the rhetorical register they were trained on, not the underlying rigor. Making them robust requires exactly what the arXiv paper implies: adversarial training, multiple reviewers, strict protocols, and constant testing for rhetorical reward-hacking.
Even strict review did not fix the problem. The authors found that a stricter protocol lowered the mean overall assessment by 1.36 points but did not consistently reduce rhetorical sensitivity. Tougher grading made reviewers harsher; it did not make them wiser.
How to read past the hack
The good news is that the same skill that protects you from human rhetoric protects you from AI rhetoric. It just needs to be applied more deliberately, because AI-generated prose can be faster, smoother, and more uniform than human prose.
First, rewrite the claim in the opposite direction. If a paper is framed as "resolving a long-standing debate," mentally reframe it as "a minor contribution to an established literature." If it is framed as "modest but rigorous," imagine it pitched as "groundbreaking." Does your evaluation change? If the conclusion flips with the framing, you were judging style.
Second, separate the evidence from the packaging. Strip away adjectives like "robust," "novel," "surprising," and "comprehensive." What is left? Is the methodology actually sound? Do the numbers support the interpretation? Is the sample large enough? Those questions do not depend on rhetoric.
Third, look for what is not there. A well-dressed weak argument often omits the strongest objection, hides the control condition, or treats correlation as causation without defending the leap. AI reviewers are bad at this. Humans do not have to be.
Fourth, use multiple judges. The arXiv paper found that the reviewer model determined the magnitude and sign of the rhetorical effect. Different reviewers had different blind spots. The same principle applies outside peer review. If one AI evaluation feels decisive, get a second one, ideally with a different prompt or model.
What this means for AI and argument
DebateAI exists because arguing well is a skill. A large part of that skill is judging arguments: separating strong claims from merely confident ones, spotting when evidence is being framed rather than presented, noticing when novelty is claimed and not earned.
The rhetoric hack tells us that AI can play both sides of this game. It can generate the polished argument and then, if asked, evaluate it—and be influenced by the polish it generated. It can be the debater and the judge, with the same stylistic bias running through both roles.
That does not make AI debate partners useless. It makes them sparring partners, not referees. They can force you to clarify, object, defend, and revise. They can simulate a hostile audience without the ego or fatigue. But they cannot yet be trusted to score the match on their own.
The skill we need to cultivate is the same one the experiment reveals is missing: the ability to read past the rhetoric and ask what would change if the prose were stripped away.
The harder skill
In the long run, the question is not whether AI can argue. It is whether we can still tell the difference between a well-argued claim and a well-dressed one.
That skill is harder than it sounds because rhetoric is not cosmetic. It shapes what we notice, what we remember, what we trust. A claim wrapped in the language of rigor feels rigorous. A finding framed as a breakthrough feels important. An argument that acknowledges its limits feels honest—and may therefore be more persuasive, even when the limits are manufactured.
The defense is procedural. Do not let one AI judge another AI without a human in the loop. Do not trust a single review score, whether the reviewer is carbon or silicon. Ask what the argument would look like rewritten in the opposite rhetorical direction. If the conclusion flips, you were not evaluating substance.
AI can help us think, but only if we keep doing the hardest part ourselves: deciding what counts as a good reason.
Related Posts

The Peer Pressure Machine: How AI Falls for Bad Arguments Under Pressure
New research shows GPT-4o can be talked into abandoning correct answers after just three turns of misleading persuasion. What AI's peer-pressure problem reveals about the skill of thinking for yourself.

The Explanation Trap: How AI Rationales Make Us Stop Thinking for Ourselves
New research shows that AI-generated rationales can degrade human judgment and cause cognitive atrophy. Why explanations that feel like reasoning may be the most dangerous AI output of all.

The Gish Gallop Machine: How AI Wins Arguments by Overwhelming You
AI chatbots are changing minds with a classic debater's trick: more claims than anyone can check. What the Gish gallop is, why AI is perfect at it, and how to defend yourself.
You just read the argument. Can you make one?
The AI takes the other side, every time. Three rounds, one scored verdict.
Argue today's Daily