
The problem of detecting AI-tool use in peer assessment is proving problematic.Credit score: BrianAJackson/iStock by way of Getty
It’s virtually not possible to know whether or not a peer-review report has been generated by synthetic intelligence, in keeping with a research that put AI-detecting instruments to the take a look at.
A analysis group primarily based in China used the Claude 2.0 massive language mannequin (LLM), created by Anthropic, an AI firm in San Francisco, California, to generate peer-review stories and different sorts of documentation for 20 revealed cancer-biology papers from the journal eLife1. The journal’s writer makes papers freely obtainable on-line as ‘reviewed preprints’, and publishes them alongside their referee stories and the unique unedited manuscripts.
The authors fed the unique variations into Claude and prompted it to generate referee stories. The group then in contrast the AI-generated stories with the real ones revealed by eLife.
The AI-written evaluations “seemed skilled, however had no particular, deep suggestions”, says Lingxuan Zhu, an oncologist on the Southern Medical College in Lianyungang, China, and a co-author of the research. “This made us notice that there was a significant issue.”
The research discovered that Claude might write believable quotation requests (suggesting papers that authors might add to their reference lists) and convincing rejection suggestions (made when reviewers assume a journal ought to reject a submitted paper). The latter functionality raises the chance of journals rejecting good papers, says Zhu. “An editor can’t be an professional in every thing. In the event that they obtain a really persuasive AI-written unfavourable assessment, it might simply affect their determination.”
The research additionally discovered that almost all of the AI stories fooled the detection instruments: ZeroGPT erroneously labeled 60% as written by a human, and GPTzero concluded this for greater than 80%.
Differing opinions
A rising problem for journals is the truth that LLMs could possibly be utilized in some ways to provide a referee report. What’s deemed an ‘acceptable’ use of AI additionally differs relying on whom you ask. In a survey of some 5,000 researchers performed by Nature earlier this 12 months, 66% of respondents mentioned it wasn’t acceptable to make use of generative AI to create reviewer stories from scratch. However 57% mentioned it was acceptable to make use of it to assist with peer assessment by getting it to reply questions on papers.
And though AI-detection instruments are bettering, they battle to find out how a lot of a doc has been generated utilizing AI. An evaluation revealed final 12 months of referee stories that had been submitted to 4 computer-science conferences estimated that 17% had been considerably modified by chatbots2. It’s not clear, nonetheless, whether or not the referees used AI to enhance the stories or to put in writing them completely.
Nature Index 2025 Analysis Leaders
Jeroen Verharen, a neuroscientist on the agency iota Biosciences in Alameda, California, says he’s shocked that the AI detectors utilized by Zhu and his group weren’t higher at recognizing the AI-written referee stories.
However he provides that AI-written stories and related supplies are unlikely to develop into a widespread downside. If reviewers don’t need to assessment, he says, “they’d simply say no”.
Conversely, Mikołaj Piniewski, a hydrologist on the Warsaw College of Life Sciences, argues that it’s a rising subject. He says he has already acquired referee stories that he suspects had been written by AI.
“LLMs are more and more being utilized by peer reviewers, though that is hardly ever disclosed,” he says. “After I spoke to my colleagues within the subject of hydrology, it grew to become clear that every of us had encountered not less than one such case as an writer up to now two years. A minimum of one of many assessment stories we acquired seemed very suspicious, and the AI-detection instruments we used flagged it as doubtlessly generated by LLMs.”
Piniewski provides that he’s positive some journal editors are accepting AI-generated referee stories, unwittingly or in any other case. He suggests {that a} world scarcity of peer reviewers could possibly be inflicting some editors to be extra lenient than they need to be. “I’m afraid it’s largely pushed by comfort,” he says.

