AI Stress Points in Scholarly Publishing – Part II
Researchers who use AI tools, and those who investigate them, take the conversation beyond the editorial office.

Image: Luciana Machado, using Canva Pro elements
SHARE THIS ARTICLE
Welcome back to the first SIA Special Report. In Part I, we focused on how the growing use of AI across publishing workflows is creating new strains on editorial processes, attribution, and the scholarly record.
What Part II Covers
We draw on insights from interviews with researchers who use AI tools and investigate their use in research. We conducted original interviews with Konradin Metze, Pathology Professor at the State University of Campinas, who researches computational tools; Mike Thelwall, Professor of Data Science at the University of Sheffield and lead of the ESRC-funded project AI Peer; and Kent Anderson, founder of Caldera Publishing Solutions and of the blog The Scholarly Kitchen, and co-author of the forthcoming How the Internet Disrupted Science.
We have also stayed attentive to what the research community is saying, so you will find our interpretations of some of the discussions unfolding at recent conferences and community events — from the 9th World Conference on Research Integrity (pages 6–45) to the European Association of Science Editors (EASE) Summer Symposium and the PurePub conference. As always, we don't aim to cover everything, since these themes are evolving quickly and new AI tools and version updates keep being released. Full interviews are available in the HIKE forum, where SIA subscribers can join the larger discussion about trust, research assessment, and community.
Retractions, Metadata, and Research Assessment
In 2023, as the ChatGPT hype was beginning, Metze and a student tested whether it could help them find relevant literature. While the ChatGPT delivered references whose titles looked very helpful, none of them existed. "This made me very angry," Metze told us, "but I also became interested in learning how this could happen." While some researchers use specialized AI tools to automate workflows, accelerate data analysis, and synthesize vast amounts of literature, others seem to have expectations for generalist LLMs that do not align with their function. A much-publicized analysis found that fabricated references in the biomedical literature have risen sharply, roughly 12-fold over two years, reaching about one in 277 papers in early 2026. When used in research designs that aim to produce valid, reliable, and reproducible results, much of the variability seen in AI outputs may be caused by rapid model updates or system-level changes, rather than purely algorithmic inconsistency.
“Every chatbot output must be double checked if we want to continue with the standards of good scientific practice.”
Konradin Metze, Professor at UNICAMP
This is why Metze urges educators to teach critical thinking when discussing generative AI, and he is not alone in advocating for a focus on AI literacy in higher education. In a landmark study in 2025, Metze found that LLMs correctly identified fewer than half of retracted papers from a curated reference list. The language used in retraction notices is not standardized, and for now, LLMs do not reliably capture unstructured metadata such as retraction flags. Does the problem reside primarily in model design? Will integration with trusted databases (e.g., Crossref, PubMed retraction notices, Retraction Watch Database) solve this issue? Or do we still need overburdened editors and reviewers to conduct reference validation and DOI checking?
Whether the inability of LLMs to weed out retracted studies, as Metze documented, arises from name disambiguation problems, hallucination, or both, it could be important for policy discussions around AI use in administrative contexts. In an earlier study, Metze also found that the less published on a subject, the more errors ChatGPT produced. This effect could reinforce existing visibility disparities between major and niche research fields, similar to citation concentration effects already observed in the field of bibliometrics.
Thelwall's research points the same way, as he found that ChatGPT failed to flag any of 217 high-profile retracted or otherwise concerning academic papers, and since even the most publicized retractions went unrecognized, he told us, less visible retractions "will almost certainly be completely ignored by LLMs unless they are specifically asked to check if the article has been retracted."
As lead of the AI Peer project, Thelwall is testing whether LLMs can reliably review academic work and contribute to the UK's Research Excellence Framework and journal peer review. The use of scientometric indicators (for the evaluation and mapping of scientific fields, exploring research themes and collaboration clusters, and identifying gaps and future trends) may become the link to a new way of working with LLMs when iterative correction is key. A recent example of this was the development of a tool that helps researchers report details of a toxicological study according to the legal standards needed for the study to be used in regulatory decision making.
“I think that we are overdue a serious rethink about the purpose of publishing, academics, and disciplines in the era of LLMs and potentially cheap publishing.”
Mike Thelwall, Professor of Data Science at the University of Sheffield
Community, Integrity, and Trust
Recently, Ben Kaube, an entrepreneur who invests in scholarly-communications startups, ran a word-frequency analysis of posts on The Scholarly Kitchen blog and found that "community" and "integrity" have been rising steadily over the last four years. Writing in the same blog about the first PurePub conference, Laura Harvey and Adam Hyde (founder of Pure Science, the publisher-focused automation platform behind the event) reported that 75% of polled attendees said their organizations are still at an experimental stage on the AI adoption curve. This is perhaps why the concepts "AI-assisted research integrity" and "trust and governance" were closely clustered in a knowledge graph of the sessions.
“Journals are collections of communities that come together with a common purpose.”
Ian Mulvany, Chief Technology Officer at BMJ
Publishers seem to have finally realized that the PDF is outdated and not fit for purpose in the new AI-mediated ecosystem of structured content, metadata, permissions, provenance, APIs, and connectors. At the same conference, Ian Mulvany, Chief Technology Officer at BMJ, reminded us that journals are collections of communities that come together with a common purpose, which is to try to understand a field or discipline. Whether mainstream publishing still honors that purpose is another question. Speaking directly with us, Anderson argued the industry has lost the thread of the conversation with the communities it serves in its pursuit of scale.
While some guidance exists regarding AI use disclosure at the journal, institutional, or editorial level (see Part I of this report), the shape these disclosures should take at the author level has not been agreed upon.
A participatory, community-driven initiative aimed at the development of an internationally recognized Reporting Standard for AI Disclosure in Research was recently discussed at the 9th World Conference on Research Integrity (read more about the Vancouver Standard on page 35). Participants examined features that should be expected of such a reporting standard — one that would enable disclosures, articles, and other scholarly outputs to be compared with one another. These were prioritized and weighed against the goals for the standard, namely the ability to align and communicate expectations across the research ecosystem and destigmatize the use of AI in research. The main features were:
ease of understanding
structured format
machine readability
flexibility
ability to disclose multiple tools/uses in a work
usability
Benchmarking, Multi-Agent Workflows, and Vibe Coding
Researchers and developers of AI tools have established methodologies that build agentic workflows from high-performing generative AI tools and benchmark them against human-produced work. These workflows and their use in evaluating research were heavily discussed at the latest European Association of Science Editors (EASE) Summer Satellite Event. Mark Hahnel, founder of Figshare and VP of Open Research at Digital Science, presented preprints.ai, a tool in which different agents work at different steps of the process, or cooperate, to produce a grading system conceived as a trust marker for preprints. But is this profusion of solutions drowning out what researchers need to get the job done as efficiently as possible?
At the same event, Natalie Khalil, CEO and co-founder of the multi-agent research paper review platform Reviewer3, spoke about ReviewBench, an attempt to benchmark human and AI peer-review reports, assessing comment structure, critique typing, and claim mapping. More extreme still, Copenhagen Business School Associate Professor Christian Hendriksen argues in a recent op-ed that it might soon constitute questionable research practice not to let an LLM review a paper or assist with peer review.
César Hidalgo went a step further and started in early 2026 the Journal for AI-Generated Papers, an open-access community journal designed specifically for AI-generated and co-authored academic work, with reviews by Reviewer3. Previously, the preprint server aiXiv welcomed both human-authored and AI-generated research papers, which were submitted, critiqued by AI reviewers, and revised multiple times to build up methodological rigor. But what can these experiments teach us? For Anderson, the prior question is whether a paper can be retracted at all, or corrected, or even cited, once it has been "atomized, tokenized, and weighted" inside an LLM.
In his investigation into Rachel So, an AI agent presented with a human name and a title, Anderson uncovered that a PeerJ Computer Science editor had invited the AI agent to peer review an article. The creators of Project Rachel defended their decision not to disclose the deception, citing the experiment's observational integrity; Anderson classified the project as research fraud.
LLMs may offer researchers a viable way to build custom software and even write papers through vibe coding, but applications built using AI-generated code face debugging, maintenance, and update challenges, and, worryingly, LLMs can quietly shift our attitudes on important social issues. As Anderson put it to us, generating lines of code is not the same as developing software with a purpose.
“Retraction, correction, and citation are about accountability, responsibility, transparency, and the ability to govern systems with laws and analyses.”
Kent Anderson, Caldera Publishing Solutions
Whether in failures such as zombie citations and AI agents posing as researchers, or in genuine gains in access and efficiency, our reporting pointed to the same conclusion: none of these tools eliminates the need for human judgment, and it is still unclear if any trust marker will ever be able to stand in for the trust that communities build.
As society at large and researchers at the edge of innovation grapple with these questions, where can knowledge be verified?

Subscribe to read our full interviews with Konradin Metze, Mike Thelwall and Kent Anderson in SIA's expert-led forum.
Rights and Permissions
© 2026 Science Integrity Alliance.
This article is open access, published under a Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. You may share and adapt it for noncommercial purposes, with credit to the author and to REACH. Images, illustrations, logos and other third-party material are not covered by this license unless the caption says otherwise. Permission for these must be sought from the rights holder.
Cite As
Maria Machado, Maryam Sayab, Paul Whaley, Luciana Machado. AI Stress Points in Scholarly Publishing – Part II. REACH 2026;4(April-June):80-83.
REACH is the quarterly digital magazine of the Science Integrity Alliance, a coalition of more than 25 partners working to strengthen research integrity. Editor's Choice articles are open to everyone, and subscribers make that possible. In return, the SIA Subscription brings you together with people committed to improving research culture and transparency: full issues of REACH, discussion and practical advice in HIKE (our expert-led forum), bonus episodes of the SIAcast podcast, and the chance to participate in creating ATLAS, our living encyclopedia of key research concepts.


