Dominique de Roo of De Gruyter Brill on Truth in the Age of AI. Listen to Episode 2 of Upstream by Integra.

Listen Now
Blog Aug 24, 2026 | Research Integrity

Beyond the Manuscript: From Document Review to Pattern Recognition 

4

sruthi.santhakumar Marketing Manager

Research integrity is entering a new phase. As publication fraud becomes increasingly coordinated, the challenge is no longer simply evaluating individual manuscripts. It is recognising patterns of behaviour that only emerge across submissions, journals, and publishing ecosystems. 

Fraudulent submissions rarely arrive alone. Increasingly, neither should research integrity investigations. As paper mills evolve into coordinated, multi-actor operations, screening a manuscript in isolation is no longer enough. Journals need to see the network behind it. But paper mills are only the motivating example here, not the real subject. The real question is no longer simply whether a manuscript is sound. It is what that manuscript reveals when viewed alongside everything else moving through the publishing ecosystem. That shift changes what journals need from both technology and people. 

AI does not solve the research integrity challenge by replacing judgment. It helps preserve judgment for the questions only people can answer. 

From Evaluating Documents to Recognising Behaviour 

COPE’s decision to relaunch its Systematic Manipulation of the Publication Process Working Group makes the shift explicit. The group’s own framing is telling: paper mills haven’t disappeared, they’ve evolved, growing new capabilities as coordinated misconduct finds new tools, AI chief among them. That reframing matters for anyone responsible for manuscript screening, because it points to a gap that document-level checks were never designed to close. 

Most editorial screening, however sophisticated, is built to answer one question at a time: is this manuscript sound? Are the references legitimate? Do the images show manipulation? Are the ethics disclosures complete? Is the language ready for peer review? 

Those are the right questions, and asking them well matters enormously. A large share of genuine quality problems are caught exactly this way. But organised fraud is designed to survive that kind of scrutiny. A paper mill doesn’t submit one weak manuscript and hope it slips through. It submits many, spread across journals, editors, and special issues, each one individually plausible, each one clean enough to pass a standalone check. The tell isn’t in any single manuscript. It’s in what happens when you look at several of them together. 

This is where “pattern” starts to mean something much bigger than fraud detection. Once you start looking across submissions rather than within one, all sorts of signals become visible that a single manuscript record could never show on its own: repeated reviewer recommendations that show up across unrelated papers, recurring author networks, citation clusters that reinforce each other’s standing, institutional relationships that don’t quite add up, guest editor behaviour that deviates from a journal’s usual submission profile, submission timing that only makes sense as coordination, image reuse across supposedly unconnected studies, and even language or AI-generation fingerprints that recur across manuscripts with no obvious connection to one another. 

None of these patterns are visible inside a single manuscript. That, more than any specific fraud typology, is this article’s central insight, and it’s why treating research integrity as a document-review problem was always going to reach a ceiling. 

Why the Ceiling Is Real, Not Theoretical 

This is harder than it sounds for most editorial offices to act on. Submission volumes keep climbing, and editorial staff are already stretched across technical checks, ethics disclosures, and language readiness before a manuscript has even reached scientific evaluation. 

Information is no longer the scarce resource in scholarly publishing. Editorial attention is. 

It isn’t that editors lack skill or diligence. It’s that human judgment, however sharp, can only be pointed at so many things in a given week, and fraud tends to organise itself around exactly the places that attention can’t reach. That scarcity shows up in exactly the places fraud likes to hide: in the manuscripts that get less scrutiny simply because there isn’t time to give every submission the same level of attention. Document-level screening was built to catch fabrication, not coordination, and an overstretched editorial office rarely has the bandwidth to do both manually. 

The Anatomy of a Coordinated Operation 

Coordinated manipulation typically combines multiple weak signals that reinforce one another. False or purchased authorship, fabricated reviewer identities, manipulated citation patterns, reused or altered images, compromised special issues used as low-scrutiny entry points, and unusual submission timing are each, on their own, individually concerning. Taken together, they reveal coordinated behaviour rather than isolated misconduct — something closer to infrastructure than to a single bad actor. 

From Document-Level Screening to Network-Level Investigation 

This is where the more useful distinction sits: not “screening versus no screening,” but document-level screening versus network-level investigation. Document-level screening asks, is this manuscript trustworthy? Network-level investigation asks a different question: does this manuscript behave like everything else we know about trustworthy research? Has this reviewer email appeared on other submissions under different names? Do these citation patterns match known citation-cartel behaviour? Is a guest editor’s special issue receiving submissions at a volume inconsistent with genuine research output? 

None of that is visible from inside a single manuscript record, and none of it is work an editorial office can reasonably do on top of an already full screening workload. It requires dedicated investigative capacity — connecting submissions to each other across time, across editors, and increasingly across publishers — which is exactly the kind of cross-industry coordination COPE’s working group and initiatives like the STM Integrity Hub are being built to support. 

The real question is no longer whether journals should automate. It is what automation should enable. At minimum, it needs to do four things: create consistency across every submission regardless of how busy a given week is, reduce the fatigue that comes from repetitive first-pass checks, surface signals that a human reviewer would otherwise have to notice unaided, and help prioritise which cases deserve deeper investigation. Get those four things right, and you preserve editorial judgment rather than replace it. Get them wrong, and automation just adds noise to an already strained system. 

Building the Next Generation of Research Integrity 

Research integrity increasingly depends on two complementary forms of capability. 

The first capability is consistency. Technical compliance, reporting standards, disclosure requirements, language quality, citation validation, image checks, and other document-level signals need to be evaluated for every submission, every time. This is the kind of work that automation is increasingly well suited to perform, ensuring consistency while reducing the routine workload that consumes editorial attention. 

The second capability is investigation. When patterns emerge that suggest coordinated manipulation rather than isolated problems, human expertise becomes indispensable. Understanding reviewer networks, analysing authorship relationships, interpreting citation behaviour, and investigating potential paper mill activity all require experience, context, and judgment that extend well beyond document screening. 

Technology and Expertise Are Complementary 

These two capabilities should not be viewed as alternatives. Automation creates consistency, surfaces signals, and helps direct limited editorial attention. Human investigators provide the discernment needed to determine whether those signals represent genuine misconduct or simply unusual but legitimate research practices. 

The Bigger Shift 

Paper mills didn’t get more sophisticated by accident. They adapted to a system built to catch one bad manuscript at a time, and they organised around its blind spots, including the very real constraints on editorial attention that make consistent, submission-by-submission vigilance hard to sustain. Closing those blind spots means editorial teams, publishers, and the wider integrity community treating fraud detection as a pattern-recognition problem, not a document-review problem, and building the capacity, automated and human, to work at both levels at once. 

Document-level screening remains the foundation. It’s no longer the whole answer, but it’s what makes the harder, network-level work possible at all. 

The manuscript is still the unit of editorial decision. Increasingly, however, it is no longer the unit of investigation. As organised misconduct becomes more coordinated, the future of research integrity will belong to publishers that can understand not only what a manuscript says, but what it reveals when viewed as part of a much larger pattern. 

From Insight to Practice 

The future of research integrity will not be defined by AI alone, nor by human expertise alone. It will be defined by how effectively the two are combined. 

Publishers are increasingly looking to combine these complementary capabilities within existing editorial workflows. At Integra, this approach is reflected in EditorialPilot, our AI-assisted manuscript assessment platform, and our peer review management services, where experienced research integrity specialists investigate cases requiring deeper scrutiny beyond automated screening. 


Recent Blogs

Where Educational Value Moves in the Age of AI
AI in Education

Where Educational Value Moves in the Age of AI

Beyond the Page: The Discipline of Enterprise Account Leadership
Beyond The Page

Beyond the Page: The Discipline of Enterprise Account Leadership

Beyond the Page: A Conversation with EditorialPilot
EditorialPilot

Beyond the Page: A Conversation with EditorialPilot

Want to
Know More?