"Partnership Announcement 2024"
Blog
博客

AI Did Not Break Research Integrity

It Turned On the Lights.

Caption

Research integrity has suddenly become one of the big conversations in scholarly publishing. Paper mills, manipulated images, fake reviewers, fabricated data, questionable citations, authors who may not actually be authors. And now, inevitably, AI.

It is tempting to put all of this together and conclude that generative AI has created a research integrity crisis.

There is only one problem with that argument.

The dates do not cooperate.

Most of these problems were around long before ChatGPT entered the vocabulary. AI has undoubtedly changed the situation, and not always for the better. But it may be more useful to think of AI as an accelerant rather than the source of the fire.

In fact, rather ironically, the same technology that can make research misconduct easier to scale is also helping publishers discover just how much of it may already have been there.

So perhaps we are asking the wrong question.

The cracks were already there

Scholarly publishing has always operated on a considerable amount of trust.

An editor receives a manuscript and assumes that the research actually happened, the data exists, the authors are who they say they are, the images are genuine and the references support the claims being made. The reviewer is also assumed to be a real person with the expertise and independence they claim to have.

For the vast majority of researchers, those assumptions are perfectly reasonable.

The difficulty is that scholarly publishing does not operate on trust alone. It also operates on incentives.

Publications help careers. Citation counts matter. Research output affects institutional rankings. Funding and promotion can depend on publication records. None of this is inherently wrong. But whenever a system attaches considerable value to something, somebody eventually works out how to manufacture it.

Paper mills are the rather depressing proof.

COPE and STM published research drawing on more than 53,000 submissions across six publishers. In the journals examined, suspected paper-mill submissions ranged from around 2% to as high as 46%.

The 46% figure should not be misunderstood. It certainly does not suggest that nearly half of scholarly publishing is fraudulent. It shows that particular journals can be targeted extraordinarily heavily.

More importantly for the present argument, the research predates the explosion of generative AI.

AI did not create the cracks. It turned on the lights.

What AI really changed was the economics

Producing a convincing fraudulent manuscript used to involve a fair amount of work. Somebody had to write the text, invent or manipulate the research, assemble the references, prepare the images and perhaps manufacture reviewer identities and reports.

There was, unfortunately, a certain amount of craftsmanship involved.

Generative AI changes the economics considerably.

Text can now be produced or rewritten in seconds. Images can be generated or altered. References can be produced, although with the occasional inconvenience that some of them may not exist. Reviewer reports can be created almost instantly.

None of this gives people a new reason to cheat. It simply makes some forms of cheating faster and cheaper.

Paper mills have acquired better machinery.

This distinction matters because it changes the problem we need to solve. If AI were the cause of research misconduct, detecting AI might be a reasonable response. If AI is primarily making existing misconduct easier to scale, detecting AI alone will not get us very far.

Publishers, fortunately, have technology too

There is another side to this story.

The STM Integrity Hub, has become a shared infrastructure through which publishers can identify potentially problematic submissions. STM has also reported screening more than 125,000 manuscripts a month and identifying roughly 1,000 suspected paper-mill submissions monthly.

Springer Nature has contributed technology originally developed to detect AI-generated nonsense text. Other systems look for duplicate submissions, suspicious patterns and a growing collection of integrity signals.

There is something wonderfully circular about this.

We are developing technology that can produce questionable content at extraordinary speed, while developing other technology to identify the questionable content produced by the first technology.

Anyone who has worked in cybersecurity will recognise the business model.

The danger is that research integrity becomes the same sort of arms race. Publishers improve detection. Bad actors learn what is being detected and adapt. Publishers improve again.

That may be unavoidable to some extent. But it would be unfortunate if that became our entire strategy.

Perhaps AI is the wrong suspect

There is a more uncomfortable part of the discussion that technology cannot solve quite so neatly.

Why does somebody buy authorship on a paper?

Why does a paper mill have customers?

Why would a researcher risk a career to manufacture research?

Eventually, we arrive at incentives.

In many academic systems, publication is not simply a way of communicating research. It is connected to promotion, funding, reputation, institutional targets and career progression.

AI did not invent publish or perish.

It simply gave publish or perish a faster computer.

Unless those underlying incentives are considered, there is a limit to what another detection system can achieve.

I am not sure “Was AI used?” is the right question

A great deal of attention is currently being given to determining whether AI was involved in producing a manuscript. I am not convinced this will remain a particularly useful question.

Researchers will use AI.

They will use it to improve English, translate material, analyse information, write code, organise notes and restructure prose. Some uses will be acceptable, some will require disclosure, and some clearly will not be acceptable.

But the mere presence of AI tells us surprisingly little about whether the underlying research can be trusted.

A paper written entirely by a human can contain fabricated data. A researcher who used an AI tool to improve the English in a perfectly sound manuscript has not suddenly invalidated the science.

The more useful questions are rather older-fashioned ones.

Who are the authors? Can their identities and affiliations be verified? Does the underlying data exist? Are the images authentic? Do the citations actually support the claims? Has the manuscript appeared elsewhere? Is the reviewer who they claim to be? Was AI used, and if so, where and for what?

The distinction may seem subtle, but I think it is important.

We stop trying simply to detect AI and start trying to verify research.

From declared trust to verifiable trust

This may be the more interesting change taking place.

For a very long time, scholarly publishing has operated largely on declared information.

The author tells us who they are.

The author provides an affiliation.

The author says the data exists.

The manuscript presents an image as genuine.

A citation is offered in support of a claim.

A reviewer declares that there is no conflict of interest.

The researcher discloses how AI was used.

Again, most of those declarations will be entirely legitimate. The answer is not to treat every researcher as a potential fraudster. Apart from being unfair, that would make scholarly publishing practically impossible.

But perhaps we can do better than declaration alone.


 

Identity can increasingly be checked against external evidence. Data can have provenance. Images can be examined for manipulation or duplication. Citations can be checked not merely for existence but for whether they support the statement being made. Reviewer identities, expertise and conflicts can be assessed. AI use can be disclosed in a structured way and considered in context.

No single one of these checks proves that a paper is trustworthy.

That is precisely the point.

Research integrity probably does not need one enormous machine that produces a green tick or a red cross. It needs many smaller signals, each answering a specific question and showing the editor the evidence behind the answer.

Taken together, those signals begin to create something rather more useful.

A chain of trust.

Or, perhaps more accurately, a move from declared trust to verifiable trust.

Verification has its own risks

There is an obvious danger in all this.

Once we have the ability to score things, we have a remarkable tendency to believe the score.

That would be a mistake.

Researchers whose first language is not English may legitimately use AI extensively in preparing their prose. Academic writing is often formulaic. Interdisciplinary research can look unusual precisely because it crosses established boundaries. New areas of research may have patterns that differ substantially from the historical data used to build a detection system.

False positives are not an inconvenience when reputations and careers are involved.

A signal is not evidence of misconduct, and a score is certainly not a verdict.

The useful role for technology is to tell an editor, “You might want to look at this.”

What happens next still requires judgement.

This is already beginning

The industry is moving towards more structured approaches to AI disclosure. STM, COPE, the International Science Council and the Global Young Academy have been working on a common framework for reporting the use of AI in research, including work towards what has been described as a Vancouver Standard for AI reporting.

That may sound like another publishing standard, and scholarly publishing is not exactly suffering from a shortage of those.

But the underlying idea is important.

Disclosure is really about provenance.

What was done? By whom? Using what? Where did technology play a role?

Once you start asking those questions, the conversation becomes much larger than AI.

It becomes a conversation about how trust itself is established in scholarly communication.

Perhaps AI has done us a favour

That is an odd thing to say about a technology frequently blamed for making the research integrity problem worse.

But it may nevertheless be true.

AI has made it easier to produce questionable content at scale. That deserves serious attention.

At the same time, it has exposed assumptions in scholarly publishing that were becoming increasingly difficult to sustain at today’s submission volumes. And it is providing some of the tools that could help us replace those assumptions with evidence.

The objective should not be to create a publishing system in which every author is treated as a suspect. It should be the opposite: allow legitimate research to move through the system efficiently, while giving editors much better evidence about the relatively small number of submissions that genuinely deserve closer attention.

That seems considerably more useful than trying to determine whether a machine thinks another machine wrote a paragraph.

AI did not invent paper mills. It did not invent publication pressure, manipulated research, questionable citations or weaknesses in peer review.

It simply made some of those weaknesses much harder to ignore.

So I would not ask:

How do we stop AI from damaging research integrity?

I would ask:

Now that AI has shown us where the system is vulnerable, what are we going to do about it?

That is a much more interesting conversation.

添加评论
Gain better visibility and control over your entire processes.
Retain control over your content; archive and retrieve at will.
Achieve a 20% cost-saving with our AI-based publishing solution.