"Partnership Announcement 2024"
Blog
Blogs

Dear Reference, Why Am I Still Editing You?

Caption

We recently looked at 100 articles containing close to 1000 author queries.

Approximately 80% of those queries were related to references.

Missing page numbers. Incorrect years. Journal abbreviations. Volume and issue numbers. Author names. DOIs. Publisher details. Reference formatting.

That number made me stop.

Not because references are unimportant. Quite the opposite.

But because much of the information we are asking authors to provide already exists somewhere else.

So perhaps the question is not: How can we make reference editing more efficient?

Perhaps it is: Why are we doing so much of it in the first place?

I decided to ask a reference.

Me: We need to talk.

Reference: Again?

Me: Your journal title is wrong.

Reference: Is it?

Me: The author supplied the full title. This journal requires the abbreviated title.

Reference: Anything else?

Me: Your page range needs checking. The DOI is missing. Your issue number may need to be removed. And one author’s initials do not match.

Reference: Can I ask you something?

Me: Go ahead.

Reference: I was published in 2012. I have a DOI. I am indexed. Crossref knows who I am. PubMed knows who I am. Why are you asking the author?

That is a fair question.

Are we asking authors questions that databases can answer?

The traditional publishing workflow makes perfect historical sense.

Authors supplied references as text. Publishers checked them, corrected them and converted them into journal style.

The manuscript was the source.

But the world around that workflow has changed.

For a journal article with a DOI, Crossref may already hold structured metadata including the title, authors, journal, publication information and DOI. In biomedical publishing, PubMed provides another important source of bibliographic information.

Yet we can still find ourselves asking:

Please provide the missing volume number for Reference 26.

There is something slightly strange about asking an author in 2026 to investigate the bibliographic details of a paper published in 2004 when machines can retrieve those details in milliseconds.

We are asking the author to reconstruct history.

Then we check their reconstruction against a database containing the history.

There may be a shorter route.

Perhaps we have confused the reference with its appearance

Style Guide: Excuse me.

Me: I thought you might turn up.

Style Guide: References must be consistent.

Absolutely.

Consistency matters because it improves readability. A reader should be able to recognise authors, titles, dates and publication information without deciphering a new structure every time.

But there is an important distinction.

The bibliographic data and the way we display that data are not the same thing.

If we know that something is an author name, article title, journal, year, volume, issue, page range and DOI, software can render that information according to almost any journal style.

So why are we manually correcting the presentation when we could first establish the underlying data and then generate the presentation?

And does the reader really care?

This is where the question becomes slightly uncomfortable.

Does the researcher reading an article care whether the journal title is abbreviated?

Does the reader care whether an issue number appears?

Does a missing full stop materially affect their understanding of the science?

Probably not very much.

What they almost certainly care about is:

Is this the right paper?

Can I find it?

Does it actually support the statement I have just read?

The first two can increasingly be helped by identifiers and trusted metadata.

The third is where things become much more interesting.

Because it is perfectly possible to have a beautifully formatted, completely accurate reference that has absolutely no business being attached to the sentence beside it.

We may spend five minutes fixing its punctuation without spending time asking whether it supports the claim.

That seems like an odd use of increasingly scarce editorial attention.

The author still has an important job

This does not mean blindly replacing every author-supplied reference with whatever Crossref or PubMed returns.

Metadata contains errors. Not everything has a DOI or PMID. Books, chapters, datasets, standards, websites, conference proceedings and older publications can be complicated. Different sources may contain different metadata.

And sometimes the author has simply cited the wrong paper.

The author therefore remains responsible for something much more important than punctuation:

Did I cite the work I intended to cite?

Perhaps that should become the principle.

The author supplies the intellectual connection.

The publishing system resolves the bibliographic identity.

Trusted sources supply or verify the metadata.

The journal style determines how it is displayed.

And humans become involved when those things do not agree.

That changes the workflow from check everything to check the exceptions.

What would happen to author queries?

Imagine that Reference 42 arrives without a page range.

Today we might query the author.

Tomorrow the system finds the DOI, retrieves the metadata and fills the gap.

No query.

Another reference contains the wrong year.

The DOI resolves confidently to a publication with a different year.

Instead of asking the author to type the year, perhaps we ask the only question that matters:

“We believe this is the article you intended to cite. Please confirm.”

Another reference cannot be identified at all.

Now we query the author.

That is a useful query because we are asking the author something the system genuinely cannot determine.

This is not about removing humans.

It is about stopping humans from doing work that machines can do perfectly well.

What could we do with the time?

This, for me, is the more important question.

If 80% of our author queries concern references, what happens if we can remove even half of them?

Authors spend less time answering queries.

Copyeditors spend less time raising and resolving them.

Production cycles become shorter.

But more importantly, editorial attention can move somewhere more valuable.

Instead of:

Please provide the issue number.

We can ask:

Does this reference support the claim?

Has this article been retracted or corrected?

Is the author citing the original research?

Is an important piece of evidence missing?

Is the citation being represented accurately?

Those are questions about scholarship rather than typography.

And they are much harder.

Maybe the reference needs a new job description

Reference: So you do not want me anymore?

Me: I did not say that.

Reference: It sounded like it.

Me: I want you to do something more useful.

Reference: Such as?

Me: Stop being a carefully formatted string of text and start being a verifiable connection between a claim and its evidence.

Reference: That sounds like considerably more work.

Me: For you, yes. Hopefully less for us.

For years, publishing has put enormous effort into making references consistent.

There was a good reason for that. Consistency improves readability, and historically the reference itself was our primary bibliographic record.

But the infrastructure has changed.

We now have DOIs, PMIDs, Crossref, PubMed and other structured sources capable of identifying large parts of the scholarly record.

Perhaps our workflows need to change with them.

The author should tell us what they intended to cite.

Trusted scholarly infrastructure can help establish what that publication actually is.

The journal can decide how it should look.

And our editors can concentrate on the question that really matters:

Does it belong there?

Reference: So no more queries about my commas?

Me: Fewer.

Reference: I will take that.

And perhaps authors will too.

Gain better visibility and control over your entire processes.
Retain control over your content; archive and retrieve at will.
Achieve a 20% cost-saving with our AI-based publishing solution.