Tag Archives: LLMs

Keeping up with pollination: using ChatGPT as a research alert

There is too much science.

That is not a complaint about the quantity of research being done. Quite the opposite: it is remarkable how much good work is now being published across the world. But the sheer volume creates a serious practical problem. How can any researcher keep track of what is relevant to their field, let alone read and properly absorb it?

For someone interested in pollination, the difficulty is compounded by the extraordinary breadth of the subject. Pollination is not confined to a single academic discipline, or even to a small cluster of them. It cuts across botany, zoology, ecology, evolution, conservation biology, agriculture, economics, food security, national security, palaeontology, biogeography, genetics, behaviour, climate science and many other areas.

A paper that changes how I think about pollination might appear in a specialist plant journal, an entomological publication, or journals covering conservation, agricultural economics, or palaeontology. It might concern the structure of ecological networks, the nutritional quality of crops, the evolution of flowers, pesticide regulation, the movements of migratory birds or the fossil record of insects. It may not even use the word “pollination” prominently in its title or abstract.

That breadth is one of the great attractions of a subject that has kept me fascinated for 35 years, and which I tried to capture in my book Pollinators & Pollination: Nature and Society. It is also what makes keeping up with it so challenging.

The limitations of conventional alerts

There are, of course, many ways to receive information about new research. Journals send tables of contents. Google Scholar provides alerts based on keywords or authors. ResearchGate regularly tells me that someone has cited one of my publications.

These services have their place, but they tend to provide a rather narrow window onto the literature.

For example, Google Scholar can alert researchers to newly indexed material matching a search query. This is useful, but it is neither a comprehensive record of everything published nor a carefully curated selection: broad searches produce noise, while narrow ones inevitably miss relevant work.

A citation alert from ResearchGate tells me about research connected to work that I have already published. That can be valuable, and occasionally flattering, but it inevitably looks backwards. It shows me the expanding wake of my own research rather than offering a broad view of where the subject is moving.

Keyword alerts have a different problem. They can generate large quantities of material with very little discrimination. A search for “pollination”, “pollinator” or “plant–pollinator interactions” will retrieve many relevant papers, but also conference notices, marginally related studies, duplicate records and work of highly variable importance.

More restrictive searches reduce the noise but risk excluding the unexpected paper that turns out to be especially interesting.

What I need is not simply a larger stream of titles. I want something closer to an informed research assistant: a system that could search widely, exercise some judgement, explain why particular items might matter, and alter its approach in response to my comments.

I have therefore been experimenting with ChatGPT’s scheduling function as a weekly research-alert system.

A scheduled conversation

My current instruction is for ChatGPT to provide a shortlist every Friday morning of the most worthwhile new papers, preprints, reports and substantive analyses relating to plant–pollinator interactions, pollination networks, bird pollination and related broader biodiversity topics.

I have asked it to include no more than ten items, and fewer when the available material is weak. That final qualification is important: I do not need ten references merely to fill ten spaces, I’d rather receive four genuinely interesting papers than a padded list containing six that I will never read.

For each item, the system can provide the citation, a short account of the main finding and an explanation of why it may be relevant to my interests. It can also distinguish between peer-reviewed research, preprints, reports and other forms of analysis.

In that sense, the alert is already more useful than a conventional automated search. But the most important difference is that it is a conversation.

I can tell it that a particular paper was especially useful and ask it to look for more work of that kind. I can point out that another item was only marginally relevant. I can ask it to widen its search into palaeontology, ecological economics or agricultural policy, or to pay closer attention to a particular taxonomic group.

I can also tell it what not to do. When a major European pollinator-research white paper appeared, for example, the alert quite reasonably identified it as relevant. But I was one of its co-authors and did not need an artificial intelligence system to introduce it to me as a new discovery. I could therefore instruct the system to recognise my own publications and either omit them or flag them only when they were strategically relevant.

That adaptability is difficult to reproduce with conventional keyword alerts. The scheduled search is therefore not a fixed filter, it can be refined as my interests, projects and frustrations change.

From retrieval to assessment

The distinction between finding research and assessing it is also important.

A long list of newly published papers transfers the problem of selection from the search engine to the researcher. A useful alert should do more than retrieve documents. It should offer some preliminary judgement about novelty, relevance and significance.

Does a paper introduce a genuinely new idea, or does it repackage a familiar concept in new terminology? Is a striking conclusion supported by a strong study design? Does a paper matter because of its empirical results, its methods, its conceptual framework or its policy implications? Is it directly relevant to my work, or merely adjacent to it?

ChatGPT cannot answer such questions infallibly. Nor should its assessment be accepted without scrutiny. But it can help to triage the literature and identify which papers deserve closer attention.

This is particularly valuable outside one’s immediate specialism. I can usually make a rapid initial judgement about a field study of flower visitors or a paper on pollination networks. I may need more assistance in deciding whether a new economic analysis, remote-sensing method or palaeontological reconstruction is likely to be important.

The purpose is not to delegate scientific judgement, it is to direct that judgement more efficiently.

A few necessary cautions

There are obvious limitations.

An AI-generated research brief is only as good as the literature it can locate and the instructions it has been given. It may miss relevant work, misunderstand a paper, exaggerate novelty or place too much weight on a fashionable topic.

There is also a danger of creating an intellectual echo chamber. If I repeatedly tell the system what interests me, it may become increasingly good at confirming those interests while overlooking research that sits outside them. A useful alert needs some permission to surprise.

For that reason, I think the instructions should combine a clear core remit with an explicit request to include occasional unexpected items that could change how the field is viewed.

The summaries themselves must also remain starting points. Important papers need to be read, methods inspected and conclusions considered in relation to the evidence. An articulate paragraph generated by ChatGPT is not a substitute for engaging with the original study.

A better kind of alert

Used carefully, however, scheduled ChatGPT alerts offer something that conventional notification services do not: an iterative relationship between the researcher and the search process.

The system can learn that I am interested not only in pollination as an ecological interaction, but also in its agricultural, economic, evolutionary, historical and political dimensions. It can respond when I say that a particular area is becoming more important to a project. It can explain its selections, accept criticism and search differently the following week.

That makes it less like an electronic table of contents and more like a research assistant who arrives on Friday morning with a small pile of papers and a cup of coffee and says: “These are the ones I think you should look at, and this is why.”

Like any research assistant, it will sometimes be wrong. It may overlook something obvious, become overenthusiastic about a weak paper or misunderstand why a study matters. But it can be corrected. The conversation continues, and the next briefing can be better.

For a field as broad and rapidly expanding as pollination research, that may be the most useful feature of all.

I will finish with an example. Today’s alert arrived at the scheduled 08:00 (though I had to make my own coffee…) and included four items for me to look at; this was the fourth, which I’m reproducing verbatim:

4. Agarwala et al. — Biodiversity loss will decrease the future creditworthiness of nations

The authors incorporate projected losses of wild pollination, tropical timber and marine fisheries into sovereign credit-risk models for 23 countries. Their scenarios suggest that partial ecosystem collapse could substantially increase national borrowing costs and expose a major omission in conventional credit ratings.

Why it matters: This is one of the clearer attempts to translate pollination loss from ecosystem-service valuation into a decision framework already used by governments and financial markets. For ecological-economics work, the methodological bridge—from spatial ecosystem-service scenarios through macroeconomic models to sovereign credit ratings—is more important than the headline monetary totals.

The ecological assumptions deserve scrutiny. “Wild pollination” is compressed into a tractable economic input, so uncertainty in crop dependence, substitution, adaptation and spatial mismatch may be obscured once propagated through the financial model.

Assessment: Strategically consequential interdisciplinary analysis, not new pollination biology.

The paper actually appeared on 4th June but I missed it, even though it cites this paper by Millard et al., in which I was involved, and presumably appeared on a ResearchGate citation alert last month. ChatGPT rightly spotted that it was something which would interest me and that’s what’s so fascinating about this approach to keeping up with the literature – this system is intelligent enough to “understand” my needs, even if it’s not conscious of precisely why it understands them.

As always, I’d be pleased to hear your views on this topic – feel free to comment below.

Could LLMs like ChatGPT ever replace part of the academic peer-review process?

Recently, I made a comment over on Steve Heard’s Scientist Sees Squirrel blog:

I have never published a paper that’s not been improved, to some degree, by peer review, and broadly the system works. But I do wonder if it’s sustainable in the long-term and whether in the future LLMs might actually be a more effective way of assessing manuscripts. I recognise that’s (currently) a controversial statement to make – but having recently run a few of my own manuscripts through ChatGPT and asked for its “opinion”, I can honestly say that the feedback has improved not just the writing but also the framing and focus of the work. It’s also picked up weaknesses and errors that I had otherwise missed.

That initiated an email conversation with Steve which resulted in me running a short experiment with ChatGPT model 5.5. I first loaded up the original manuscript that I’d submitted to a journal of this paper on pollinator effectiveness. I then asked ChatGPT to write a review of the manuscript as though it was a peer reviewer of the journal. Which it did – in some detail – in 28 seconds! If anyone is interested I can send them that review, but it’s the next bit that I think is especially interesting.

After ChatGPT had completed the review, I then uploaded the actual peer reviews I’d received from the journal, plus the editor’s comments, and asked it to summarise the degree to which its review agreed with those I had received.

Below I have copied the response, but it’s the take-home message, summarised in the graphic above, which I think is the most intriguing aspect: there was a 55% agreement with Reviewer 1, a 70% agreement with Reviewer 2, and a 65% overall agreement between ChatGPT’s assessment of the manuscript and that of the Editor.

That may not sound like a large overlap, but note that it’s the level of agreement between the two reviewers was only 65%! In my not inconsiderable experience (>130 peer reviewed papers and chapters, hundreds of peer reviews of my own) this is fairly typical – reviewers ALWAYS pick up on different strengths and weaknesses in a manuscript!

What particularly impressed me about ChatGPT’s response below, is that it went far beyond simply looking at the language being used, it provided a critique of how the study was conducted, how data were analysed, and the conclusions drawn. Which is just what we expect a human reviewer to do! Not only that, but it was more critical of the work than the reviewers and the editor.

So where does this leave us? Could Large Language Models (LLMs) like ChatGPT replace part of the academic peer-review process? Clearly there need to be guide rails and human agency involved in final decisions, but I think that, yes, they may well have a role to play in the not-too-distant future. As always, I’d be interested in your comments.

If you want to dive into this more deeply, here’s the verbatim account of what ChatGPT wrote (including its emboldened parts) when I asked how much agreement there was between its assessment and those of the other two reviewers and the editor:

There is a fair amount of overlap, but not a perfect match.

I would say my review was moderately similar in substance, but noticeably more critical in tone and in editorial recommendation. The strongest areas of convergence were on the paper’s core strengths and on one key statistical concern. The biggest mismatch is that the actual editor and reviewers were much more positive overall, and most of their requested changes are closer to clarification, presentation, and framing than to a deep challenge to the paper’s publishability.

The clearest points of agreement are these. First, I highlighted the manuscript’s strengths as the multi-year design, the focus on different functional groups, and the interesting role of Lepidoptera; that aligns very closely with the guest editor, who explicitly says that “the roles of lepidopterans, the inclusion of flight distance measures, and the multi-year study are particularly notable,” and with Reviewer 2, who lists the multi-year field study, multiple pollinator groups, and the unexpected importance of Lepidoptera as strengths. Second, I flagged the Poisson treatment of an ordinal flight-distance variable as statistically weak; Reviewer 1 independently makes essentially the same point and explicitly recommends ordinal regression instead. Third, I thought the discussion was somewhat over-extended and not always tightly aligned with the evidence; Reviewer 2 likewise suggests embedding the work better in the broader literature and moving some material out of the conclusion into the discussion.

There is also a partial overlap around the manuscript’s central interpretation. I argued that the paper sometimes pushes too hard from pollen deposition and coarse movement classes to larger claims about true generalisation, outcrossing value, and broader evolutionary or environmental buffering. The editor does not object to those broad claims in the same way; instead, the editor actually asks for the discussion to connect the results more explicitly to the theoretical statement that plants should generalise on equally effective pollinators under dynamic environments, and even invites discussion of anthropogenic disturbance and future interactions. In other words, where I wanted that interpretive arc toned down, the editor mostly wanted it better developed and better supported.

The main differences are quite important.

First, I was more sceptical about inference than the real reviewers were. I was worried that the manuscript moved too quickly from non-significant or coarse flight-distance patterns to statements about butterfly superiority, reduced geitonogamy, and broader adaptive significance. Reviewer 1 only objected narrowly to the modelling choice; neither reviewer makes my broader inferential critique, and the editor seems broadly comfortable with the interpretive direction so long as it is supported with more descriptive data and clearer exposition.

Second, I was more concerned about the coarse pooling of visitor taxa. I thought the manuscript should be more cautious because “butterflies,” “hoverflies,” “bumblebees,” and especially “Others” are broad bins. The actual comments do touch this, but in a softer, more taxonomic way: the editor asks whether there is a species or morphotype list, and Reviewer 2 asks for more detail on the butterfly taxa and justification for putting solitary bees into “Others.” That is related to my concern, but it is narrower and more presentational than the conceptual objection I raised.

Third, I did not anticipate how much of the real feedback would focus on descriptive content and figures. The editor repeatedly asks for more floral biology data, clearer results presentation, more absolute numbers, more behavioural quantification, clearer site description, and better figure design and consistency. Reviewer 2 similarly asks for effect sizes, better figures, units, and clearer display of the underlying data distribution. I mentioned presentational issues only in a general way; the actual reports are much more detailed and concrete here.

Fourth, I did not raise the absence of random effects, whereas Reviewer 2 does. That is probably the single most important substantive point that I missed. Reviewer 2 specifically asks why insect individual and/or observation day were not modelled as random effects.

Fifth, my editorial recommendation was harsher. I recommended major revision. The actual reports read much more like minor-to-moderate revision. The editor calls it “an excellent piece of work” and says it is well executed with robust methodology; Reviewer 2 says they have “little to criticise,” mostly editorial suggestions plus a few statistical and figure issues; Reviewer 1 has only a short list of points. That collective tone is clearly more encouraging than mine.

So, in condensed form:

  • Where I matched well: multi-year dataset as a major strength; Lepidoptera result as a notable contribution; the Poisson-for-ordinal issue; some concern that discussion/framing needed tightening.
  • Where I was too severe: I treated the interpretive overreach as a major obstacle, whereas the real reviewers mostly saw a strong paper needing clarification and enrichment rather than retrenchment.
  • What I underplayed or missed: requests for more descriptive natural history and floral biology data, much more detailed figure/presentation criticism, and the random-effects issue.

My bottom-line judgement is that my review was directionally right on some key scientific points, especially the statistical issue and the general strengths, but too pessimistic about the paper’s editorial position. If I had seen these reports first, I would have shifted from “major revision because the evidence-conclusion fit is shaky” to something more like “moderate revision: good paper, but it needs clearer descriptive support, better figure presentation, and tighter statistical justification. If revised along those lines, I think the manuscript could become a credible and interesting contribution”.

Here is a concise version of the editorial decision in journal-review language:

This manuscript addresses an interesting and worthwhile question using a useful multi-year dataset and a stronger-than-usual measure of pollinator performance. However, the current version over-interprets some of the results, particularly those concerning movement behaviour, outcrossing implications, and the broader adaptive significance of generalised pollination. The statistical treatment of flight distance is also not fully convincing. I therefore recommend major revision. The manuscript has clear potential, but its conclusions need to be more tightly aligned with what the data actually demonstrate.