Skip to main content
  • Home
  • About
  • Who we are
  • News
  • Events
  • Publications
  • Search
  • Independent Science for Development CouncilISDC
    • Home
    • Who we are
    • News
    • Events
    • Publications
    • Featured Projects
      • Inclusive Innovation
        • Agricultural Systems Special Issue
      • Proposal Reviews
        • 2025-30 Portfolio
        • Reform Advice
      • Foresight & Trade-Offs
        • Megatrends
      • QoR4D
      • Comparative Advantage
  • Standing Panel on Impact AssessmentSPIA
    • Who We Are
      • SPIA Team
      • Our Mandate
      • Impact Assessment Focal Points
      • SPIA Affiliates Network
    • What We Do
      • Country Studies
        • Community of Practice
      • Causal Impact Assessment
      • Use of Evidence
    • Where We Work
      • Bangladesh Study
      • Ethiopia Study
      • Uganda Study
      • Vietnam Study
    • Resources
      • Publications
      • News & Blogs
        • Blog Series on Qualitative Methods for Impact Assessment
      • Data Resources
        • Stocktake Data
        • Replication Files
        • SPIA-emLab Agricultural Interventions Database
      • Webinars
      • Events
  • Evaluation
    • Who we are
    • News
    • Events
    • Publications
    • Evaluations
      • CGIAR Center External Reviews (CER)
      • Learning on CGIAR's Ways of Working
      • Science Group Evaluations
      • Platform Evaluations
        • CGIAR Genebank Platform Evaluation
        • CGIAR GENDER Platform Evaluation
        • CGIAR Excellence in Breeding Platform
        • CGIAR Platform for Big Data in Agriculture
    • Framework and Policy
      • Evaluative Learning Hub
      • Evaluation Method Notes Resource Hub
      • Management Engagement and Response Resource Hub
      • Evaluating Quality of Science for Sustainable Development
      • Evaluability Assessments – Enhancing Pathway to Impact
      • Evaluation Guidelines
  • Independent Science for Development CouncilISDC
  • Standing Panel on Impact AssessmentSPIA
  • Evaluation
Back to IAES Main Menu
  • Home
  • Who we are
  • News
  • Events
  • Publications
  • Featured Projects
    • Inclusive Innovation
      • Agricultural Systems Special Issue
    • Proposal Reviews
      • 2025-30 Portfolio
      • Reform Advice
    • Foresight & Trade-Offs
      • Megatrends
    • QoR4D
    • Comparative Advantage
News

AI Meets the Human Expert: Rethinking Scientific Review Processes

You are here

  • Home
  • Independent Science for Development CouncilISDC
  • News
  • AI Meets the Human Expert: Rethinking Scientific Review Processes

This is the fourth blog on AI use from the Independent Advisory and Evaluation Service. Scroll to the end for other blogs and links. 

 

Scientific reviews are about more than finding and checking information. They require deciding what matters.

That distinction is becoming increasingly important as Artificial Intelligence (AI), and particularly large language models (LLMs), become better at reading, organizing, and synthesizing vast amounts of information. If an LLM can review hundreds of pages in minutes (or even seconds), where does that leave human scientific expertise in the Independent Science for Development Council (ISDC) advisory processes?

For ISDC, this is not a hypothetical question. ISDC provides independent scientific advice to CGIAR across a complex research Portfolio of 13 Programs and Accelerators. This includes reviewing evidence, identifying connections and gaps, and determining which issues matter most for CGIAR decision makers today and in the future.

Earlier this year, ISDC decided to test LLMs. We wanted to understand where LLMs could add value to expert reviews and where human judgment remained essential—not whether LLMs could replace human expertise. An early concept note for this research was shared during the 24th CGIAR System Council Meeting.

Putting AI Alongside the Human Expert

Case Study 1 grew out of a live ISDC strategic document review. ISDC members and subject matter experts reviewed CGIAR’s 13 Programs and Accelerators as part of a broader Portfolio-level assessment. Each completed a structured review using a template based on specified source documents and their scientific expertise.

At the same time, an LLM completed the same template using the same source documents and it synthesized the reviews and identified cross-cutting themes across the Portfolio. The source documents included the Quality of Research for Development in the CGIAR Context  (QoR4D) as framing for the analysis. The process gave ISDC an opportunity to examine where an LLM might strengthen an existing advisory process and where it might fall short.

The comparison raised some practical questions:

  • Can LLM spot patterns across 13 Programs and Accelerators that may be difficult for an individual reviewer to identify? 
  • Can LLM make a review more focused or easier to synthesize? 
  • What happens when reviewing requires knowledge that is not contained in the documents such as scientific experience, institutional context, diverse perspectives, or judgment about what is actually necessary?

Taking Away the Labels

We wanted participants to assess the reviews based on their content, rather than their assumptions and perceptions about AI.

The study therefore used a three-stage approach. In the first stage, ISDC members (N = 8*) compared the LLM review with their own expert review and the resulting synthesis. Because participants were assessing LLM output against work they had personally produced, this stage was explicitly recognized as potentially subject to reviewer bias. The questionnaire therefore examined not only perceived overall quality, but where expert knowledge added value beyond the LLM contribution.

Stage 2 included a broader sample (N = 26) of the same ISDC members, IAES staff and subject matter experts who completed a blind comparison questionnaire. They assessed two reviews with one produced by a human expert and one by the LLM without knowing which was which. Only after completing the comparison were the sources revealed for stage 3, allowing participants to reflect on whether knowing who, or what, produced each review changed their perceptions. 

While the questionnaire responses for stages 2 and 3 from ISDC participants occurred live during an in-person meeting (June 2026), the other respondents completed in June-July 2026. The findings from the ISDC meeting were shared live and with a subsequent facilitated ISDC discussion that provided additional interpretation of the results and surfaced issues related to prompting, synthesis, task suitability, confidentiality, governance, and human accountability. Together, the staged surveys and facilitated discussion provided both quantitative and qualitative evidence for assessing where LLMs may appropriately support ISDC advisory processes. 

Some questions asked included: 

  • Which review showed greater strategic insight? 
  • Which seemed more expert informed? 
  • Which better identified Portfolio-level implications? 
  • Which was more focused? 
  • Which better judged what not to emphasize? 
  • And which would be more useful for decision making?

Although the intent was not a contest, the results did not produce a straightforward winner.

Different Strengths, No Simple Winner

Case study 1 began to reveal a more interesting division of labor. Humans and the LLM brought different capabilities to the review process. The findings also raised important questions about context, efficiency, synthesis, diversity of perspectives, transparency, and responsibility for final scientific judgment.

Some of the results surprised us. 

Others reinforced why independent scientific advice still depends on human expertise and experience.

Two findings give a glimpse of what we found: In the blind comparison, 73 percent of participants judged the LLM review to be more focused than the human expert review. And after revealing which review was produced by LLM, none of the 25 respondents said that knowing its source reduced its credibility. 

But these findings tell only part of the story. Human reviewers showed different strengths, particularly where expertise, context, and judgment mattered. The full case study 1 report will discuss these differences and what they could mean for the future use of LLMs in independent science advice.

What’s Next: Scientific Proposal Review (Case Study 2)

As we finalize Case Study 1, Case Study 2 has launched and is examining the potential role of LLMs in CGIAR’s scientific proposal reviews, which is a primary role of ISDC’s scientific advice.

Proposal reviews present a different challenge. They require reviewers not only to identify and organize information but also to assess them across the CGIAR QoR4D’s four pillars: scientific quality, relevance, legitimacy, and effectiveness. A scientific proposal can be comprehensive, polished, and convincing yet still miss what an experienced scientist recognizes as quality research.

The question, then, is not simply whether LLMs can support independent science advice but is where they add value and where human expertise and judgment remain essential.

Case study 1 provides some initial evidence by comparing perceptions of LLM and human expert reviews. Case study 2 takes the next step, testing LLMs in scientific proposal reviews retrospectively.

Both case study findings will be reported for the 26th Meeting of the CGIAR Council and will be available online in early December with a summary blog to follow. ISDC also intends to submit an adapted paper to a peer-review publication in early 2027. 

* Because of gaps in ISDC member expertise and only seven members (full ISDC membership is eight), two subject matter experts from the IAES roster were included in the ISDC review.

 

READ THE OTHER BLOGS IN IAES'S AI SERIES:

  • Opening the series: What’s a Human to Do? Spotting Valid AI Use Cases in CGIAR’s Advisory Bodies 
  • Accelerating Use of Evidence: Bespoke Messaging for Funders Supported by AI 
  • Working Smarter, Not Just Faster: How is CGIAR’s Evaluation Function Using AI Tools?

Share on

Science ISDC
Sep 15, 2026

Written by

  • Amy R. Beaudreault

    Lead, ISDC Secretariat
  • Glenn Fitzgerald

    ISDC MEMBER

Related News

Posted on
03 Aug 2026
by

Join ISDC: Vacancies for Up to Two New Members

Posted on
08 Jul 2026
by
  • Ines Gonzalez de Suso
  • Magali Garcia

La primera reunión presencial del ISDC en 2026 en Lima: conectando con la ciencia del CIP (y sus papas)

Posted on
30 Mar 2026
by
  • Domagoj Vrbos
  • Lesley Torrance

A Lasting Impact: Reflecting on Nompumelelo Obokoh’s Contributions to ISDC

More News

Related Publications

Assessments & Commentaries
Science ISDC
Issued on 2026

Closing the Advisory Loop: ISDC Reflections on the CGIAR 2025–2030 Portfolio Inception and Implementation

Strategic & Synthesis Studies
Science ISDC
Issued on 2025

Brief Primer on Tools to Prioritize a Research and Innovation Portfolio in a Complex System Involving Multiple Decision Makers

ISDC review cover page
Assessments & Commentaries
Science ISDC
Issued on 2025

ISDC Review of 2025-2030 Research and Innovation Inception Reports

More publications

CGIAR Independent Advisory and Evaluation Service (IAES)

Alliance of Bioversity International and CIAT
Via di San Domenico,1
00153 Rome, Italy
  • IAES@cgiar.org
  • (39-06) 61181

Follow Us

  • LinkedIn
  • Twitter
  • YouTube
JOIN OUR MAILING LIST
  • Terms and conditions
  • © CGIAR 2026

IAES provides operational support as the secretariat for the Independent Science for Development Council and the Standing Panel on Impact Assessment, and implements CGIAR’s multi-year, independent evaluation plan as approved by the CGIAR’s System Council.