Skip to main content

Key challenges of using AI for evaluation and evidence synthesis

Monday 24 – Wednesday 26 March 2025

WP3518 event image

Despite its significant potential, participants identified several barriers to the adoption of AI in evidence synthesis. These included reservations around the transparency and quality of outputs, a shortfall in appropriate tools and the skills required to use them, concerns around data security and privacy, fears that AI may compound existing geographic and methodological biases, and concerns around the social and environmental impact of AI. More broadly participants recognised that the fundamental conceptualisation and framing of AI remains in flux, with a greater understanding required of its appropriate role and limitations. The production of a ‘dystopian’ theory of change to map out the full range of potential negative outcomes of using AI in evaluation and evidence synthesis was suggested, with participants identifying a series of actions designed to prevent the most severe of these from occurring.

“AI is not a silver bullet”

Transparency. The importance of ensuring that the outputs of AI are traceable for audiences of evaluation and synthesis products was emphasised. Participants warned that too often AI tools represent a ‘black box’, obscuring our understanding of their mechanics, preventing replicability, and making effective human oversight challenging.

Developing principles around transparency was seen as critical to ensure that AI tools are developed and utilised in a responsible way with a focus on public good.

Quality. Although AI tools are capable of achieving a high degree of accuracy when undertaking specific tasks, the potential for AI to hallucinate and falsify information was seen as undermining trust in AI technologies to produce valid and reliable results. Additional concerns were expressed that AI software may miss important outliers during analysis, filtering out complexity to present an artificially clean summary of data.

Participants recommended setting out a clear definition of what is meant by ‘quality’ in the context of AI, creating robust measures against which AI tools can be assessed, and maintaining scientific rigor in AI-assisted synthesis via quality assurance and clear communication to users of uncertainties and limitations.

Tool applicability. Although open-source AI can perform a small number of discrete tasks to a high standard, the capability of off-the-shelf tools to undertake synthesis-specific functions is often limited. Meanwhile, the pace at which technology is developing means that it can often be unclear which tools have been rigorously tested. Although the development of in-house AI has the potential to result in more applicable tools this can often be resource-intensive, while working on tool development in silos risks significant duplication of efforts.

Participants issued a call for researchers to consider what collective action might be taken to shape the market in this area, and for organisations to work more closely together to make the process of developing bespoke AI tools more efficient.

Skills and capability. Uneven levels of awareness and engagement with existing AI tools have led to a high degree of variance in the skills and capabilities of researchers within and across organisations. Meanwhile, the pace at which the field is advancing has resulted in challenges not only in training researchers to use AI but in keeping their skills up to date. Further, placing AI tools in the hands of evidence users without equipping them with an understanding of their limitations risks amplifying unsound or incorrect information.

Ensuring the systematic rollout of AI tools and complementing this with a focus on training and capacity-building among both researchers and evidence users was seen as key to addressing these concerns.

Data security. Concerns around the use of AI to analyse sensitive data has led to researchers typically limiting its use to open-access documentation, while fears around intellectual property infringement have further narrowed the scope of information researchers are willing to expose to AI tools.

A drive to encourage governments and development organisations to broaden access to the evaluation data that they hold was recommended, while participants suggested funding future research only on the condition that data and results are publicly accessible, discoverable, and machine-readable.

Bias. Participants emphasised that any form of bias in the underlying data analysed by AI tools will be reproduced in their outputs, with the results of AI-assisted evaluation and evidence synthesis replicating – or even compounding – the uneven distribution of current evidence across geographies, themes, and methodologies. Meanwhile, as AI models are typically trained using data derived from randomised control trials largely emanating from health literature, such models may struggle to properly capture social science and evaluation publications that may rely more heavily on qualitative or mixed methods.

The importance of developing tools that can accurately interpret a broad range of evaluation literature was underscored, as was the foundational importance of expanding the current evidence base to include evidence produced in the global south and addressing a wider span of thematic issues.

Resources and funding. The time and money required to design, set up, and roll out high-quality AI tools was flagged by participants as a key constraint, with processes for trialling and capacity building requiring additional resource.

Developing sustainable ways of financing innovation in this area was identified as crucial, while communicating the value of AI in evaluation and evidence synthesis to funders was positioned as central to securing future investment.

Environmental and societal impact. Broader issues were also raised about the wider negative impacts of AI technologies. These included fears over the intensive use of resources including water and energy, as well as concerns over the working conditions of those tasked with tagging and annotating large volumes of data. Meanwhile, reservations were aired over the prospect of governments and development organisations designing AI tools for evaluation and evidence synthesis alongside large technology companies who may have vested economic and political interests.

Options proposed for mitigating these concerns included prioritising smaller-scale AI models and ‘green AI’, while a suggestion was offered that these preferences might be enshrined in industry-wide guidelines.

Conceptualisation of AI. Participants also discussed the wider challenges involved in appropriately framing AI in the context of evaluation and evidence synthesis. While acknowledging the gravity of potential issues with hallucination, bias, and reliability, the considerable advantages of using AI in a targeted and pragmatic manner were emphasised. This was encapsulated by the proposition that AI tools should be thought of in a similar way to a human possessing significant aptitude in some areas but lacking skills in others.

Participants advocated for the adoption of an informed and nuanced view of the advantages and risks involved in AI utilisation, noting the importance of putting AI to good use while being alive to its shortcomings.

Want to find out more?


Sign up to our newsletter