The development of effective AI tools for evidence synthesis offers an opportunity to broaden the evaluative evidence base, reduce inequalities, and facilitate more equitable access to data. However, if improperly managed, technological advances have the potential to enhance disparities in evaluation evidence between the global north and south, between information in English and in other languages, and between quantitative data derived from randomised control trials or quasi-experimental studies and more traditional or qualitative information. Participants discussed a range of strategies to ensure that the integration of AI in evaluation and evidence synthesis has a positive impact upon data equity.
“It’s about creating meaningful and relevant insights from a diverse body of material”
Global north and south partnership and engagement. Participants highlighted a fundamental gap in understanding of AI capabilities outside of the global north. Although significant technical knowledge and skills exist outside of the global north there is limited awareness of and engagement with these. Increased collaboration with partners in the global south is required to develop paths to cooperation, investment, and interoperability across different contexts. Meanwhile, models built to suit the requirements of researchers in the global north are typically trained using data and documentation derived from traditionally dominant geographies. This leads to a marginalisation of evaluative data gathered elsewhere and a reduction in their wider accuracy and validity. Participants emphasised the importance of making use of a diversity of training data to minimise the risk of biased outputs, and of developing tools that are better able to integrate local and indigenous knowledge into synthesis platforms and repositories.
Language capabilities. The majority of AI tools typically perform poorly when reading and interpreting languages other than English, resulting in the over-representation of English-language data sources in their outputs. Although the challenges involved in addressing this issue are substantial, participants noted ongoing work aimed at developing large language models (LLMs) capable of understanding languages other than English together with an increased focus on the production of minority language datasets which can be used to train AI models. The necessity to bring this work together under a common governance structure was emphasised, as was the importance of establishing minimum standards for training AI tools using non-English language data.
Machine readability. Due to their limitations in interpreting qualitative data and information written in languages other than English, current AI tools typically struggle to account for informal, indigenous, or community-generated knowledge that forms a key part of the evidence base in many geographic contexts – particularly in the global south. As a result, policy dialogues and exchanges that are often unique to specific settings risk being overlooked despite their centrality to decision making. Active methods for ensuring that AI models can understand and learn from such evidence were called for, with participants noting that evaluators in the north are currently missing out on important innovations emerging from these contexts.
Systems and incentives. Many AI models draw heavily on data published in academic journals. However, systems for publishing articles via these channels can marginalise evidence emerging from the global south, while leading journals often require readers to pay for access. This represents an important barrier to the ability of AI tools to reflect vital and innovative bodies of knowledge emerging from contexts outside of the global north, and creates inequalities in data access. Participants advocated for open-access publication as an effective route to ensuring greater equality in the data that AI models can draw on, while highlighting its role in widening access to evaluation data more broadly. Participants also suggested that current systems which incentivise researchers to publish data in paid-for journals should be revised, with researchers instead encouraged to pursue the more impactful and equitable dissemination of evidence via other means.