Skip to main content

Principles and standards

Monday 24 – Wednesday 26 March 2025

WP3518 event image

With AI tools being developed and rolled out at pace, participants voiced concerns that the necessary focus on standards and rigor for utilising AI in evaluation and evidence synthesis may be lost. To safeguard against this, and to build and maintain trust in the use of AI, participants advocated for the creation of cross-organisational principles and guidelines and debated the validity of imposing standards for the responsible use of AI.

“There’s a need for both principles and standards to engender trust in the use of AI”

Fundamental principles. Participants agreed on a series of high-level principles including transparency and accountability, fairness and inclusivity, data protection and privacy, and validity and reliability. Participants also underlined the fundamental importance of agility and collaboration, while also advocating for the adoption of a human rights-focused approach to the use of AI which places dignity and rights at the centre of technological development.

Drawing on existing frameworks. Although a number of guidance and policy documents covering the use of AI in evaluation exist, many of these are institution-specific while awareness of those that are publicly available is low. Meanwhile, more established frameworks for evaluation and evidence synthesis have often not been adapted to account for the emergence of AI technologies. Participants suggested taking stock of pre-existing guidelines covering the use of AI as a starting point in developing sector-wide policies, and discussed the potential to adapt and modify wider guidance for evaluation and evidence synthesis to cover developments in this area.

Guidelines or standards? Participants distinguished between the development of broad guidelines and the imposition of more rigorous standards, discussing the relative merits and disadvantages of each. Although standards can set clear boundaries and help the market tailor tools to the specific needs of the evaluation community, concerns were raised over their potential to dampen innovation and to quickly become outdated, while also requiring a high level of consensus to be considered legitimate. In contrast, more general guidelines were seen as both flexible and adaptable and capable of being developed with greater efficiency, while broad agreement on the principles that should underpin the use of AI in evaluation and evidence synthesis is already in existence.

Supporting different actors at different stages. While participants voiced a desire to work towards the production of comprehensive guidance covering the use of AI in evaluation and evidence synthesis, the importance of tailoring this to those operating at different points in the evidence ecosystem was flagged. Specific guidance designed to suit the needs of implementors, brokers, and evidence users, accompanied by supportive checklists and tools for their practical application, was seen as having the potential to allow those at different points in the research cycle to maximise the utility of AI. Meanwhile, the value of producing uniform guidance for suppliers bidding for, as well as commissioning, evaluative work was also highlighted. Many tenders are currently unclear on how AI might be used as part of project delivery, and a detailed set of guidelines may allow bidders to be more transparent about their planned use of AI.

Want to find out more?


Sign up to our newsletter