For the complete documentation index, see llms.txt. This page is also available as Markdown.

Resources for Funders

15 Questions for Funders to Assess an Applicant’s AI Evaluation Maturity

Purpose

This tool helps funders assess the evaluation maturity of organizations building or implementing generative AI applications in the development sector. It is a ready-to-use adaptation of the Generative AI Evaluation Playbook for the Development Sector. Having a response that reveals the evaluation practices are underdeveloped does not mean the project should not be funded. It does however suggest that the applicant should close those gaps before being considered impactful, and before large scale public procurement.


15 questions funders can ask applicants

Level 1: Does the AI system perform as intended?
#
Goal
Question
Example Answers
Playbook sections for further reference

1

Determine if the applicant has identified the AI system behaviors and performance that matter most for their use case

Have you come up with a rubric to assess your AI system’s performance?

What factors do you assess to determine if your AI product is performing as intended?

Well developed:

  • Identifies and uses 3-5 dimensions most relevant to their use case

  • Metrics for each dimension have quantitative thresholds they aim to meet before deployment. They have reasoned why these thresholds are appropriate for their industry.

  • E.g. “We have 4 dimensions—accuracy, safety, robustness, and linguistic consistency. We define safety as X which is line with standard Y. We measure performance before release and after against metrics associated with these dimensions.”

Underdeveloped:

  • E.g. “We don’t measure internally but users keep coming back. It’s accurate enough.”

  • E.g. “We have expert users who give us feedback”

How to identify dimensions most relevant to your use case: Decide on an evaluation rubric

How to identify metrics for each dimension: Decide on metrics

2

Assess the quality of their golden dataset, and thoughtfulness in design

Have you developed a golden dataset? How was it developed?

Well developed:

  • Has a golden dataset that is representative of real-world use

  • E.g. “We developed our own using a select group of experts who are representative of our clinics across the country. We continue to add to our golden dataset based on new questions we get from live use.”

  • E.g. “We started with a public dataset for a use case similar to ours, and have customized further based on new questions from our users.”

Underdeveloped:

  • E.g. “We came up with a golden dataset based on LLM generated answers and queries, and got an expert to validate.”

How to create a golden dataset that is representative of real-world use: Develop a golden dataset

3

Assess if applicant has mitigated potential for harm and malicious use

Have you done any adversarial testing?

Well developed:

  • Well defined and ranked red lines with failure modes documented

  • Red team composed of domain expert and engineers

  • Red teaming exercise before significant new releases and long context windows

Underdeveloped:

  • We have prompts that prevent the model from malicious use or harmful output

  • We have tested edge cases

How to conduct adversarial testing: Red-teaming

Level 2: Does the overall product engage and retain users?
#
Goal
Question
Example Answers
Playbook sections for further reference

1

Assess how well the applicant knows and tracks their user journey

Please describe your user journey and how you track it

Well developed:

  • Well developed:

    • Defined user stages such as user activation and retention.

    • Has identified and instrumented events in the app that signal user journey through these stages

    • Has automatically generated and monitored metrics for user stages

    • Can identify and overcome friction points

Underdeveloped:

  • No digital instrumentation to automatically track user journey

  • Reports minimal metrics of product use such as

How to identify user stages: The user funnel: track the journey across Levels 1-4,

Define the user funnel and metrics

How to identify metrics for each stage: What is the “Product” being evaluated?

How to automate metrics: Automate metrics and analyze trends

How to identify and overcome fraction points: Identify frictions and design improvements

2

Assess how methodical the applicant is when they roll out new product features or upgrades

How do you evaluate if a new feature you rolled out drove engagement?

Well developed:

  • Uses A/B testing or dynamic assignment with preset success metrics

Underdeveloped:

  • E.g. “We roll out new features and watch the metrics from the user journey”

3

Assess the applicant’s ability to troubleshoot

What do you do when the product is not being used the way you thought it would be?

Well developed:

  • Uses qualitative (interviews, observations) or quantitative methods (surveys) to understand product use

  • Uses learnings and insights to overcome friction points and barriers

Underdeveloped:

  • Cannot describe how they overcame product challenges in a systematic way

How to troubleshoot: Why Aren't Users Engaging?

Level 3: Does the product change users' thoughts, feelings, knowledge, and behavior towards the development outcome?
#
Goal
Question
Example Answers
Playbook sections for further reference

1

Assess if the applicant can explain how a user’s thoughts, knowledge, or behavior is meant to impact the development outcome

What user thought, knowledge or behavior is your product expected to change?

Well developed:

  • Can identify 1-2 relevant user behavioral and cognitive measures that aligns with the theory of change

  • Can list at least one user affect or behavior indicative of harm

Underdeveloped:

  • Cannot satisfactorily name user behavior or cognition that is expected to change with product use

  • Gives AI metric such as hallucination as user harm

How to develop a Theory of Change: The Foundation: Start with formative research, a theory of change, and subgroup identification

How to identify user behavioral and cognitive measures: Identify outcome metrics

How to identify potential harm: Define guardrail metrics and measure potential harm

2

Assess whether the applicant has chosen meaningful user evaluation metrics

How do you measure if users' thoughts, behavior, or knowledge are changing with use?

Well developed:

  • Combines at least one behavioral or trace metric from user interaction with short, well-timed, survey

Underdeveloped:

  • Has not defined how to measure the user’s behavior or knowledge

3

Assess if the applicant has validated user interaction tracers as proxy measures for user feelings, knowledge and behavior

How do you know if the user interaction tracers is reliable and sufficient for user evaluation?

Well developed:

  • Can report external checks (focus groups, expert validation) to ensure interaction tracers reflect real-world measures

Underdeveloped:

  • Highlights only selected user stories

How to connect to the broader program ToC: Why Aren’t Thoughts, Feelings, and Behavior Changing?

Level 4: Does your product use improve development outcomes?
#
Goal
Question
Example Answers
Playbook sections for further reference

1

Assess if the applicant is ready for impact evaluation

What is the development outcome you are trying to impact? What makes you confident that you are achieving it?

Well developed:

  • Can give a specific and measurable development outcome

  • Has robust results from the other levels and can logically connect them to the expected development outcome

Underdeveloped:

  • Uses vague goals such as improvements in “wellbeing” and “access” as intended development outcomes

  • Reports product use, user satisfaction and anecdotes as readiness for impact evaluation

How to conduct impact evaluation: What is the “intervention” being evaluated?

2

Assess whether the applicant collects cost and implementation data needed for scale up decision

Do you track how much it costs per user per month to use your product?

How do you understand if your product is deployed and implemented well?

Well developed:

  • Can report technology costs, acquisition costs, training needs, staff time, maintenance costs from large scale deployment or alongside outcomes data

  • Can track implementation fidelity and delivery quality so that evaluators can explain why outcomes did or did not change

Underdeveloped:

  • Cannot report what costs and other resources are needed to operate product/services at large scale.

  • Cannot monitor if the intervention is delivered as intended

How to do AI-specific impact evaluations: Key design considerations for AI-specific impact evaluations

3

Assess if the applicant has identified an evaluation partner that can manage outcome evaluation risks such as attrition, spillovers and product dynamism

How will your research partner manage model or product changes, external dependencies, spillovers and user drop offs?

Well developed:

  • Uses version control or freezes product or limits changes

  • Can measure exposure through product to limit spillovers

  • Uses level 2 data to help evaluators generate necessary sample sizes

Underdeveloped:

  • Cannot control model and product versions during study period

  • Does not have up-to-date level 1 or 2 data

How to do AI-specific impact evaluations: Key design considerations for AI-specific impact evaluations

Resourcing: Do you have the team and resources for effective evaluations?
#
Goal
Question
Example Answers
Playbook sections for further reference

1

Assess if the applicant has a cross-functional team that can effectively manage L1-L4 evaluations

Who conducts the various levels of evaluation?

Well developed:

  • Names roles such as engineer or data science lead for model evaluation, error analysis etc.

Underdeveloped:

  • Cannot name roles with primary responsibility for different evaluation levels.

  • Traditional M&E team handles evaluations without domain experts or engineering/product team members

How to build a cross-functional team: Building the team

2

Assess level of cross-functional collaboration

Can you give some examples of how different members of your team work together to conduct a particular evaluation?

Well developed:

  • Engineers working with domain experts to create golden datasets and evaluate model outputs.

  • Can give shared definition of success understood by all team members at various levels of evaluation (L1 accuracy, L2 engagement)

  • Regular use of dashboards and other tools for key metrics accessibility and joint review by all with escalation pathways for safety issues, user harm and drop offs

Underdeveloped:

  • Teams review their metrics in isolation

  • Product changes without involvement of various evaluation team members

How to ensure cross-functional collaboration: Best practices for cross-level collaboration

3

Assess if the team understands their blind spots and when to bring in external help

What evaluation skill gaps does your team have and are you trying to hire or partner with external consultants?

Well developed:

  • E.g. “we’re great on tracking user engagement, but do not know how to do sentiment analysis and so are working with behavioral scientists at research institution X to conduct it and recommend what to change in our product.”

Underdeveloped:

  • Cannot identify gaps, or identifies gaps across many levels

What are Minimum Viable Evaluations (MVE) across levels: Minimum Viable Evaluations

Who’s most involved at each level:

Level 1 Model,

Level 2 Product,

Level 3 User,

Level 4 Impact

Last updated

Was this helpful?