> For the complete documentation index, see [llms.txt](https://eval.playbook.org.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://eval.playbook.org.ai/additional-resources/resources-for-funders.md).

# Resources for Funders

## Purpose

This tool helps funders assess the evaluation maturity of organizations building or implementing generative AI applications in the development sector. It is a ready-to-use adaptation of the Generative AI Evaluation Playbook for the Development Sector. Having a response that reveals the evaluation practices are underdeveloped does not mean the project should not be funded. It does however suggest that the applicant should close those gaps before being considered impactful, and before large scale public procurement.

***

### 15 questions funders can ask applicants

<details>

<summary>Level 1: Does the AI system perform as intended?</summary>

<table data-header-hidden="false" data-header-sticky><thead><tr><th width="62" valign="top">#</th><th valign="top">Goal</th><th width="123.33331298828125" valign="top">Question</th><th valign="top">Example Answers</th><th valign="top">Playbook sections for further reference</th></tr></thead><tbody><tr><td valign="top">1</td><td valign="top">Determine if the applicant has identified the AI system behaviors and performance that matter most for their use case</td><td valign="top"><p>Have you come up with a rubric to assess your AI system’s performance?  </p><p> </p><p>What factors do you assess to determine if your AI product is performing as intended? </p></td><td valign="top"><p><strong>Well developed:</strong> </p><ul><li>Identifies and uses 3-5 dimensions most relevant to their use case </li><li>Metrics for each dimension have quantitative thresholds they aim to meet before deployment. They have reasoned why these thresholds are appropriate for their industry. </li><li>E.g. “We have 4 dimensions—accuracy, safety, robustness, and linguistic consistency. We define safety as X which is line with standard Y. We measure performance before release and after against metrics associated with these dimensions.” </li></ul><p></p><p><strong>Underdeveloped:</strong>  </p><ul><li>E.g. “We don’t measure internally but users keep coming back. It’s accurate enough.” </li><li>E.g. “We have expert users who give us feedback” </li></ul></td><td valign="top"><p>How to identify dimensions most relevant to your use case: <a href="https://eval.playbook.org.ai/model-behaviour/how-to-evaluate/1.-decide-on-an-evaluation-rubric">Decide on an evaluation rubric</a> </p><p> </p><p>How to identify metrics for each dimension: <a href="https://eval.playbook.org.ai/model-behaviour/how-to-evaluate/2.-decide-on-metrics">Decide on metrics</a> </p></td></tr><tr><td valign="top">2</td><td valign="top">Assess the quality of their golden dataset, and thoughtfulness in design</td><td valign="top">Have you developed a golden dataset? How was it developed?</td><td valign="top"><p><strong>Well developed:</strong> </p><ul><li>Has a golden dataset that is representative of real-world use </li><li>E.g. “We developed our own using a select group of experts who are representative of our clinics across the country. We continue to add to our golden dataset based on new questions we get from live use.” </li><li>E.g. “We started with a public dataset for a use case similar to ours, and have customized further based on new questions from our users.” </li></ul><p> </p><p><strong>Underdeveloped:</strong> </p><ul><li>E.g. “We came up with a golden dataset based on LLM generated answers and queries, and got an expert to validate.” </li></ul></td><td valign="top">How to create a golden dataset that is representative of real-world use: <a href="https://eval.playbook.org.ai/model-behaviour/how-to-evaluate/3.-develop-a-golden-dataset">Develop a golden dataset</a>  </td></tr><tr><td valign="top">3</td><td valign="top">Assess if applicant has mitigated potential for harm and malicious use</td><td valign="top">Have you done any adversarial testing?</td><td valign="top"><p><strong>Well developed:</strong> </p><ul><li>Well defined and ranked red lines with failure modes documented </li><li>Red team composed of domain expert and engineers </li><li>Red teaming exercise before significant new releases and long context windows </li></ul><p></p><p><strong>Underdeveloped:</strong> </p><ul><li>We have prompts that prevent the model from malicious use or harmful output </li><li>We have tested edge cases </li></ul></td><td valign="top">How to conduct adversarial testing: <a href="https://eval.playbook.org.ai/model-behaviour/how-to-evaluate/6.-red-teaming">Red-teaming</a> </td></tr></tbody></table>

</details>

<details>

<summary>Level 2: Does the overall product engage and retain users?</summary>

<table data-header-hidden="false" data-header-sticky><thead><tr><th width="62" valign="top">#</th><th valign="top">Goal</th><th valign="top">Question</th><th valign="top">Example Answers</th><th valign="top">Playbook sections for further reference</th></tr></thead><tbody><tr><td valign="top">1</td><td valign="top">Assess how well the applicant knows and tracks their user journey</td><td valign="top">Please describe your user journey and how you track it</td><td valign="top"><p><strong>Well developed:</strong> </p><ul><li><p>Well developed:  </p><ul><li>Defined user stages such as user activation and retention. </li><li>Has identified and instrumented events in the app that signal user journey through these stages </li><li>Has automatically generated and monitored metrics for user stages  </li><li>Can identify and overcome friction points </li></ul></li></ul><p> </p><p><strong>Underdeveloped:</strong> </p><ul><li>No digital instrumentation to automatically track user journey </li><li>Reports minimal metrics of product use such as</li></ul></td><td valign="top"><p>How to identify user stages: <a href="https://eval.playbook.org.ai/getting-started/building-the-infrastructure#id-2.-the-user-funnel-track-the-journey-across-levels-1-4">The user funnel: track the journey across Levels 1-4,</a> </p><p><a href="https://eval.playbook.org.ai/product-analytics/how-to-evaluate/how-is-level-2-evaluation-performed#define-the-user-funnel-and-metrics">Define the user funnel and metrics</a> </p><p> </p><p>How to identify metrics for each stage: <a href="https://eval.playbook.org.ai/product-analytics/level-2-product-evaluation/what-is-the-product-being-evaluated">What is the “Product” being evaluated?</a> </p><p> </p><p>How to automate metrics: <a href="https://eval.playbook.org.ai/product-analytics/how-to-evaluate/how-is-level-2-evaluation-performed#automate-metrics-and-analyze-trends">Automate metrics and analyze trends</a> </p><p> </p><p>How to identify and overcome fraction points: <a href="https://eval.playbook.org.ai/product-analytics/how-to-evaluate/how-is-level-2-evaluation-performed#identify-frictions-and-design-improvements">Identify frictions and design improvements</a> </p></td></tr><tr><td valign="top">2</td><td valign="top">Assess how methodical the applicant is when they roll out new product features or upgrades</td><td valign="top">How do you evaluate if a new feature you rolled out drove engagement?</td><td valign="top"><p><strong>Well developed:</strong>  </p><ul><li>Uses A/B testing or dynamic assignment with preset success metrics </li></ul><p> </p><p><strong>Underdeveloped:</strong> </p><ul><li>E.g. “We roll out new features and watch the metrics from the user journey”</li></ul></td><td valign="top">How to conduct A/B testing: <a href="https://eval.playbook.org.ai/product-analytics/how-to-evaluate/methods-for-experimentation-a-b-testing-and-beyond">Methods for experimentation: A/B testing and beyond</a> </td></tr><tr><td valign="top">3</td><td valign="top">Assess the applicant’s ability to troubleshoot</td><td valign="top">What do you do when the product is not being used the way you thought it would be?</td><td valign="top"><p><strong>Well developed:</strong>  </p><ul><li>Uses qualitative (interviews, observations) or quantitative methods (surveys) to understand product use </li><li>Uses learnings and insights to overcome friction points and barriers </li></ul><p> </p><p><strong>Underdeveloped:</strong> </p><ul><li>Cannot describe how they overcame product challenges in a systematic way </li></ul></td><td valign="top">How to troubleshoot: <a href="https://eval.playbook.org.ai/product-analytics/how-to-evaluate/why-arent-users-engaging">Why Aren't Users Engaging?</a> </td></tr></tbody></table>

</details>

<details>

<summary>Level 3: Does the product change users' thoughts, feelings, knowledge, and behavior towards the development outcome?</summary>

<table data-header-hidden="false" data-header-sticky><thead><tr><th width="62" valign="top">#</th><th valign="top">Goal</th><th valign="top">Question</th><th valign="top">Example Answers</th><th valign="top">Playbook sections for further reference</th></tr></thead><tbody><tr><td valign="top">1</td><td valign="top">Assess if the applicant can explain how a user’s thoughts, knowledge, or behavior is meant to impact the development outcome</td><td valign="top">What user thought, knowledge or behavior is your product expected to change?</td><td valign="top"><p><strong>Well developed:</strong>  </p><ul><li>Can identify 1-2 relevant user behavioral and cognitive measures that aligns with the theory of change </li><li>Can list at least one user affect or behavior indicative of harm </li></ul><p> </p><p><strong>Underdeveloped:</strong>  </p><ul><li>Cannot satisfactorily name user behavior or cognition that is expected to change with product use </li><li>Gives AI metric such as hallucination as user harm </li></ul></td><td valign="top"><p>How to develop a Theory of Change: <a href="https://eval.playbook.org.ai/getting-started/building-the-infrastructure#id-1.-the-foundation-start-with-formative-research-a-theory-of-change-and-subgroup-identification">The Foundation: Start with formative research, a theory of change, and subgroup identification</a> </p><p> </p><p>How to identify user behavioral and cognitive measures: <a href="https://eval.playbook.org.ai/user-experience/how-to-evaluate/descriptive-analysis">Identify outcome metrics</a> </p><p> </p><p>How to identify potential harm: <a href="https://eval.playbook.org.ai/user-experience/how-to-evaluate/defining-guardrail-metrics-measuring-potential-harm">Define guardrail metrics and measure potential harm</a> </p></td></tr><tr><td valign="top">2</td><td valign="top">Assess whether the applicant has chosen meaningful user evaluation metrics</td><td valign="top">How do you measure if users' thoughts, behavior, or knowledge are changing with use?</td><td valign="top"><p><strong>Well developed:</strong> </p><ul><li>Combines at least one behavioral or trace metric from user interaction with short, well-timed, survey </li></ul><p> </p><p><strong>Underdeveloped:</strong> </p><ul><li>Has not defined how to measure the user’s behavior or knowledge </li></ul></td><td valign="top">How to conduct A/B testing: How to improve selected metrics: <a href="https://eval.playbook.org.ai/user-experience/how-to-evaluate/why-arent-thoughts-feelings-and-behavior-changing">Consider conducting experiments to improve the selected key metrics and running process evaluations</a> </td></tr><tr><td valign="top">3</td><td valign="top">Assess if the applicant has validated user interaction tracers as proxy measures for user feelings, knowledge and behavior</td><td valign="top">How do you know if the user interaction tracers is reliable and sufficient for user evaluation?</td><td valign="top"><p><strong>Well developed:</strong>  </p><ul><li>Can report external checks (focus groups, expert validation) to ensure interaction tracers reflect real-world measures </li></ul><p> </p><p><strong>Underdeveloped:</strong> </p><ul><li>Highlights only selected user stories </li></ul></td><td valign="top">How to connect to the broader program ToC: <a href="https://eval.playbook.org.ai/user-experience/how-to-evaluate/user-privacy-and-security">Why Aren’t Thoughts, Feelings, and Behavior Changing?</a> </td></tr></tbody></table>

</details>

<details>

<summary>Level 4: Does your product use improve development outcomes?</summary>

<table data-header-hidden="false" data-header-sticky><thead><tr><th width="62" valign="top">#</th><th valign="top">Goal</th><th valign="top">Question</th><th valign="top">Example Answers</th><th valign="top">Playbook sections for further reference</th></tr></thead><tbody><tr><td valign="top">1</td><td valign="top">Assess if the applicant is ready for impact evaluation</td><td valign="top">What is the development outcome you are trying to impact? What makes you confident that you are achieving it?</td><td valign="top"><p><strong>Well developed:</strong>  </p><ul><li>Can give a specific and measurable development outcome </li><li>Has robust results from the other levels and can logically connect them to the expected development outcome </li></ul><p> </p><p><strong>Underdeveloped:</strong> </p><ul><li>Uses vague goals such as improvements in “wellbeing” and “access” as intended development outcomes </li><li>Reports product use, user satisfaction and anecdotes as readiness for impact evaluation </li></ul></td><td valign="top">How to conduct impact evaluation: <a href="https://eval.playbook.org.ai/social-impact/level-4-impact-evaluation/what-is-the-intervention-being-evaluated">What is the “intervention” being evaluated?</a>  </td></tr><tr><td valign="top">2</td><td valign="top">Assess whether the applicant collects cost and implementation data needed for scale up decision</td><td valign="top"><p>Do you track how much it costs per user per month to use your product? </p><p> </p><p>How do you understand if your product is deployed and implemented well? </p></td><td valign="top"><p><strong>Well developed:</strong>  </p><ul><li>Can report technology costs, acquisition costs, training needs, staff time, maintenance costs from large scale deployment or alongside outcomes data </li><li>Can track implementation fidelity and delivery quality so that evaluators can explain why outcomes did or did not change </li></ul><p> </p><p><strong>Underdeveloped:</strong> </p><ul><li>Cannot report what costs and other resources are needed to operate product/services at large scale. </li><li>Cannot monitor if the intervention is delivered as intended </li></ul></td><td valign="top">How to do AI-specific impact evaluations: <a href="https://eval.playbook.org.ai/social-impact/how-to-evaluate/key-design-considerations-for-ai-specific-impact-evaluations">Key design considerations for AI-specific impact evaluations</a> </td></tr><tr><td valign="top">3</td><td valign="top">Assess if the applicant has identified an evaluation partner that can manage outcome evaluation risks such as attrition, spillovers and product dynamism</td><td valign="top">How will your research partner manage model or product changes, external dependencies, spillovers and user drop offs?</td><td valign="top"><p><strong>Well developed:</strong>  </p><ul><li>Uses version control or freezes product or limits changes </li><li>Can measure exposure through product to limit spillovers </li><li>Uses level 2 data to help evaluators generate necessary sample sizes </li></ul><p> </p><p><strong>Underdeveloped:</strong> </p><ul><li>Cannot control model and product versions during study period </li><li>Does not have up-to-date level 1 or 2 data</li></ul></td><td valign="top">How to do AI-specific impact evaluations: <a href="https://eval.playbook.org.ai/social-impact/how-to-evaluate/key-design-considerations-for-ai-specific-impact-evaluations">Key design considerations for AI-specific impact evaluations</a> </td></tr></tbody></table>

</details>

<details>

<summary>Resourcing: Do you have the team and resources for effective evaluations?</summary>

<table data-header-hidden="false" data-header-sticky><thead><tr><th width="62" valign="top">#</th><th valign="top">Goal</th><th valign="top">Question</th><th valign="top">Example Answers</th><th valign="top">Playbook sections for further reference</th></tr></thead><tbody><tr><td valign="top">1</td><td valign="top">Assess if the applicant has a cross-functional team that can effectively manage L1-L4 evaluations</td><td valign="top">Who conducts the various levels of evaluation?</td><td valign="top"><p><strong>Well developed:</strong>  </p><ul><li>Names roles such as engineer or data science lead for model evaluation, error analysis etc. </li></ul><p></p><p><strong>Underdeveloped:</strong> </p><ul><li>Cannot name roles with primary responsibility for different evaluation levels. </li><li>Traditional M&#x26;E team handles evaluations without domain experts or engineering/product team members</li></ul></td><td valign="top">How to build a cross-functional team: <a href="https://eval.playbook.org.ai/getting-started/building-the-team">Building the team</a>  </td></tr><tr><td valign="top">2</td><td valign="top">Assess level of cross-functional collaboration</td><td valign="top">Can you give some examples of how different members of your team work together to conduct a particular evaluation?</td><td valign="top"><p><strong>Well developed:</strong> </p><ul><li>Engineers working with domain experts to create golden datasets and evaluate model outputs. </li><li>Can give shared definition of success understood by all team members at various levels of evaluation (L1 accuracy, L2 engagement) </li><li>Regular use of dashboards and other tools for key metrics accessibility and joint review by all with escalation pathways for safety issues, user harm and drop offs </li></ul><p> </p><p><strong>Underdeveloped:</strong> </p><ul><li>Teams review their metrics in isolation  </li><li>Product changes without involvement of various evaluation team members </li></ul></td><td valign="top">How to ensure cross-functional collaboration: <a href="https://eval.playbook.org.ai/getting-started/building-the-team#best-practices-for-cross-level-collaboration">Best practices for cross-level collaboration</a> </td></tr><tr><td valign="top">3</td><td valign="top">Assess if the team understands their blind spots and when to bring in external help</td><td valign="top">What evaluation skill gaps does your team have and are you trying to hire or partner with external consultants?</td><td valign="top"><p><strong>Well developed:</strong> </p><ul><li>E.g. “we’re great on tracking user engagement, but do not know how to do sentiment analysis and so are working with behavioral scientists at research institution X to conduct it and recommend what to change in our product.”  </li></ul><p> </p><p><strong>Underdeveloped:</strong> </p><ul><li>Cannot identify gaps, or identifies gaps across many levels  </li></ul></td><td valign="top"><p>What are Minimum Viable Evaluations (MVE) across levels: <a href="https://eval.playbook.org.ai/additional-resources/minimum-viable-evaluations">Minimum Viable Evaluations</a> </p><p> </p><p>Who’s most involved at each level:  </p><p><a href="https://eval.playbook.org.ai/model-behaviour/level-1-module-evaluation/why-is-this-level-of-evaluation-important">Level 1 Model,</a> </p><p><a href="https://eval.playbook.org.ai/product-analytics/level-2-product-evaluation/why-is-this-level-of-evaluation-important">Level 2 Product,</a> </p><p><a href="https://eval.playbook.org.ai/user-experience/level-3-user-evaluation/why-is-this-level-of-evaluation-important">Level 3 User,</a> </p><p><a href="https://eval.playbook.org.ai/social-impact/level-4-impact-evaluation/why-is-this-level-of-evaluation-important">Level 4 Impact</a> </p></td></tr></tbody></table>

</details>
