# AI Evaluation Playbook

## Introduction

- [About this playbook](https://eval.playbook.org.ai/overview/readme.md)
- [About this playbook](https://eval.playbook.org.ai/overview/readme-1.md)
- [The Process Behind it](https://eval.playbook.org.ai/overview/the-process-behind-this-playbook.md)
- [How to Contribute to the Playbook](https://eval.playbook.org.ai/overview/how-to-contribute-to-the-playbook.md): How to provide feedback, suggest edits, or contribute to the AI Evaluation Playbook.
- [Building Blocks for GenAI Evaluation](https://eval.playbook.org.ai/getting-started/building-blocks-for-genai-evaluation.md)
- [Building the Team](https://eval.playbook.org.ai/getting-started/building-the-team.md)
- [Building the Infrastructure](https://eval.playbook.org.ai/getting-started/building-the-infrastructure.md)
- [Frequently Asked Questions](https://eval.playbook.org.ai/additional-resources/frequently-asked-questions.md)
- [Tools & Templates](https://eval.playbook.org.ai/additional-resources/additional-resources.md)
- [Minimum Viable Evaluations](https://eval.playbook.org.ai/additional-resources/minimum-viable-evaluations.md)
- [Glossary](https://eval.playbook.org.ai/additional-resources/glossary.md)
- [Using the Playbook with AI Tools](https://eval.playbook.org.ai/additional-resources/using-the-playbook-with-ai-tools.md)

## L1 - Model Evaluation

- [Overview](https://eval.playbook.org.ai/model-behaviour/level-1-module-evaluation/overview.md): Does the AI system perform as intended?
- [Who is most involved in this level of evaluation?](https://eval.playbook.org.ai/model-behaviour/level-1-module-evaluation/why-is-this-level-of-evaluation-important.md)
- [What is the “AI system” being evaluated?](https://eval.playbook.org.ai/model-behaviour/level-1-module-evaluation/what-is-the-ai-system-being-evaluated.md)
- [What is the Minimum Viable Evaluation for Level 1?](https://eval.playbook.org.ai/model-behaviour/level-1-module-evaluation/what-is-the-minimum-viable-evaluation-for-level-1.md)
- [How is Level 1 evaluation performed?](https://eval.playbook.org.ai/model-behaviour/how-to-evaluate/how-is-level-1-evaluation-performed.md)
- [Decide on an evaluation rubric](https://eval.playbook.org.ai/model-behaviour/how-to-evaluate/1.-decide-on-an-evaluation-rubric.md)
- [Decide on metrics](https://eval.playbook.org.ai/model-behaviour/how-to-evaluate/2.-decide-on-metrics.md)
- [Develop a golden dataset](https://eval.playbook.org.ai/model-behaviour/how-to-evaluate/3.-develop-a-golden-dataset.md)
- [Scoring & error analysis](https://eval.playbook.org.ai/model-behaviour/how-to-evaluate/4.-scoring-and-error-analysis.md)
- [Automate your evaluations](https://eval.playbook.org.ai/model-behaviour/how-to-evaluate/5.-automate-your-evaluations.md)
- [Red-teaming](https://eval.playbook.org.ai/model-behaviour/how-to-evaluate/6.-red-teaming.md)

## L2 - Product Evaluation

- [Overview](https://eval.playbook.org.ai/product-analytics/level-2-product-evaluation/overview.md): Does the overall product engage and retain users?
- [Who is most involved in this level of evaluation?](https://eval.playbook.org.ai/product-analytics/level-2-product-evaluation/why-is-this-level-of-evaluation-important.md)
- [What is the “Product” being evaluated?](https://eval.playbook.org.ai/product-analytics/level-2-product-evaluation/what-is-the-product-being-evaluated.md)
- [What is the Minimum Viable Evaluation?](https://eval.playbook.org.ai/product-analytics/level-2-product-evaluation/what-is-the-minimum-viable-evaluation.md): We recommend using commercial platforms, when feasible, to track user metrics and automate experiments.
- [How is Level 2 evaluation performed?](https://eval.playbook.org.ai/product-analytics/how-to-evaluate/how-is-level-2-evaluation-performed.md)
- [Methods for experimentation: A/B testing and beyond](https://eval.playbook.org.ai/product-analytics/how-to-evaluate/methods-for-experimentation-a-b-testing-and-beyond.md)
- [Connection with other levels](https://eval.playbook.org.ai/product-analytics/how-to-evaluate/connection-with-other-levels.md)
- [Why Aren’t Users Engaging?](https://eval.playbook.org.ai/product-analytics/how-to-evaluate/why-arent-users-engaging.md)

## L3 - User Evaluation

- [Overview](https://eval.playbook.org.ai/user-experience/level-3-user-evaluation/overview.md): Does the product change users' thoughts, feelings, knowledge and behaviour towards the development outcome?
- [Who is most involved in this level of evaluation?](https://eval.playbook.org.ai/user-experience/level-3-user-evaluation/why-is-this-level-of-evaluation-important.md)
- [Who is the “User” being evaluated?](https://eval.playbook.org.ai/user-experience/level-3-user-evaluation/who-is-the-user-being-evaluated.md)
- [What is the Minimum Viable Evaluation?](https://eval.playbook.org.ai/user-experience/level-3-user-evaluation/what-is-the-minimum-viable-evaluation.md)
- [How is Level 3 evaluation performed?](https://eval.playbook.org.ai/user-experience/how-to-evaluate/how-is-level-3-evaluation-performed.md)
- [Identify outcome metrics](https://eval.playbook.org.ai/user-experience/how-to-evaluate/descriptive-analysis.md)
- [Define guardrail metrics and measure potential harm](https://eval.playbook.org.ai/user-experience/how-to-evaluate/defining-guardrail-metrics-measuring-potential-harm.md)
- [Consider conducting experiments to improve the selected key metrics and running process evaluations](https://eval.playbook.org.ai/user-experience/how-to-evaluate/why-arent-thoughts-feelings-and-behavior-changing.md)
- [Why Aren’t Thoughts, Feelings, and Behavior Changing?](https://eval.playbook.org.ai/user-experience/how-to-evaluate/user-privacy-and-security.md)

## L4 - Impact Evaluation

- [Overview](https://eval.playbook.org.ai/social-impact/level-4-impact-evaluation/overview.md): Do users with access to the product improve development outcomes?
- [Who is involved in this evaluation?](https://eval.playbook.org.ai/social-impact/level-4-impact-evaluation/why-is-this-level-of-evaluation-important.md)
- [What is the “intervention” being evaluated?](https://eval.playbook.org.ai/social-impact/level-4-impact-evaluation/what-is-the-intervention-being-evaluated.md)
- [Minimum Viable Evaluation](https://eval.playbook.org.ai/social-impact/level-4-impact-evaluation/minimum-viable-evaluation.md)
- [How is Level 4 evaluation performed?](https://eval.playbook.org.ai/social-impact/how-to-evaluate/how-is-level-4-evaluation-performed.md)
- [A Quick Primer on Impact Evaluation Methods](https://eval.playbook.org.ai/social-impact/how-to-evaluate/a-quick-primer-on-impact-evaluation-methods.md)
- [Key design considerations for AI-specific impact evaluations](https://eval.playbook.org.ai/social-impact/how-to-evaluate/key-design-considerations-for-ai-specific-impact-evaluations.md)
- [Common pitfalls to avoid](https://eval.playbook.org.ai/social-impact/how-to-evaluate/common-pitfalls-to-avoid.md)
- [Process Evaluation: Why Aren’t Outcomes Changing?](https://eval.playbook.org.ai/social-impact/how-to-evaluate/process-evaluation-why-arent-outcomes-changing.md)

## Level Linkages

- [Overview](https://eval.playbook.org.ai/level-linkages/linkage-across-levels/overview.md)
- [Risk assessment and mitigation](https://eval.playbook.org.ai/level-linkages/linkage-across-levels/risk-assessment-and-mitigation.md)
- [Data protection](https://eval.playbook.org.ai/level-linkages/linkage-across-levels/data-protection.md)
- [Process Evaluations](https://eval.playbook.org.ai/level-linkages/linkage-across-levels/process-evaluations.md)
- [Do I need a Process Evaluation?](https://eval.playbook.org.ai/level-linkages/linkage-across-levels/process-evaluations/do-i-need-a-process-evaluation.md)
- [What does it take to do a process evaluation?](https://eval.playbook.org.ai/level-linkages/linkage-across-levels/process-evaluations/what-does-it-take-to-do-a-process-evaluation.md)
