Skip to field manual content
EvalCommunity Field Manual

Responsible AI for Monitoring & Evaluation

A practical, human-centred field manual for using AI in M&E while protecting evidence quality, ethics, contextual understanding, and professional accountability.

Download the Field Manual Scrolls to the risk ladder section. Scrolls to the interactive risk checker.
5risk levels for responsible AI use
7M&E cycle stages covered
10ready-to-use practitioner tools
100%browser-based guidance and checklists

Professional commitments

What responsible AI means in M&E practice

This manual is built for real projects: rushed timelines, imperfect data, sensitive contexts, complex stakeholders, and decisions that matter.

01

Human accountability

AI may support a task, but responsibility for analysis, interpretation, recommendations, and final outputs remains with qualified people.

02

Evidence connection

Every important claim should be traceable to source material, not merely to a convincing AI-generated paragraph.

03

Proportional safeguards

The more a task affects meaning, judgement, people, resources, or decisions, the stronger the review process should be.

04

Data care

Personal, confidential, political, commercial, or community-sensitive information requires explicit permission and appropriate protections.

05

Context awareness

AI outputs should be checked for missing cultural, political, historical, organisational, and local context.

06

Transparent records

Teams should be able to explain what AI was used for, what was checked, what changed, and who approved the final work.

Decision framework

The M&E AI risk ladder

Use this ladder to decide whether a task is routine, requires safeguards, or should be avoided without formal approval.

Level1

Routine support

AI improves presentation, wording, or organisation without changing meaning.

ProofreadingFormattingAgendas
Level2

Evidence organisation

AI groups, condenses, or structures information, while humans decide what it means.

TablesSummariesClustering
Level3

Analytical assistance

AI suggests codes, themes, patterns, or comparisons that may shape analysis.

CodesThemesPatterns
Level4

Evaluative interpretation

AI affects findings, explanations, conclusions, recommendations, or performance judgements.

FindingsCausalityAdvice
Level5

Normally inappropriate

AI replaces professional judgement, stakeholder validation, or handles sensitive data without approval.

Final judgementSensitive dataSimulated voices

Workflow map

Responsible AI across the M&E cycle

AI can appear at many stages. The goal is not blanket permission, but thoughtful use with suitable controls.

ScopingAgree allowed uses, disclosure expectations, data boundaries, and review time.
DesignDraft questions, logic models, and indicators, then adapt them to context.
CollectionImprove instruments while checking language, consent, power, and accessibility.
ManagementOrganise material with attention to confidentiality, labelling, and source links.
AnalysisUse AI for support, not replacement. Keep humans close to the evidence.
ReportingCheck every claim, caveat, reference, and recommendation before publishing.
LearningUse AI to prepare discussion, not to substitute stakeholder interpretation.

Quality discipline

Review AI-assisted work before it shapes decisions

The most dangerous output is often the one that sounds polished enough to avoid challenge. Quality checks should be built into the workflow.

Core review questions

  • Is the output supported by the source material?
  • What has been omitted, simplified, or overstated?
  • Are minority, dissenting, or unexpected views still visible?
  • Are causal claims justified by the evidence?
  • Could the evaluator defend the final conclusion without the AI output?

Verification methods

  • Compare AI outputs with original sources.
  • Spot-check summaries, codes, and grouped responses.
  • Map important claims to supporting evidence.
  • Use peer challenge for higher-risk interpretation.
  • Validate sensitive findings with appropriate stakeholders.

Ethics, equity, and power

Protect people, context, and voice

M&E work often involves sensitive realities. AI should not flatten lived experience, expose confidential information, or turn complex contexts into generic findings.

Sensitive information

Check consent, confidentiality, security, and organisational rules before using AI with transcripts, personal data, vulnerable groups, or protected information.

Equity review

Look for missing perspectives, dominant assumptions, biased language, and summaries that erase the experience of smaller or less powerful groups.

Stakeholder meaning

AI may support preparation, but it should not replace participatory interpretation, validation, or locally grounded sense-making.

Governance in practice

Set rules before the work gets complicated

Responsible AI use is easier when expectations are agreed at the start of a project. This section helps teams move from informal experimentation to clear working practice.

Project start-up questions

  • Which AI uses are allowed, conditional, or prohibited?
  • What kinds of data must never be entered into AI tools?
  • Who approves higher-risk uses?
  • How will AI use be disclosed to clients, funders, or participants?

Minimum governance roles

  • Practitioner: records and checks AI-supported work.
  • Project lead: approves risk level and review process.
  • Reviewer: challenges analysis, evidence links, and recommendations.
  • Data or ethics lead: advises on sensitive information.

Escalation triggers

  • The task involves participant-level or identifiable data.
  • The output may influence findings or recommendations.
  • The context is politically, socially, or commercially sensitive.
  • The team cannot explain how the output was checked.

Field manual tools

Practical instruments for M&E teams

These tools turn the guide into an operational field manual: something a team can use before, during, and after AI-supported work.

1. AI use register

Keep a project log of tool used, task, data type, risk level, review method, disclosure decision, and responsible person.

2. Evidence-to-claim map

Link each major finding, conclusion, or recommendation to the evidence that supports it and note any AI involvement.

3. Sensitive data decision tool

Check whether data is public, internal, confidential, personal, politically sensitive, or protected by consent conditions.

4. Professional judgement check

Confirm that the practitioner can explain the finding, uncertainty, alternative explanations, and limits without relying on AI wording.

5. Equity and voice review

Check whether smaller groups, dissenting views, local categories, and marginalised perspectives have been preserved.

6. Disclosure builder

Prepare clear language explaining what AI was used for, what it was not used for, and how human review was maintained.

Tool 1

AI use register

A simple register gives the team a visible record of where AI was used, what data was involved, what risk level was assigned, and who checked the output.

When to use it

  • At the start of any AI-supported project.
  • Whenever AI is used to draft, summarise, classify, code, analyse, translate, or prepare evaluation material.
  • Before delivery, to confirm that all AI-supported outputs have been reviewed.

Minimum fields

  • Date and person responsible.
  • AI tool used and task supported.
  • Data type and sensitivity level.
  • Risk level and review method.
  • Disclosure decision and approval notes.
FieldPromptExample
Task supportedWhat did AI help with?Organised interview notes into preliminary themes.
Data typeWhat material was entered or referenced?Anonymised notes; no names or direct identifiers.
Risk levelWhich risk level applies?Level 3 because AI suggested analytical themes.
Review routeWho checked it and how?Evaluator compared themes with source notes and revised them.

Tool 2

Evidence-to-claim map

This tool prevents polished AI-assisted writing from drifting away from the evidence. It links each major claim to the source material that supports it.

How to build it

  • List every finding, conclusion, and recommendation that may influence decisions.
  • Add the evidence source, page, transcript reference, dataset, or observation note.
  • Mark whether AI contributed to summarising, grouping, drafting, or wording the claim.
  • Record caveats, contradictory evidence, and confidence level.

Review questions

  • Can the claim be defended without relying on AI wording?
  • Has the source evidence been checked by a person?
  • Are minority or dissenting views visible?
  • Is uncertainty preserved rather than smoothed away?
ClaimEvidence sourceAI involvementReviewer decision
Participants reported improved access to support.Interview set A, questions 4-6; service logs.AI summarised recurring points.Accepted after source check; caveat added for rural respondents.
Programme efficiency improved.Budget records and delivery timeline.No AI used for calculation.Accepted; calculation checked manually.

Tool 3

Sensitive data decision tool

Use this section before placing any evaluation material into an AI tool. The safest decision is made before data is copied, uploaded, pasted, or summarised.

Green

  • Public information.
  • No personal identifiers.
  • No confidential project details.
  • AI use allowed with normal review.

Amber

  • Internal material.
  • Potentially sensitive context.
  • Indirect identifiers possible.
  • Use only with approval and safeguards.

Red

  • Personal or participant-level data.
  • Protected, political, or vulnerable-group information.
  • Confidential client or funder material.
  • Do not use without formal approval and a secure tool route.

Decision steps

  • Classify the material before using AI.
  • Remove identifiers where possible.
  • Check consent and contractual restrictions.
  • Use approved tools only for sensitive material.

Stop triggers

  • You cannot explain where the data will be processed.
  • You do not know whether the tool trains on user inputs.
  • The data includes children, vulnerable groups, political exposure, or protected characteristics.
  • The client or ethics process prohibits AI processing.

Tool 4

Professional judgement check

This check confirms that the evaluator, not the AI system, owns the analysis, reasoning, limits, and final interpretation.

Explainability

  • Can the evaluator explain the finding in their own words?
  • Can they identify the evidence that supports it?
  • Can they explain why alternatives were not selected?

Uncertainty

  • Are caveats visible?
  • Are confidence levels proportionate?
  • Has AI language made weak evidence sound stronger than it is?

Accountability

  • Who approved the final wording?
  • Who checked the evidence?
  • Would the team defend the conclusion without showing the AI output?

Pass / revise / escalate

Pass when the evaluator can defend the finding from source evidence. Revise when the output is plausible but overstates, omits, or generalises. Escalate when the output influences major decisions, uses sensitive material, or cannot be verified.

Tool 5

Equity and voice review

AI summaries can flatten difference. This review checks whether smaller groups, dissenting views, local categories, and marginalised perspectives remain visible.

What to check

  • Which voices are most visible in the AI-assisted output?
  • Which groups, locations, languages, or roles are less visible?
  • Were dissenting views merged into the majority view?
  • Did AI replace local terms with generic development language?

Corrective actions

  • Return to source material for underrepresented groups.
  • Separate majority, minority, and divergent findings.
  • Preserve local words and concepts where appropriate.
  • Validate sensitive interpretations with suitable stakeholders.
Review areaQuestionAction if risk appears
RepresentationAre small groups still visible?Add disaggregated notes or separate findings.
PowerDid official voices dominate community voices?Rebalance evidence and cite source groups clearly.
LanguageWere local meanings converted into generic wording?Restore local terminology with explanation.

Tool 6

Disclosure builder

Disclosure should be clear, proportionate, and honest. It should say what AI did, what it did not do, and how human review was maintained.

Low-risk wording

AI was used to support editing, formatting, or readability. The evaluation team reviewed the final text and retained responsibility for meaning and accuracy.

Analysis-support wording

AI was used to assist with organising material and identifying possible patterns. The evaluation team checked outputs against source evidence and made all analytical judgements.

Sensitive-use wording

Where AI-supported processing involved sensitive material, this was handled only through approved systems and review procedures. Human reviewers verified outputs before use.

Disclosure checklist

  • Identify the task AI supported.
  • State that humans retained responsibility for findings and recommendations.
  • Explain the review process.
  • Avoid naming tools where procurement, client, or security rules require different wording.
  • Do not disclose participant information or confidential details in the disclosure text.

Interactive toolkit

AI task risk checker

This lightweight tool runs entirely in the browser. It does not call an AI service, store data remotely, or send information outside the page.

Classify a planned AI use

Select the closest options. Use the result as a discussion aid, not as automatic approval.

Level 1

Routine support

AI can be used for presentation support when evidence, meaning, and judgement remain unchanged.

  • Keep the human author accountable for the final output.
  • Check that wording changes do not alter meaning.
  • Do not enter sensitive data unless approved for the chosen tool.

Reminder: Use the result as a prompt for professional judgement, not as automatic approval.

Templates

Reusable wording and records

Adapt these short templates to your project, client, funder, or organisational policy.

Project: AI tool used: Task supported: Data type: Risk level: Human reviewer: Evidence checked: Disclosure decision: Approval / notes:

Operating rhythm

Make responsible use routine

Simple team habits reduce risk more effectively than one-off rules that nobody revisits.

MomentActionOwnerRecord
Project startAgree allowed, conditional, and prohibited AI uses.Project leadAI use plan
Before using AIClassify the task, data sensitivity, decision influence, and review route.PractitionerRisk checker result
During analysisKeep evidence links and preserve dissenting or minority views.AnalystEvidence-to-claim map
Before deliveryReview findings, claims, caveats, recommendations, and disclosure wording.ReviewerQA checklist
After deliveryCapture lessons about what AI helped, harmed, or should change next time.TeamLearning note
We use cookies to improve your experience on our website. By browsing this website, you agree to our use of cookies.

Cart

Your cart is currently empty.

LAND A JOB REFERRAL IN 2 WEEKS (NO ONLINE APPS!)

Sign Up & To Get My Free Referral Toolkit Now: