⇓ More from ICTworks

How to Evaluate Artificial Intelligence Systems for LMIC Usage

By Wayan Vota on September 22, 2026

Artificial Intelligence evaluation

Are we building AI systems backwards, evaluating them wrongly, and then wondering why our multi-million dollar algorithms fail to deliver the outcomes we promised donors?

After analyzing the latest research from the Center for Global Development alongside recent humanitarian AI deployments, I’m convinced that our sector’s approach to AI evaluation is fundamentally flawed. And that is killing good projects before they even have a chance to succeed.

The problem isn’t that humanitarian AI lacks potential.

The problem is that we keep jumping straight to impact evaluations before ensuring our AI systems function properly—a costly mistake that’s become endemic across the sector.

Sign Up Now for more digital development insights 

We’re Measuring the Wrong Thing

Asking “Does this AI chatbot reduce infant mortality?” before confirming “Does this chatbot provide medically accurate responses?” is professional malpractice.

Yet this is exactly what humanitarian organizations do when they rush into randomized controlled trials for AI systems that haven’t been properly validated at the technical level. Spoiler alert: this backwards approach is why so many AI pilots never make it to scale.

We have over 53 published ethical AI guidelines in humanitarian contexts, yet a 2024 Wilton Park report found that many AI applications are still being deployed without sufficient evidence of their effectiveness, feasibility, or ethical implications.

This is actively harmful. When poorly evaluated AI systems fail in the field, they don’t just waste resources. They erode community trust, endanger vulnerable populations, and set back the entire humanitarian technology agenda.

The Four-Level Solution That Works

ai evaluation levels

The Center for Global Development’s new evaluation framework offers a systematic alternative to our current chaos. Instead of jumping to impact measurement, they propose four sequential evaluation levels that build credible evidence step by step.

Here’s how humanitarian organizations should implement this framework:

Level 1: Model Evaluation – Does Your AI Behave Correctly?

Before deploying any AI system, you must verify it performs its intended function accurately and consistently. For humanitarian AI, this means rigorous testing against expert-validated responses.

Take medical AI chatbots serving refugee populations.

Level 1 evaluation requires creating reference datasets where medical experts provide ideal responses to common health queries. The AI’s outputs are then scored against these references for accuracy, empathy, and appropriateness. If your AI chatbot recommends dangerous treatments or provides culturally insensitive advice, you’ll discover this during controlled testing—not after deployment in a refugee camp.

The key insight from CGD’s research is that unlike human-operated services, AI allows rapid modification and testing cycles. Organizations should leverage this advantage by conducting multiple evaluation rounds with different prompts and configurations before settling on a final model.

Level 2: Product Evaluation – Are People Using It?

Technical accuracy means nothing if your target users abandon the system after first contact. Level 2 evaluation focuses on engagement metrics and user experience optimization.

Humanitarian organizations need to track user interaction patterns systematically.

  • How long do displaced persons spend interacting with your AI information system?
  • Do they return for follow-up queries?
  • Are usage patterns consistent across different demographic groups?

Tools like IDinsight’s Experiments Engine and Agency Fund’s Evidential platform enable rapid A/B testing of different interface designs, conversation flows, and content strategies. This level of evaluation often reveals critical barriers—like language preferences or cultural communication norms—that technical testing misses entirely.

Level 3: User Evaluation – Is It Changing Behavior Appropriately?

Even high engagement doesn’t guarantee meaningful outcomes. Level 3 evaluation assesses whether AI interaction produces the cognitive, affective, and behavioral changes needed for downstream impact.

For AI-powered health education systems, this means measuring whether users gain accurate knowledge, feel confident about health decisions, and take appropriate preventive actions. These intermediate outcomes serve as early indicators of eventual health impact while remaining faster and cheaper to measure than clinical endpoints.

This evaluation stage typically combines usage analytics with user surveys, qualitative interviews, and behavioral tracking. The goal is confirming that your AI system creates the psychological and behavioral preconditions necessary for ultimate impact.

Level 4: Independent Impact Evaluation – Does It Improve Lives?

Only after confirming technical reliability, user engagement, and behavioral change should organizations invest in randomized impact evaluations. At this stage, you’re measuring whether AI-enhanced interventions produce better development outcomes than traditional approaches.

The Brazilian state of Espírito Santo exemplifies this approach.

They piloted the Letrus AI writing platform with rigorous evaluation by J-PAL researchers, finding that students wrote more essays, received higher-quality feedback, and scored better on national writing tests. Based on this evidence, they expanded statewide—a decision justified by systematic evaluation across all four levels.

Why Sequential Evaluation Saves Money and Lives

This framework prevents the expensive mistakes I’ve seen across the humanitarian sector. Consider what happens when organizations skip early evaluation levels:

  • Without Level 1 evaluation, you deploy AI systems with unknown accuracy rates. The International Review of the Red Cross warns that biased AI systems can perpetuate discrimination against already vulnerable populations—discrimination that goes undetected without proper technical evaluation.
  • Without Level 2 evaluation, you build systems that technically work but nobody uses. This is particularly problematic in humanitarian contexts where digital literacy varies widely and user interface assumptions often prove incorrect.
  • Without Level 3 evaluation, you deploy systems that people use but that don’t change behavior in meaningful ways. Users might enjoy interacting with your AI health advisor while continuing harmful practices.

Following this sequential approach transforms abstract discussions about AI reliability into concrete, evidence-based decisions that can inform both program design and policy development.

Filed Under: Management
More About: , , ,

Written by
Wayan Vota co-founded ICTworks. He also co-founded Technology Salon, Career Pivot, MERL Tech, ICTforAg, ICT4Djobs, ICT4Drinks, JadedAid, Kurante, OLPC News and a few other things. Opinions expressed here are his own and do not reflect the position of his employer, any of its entities, or any ICTWorks sponsor.
Stay Current with ICTworksGet Regular Updates via Email

Leave a Reply

*

*