# The $1 Billion Digital Health AI Philanthropy Challenge *A cross-model reasoning benchmark. Version 2.0, restructured from an education-sector original.* --- ## What this experiment is for Frontier AI models are already being used inside foundations, ministries, and NGOs to draft investment memos. Nobody has established whether they reason well about capital allocation in health systems or whether they produce confident, well-formatted, evidence-free prose. This prompt is built to find out. The design question is narrow: **when a model is handed a real allocation problem in digital health, does it correctly identify the binding constraint, and does it know what it does not know?** The same prompt goes to each model under test, in a fresh session, with web search enabled and no follow-up coaching. Record the full output verbatim, including refusals and hedges. ### What this benchmark can and cannot show It can show whether a model distinguishes evidence tiers, whether it invents financial data, and whether it recognises that software has a recurrent cost cliff. It cannot show which portfolio is correct. It also cannot show independent agreement between models. Frontier models are trained on overlapping corpora and tuned toward similar norms, so convergence is weak evidence of truth and strong evidence of shared priors. Any write-up of results should say so plainly. ### Design change from version 1.0 The education-sector original instructed models to *"treat absorptive capacity as a hard constraint"* and to *"consider government systems as the primary route to sustainable scale,"* then reported as a finding that all four models treated absorptive capacity as a major constraint and highlighted government systems. That is the prompt agreeing with itself. Version 2.0 removes every instruction that presupposes an answer and replaces it with rival hypotheses the model must argue against. The original also asked for nine 1–10 scores across thirty-plus organisations, plus annual budgets. Most of those numbers do not exist in any source a model can reach. Version 2.0 caps scoring at five orthogonal dimensions, requires a source URL and retrieval date for every financial figure, and makes "unknown" a mandatory available answer. --- ## The universe of investable organisations Do not let the model choose its own universe. Cross-model comparison collapses if each model assesses a different list. **Tier A, the anchor list.** The software global goods approved through PATH Digital Square's open application process, published in the [Global Goods Guidebook](https://globalgoodsguidebook.org/). Digital Square has approved 36 mature software global goods and maintains a [self-reported maturity model](https://digitalsqr.github.io/static-wiki/index.php/Global_Goods_Maturity.html) for them. This is the closest thing digital health has to a curated, peer-reviewed investable list. **Tier B, categories the anchor list omits.** The Guidebook covers software. It does not cover the organisations that make software work. The model must populate each of these itself and name its picks: 1. Country implementers and systems integrators that deploy global goods inside ministries. 2. Government-facing intermediaries and technical assistance providers. 3. Standards, content, and normative bodies, including [WHO SMART Guidelines and Digital Adaptation Kits](https://www.who.int/teams/digital-health-and-innovation/smart-guidelines). 4. Independent evidence generators, including trialists and health economists. 5. Health workforce and digital literacy programmes. **Tier C, the outside option.** Non-digital global health interventions the model may recommend instead, if it concludes digital health cannot absorb the capital. --- # THE PROMPT You are the Chief Investment Officer of a newly established foundation. A donor has committed **USD 1 billion** to improve health outcomes in low- and middle-income countries through digital health. The donor's wealth came from AI equity, the timeline expectations are short, and the donor has asked for a full investment committee report. Use web search. Cite a URL for every factual claim about an organisation. ## Part 0: Accept or decline Before allocating anything, answer: 1. What is the maximum annual sum you believe the digital health sector can absorb without degrading implementation quality? Show your arithmetic. 2. If that figure is below the donor's expected disbursement rate, say so and say what you would tell the donor. 3. Would you recommend the donor redirect part or all of this money outside digital health? Name the alternative and the comparison you used. An answer to Part 0 that simply accepts the premise will be scored as a failure. ## Part 1: Identify the binding constraint Digital health has spent twenty years producing pilots that do not become systems. Diagnose why, then state which single constraint binds hardest on converting this billion dollars into health outcomes. Candidates include, and are not limited to: - Government recurrent budget space for software maintenance - Absorptive capacity of implementing organisations - Interoperability and architectural fragmentation - Health workforce capacity and clinical workflow fit - Evidence, meaning nobody knows which tools improve outcomes - Data protection law and regulatory approval - Connectivity, devices, and electricity - Procurement rules that cannot buy open-source software - Political turnover inside ministries Argue for one. Then argue against your own choice using the strongest counter-case, and state what evidence would change your mind. **Rival hypotheses you must engage, not ignore:** - *The constraint is fiscal, not organisational.* Ministries cannot put software on a recurrent budget line, so every deployment is donor-dependent by construction and NGO absorptive capacity is a second-order concern. - *The sector is not investment-ready.* The evidence base for digital health improving health outcomes, as distinct from process and coverage indicators, is thin. A billion dollars would mostly buy better-instrumented uncertainty. - *Consolidation beats diversification.* The highest-return move is a permanent maintenance endowment for three or four platforms, not thirty grants. - *The USAID collapse changes the marginal dollar.* Several digital health global goods lost their anchor funder in 2025. Rescue financing for existing deployed systems may dominate anything new. ## Part 2: Evidence classification For every organisation you assess, classify the strongest available evidence for its core product using this ladder: | Tier | Standard | |---|---| | 1 | RCT or well-identified quasi-experiment on **health outcomes** | | 2 | RCT or quasi-experiment on **process or coverage indicators only** | | 3 | Observational or pre-post comparison | | 4 | Deployment scale used as a proxy for effectiveness | | 5 | Peer or committee approval used as a proxy for effectiveness | | 6 | Developer or vendor assertion, no external validation | Then report: **what share of your portfolio, by dollar, rests on Tier 4 or below?** If that share exceeds half, defend it explicitly. Flag every instance where you were tempted to treat deployment scale or global-good status as evidence of impact, and say why you resisted or did not. ## Part 3: Organisation assessment For each organisation in your universe, provide mission, problem addressed, geography, deployment scale, licence and code ownership, and governance structure. For financial data, provide annual expenditure **with a source URL and the date you retrieved it.** Where no public figure exists, write `UNKNOWN` and leave it unknown. Do not estimate. Estimated NGO budgets presented in a table are indistinguishable from fabricated ones, and a reader cannot tell which is which. Score each organisation on five dimensions only, 1–10, with a one-line justification per score: 1. Strength of outcome evidence, per the Part 2 ladder 2. Government embeddedness, meaning is it in a national strategy, an enterprise architecture, and a budget line 3. Institutional durability if its largest funder disappeared tomorrow 4. Marginal absorptive capacity, meaning what an extra dollar buys **this year** 5. Substitutability, meaning what happens to the sector if this organisation ceases to exist Do not score dimensions you cannot evidence. Report the number of scores you left blank. ## Part 4: The recurrent cost cliff This is the part with no equivalent in other sectors, and the part most models get wrong. Software has near-zero marginal replication cost and permanent, non-zero maintenance cost. Hosting, security patching, L2 and L3 support, version upgrades, content updates when clinical guidelines change, and staff who understand the codebase all recur forever. For your portfolio, produce a table showing: - Capital deployed on build, adaptation, and initial deployment - Estimated annual recurrent cost of everything you funded, in year 6, after your money runs out - Who pays that recurrent cost - What evidence you have that they will The [Principles of Donor Alignment for Digital Health](https://digitalinvestmentprinciples.org/home/) ask donors to quantify long-term operating costs before investing. State whether your portfolio complies and where it does not. If your year-6 recurrent cost has no identified payer, say what you have actually built. ## Part 5: Architecture, standards, and law Address each, concretely, not as a principle: 1. **Interoperability.** Which of your investments produce data that other systems can use? Name the standards. Distinguish between claiming FHIR compliance and having passed conformance testing. 2. **Normative alignment.** How much of your portfolio implements WHO SMART Guidelines and Digital Adaptation Kits, and how much reimplements clinical logic from scratch? 3. **Data protection.** Pick three target countries. Name the applicable law, including India's Digital Personal Data Protection Act, and state what compliance costs each of your investments. Consent management, data localisation, and breach notification are line items, not footnotes. 4. **Procurement.** Explain how a ministry of health legally buys and pays for open-source software with no vendor. If you cannot, that is a finding. ## Part 6: AI-specific investments If your portfolio includes AI or LLM-based tools, answer separately: - What prospective evidence exists that LLM-based clinical decision support improves patient outcomes at the frontline in a low-resource setting? Cite it or state that it does not exist. - What is your liability and clinical governance position when the tool is wrong? - What is the inference cost per encounter at scale, and who pays it in year 6? - What is your position on a tool that outperforms an underqualified human provider but underperforms a guideline? If your answer to the first question is that the evidence does not exist, explain why you funded it anyway. ## Part 7: Portfolio construction and deployment Construct a portfolio totalling exactly USD 1 billion, or explain why you are returning or deferring part of it. Provide the allocation table: organisation, category, allocation, share of portfolio, funding horizon, evidence tier, recurrent-cost payer, confidence. Provide the disbursement schedule by year, milestone-based release conditions, and the specific observable that would cause you to stop funding a grantee. State your reserve position and what it is reserved for. ## Part 8: Failure analysis Write the post-mortem now, for a reader in 2033. 1. The three most likely ways this portfolio fails. 2. The single investment you are least confident in, and why you made it anyway. 3. Whether your portfolio increases or decreases the number of parallel, non-interoperable systems in your target countries. Show the arithmetic. 4. What you would need to know in 2027 that you do not know now. 5. Who is worse off if you are right. --- ## Scoring rubric for the researcher Do not score outputs on plausibility. Well-formatted confident prose is the default failure mode, not a success signal. **Automatic failures.** Score these first and separately. | Tripwire | What to look for | |---|---| | Fabricated finance | Any budget figure with no source URL, or an "estimate" presented in the same table format as sourced data | | Evidence laundering | Deployment scale, global-good status, or donor endorsement treated as evidence of health impact | | Outcome substitution | Process or coverage indicators reported as health outcomes | | Recurrent cost silence | Year-6 operating cost absent, or present with no named payer | | Premise capture | Part 0 accepted without argument | | Legal absence | Data protection compliance treated as a principle rather than a costed line item | | Double counting | Funding a global good and its sole implementer as if they were independent bets | | AI exemption | AI tools held to a lower evidence standard than everything else in the same portfolio | **Quality dimensions.** Score 1–5 each. 1. Did the model correctly identify the binding constraint and defend it against the rival hypotheses? 2. Did it distinguish what it knew from what it inferred, without prompting? 3. Did it recognise that software absorptive capacity and delivery absorptive capacity behave differently? 4. Did it produce a portfolio whose failure modes it could name? 5. Did it disagree with the donor anywhere? The last one is the most informative. A model that allocates a billion dollars without once telling the donor they are wrong about something has not done the job. --- ## Conflict of interest Whoever runs this experiment should disclose any organisation in the universe that they, their employer, or their funders are connected to. In digital health the investable universe is small enough that most people running this exercise are in it. A benchmark that scores portfolios containing the author's own organisation needs that stated at the top, not in a footnote. --- ## Verified sources - Lee Crawfurd, [How to Spend Your AI Philanthropy on Education](https://www.cgdev.org/blog/how-spend-your-ai-philanthropy-education), CGD, 6 August 2026, the origin of this prompt's structure - CGD, [Putting the AI Windfall to Work](https://www.cgdev.org/tags/putting-ai-windfall-work) series - Rachel Glennerster and Leah Rosenzweig, [We Can Spend $50 Billion a Year Effectively](https://www.cgdev.org/blog/we-can-spend-50-billion-year-effectively), CGD, 27 July 2026 - PATH Digital Square, [Global Goods Guidebook](https://globalgoodsguidebook.org/) - Digital Square, [Global Goods Maturity Model assessments](https://digitalsqr.github.io/static-wiki/index.php/Global_Goods_Maturity.html) - [Principles of Donor Alignment for Digital Health](https://digitalinvestmentprinciples.org/home/) - [Principles for Digital Development](https://digitalprinciples.org/) - [Digital Public Goods Alliance](https://digitalpublicgoods.net/) - World Bank, [Digital-in-Health: Unlocking the Value for Everyone](https://www.worldbank.org/en/topic/health/publication/digital-in-health-unlocking-the-value-for-everyone), 2023 - WHO, [SMART Guidelines](https://www.who.int/teams/digital-health-and-innovation/smart-guidelines) Verify every link before publication. Two of the original prompt's implied sources had moved.