
Billions in new philanthropic AI wealth are heading toward international development. I wanted to know what the frontier AI models had to say about digital health absorbing that money.
I wrote a test and handed the identical prompt to OpenAI’s GPT, Anthropic’s Claude, and Google’s Gemini: you are the chief investment officer of a new foundation, a donor has given you $1 billion for digital health in low- and middle-income countries, and they want it deployed fast.
The sector usually answers this question by asking for more coordination.
PATH’s Digital Square argued in its March 2025 global goods ecosystem report that better market shaping and pooled financing would fix the funding problem, and the 30 donors who endorsed the Principles of Donor Alignment for Digital Health have been asking each other since 2018 to quantify long-term operating costs before investing.
I built the scoring rubric to catch the opposite behavior: fabricated finance, deployment scale passed off as health impact, year-six operating costs with no named payer, and absence claims made without naming the searches.
Sign Up Now for more digital health funding analysis
All Three Refused the Donor’s Timeline
Three portfolios, one shared conclusion: the sector cannot take $200 to $333 million a year.
- ChatGPT declined 75% of the mandate, staging $250 million into digital health and sending $500 million to Gavi.
- Claude accepted the billion and redirected $400 million to Against Malaria Foundation, Helen Keller Intl, and Malaria Consortium.
- Gemini stretched the window to seven years and capped year one at $145 million.
That consensus is worth exactly as much as the evidence underneath it, which is where the submissions stop resembling each other.
Gemini’s Report Contains Zero URLs
I ran a character match across all three documents. GPT cited 34 sourced values in a numeric ledger. Claude cited 44 unique URLs. Gemini cited none. Gemini produced 5,390 words, six organizational scorecards, and nine numbered tables without a single link to anything.
- Base software steward capacity of $42 million a year.
- OpenMRS core budget of $2.8 million.
- Non-digital cost-effectiveness of $27 per DALY averted.
Each is tagged ESTIMATE with a range, which reads like discipline until you notice there is nothing on the other end of the tag. It gets worse under checking.
Gemini’s sixth in-depth organizational assessment is “Swinnovations / Frontier AI-CDS,” scored across five dimensions and anchoring a $100 million portfolio line. That’s a total hallucination. Gemini also triggered three automatic failures in the rubric: fabricated finance, unsearched absence, and format compliance without content.
I am continuously surprised that the company who literally wrote the LLM defining paper, cannot field a decent LLM.
Claude Lost on a Paper It Never Found
Claude wrote the sharpest document. It identified the binding constraint as the disappearance of the payer for the platform layer, backed by live evidence: the HISP Centre launched the DHIS2 Shared Services Fee on 20 May 2026, a voluntary annual contribution, because donors who once funded core platform work moved their money to country systems.
It pulled Kenya’s own health sector budget report showing absorption rates of 90%, 87%, and 84% across three years, then observed that a ministry which cannot spend what it holds does not gain capacity when a foundation sends more. It caught itself laundering donor endorsement as government embeddedness and said so unprompted. It told the donor the test itself was defective.
Then it proposed $20 million to fund a randomized trial of LLM clinical decision support at Penda Health in Nairobi, calling Penda “the strongest platform in the world for that trial,” and filed an absence claim that no prospective study had measured LLM decision support against patient health outcomes in a low-resource setting.
Penda Health tested LLM decision support across the same 16 clinics. Treatment failure occurred in 2.2% of the AI-assisted arm and 2.0% of the control arm, with an adjusted odds ratio of 0.77 and a P value of 0.13.
GPT found it and built its entire AI position on the null result. Claude named its search queries, listed its sites, and missed a Nature Medicine paper involving the exact operator it recommended. Naming your searches does not make the absence claim true.
Claude also let two portfolio numbers contradict each other. Its cover recommendation says $385 million to digital health. Its Part 7 table totals $430 million. Its headline evidence exposure, 71% at Tier 4 or below, uses the smaller denominator. The document flags the discrepancy in a footnote and never resolves it.
GPT Won on Arithmetic
I checked GPT’s numbers against the primary sources. Its HISP figure runs $20.174 million in 2024 funding times 74% for the core platform. I pulled the HISP Centre annual report PDF and found both values on the same page. Every calculation in its arithmetic appendix reproduces. Every absence claim carries a search log with queries and sites.
The more telling move is what GPT refused to compute.
It took Gavi’s stated 2026 to 2030 target of $11.9 billion and 8 million lives saved, divided them, and then declined to apply the resulting $1,487.50 per life to its own $500 million allocation, because the Gavi figure is a prospective model and marginal dollars are not average dollars.
Claude ran the same division on GiveWell’s $5,500 per life and converted $400 million into 72,727 deaths averted. That is the linear-marginality error, caveated but committed.
GPT also put no AI deployment money in the portfolio at all, capping AI at $5 million inside a $20 million evidence line. On a rubric with an explicit AI exemption failure, that is the cleanest pass available.
Where the Winner Is Still Wrong
GPT’s year-six recurrent cost is UNKNOWN on every single line. It calls that a stop condition instead of an omission, and it names an intended payer and an evidence-of-payer column for each. A strict reading of the rubric’s fourth automatic failure catches it anyway.
Claude did the harder thing here: it produced $88 million, labeled the 20% to 25% recurrent-to-capital ratio as unsourced, and then said plainly that on $39 million a year it had built a six-year subsidy for software that will need the same subsidy again in 2032 from a donor it cannot name.
GPT also funds PATH Digital Square $10 million while funding five of the global goods Digital Square approved, which is close to the rubric’s double-counting failure. And it sends the donor to interview Ambrose Agweyu as lead author of the Kenya trial. Agweyu led a companion safety paper in Nature Health. Bilal Mateen at PATH is the corresponding author on the trial itself.
None of the three could say what a marginal digital health dollar buys in health.
GPT wrote that if the board requires a demonstrated cost-per-life-saved advantage before investing, the correct digital allocation today is zero. Claude wrote that on the only metric either party could compute, doing nothing for five years beats its own portfolio.
Lessons for Incoming Money
What separated these three had nothing to do with digital health. It was whether the model would show a number’s provenance when nobody was checking. Two did. One produced an investment memo that dissolves on contact with a search engine, and it was the fastest to read and the easiest to believe.
Digital health has a version of the same problem.
Five of the six largest organizations in this sector publish no product-level expenditure, so no funder can answer what an additional dollar buys this year. HISP asking its users to pass a hat in May 2026 is what that looks like from the inside.
The Gates Foundation and Wellcome are already paying for the missing evidence, with $60 million for AI decision support evaluations led by organizations registered in Africa and Asia. That is roughly one twentieth of what a single AI philanthropy could distribute in a year under a 5% payout rule.
What to Do Now?
Digital health builders, please publish your product-level annual expenditure before the money arrives and someone asks. And when a model hands you a table with no source lines, read the source lines first and the argument never.
Finally, if you want to apply for open AI-for-good funding, try ImpactOpen, which tracks new philanthropic AI funding opportunities – 20 of them open right now. No login, no paywall, and the full dataset downloads as a CSV.

