{"slug": "ai-governance-work-needs-much-better-monitoring", "title": "AI governance work needs much better monitoring", "summary": "A new report from Berlin-based think tank Future Matters finds that 13 of 22 major earning-to-give donors cannot tell whether AI governance work achieves anything, including four of six who work at frontier AI labs. The author argues that this 'illegibility' is a choice donors make by not asking the right questions, and calls for better monitoring, evaluation, and learning (MEL) in AI governance and other hard-to-measure areas.", "body_md": "*Linkpost for my Substack piece, adapted a reasonable amount for EA specifically. *\n\n**Almost all EA projects would benefit from better monitoring, but AI governance most of all, in my experience.**\n\nMost donors in EA never find out whether their grants worked. They model the impact before the money goes out, but don’t check if the models were accurate. *Evaluating impact* is hard, but *monitoring grant progress* is simple: agree indicators in writing before you send the money, ask the grantee to put a probability on each, and score them once a year.\n\nCharities usually send grant reports anyway. Unless you ask for something specific, they rarely contain the most useful information.\n\nA recent report to a client touted the project’s success: it had made 20 policy recommendations. When pressed, the grantee responded that only 20% had been implemented, even partially.\n\nA 20% implementation rate may or may not be a good outcome. Either way, it wasn’t the outcome they reported, or what they were asked to report by the donor.\n\nIronically, this is pervasive in ‘evidence-based giving’, and particularly in AI governance work.\n\nWhen it comes to individual donors, often nothing.\n\n‘Evidence-based giving’ typically applies before the grant. We are very good at modelling what a grant might achieve. We’re surprisingly bad at checking what it did.\n\nThis varies by sector. Frontline global health interventions routinely monitor outputs - clinic visits, vaccines distributed - and typically push further to monitor outcomes and impact, because the whole sector expects it.\n\nIn harder-to-measure areas, with no strong standard for self-monitoring, it’s the Wild West. I see policy, advocacy and governance interventions struggle the most.\n\nBerlin-based think tank Future Matters recently [found](https://future-matters.org/updates/legibility-in-the-coming-wave-anonymous-philanthropists-barriers-to-funding-ai-governance/) that 13 out of 22 major earning-to-give donors couldn’t tell whether AI governance work achieves anything, including four out of six who work at frontier AI labs.\n\nIf even the people closest to the technology aren’t sure what their grants are achieving, what hope do the rest of us have?\n\nFuture Matters blames legibility. I’d go further: **illegibility is a choice donors make, by not asking the right questions at the right time.**\n\nThe effective giving space currently inverts best practice: monitoring is often most rigorous where uncertainty is lowest (bednets), and often absent where it’s highest (policy and advocacy).\n\nMonitoring, evaluation and learning (MEL) is how we understand what is working, to improve both projects and funders.\n\nA useful MEL plan starts with its purpose, and what questions it seeks to answer.\n\nAs a project, you might want to know whether your services are effective and reach the right people; how to improve them; or what you can use to fundraise.\n\nAs a funder, you have a stake in whether your individual grantees are succeeding. However, there are many other benefits to you: improving your funding decisions; [managing risk](https://fundinganthropalypse.com/p/most-donors-get-risk-wrong?r=4b7xoz) across your portfolio; and making a meaningful difference to the problem (even more than a single project, as you hold a variety of bets). You might also want to share lessons with the wider ecosystem or even build the field.\n\nDefining which of these questions matters focuses our activities and makes our MEL plans useful.\n\nMEL is the *only way* to answer these questions systematically. Without it, we risk wasting funding and time. In the absence of evidence, projects can go astray and donors can [withhold grants](https://future-matters.org/updates/legibility-in-the-coming-wave-anonymous-philanthropists-barriers-to-funding-ai-governance/#:~:text=donors%20are%20not%20confident%20they%20will%20give%20to%20policy%20advocates%20in%20AI%20gov%C2%ADer%C2%ADnance). **Consequently, we will** **struggle to solve the life-or-death problems we are working on.**\n\nMonitoring is the ongoing tracking of execution at the programme level. Evaluation is the wider assessment of what was achieved.\n\nAlthough this post focuses on monitoring, because evaluation is much more resource-intensive and difficult, effective giving is also surprisingly weak at evaluating its grants. Even GiveWell, the field-leader in terms of rigour and transparency, has conducted very few retrospective evaluations, only starting this process in [2025](https://blog.givewell.org/2025/07/09/what-weve-learned-from-our-first-lookbacks/).\n\nInterestingly, when they did, they found that some grants had performed much better than modelled and some had performed worse.\n\nI think this is notable in a field that traditionally looks for RCTs from its grantees. We very rarely apply the same rigourous measurement to our own grantmaking.\n\nIf we do more *post hoc* evaluation in future, it is likely that we can check how well calibrated our models were and improve them over time. I won't dwell on this in this post but would like to see more of this work, especially outside global health, which tends to be stronger on all aspects of MEL than other cause areas.\n\nOne motivation is noble: not overburdening the project. This has been a positive shift away from intensive, performative reporting that almost nobody read anyway.\n\nThis change has been driven by two very different camps: cost-effective giving at one end and [trust-based philanthropy](https://www.nptrust.org/philanthropic-resources/philanthropist/trust-based-philanthropy-a-primer-for-donors/) at the other.\n\nThe problem is that each risks having the right diagnosis and the wrong cure. **The fact that traditional reporting is** **too burdensome and not useful doesn’t mean that you shouldn’t ask anything at all.**\n\nAny strong organisation should actively want to know whether its work is effective. If it is happy to run purely on anecdotes and intuition, that’s a red flag.\n\nA second reason to forgo monitoring and evaluation is the belief that some interventions are too hard to measure.\n\nThis is half right. *Evaluating the impact* of policy, advocacy and governance work is extremely difficult. You can’t usually run RCTs, so counterfactuals are tricky. Results can take years to arrive. Several groups will usually work on the same issue, without a clean way to assign the credit. As Future Matters points out, [publicly claiming attribution can burn](https://future-matters.org/updates/legibility-in-the-coming-wave-anonymous-philanthropists-barriers-to-funding-ai-governance/#article__heading_2:~:text=A%20great%20deal,and%20the%20official.) your most important relationships.\n\nThis is particularly acute in AI safety and governance, where it isn’t even obvious what makes a good regulation. If reasonable people disagree about whether slowing down AI development is net positive, no monitoring and evaluation plan will solve that. It’s just the risk you take by funding this work.\n\nThere are nonetheless proxy indicators of effectiveness - certainly enough to distinguish organisations from one another in the short-term. For example, bills sometimes use language drawn directly from a project’s publications; lawmakers sometimes acknowledge conversations with particular groups; and it can be surprisingly effective to ring congressional staff and ask how influential an advocacy project was.\n\nIn addition, *monitoring* policy work is really not that complex, even when impact measurement is hard. Any MEL professional would find it routine: ask an organisation what it plans to do and how likely it is to succeed, and then track whether it did.\n\nThe best projects I have assessed could point me clearly to language in a bill that passed into law, and show how it was derived or quoted from their work, which is a pretty serious achievement. T**hese organisations should be rewarded with more funding** **and support.**\n\nA strong monitoring plan includes a handful of indicators, with the project assigning each a probability of success. These can be scored once a year, with any failures left in. It should take the project around half a day to complete and track things it wants to know anyway.\n\nMy advisory, Ultra Philanthropy, has an [in-house MEL specialist](https://ultraphilanthropy.org/team#:~:text=Olivia%20Kaye%20is%20an%20established%20research%2C%20monitoring%20and%20evaluation%20consultant%2C%20with%2012%20years%27%20experience%20advising%20complex%20development%20programmes%20in%20challenging%20environments) with 12 years’ experience advising development programmes. We have worked successfully with some admirably responsive, transparent organisations to produce these sorts of plans, even in thorny areas of advocacy.\n\nFor example, we worked with a climate advocacy group in 2025 to define the following:\n\nIt would have been very easy for the project to focus on controllable activities (‘we’ll have 50 meetings with lawmakers’), or to claim that their work was powerful magic that defied tracking (a depressingly common claim, which I have universally found to be false).\n\nInstead, its plan gave our client (and the organisation itself) a clear, measurable monitoring framework. Even though the impact of these legislative achievements remains uncertain (laws still have to be enforced, might have unintended consequences, take a long time to have an effect, etc.), it’s easy to see whether or not they moved the ball down the field.\n\nUnfortunately, this clarity is extremely rare in the advocacy and policy work that I encounter, especially in AI safety and governance.\n\nAn aspect of our monitoring that goes beyond the norm is assigning probabilities to each output.\n\nOne benefit is that this exposes overconfident fundraisers. Often, fundraising teams send donors several impressive milestones that a grant could enable. Ask their programme team to put a success percentage against each, though, and it’s clear how likely you are to see them.\n\nIt also prevents [sandbagging](https://www.investopedia.com/terms/s/sandbag.asp). If a project suggests indicators that it has a 95% chance of achieving, we look more closely at its ambition and risk-tolerance, and check whether those outcomes are actually meaningful.\n\nSometimes success is both highly likely and meaningful - ‘we’re 85% confident of reaching 10,000 people with oral rehydration solution this year’. But this can also be playing it safe, or not focusing on the most important changes.\n\nWielded together, these opposing concerns create something exciting: a portfolio that blends solid gains with stretching, ambitious goals, which tell the donor and the project how close they are to getting the hard thing done.\n\nDonors then need to focus on calibration, not merely success rate. An ambitious miss is usually more useful than a 100% hit rate. It’s essential if we are going to [embrace more risk](https://fundinganthropalypse.com/p/most-donors-get-risk-wrong?r=4b7xoz) in our giving.\n\nFor most large donors (>$1m/year), it makes sense to engage a MEL consultant or an advisory with in-house expertise to do this work, since it usually involves negotiating with grantees, and a reasonable amount of data collection and analysis across a portfolio.\n\nFor donors wanting to implement this themselves, however, there are some things to watch out for:\n\nEstablishing a consistent process to monitor your grants and learn from them is essential if you are about to become a significant philanthropist.\n\nThe [Funding Anthropalypse](https://fundinganthropalypse.com/p/what-is-the-funding-anthropalypse?r=4b7xoz) will create many large donors who have little or no experience in grantmaking or MEL. This makes feedback loops a high priority, so that you can improve your grants, [risk-tolerance](https://fundinganthropalypse.com/p/most-donors-get-risk-wrong?r=4b7xoz) and general understanding over time.\n\nYou may also give a significant amount to AI safety and governance work, which is much more uncertain and harder to measure than many cause areas, with greater risk of accidental harm. This makes effective monitoring even more important.\n\nUse the pointers above to shape your monitoring strategy, and talk to a professional if you can. You can also defer to evaluators, advisories or funds you trust to do this work for you.\n\n**Illegibility is a** **choice.** Make a different one.\n\n*Jack Lewars is the founder of Ultra Philanthropy, an independent advisory that helps major donors give for maximum impact, and is the fund manager of its Mid-Stage Global Health Fund. He advises donors giving up to nine figures a year, and is Chair of Trustees at High Impact Athletes. **Talk to him about your giving**.*\n\n*Thanks to Olivia Kaye and Samuel Verbi for feedback on the draft. I used Claude to help structure my thoughts and to suggest improvements and flag gaps, as well as for proofreading; all views, primary drafting and final edits are mine.*", "url": "https://wpnews.pro/news/ai-governance-work-needs-much-better-monitoring", "canonical_source": "https://www.lesswrong.com/posts/EGh9qkCH94SGG6Kvw/ai-governance-work-needs-much-better-monitoring", "published_at": "2026-08-11 17:42:45+00:00", "updated_at": "2026-08-11 18:10:48.345451+00:00", "lang": "en", "topics": ["ai-policy", "ai-ethics"], "entities": ["Future Matters", "AI governance"], "alternates": {"html": "https://wpnews.pro/news/ai-governance-work-needs-much-better-monitoring", "markdown": "https://wpnews.pro/news/ai-governance-work-needs-much-better-monitoring.md", "text": "https://wpnews.pro/news/ai-governance-work-needs-much-better-monitoring.txt", "jsonld": "https://wpnews.pro/news/ai-governance-work-needs-much-better-monitoring.jsonld"}}