Skip to content
News

To Thrive in a Landscape of Swans and Rhinos, Evaluators and Funders Must Work Together

By Kate Callahan

When I read Steffen Bohni Nielsen and Sebastian Lemire’s recent article, “What Happens When a Black Swan Meets a Gray Rhino?,” I recognized RFA on almost every page.

The authors name what evaluators have been feeling for more than a year, but haven’t had a clean way to talk about. The federal funding collapse under DOGE was a Black Swan: sudden, disorienting, impossible to plan for. Generative AI is a Gray Rhino: visible for a decade, chronically under addressed, now impossible to ignore.

Most of us in the evaluation field are experiencing both dynamics at once.

Like many organizations our size, Research for Action (RFA)  grew for over a decade on the strength of federal evaluation contracts. But as we all know, that well of funding has narrowed considerably. What’s more, larger national firms, having lost federal projects that once anchored them, are moving into local, place-based evaluation alongside smaller firms.

We are all fighting for the same shrinking pool.

Given this new environment, RFA, like many of our smaller colleagues, is spending enormous energy competing to hold our ground—energy that can go toward addressing a tough question raised in the article: what does generative AI mean for how we work, who we train, and what we’re worth to the change agents who partner with us?

That question isn’t conceptual for RFA and our field colleagues. It’s existential: Evaluators that apply the strongest balance of human and AI contributions will be the ones that survive.

Any field grappling with AI has the same responsibility: understanding what is uniquely human about its work, and naming that clearly. For research and evaluation, that work shows up in design, not just execution. For instance, someone still must decide which questions are worth asking, and whose experiences should shape those questions. Those decisions require positionality, an honest accounting of whose perspective is centered and whose is missing, and it requires judgment about what counts as evidence, how much weight the evidence deserves, and what it means in context.

Relationships remain essential, as well. Knowing whether a finding from one place will hold up somewhere else, what researchers call external validity, depends on understanding social, political, and cultural contexts—an understanding that comes from connection to the people most affected, not from data alone.

What excites me about RFA’s incorporation of AI is how it can shorten the distance between research and practice. We’ve built AI into tools for transcription and data visualization; work that used to take days now happens in hours. As an applied research organization, we’ve always wanted our strategies to be useful. AI is taking that commitment to a level we couldn’t reach before. A finding that once lived in a static report, written for one audience, can now reach a school board, a parents group, and state legislators with language that make sense to them, and in a format they’ll actually use—such as a short video, a one-page brief, or an interactive dashboard.

At the same time, we are careful about which AI tools we use and what data goes into them, especially when the people we work with could be harmed if their information—such as a student’s disciplinary or special education record, a family’s immigration status, or a health condition disclosed in an interview—becomes exposed. Knowing when not to use AI is part of the same judgment this work has always required.

Amid this challenging environment, what role can foundations play as partners with evaluators?

Not surprisingly, foundations are reasonably cautious about investing in an evaluation sector that has not answered a critical question raised by the article: what can we offer that AI cannot? I’d argue that this caution is not only about AI, as it also reflects years of disappointment in the impact of evaluation and research, funding a great deal of rigorous work that produced high-quality findings and little change.

I also believe this caution reflects a history in which evaluation was too often done to marginalized communities rather than conducted with them, treating people as subjects of measurement rather than as partners in defining what mattered and what should happen next. But I also would emphasize that evaluators cannot fully answer the AI question, or repair that older history, without the time, training, and general operating support that only a real investment from foundations can provide.

Someone must move first. Evaluators understand, in tangible and substantive ways, what judgment, positionality, and context bring to their work. As I noted earlier, evaluators who understand the interplay and relationship between these human factors and AI will be the ones that make it. Once we demonstrate this relationship, funding the retraining and infrastructure to support AI becomes a far easier decision.

I’d like to hear from other evaluators, funders, and researchers about what this moment looks like from where you stand. I also welcome the chance to talk with you about this piece, or about where you see the philanthropic and evaluation sectors headed.

As evaluators, we’re all dealing with the black swan and gray rhino in different ways. We might not be able to tame them, but we can collectively figure out a new reality that reinforces the enduring value of our work.

Kate Callahan is Executive Director of Research for Action