Beyond the Spreadsheet: Why the Numbers Foundations Count Are Often the Wrong Ones
The Report That Cannot Be Filed
Maria Delgado has worked in community health outreach in South Texas for eleven years. She knows her neighborhoods. She knows which families have not seen a doctor in four years and which grandmothers are quietly managing diabetes without medication. She knows when something is shifting—when trust is building, when a block is beginning to function differently, when the invisible architecture of a community is being repaired.
What she cannot do, she says, is put any of that in a grant report.
"They want numbers," she explains. "How many people did you serve? How many referrals did you make? How many participants completed the program? Those things I can count. But the reason this neighborhood is different than it was three years ago—I can't put that in a spreadsheet. And so the funder never really knows what their money did."
Delgado's frustration is widely shared among nonprofit practitioners across the United States. As foundations have invested in increasingly rigorous measurement frameworks over the past two decades, a quiet crisis has developed at the intersection of accountability and understanding. The tools built to demonstrate impact have, in many cases, become obstacles to comprehending it.
The Metrics Philanthropy Built—and What They Miss
The shift toward quantitative evaluation in American philanthropy was not arbitrary. It emerged from legitimate concerns about accountability, comparability, and the responsible stewardship of charitable resources. Logic models, theory-of-change frameworks, and outcome-tracking dashboards were developed to give foundations confidence that their investments were producing measurable results.
The problem is not measurement itself. The problem is the assumption, embedded in most current evaluation practice, that the most important dimensions of community change are the ones most amenable to counting.
Consider what standard grant metrics typically capture: number of individuals served, number of program completions, pre- and post-survey scores on specific indicators, percentage of participants who meet defined benchmarks. These figures are not meaningless. They provide a useful baseline of organizational activity. But they are systematically blind to several of the most consequential forms of social transformation.
They do not capture shifts in community narrative—the moment when a neighborhood that has been defined by its problems begins to define itself by its assets. They do not capture relational density—the thickening web of trust and mutual obligation that makes communities resilient in ways that no intervention can directly produce. They do not capture the counterfactual—what would have happened to a family, a block, a town in the absence of a particular investment. And they almost never capture failure in useful ways, because grant reporting structures create powerful incentives for nonprofits to present the most favorable interpretation of available data.
What Frontline Workers Actually See
Ask the people doing community development work what signals genuine progress, and the answers are consistently qualitative, relational, and contextual.
A youth program director in Detroit describes the moment she knew her organization's work was taking hold: not when graduation rates improved, but when parents who had never spoken to each other started meeting at the building on Saturday mornings, without any program prompting them. "That wasn't in our outcomes framework," she says. "But it was the most important thing that happened that year."
A housing counselor in rural Appalachia points to a different kind of indicator: the decrease in the number of calls she receives from clients in crisis. "When things are working, people don't call me as much because they're calling each other. Their neighbors. Their family. That's what stability looks like. But in my grant report, fewer calls just looks like lower engagement."
A literacy coordinator in rural Mississippi describes how she evaluates her own program's success: not by reading scores alone, but by watching whether the parents of her students begin attending community meetings, writing letters to local officials, or speaking up in conversations they previously avoided. "Reading is the door," she says. "What I'm actually building is agency. And there's no rubric for that."
These observations are not anecdotal noise. They point toward a category of change—what researchers sometimes call second-order effects—that standard metrics are structurally unable to detect.
The Case for Narrative Evaluation
A growing body of evaluation theory argues that the solution is not to abandon quantitative measurement but to supplement it with rigorously collected qualitative evidence. Narrative evaluation, developmental evaluation, and most-significant-change methodologies all offer frameworks for capturing the human texture of social transformation without sacrificing analytical discipline.
The most-significant-change approach, for example, asks program participants and community members to identify and describe the most meaningful change they have experienced as a result of a program or initiative. These stories are then reviewed and discussed by stakeholders at multiple levels of an organization, creating a deliberate process for surfacing the kinds of change that numbers miss while maintaining a structured approach to analysis.
Developmental evaluation, pioneered by evaluation theorist Michael Quinn Patton, is designed specifically for complex social interventions operating in dynamic environments—precisely the conditions that characterize most community development work. Rather than measuring outcomes against predetermined indicators, it tracks how programs evolve, what they learn, and how they adapt, treating the process of change as itself a meaningful object of inquiry.
What these approaches share is a fundamental reorientation: from asking whether a program achieved what it predicted to asking what actually happened and why it matters.
What Foundations Would Have to Give Up
Adopting more qualitative, narrative-centered evaluation practices is not costless for foundations. It requires giving up something that quantitative metrics provide in abundance: the ability to compare organizations against each other on standardized scales.
The appeal of a common measurement framework is understandable. If every grantee reports using the same indicators, program officers can theoretically assess relative performance, identify outliers, and make allocation decisions based on comparative data. This logic drives much of the current enthusiasm for collective impact frameworks and sector-wide measurement initiatives.
But the comparison it enables is often illusory. A youth program serving young people experiencing homelessness in Los Angeles and a youth program serving students in suburban Ohio are not meaningfully comparable on a standardized rubric, regardless of how precisely that rubric is constructed. The contexts are different. The populations are different. The definitions of success are different. And the act of forcing both programs into the same measurement container distorts both.
Foundations willing to release the comfort of false comparability gain something more valuable in return: an honest understanding of what their investments are actually producing in the specific communities where they work.
Toward a Fuller Accounting
At Lunt Foundations, the belief that communities are complex, living systems—not programs to be optimized—shapes how we think about evidence and evaluation. Numbers have their place. They tell part of a story. But the stories themselves—told by the people who live them, in language that captures their full texture—are the most reliable evidence we have that something real is changing.
Maria Delgado's neighborhood in South Texas is different than it was. She knows it. The families there know it. The challenge for philanthropy is to build the capacity to know it too—not despite the messiness of qualitative evidence, but because of the honesty it demands.
Measuring what matters begins with being willing to ask different questions. The answers, it turns out, have been available all along.