About this research. This page is an accessible companion to my doctoral thesis, which examined IMM's development between 2011 and 2021 based on 106 semi-structured interviews conducted in 2020 with investors, standard-setters, intermediaries, and practitioners, alongside documentary analysis spanning the decade. It is derived from that research, but written differently: condensed, reorganised for a practitioner and field-building audience, and without the theoretical apparatus and full evidentiary detail of the original. The analysis ends in 2021, and the field has continued to evolve since. The perspective is field-level rather than organisational: it explains how the field developed, not how any particular organisation should practise. The interpretations are mine; the fuller evidentiary base sits in the thesis itself, which is openly accessible here.
My doctoral research examined how IMM emerged as a field, what shaped its evolution, and why it developed in the way it did. The perspective is institutional: rather than asking which measurement approaches are best, it asks how a field forms and holds together when multiple approaches, actors, and expectations are in play — tracing the interplay of infrastructure, discourse, and contestation across the decade.
The full literature review, methodology, and evidentiary base are in the thesis.
This page presents the central findings in an accessible form. It traces IMM's evolution through three overlapping phases, each responding to a different problem. It explains why IMM developed as an interstitial field, positioned between several established domains, and why that position matters for how the field's complexity should be understood. It maps where the field's disagreements concentrated, and shows how those disagreements generated some of its most significant innovations. And it examines the mechanisms through which the field held together — as well as what those same mechanisms cost, in the form of a persistent accountability gap.
The central finding is that IMM developed and matured over this decade through a different mode than settlement. The field did not converge on a single framework or definition, and the evidence suggests it was never going to: by 2020, 89% of surveyed impact investors used at least one external framework — where the field's first annual survey had found 85% relying on proprietary systems — yet most continued to run proprietary approaches alongside the shared ones, and the field's definitional debates persisted throughout. Widespread adoption arrived without convergence.
What developed instead is what I conceptualise in the thesis as generative plurality: a condition in which contested ideas, practices, and standards coexist as productive elements of field identity, held together by coordination mechanisms that enable diversity without fragmentation. On this account, field maturity is better understood as adaptive capacity than as normative settlement.
IMM did not develop through steady accumulation toward a settled model. It evolved through three overlapping phases, each oriented toward a different problem, and each revealing limitations that the next phase had to address. Select a phase to read its story.
If impact investing was to be taken seriously, it needed more than intention and anecdote; it needed ways to classify, measure, and compare social and environmental performance. The solution was borrowed from adjacent fields: IRIS provided a taxonomy of standardised indicators, and GIIRS sought to rate enterprises and funds in a manner analogous to credit ratings in financial markets. These were significant achievements — they gave a nascent field a focal point and a common vocabulary, and the infrastructure was largely funded by philanthropy.
But the limits surfaced quickly. Standardised metrics supported comparability while flattening context: they could count outputs more easily than they could capture significance, contribution, or variation in stakeholder experience. Many investors responded by running customised metrics alongside the standardised ones. Once that trade-off between comparability and contextual relevance was acknowledged, plurality stopped looking like an accident of immaturity. It was built into the work.
The label Rational denotes the phase's prevailing orientation: a positivist confidence, carried over from the fields the new infrastructure had borrowed from, that the right metrics and methods would progressively improve over time.
The questions were no longer only about metrics; they were increasingly about process, interpretation, participation, and coordination — how decisions about measurement should be made, and whose perspectives count. Multi-stakeholder processes proliferated: the G8 Impact Measurement Working Group, the GECES subgroup in Europe, and practical guidance from networks such as EVPA.
The pivotal initiative was the Impact Management Project, launched in 2016 as a deliberately collaborative and time-bound effort to build shared fundamentals. The IMP carried a critical conceptual shift — from impact measurement as retrospective assessment to impact management as a continuous, decision-oriented process — and its Five Dimensions of Impact offered a common vocabulary that different actors could interpret according to context and purpose.
The phase broadened the field's boundaries, bringing evaluative, managerial, and stakeholder logics alongside the financial logic that had dominated early on. This period also saw approaches that departed from the metrics-first orientation of the first phase — most notably Acumen's Lean Data, which prioritised direct customer and beneficiary insight over standardised indicators.
The emphasis shifted from measurement as a standalone exercise toward embedding impact in management, governance, and field-level interoperability. The Operating Principles for Impact Management and the SDG Impact Standards codified expectations around governance, process, and verification. IRIS evolved into IRIS+, transforming from a static metrics catalogue into a modular system linking IRIS metrics to the Five Dimensions and the SDGs.
The IMP's Structured Network brought the field's major standard-setting organisations — fifteen of them, including B Lab, GIIN, GRI, IFC, OECD, PRI, SASB, UNDP, and UNEP — into coordination for the first time, working together across practice, performance, and benchmarking workstreams rather than in parallel.
The period also saw IMM consolidate as a professional practice. Specialised impact-management functions and roles emerged within investment organisations; impact due diligence and portfolio-level approaches became standard components of practice; and independent verification services such as BlueMark and Tideline appeared in response to the new standards. Performance data lagged behind: although IRIS's first data report was launched in 2011, it took roughly a decade for sector-specific performance data — and a methodology for comparing impact results — to become available.
Adoption widened dramatically over this period, producing the 89% adoption figure noted in the overview. Yet the same period saw impact washing emerge as a headline concern, because the flexibility that had enabled broad participation also created space for actors to claim alignment without changing practice.
| Rational2011–2013 | Relational2014–2017 | Integrative2018–2021 | |
|---|---|---|---|
| Focus | Metrics, methods, and early infrastructure | Convening, coordination, and guidelines | Common language, integration, and accountability |
| Orientation | Positivist | Interpretivist | Integrative |
| Key actors | Rockefeller Foundation, GIIN, B Lab | G8 SIITF, EC/GECES, IMP, EVPA, Bridges | IMP Structured Network, IFC, UNDP |
| Infrastructure | IRIS metrics, GIIRS ratings | IMWG report, GECES report, IMP Five Dimensions | IMP Norms, OPIM, SDGIS, IRIS+ |
| Boundary work | Impact investing focus | Capital chain inclusion | Cross-sector engagement |
| Dominant logics | Financial + evaluative | + stakeholder + managerial | Various / blending |
| Core challenge | Narrow investor focus; fragmented tools | Multiple approaches; lack of common language | Balancing standardisation with flexibility; impact washing |
Adapted from the thesis (Table 11: Integrated Analysis of Phases of IMM Field Emergence and Evolution).
The arc across the three phases is the important finding. Each phase addressed the previous phase's defining problem — credibility, then legitimacy, then accountability — and each generated the problem that defined what followed. The layers accumulated rather than being cleanly replaced, which is why the field can feel crowded.
The frameworks in use today still carry the history of the problems they were built to solve. They are not neutral tools floating above the field — they embody the trade-offs of the moment in which they gained traction.
One pattern within this arc deserves note: the deliberate expansion of the field's boundaries, visible in the table above — from technical measurement for impact investors, to inclusion along the capital chain, to positioning IMM as relevant to any organisation. The expansion was strategic, bringing legitimacy and reach, but it should not be overstated. The field's centre of gravity remained with investors throughout — and within the investment community, asset managers were considerably more engaged in building and adopting the field's infrastructure than the asset owners whose mandates ultimately shape it.
The IMM field emerged at the intersection of several established domains — finance, programme evaluation, management practice, and stakeholder-oriented accountability traditions, with philanthropy and international development close at hand. Institutional theorists describe such spaces as interstitial issue fields: fields that form around a shared issue in the spaces between established domains, drawing concepts, tools, and legitimacy from each neighbouring domain without being reducible to any one of them, and without a single dominant institutional logic to organise them.
Each of those domains contributes a distinct logic of what counts as good practice:
When actors in IMM disagree about what good practice looks like, they are often not disagreeing about technical design at all. They are bringing different institutional logics to bear on the same question — which is why debates about evidence, materiality, or stakeholder voice persist even when everyone appears to be using the same words. Shared terminology does not indicate shared assumptions.
This interstitial position explains why settlement was never the available outcome. Convergence on a single standard would have required one parent logic to win — for IMM to become, in effect, a sub-discipline of financial reporting, or of evaluation, or of management. No actor or coalition had the authority to impose that resolution, and any attempt to do so would have expelled the communities whose participation made the field viable. Plurality in IMM is therefore structural rather than transitional. Many young fields experience contestation on their way to settlement; what distinguishes an interstitial field is that the contestation is replenished continuously by its parent domains, each of which keeps supplying practitioners, concepts, and expectations shaped by its own logic.
Plurality in IMM is structural rather than transitional: it reflects the field’s position at the intersection of established domains, each of which continues to supply its own logics, practitioners, and expectations. Complexity, on this reading, is not simply a coordination failure to be engineered away — though neither is it costless, as the gaps examined below make clear.
Across the decade, disagreement in IMM clustered in four areas — purposes, standards, evidence, and accountability — and each marked a point where the field's parent logics met. Mapping them matters because the field's most significant innovations each emerged from one of these areas, not from the spaces where actors already agreed.
Measurement for proving met measurement for improving.
The most fundamental disagreement concerned the purposes and scope of IMM. The recurring tension ran between measurement for proving — reporting and communicating impact externally, principally to and for investors — and measurement for improving: collecting information that organisations and their stakeholders actually use to learn and to deliver better. Practitioners across the field acknowledged that while the language of improvement was widely espoused, the proving function continued to drive most measurement in practice.
Related disputes concerned the balance between investor-centric and stakeholder-oriented approaches, and where the boundaries of accountability should sit. The field's response was not to adjudicate among these purposes but to widen the frame that contained them — and the shift to impact management, examined below, gave those plural purposes a shared home.
The pursuit of comparability met the recognition of context.
Standardisation was positioned simultaneously as the necessary foundation of field-level coherence and as a reductionist imposition that flattened the multidimensional realities of social impact. The trajectory of adoption captures how the field carried this tension rather than resolving it: the movement from 85% proprietary systems to 89% external framework use looks like standardisation's victory, but investors coalesced around several standards simultaneously — the SDGs, IRIS+, and the IMP's frameworks — while continuing to maintain customised approaches alongside them.
The field's most successful response was interoperability rather than a single winner: the evolution of IRIS into IRIS+, designed to be interoperable with other frameworks, sought to create a more adaptive infrastructure in which standards interlock rather than compete.
Positivist confidence in metrics met interpretivist attention to context and meaning.
The pressure to reduce impact to a small number of comparable indicators met the insistence that impact is irreducibly multidimensional. One expression of this tension was the field's slow movement from outputs to outcomes — from counting jobs created to asking about the quality and durability of those jobs.
Another was a pair of instructive setbacks: early attempts to import evidence hierarchies from evaluation, and GIIRS's struggles with adoption, credibility, and a sustainable business model, both demonstrated that infrastructure borrowed from adjacent fields transferred only partially to IMM's more heterogeneous terrain.
The most durable response reframed evidence around decision-usefulness: in a world of imperfect data, the job is to make the best decision possible with the information available — not to go away and prove impact after the fact. That reframing, along with Lean Data's demonstration that stakeholder perspectives could be gathered rigorously at scale, is taken up in the innovations below.
Voluntary disclosure met rising expectations of integrity.
Unlike financial accounting, where standards and independent audit lend consistency and credibility, impact reporting developed on voluntary disclosure and internally developed metrics, inviting selective disclosure and over-claiming. As impact washing concerns sharpened, newer standards — the Operating Principles and the SDG Impact Standards — embedded verification requirements, a genuine advance.
But verification practice concentrated on practice-based accountability, assessing whether organisations do what they say they do, rather than performance-based accountability, assessing whether the claimed impact occurred. The deepest questions — who defines reporting standards, how they are audited, and what consequences follow from non-compliance in the absence of regulation — remained open at the end of the study period, and remain open now. These questions are examined at length in the gaps section below.
The research describes this dynamic as productive contestation: the generative role of ongoing debate, negotiation, and conflict in shaping the evolution of the field. Three innovations illustrate the pattern. In each case, a concern from the field's margins moved to its centre through translation rather than consensus. None of these innovations fully resolved the tensions from which they emerged, but each shifted the field’s expectations of what credible practice involves.
Emerged from contestation over purposes.
The early framing of impact as something assessed after the fact, by experts, for upward reporting, came under pressure from several directions at once — too retrospective for investors who wanted decision-useful information, too extractive for enterprises and communities supplying the data, too compliance-oriented to support learning.
The language of impact management translated that dissatisfaction into a usable form, tying impact to decision-making, portfolio construction, and governance, and allowing actors with very different mandates to locate themselves in the same practice: an investor seeking primarily to avoid harm and a catalytic capital provider seeking to contribute to solutions could both see themselves in impact management without agreeing on what impact meant in every context.
The shift was carried most visibly by the IMP's widely adopted definition — identifying and considering the positive and negative effects of business actions on people and planet, then working to mitigate the negative and maximise the positive — which reframed the purpose of the entire enterprise from backward-looking assessment to a continuous, decision-oriented process. Compared to the complex methodologies proposed in earlier phases, the relative simplicity of the IMP's frameworks — the Five Dimensions and the ABC classification (Act to avoid harm, Benefit stakeholders, Contribute to solutions) — encouraged adoption well beyond IMM specialists.
By emphasising decision-usefulness and organisational learning, impact management also offered the field's most durable answer to the contestation over evidence: it created a more direct relationship between measurement and its implications for strategic and operational decisions, navigating the tension between rigour and relevance rather than resolving it.
Emerged from contestation over evidence — specifically, whose voices count.
Early IMM efforts were often criticised for their lack of engagement with end beneficiaries, and the gap became increasingly visible: impact reporting frequently ignored the voices of those most affected by investments. Lean Data, developed by Acumen, was a direct response to that critique. At its core it represents a fundamental reframing of the measurement process — elevating the perspectives of customers and end-users as critical to understanding impact, and shifting the focus from compliance and reporting to learning and value creation.
By prioritising rapid, low-cost data collection methods, such as mobile surveys and voice messaging, Lean Data enabled organisations to gather timely, actionable feedback directly from stakeholders, informing ongoing decision-making and course correction. In doing so, it challenged the field's reliance on extractive, resource-intensive measurement processes — and the power dynamics and biases inherent in the expert-driven evaluation approaches routinely used in international development.
The concern was not new. Debates about downward accountability — accountability to the people programmes are meant to serve — run back to international development in the early 1990s. What Lean Data did was translate that concern into the market vocabulary of impact investing: beneficiary voice became customer feedback, anchored in a mentality of serving the customer, and delivered through technology that made listening cheap, fast, and repeatable.
Methodologically, Lean Data is itself a synthesis — integrating principles of lean experimentation, human-centred design, and participatory evaluation. That combination demonstrated that engaging beneficiaries could be both practically feasible and insightful for investors and enterprises, countering the assumption that stakeholder-centred measurement was necessarily slow, expensive, or anecdotal.
Its downstream effects went further than method. Lean Data provoked new questions about the expectations and thresholds of impact performance reporting, and over time contributed to the development of alternative benchmarks — built on context-rich data gathered from individuals and organisations across different regions and demographics — a bottom-up complement to the field's top-down standards. At the same time, the research is clear about the limits: approaches like Lean Data remained largely voluntary, and stakeholder accountability was not yet systematically embedded as an expectation across the field by the end of the period.
Emerged from contestation over what the field kept leaving out.
The impact risk framing is a conceptual innovation that integrates previously siloed concerns — negative impacts, unintended consequences, and impact integrity — into a risk management framework. It works by blending logics: it takes the concept of risk from the financial logic and combines it with the evaluative logic's emphasis on holistic measurement, reframing the conversation around negative impacts and unintended consequences.
The effect is a shift from retrospective measurement to proactive management of potential impact failures or shortfalls. In doing so, impact risk crystallises the tension between the aspiration for positive impact and the reality of uncertainty and complexity in social impact pathways — making impact failure something to analyse and manage rather than an embarrassing exception.
Impact risk entered mainstream practice as one of the Five Dimensions of Impact, with its own defined categories of risk. That placement mattered twice over. It elevated a traditionally underexplored aspect into a core analytical category — alongside underserved stakeholders — and it aligned impact risk with established practices in financial risk management, making the concept legible and relevant to mainstream investors entering the field.
The framing also functions as a productive provocation. By recognising the inherent uncertainties, contingencies, and contextual variabilities that shape impact pathways, it demands a more adaptive and probabilistic approach to impact goal-setting, measurement, and management — a rethinking of how organisations define, measure, and communicate impact goals, and how they allocate resources and make decisions under uncertainty. Its importance has only grown as impact washing concerns have become more visible.
The honest caveats from the research: evidence on how impact risk is actually being used across different contexts remained limited by the end of the study period, and — together with the Contribution dimension — it remains among the most contested aspects of the Five Dimensions. Its significance lies less in settled practice than in what it demonstrates: previously contested aspects being integrated productively into the field rather than ignored or downplayed.
A microcosm of how the field holds competing approaches productively.
A prominent example of the field's plurality is impact monetisation — the attempt to express impact in monetary terms. Multiple techniques proliferated, each seeking legitimacy in a different institutional context: SROI emphasises stakeholder engagement and social value creation, aligning with nonprofit and community-based organisations; the Impact Multiple of Money developed by TPG's Rise Fund reflects a more quantitative, investor-oriented perspective; Impact-Weighted Accounts seek to integrate monetisation into corporate accounting and reporting; and social impact bonds put monetisation to work in outcomes-based public-sector financing.
The debate was genuine on both sides. Proponents argued that monetisation could enhance the comparability and decision-usefulness of impact data; critics cautioned against reductionism and commodification. Many practitioners remained more inclined toward scoring or rating approaches than toward aggregating impact and financial performance into a single number — and the divergence persisted.
What makes the case instructive is what the field did with that divergence. Rather than converging prematurely on a single dominant approach to impact valuation, actors chose methodologies aligned with their specific goals and capacities, and the coexistence of approaches encouraged experimentation and adaptation. The debate also surfaced productive tensions about the ethics and limits of quantifying complex social outcomes — questions a settled standard would have buried. Monetisation is generative plurality in miniature: a contested space that enriched the field precisely because it was not resolved.
Not all contestation is productive by default. Disagreement can also fragment attention, reproduce power imbalances, or allow ambiguity to become a cover for weak practice. Some forms of flexibility broaden participation; others dilute accountability. This is one of the central tensions of IMM's development: the same openness that enabled innovation also made it harder to settle terms, fix boundaries, and enforce consequences.
That does not invalidate the innovations; it means we should understand them realistically — not as final resolutions, but as negotiated responses to persistent tensions. The more useful question is not how to eliminate disagreement, but how to distinguish contestation that sharpens practice from proliferation that obscures it.
If the field did not settle, what held it together? The answer has two layers: the infrastructure the field built, and the mechanisms through which that infrastructure enabled coordination.
The research identifies a typology of five interconnected forms of institutional infrastructure, each playing a distinct role in the field's development:
These forms developed cumulatively and iteratively: early methodological standardisation laid the groundwork for the normative frameworks of the middle period, which in turn provided the foundation for more flexible, adaptive platforms like the IMP. The pattern across the decade ran from siloed to connected to blended — elements developed separately were progressively linked and hybridised, as when theory of change, once deployed as a standalone process, was absorbed into the GECES guidance via EVPA and later into the Operating Principles. The result was infrastructure that is modular and interpretively flexible — preserving rather than eliminating institutional complexity, and enabling coordination through interoperability rather than standardisation. Nor was any of it neutral: IRIS's metric categories defined what counted as reportable impact, the Five Dimensions established the field's shared vocabulary, and OPIM's verification requirement signalled its aspiration toward accountability — each decision embedding assumptions about who the field is for and what counts as legitimate practice.
The second layer is how this infrastructure worked. The research identifies three field-level mechanisms through which IMM leveraged, rather than resolved, its institutional complexity. Contestation created the conditions for questioning and experimentation; these coordination mechanisms prevented that questioning from tipping into fragmentation.
The field's most influential frameworks succeeded not because they resolved differences but because they let different communities coordinate without resolving them. The Five Dimensions of Impact are the paradigm case. They combine elements of financial logic (contribution, risk), evaluative logic (evidence of outcomes), and stakeholder logic (who is affected, and how much), providing enough structure to support shared conversation while leaving room for interpretation in context. In sociological terms they function as boundary objects: common reference points that investors, evaluators, enterprises, and policymakers can each use according to their own priorities.
Two design features explain the Dimensions' reach. First, the level of abstraction at which consensus was achieved was deliberate: going one layer deeper — into specific questions, data requirements, and sub-components — would have made agreement far more difficult. Practitioners consistently credited the convening behind the frameworks: bringing that many actors around the table, with a frame that meant something to specialists while remaining accessible to those just entering the conversation. Second, the linkages between dimensions, such as how stakeholder characteristics inform outcome thresholds, or how duration connects to impact risk, created natural bridges between assessment and action, helping shift IMM beyond a narrow measurement focus while preserving flexibility in implementation.
Boundary objects also needed boundary-spanners. The IMP — and initiatives like it — actively reframed field boundaries, brokered relationships across institutional domains, and advocated for field-level coordination, culminating in the Structured Network's coordination of fifteen standard-setters. Nor are the Dimensions an isolated case: the Operating Principles embed evaluative and managerial logics alongside financial ones, and Lean Data foregrounds the stakeholder logic that earlier phases had marginalised.
Key terms — impact management, impact performance, even impact itself — were maintained with a degree of deliberate interpretive openness. This was not sloppiness. It allowed diverse actors to participate in the field without first resolving every definitional dispute, and it supported coalition-building and experimentation during the years when the field's development depended on breadth.
Definitional flexibility functioned, in effect, as coordination infrastructure rather than as a temporary uncertainty awaiting resolution. A more rigidly defined field might have achieved greater precision in places, but at the cost of narrower participation and weaker adaptability.
What distinguishes IMM's use of ambiguity from ordinary definitional drift is that it involved strategic cultivation rather than passive tolerance. The field's most effective initiatives treated interpretive openness as a design choice — the IMP's frameworks are deliberately pitched at a level of abstraction that multiple communities could adopt — while narrowing ambiguity selectively where coordination demanded it, as verification and disclosure expectations did in the later period. The risk, examined in the gaps section, is that the same openness can become cover for weak practice when it is no longer strategic.
The travels of the term impact management itself illustrate the point. It was broad enough to be claimed by development finance institutions, private equity funds, foundations, and social enterprises alike — actors with very different mandates and definitions of impact — and that breadth is what allowed a shared, field-level conversation to form. The IMP's own language reflects the same choice: its consensus outputs were framed as norms — shared fundamentals open to interpretation in context — rather than as standards demanding uniform compliance.
Rather than one logic displacing another, or logics merging into a stable hybrid, the field developed the capacity to combine elements of different logics as circumstances required — financial and evaluative logics in the early infrastructure, stakeholder and managerial logics as the field broadened, and increasingly deliberate integration across all of them in the later period.
The blending is visible across the field's infrastructure and its debates. GIIRS combined the financial logic's rating architecture with evaluative ambitions; theory of change carried the evaluative logic from grant-funded programmes into investment strategies; the Operating Principles embed evaluative and managerial logics alongside financial ones; Lean Data joined the stakeholder logic to a market vocabulary of customer feedback; and the monetisation debate set the financial logic's drive for commensuration against evaluative and stakeholder concerns about what resists quantification. The combinations shifted with context and purpose rather than settling into a single stable form — which is what distinguishes blending from the more familiar idea of a hybrid.
The progression is visible in the infrastructure itself: IRIS+ linking metrics to the Five Dimensions, the SDGs, and evidence frameworks is logic blending made concrete, in modular components designed to interlock with what already exists rather than replace it.
Each of the three innovations above is an act of logic blending. Impact management blends managerial decision-usefulness with evaluative learning. Lean Data blends stakeholder voice with a market logic of customer feedback and a technology-enabled method. Impact risk blends the financial logic's concept of risk with the evaluative logic's holistic view of impact pathways. The pattern is consistent: the field's advances came not from choosing among its parent logics but from combining them in forms that multiple communities could recognise and use.
Taken together, these mechanisms describe how generative plurality worked in practice. Where existing accounts of field development emphasise convergence, hybridisation, or settlement as the pathways to stability, IMM demonstrates that a field can maintain both dynamism and coherence: contestation creating the conditions for questioning and experimentation, and coordination mechanisms preventing that questioning from tipping into fragmentation.
Productive contestation and generative plurality work in conjunction: contestation is the engine of the field's innovation; plurality is the condition — sustained by boundary objects, strategic ambiguity, and logic blending — that lets contested ideas, practices, and standards coexist as productive elements of field identity rather than forces of dissolution.
Field maturity, on this view, is not the moment contestation ends. It is the demonstrated capacity to sustain contestation as a resource for learning and adaptation without fragmenting.
A fair reading of this argument has to confront its costs, because the mechanisms that built the field also created its most serious vulnerability. The same ambiguity that broadens participation allows actors to align themselves with the field's discourse without substantively changing their practice. The same boundary objects that travel across communities can be adopted procedurally and ignored substantively. Flexibility that fosters inclusion can equally foster performative engagement.
These are self-assessments rather than verified results, and that is precisely the point: the juxtaposition of near-universal satisfaction with rising integrity concerns indicates a field in which claims, self-assessment, and consequence have come apart. Practitioners in the research described the missing piece plainly: what people want is the link between what an organisation says it will do and evidence that it did it — and the field has cut some of the building blocks for that link without yet enabling it.
This marks the limit of what infrastructure alone can achieve. Over the decade, IMM produced a substantial architecture: metrics, principles, taxonomies, governance standards, verification regimes, and benchmarking approaches. But infrastructure alone does not ensure meaningful use — organisations can adopt field language without embedding it, and comply procedurally without changing behaviour — and the verification regimes that emerged, a genuine advance, concentrated on whether organisations follow their stated processes rather than whether the claimed impact occurred. What remains unresolved is institutional rather than technical: who defines the standards against which impact is reported; how those reports are audited, and by whom; whether verification moves beyond assessing process toward assessing performance; and what consequences follow when claims and evidence diverge, in a field that still operates largely without regulation.
The gap is also a question of power and inclusion. The field's emphasis on multi-stakeholder engagement has sometimes obscured underlying power asymmetries: more privileged actors — investors, and those based in the Global North — have been able to shape the terms of discourse and decision-making in ways that marginalise or exclude other voices, and participants seeking Global South perspectives on field-building initiatives repeatedly encountered the same narrow circle of contacts. The asymmetry is epistemological as well as geographic: the dominance of Western, technocratic, expert-driven approaches has been critiqued for marginalising culturally diverse ways of understanding and assessing social value — including Indigenous and traditional knowledge systems that emphasise qualitative, relational, and context-specific approaches — and frameworks built in one context transfer imperfectly to others.
These dynamics complicate the field's pluralism. Methodological pluralism and interoperability can exacerbate power imbalances by enabling well-resourced actors to shape the terms of impact measurement in ways that privilege their own interests, and integrative frameworks risk transforming inherently political and ethical concerns into technical exercises. Lean Data and impact risk show that contestation can amplify previously marginalised voices — but realising that potential requires attending to how power relations shape whose interpretations prevail within ostensibly neutral frameworks.
Plurality is not generative by default. It becomes degenerative precisely when it serves as cover for conceptual looseness, selective reporting, or weak accountability — and distinguishing between the two is the field's central unfinished task.
The findings carry implications at three levels: for how fields addressing complex societal problems are built, for how IMM is practised and taught, and for where the field's own attention should go next.
For field-building. The decade's most durable infrastructure shared a design pattern: it was modular and interoperable rather than totalising. IRIS+ succeeded by aligning to the Five Dimensions and the SDGs rather than displacing them; the IMP succeeded as a deliberately time-bound, collaborative platform rather than a permanent authority; and the frameworks that travelled furthest were the ones that functioned as boundary objects — structured enough to coordinate around, open enough to interpret in context. The general lesson is to design for plurality rather than pursue premature standardisation: initiatives that attempt to settle terms before the relevant communities are at the table tend to achieve precision at the expense of participation, and the participation is what makes an interstitial field viable.
For practice and education. If plurality is structural, then navigating it is a core professional competence, not a temporary inconvenience. The hardest problems in IMM practice are interpretive and relational rather than technical: knowing what problem a given framework was designed to solve, which logic it privileges, what it makes visible and what it leaves in the background, and how to translate across communities that use the same words to mean different things. Framework selection becomes a matter of fit-for-purpose judgment rather than a search for the correct standard. The field's capability-building — its training, its guidance, its professional development — has largely been organised around technical proficiency, and has not yet caught up with the interpretive expertise its own structure demands.
For the field's next phase. The infrastructure-building agenda that defined 2011–2021 has largely succeeded on its own terms; the unfinished agenda is institutional. Closing the accountability gap will not come from another framework, because the gap is not primarily a measurement problem. It is a question of governance and consequence: whether verification moves from assessing process to assessing performance; whether reporting extends across the full range of results rather than only the successes; whether the field's governance becomes genuinely more inclusive of the people and places capital is meant to serve; and what actually happens — in incentives, in capital allocation, in credibility — when claims and evidence diverge. These questions sit downstream of the infrastructure the field spent a decade building, and they will determine whether that infrastructure ends up mattering.
The research also opens an agenda for further study, set out in full in the thesis: comparative work on other interstitial fields such as sustainable finance and ESG investing; longitudinal work on whether generative plurality can be maintained, scaled, or subverted over longer timeframes; closer study of the design and governance of field-level infrastructure; and direct engagement with Global South experiences and Indigenous ways of knowing, confronting the overrepresentation of Global North perspectives in how the field — and research about it — has developed.
The conclusion I draw from a decade of IMM's evolution is therefore not that the field needs to converge, and not that its complexity will — or should — resolve. IMM is unlikely ever to look like financial accounting, because it sits where accounting's questions meet evaluation's, management's, and the claims of those whom capital is meant to serve. The useful question is no longer when the field will settle. It is what keeps its plurality generative rather than degenerative — and that is a question about institutions, incentives, and judgment, not only about tools.