Chapter 2
Theoretical Framework and Methodology
This chapter sets out the theoretical and methodological foundation of the work. Section 2.1 locates the research question within infrastructure theory. Section 2.2 explains why a separate evaluation framework must be developed, documents the research gap, and demarcates against existing approaches. Section 2.3 presents the evaluation framework in full, from the evaluation logic through the twelve criteria to the scope boundaries. The work pursues a dual goal: it delivers a reasoned judgment on the research question and simultaneously documents evidence and conclusions with sufficient transparency for the reader to arrive at an independent, potentially divergent judgment.
2.1 What Is Infrastructure?
The research question of this work is whether Ethereum’s architecture meets the requirements that must be placed on a fundamental digital infrastructure. Before this question can be answered, it is necessary to clarify what the term infrastructure means in an academic context. In political and media discourse, the term has become inflated: broadband networks are described as digital infrastructure, education systems as social infrastructure, charging stations as charging infrastructure, and in the crypto industry every protocol with more than a handful of users goes by the name of infrastructure. Everyday usage follows an implicit heuristic, according to which any system that appears important and is used by many deserves the title of infrastructure. For an academic assessment this heuristic does not suffice, because it draws no line between infrastructure and non-infrastructure and thereby renders the research question unanswerable. The following theoretical perspectives supply the conceptual precision that makes such a line possible.
The Economic Definition
In 2012, Brett Frischmann presented, in Infrastructure: The Social Value of Shared Resources, a foundational economic definition of infrastructure.1 His central insight shifts the analytical focus: infrastructure is defined by what it enables, not by what it is. Whether a system is physical or digital, large or small, expensive or inexpensive tells us nothing, according to Frischmann, about whether it is infrastructure. What matters are three functional criteria. The first criterion is non-rivalry in consumption up to a capacity threshold: use by one actor does not preclude simultaneous use by others, as long as the system does not reach its capacity limit. The second criterion concerns the demand structure: demand for the system is driven primarily by productive activities that require it as an input, with the aggregate social benefit exceeding individual benefit. The third criterion is input character: the system serves as a factor of production for further activities and is not an end product.
The road network illustrates the three criteria. It is non-rival as long as there is no congestion, because one user’s journey does not preclude another’s until the road’s capacity limit is reached. It generates productive demand, because trade, labor mobility, and social participation build on a functioning road network and the aggregate social benefit exceeds individual travel benefit. It has input character, because without roads neither a logistics sector nor widespread retail nor a labor market with commuter flows could exist. Frischmann’s three criteria thus describe a specific economic role: infrastructure is an enabling structure whose value is realized in the activities it makes possible.
From this insight Frischmann draws a regulatory consequence that is directly relevant to the subject of this work. Because infrastructure generates its value through enabling and that value increases with the breadth of use, commons management — that is, non-discriminatory open access — is in many cases economically more efficient than restricting access through private property rights.2 Ethereum’s permissionless design, which allows every actor to use the network without the permission of a gatekeeper, is a technical implementation of this argument. The connection between Frischmann’s infrastructure theory and blockchain systems is not an analogical inference drawn by this work: James Grimmelmann and A. Jason Windawi showed in 2023, in the William & Mary Law Review, a respected American law review, that blockchains meet Frischmann’s infrastructure criteria and simultaneously function as semicommons — that is, as systems with intertwined private and shared resource use.34 The present work builds on this academically established connection.
The Sociological Extension
Frischmann’s definition is economic and works for systems whose value creation is measurable. Susan Leigh Star introduced in 1999, in The Ethnography of Infrastructure, a perspective that goes beyond the economic determination of function.5 Her central insight is: infrastructure is relational. It arises not from objective properties of a system, but in relation to organized practices. What for the engineer maintaining the fiber-optic network is an object of work is, for the user conducting a video conference over that network, invisible infrastructure. What for the inhabitant of an industrialized country is self-evident infrastructure can, for the inhabitant of another country, be an unattainable good. Star identified a further property that belongs to the standard repertoire of infrastructure research: infrastructure becomes visible upon breakdown. As long as it functions, it is invisible. No one thinks about the power grid until the electricity fails, and no one thinks about the Domain Name System until a website becomes unreachable.
Star’s perspective alters the research question of this work in a way that shapes the entire evaluation framework. If infrastructure is relational, then the question of whether Ethereum is infrastructure cannot be answered in binary terms, because the answer depends on for whom and in what application context it is posed. For a DeFi developer whose entire business logic is built on Ethereum smart contracts, Ethereum is already infrastructure today in Star’s sense, because the system is an invisible precondition for their organized practice. For a user making a single transaction, it is more of a tool. This context-dependency is precisely what the meta-pattern operationalizes that shapes the entire evaluation logic of this work: the degree of fulfillment must match the claim. The evaluation framework does not assess whether Ethereum is infrastructure in an absolute sense, but whether it meets the requirements that must be placed on a fundamental digital infrastructure — that is, context-sensitive and calibrated to the claim.
Digital Infrastructure as a Category in Its Own Right
The theory discussed so far describes predominantly physical infrastructure, and transferring it to digital systems requires a conceptual foundation of its own. Ole Hanseth and Kalle Lyytinen defined information infrastructures in 2010 as shared, open, heterogeneous, and evolving sociotechnical systems, and identified two fundamental design problems that distinguish digital from physical infrastructure.6 The bootstrap problem describes the challenge that an infrastructure must be immediately useful to early users in order to gain the momentum that enables later network effects. The adaptability problem describes the challenge that local design decisions must accommodate unlimited growth and functional uncertainty, because the future forms of use of the infrastructure are not foreseeable at the time of design. Tilson, Lyytinen, and Sørensen argued in the same year that digital infrastructure constitutes its own analytical category, with its own paradoxes — in particular the tension between openness and control — and its own research questions arising from the intertwining of technical and institutional dynamics.7
Both problems manifest in the subject of this work. Ethereum had to start in 2015 — financed through a crowdsale of 18.3 million US dollars — with a concrete use case that was immediately useful to early users: the possibility of deploying and executing smart contracts. The EVM, Ethereum’s virtual machine, provided from day one a programmable execution environment that went beyond pure payment transactions and thus created the foundation for the network effects that sustain the system today. The adaptability problem manifests in the roadmap complexity that Chapter 5 of this work describes in detail: the system must evolve to meet growing demands for scaling, data availability, and user access, without breaking the existing user base and the applications built on the system.
The Conceptual Ground of the Evaluation Criteria
The three perspectives together define the analytical framework within which the research question stands. Frischmann supplies the economic foundation: infrastructure is an enabling structure whose open access can be justified economically. Star supplies the sociological deepening: infrastructure is relational and context-dependent, which rules out a binary answer to the research question and requires an assessment calibrated to the claim. Hanseth and Lyytinen supply the information-technological specification: digital infrastructure is subject to its own design problems that must be taken into account when assessing a concrete system. Against this backdrop the question of whether Ethereum is fundamental digital infrastructure can be posed as a question that is academically tractable: it asks whether a concrete system exhibits the properties that real infrastructures share. The three perspectives thereby supply the conceptual ground from which the evaluation criteria are analytically derived. The interplay of the economic, the relational, and the information-technological determination of infrastructure is what constitutively characterizes fundamental digital infrastructure. Chapter 3 then tests the criteria derived in this way against seven real infrastructures. The theory supplies the framework and the derivation. The seven reference systems supply the empirical test.
For a work written in German, the tradition of large technical systems is additionally relevant. Renate Mayntz conceived large technical systems in 1993 as components of function-specific infrastructure systems and identified the tension between control and self-organization as a central research problem.8 The tension describes a fundamental conflict inherent in every infrastructure: the larger and more complex the system, the more difficult central control becomes, and the stronger the system’s own dynamics. Ethereum’s governance model, which operates in a decentralized manner without a formal governing authority and rests on the principle of Rough Consensus, is an extreme case of this tension. There is no central authority that can mandate protocol changes. Every change must be voluntarily adopted by a majority of node operators. Whether this model meets the requirements of adaptive governance that must be placed on fundamental infrastructure is one of the twelve assessment dimensions of the evaluation framework of this work.
An institutional-economics perspective complements the framework. Sinclair Davidson, Primavera De Filippi, and Jason Potts argued in 2018 that blockchain represents an institutional innovation that extends Oliver Williamson’s governance framework beyond the dichotomy of markets and hierarchies.9 Ethereum can be read in this perspective as a system that realizes a third form of governance: protocol-based coordination through algorithmic rules, which are neither negotiated in markets nor ordered hierarchically, but arise from the consensus of network participants.
2.2 Why a Separate Evaluation Framework?
The Research Gap
The infrastructure theory set out in the preceding section offers a mature conceptual apparatus for posing the question of infrastructure suitability. For blockchain systems in general and Ethereum in particular, this apparatus has not yet been systematically applied. Svein Ølnes and Arild Jansen transferred in 2018 and 2021 Star and Ruhleder’s relational infrastructure definition and Hanseth and Lyytinen’s information infrastructure theory to blockchain in general, and showed that blockchain systems can be conceived as information infrastructures.1011 Their work identifies properties that blockchain shares with the information infrastructures described in IS research — such as openness and heterogeneity of users — and concludes from this that the vocabulary of infrastructure research is applicable to blockchain. The works describe blockchain as infrastructure. Whether a concrete system meets the requirements of fundamental infrastructure, they do not examine: what is missing are operationalized criteria, thresholds, and an assessment mechanism. Kelsie Nabben analyzed Web3 in 2023 as self-infrastructuring — that is, as a process in which the community itself produces and continuously transforms the infrastructure it uses.12 Her analysis works with the vocabulary of infrastructure studies and provides insights into the emergence dynamics of decentralized infrastructure, but aims at process description rather than suitability assessment. Paul Dylan-Ennis, Donncha Kavanagh, and Luis Araujo examined Ethereum in 2023 as a sociotechnical project with three imaginaries — three imaginative worlds that drive the Ethereum community: Ethereum as a world computer, as a decentralized financial infrastructure, and as a coordination protocol.13 Their sociological analysis sheds light on the self-description and internal tensions of the community. It does not deliver an assessment of the architecture or of infrastructure suitability on the basis of verifiable criteria.
The gap that this work fills is thus specific: there exists no framework-based assessment that defines operationalized criteria, sets transparent thresholds, and applies a reproducible evaluation procedure to examine whether Ethereum meets the requirements of a fundamental infrastructure. The works mentioned contribute preliminary work on which this work builds, but none of them answers the research question.
Demarcation against Existing Blockchain Assessment Approaches
The need for a separate evaluation framework becomes clearer when one considers the existing assessment approaches from the blockchain literature. What the approaches share is a common thrust: they measure decentralization or organizational suitability, but they do not assess infrastructure properties in the infrastructure-theoretical sense.
In 2017, Balaji Srinivasan and Leland Lee proposed the Nakamoto Coefficient, which measures the degree of decentralization of a system as the minimum number of actors that would need to collude to compromise the system.14 The coefficient captures an important property, but says nothing about infrastructure suitability: SWIFT, the global financial messaging system, has a Nakamoto Coefficient close to one, because it is operated by a single organization, and is nonetheless fundamental infrastructure of international payments. In 2018, Adem Efe Gencer and colleagues presented an empirical decentralization measurement via mining distribution and network topology that quantifies the operational degree of decentralization, but captures neither governance dimensions nor institutional embeddedness.15 The most recent systematization of the research strand, a SoK study (Systematization of Knowledge) on measuring blockchain decentralization published in 2025, inventories all existing decentralization metrics and makes visible that the entire research strand is oriented toward decentralization measurement.16 The question of whether a system meets the requirements of infrastructure is addressed by none of the metrics.
Alongside decentralization measurement there are approaches that assess blockchain suitability from an organizational perspective. In 2023, Spencer-Hicken and colleagues developed a Blockchain Feasibility Assessment for organizations that evaluates whether a company should deploy blockchain technology. It thus makes an adoption decision and does not assess infrastructure suitability.17 In 2021, Jensen and colleagues examined the governance decentralization of DeFi protocols, which operate at the application layer, not at the base layer that this work assesses.18 In 2023, Pincheira and colleagues analyzed the infrastructure costs of blockchain systems, using the term infrastructure in the IT-technical sense — as a generic term for servers, storage, and network resources — rather than in the social sense of an enabling structure in Frischmann’s sense.19
The conclusion from the stock-taking is clear: the existing approaches ask how decentralized a system is or whether blockchain is suitable for a particular organization. The research question of this work is whether a system meets the requirements that must be placed on fundamental infrastructure. This work answers the question by means of an evaluation framework that operationalizes infrastructure properties, grounded in infrastructure theory and validated against real infrastructures. An evaluation framework that originates from the blockchain literature and aggregates decentralization metrics cannot answer the research question, because it measures the wrong property.
The Methodology of Constructing the Evaluation Framework
The work validates its evaluation framework against seven reference infrastructures: power grid, Internet (TCP/IP), road network, SWIFT, GPS, cloud infrastructure, and legal system. The seven cases were chosen because they cover five infrastructure types: physical (power grid, road network), digital-open (Internet/TCP-IP), digital-closed (SWIFT, GPS), hybrid (cloud infrastructure), and institutional (legal system). This heterogeneity represents the breadth of the infrastructure concept and strengthens the validation force of the selection, because a criterion that is found across such diverse systems can be regarded with greater plausibility as constitutive of infrastructure.
The analytical approach was deliberately chosen because there is no established evaluation framework for assessing cryptographic infrastructure. Existing assessment approaches from the telecommunications or energy sector are tailored to institutional infrastructure operated by central organizations and embedded in regulatory frameworks, such that their transfer to a permissionless system without a central operating organization would, in the work’s assessment, produce categorical misassumptions. The work therefore develops its evaluation framework independently, by means of analytical deduction. It decomposes the concept of fundamental digital infrastructure — as determined by the theoretical framework of the preceding section — into the constitutive properties that a system must exhibit in order to satisfy the claim. It operationalizes each of these properties as a verifiable criterion with its own definition, its own indicators, and its own thresholds.
The procedure follows the Conceptual Framework Analysis described by Yosef Jabareen. This qualitative method for framework construction proceeds in systematic phases from the identification of relevant conceptual and theoretical sources through the extraction of the load-bearing concepts to their integration into a coherent instrument.20 The test of an evaluation framework constructed in this way is not its origin from a set of cases, but its internal coherence, the completeness of its dimensions, and its analytical productivity — that is, the question of whether it makes the distinctions that matter for the assessment. Each of the twelve criteria is self-sustaining through its own conceptual derivation, and none claims validity on the basis of the frequency with which a property appears in a sample of real systems.
Because an analytically derived evaluation framework runs the risk of being constructed past the reality of established infrastructure, Chapter 3 validates the twelve criteria. It tests them against seven structurally heterogeneous reference infrastructures and asks whether the deductively obtained requirements withstand confrontation with systems that have historically established themselves as fundamental infrastructure. This test is not a retrospective illustration, but the empirical stress test of the evaluation framework, whose results are fully documented in Chapter 3.
2.3 The Evaluation Framework
The evaluation framework is the instrument by which the research question is answered.
2.3.1 Evaluation Logic
Three fundamental principles determine the architecture of the evaluation framework. The first principle is the conditional structure of the assessment. The research question is formulated conditionally: it asks whether the architecture fulfills the claim it raises. The result is a spectrum that ranges from “Suitable” through “Suitable with Conditions” and “Suitable under Substantial Conditions” to “Conditionally Suitable” and “Not Suitable”. Each category is defined, and assignment to a category follows from explicit rules laid down in the evaluation mechanism (M3 cascade).
The second principle is the meta-pattern introduced in Section 2.1: the degree of fulfillment must match the claim. It is confirmed in the validation against the seven reference infrastructures in Chapter 3. GPS offers maximum inclusivity, because any receiver worldwide can use the signal, but no neutrality, because the system is under the unilateral control of the US government, which can selectively switch off or throttle the signal. SWIFT offers global coordination in financial messaging, but no open access and no censorship resistance, because states can exclude individual participants from the system. Both systems are fundamental infrastructure, even though neither of them fulfills all conceivable infrastructure properties at maximum degree. The meta-pattern prevents the fallacy that Ethereum must fulfill every criterion at maximum degree in order to qualify as infrastructure. At the same time it prevents the reverse fallacy that any arbitrary degree of fulfillment is acceptable: the degree of fulfillment must match the specific claim, and the claim that Ethereum makes is that of a neutral, permissionless, censorship-resistant infrastructure.
The third principle is a deliberate design decision: the assessment is hierarchical-rule-based, not additive. The work does not use a weighted scoring model in which each criterion receives a percentage value and the overall assessment follows from the sum of the weighted individual values. Instead, a rule-based hierarchy operates in which Critical Conditions function as a threshold: if a Critical Condition is not met at the required level, this limits the best possible overall assessment, regardless of how well the other criteria perform. In an additive model, a system that stands at “Open” on Neutrality — a Critical Condition — but reaches “Met” on all other eleven criteria could still achieve a good overall result. This contradicts the infrastructure logic: a system that is not neutrally accessible cannot fulfill the claim of a fundamental infrastructure, regardless of its performance in other dimensions. The hierarchy reflects this logic. It trades accumulation sensitivity — the ability to aggregate many small limitations into a worse overall result — for reproducibility and transparency: a second researcher assessing the same material arrives at the same result, because the rules are explicit and require no implicit weighting decisions.
2.3.2 The Three-Level Logic
Each of the twelve criteria is examined at three levels. The protocol level asks what the architecture makes possible. The operational level asks what of that is actually realized. The infrastructure level asks whether the interplay of architectural possibility and operational reality fulfills the infrastructure claim. This three-level logic is the central analytical tool of the entire work. It prevents the assessment from stopping at the architecture and ignoring operational reality, which would lead to overassessment, because every well-designed protocol exhibits the right properties at the architectural level. Conversely, it rules out assessing operational reality without taking architectural potential into account, which would lead to underassessment, because operational deficits may be remediable through protocol changes. The infrastructure level synthesizes both perspectives and forces the question that is decisive for the infrastructure claim: is what is architecturally possible and operationally realized sufficient for the claim?
An example makes the three levels tangible. Censorship resistance — one of the criteria in the Critical Conditions category — operates at all three levels with different findings. At the protocol level, Ethereum’s architecture enables any valid transaction to be included in a block: the protocol itself contains no mechanism that discriminates transactions by content or sender. At the operational level, a divergent reality appears: the large majority of blocks is produced by a small number of builders, who could theoretically exclude transactions, and a considerable share of blocks is built in compliance with OFAC sanctions lists. At the infrastructure level, the decisive question arises: does the protocol-level possibility of inclusion suffice, or must operational censorship resistance be present for the system to fulfill the claim of a neutral infrastructure? The answer, which Chapter 4 justifies in detail, depends on whether the operational discrepancy undermines the infrastructure claim or whether mitigation mechanisms — such as protocol-level inclusion lists (FOCIL) — can close the gap between protocol design and operational reality. The final assessment follows the principle of the weakest link: the level with the lowest degree of fulfillment determines the overall assessment of the criterion.
2.3.3 The Three Building Blocks of the Evaluation Framework
The evaluation framework consists of three interlocking building blocks: the Implementation Status Taxonomy (M1), the thresholds for the twelve criteria (M2), and the Criteria Hierarchy and Assessment Mechanism (M3 cascade).
Implementation Status Taxonomy (M1)
The taxonomy classifies all Ethereum features according to their implementation maturity in five stages. Features with the status IMPL (implemented) are active on mainnet and in productive use, stable for at least three months and free of critical bugs. Features with the status DEPL (in deployment) are implemented in client releases and active on testnets, with an announced mainnet date. Features with the status PLAN (planned) have an accepted EIP, are in Review or Last Call status, are on the official roadmap, and have an estimated time horizon of less than 24 months. Features with the status RES (research) are mentioned in roadmap documents and are the subject of active research, but without a final specification. Features with the status DEP (deprioritized) are explicitly deferred or abandoned.
The boundary between DEPL and PLAN is the central epistemic dividing line of the work. IMPL and DEPL are treated as present or secured, because their implementation is operationally confirmed or immediately imminent. PLAN and RES are treated as conditional, with explicit uncertainty marking, because their realization depends on future governance decisions, technical feasibility, and community consensus. Everything below DEPL is projection, and the taxonomy indicates the degree of projection. Five stages were chosen because they map the relevant epistemic gradations without generating unnecessary granularity: three stages (implemented, planned, not planned) would flatten the difference between a feature in testnet deployment and a feature in early research, which is decisive for the robustness of the assessment. The taxonomy produces a certainty gradient that determines the entire work and is operationalized in the integrated robustness assessment in Section 6.1 according to the maturity level of the load-bearing features.
Thresholds: Twelve Criteria in Three Categories (M2)
The twelve criteria are divided into three categories that play different roles in the hierarchy of the evaluation framework. Each criterion is assessed on a four-stage scale. “Met” means the requirement is fully implemented and there are no structural limitations. “Met with Qualification” means the requirement is fundamentally met, with identified limitations that do not fundamentally undermine the claim. “Conditionally Met” means fulfillment depends on conditions that are not yet secured, but for which a recognizable path exists. “Open” means the requirement is not met and no operational path to fulfillment exists. The criterion assessment results from the argumentative synthesis of the indicators along the three-level logic. The indicators provide the evidence base, but the assessment does not follow a mechanical counting rule, because the indicators of a criterion carry different weight for the infrastructure claim. A counting rule that would treat, say, three out of four indicators at “Met” as sufficient for an overall “Met” would treat all indicators as equivalent, which is substantively untenable. The argumentative synthesis instead enforces a structured justification that is transparently documented in each individual assessment. The reader can assess the evidence and the conclusion independently. If an indicator does not change between two assessment timepoints, its level is retained. An unchanged indicator cannot fall below the level it achieved in the initial state.
The following paragraphs introduce the twelve criteria in the order of their hierarchical function. First come the Critical Conditions, which form the foundation of the infrastructure claim and whose non-fulfillment limits the best possible overall assessment. Then follow the Structural Conditions, which determine the load-bearing capacity of the system. Finally come the Qualitative Criteria, which differentiate the degree of suitability within the category reached.
Security and Trust Load (I.2) operates on two dimensions. The first dimension asks whether attacks on Ethereum are economically infeasible. Economic security is measured by the cost of a 34-percent attack, for which the threshold for “Met” stands at more than 30 billion US dollars, derived from the CoinMetrics study “Breaking BFT” of 2024. The 34-percent threshold follows directly from the blocking minority of the Casper-FFG consensus algorithm: an attacker controlling 34 percent of staked ETH can block finalization. The second dimension asks how much residual trust in third parties the user must place in order to use the system securely. This dimension is complementary to criterion II.3 (Independent Verifiability): whereas II.3 asks whether the possibility of verification exists, I.2 asks how much residual trust remains despite that possibility. The answers can diverge, because a system that enables verification can nonetheless generate a high trust load if the majority of users access the system through centralized intermediaries such as RPC providers.
Minimal Load-Bearing Guarantees (I.4) asks whether Ethereum offers a minimum set of guarantees that are maintained even under stress. These include Finality — the irreversibility of finalized transactions — Liveness — the ability to continue producing blocks under stress — a degradation mode in the event of a validator outage, and censorship resistance under attack, which is co-assessed via a cross-reference to criterion II.1. The time-to-finality threshold of under 15 minutes for “Met” derives directly from the Beacon Chain specification, which provides for two epochs of 32 slots each at 12 seconds per slot — approximately 12.8 minutes — for finalization.
Neutrality and Censorship Resistance (II.1) is the criterion with the largest indicator set in the evaluation framework. Seven indicators examine neutrality at two levels: at the transaction level, whether individual transactions are censored or delayed by builders, relays, or regulation, and at the system level, whether a state or group of actors can instrumentalize Ethereum as a whole. The OFAC compliance rate of blocks, builder concentration, the existence of a protocol-level censorship resistance mechanism, the maximum censorship delay, the geopolitical jurisdictional diversity of validators, the market share of the largest liquid staking provider, and the cumulative concentration of the three largest staking providers — regardless of whether they appear as a liquid staking protocol or as a custodial service — are each measured individually and synthesized in the three-level logic. The 50-percent threshold for OFAC-compliant blocks as the boundary to “Met” goes back to Uri Klarman’s watershed analysis. The 33-percent and 66-percent thresholds for staking concentration are derived from the blocking minority and the supermajority of the Casper-FFG consensus algorithm.
Functional Irreplaceability (I.1) targets depth of entrenchment, measured by network effects, liquidity concentration, and ecosystem lock-in. It explicitly does not ask about substitutability by alternative systems, because the work assesses Ethereum in isolation against its own claim, not against competitors. Three indicators examine it: dominance in DeFi Total Value Locked across all Layer-1 systems, the share of global stablecoin issuance, and the size of the developer ecosystem as measured by public repository activity. The thresholds are set such that dominance above 50 percent reaches “Met” and a drop below 30 percent triggers “Open”. The third indicator — the developer ecosystem — is transparently normatively grounded, because the data source, the Electric Capital Developer Report, measures public repository activity as a proxy for developer network effects: an off-chain metric with greater measurement uncertainty than the on-chain verifiable TVL and stablecoin data.
Coordination Function (I.3) asks whether Ethereum provides a coordination service that could not be rendered without the system, or only at significantly higher cost. The criterion name and the overarching pattern from the infrastructure contextualization are oriented toward the function of real infrastructures: power grids coordinate supply and demand, SWIFT coordinates financial messages between banks, the Internet coordinates packet routing between networks. Three indicators operationalize this question. The first is the coordination scope, which measures the breadth of the coordination service, from pure settlement function to settlement, execution, verification, and data availability. The second is Layer-2 settlement dependency, which measures how many Layer-2 systems use Ethereum as a coordination point. The third is cross-Layer-2 composability, which captures the coherence of the coordination, because Ethereum’s strongest coordination property — atomic composability within a single transaction — is lost at the Layer-2 level and thereby fragments the coordination service.
Long-Term Stability (III.1) asks whether Ethereum can operate over decades without inherent design decisions destabilizing the system. This criterion holds a special position within the Structural Conditions as the most critical long-term condition, because without a solution for State growth and client diversity, long-term load-bearing capacity cannot be guaranteed. The five indicators measure client diversity on the execution layer, client diversity on the consensus layer, the existence of a State growth mitigation solution, the stability of the last five protocol upgrades, and the concentration of the ten largest validators. The boundary to “Met” is formed by the 33-percent threshold of the blocking minority from Casper FFG.
The Qualitative Criteria differentiate the degree of suitability within the category reached. They decide the how, not the whether.
Open Generativity (II.2) asks whether anyone can build and innovate on Ethereum permissionlessly — whether, that is, new applications, protocols, and business models can build on the system without a gatekeeper. The three indicators measure permissionless deployment (can anyone deploy without permission?), atomic composability (can smart contracts interact with one another within a single transaction?), and open-source tooling (are the most important development tools open-source?). The threshold between “Met” and “Met with Qualification” lies at documented Gas limitations and cross-Layer-2 restrictions that fragment composability.
Independent Verifiability (II.3) asks whether the correctness of smart contracts and protocol State can be verified without trust in third parties. This criterion represents the supply side to Security and Trust Load (I.2), whose complementarity was explained there. The indicators measure the availability of formal verification tools, the existence of a formal EVM specification, and the prevalence of security audits as standard practice before the publication of DeFi protocols.
Low-Threshold Inclusivity (II.4) asks whether participation in Ethereum as a user, developer, validator, or verifier is possible without prohibitive barriers. With five indicators, this criterion carries the second-largest indicator set in the evaluation framework alongside Long-Term Stability (III.1), because inclusivity operates at the intersection of several barrier types: the hardware requirements for operating a full node, the hardware requirements for operating a validator, staking access below the 32-ETH threshold, Layer-2 transaction costs compared to conventional transaction costs, and usability without cryptographic prior knowledge. The last indicator is tied to the implementation status of Account Abstraction (EIP-7702) and measures the complexity barrier: can a user who is able to operate online banking and app stores, but has never interacted with cryptocurrencies, use Ethereum? The grounding of this indicator is transparently normative, which means the reader may weight its significance differently.
Adaptive Governance (III.2) asks whether Ethereum can adapt to changed requirements without the adaptation process itself destabilizing the system. The three indicators measure the functionality of the EIP process as a documented governance mechanism, the capacity for hotfixes in the event of critical bugs within 48 to 72 hours, and the ability to resolve contentious decisions without a chain split, which has been empirically demonstrated since the DAO fork of 2016.
Sovereign Portability (III.3) asks whether Ethereum is free from proprietary dependencies that could produce vendor lock-in. The indicators examine whether the complete blockchain history can be exported, whether the system is entirely open-source with multiple independent clients, and whether assets can be moved between layers without trust in individual bridges. The threshold between “Met” and “Met with Qualification” lies at practical barriers to data export and at trust assumptions in bridge design.
Hardware Agnosticism (III.4) asks whether Ethereum can be operated on heterogeneous hardware without specific hardware requirements limiting participation. The three indicators measure cloud independence (no single provider above 50 percent), geographical distribution (nodes in at least ten countries each with more than one percent), and architecture support (all major clients on x86 and ARM). This criterion holds a special position among the twelve. None of the seven reference systems exhibits the specific risk of cloud concentration in the form that exists with Ethereum: the power grid and the road network are physically distributed. The Internet is architecturally decentralized. SWIFT and GPS operate on proprietary infrastructure. Ethereum, by contrast, runs on consumer hardware and cloud servers, and a concentration of node infrastructure on a small number of cloud providers creates a physical dependency that undermines decentralization at the logical layer. The criterion is motivated by the specific claim: it exists because the subject requires it, and it functions as a precondition for several other criteria, because client diversity (III.1) presupposes hardware independence and validator operation on consumer hardware must remain possible in order to secure inclusivity (II.4) and jurisdictional diversity (II.1). The work marks this criterion transparently as a special case whose justification derives from the architectural requirements of the subject, which no analogue in any of the seven reference infrastructures exhibits.
Criteria Hierarchy and Assessment Mechanism (M3 Cascade)
The M3 cascade converts the twelve individual assessments into an overall judgment. The twelve criteria are organized into three hierarchical levels that determine the evaluation logic. At the first level stand the three Critical Conditions (I.2, I.4, II.1), which form the foundation of the infrastructure claim. At the second level stand the three Structural Conditions (I.1, I.3, III.1), which determine the load-bearing capacity of the system. At the third level stand the six Qualitative Criteria (II.2, II.3, II.4, III.2, III.3, III.4), which differentiate the degree of suitability.
The cascade examines in an explicit sequence of five steps. In the first step it is checked whether at least one Critical Condition stands at “Open”. If so, the overall judgment is capped at “Conditionally Suitable”, regardless of all other assessments. In the second step it is checked whether at least one Critical Condition stands at “Conditionally Met”. If so, the overall judgment is limited to “Suitable under Substantial Conditions”. In the third step the Structural Conditions are examined according to the number of their weak ratings. A condition is considered weak if it stands at “Met with Qualification”, “Conditionally Met”, or “Open”. If none of the three Structural Conditions is weak, the top category is attainable, provided the Critical Conditions also stand at least at “Met”. If exactly one stands weak, the judgment is limited to at most “Suitable with Conditions”. If two or more stand weak, it is limited to at most “Suitable under Substantial Conditions”. The capping checks of the first three steps take precedence over the positive definitions of the categories: if a cap applies, it determines the highest possible category, and the positive definition assigns the judgment within that upper limit. In the fourth step the Qualitative Criteria determine the degree within the suitability category reached, according to a degree scale of Good (all six qualitative criteria at least at “Met with Qualification”) through Satisfactory (four or five) and Adequate (two or three) to Poor (fewer than two). In the fifth step the overall judgment is formulated. Identified tensions between criteria enter the cascade by capturing for each tension the mitigation path with its implementation status according to the M1 taxonomy, and the degree of fulfillment of the affected criterion thus determined flows into the responsible step.
The cascade rule in the second step was added during the application of the evaluation framework. In the original version, the cascade provided for only two states for critical conditions: “Open” as the capping threshold and all levels above it as passable. The application of the evaluation framework to the current state identified a situation that the binary distinction did not capture. A Critical Condition at “Conditionally Met” is qualitatively a different state from one at “Met with Qualification”, because the core property of the infrastructure claim is not yet secured at the required level on the protocol side. The hierarchy defines Critical Conditions as the foundation of the infrastructure claim. A foundation at “Conditionally Met” means that the path to fulfillment is recognizable but not secured, while a foundation at “Met with Qualification” means that the property is fundamentally present, with documented limitations. The cascade must reflect this distinction, and the addition does so through an independent judgment category. The addition is formulated generically: it applies to every application of the evaluation framework in which a Critical Condition stands at “Conditionally Met”, regardless of which criterion is affected and which system is being assessed.
The cascade produces five possible judgment categories, whose definitions clarify what a category means and what it excludes. The degree is determined for each category reached by the Qualitative Criteria. It differentiates suitability within the category reached without shifting the category itself. “Suitable” means that all Critical and Structural Conditions stand at least at “Met”. It expressly does not mean operational readiness for deployment, global scaling, or completeness, because these factors describe market and adoption dynamics, not infrastructure properties. “Suitable with Conditions” means that at least one Critical Condition or exactly one Structural Condition exhibits a limitation, as long as no capping rule of a higher level applies. The limitation qualifies the claim without fundamentally undermining it. “Suitable under Substantial Conditions” means that at least one Critical Condition stands at “Conditionally Met”: the core property of the infrastructure claim is not yet secured at the required level, and suitability is tied to named, verifiable conditions. “Conditionally Suitable” means that at least one Critical Condition stands at “Open” and a fundamental requirement is missed, for which, however, a recognizable path to fulfillment exists. “Not Suitable” means that fundamental requirements are missed without a recognizable path to fulfillment.
2.3.4 Scope and Boundaries
The evaluation framework operates within explicit boundaries that define its informative value and demarcate its applicability.
The first boundary is the oracle boundary. The evaluation framework assesses the on-chain state in full — that is, all properties that are measurable within the protocol. What lies outside the protocol cannot be assessed by the evaluation framework in principle. The quality of data that enter the system is a central example: a smart contract can correctly execute the logic of an insurance payout, but if the data input that reports the triggering event is manipulated, the correct execution is of no use. The regulatory environment in which Ethereum operates, and the social acceptance that determines its actual use as infrastructure, are further variables that are not measurable within the protocol and thus lie outside the scope of assessment. The oracle boundary is a restriction in favor of methodological rigor: the evaluation framework assesses what it can assess and documents what it cannot.
The second boundary is the application-layer dividing line. Applications built on Ethereum are treated only insofar as they influence the infrastructure itself. The dividing line follows a question: does the application influence infrastructure suitability? A liquid staking provider whose market share touches the consensus thresholds of the protocol is relevant, because its concentration affects the neutrality of the infrastructure. An NFT marketplace is not. The dividing line is applied where a finding decides its location — for example, at the oracle boundary in Chapter 4.
The third boundary is the current state/target state dichotomy. The work assesses two temporal planes with the same evaluation framework. The current state describes what Ethereum is in the first quarter of 2026, and rests on operational evidence — that is, features with the taxonomy status IMPL. Features with the status DEPL are included insofar as they already shape the running system, but carry no assessment on their own. The target state describes what Ethereum can become on the assumption of complete roadmap implementation, and rests on architectural coherence with graduated implementation maturity — that is, features with the status PLAN or RES. The dichotomy is a methodological decision: Ethereum is an evolving system whose current properties differ from the properties the roadmap holds out. Without the dichotomy, the work would either have to assess the current state and ignore the roadmap — which would obscure the system’s dynamics — or take the target state as given, which would conceal the epistemic uncertainty. The dichotomy allows the reader to distinguish between what is operationally confirmed today and what the roadmap holds out, and to assess the robustness of this undertaking on the basis of the taxonomy stages. It shows that the leap between the current state and the target state is tied to concrete implementations whose robustness Chapter 6 analyzes separately.
Infrastructure suitability is a multidimensional property that resists reduction to a single metric. The work has therefore chosen a qualitative approach. A quantitative assessment would require weightings between the criteria that would either have to be set arbitrarily or statistically derived from a larger sample. Both are not possible for the subject: arbitrary weightings would produce a false precision that conceals the actual uncertainty, and statistical derivation fails because the population of cryptographic infrastructure systems is too small for statistical methods. The completeness of the twelve criteria is reflected on in Section 6.3, where four properties are identified that the evaluation framework cannot measure: permissionless composability, cryptographic system State verification, decentralized self-development capacity, and privacy as an architectural infrastructure property. These four documented boundaries of the evaluation framework mark the point at which validation against the reference material reaches its limit and the subject exhibits properties that none of the seven infrastructures shows in comparable form.
That the criterion assessment rests on argumentative synthesis and does not follow a counting rule was justified in the section on evaluation logic. The methodological justification for this procedure lies in the nature of qualitative application of the evaluation framework: the strength of a rule-based but interpretive assessment lies in the theory-guided weighing of the indicators, not in their mechanical aggregation into a point value. The transparency of the procedure enables the reader to independently trace each individual assessment on the basis of the documented evidence and the reasoned conclusion, and, if desired, to arrive at a divergent judgment. This dual goal — introduced in the chapter introduction — is a methodological strength that defuses a frequent criticism of qualitative research: the work is useful even if the reader does not agree with the judgment, because it provides the foundation for an independent judgment.
A final methodological boundary is the survivorship bias in the validation base: the seven reference infrastructures are recognized systems. Failed ones are not included, and the validation therefore supports only the reading of the twelve criteria as necessary conditions. Section 3.3 elaborates on this boundary. Chapter 3 documents the validation of the twelve criteria and the methodological boundaries of this procedure. The complete individual analyses of the seven infrastructures are in Section 3.1, the consolidation and validation results in Section 3.2.
- Frischmann, Brett M. (2012): Infrastructure: The Social Value of Shared Resources. Oxford University Press. ↩
- Frischmann's argument builds on the empirically substantiated finding of Elinor Ostrom (1990) that communities can successfully manage shared resources themselves, without privatization or state control. Ostrom's eight design principles for sustainable commons governance are applicable to Ethereum's governance model. A systematic application to blockchain systems has been undertaken by Rozas et al. (2021). — Ostrom, Elinor (1990): Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press. — Rozas, David / Tenorio-Fornés, Antonio / Díaz-Molina, Silvia / Hassan, Samer (2021): When Ostrom Meets Blockchain: Exploring the Potentials of Blockchain for Commons Governance. SAGE Open, 11(1), pp. 1–14. DOI: 10.1177/21582440211002526. ↩
- Grimmelmann, James / Windawi, A. Jason (2023): Blockchains as Infrastructure and Semicommons. William & Mary Law Review, 64(4), pp. 1097–1129. ↩
- Cf. also Dimitropoulos, Georgios (2020): The Law of Blockchain. Washington Law Review, 95(3), pp. 1117–1200, who conceives of blockchain as "infrastructural commons". ↩
- Star, Susan Leigh (1999): The Ethnography of Infrastructure. American Behavioral Scientist, 43(3), pp. 377–391. ↩
- Hanseth, Ole / Lyytinen, Kalle (2010): Design Theory for Dynamic Complexity in Information Infrastructures: The Case of Building Internet. Journal of Information Technology, 25(1), pp. 1–19. ↩
- Tilson, David / Lyytinen, Kalle / Sørensen, Carsten (2010): Research Commentary — Digital Infrastructures: The Missing IS Research Agenda. Information Systems Research, 21(4), pp. 748–759. ↩
- Mayntz, Renate (1993): Große technische Systeme und ihre gesellschaftstheoretische Bedeutung. Kölner Zeitschrift für Soziologie und Sozialpsychologie, 45(1), pp. 97–108. ↩
- Davidson, Sinclair / De Filippi, Primavera / Potts, Jason (2018): Blockchains and the Economic Institutions of Capitalism. Journal of Institutional Economics, 14(4), pp. 639–658. DOI: 10.1017/S1744137417000200. ↩
- Ølnes, Svein / Jansen, Arild (2018): Blockchain Technology as Infrastructure in Public Sector: An Analytical Framework. In: Proceedings of the 19th Annual International Conference on Digital Government Research (dg.o '18). ACM, pp. 1–10. DOI: 10.1145/3209281.3209293. ↩
- Ølnes, Svein / Jansen, Arild (2021): Blockchain Technology as Information Infrastructure in the Public Sector. In: Reddick, C.G. / Rodríguez-Bolívar, M.P. / Scholl, H.J. (eds.): Blockchain and the Public Sector. Public Administration and Information Technology, vol. 36. Springer, Cham, pp. 19–46. DOI: 10.1007/978-3-030-55746-1_2. ↩
- Nabben, Kelsie (2023): Web3 as 'Self-Infrastructuring': The Challenge Is How. Big Data & Society, 10(1). DOI: 10.1177/20539517231159002. ↩
- Dylan-Ennis, Paul / Kavanagh, Donncha / Araujo, Luis (2023): The Dynamic Imaginaries of the Ethereum Project. Economy and Society, 52(1), pp. 87–109. DOI: 10.1080/03085147.2022.2131280. ↩
- Srinivasan, Balaji S. / Lee, Leland (2017): Quantifying Decentralization. Preprint/Blogpost. URL: https://news.earn.com/quantifying-decentralization-e39db233c28e. ↩
- Gencer, Adem Efe et al. (2018): Decentralization in Bitcoin and Ethereum Networks. In: Financial Cryptography and Data Security. Springer, pp. 439–457. ↩
- Ovezik, Christina / Karakostas, Dimitris / Milad, Mary / Kiayias, Aggelos / Woods, Daniel W. (2025): SoK: Measuring Blockchain Decentralization. arXiv:2501.18279. Also published in: Fischlin, M. / Moonsamy, V. (eds.): Applied Cryptography and Network Security (ACNS 2025). Lecture Notes in Computer Science, vol. 15825. Springer, Cham. DOI: 10.1007/978-3-031-95761-1_7. ↩
- Spencer-Hicken, Scott et al. (2023): Blockchain Feasibility Assessment: A Quantitative Approach. Frontiers in Blockchain, 6. DOI: 10.3389/fbloc.2023.1155125. ↩
- Jensen, Johannes R. / von Wachter, Victor / Ross, Omri (2021): How Decentralized is the Governance of Blockchain-based Finance: Empirical Evidence from four Governance Token Distributions. In: Proceedings of the 29th European Conference on Information Systems (ECIS 2021). ↩
- Pincheira, Miguel et al. (2023): An Infrastructure Cost and Benefits Evaluation Framework for Blockchain-Based Applications. Systems, 11(4), 184. DOI: 10.3390/systems11040184. ↩
- Jabareen, Yosef (2009): Building a Conceptual Framework: Philosophy, Definitions, and Procedure. International Journal of Qualitative Methods, 8(4), pp. 49–62. ↩