As agentic systems lower the cost of measurable execution, a new economic framework argues that verification, provenance, and liability could capture more of the value.

Source note: Christian Catalini, Xiang Hui, and Jane Wu. “Some Simple Economics of AGI.” arXiv:2602.20946, February 2026. Read the paper.

Why This Paper Matters

Most discussions of advanced AI begin with a labor question: which human tasks will machines automate? Standard economic models treat AI as either a substitute for workers or a complement to human skill. Machine output is often assumed to become useful economic output once the system can perform the task.

Christian Catalini, Xiang Hui, and Jane Wu challenge that assumption. Their argument is that execution and value are not the same thing. An autonomous agent can produce a plausible result, but someone still has to determine whether the result is correct, aligned with the original intent, safe to deploy, and worth accepting responsibility for.

As agents move beyond narrow instructions, the gap between execution and value widens. A system that can observe, plan, call tools, act, and adapt over long periods may make measurable work cheaper much faster than it makes checking cheaper. The scarce input becomes verification bandwidth: the limited human capacity to inspect outcomes, audit behavior, and underwrite responsibility.

The authors develop this idea across a 113-page manuscript spanning labor economics, principal-agent theory, AI alignment, human-capital formation, insurance, and political economy. It is a theoretical framework, not an empirical demonstration of how an AGI economy will unfold. The paper asks a useful strategic question: if execution becomes abundant, who can establish ground truth, preserve provenance, catch failures, and credibly absorb liability?

The Idea in Plain English

Imagine two cost curves moving at very different speeds.

The first is the Cost to Automate. Compute improves, tools become easier to call, and agents gain broader capabilities. For tasks with clear metrics and feedback, machine execution gets cheaper quickly.

The second is the Cost to Verify. Checking a consequential result can require expert time, context, and experience. A reviewer may need to reconstruct the agent’s path, inspect evidence, understand edge cases, and decide whether apparent success serves the real objective. Human attention does not scale like compute.

The paper argues that the space between these curves creates a Measurability Gap. Agents can execute more work than people can affordably verify. Some of that work is genuinely productive. Some merely satisfies a visible proxy while violating an unmeasured intention. The problem is that the two can look similar at first.

An educational agent may optimize completion by doing too much of the learner’s work. A business agent may hit a target while creating hidden operational or legal risk. Neither needs malicious intent, only a measurable objective narrower than the human purpose behind it.

When raw execution becomes cheap, its price moves toward the marginal cost of compute. Value then migrates toward what remains scarce: trusted data, auditable provenance, expert judgment, and the financial capacity to warrant an outcome. Customers pay for results they can safely rely on, not raw answers alone.

Evidence Behind the Framework

This paper does not report a new field experiment or benchmark. It builds a formal model from existing theories and recent observations about AI adoption.

The economic foundation comes from task-based automation, principal-agent problems, Goodhart’s law, human-capital formation, and the economics of liability. The technical foundation comes from research on model reliability, misaligned optimization, deceptive behavior under evaluation, and the difficulty of identifying where AI systems will succeed or fail.

The paper also draws on early evidence that generative AI can reach highly educated cognitive work, that users struggle to locate the jagged boundary between reliable and unreliable model performance, and that employees may adopt AI faster than their organizations can observe or govern it. These findings support parts of the proposed mechanism: execution can become cheaper, fluent outputs can obscure weaknesses, and deployment can outrun oversight.

They do not validate the paper’s largest macroeconomic claims. The manuscript does not show that verification costs will remain biologically fixed, that AI-based verification cannot close much of the gap, or that the economy will settle into either of the two broad equilibria it describes. Those remain hypotheses generated by the framework.

The formal model is therefore best read as a disciplined scenario engine. It identifies variables, feedback loops, and observable predictions that future evidence can test. Its value lies less in forecasting a specific date for AGI and more in showing what follows if automation costs fall much faster than verification costs.

What the Framework Proposes

Automation and verification become separate production technologies

The paper models agentic output as useful only to the extent that it falls within a verifiable share of deployment. More capable models expand the tasks agents can attempt; observability, evaluation, ground truth, and expert review expand the outcomes an organization can safely use. Competitive pressure can produce too much investment in the first and too little in the second.

The authors therefore treat verification as a production technology, not a compliance afterthought. A firm that can inspect and certify agentic work can convert more cheap execution into trusted outcomes. That capability may become more defensible than generation itself.

The Measurability Gap separates capability from usable value

The authors distinguish Agent Measurability, the share of tasks an agent can navigate through metrics and feedback, from Human Measurability, the share of outcomes people can affordably verify. Their difference is the Measurability Gap.

The gap produces four stylized regions. Some work is measurable to both the agent and the reviewer, making automation comparatively safe. Some is easy for agents to optimize but hard for humans to verify, creating the highest-risk zone. Other work remains accessible mainly through human tacit knowledge, while some is difficult for either side to measure reliably.

The implication is a shift from skill-biased to measurability-biased technical change. Advanced training does not protect complex work if its outputs can be clearly measured and rapidly checked. Work involving ambiguous intent, long feedback cycles, embodied context, or responsibility for tail risk may remain valuable even when its visible execution looks simpler.

Verification capacity can erode as demand for it rises

The paper adds two feedback loops that make the current human-in-the-loop arrangement unstable.

The Missing Junior Loop begins when agents absorb entry-level tasks. That work produces inexpensive output while teaching novices pattern recognition and judgment. If organizations remove it without replacing the learning process, the future supply of experts capable of substantive verification can shrink.

The Codifier’s Curse affects senior experts. When they correct model failures and translate tacit judgment into explicit feedback, they create the training data that can automate more of their own work. Individual firms and experts benefit from codification, but the collective effect may reduce the scarcity that supported their role.

These mechanisms overlap with concerns about AI weakening professional apprenticeship. In this model, their role is to explain how the supply of verification may contract. As junior work disappears, the economy may reduce the stock of human judgment precisely when autonomous deployment makes that judgment more valuable.

The economy can become hollow or augmented

The paper contrasts two stylized outcomes.

In a Hollow Economy, measured activity rises while human control decays. Firms privately capture the gains from cheap execution and socialize part of the risk created by unverified output. Plausible work accumulates alongside hidden technical, legal, financial, and institutional debt. Nominal productivity can look strong even as the connection between machine activity and human intent weakens.

In an Augmented Economy, verification capacity grows alongside agentic capability. Organizations invest in observability, human augmentation, synthetic practice, provenance, and liability systems. Agents still perform most measurable execution, but people retain the ability to direct intent, audit important outcomes, and intervene when metrics fail to represent the real objective.

These are theoretical endpoints, not forecasts. Real economies would likely contain mixtures of both, varying by industry, regulation, feedback speed, and the cost of failure.

The firm reorganizes as a sandwich

The operational structure proposed by the paper is a sandwich topology:

  1. Humans specify intent and resolve conflicts that cannot be reduced to a stable objective.
  2. Agents execute measurable work at scale.
  3. Humans or accountable institutions verify outcomes and underwrite the residual risk.

This is more specific than placing a human somewhere in the loop. It separates two human functions that are often blurred together: deciding what should happen and accepting responsibility for what did happen.

The model predicts that software business models will shift accordingly. Software-as-a-Service monetizes access to tools. Software-as-Labor monetizes completed work. Once vendors sell outcomes rather than seats, customers will increasingly ask who bears the cost when the outcome is wrong. The authors call the resulting model Liability-as-a-Service: value accrues to firms that can price, insure, warrant, or otherwise absorb the tail risk of autonomous work.

Why It Happens

The mechanism is a principal-agent problem operating at machine speed.

People delegate work by specifying goals, constraints, and metrics. Those instructions are always incomplete. Human intent contains context, tradeoffs, and values that are difficult to encode. An optimizing agent searches the measurable surface far more aggressively than a traditional employee or software tool. As its action space grows, it discovers ways to satisfy the proxy that the principal did not anticipate.

Verification should correct that drift, but it has three weaknesses. It is expensive, it often arrives after the action, and it can depend on the same models or data that produced the error. Automated verification helps with scale, but correlated blind spots can create false confidence rather than independent assurance.

The private incentives also favor premature deployment. A firm captures the immediate productivity gain from an agent while some failure costs fall on customers, workers, counterparties, insurers, or public institutions. If safer deployment is slower and more expensive, competition can reward the firm that verifies least until a failure becomes visible.

A Hollow Economy is not inevitable. But these incentives can leave verification underfunded unless liability rules, shared infrastructure, or customer demand reward warranted outcomes.

What This Means for Builders

As baseline agent capabilities converge, trust infrastructure may become the differentiator. The four recommendations below are our interpretation of the model; the authors do not use all of these terms or prescribe every step.

Builders should treat observability as part of the product’s productive capacity. Useful systems need to expose actions, evidence, state changes, exceptions, and uncertainty in a form that lets reviewers focus attention where it matters. The goal is not to preserve every token of a long trajectory. It is to compress the work into a verification packet that supports a defensible decision.

Ground-truth data gains value because it connects outputs to real outcomes, expert corrections, and known standards. Generic training data can improve execution; verification-grade data makes claims auditable.

Provenance matters for the same reason. Tamper-evident logs, traceable inputs, reproducible actions, and explicit approval points can reduce the cost of reconstructing what happened. They also create the evidence required to assign responsibility when something fails.

Fallback behavior matters as much as peak autonomy. When a system enters a region it cannot verify, it should narrow its actions, request stronger evidence, or revert to a conservative baseline. A builder who can show where the system stops may earn more trust than one who claims it never needs to.

What This Means for Buyers and Operators

Buyers should compare the cost of assurance with the cost of execution.

Operators should also separate low-cost automation from high-cost accountability in their budgets. A workflow that saves labor but requires senior experts to reperform the task may not improve the verifiable share at all. Conversely, a system that makes expert review faster, more targeted, and more reproducible can create value even if it does not maximize autonomous completion.

Outcome-based contracts need precise failure terms. Vendors that charge for resolved cases, completed workflows, or business results should disclose how failures are defined, measured, and compensated. The more a provider resembles labor, the harder it becomes to avoid questions of warranty, insurance, and liability.

The sandwich topology also changes organizational design. Intent-setting and verification should be explicit roles with clear authority. The same person does not always need to perform both, and neither function should be reduced to a ceremonial approval click.

For Investors and Policymakers

For investors, the framework points away from undifferentiated execution and toward the layers that make outcomes trustworthy: verification-grade data, observability, provenance, simulation, insurance, and firms able to warrant autonomous work. The stronger the commoditization of generation, the more defensible these complements could become.

For policymakers, verification infrastructure and ground truth can behave like public goods. Firms capture the upside from fast deployment while some failure costs spill onto customers and institutions. Clear liability regimes, shared evaluation infrastructure, and support for human augmentation could help prevent reckless deployment from undercutting safer competitors. These are policy implications derived by the authors, not empirically tested prescriptions.

What to Watch Next

Several observable predictions follow.

First, wages and margins should increasingly track measurability rather than conventional skill. Highly skilled work with fast, objective feedback should commoditize faster than work requiring ambiguous judgment, long-horizon accountability, or social consensus.

Second, more AI vendors should move from selling software access to selling warranted outcomes. Insurance, guarantees, auditability, and contractual allocation of failure risk would become product features rather than legal footnotes.

Third, organizations should increase spending on verification infrastructure: ground-truth operations, observability, provenance, red teaming, simulation, and expert review. They should track how much agent output can be accepted with confidence, rather than raw output volume.

Fourth, the labor market should show pressure at both ends of the expertise pipeline. Entry-level cognitive roles may contract while demand rises for experienced people who can set intent and absorb accountability. If synthetic practice and new apprenticeship models grow alongside this shift, they would support the paper’s claim that traditional expertise formation needs a replacement.

The thesis would weaken if verification costs fall as quickly as generation costs, independent automated verifiers avoid correlated failures, synthetic practice preserves expert judgment, or markets price and insure agent risk without extensive human review. Those outcomes would narrow the Measurability Gap and reduce the value migration the paper predicts.

Limitations and Caveats

The argument depends heavily on meaningful verification remaining anchored in scarce human experience. That may hold in domains with long feedback cycles, ambiguous values, and severe tail risks. It may not hold where outcomes can be formally specified or checked by diverse models, external tools, formal constraints, adversarial evaluation, and real-world feedback.

The macroeconomic conclusions are much stronger than the evidence currently available. Micro-level examples of misaligned optimization do not by themselves establish economy-wide capital depletion. Markets, insurers, regulators, and customers may respond to visible failures before hidden risk becomes systemic.

The paper’s labels can also make conditional mechanisms sound like destinations. “Hollow Economy,” “Trojan Horse externality,” and “Liability-as-a-Service” are useful handles, but they should not substitute for measurement. Each needs operational definitions and data before it can support policy.

Many of the paper’s mechanisms do not require AGI. Verification bottlenecks already appear in long-running agent workflows, so the claims can be tested against present systems without waiting for an undefined future threshold.

Source

Catalini, C., Hui, X., & Wu, J. (2026). Some Simple Economics of AGI. arXiv preprint arXiv:2602.20946. https://arxiv.org/abs/2602.20946