Quantitative stress model · July 2026

Agentic Execution Risk

A logistic growth model with differentiated error rates and market saturation for autonomous AI agent proliferation.

Reject the assumptions. Replace them with better ones. Then run the model again.
VALO Research Group · Njål Gaute Solland, Project Lead · Agentic Execution Risk Observatory
Purpose

A stress model, not a reassurance model.

The report tests whether autonomous execution may scale faster than the systems designed to govern it. It uses a logistic population model rather than unconstrained exponential growth, and explicitly exposes the assumptions that drive the result.

The model is designed to be challenged. Growth rate, carrying capacity, failure differentials and under-reporting factors are assumptions, not facts.

Baseline model

Four assumptions drive the stress surface.

20M baseline agentsBottom-up estimate across developer, enterprise, trading and consumer tool-use populations.
1B carrying capacityA logistic ceiling representing a global addressable operator population, not an unconstrained exponential projection.
0.01% simple-agent failure rateIllustrative quarterly critical-failure rate for single-action systems.
0.30% multi-agent failure rateIllustrative 30× differential reflecting chained complexity, permissions and cascading failure.
Scenarios

Population growth and complexity interact.

The report evaluates doubling periods of 7, 14, 30, 60 and 90 days across 3-, 6-, 9- and 12-month horizons. Short doubling periods quickly hit the logistic ceiling, while slower scenarios remain materially below saturation.

The key structural claim is that incident exposure is driven by both agent count and the rising share of multi-agent workflows. The model therefore increases the weighted failure rate as multi-agent share grows.

30d
Illustrative central scenarioThe report treats a 30-day doubling period as its planning baseline.
30×
Complexity differentialMulti-agent workflows are modeled as materially riskier than simple agents.
Evidence discipline

Observed incidents are treated as a floor.

The report uses public incident data as the visible minimum rather than as a complete measure of real-world failures. It explicitly models an under-reporting range and separates documented observations from inferred totals.

That distinction matters: the model should not be read as a precise forecast of future incidents. It is a sensitivity analysis showing how execution risk behaves when agent population, autonomy and workflow complexity grow together.

Operational implication

Governance capacity has to scale before consequence.

The report argues for bounded authority, human review for high-consequence multi-agent workflows, reversibility where technically possible, sandbox testing and better incident reporting.

Status: quantitative stress model and research artifact. Scenario probabilities, carrying capacity, error rates and under-reporting factors are assumptions requiring continued empirical revision. The underlying artifact points to a public model repository and raw-data workflow for challenge and rerun.