Research library · Full memo
Locally Owned Reasoning Infrastructure: Economics and Institutional Implications
This memo reviews evidence current through 10 August 2026 on economics, organizational design, assurance, interoperability, and institutional structure. It is an analysis of system implications, not a market forecast or investment thesis.
Executive assessment
Locally owned reasoning infrastructure is best understood as a change in the boundary of the firm, household, or professional practice—not as a claim that all computation should move onto a personal device. It moves some context, memory, workflow state, policy, and execution closer to the person or institution that is accountable for them. Models and compute can remain local, remote, or hybrid. The economic question is therefore not simply “local model or cloud API?” It is which assets must remain portable and governable, which execution should be bought as a service, and where the costs of integration, verification, and switching should sit.
The stronger AIOS thesis examined in this memo is that increasingly capable local models, mature scaffolding, durable local context, and selective access to frontier models could satisfy a large share of ordinary personal and organizational reasoning needs. If substantially correct, the first-order effect would be to move accumulated context, expertise, memory, workflow, and decision authority from remote applications into person- or institution-owned systems. The second-order effects could include thinner centralized application layers, less routine remote inference, more private cognitive work, widespread domain-specific reasoning systems, new peer knowledge institutions, and broader access where bandwidth or cloud services are constrained. These are conditional implications—not current measurements—but lack of present system-level evidence is not evidence that the implications are negligible.
AIOS also treats useful intelligence as a system-level outcome rather than a property of model weights alone. Purpose framing, composed context, distinct cognitive modes, layered memory, relational metadata, provenance, bounded judgment, planning, verification, and post-response integration may each supply modest gains that compound into fewer retries, lower inference burden, and lower hallucination exposure per accepted task. Dynamic routing could then assign a local or frontier model according to the cognitive operation, capability need, privacy, latency, cost, and sensitivity. Present research measures several components of this mechanism separately; it does not establish their cumulative effect in one system. Economically, that unmeasured complementarity is central because it could either offset agentic rebound or, if the scaffold adds calls and coordination failures, intensify it.
The strongest empirical result is a divergence between unit price and task cost. At a fixed benchmark capability, public API prices fell extraordinarily quickly through 2024. Stanford’s 2025 AI Index reported that the cheapest system exceeding a GPT-3.5-level MMLU threshold fell from about $20 to $0.07 per million tokens between November 2022 and October 2024. An Epoch AI reconstruction found rates ranging from roughly 9-fold to 900-fold per year depending on task and performance threshold. These are retail-price/benchmark measurements, not estimates of provider cost, and they exclude most of the orchestration surrounding a useful organizational task.
Agentic systems create a plausible rebound mechanism. A software agent may call a model repeatedly, read and reread long contexts, invoke tools, branch, recover from failures, and ask stronger models to review weaker ones. Recent preprints report orders-of-magnitude differences in token use across agent designs and runs. One benchmark study found tool-using designs making 9.2 times as many model calls as a chain-of-thought baseline on average; another found agentic coding trajectories using about three orders of magnitude more tokens than single-turn code tasks. These findings are domain- and scaffold-specific, and they do not establish an economy-wide rebound elasticity. They do establish that falling price per token cannot be treated as falling cost per completed, verified task.
On-device inference is now technically credible for bounded workloads. Peer-reviewed systems work demonstrates large speed and energy gains from neural-processing-unit execution of small models on a narrow set of phones. Standardized client benchmarks now include 1B–8B parameter language models. But there is not yet a general empirical basis for claiming that local inference is always cheaper, greener, more private, or more reliable. The relevant comparison includes amortized hardware, engineering and update costs, thermals, battery wear, utilization, model capability, human review, and the avoided costs of latency, network dependence, data transfer, and compliance. Repeated, stable, privacy-sensitive tasks on already-owned hardware are the most favorable local case; bursty frontier workloads and high-throughput shared services are the most favorable remote case. Hybrid execution is the defensible default scenario.
Upstream infrastructure remains concentrated even when downstream knowledge is locally owned. The International Energy Agency estimates that data centers consumed about 415 TWh in 2024 and projects roughly 945 TWh in 2030 in its base case. Its 2026 update estimates growth of 17% in global data-center electricity use during 2025 and about 50% for AI-focused facilities, while projecting about 950 TWh by 2030. The United States, China, and Europe account for most current demand, and new facilities cluster at particular grid nodes. The US Federal Trade Commission, UK Competition and Markets Authority, and OECD all document capital intensity, hyperscaler concentration, technical and commercial switching barriers, and contractual ties between cloud firms and model developers. Local ownership of canonical files and accepted ground can improve bargaining position and continuity, but it does not by itself decentralize semiconductor fabrication, frontier training, cloud capacity, or electricity demand.
Switching costs have at least three layers. Technical switching costs arise from data movement, interfaces, orchestration, identity, and provider-specific services. Behavioral switching costs arise when prompts, tool schemas, evaluation thresholds, and error-handling routines must be retuned for a different model or version. Semantic switching costs arise when an institution’s meanings—annotations, exceptions, relationships, provenance, tacit categories, review rules, and authority boundaries—are embedded in a provider’s representations or workflows. Regulators and interoperability protocols address the first layer most directly. A 2026 peer-reviewed experiment showing that user prompt adaptation accounted for about half the performance gain from a model upgrade on a bounded image-replication task is evidence for behavioral complementarity, not a complete measure of semantic lock-in. “Semantic switching cost” should therefore be treated as a useful analytic construct whose measurement remains immature.
Productivity evidence argues against both frictionless adoption and categorical pessimism. Randomized field experiments report material gains on selected tasks: a pooled 26% increase in completed tasks across three developer experiments; faster and higher-quality work for consultants on tasks inside a model’s capability frontier; and individual product-development workers with AI reaching the judged output quality of unaided two-person teams in one firm experiment. The same research also finds uneven adoption, larger gains for less-experienced workers, performance losses outside the capability frontier, and limited change in meeting behavior or coordinated work. Danish administrative data show little detectable effect on earnings or hours in the first two years despite adoption. The coherent interpretation is that AI is often a complement to task redesign, evaluation, training, and organizational capital. It is not yet an observed autonomous substitute for those complements.
Local custody can make audit evidence easier to retain, inspect, and migrate, but auditability is a property of the whole sociotechnical process. Regulated assurance requires versioned inputs, sources, model and configuration identifiers, tool calls, file changes, identities, authorizations, tests, exceptions, approvals, and post-deployment monitoring. Neither locality nor a transcript proves correctness, fairness, safety, or compliance. Recent NIST work finds deployed-AI monitoring fragmented across operational, human, security, compliance, and societal concerns; the EU AI Act and sectoral guidance impose or propose lifecycle documentation, logging, human oversight, risk management, and monitoring obligations. A local-first architecture can support these obligations, but adjacent standards and studies do not validate any particular implementation.
Interoperable protocols are promising but incomplete institutional infrastructure. The Model Context Protocol standardizes a way for model applications to expose tools, resources, and prompts. Agent2Agent addresses capability discovery and task exchange between independent agents. W3C Verifiable Credentials, RO-Crate, C2PA, and Sigstore offer building blocks for issuer claims, research-object packaging, content provenance, and signed software attestations. Together they make an attestable domain package technically plausible: ordinary files plus a content-addressed manifest, provenance, schemas, licenses, tests, policy constraints, signatures, and issuer credentials. They do not make the package true, current, safe, or “certified” unless an accountable certification scheme defines claims, test methods, issuers, surveillance, revocation, and liability.
The “AI as electricity” analogy is useful only in a narrow economic sense. Both can be pervasive enabling inputs whose productivity contribution depends on complementary capital and organizational redesign; both create large infrastructure and grid effects. The analogy breaks at the point of the delivered service. Within a specified grid product and quality standard, electricity is metered as a largely fungible commodity. A reasoning output is probabilistic, versioned, context-dependent, and entangled with a user’s data, policies, tools, and authority. Electricity does not silently change its semantics after a provider update. AI is therefore closer to an evolving cognitive production system that runs on concentrated digital and electrical infrastructure than to electricity itself.
1. Method and evidence discipline
This memo prioritizes work published or materially updated from August 2024 through August 2026. It emphasizes peer-reviewed studies, official statistics, regulator investigations, and adopted standards. Preprints, working papers, vendor-authored studies, draft guidance, and scenarios are labeled. Older work appears only in a separate foundational-lineage section.
Five analytic labels are used:
- Observed measurement: a historical price, metered energy value, administrative record, field-experiment outcome, benchmark result, or regulator evidence about actual conduct.
- Modeled estimate: a value reconstructed from measured inputs through an explicit model, such as bottom-up server-energy accounting.
- Scenario or projection: a conditional forward result. It is not a measurement of what will occur.
- Interpretation or design hypothesis: synthesis developed in this memo. It may be plausible without having been tested as a complete system.
- Conditional implication: a first- or second-order consequence that would follow if a stated architectural thesis is substantially correct. It is neither a present measurement nor a forecast by itself.
Replication note. Peer review is not replication. No empirical result below is treated as independently replicated unless the text says so. Results based on one paper, one intervention, one firm, or one benchmark suite are identified in their limitation paragraphs; where a finding is especially easy to overgeneralize and no independent replication was located, that absence is stated directly.
Benchmark results are not silently generalized to institutions. A model score is not a completed-work outcome; a completed-work count is not quality-adjusted productivity; an energy estimate per query is not a provider’s fleet total; a signed artifact is not verified knowledge; and a regulation is not evidence that compliance has been achieved.
2. The economic object: ownership, execution, and complements
“Locally owned reasoning infrastructure” combines several choices that should be analyzed separately:
- Custody: who holds the canonical files, memory, annotations, provenance, policies, and relationship data.
- Control: who can inspect, version, authorize, modify, export, or retire those assets.
- Execution location: whether a particular inference or tool operation runs on a client device, local server, private cloud, or public service.
- Model sourcing: whether model weights are locally controlled, remotely served, or selected dynamically.
- Assurance: how behavior, changes, exceptions, and human decisions are evaluated and recorded.
- Interoperation: how tools, agents, and knowledge packages exchange capabilities and evidence.
These dimensions are often conflated. Canonical files and accepted ground can remain person-controlled while a remote model performs a difficult inference. Conversely, a local model can still create lock-in if the surrounding memory store, embeddings, schemas, or workflow engine cannot be exported intelligibly. “Local-first” is thus primarily a claim about default custody, inspectability, and continuity; it is not synonymous with offline-only inference or self-sufficiency from industrial compute.
The relevant unit of economic analysis is the accepted task outcome, not the token. Model weights are one productive input alongside context quality, memory, metadata, tools, workflow design, verification, exact integration, and human judgment. A weaker or smaller model embedded in a better system may have a lower cost per accepted outcome than a stronger model requiring repeated calls and rework; the reverse may hold when scaffolding overhead exceeds its gains. A minimal comparison is:
Effective cost per accepted task = compute and service charges + amortized hardware + engineering and operations + network and movement costs + evaluation and human review + expected failure/rework + compliance and continuity costs.
For local execution, the dominant term may be a sunk device at high utilization, or it may be scarce engineering time maintaining model variants across heterogeneous hardware. For remote execution, the token bill may be small while retained-data review, outage exposure, contractual restrictions, or provider-specific workflow dependencies dominate. Any comparison that omits the acceptance criterion and review burden is incomplete.
3. Inference prices, agentic demand, and the rebound question
3.1 What has been measured
Research question and method. How quickly has the posted price of a minimum level of model capability fallen? The Stanford AI Index 2025 uses public API prices and benchmark thresholds to compare the cheapest available model at a fixed capability. It reports that the price of a system scoring at least 64.8 on MMLU fell from $20 per million tokens in November 2022 to $0.07 in October 2024, a decline greater than 280-fold. For a greater-than-50% threshold on GPQA, the minimum observed price fell from $15 per million tokens in May 2024 to $0.12 in December 2024. These are unusually rapid declines.
The underlying Epoch AI analysis combines public prices with benchmark observations, assumes a 3:1 mix of input to output tokens, selects the cheapest model above each task-specific threshold, and fits log-linear trends. Across six benchmark/threshold combinations, it estimates declines from roughly 9-fold to 900-fold per year, with a median around 50-fold over the full observation period and faster estimates after January 2024. Epoch makes code and observations available, which improves inspectability.
Limits and relevance. These series measure posted retail prices at selected benchmark levels. They do not identify marginal provider cost, discounts, capacity reservations, batching, caching, latency, reliability, output length, orchestration, or review. Benchmark coverage is partial and vulnerable to contamination and saturation. Reasoning models were excluded from the initial Epoch reconstruction. Selecting a moving “cheapest adequate model” also combines engineering progress, new entry, and pricing strategy. The series is directly relevant to the price of substitutable model calls, but only indirectly relevant to the cost of an accepted organizational result.
A later preprint, Gundlach et al., “The Price of Progress”, asks how fixed-performance inference price changed and how much of the change can plausibly be separated into competition, hardware, and algorithmic efficiency. It reconstructs performance-adjusted price trends with a larger historical combination of Artificial Analysis pricing and Epoch benchmark data. It estimates price declines of approximately 5–10-fold per year for fixed performance in knowledge, reasoning, mathematics, and software tasks. Restricting parts of the analysis to openly available models and accounting for hardware-price improvement, the authors estimate algorithmic efficiency gains near three-fold per year.
Limits and relevance. This is a preprint. Its decomposition depends on model coverage, benchmark comparability, posted prices, hardware assumptions, and fitted functional forms. It also distinguishes cheap fixed performance from frontier performance: the cost of evaluating or serving the most capable available system need not fall on the same trajectory. The decomposition is relevant to long-run sourcing assumptions, but is too uncertain to serve as a budgeting forecast.
3.2 Cost of a correct answer
Price per token is a poor sufficient statistic when models vary in accuracy and verbosity. The preprint Erol et al., “Cost-of-Pass”, defines the expected monetary cost of obtaining a correct solution and compares models and inference strategies on quantitative and knowledge benchmarks. It finds that model rankings change when accuracy is combined with price, and that majority voting or self-refinement often supplies too little marginal accuracy to justify its added cost in the evaluated settings.
Limits. “Correct” is benchmark-defined and often easier to determine than acceptance in professional work. API prices and model behavior change rapidly. The metric does not include latency, integration, review, or downstream damage from an accepted error. Its main relevance is conceptual: even a narrow task should be priced per accepted outcome rather than per token.
3.3 Agentic token demand
The preprint Kim et al., “The Cost of Dynamic Reasoning”, compares chain-of-thought, ReAct, Reflexion, LATS, and LLMCompiler agent designs using Llama 3.1 8B and 70B models on question-answering and interactive benchmarks. Tool-using agents made 9.2 times as many model calls as the chain-of-thought baseline on average; LATS averaged 71 calls per request. On HotpotQA, selected configurations increased measured GPU energy from 0.32 Wh for a single 8B call to 22.76 Wh for LATS and 41.53 Wh for Reflexion; corresponding 70B figures rose from 2.55 Wh to 158.48 Wh and 348.41 Wh. Accuracy improvements showed diminishing returns.
Method and sample. The authors run five designs on A100 GPUs, measure power with NVIDIA tooling, and use 50-question samples for key design comparisons. They extrapolate several illustrative fleet scenarios.
Limits and relevance. This is a preprint using older model families, selected configurations, no production batching, and GPU-only energy. Its fleet totals are scenarios, not observed consumption. The measured per-workflow result is nevertheless important: architecture can change energy and call counts by two orders of magnitude even before a task leaves the benchmark.
The preprint Bai et al., “How Do AI Agents Spend Your Money?”, analyzes eight frontier models on 500 SWE-bench Verified software-maintenance tasks. It reports that agentic coding uses roughly a thousand times the tokens of code-reasoning or chat-style requests, that repeated runs of the same task can differ by up to 30-fold in token use, and that higher consumption does not monotonically improve success. Models are also poor at predicting their own eventual cost.
Limits and relevance. The result is specific to software maintenance, the selected scaffold, benchmark, and current model behavior. Tokens are not identical to billed tokens when caches and discounts apply, and repository-scale contexts are unusually large. It is direct evidence of high within-task cost dispersion and weak self-budgeting, not an estimate of typical office-agent use.
3.4 Rebound: supported mechanism, unmeasured elasticity
The rebound hypothesis is that cheaper inference encourages longer contexts, more attempts, larger models, more users, more background agents, and more verification, so aggregate resource use grows even as unit cost falls. The IEA’s 2026 assessment reports both rapid improvement in energy per AI task and rising aggregate consumption; it estimates 2025 electricity growth of 17% for data centers overall and about 50% for AI-focused facilities. This coexistence is consistent with rebound.
It is not, by itself, a causal estimate of rebound. Aggregate growth also reflects training, new capacity, conventional cloud workloads, demand growth, model scaling, redundancy, and geographical changes. No source reviewed here estimates a stable economy-wide elasticity linking an exogenous fall in inference price to total agentic tokens or electricity use. The safe statement is:
Falling price and energy per inference are real; agent designs can multiply inference demand; and aggregate data-center energy is rising. The extent to which the second causally offsets the first remains unresolved.
The AIOS system-level thesis introduces a countervailing mechanism. Better context composition, memory selection, provenance, bounded actions, verification, and routing could reduce failed attempts, oversized prompts, repeated retrieval, and unnecessary frontier calls. Those gains could lower tokens and energy per accepted task even when the selected model is unchanged. But the components can also add orchestration steps, review calls, caches, indexes, and background work. No reviewed study estimates the cumulative net effect of this full bundle. It should therefore be represented as a potentially important efficiency complement—and possible rebound counterforce—not assumed savings.
A vendor-authored, bottom-up perspective from Microsoft researchers, Oviedo et al., “Energy Use of AI Inference”, estimates a median of 0.34 Wh for a frontier-model query under an at-scale H100 configuration and about 4.32 Wh in a scenario with 15 times more test-time tokens. It argues that batching, hardware utilization, quantization, model routing, and hardware improvement can reduce energy by 8–20 times in combined scenarios.
Limits. The figures are modeled estimates rather than a provider fleet disclosure. Configuration, query, facility overhead, hardware, and utilization assumptions strongly affect results. The paper usefully demonstrates sensitivity to test-time compute and deployment efficiency; it should not be quoted as a universal “energy per prompt.”
4. On-device economics and the case for hybrid execution
4.1 Technical evidence
The strongest recent peer-reviewed systems evidence is Xu et al., “Fast On-device LLM Inference with NPUs”, published at ASPLOS 2025. The research question is whether commodity smartphone neural processing units can overcome the memory movement and operator constraints that make local language-model inference slow and energy-intensive.
Method and sample. The authors implement an NPU-oriented inference system and test it on two Android phones: a Snapdragon 8 Gen 3 device with 24 GB of memory and a Snapdragon 8 Gen 2 device with 16 GB. They evaluate five models from 1.8B to 7B parameters across language understanding, long-context, question-answering, and mobile task benchmarks, compare five CPU/GPU baselines, repeat runs three times, and sample device power at 100 ms intervals.
Principal finding. Depending on model and baseline, the system reports 7.3–18.4 times CPU speedups, 1.3–43.6 times GPU speedups, 1.9–59.5 times lower energy, less than one percentage point of accuracy loss, and 1.4–32.8 times end-to-end speedups. Prefill can exceed 1,000 tokens per second in favorable cases.
Limitations and relevance. Two Qualcomm phones, five relatively small models, a specialized runtime, and selected benchmarks cannot establish general device economics. Results are strongest for prefill; memory, decoding, thermals, battery aging, hardware fragmentation, and longer production sessions remain constraints. No cloud total-cost comparison is attempted, and no independent replication of the paper’s exact speed and energy claims was identified. The paper demonstrates feasible and potentially efficient local execution for bounded models, not a universal economic advantage.
MLCommons’ MLPerf Client and MLPerf Mobile v6.0 extend standardized measurement to client language models. Client tests include Llama 3.1 8B and Phi-3.5 Mini task suites with accuracy constraints; Mobile v6.0 includes Llama 3.2 1B/3B and Llama 3.1 8B on Android. Power methodology remains less mature than performance measurement, and some devices cannot run the largest or longest components. These are capability and performance benchmarks, not evidence of lower lifetime cost.
4.2 A workload-level break-even model
For a workload class \(w\), a useful break-even condition is:
\[ C_{local,w} = \frac{H_w}{N_w} + E_w + O_w + Q_w + F_w \]
\[ C_{remote,w} = P_w + X_w + L_w + G_w + Q'_w + F'_w \]
where \(H\) is the allocable hardware and deployment cost, \(N\) accepted tasks over its useful life, \(E\) incremental energy, \(O\) local operations and updates, \(P\) remote inference price, \(X\) data movement and egress, \(L\) latency and availability cost, \(G\) governance/contractual burden, \(Q\) evaluation and human review, and \(F\) expected failure and rework. The symbols are not intended to suggest that all terms are easily monetized. They identify omitted variables in simple token-price comparisons.
Interpretation. Local execution is favored by high repeated volume, a stable narrow task, acceptable small-model performance, already-owned hardware, strict latency or offline requirements, sensitive inputs, and expensive network or contractual review. Remote execution is favored by irregular demand, need for frontier capability, very long contexts, shared high-utilization hardware, rapid model upgrades, and organizations that cannot maintain heterogeneous runtimes. A local-first knowledge layer with dynamic routing can preserve option value: ordinary files and evidence remain controlled while execution moves according to task requirements.
4.3 Externalities and distribution
Moving inference to a device can reduce network traffic and data-center demand for the displaced task, but it can also lower utilization, duplicate model storage, accelerate device replacement, or move electricity to less efficient hardware and grids. Conversely, hyperscale facilities may achieve higher accelerator utilization and lower cooling overhead while creating concentrated grid, water, land, and transmission burdens. Comparative environmental claims require a lifecycle boundary, counterfactual utilization, regional electricity mix, embodied hardware, and measured task acceptance. “Local is greener” and “cloud is greener” are both unsafe as general propositions.
5. Capital, energy, and geographic concentration
5.1 Current and retrospective estimates
Research question and method. The IEA asks how much electricity data centers currently use, where that demand is concentrated, and how equipment, facility, and deployment trends could change it. Its global model combines historical and projected server shipments—including accelerated servers—with power, utilization, cooling, facility, and regional assumptions, supplemented by company reporting, project pipelines, geospatial analysis, and consultation. It is a bottom-up accounting model rather than a comprehensive meter census.
The IEA’s 2025 Energy and AI report estimates that data centers used about 415 TWh of electricity in 2024, or roughly 1.5% of global consumption, after growing about 12% annually over the prior five years. The IEA’s 2026 update estimates approximately 485 TWh in 2025 and year-over-year growth of 17%, with AI-focused data centers growing about 50%. These are retrospective modeled global estimates built from equipment, utilization, and facility data; no complete global meter census exists.
The IEA also reports strong geographic concentration. The United States, China, and Europe account for the large majority of current data-center electricity demand, while facilities cluster much more tightly than national electricity totals suggest. Power density has risen quickly, placing connection, generation, transmission, and cooling constraints on particular localities.
Capital concentration is similarly pronounced. The 2026 IEA report states that five large technology firms spent more than $400 billion in capital expenditure in 2025 and announced plans implying a further large increase in 2026. Corporate spending and announced plans are observable, but allocation specifically to AI data centers is estimated and plans can be revised.
Limitations and relevance. Category boundaries between AI-focused and conventional facilities are estimated, provider disclosure is incomplete, and utilization changes rapidly. The estimates are directly relevant to infrastructure concentration and grid planning; they cannot allocate energy reliably to one user response, one agent architecture, or one local-versus-remote counterfactual.
5.2 Projections, not measurements
The IEA’s central scenario reaches about 945–950 TWh of data-center electricity demand in 2030, close to twice 2024–2025 levels, with AI-focused demand roughly tripling. This is conditional on deployment, hardware, utilization, and efficiency assumptions.
The 2024 Lawrence Berkeley National Laboratory US Data Center Energy Usage Report estimates that US data centers accounted for about 4.4% of electricity in 2023 and develops scenarios reaching 6.7–12% in 2028. A 2026 LBNL update uses planned equipment shipments, per-device energy, and cooling simulations to construct 2030 scenarios of 9.5–15.3% of US electricity, with 11.8% as its central case. These are increasingly detailed bottom-up scenarios, not guaranteed demand.
5.3 Institutional consequences
Concentrated capital and energy create at least five institutional effects:
- Bargaining asymmetry. Access to accelerators, capacity reservations, model endpoints, and favorable pricing can depend on scale and long contracts.
- Vertical leverage. Cloud firms can combine infrastructure, models, identity, productivity software, and distribution, making one layer’s advantage reinforce another.
- Local externalities. Nationally modest electricity shares can produce acute grid, land, water, noise, and ratepayer effects around a cluster.
- Continuity risk. A change in provider policy, model retirement, region availability, or contract terms can affect many dependent workflows simultaneously.
- Regulatory coupling. Competition, energy planning, environmental permitting, cybersecurity, data protection, and sector assurance become interdependent rather than separable policy domains.
Locally controlled knowledge does not remove these dependencies, but it can reduce the amount of organizational meaning that must be stranded when execution providers change. That is an option-value argument, not proof of lower day-to-day cost.
6. Cloud switching, lock-in, and semantic switching costs
6.1 Regulatory evidence on technical and commercial barriers
The UK Competition and Markets Authority’s cloud-services market investigation, finalized in July 2025, drew on provider submissions, customer evidence, financial analysis, and technical examination. It found highly concentrated infrastructure-as-a-service and platform-as-a-service markets; substantial sustained returns for the largest providers; and adverse effects from egress fees, technical barriers to switching and multicloud use, and Microsoft software-licensing practices. The inquiry examined committed-spend agreements but did not find that their current forms harmed competition. Customers reported that switching or using multiple clouds could require architecture changes, retraining, duplicated operations, and data movement.
Limitations and relevance. The inquiry concerns UK cloud markets and the evidence period preceding the final report. Product boundaries and remedies are jurisdiction-specific. It is authoritative evidence that practical switching is much more than copying data, and that nominal multicloud availability does not imply frequent or economical switching.
The OECD’s 2025 competition paper on cloud computing synthesizes market studies across member jurisdictions. It reports that the two largest providers can account for up to 80% in some major OECD markets, and identifies economies of scale and scope, capital intensity, chip access, proprietary APIs, software dependencies, discounts, and egress costs as reinforcing barriers. The OECD’s 2026 report on AI markets maps similar dependencies across chips, cloud, models, data, energy, talent, and downstream applications.
Limitations. Both reports synthesize heterogeneous jurisdictions and market definitions. They identify mechanisms and recurring findings rather than estimating one universal concentration or switching coefficient.
The FTC’s January 2025 staff report on AI partnerships uses compulsory information requests concerning Alphabet–Anthropic, Amazon–Anthropic, and Microsoft–OpenAI relationships. It documents equity and revenue-sharing interests, consultation and control rights, exclusivity provisions, cloud-spending commitments, and exchanges of information or resources. Staff identifies possible switching costs, input foreclosure, and access to sensitive technical and business information.
Limitations and relevance. This is a competition staff study, not a judicial finding that any arrangement is unlawful or caused a specified price effect. It directly supports the institutional point that model and cloud markets are linked through contracts, not merely technical performance.
The EU Data Act, applicable from September 2025, requires cloud and edge providers to reduce switching obstacles, provide export and interface support, and—subject to staged implementation—eliminate switching charges. It is an important legal response to technical and commercial portability. It does not ensure equivalent model behavior or preserve an organization’s tacit meanings.
6.2 A three-layer switching model
| Layer | What must move | Typical friction | What current policy/protocols address |
|---|---|---|---|
| Technical | Files, databases, identities, APIs, workflows, compute images | Egress, proprietary services, incompatible interfaces, downtime | Cloud-switching rules, open interfaces, container and tool protocols |
| Behavioral | Prompts, tool descriptions, schemas, thresholds, recovery logic, evaluations | Model-specific sensitivities, version drift, different errors and verbosity | Evaluation suites, adapters, prompt migration, conformance tests |
| Semantic/organizational | Categories, annotations, provenance, exceptions, relationships, memory, decision rights | Tacit knowledge, lost context, representation mismatch, changed authority | Only partially addressed by portable files, explicit metadata, governance and human review |
The term semantic switching cost is used here for the third layer and for the portion of the second layer caused by meaning rather than syntax. It is not yet a standardized economic statistic.
A recent peer-reviewed study offers a partial empirical anchor. Jahani et al., “Prompt Adaptation as a Dynamic Complement in Generative AI Systems”, published in Information Systems Research in 2026, asks how users’ prompts adapt when the underlying generative model improves.
Method and sample. Two preregistered experiments assigned 3,750 participants to image-replication and open logo-design tasks, producing about 37,000 prompts. The design compared DALL-E versions, replayed prompts across models, and separated gains due to the model from gains due to changed user behavior.
Principal finding. On the bounded replication task, prompt adaptation accounted for about half of the performance improvement from DALL-E 2 to DALL-E 3. On the open creative task, user adaptation accounted for about 7%, with more than 90% attributed to the model upgrade. Automated prompt rewriting harmed the bounded task and provided only a small, statistically uncertain benefit on the open task.
Limitations and relevance. The study concerns text-to-image generation, one product transition, short sessions, and tasks without durable organizational memory. Some authors have Microsoft/OpenAI relationships, and no independent replication of the estimated adaptation shares was identified. It shows that model performance and user routines can be dynamic complements, especially where an output target is precise. It does not quantify the cost of migrating a file-native knowledge system or establish semantic lock-in as a general market metric.
Emerging preprints on prompt migration and cross-model prompt drift report that structured adaptation can recover performance after model changes, but current results are case-specific and not an adequate foundation for a universal migration-cost claim.
6.3 Why ordinary files help but do not solve semantics
Ordinary local files can lower switching cost when they remain authoritative, documented, versioned, and usable without a particular model. Companion metadata can preserve provenance and relationships without forcing all meaning into a provider’s hidden state. Exact read/write/edit operations can make state transitions inspectable. These are design hypotheses with strong portability logic.
They do not eliminate switching costs. Embeddings must be regenerated; retrieval rankings change; tool descriptions interact differently with different models; long-lived annotations accumulate local conventions; and a new model may obey the same schema with a different distribution of errors. Portability therefore requires behavioral regression tests and semantic stewardship, not just export.
7. Productivity complements and organizational design
7.1 What recent field evidence establishes
Cui et al., “The Effects of Generative AI on High-Skilled Work”, published in Management Science in 2026, asks whether access to an AI coding assistant changes the output of professional software developers.
Method and sample. The authors pool three randomized field experiments at Microsoft, Accenture, and a Fortune 100 firm, covering 4,867 developers. Treatment was access to the assistant; the primary outcome was completed work items in operational repositories rather than a laboratory coding score.
Principal finding. The pooled estimate is a 26.08% increase in completed tasks, with a standard error of 10.3%. Adoption and effects were larger among less-experienced developers. Results varied across the three firms, and individual experiments were noisier than the pooled number.
Limitations and relevance. Completed tasks are not a complete quality-adjusted productivity measure; work-item definitions, code review, task selection, and organizational context differ. The study covers an early tool and developer work at large firms, including the product’s vendor. Its pooled design supplies within-paper, multi-site replication, but no independent replication of the precise effect was identified. It is strong evidence of task-level gains under some conditions, not a general 26% firm-productivity parameter and not evidence about local-first architectures.
Dell’Acqua et al., “The Cybernetic Teammate” and “Navigating the Jagged Technological Frontier” are often discussed together but concern distinct experiments. The peer-reviewed Organization Science “jagged frontier” study asks how GPT-4 affects knowledge workers on tasks inside and outside the system’s capabilities.
Method and sample. A preregistered randomized experiment assigned 758 Boston Consulting Group consultants to no-AI, AI, or AI-plus-prompt-overview conditions. Participants completed realistic consulting tasks selected to lie inside the model’s capability frontier and a task selected outside it.
Principal finding. For inside-frontier work, AI users completed 12.2% more tasks, worked 25.1% faster, and produced higher judged quality. For the outside-frontier task, AI users were 19 percentage points less likely to reach the correct answer. Lower baseline performers received larger inside-frontier gains.
Limitations and relevance. Data were collected in 2023 using an older model. Participants were consultants, the task battery was curated, and the outside-frontier inference rests heavily on one problem type. No independent replication of the precise inside-/outside-frontier effects was identified. The study’s durable contribution is not its precise percentages but the interaction between task fit, reliance, and evaluation.
The separately published “Cybernetic Teammate” study, revised in 2026, asks whether generative AI changes the performance and functional composition of individuals and teams doing product-innovation work.
Method and sample. A preregistered field experiment randomly assigned 791 Procter & Gamble professionals, spanning commercial and research-and-development functions, to work individually or in two-person teams, with or without AI, on real product-development problems. Outputs were evaluated by expert judges.
Principal finding. Individuals using AI produced outputs judged comparable in quality to unaided human teams. AI also reduced the distinction between commercial and technical proposals: individuals generated more balanced solutions across functional domains. Participants reported more positive emotional experience.
Limitations and relevance. The setting is one firm, one workshop-like innovation process, and judged proposals rather than implemented products or long-run outcomes. Several authors were company employees; the article reports research funding from Harvard Business School. No independent replication of the individual-versus-team result was identified. The result supports the possibility that AI can bridge functional vocabularies; it does not show that teams can generally be removed, that knowledge becomes organization-independent, or that outputs remain diverse over time.
A large vendor-affiliated working paper, Dillon et al., “Shifting Work Patterns with Generative AI”, asks whether integrated generative AI changes how office workers allocate time across email, documents, and meetings.
Method and sample. A six-month randomized field experiment across 66 firms offered an integrated assistant to about half of 7,137 knowledge workers. The study uses telemetry and surveys, analyzing both assignment and actual usage.
Principal finding. In the paper’s November 2025 revision, the 80% of treated workers who used the tool in the second half of the experiment spent about two fewer hours per week on email and reduced work outside regular hours. Apart from these individual time savings, the authors did not detect changes in the quantity or composition of tasks from individual-level access.
Limitations and relevance. Several authors are Microsoft researchers and the intervention is a Microsoft product. Access is not use, behavior is not output quality, and early adoption may differ from equilibrium practice. No independent replication of the reported time-allocation effects was identified. The contrast between substantial changes in individual email work and little change in meetings is consistent with coordinated routines requiring organizational redesign rather than a personal tool alone.
Humlum and Vestergaard, “Large Language Models, Small Labor Market Effects”, is a working paper that asks whether rapid chatbot adoption translated into earnings and hours effects in highly exposed occupations.
Method and sample. The authors link representative surveys of adoption and employer encouragement to Danish administrative records for workers and workplaces across 11 exposed occupations, and use difference-in-differences and event-study designs.
Principal finding. Over the first two years after widespread chatbot availability, the estimates for earnings and recorded hours are close to zero and sufficiently precise to rule out effects much larger than about 2%. The study also finds adoption, task restructuring, and occupational movement, particularly where employers encouraged use.
Limitations and relevance. Denmark’s labor institutions and wage setting may not generalize. Earnings and hours are lagging, coarse outcomes; quality, consumer surplus, work intensity, and unrecorded time savings may change first. The study is a useful counterweight to extrapolating task experiments directly into aggregate labor-productivity or displacement claims.
7.2 Complements, not an automatic production function
Taken together, these studies support four propositions:
- AI can raise throughput or judged quality on selected, well-scoped tasks.
- Effects vary by baseline skill, task position relative to model capability, adoption, and evaluation.
- Personal work patterns may change before meetings, approval chains, incentives, and cross-functional coordination.
- Short-run task gains do not automatically appear in earnings, hours, firm productivity, or macroeconomic statistics.
The required complements include training, task decomposition, reliable source access, evaluation criteria, workflow integration, review allocation, incident handling, and incentives. Durable files and explicit process state may make these complements easier to build and preserve. They can also add overhead. Whether the balance is positive is an empirical question for each workflow.
7.3 Standardization with local variation
Organizations face a real design tension. Too little standardization produces incomparable records, duplicated integrations, unsafe permissions, and weak assurance. Too much central standardization erases domain distinctions, creates brittle universal prompts, and pushes tacit exceptions into shadow practices.
A defensible architecture separates a common control plane from locally governed semantic content:
| Common organizational grammar | Domain-owned variation |
|---|---|
| Identity, authorization, timestamps, versioning | Vocabulary, rubrics, exceptions, annotations |
| Provenance and evidence references | Domain source selection and relationship meaning |
| Tool-call and file-change records | Task decomposition and review criteria |
| Minimum evaluation and incident schema | Risk thresholds beyond the minimum |
| Export, package, signature, and retention formats | Local model choice and context composition |
| Human approval and escalation hooks | Allocation of expert authority |
This is an institutional design proposal, not a measured optimum. The recent field evidence makes it plausible: AI can cross functional vocabularies, but coordinated practices change slowly and errors depend on task context. The correct object of standardization is therefore often the evidence and control contract, not a single universal way of reasoning.
Local variation also requires governance. Without shared minimum controls, “local autonomy” can become inconsistent risk acceptance, inaccessible records, or model proliferation. Without local authority, “enterprise alignment” can become semantic flattening and dependency on a central team that cannot evaluate every domain. Successful variation is bounded, testable, and revocable.
8. Auditability and regulated-industry assurance
8.1 What authorities currently require or recommend
The EU AI Act, in force since August 2024 with phased application, provides a legal reference point for high-risk systems. Its framework includes risk management, data governance, technical documentation, automatic record-keeping, transparency to deployers, human oversight, accuracy, robustness, cybersecurity, quality management, conformity assessment, and post-market monitoring. Requirements differ by role and use case, and implementation dates and supporting standards continue to develop. The Act is a legal obligation in scope, not an assurance method by itself.
The European Commission’s general-purpose AI obligations, applicable in stages from August 2025, add technical documentation, downstream information, copyright-policy, and training-content-summary duties, with risk evaluation, incident reporting, and security measures for systemic-risk models. Local downstream custody does not remove provider or deployer duties.
NIST AI 800-4, “Challenges to Monitoring Deployed AI Systems”, asks what makes post-deployment monitoring difficult in practice.
Method and sample. NIST combined a literature review with three practitioner workshops held during 2025, then thematically coded challenges across functionality, operations, human factors, security, compliance, and large-scale impacts.
Principal finding. Monitoring is fragmented across disciplines and tools. Organizations struggle to determine what to log, connect technical indicators with human and compliance effects, manage scale, and relate monitoring evidence to audit and governance processes.
Limitations and relevance. This is an authoritative exploratory report, not a controlled trial or a certification standard. Workshop participants are not a representative sample of all deployers. It directly supports a lifecycle view: pre-deployment evaluation alone is inadequate, and a useful record must join technical, operational, and human evidence.
NIST’s 2026 “Building Evaluation Probes into Agentic AI” project is ongoing research, not an adopted standard. Its prototype research-agent pipeline emits machine-readable records and probes citation faithfulness, completeness, and sufficiency. The project is relevant because evaluation is embedded into the workflow rather than appended after the fact. Automated judges and probes remain fallible and require external validation.
The US Food and Drug Administration’s January 2025 draft guidance for AI-enabled medical devices proposes a total-product-lifecycle approach: intended-use specification, data management, bias analysis, validation, transparency, risk management, and post-market performance monitoring. It is explicitly draft, nonbinding guidance. FDA’s subsequent final guidance on predetermined change-control plans addresses how certain planned model changes can be specified, validated, and controlled. These materials are sector-specific; their broader lesson is that change management and monitoring are assurance objects, not administrative afterthoughts.
The Bank for International Settlements’ 2025 report on central-bank AI governance synthesizes practices among member central banks in the Americas. It recommends adaptive governance, clear accountability, human oversight, inventories, tiered risk management, and three lines of defense. It is an institutional practice report rather than causal evidence, but it shows why confidential data, operational resilience, and public legitimacy make governance inseparable from deployment design.
8.2 The minimum evidence chain
For consequential reasoning workflows, an auditable event should be able to identify, at an appropriate level of disclosure:
- the task, intended use, risk class, and applicable policy version;
- authenticated actors, roles, permissions, delegations, and human decision authority;
- input file versions and source references, or privacy-preserving commitments where raw retention is inappropriate;
- model/provider, version or immutable identifier, configuration, routing decision, and relevant safety settings;
- context assembly and retrieval results sufficient to reconstruct what evidence was made available;
- tool calls, parameters, returned status, and affected resources;
- exact file additions, edits, deletions, and post-processing transformations;
- evaluation results, confidence or uncertainty where meaningful, exceptions, and failed attempts;
- human reviews, overrides, approvals, and reasons at mandated control points;
- timestamps, retention controls, tamper evidence, incident linkage, and later corrections.
Full hidden reasoning traces are neither necessary nor sufficient. They may be unavailable, misleading as explanations, sensitive, or too costly to retain. Assurance should prioritize observable inputs, evidence, actions, tests, and accountable decisions.
8.3 Locality: an assurance affordance, not an assurance result
Local custody can improve the chain of evidence by keeping canonical sources, annotations, diffs, and decisions under an accountable institution’s retention and access controls. It can reduce dependence on a provider’s log format or retention policy, and make inspection possible after a service or model is retired.
The same local environment can be poorly secured, silently modified, incompletely logged, or governed by one person without separation of duties. Remote services may supply stronger physical security, immutable logging, validated controls, and professional operations. The correct claim is conditional:
Local-first custody can make evidence continuity and inspection easier when it is paired with identity, authorization, tamper evidence, retention, evaluation, and independent review. Locality alone proves none of those properties.
For regulated sectors, assurance must attach to a defined intended use and deployed configuration. A general-purpose model certificate cannot substitute for workflow validation; a workflow log cannot substitute for clinical, financial, safety, or legal performance evidence; and a management-system certification cannot guarantee each output.
9. Interoperable agents and attestable domain packages
9.1 Protocol layers
The current Model Context Protocol specification, revision 2026-07-28, defines JSON-RPC-based, stateless requests through which model-facing applications can discover and invoke tools, read resources, and use prompts. It also defines an opt-in extension framework, including long-running tasks. Its principal economic value is interface reuse: a capability can be exposed once rather than integrated independently with every model application. Its security value depends on authentication, authorization, consent, validation, and deployment practice outside the core message grammar. The breaking move from a stateful 2025 core to the stateless 2026 revision also illustrates that the protocol layer itself still carries migration cost.
The Agent2Agent protocol, transferred to the Linux Foundation in June 2025, addresses communication between independent agents. Agent Cards advertise capabilities; task, message, artifact, and status structures support longer-running exchanges; authentication is intended to use standard web mechanisms. It complements rather than replaces tool protocols: one addresses agent-to-agent delegation, the other agent-to-resource access.
The Ecma NLIP standards suite, approved in December 2025, and several developing IEEE and ITU work items reflect parallel efforts to standardize agent communication. The NIST AI Agent Standards Initiative, launched in February 2026, explicitly treats interoperability, security, identity, and evaluation as unfinished work.
Protocol specifications establish message and capability contracts. They do not establish:
- that two agents attach the same meaning to a field;
- that a capability claim is true or current;
- that delegation is authorized for the particular data and consequence;
- that the receiving agent preserves provenance and policy;
- that an artifact is safe, accurate, or fit for purpose;
- that a transaction is auditable across organizational boundaries.
Open protocols can reduce integration and some switching costs while increasing the number of callable endpoints and trust boundaries. Security, semantic profiles, test suites, and institutional accountability must mature alongside adoption. Claims of a settled universal agent layer are premature.
9.2 Building blocks for portable knowledge packages
Several standards provide reusable components without being designed as one universal “knowledge package” format:
- RO-Crate 1.3, released as a community recommendation in June 2026, uses JSON-LD to describe files, entities, people, software, actions, and relationships in a portable research object. It is a backward-compatible update to the stable 1.2 series.
- W3C Verifiable Credentials Data Model 2.0, a Recommendation since May 2025, represents cryptographically verifiable issuer claims about subjects, with status and selective-disclosure mechanisms available through related specifications.
- C2PA 2.2, released in May 2025, defines signed content-credential manifests, assertions, actions, ingredients, trust lists, timestamps, and validation for digital-media provenance.
- Sigstore bundles package the material needed to verify a software-artifact signature and transparency-log evidence. Sigstore’s model illustrates identity-bound, transparent attestations rather than domain-knowledge certification.
A possible attestable domain package could combine:
- ordinary, content-addressed files;
- a versioned manifest describing entities and relationships;
- source citations, acquisition dates, licenses, and transformation provenance;
- schemas and declared semantic profiles;
- intended uses, exclusions, policy constraints, and risk classification;
- executable conformance tests and benchmark claims with environments;
- issuer identity, signature, credential, timestamps, and status/revocation endpoints;
- review history and accountable human approval;
- privacy partitions separating distributable content from local overlays.
This is a design scenario assembled from adjacent standards, not an adopted cross-domain standard or demonstrated market institution.
9.3 Peer-to-peer distribution and institutional trust
Content addressing permits packages to be cached or exchanged peer to peer while a digest continues to identify the same bytes. Local overlays can preserve private annotations while shared base packages remain verifiable. This could reduce repeated downloads, improve resilience, and let professional communities distribute bounded corpora, methods, or evaluation suites without centralizing every use event.
The hard problems are institutional rather than transport-level:
- Who is authorized to make which claims?
- What evidence and testing support a claim?
- How are corrections, expiry, and revocation propagated to offline peers?
- How are licenses, privacy restrictions, and confidential sources enforced?
- What prevents a signed but obsolete or malicious package from gaining misplaced authority?
- Which party bears liability when a package, model, integration, or user decision fails?
- How are forks handled when reasonable experts disagree?
“Certified knowledge” is therefore unsafe terminology unless a named scheme has an accredited or otherwise accountable issuer, explicit scope, test method, surveillance, appeals, revocation, and liability. A cryptographic signature establishes integrity and a binding to an issuer; a verifiable credential establishes that an issuer made a claim; provenance establishes a history claim. None establishes truth.
10. Evaluating the “AI as electricity” analogy
10.1 What the analogy explains
The OECD’s 2025 assessment, “Is Generative AI a General Purpose Technology?”, evaluates generative AI against economic criteria such as pervasiveness, continuing technical improvement, and innovation complementarities. It concludes that generative AI plausibly displays important general-purpose-technology characteristics, while diffusion, productivity, distribution, and institutional effects remain uncertain.
Method and limitations. This is a structured evidence synthesis, not a causal test that AI has already produced general-purpose productivity growth. The technology and evidence base are changing rapidly, and GPT classification does not specify timing, magnitude, beneficiaries, or governance.
The analogy to electrification is helpful in three respects:
- Complementarity and delay. Installing motors did not immediately deliver the productivity of redesigned factories. Similarly, adding a model endpoint does not redesign decisions, evidence flows, incentives, or accountability.
- Pervasive input. Reasoning services may enter many activities without being a single final product, making direct sector classification difficult.
- Infrastructure externalities. Generation, transmission, siting, and local grid constraints have direct counterparts in the electricity actually used by digital infrastructure.
Acemoglu’s “The Simple Macroeconomics of AI”, published online in August 2024 and in a January 2025 journal issue, asks what current task-exposure and cost-saving evidence implies for aggregate productivity.
Method. It combines a task-based macroeconomic framework with estimates of AI exposure, adoption, cost savings, and the share of tasks likely to be profitably automated.
Principal finding. Under the paper’s assumptions, AI raises total factor productivity by no more than roughly 0.66% over ten years, with smaller baseline estimates; distributional effects may be adverse if capital captures gains or if low-productivity uses expand.
Limitations and relevance. This is a model-based scenario, not a forecast fact. Results depend on exposure measures, profitable task shares, cost savings, complementarity, new-task creation, and the period over which reorganization occurs. It is valuable because it prevents micro-level task gains from being arithmetically treated as macro productivity.
10.2 Where the analogy fails
| Electricity | Reasoning infrastructure |
|---|---|
| Metered, largely fungible commodity within specified grid products and quality standards | Heterogeneous output whose usefulness depends on task, evidence, context, and evaluation |
| Stable physical behavior at the point of use | Probabilistic and version-dependent behavior; providers can change models and policies |
| Commodity normally does not inspect or transform customer meaning | Service processes confidential data, retrieves sources, invokes tools, and can alter records |
| Quality is largely independent of local organizational vocabulary | Performance depends on local categories, exceptions, memory, prompts, tools, and authority |
| Mature metering, safety codes, common-carrier institutions, and grid governance | Immature agent protocols, semantic profiles, assurance practice, and cross-provider accountability |
| Switching electricity supplier rarely redesigns the appliance’s meaning | Switching models can require prompt, evaluation, workflow, and semantic adaptation |
| One unit cannot be copied at near-zero marginal cost | Model software and files are partly non-rival, while chips, latency, energy, and attention remain rival |
The “grid versus local generation” metaphor is also incomplete. A locally owned reasoning environment is not merely a generator behind the meter. Its important asset may be canonical context, provenance, memory, and authority—assets for which electricity has no analogue.
10.3 Safe use of the analogy
The analogy should be limited to a claim about diffusion and complements: a broadly applicable technical input can require long periods of reorganization and complementary investment before aggregate productivity appears, while creating concentrated infrastructure requirements. It should not be used to imply that intelligence is a homogeneous utility, that reasoning quality can be metered like energy, that provider switching is semantically neutral, or that regulation can be copied from electricity markets without substantial modification.
For the AIOS thesis, the analogy’s strongest implication is organizational reconfiguration rather than utility equivalence. Frontier and local models are inputs into a larger production system whose productive capital includes purpose, files, memory, relationships, scaffolding, evaluation, and authority. If those complements become portable and locally owned, model services may diffuse widely without owning the durable intelligence layer. That possibility resembles the complement-intensive reorganization associated with electrification; it does not make reasoning output electricity or imply the disappearance of centralized generation-like infrastructure.
11. Institutional scenarios—not forecasts
The following scenarios expose tradeoffs. They are not probability-weighted forecasts.
Scenario A: cloud-dominant execution with portable local state
Canonical files, metadata, evaluations, and approvals stay under user or institutional custody; most inference is remote and routed among providers. This exploits frontier capability and shared utilization while preserving some exit option. Its risks are behavioral drift, provider concentration, confidential-context exposure, and the operational difficulty of maintaining cross-model regression tests.
Scenario B: device and private-edge specialization
Small models handle high-frequency classification, retrieval, drafting, redaction, and file operations on devices or private servers; remote systems handle exceptions and capability-intensive review. The economic case rests on repeat volume, latency, privacy, and already-owned capacity. Its risks are model fragmentation, patching, hardware churn, uneven assurance, and silent quality differences between local and escalated paths.
Scenario C: federated professional knowledge packages
Communities, firms, or public institutions publish versioned, signed packages containing sources, methods, policies, tests, and declared scopes. Peers retain private overlays and choose local or remote execution. The institutional rationale is provenance, reuse, and contestable authority rather than sale of “intelligence.” Its risks are false confidence in signatures, governance capture, revocation failure, liability ambiguity, licensing conflicts, and fragmentation among semantic profiles.
Across all three scenarios, the durable institutional asset is the ability to preserve meaning and evidence independently of a particular execution vendor. Whether that asset is worth its maintenance cost remains workload- and institution-specific.
12. Four-layer examination of the AIOS architectural thesis
This section places the AIOS architectural thesis beside the evidence without collapsing one into the other. AIOS is treated as a complementary reasoning and knowledge architecture: it may call centralized frontier models when their capabilities are needed while moving durable context, expertise, memory, workflow state, and authority into local person- or institution-owned systems. Each subsection distinguishes four layers:
- what current evidence establishes;
- the AIOS architectural thesis;
- first- and second-order implications if the thesis is substantially correct;
- conditions, uncertainties, counterforces, and discriminating tests.
The purpose is not to make absence of evidence do the work of counterevidence. It is to preserve the scale of the proposed transition while making its causal chain and empirical burden explicit.
12.1 Bounded semantic judgment, deterministic effects, and human authority
1. Current evidence. Assurance standards and regulatory guidance consistently separate accountable roles, intended use, authorization, validation, logging, human oversight, and post-deployment monitoring. The field experiments reviewed above show that models can improve selected judgments while also inducing confident error outside their capability frontier. Exact state changes, identity checks, and access controls remain conventional deterministic-software strengths. This evidence supports explicit control boundaries; it does not compare complete “pre-AI” and “AI-native” production architectures.
2. AIOS architectural thesis. Language models should make bounded semantic judgments within a declared context and action space. Deterministic code should establish identity and scope, validate schemas and permissions, and guarantee the exact mechanical effect of an authorized operation. Humans should retain purpose, exception authority, and consequential decision rights. Model output is therefore not merely prose for downstream code to scrape; it can be a typed, reviewable judgment inside a controlled state transition.
3. Implications if substantially correct. The first-order implication is a different division of computational labor. Models contribute flexible interpretation where exhaustive rules are costly, while deterministic components make permissions and effects inspectable. This could lower the expected cost of failures, review, and integration, particularly for repeated workflows that mix ambiguous evidence with exact file changes. It could also make model substitution easier because the model is not the canonical store of identity, authority, or state.
The second-order implication is an application architecture organized around local evidence, policies, and state transitions rather than around a remote application’s database and workflow engine. If model judgments can be bounded reliably, more software functions may become compositions of local knowledge, semantic proposals, deterministic constraints, and human ratification. That could reduce dependency on centralized application logic while increasing demand for evaluation packs, policy schemas, local operations, and assurance tooling.
4. Conditions, counterforces, and tests. Deterministic code guarantees only properties correctly captured in its specification and implementation. It cannot guarantee that a semantic judgment is substantively correct, that an identity record is accurate, or that an authorized action is wise. Human approval can be overloaded or ceremonial. Separation adds schemas, interfaces, latency, and maintenance, and observed gains might be attributable to ordinary software modularity rather than to a distinct architectural regime.
A direct test should compare human-only, unrestricted model, free-text-output-plus-parser, typed bounded-judgment, and bounded-judgment-plus-human-approval designs on identical workflows. Outcomes should include accepted-task quality, unauthorized effects blocked, semantic false positives and negatives, review time, recovery, audit reconstruction, and maintenance effort. The comparison must hold model access and inference budgets constant.
12.2 Coordinated intelligence and distinct cognitive modes
1. Current evidence. Outcomes from model-based systems vary with prompts, retrieved evidence, tools, scaffolds, evaluation, and human use; model capability alone does not determine the result. Agent studies also show that added coordination can help or hurt depending on task structure and budget.
Kim et al., “Capable Language Models Can Outgrow the Benefits of Collaboration”, published in Nature Machine Intelligence in July 2026, asks when multi-agent coordination outperforms a strong single agent.
Method and sample. The authors hold prompts, tools, and per-system compute ceilings constant across 260 configurations: a single agent and four coordination structures, three model families, and six benchmarks spanning browsing, financial research, planning, workplace tasks, software maintenance, and terminal work.
Principal finding. Single-agent baseline performance was the most statistically robust predictor of whether coordination helped. The fitted system selected the best architecture in 87% of held-out within-domain configurations. Multi-agent designs incurred 1.6–6.2 times as many realized reasoning turns as the single-agent baseline despite matched compute ceilings. An approximately 45% single-agent-performance threshold helped predict the direction of multi-agent gains in an additional validation, but the underlying interaction did not survive the paper’s strongest cluster-robust correction and is presented as a selection rule rather than a universal scaling law.
Limitations and relevance. The six benchmarks do not cover long-lived file-native domains or consequential institutional workflows; architecture-specific prompts were not tuned; absolute cross-domain prediction was poor; and several tool-overhead patterns did not survive conservative inference. The study was funded by Google, many authors were Google employees, and no independent replication was identified. It is direct evidence that coordination structure matters and that more agents are not inherently better. It does not test sequential cognitive modes operating over stable local files.
2. AIOS architectural thesis. Useful intelligence is a system-level outcome, not merely a property of model weights. A user-visible act of reasoning can emerge from purpose framing, composed context, files, annotations, relational metadata, provenance, multiple memory resolutions, distinct modes such as thinking, writing, planning, editing, and review, bounded model judgments, deterministic operations, verification, post-response integration, and human authority. These functions need not be collapsed into a single persistent agent or a single undifferentiated context. Stable files can carry state between modes, while each mode receives only the evidence and permissions it needs. The central cumulative hypothesis is that many individually modest improvements can interact to reduce inference burden, retries, and hallucination exposure per accepted task.
3. Implications if substantially correct. At the personal level, intelligence infrastructure becomes cumulative: a person’s files, distinctions, annotations, corrections, and methods can improve future work without being locked inside one conversational history or provider. The model remains important but becomes one replaceable participant in a longer-lived cognitive system.
At the organizational level, different functions could share evidence and control contracts while retaining domain-specific reasoning contexts. Work could be distributed across foreground interaction, non-blocking analysis, specialist review, and exact integration without placing one autonomous agent above the whole organization. If role separation reduces context interference and makes review more independent, increasingly capable models could turn mature local scaffolding into a capability multiplier rather than merely a coordination overhead.
Second-order effects could include durable organizational memory that survives staff and vendor changes, new forms of delegation to local systems, and a shift in competitive advantage from exclusive application data silos toward accumulated domain methods and evaluation quality. They could also permit a smaller local model with excellent domain scaffolding to satisfy tasks that would otherwise require repeated frontier calls. If purpose, context, memory, provenance, planning, verification, and integration reduce different sources of failure, their combined economic value could be superadditive: fewer retries reduce tokens, review, latency, and rework together. That cumulative effect is a core implication of the thesis, not a finding established by the component studies.
4. Conditions, counterforces, and tests. The proposition that intelligence “has no single center” is an architectural and philosophical thesis, not a measurement. Multiple modes can discard shared evidence, repeat tokens, propagate inconsistent summaries, or obscure responsibility for integration. A strong single model with a well-structured context may outperform decomposition on short or tightly coupled problems. Any benefit must be distinguished from simply spending more inference and human attention.
A factorial study should vary role separation, context separation, independent review, stable file-mediated state, and total inference budget independently. It should include short and long tasks, parallelizable and sequential tasks, conflicting evidence, and inside- and outside-frontier work. Measures should include final quality, calibration, contradiction, diversity, latency, token and energy use, audit reconstruction, and human integration effort.
12.3 Bounded autonomy through menus, permissions, and escalation
1. Current evidence. Security engineering and regulated assurance support least privilege, authenticated roles, explicit consent, controlled changes, and human oversight. Current agent protocols standardize messages and capabilities but do not themselves establish that a delegation is authorized or safe. There is little direct evidence identifying an optimal action-space size for general-purpose agents.
2. AIOS architectural thesis. Bounded menus and permissions can provide an intermediate design between rigid pipelines and unrestricted autonomy. A model may choose among semantically meaningful actions, ask for clarification, or propose an exception, while deterministic controls constrain resource access and exact effects. The menu is not intended to prescribe every reasoning step; it defines the legitimate action surface and escalation path.
3. Implications if substantially correct. The first-order implication is practical autonomy that is both more adaptable than a fixed workflow and more governable than an unconstrained agent. Institutions could delegate a wider range of ordinary work while preserving visible decision rights and preventing entire classes of unauthorized action.
The second-order implication is a new governance layer for everyday software: permissions would attach not only to users and files but to semantic acts, risk classes, and transitions. Local domain owners could encode their own thresholds within organization-wide minimum controls. This could support decentralized variation without surrendering accountability to either a central platform or a free-ranging agent.
4. Conditions, counterforces, and tests. A menu may omit a safe novel action, force an ill-fitting category, or transfer hidden power to its designer. Formal authorization can still permit harmful sequences; prompt injection, confused-deputy behavior, information leakage, and collusion across tools remain. Menus and policies also require upkeep as work changes.
Identical workflows should be randomized across fixed pipelines, bounded action menus, and open-ended agents with matched tools and budgets. Evaluation should include completion, unsafe attempts, unauthorized effects actually blocked, false refusals, escalation quality, user workarounds, and the long-run cost of changing permissions.
12.4 Local domain systems and selective frontier-model use
1. Current evidence. Recent systems research demonstrates technically viable on-device inference for bounded small-model workloads, and standardized client benchmarks now test language models up to 8B parameters. Public inference prices have fallen rapidly, while frontier capability and shared cloud utilization remain economically attractive for difficult or bursty work. Competition investigations establish that remote cloud and model dependencies can carry technical, contractual, and semantic switching costs.
2. AIOS architectural thesis. Self-contained local domains can hold the canonical files, accepted ground, relationships, memory, evaluation criteria, and authority needed for a large share of ordinary reasoning. Increasingly capable local models and mature scaffolding may handle many repeated personal and organizational tasks; remote frontier models remain available selectively for difficult inference, current external knowledge, or independent review. Model selection can be dynamic by cognitive movement—the kind of step being performed, such as retrieval, classification, drafting, planning, review, or integration—as well as by capability, privacy, latency, cost, sensitivity, and assurance need. Local custody and local execution are separable choices.
3. Implications if substantially correct. The first-order effect would be a large change in the mix of execution rather than the disappearance of remote AI. High-frequency retrieval, classification, drafting, planning, annotation, redaction, and file operations could occur locally; only exceptions or capability-intensive tasks would cross the network. The accepted-task cost could fall where hardware is already owned, workloads repeat, and local context avoids repeated transmission and reconstruction.
The second-order effect would be a change in the boundary of application infrastructure. Remote services could become less responsible for holding every user’s durable state, memory, and workflow. Some centralized application functions might be replaced by local orchestration over ordinary files, while remaining remote services specialize in frontier inference, synchronization, identity, software and model distribution, backup, or cross-party coordination. This would reduce dependence on particular application providers even if total compute demand continued to grow.
A hybrid system also creates option value. When frontier-model prices or capabilities change, canonical context does not have to migrate with them. When a local model improves, more tasks can move onto already-governed state without rebuilding the entire knowledge environment. Dynamic routing can reserve expensive or less-private frontier calls for cognitive movements where their marginal capability is worth the cost and risk, while keeping routine or sensitive movements local. Repeated substitution at the execution layer could make model competition more effective than nominal API compatibility alone.
4. Conditions, counterforces, and tests. “A large share” needs a declared task denominator. Ordinary work often depends on live external facts, other people’s authority, specialized tools, and frontier capability. A local system can appear sufficient because difficult cases are silently deferred to humans or remote services. Hardware heterogeneity, patching, model updates, backup, and evaluation can outweigh saved inference charges. Cheaper private execution can also cause rebound through more background work.
A prospective workload study should inventory tasks before classifying outcomes as local-only, local with external data, hybrid with remote inference, remote-first, or human-only. It should measure acceptance quality, review, latency, privacy exposure, failure, and total cost by risk and task type. Local knowledge custody and local model execution must be varied independently.
12.5 Private local cognition and person-owned intelligence
1. Current evidence. Local execution can keep selected raw inputs off remote inference services, and local custody can give users direct control over files and retention. Existing research does not measure the full privacy outcome of a mature personal reasoning environment. Regulators and assurance reports show that privacy depends on lifecycle data flows, access control, logs, third parties, and incident response—not location alone.
2. AIOS architectural thesis. A person should be able to accumulate durable context, memory, annotations, relationships, drafts, and methods in an inspectable local system that serves the person rather than a remote application’s account model. Private reasoning can remain private by default, with deliberate disclosure to models, collaborators, or institutions. Human purpose and authority remain upstream of the system.
3. Implications if substantially correct. The first-order effect is cognitive continuity and a stronger personal exit right. A person could change models or applications without abandoning the accumulated structure through which they think and work. Sensitive exploration, incomplete drafts, and private reflection could occur without routine transmission to a central service, expanding the practical space for unobserved cognition.
Second-order effects could be institutionally significant. Person-owned cognitive capital may alter bargaining relationships with employers and platforms; make expertise more portable across projects; support long-horizon learning and memory; and create demand for new norms around inheritance, agency, deletion, search, and selective disclosure. Personal systems could collaborate by exchanging bounded packages rather than exposing entire histories. Private local cognition could become a practical civil-liberties affordance, especially where central services are surveilled, censored, unreliable, or unavailable.
4. Conditions, counterforces, and tests. Endpoint compromise, coercive device access, uncontrolled copies, weak key management, and household or workplace surveillance can defeat local privacy. Portability can conflict with confidentiality, employer ownership, professional duties, or another person’s rights. Accumulated memory may entrench errors or become burdensome to curate. Effective ownership requires comprehensible export, deletion, backup, recovery, permissions, and the ability to operate without the original vendor.
Evaluation should trace actual data flows and retention under realistic tasks; test compromise, loss, shared-device use, deletion, selective disclosure, and provider exit; and compare user comprehension and control with well-governed remote alternatives. “Private” should describe measured access and flow properties, not merely physical location.
12.6 Organizational intelligence and regulated use
1. Current evidence. Productivity studies show that AI gains depend on task fit, worker experience, adoption, evaluation, and organizational complements. Coordinated routines change more slowly than individual work. The EU AI Act, NIST, FDA, BIS, and management-system standards emphasize lifecycle governance, evidence, accountability, change control, and monitoring. None validates a specific local-first implementation.
2. AIOS architectural thesis. An organization can standardize identity, authorization, provenance, evaluation, incident handling, and evidence formats while allowing departments and professions to control local vocabulary, sources, annotations, rubrics, exceptions, and authority. Domain systems remain self-contained enough to preserve meaning and confidentiality, but interoperable enough to exchange bounded evidence and requests.
3. Implications if substantially correct. The first-order organizational effect would be distributed intelligence with common controls. Domain experts could build and retain operational knowledge without waiting for one central application or AI team to encode every exception. Institutional memory could reside in inspectable files, relationships, and evaluation packs rather than in scattered conversations or one provider’s hidden state.
For regulated use, local evidence continuity could support reconstruction of what sources, model, policy, permissions, tool calls, file changes, evaluations, and approvals produced a consequential outcome. Local domains could be validated for defined intended uses and updated through controlled change packages. Frontier models could still be used when allowed, but their judgments would enter a retained evidence chain rather than become the whole system of record.
Second-order effects could include new professional roles for domain-package stewardship, model and workflow assurance, semantic migration, and independent review. Organizational boundaries may become more permeable where bounded domain packages can be shared without surrendering all internal context. Regulators may be able to inspect evidence contracts and declared transitions rather than rely only on provider-level model documentation.
4. Conditions, counterforces, and tests. Local autonomy can become inconsistent risk acceptance or inaccessible records; central standardization can flatten domain meaning. Detailed logs create privacy, security, litigation, and surveillance risks. Human authority must be real, trained, and resourced. A local workflow still needs substantive performance validation for its intended use, separation of duties, incident response, and ongoing monitoring.
Multi-institution pilots should compare a common evidence/control plane with centralized semantic standardization and with uncoordinated local practice. Measures should include accepted outcomes, policy exceptions, audit reconstruction, domain-user adaptation, central-team burden, migration effort, incident detection, and time to correct shared and local errors.
12.7 Peer collaboration, knowledge packages, and global accessibility
1. Current evidence. MCP, A2A, and related protocols can reduce interface work; RO-Crate, Verifiable Credentials, C2PA, and Sigstore supply components for portable manifests, issuer claims, provenance, and signatures. They do not establish semantic agreement, truth, fitness, or certification. Client inference benchmarks and on-device results demonstrate some technical basis for operation outside continuous cloud access.
2. AIOS architectural thesis. Local domains could exchange content-addressed, versioned packages containing files, metadata, sources, relationships, methods, tests, policies, signatures, and declared scope. Recipients could preserve private overlays and use their own local or remote models. Peer exchange would distribute bounded knowledge and evidence without centralizing every use event.
3. Implications if substantially correct. The first-order effect would be reusable domain infrastructure. Professional communities, educational networks, public agencies, firms, and individuals could publish inspectable packages that others can run, test, annotate, and fork. Collaboration would center on evidence and declared methods rather than on access to one hosted application.
Second-order effects could include federated knowledge institutions, contestable certification, and faster diffusion of expert practice. A rural clinic, small firm, school, or independent professional could combine an inexpensive local model with a high-quality domain package and selectively obtain remote expertise. Low-bandwidth or intermittently connected communities could continue to reason over local materials. Multiple packages could preserve competing interpretations rather than forcing one global ontology.
This architecture could broaden global accessibility by lowering recurring bandwidth and service dependence, making translation and adaptation local, and allowing communities to retain authority over sensitive knowledge. It could also shift value from exclusive access to hosted data toward the credibility, maintenance, and interoperability of shared packages.
4. Conditions, counterforces, and tests. Hardware cost, electricity, language coverage, technical literacy, accessibility needs, and maintenance can reproduce global inequality. A signed package can be obsolete, biased, malicious, or outside its intended scope. Certification requires accountable issuers, explicit claims, test methods, surveillance, revocation, appeals, and liability. Forks can support pluralism or produce fragmentation and incompatible semantic profiles.
A peer pilot should test package discovery, offline verification, local overlays, correction and revocation, cross-language adaptation, license enforcement, and migration across models. Institutional evaluation should measure who can publish, challenge, and maintain packages; whether small or resource-constrained users achieve accepted outcomes; and whether the system reduces or merely relocates dependence.
12.8 Scaffolding, context, and the usable frontier
1. Current evidence. Scaffold and context design materially change model behavior, but more context and more agent steps do not reliably produce better results. Agent studies show large variation in cost and success across architectures.
Du et al., “Context Length Alone Hurts LLM Performance Despite Perfect Retrieval”, published in Findings of EMNLP 2025, asks whether long-input performance loss is only a retrieval failure. The authors test five open- and closed-source models on mathematics, question answering, and coding while making all required evidence perfectly retrievable and manipulating input length.
Principal finding. Performance declined by 13.9% to 85% as input length increased within advertised context windows, including conditions in which filler was whitespace or masked and evidence appeared immediately before the question. A short-context recitation intervention improved GPT-4o by up to four percentage points on RULER.
Limitations and relevance. These are controlled benchmark manipulations, not end-to-end institutional tasks; the wide range shows strong model/task dependence; a maximum improvement is not an average treatment effect; and no independent replication of the exact estimates was identified. The study does not imply that context should always be short. It establishes that retrieval of all relevant information and nominal context capacity are not sufficient for reliable reasoning.
2. AIOS architectural thesis. Prevailing application scaffolds may underuse frontier-model reasoning when they flatten durable knowledge into one prompt, treat model output as unstructured text, mix incompatible cognitive tasks, or obscure evidence and authority. File-native state, multiple-resolution memory, role-specific context, bounded judgments, and post-process integration may let both frontier and local models operate closer to their useful capability.
3. Implications if substantially correct. Better scaffolding could increase the amount of useful work obtainable from a given model and token budget. The gain may be especially important locally: a smaller model operating over well-curated domain state could satisfy tasks that a generic prompt cannot. As models improve, a stable local architecture could capture those gains without repeatedly rebuilding user knowledge inside new applications. A router that recognizes different cognitive movements could also avoid treating every step as a demand for the same largest model, lowering cost and exposure without imposing one local-only ceiling.
Second-order effects could compound. If scaffolding increases accepted-task capability while local model hardware improves, remote inference may become an exception path for more workflows. If exact integration and evaluation reduce rework, the relevant productivity gain would exceed the reduction in model-call price alone. Mature scaffolding could thus be an economic complement to model progress, not merely overhead around it. At sufficient scale, cumulative reductions in prompts, retries, verification failures, and unnecessary frontier routing could change both application-layer infrastructure demand and the distribution of remote inference, even though frontier training and centralized high-capability services remain essential.
4. Conditions, counterforces, and tests. Current evidence does not identify scaffolding as the sole or dominant cause of agent failure. Failures also arise from base-model limits, excess or irrelevant context, retrieval errors, unreliable tools, ambiguous objectives, bad decompositions, stale knowledge, evaluation mismatch, unsafe delegation, and compounding errors. More elaborate scaffolding can suppress performance, multiply calls, and make debugging harder.
Preregistered ablations should vary context sufficiency, irrelevant-context load, memory resolution, role separation, scaffold, tool reliability, feedback, and token budget independently across frontier and smaller models. Claims that a scaffold unlocks a model’s latent capability require resource-matched improvements across task classes, transparent failure analysis, and independent replication. The full routing benchmark belongs to a dedicated investigation; for this memo, the required economic output is narrower: accepted-task cost and quality by route, including retries, review, latency, energy, privacy exposure, and switching effects.
12.9 Integrated first- and second-order implication map
The table preserves the scale of the thesis while keeping every consequence conditional.
| Domain | First-order implication if the thesis holds | Second-order implication | Key conditions and counterforces |
|---|---|---|---|
| Personal intelligence | Durable memory, context, methods, and model-assisted work remain person-controlled | Portable cognitive capital; stronger exit rights; private space for thought | Endpoint security, usable ownership, curation, rights in shared data |
| Organizational intelligence | Domain systems combine local semantics with common evidence and controls | Distributed institutional memory; less dependence on central application teams | Governance, evaluation, cross-domain coordination, correction processes |
| Application infrastructure | More state, workflow, and knowledge orchestration occurs locally | Central services become thinner or more specialized; switching becomes more credible | Synchronization, identity, backup, updates, collaboration still require infrastructure |
| Remote inference | Routine bounded work shifts toward devices or private edge; frontier calls become selective | Lower dependency on remote inference for ordinary work, even as frontier access remains valuable | Local capability, utilization, accepted-task cost, rebound, current external data |
| System efficiency and routing | Context, memory, verification, and model choice reduce avoidable calls and rework | Cumulative complementarity may lower infrastructure required per accepted task | Component gains may not add; routing errors and orchestration can increase cost and failure |
| Data centers and training | Little direct first-order change to frontier training requirements | Demand composition may change; total demand can still rise with expanded use | Chip and energy concentration, training scale, remote exception work, rebound |
| Regulated activity | Local evidence chains join model judgments to sources, controls, exact effects, and approvals | Workflow-level assurance and more portable supervisory evidence | Intended-use validation, tamper evidence, minimization, monitoring, accountable review |
| Peer collaboration | Signed packages and private overlays enable bounded exchange | Federated professional and public knowledge institutions | Certification scope, revocation, liability, semantic profiles, governance capture |
| Global accessibility | Local packages and models reduce recurring connectivity dependence | Wider access to domain reasoning and locally governed adaptation | Hardware, electricity, languages, disability access, maintenance, institutional capacity |
| Energy and resource use | Some remote workloads move to heterogeneous local hardware | Central energy may fall for displaced tasks or rise overall through rebound and new use | Lifecycle boundaries, device turnover, grid mix, utilization, inference intensity |
The most consequential implication is therefore not that centralized models or data centers disappear. It is that the durable intelligence layer—the purpose, context, expertise, memory, relationships, workflow, evidence, verification, and authority that make model calls useful—could migrate toward people and institutions. Dynamic local/frontier models would then serve a more portable system of intelligence rather than own the entire application relationship.
12.10 Minimum empirical program
A compact program capable of testing the architecture and its implications would include:
- Architecture trial: human-only, direct-model, conventional model-output-plus-parser, bounded-judgment/deterministic-effect, and role-separated designs on the same real workflows and resource budgets.
- Locality trial: local-only, remote-only, and hybrid execution while holding the canonical file representation and accepted ground constant, separating custody effects from model-capability effects.
- Representative workload census: prospectively sample personal and organizational tasks, declare acceptance criteria and risk classes, and measure the share handled locally, selectively remotely, or by humans without assuming the answer.
- Migration challenge: change the model and execution provider mid-study; measure export completeness, regression failures, retuning time, stranded semantics, and interruption.
- Assurance challenge: inject ambiguous authority, prompt attacks, stale sources, package revocation, and log-tampering attempts; test prevention, detection, reconstruction, and human response.
- Infrastructure accounting: measure accepted-task cost, endpoint and server energy, network and storage, maintenance labor, review, rework, recovery, and centralized services over months rather than extrapolating from token prices.
- Privacy and sovereignty audit: trace data flows, retention, access, deletion, keys, backups, provider dependencies, and operation under disconnection or service withdrawal.
- Peer-package pilot: test publication, offline verification, private overlays, cross-language adaptation, correction, revocation, licensing, and accountable certification across more than one institution.
- Independent replication: publish tasks, schemas, evaluation packs, failure taxonomies, system boundaries, and analysis code where confidentiality permits; repeat across domains, organizations, model families, and hardware classes.
12.11 Interpretive boundary
Current evidence establishes mechanisms and boundary conditions: local inference is feasible for bounded workloads; scaffolding changes performance and cost; productivity depends on complements; portability is not the same as semantic continuity; auditability requires a whole evidence chain; and centralized infrastructure is concentrated. It does not yet measure AIOS as a complete system.
The appropriate conclusion is not that the larger implications are excluded. It is that they form a coherent conditional research program. If local model capability, domain scaffolding, portable state, assurance, and user governance mature together, substantial movement of ordinary reasoning and durable intelligence toward local systems is economically and institutionally plausible. The scale, distribution, and net benefits depend on the empirical conditions set out above.
13. Foundational lineage: indispensable older work
The following sources predate the main research window and are included only because they frame the complementarity and assurance questions.
- Paul David, “The Dynamo and the Computer” (1990) uses historical comparison to explain why general-purpose technologies may show delayed productivity effects. Electrification required complementary factory redesign and capital replacement. It is a historical analogy, not an empirical model of AI.
- David and Wright, “General Purpose Technologies and Surges in Productivity” (1999) studies the manufacturing-productivity surge associated with electric unit drive and factory reorganization. It supports attention to production-system redesign rather than the input alone.
- Brynjolfsson and Hitt, “Beyond Computation” (2000) synthesizes firm-level evidence that computer investment is complementary to organizational practices and intangible capital. The technologies and firms are historical, but the complementarity framework remains relevant.
- Brynjolfsson, Rock, and Syverson, “The Productivity J-Curve” (2021) develops a model in which unmeasured intangible investment initially depresses observed productivity and later raises it. An empirical adjustment for computer hardware and software intangibles raises the estimated US TFP level by 15.9% by 2017. This is evidence about accounting and prior digital technologies, not a prediction that AI will follow the same curve.
- NIST AI 600-1, Generative AI Profile (July 2024) adapts the voluntary AI Risk Management Framework to generative-AI risks, emphasizing governance, content provenance, evaluation, incident disclosure, and third-party risks. It is a risk-management profile rather than a certification or performance study.
- ISO/IEC 42001:2023 defines a certifiable AI management system. Certification concerns an organization’s management processes within scope; it does not certify that a model is accurate or that every output is safe.
14. Conclusions
Strongest supported conclusions
- Fixed-capability retail inference became dramatically cheaper through 2024. The direction and rough order are supported by reproducible benchmark-price reconstructions. The series does not measure complete task cost or provider marginal cost.
- Agent design can dominate token, call, and energy use. Recent benchmark studies show one- to three-order-of-magnitude variation across workflows and runs. Evidence is strongest for selected coding and research-style tasks, not economy-wide use.
- On-device language-model inference is technically viable for bounded models and tasks. Peer-reviewed smartphone results and standardized client benchmarks support feasibility. Total-cost, lifecycle, and cross-workload economic superiority remain unmeasured.
- AI infrastructure is capital-, energy-, and geographically concentrated. IEA/LBNL estimates and competition-authority investigations consistently show rapid demand growth, clustered grid effects, hyperscaler scale economies, and contractual/technical dependence.
- Cloud portability is not semantic portability. Regulatory evidence establishes material technical and commercial switching barriers. Experimental evidence shows that model upgrades can require user adaptation. Full semantic switching cost is conceptually important but not yet directly measured.
- Productivity gains are conditional on task fit and complements. Randomized field studies find meaningful gains on selected tasks and losses or weak effects elsewhere. Administrative labor-market evidence does not yet show corresponding broad earnings or hours effects.
- Auditability is a system property. Local custody can improve evidence continuity, but only lifecycle logging, identity, authorization, evaluation, tamper evidence, monitoring, and accountable human review create an assurance basis.
- Open protocols reduce plumbing, not trust or meaning. MCP, A2A, and related efforts can lower interface costs. Semantic alignment, safe delegation, identity, conformance, and cross-organizational accountability remain open.
- Attestable domain packages are technically plausible; certified truth is not. Existing provenance, credential, research-object, and signing standards can bind files and claims to issuers. They do not validate knowledge without an explicit assurance institution.
- The electricity analogy is useful for complementarity and infrastructure, not for the nature of reasoning. AI outputs lack the homogeneity, stability, metering, and mature institutional framework of electricity.
- Bounded model judgments, deterministic effect controls, and explicit human authority are directionally consistent with current assurance practice. This supports evaluating the division of labor, but no comparative field evidence reviewed here establishes it as a generally superior architecture.
- More agents or roles are not monotonically better. A recent vendor-involved, peer-reviewed controlled study found that the value of coordination depended on task and base-model capability and imposed substantial reasoning-turn overhead. It does not directly test file-centered, role-specific workflows.
- Useful intelligence is economically a system-level outcome. Model capability, context, tools, scaffold, evaluation, and human use jointly affect accepted-task cost and quality. Current studies establish several component effects, but not the cumulative effect of the complete AIOS bundle.
Unresolved or contradictory evidence
- Unit inference prices and energy per task are falling while aggregate tokens, capital spending, and electricity use rise. The causal rebound elasticity is unknown.
- Large task-level productivity gains coexist with near-zero early effects on hours and earnings. Possible explanations include adoption lags, organizational complements, measurement limits, rent allocation, and time saved being absorbed rather than translated into output.
- Small models can be highly efficient locally, but the capability gap, engineering burden, hardware lifecycle, utilization, and human-review effects are not jointly measured.
- AI can narrow performance gaps and bridge functional vocabularies, yet dependence outside the capability frontier can reduce correctness and common models may homogenize outputs. Long-run expertise and diversity effects remain uncertain.
- More detailed logs improve reconstructability but can increase privacy, security, discovery, and surveillance risks. Appropriate minimization and retention are unresolved by “log everything.”
- Open agent protocols are proliferating faster than evidence about secure multivendor deployment, semantic conformance, or liability.
- Provenance can distinguish an intact signed package from an altered one, but communities still lack general methods for certifying accuracy, freshness, and fitness across contested domains.
- No comparative evidence reviewed here establishes that a central agent is inferior to coordinated role-specific components under equal model, context, inference-budget, tool, and evaluation conditions.
- The optimal boundary among model judgment, deterministic enforcement, and human review is likely task- and risk-dependent; its accuracy, latency, and maintenance tradeoffs have not been jointly measured.
- It is unresolved whether purpose framing, composed context, cognitive modes, layered memory, relational metadata, provenance, planning, verification, and integration produce cumulative savings or whether their orchestration overhead offsets the individual gains.
- Dynamic model routing could reduce cost, latency, privacy exposure, and unnecessary frontier use, but routing errors may allocate an under-capable model to a consequential step or add evaluation overhead that erases the savings.
Possible AIOS connection points for later consideration
These are design implications to evaluate, not claims that the reviewed work validates AIOS.
- Canonical local files as an exit asset. File-native custody may lower the semantic portion of switching cost if files, annotations, relationships, and history remain intelligible without the current model or orchestration provider.
- Execution routing rather than execution ideology. Local, edge, and remote models could be selected by cognitive movement, capability, sensitivity, latency, cost, energy, and assurance requirements while canonical files and accepted ground remain person-controlled.
- Cumulative complementarity accounting. Purpose, context, memory, relationships, provenance, planning, verification, and integration should be evaluated both separately and as a system. The relevant question is whether their interaction reduces accepted-task inference, retries, hallucination exposure, and review—not whether each component improves one benchmark in isolation.
- Cost-aware agents. Budgets should attach to accepted outcomes and expose token/call ceilings, repeated-context use, escalation paths, cache assumptions, latency, and review cost. Agents should not be trusted to predict their own total cost without external controls.
- Foreground/background separation. Short user-visible responses with bounded downstream work may reduce interactive latency, but background work needs explicit budgets, cancellation, evidence, and completion criteria to avoid invisible rebound.
- A common evidence grammar with local semantic overlays. A recurring Why–How–What structure might aid human inspection across domains, provided it is treated as a navigational grammar rather than proof that different domains share one ontology.
- Exact file-operation records. Deterministic read/write/edit/search operations, pre/post hashes, diffs, tool results, and human approvals could support reconstruction more reliably than prose transcripts alone.
- Portable evaluation packs. Each domain could retain regression cases, source expectations, prohibited actions, acceptance rubrics, and model-version comparisons alongside its files.
- Multiple-resolution memory with provenance. Summaries should point to source states and disclose compression or expiration, so portability does not turn lossy memory into untraceable accepted ground.
- Protocol adapters behind stable local contracts. MCP, A2A, or later protocols could be implementation choices beneath locally controlled schemas and authorization, reducing exposure to protocol or provider churn.
- Attestable domain packages. Content-addressed files, metadata, provenance, tests, signatures, issuer credentials, and revocation information could support bounded exchange. The label should describe the specific verified property, not imply certified truth.
- Human authority as an explicit control. High-impact transitions should preserve named decision rights, review records, and override paths rather than infer authorization from conversational fluency.
Claims that would be unsafe to make
The formulations below collapse conditional implications into universal or already-demonstrated results. They do not exclude the more carefully stated AIOS theses examined in Section 12.
- “Local inference is always cheaper, greener, more private, more secure, or more reliable than cloud inference.”
- “Inference-price declines prove that complete reasoning tasks will approach zero cost.”
- “Agentic rebound will necessarily exceed future efficiency gains,” or the converse.
- “Current energy projections are measurements of future electricity use.”
- “AI is becoming a standardized utility equivalent to electricity.”
- “Owning files eliminates cloud, model, embedding, workflow, or semantic lock-in.”
- “Open agent protocols guarantee interoperability, safe delegation, or shared meaning.”
- “A local transcript or file history is sufficient for regulatory auditability.”
- “Cryptographic signatures, provenance, or verifiable credentials certify that knowledge is true.”
- “A management-system certificate certifies every AI output.”
- “The productivity percentage from one task experiment generalizes to firms, sectors, or the macroeconomy.”
- “AI can replace teams because an AI-assisted individual matched a team in one bounded experiment.”
- “Frontier models already exercise universally superhuman semantic judgment.”
- “All current scaffolding suppresses frontier reasoning,” or “scaffolding is the sole determinant of usable model capability.”
- “Insufficient context is the sole or principal cause of agent failure.”
- “A local-first reasoning architecture guarantees a fixed or universal reduction in centralized infrastructure” without a declared boundary, measured baseline, and workload-matched counterfactual.
- “The cumulative AIOS components necessarily reduce inference, retries, hallucination exposure, or energy in every workload.”
- “Dynamic routing can always identify the cheapest adequate model without quality or assurance loss.”
- “Local knowledge ownership makes data centers unnecessary.”
- “Adjacent research or standards validate AIOS as an implemented system.”
15. Source table
Dates identify publication, release, or material revision within the research window. Status labels distinguish peer review and authority from emerging evidence.
| Date | Source and status | Research role | Direct link |
|---|---|---|---|
| Aug. 2024 | EU Artificial Intelligence Act — law/official text | High-risk lifecycle controls, logging, documentation, oversight | Regulation (EU) 2024/1689 |
| Aug. 2024; Jan. 2025 issue | Acemoglu, The Simple Macroeconomics of AI — peer-reviewed model-based analysis | Task-to-macro productivity scenarios and distributional implications | DOI |
| Dec. 2024 | Lawrence Berkeley National Laboratory, 2024 US Data Center Energy Usage Report — authoritative modeled estimate/scenarios | US historical estimate and 2028 electricity scenarios | Report landing page |
| Jan. 2025 | US Federal Trade Commission, AI partnerships 6(b) study — regulator staff report | Cloud/model contracts, exclusivity, spending ties, possible switching costs | Report PDF |
| Jan. 2025 | US FDA, AI-enabled device software guidance — draft, nonbinding guidance | Total-product-lifecycle assurance, validation, monitoring | Draft guidance PDF |
| Jan. 2025 | Bank for International Settlements, Governance of AI Adoption in Central Banks — authoritative sector-practice report | Accountability, inventories, three lines of defense | BIS report |
| Jan. 2025 update | Sigstore Bundle Format 0.3 — open software-signing specification | Portable signature, identity, timestamp, and transparency evidence | Specification |
| Mar. 2025 | Xu et al., Fast On-device LLM Inference with NPUs — peer reviewed, ASPLOS 2025 | Smartphone NPU performance and energy measurements | DOI |
| Mar. 2025 | Cottier et al., LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks — reproducible research-organization data analysis | Benchmark-threshold retail price trends and assumptions | Epoch AI analysis |
| Apr. 2025 | Stanford HAI, AI Index Report 2025, ch. 1 — authoritative synthesis | Fixed-capability API price decline and hardware trends | Chapter PDF |
| Apr. 2025 | Humlum and Vestergaard, Large Language Models, Small Labor Market Effects — working paper | Adoption linked to Danish administrative labor outcomes | BFI page and paper |
| Apr. 2025 | Erol et al., Cost-of-Pass — preprint | Expected price per correct benchmark result | arXiv |
| Apr. 2025 | International Energy Agency, Energy and AI — authoritative modeled estimate/scenarios | 2024 global demand, infrastructure and 2030 scenarios | IEA report |
| May 2025 | Dillon et al., Shifting Work Patterns with Generative AI — vendor-affiliated working paper | Six-month randomized workplace experiment | NBER working paper |
| May 2025 | OECD, Competition in the Provision of Cloud Computing Services — authoritative policy synthesis | Concentration, technical and commercial switching barriers | OECD report PDF |
| May 2025 | W3C Verifiable Credentials Data Model 2.0 — W3C Recommendation | Portable issuer claims and verification model | Recommendation |
| May 2025 | C2PA 2.2 — industry provenance specification | Signed manifests, content history, validation | Specification |
| June 2025 | Kim et al., The Cost of Dynamic Reasoning — preprint | Agent calls, measured GPU energy, diminishing returns | arXiv |
| June 2025 | Agent2Agent project transferred to Linux Foundation — open protocol/project | Capability discovery and agent task exchange | Linux Foundation announcement |
| June 2025 | OECD, Is Generative AI a General Purpose Technology? — authoritative evidence synthesis | Careful assessment of GPT criteria and uncertainty | OECD report |
| July 2025 | UK Competition and Markets Authority, cloud market investigation — regulator final report | Concentration, profitability, switching and licensing barriers | Case and final report |
| Aug. 2025 application | EU AI Act general-purpose AI obligations — official legal guidance | Provider documentation, copyright, systemic-risk, and incident duties | European Commission fact page |
| Sept. 2025 application | EU Data Act — law/official explanation | Cloud and edge switching, interfaces, export, charge removal | European Commission explainer |
| Sept. 2025; journal 2026 | Oviedo et al., Energy Use of AI Inference — vendor-authored bottom-up perspective; peer-reviewed journal version | At-scale inference energy model and test-time-compute scenarios | Microsoft Research page |
| Nov. 2025 | Gundlach et al., The Price of Progress — preprint | Performance-adjusted inference prices and efficiency decomposition | arXiv |
| Nov. 2025 | Du et al., Context Length Alone Hurts LLM Performance Despite Perfect Retrieval — peer reviewed, Findings of EMNLP; single study | Long-input degradation with controlled retrieval and filler | ACL Anthology |
| Dec. 2025 | Ecma NLIP standards suite — adopted Ecma standards | Additional agent-communication layer | Ecma announcement |
| Feb. 2026 | Cui et al., The Effects of Generative AI on High-Skilled Work — peer reviewed, randomized field experiments | Developer task completion across three firms | DOI |
| Feb. 2026 | NIST AI Agent Standards Initiative — official initiative, early-stage | Secure interoperability, identity, evaluation and standards coordination | NIST announcement |
| Mar. 2026 | Dell’Acqua et al., Navigating the Jagged Technological Frontier — peer reviewed, preregistered experiment | Inside/outside-frontier productivity and error | DOI |
| Mar. 2026 | NIST AI 800-4, Challenges to Monitoring Deployed AI Systems — authoritative workshop/literature study | Monitoring gaps across technical and institutional layers | NIST report page |
| Apr. 2026 | Bai et al., How Do AI Agents Spend Your Money? — preprint | Token use and cost variability in coding-agent trajectories | arXiv |
| Apr. 2026 | International Energy Agency, Key Questions on Energy and AI — authoritative retrospective estimates and forward scenarios | 2025 growth, capex, density, and 2030 central scenario | IEA report |
| Apr. 2026 | Jahani et al., Prompt Adaptation as a Dynamic Complement in Generative AI Systems — peer reviewed, preregistered experiments | User adaptation across model versions and task types | DOI |
| May 2026 | NIST, Building Evaluation Probes into Agentic AI — ongoing official research | Embedded evaluation and machine-readable agent audit trails | NIST project |
| June 2026 | Dell’Acqua et al., The Cybernetic Teammate — peer reviewed; firm-involved field experiment | Individuals/teams, functional silos, judged innovation output | DOI |
| June 2026 | LBNL, United States Data Center Energy Usage Report, 2025 — authoritative bottom-up scenarios | Planned hardware and cooling model for 2030 US demand | Report landing page |
| June 2026 | RO-Crate 1.3 — community recommendation | Portable files, entities, provenance, relationships | Specification |
| June 2026 | MLCommons, MLPerf Mobile v6.0 — industry benchmark release | Standardized Android LLM performance and accuracy | Release |
| July 2026 | Kim et al., Capable Language Models Can Outgrow the Benefits of Collaboration — peer reviewed; vendor-funded/involved controlled benchmark study | Matched-budget single-/multi-agent architecture comparison | Nature Machine Intelligence |
| July 2026 | OECD, Artificial Intelligence Markets — authoritative policy synthesis | Cross-layer concentration and vertical dependencies | Full report |
| July 2026 | Model Context Protocol, 2026-07-28 revision — open protocol specification | Stateless model/application access to tools, resources, prompts, and extensions | Specification |
| Current through Aug. 2026 | MLCommons, MLPerf Client — industry benchmark specification | Client LLM task, performance, and accuracy methodology | Benchmark documentation |
Foundational sources outside the main research window
| Date | Source and status | Research role | Direct link |
|---|---|---|---|
| 1990 | David, The Dynamo and the Computer — peer-reviewed historical perspective | Electrification delay and complementary reorganization | Bibliographic record |
| 1999 | David and Wright, General Purpose Technologies and Surges in Productivity — historical economic research | Factory electrification and production-system redesign | Oxford repository |
| 2000 | Brynjolfsson and Hitt, Beyond Computation — peer-reviewed evidence synthesis | IT, organizational practice, and intangible complements | DOI |
| 2021 | Brynjolfsson, Rock, and Syverson, The Productivity J-Curve — peer-reviewed model and empirical application | Intangible investment and delayed measured productivity | DOI |
| 2023 | ISO/IEC 42001 — international management-system standard | Certifiable organizational AI management process | ISO record |
| July 2024 | NIST AI 600-1, Generative AI Profile — voluntary official risk-management profile | Generative-AI risk, evaluation, provenance, incident practices | NIST PDF |