AIOS Proresearch
AIOS Intelligence System · Research Overview

Research library · Full memo

Locally Owned Reasoning Infrastructure: Economics and Institutional Implications

This memo reviews evidence current through 10 August 2026 on economics, organizational design, assurance, interoperability, and institutional structure. It is an analysis of system implications, not a market forecast or investment thesis.

Executive assessment

Locally owned reasoning infrastructure is best understood as a change in the boundary of the firm, household, or professional practice—not as a claim that all computation should move onto a personal device. It moves some context, memory, workflow state, policy, and execution closer to the person or institution that is accountable for them. Models and compute can remain local, remote, or hybrid. The economic question is therefore not simply “local model or cloud API?” It is which assets must remain portable and governable, which execution should be bought as a service, and where the costs of integration, verification, and switching should sit.

The stronger AIOS thesis examined in this memo is that increasingly capable local models, mature scaffolding, durable local context, and selective access to frontier models could satisfy a large share of ordinary personal and organizational reasoning needs. If substantially correct, the first-order effect would be to move accumulated context, expertise, memory, workflow, and decision authority from remote applications into person- or institution-owned systems. The second-order effects could include thinner centralized application layers, less routine remote inference, more private cognitive work, widespread domain-specific reasoning systems, new peer knowledge institutions, and broader access where bandwidth or cloud services are constrained. These are conditional implications—not current measurements—but lack of present system-level evidence is not evidence that the implications are negligible.

AIOS also treats useful intelligence as a system-level outcome rather than a property of model weights alone. Purpose framing, composed context, distinct cognitive modes, layered memory, relational metadata, provenance, bounded judgment, planning, verification, and post-response integration may each supply modest gains that compound into fewer retries, lower inference burden, and lower hallucination exposure per accepted task. Dynamic routing could then assign a local or frontier model according to the cognitive operation, capability need, privacy, latency, cost, and sensitivity. Present research measures several components of this mechanism separately; it does not establish their cumulative effect in one system. Economically, that unmeasured complementarity is central because it could either offset agentic rebound or, if the scaffold adds calls and coordination failures, intensify it.

The strongest empirical result is a divergence between unit price and task cost. At a fixed benchmark capability, public API prices fell extraordinarily quickly through 2024. Stanford’s 2025 AI Index reported that the cheapest system exceeding a GPT-3.5-level MMLU threshold fell from about $20 to $0.07 per million tokens between November 2022 and October 2024. An Epoch AI reconstruction found rates ranging from roughly 9-fold to 900-fold per year depending on task and performance threshold. These are retail-price/benchmark measurements, not estimates of provider cost, and they exclude most of the orchestration surrounding a useful organizational task.

Agentic systems create a plausible rebound mechanism. A software agent may call a model repeatedly, read and reread long contexts, invoke tools, branch, recover from failures, and ask stronger models to review weaker ones. Recent preprints report orders-of-magnitude differences in token use across agent designs and runs. One benchmark study found tool-using designs making 9.2 times as many model calls as a chain-of-thought baseline on average; another found agentic coding trajectories using about three orders of magnitude more tokens than single-turn code tasks. These findings are domain- and scaffold-specific, and they do not establish an economy-wide rebound elasticity. They do establish that falling price per token cannot be treated as falling cost per completed, verified task.

On-device inference is now technically credible for bounded workloads. Peer-reviewed systems work demonstrates large speed and energy gains from neural-processing-unit execution of small models on a narrow set of phones. Standardized client benchmarks now include 1B–8B parameter language models. But there is not yet a general empirical basis for claiming that local inference is always cheaper, greener, more private, or more reliable. The relevant comparison includes amortized hardware, engineering and update costs, thermals, battery wear, utilization, model capability, human review, and the avoided costs of latency, network dependence, data transfer, and compliance. Repeated, stable, privacy-sensitive tasks on already-owned hardware are the most favorable local case; bursty frontier workloads and high-throughput shared services are the most favorable remote case. Hybrid execution is the defensible default scenario.

Upstream infrastructure remains concentrated even when downstream knowledge is locally owned. The International Energy Agency estimates that data centers consumed about 415 TWh in 2024 and projects roughly 945 TWh in 2030 in its base case. Its 2026 update estimates growth of 17% in global data-center electricity use during 2025 and about 50% for AI-focused facilities, while projecting about 950 TWh by 2030. The United States, China, and Europe account for most current demand, and new facilities cluster at particular grid nodes. The US Federal Trade Commission, UK Competition and Markets Authority, and OECD all document capital intensity, hyperscaler concentration, technical and commercial switching barriers, and contractual ties between cloud firms and model developers. Local ownership of canonical files and accepted ground can improve bargaining position and continuity, but it does not by itself decentralize semiconductor fabrication, frontier training, cloud capacity, or electricity demand.

Switching costs have at least three layers. Technical switching costs arise from data movement, interfaces, orchestration, identity, and provider-specific services. Behavioral switching costs arise when prompts, tool schemas, evaluation thresholds, and error-handling routines must be retuned for a different model or version. Semantic switching costs arise when an institution’s meanings—annotations, exceptions, relationships, provenance, tacit categories, review rules, and authority boundaries—are embedded in a provider’s representations or workflows. Regulators and interoperability protocols address the first layer most directly. A 2026 peer-reviewed experiment showing that user prompt adaptation accounted for about half the performance gain from a model upgrade on a bounded image-replication task is evidence for behavioral complementarity, not a complete measure of semantic lock-in. “Semantic switching cost” should therefore be treated as a useful analytic construct whose measurement remains immature.

Productivity evidence argues against both frictionless adoption and categorical pessimism. Randomized field experiments report material gains on selected tasks: a pooled 26% increase in completed tasks across three developer experiments; faster and higher-quality work for consultants on tasks inside a model’s capability frontier; and individual product-development workers with AI reaching the judged output quality of unaided two-person teams in one firm experiment. The same research also finds uneven adoption, larger gains for less-experienced workers, performance losses outside the capability frontier, and limited change in meeting behavior or coordinated work. Danish administrative data show little detectable effect on earnings or hours in the first two years despite adoption. The coherent interpretation is that AI is often a complement to task redesign, evaluation, training, and organizational capital. It is not yet an observed autonomous substitute for those complements.

Local custody can make audit evidence easier to retain, inspect, and migrate, but auditability is a property of the whole sociotechnical process. Regulated assurance requires versioned inputs, sources, model and configuration identifiers, tool calls, file changes, identities, authorizations, tests, exceptions, approvals, and post-deployment monitoring. Neither locality nor a transcript proves correctness, fairness, safety, or compliance. Recent NIST work finds deployed-AI monitoring fragmented across operational, human, security, compliance, and societal concerns; the EU AI Act and sectoral guidance impose or propose lifecycle documentation, logging, human oversight, risk management, and monitoring obligations. A local-first architecture can support these obligations, but adjacent standards and studies do not validate any particular implementation.

Interoperable protocols are promising but incomplete institutional infrastructure. The Model Context Protocol standardizes a way for model applications to expose tools, resources, and prompts. Agent2Agent addresses capability discovery and task exchange between independent agents. W3C Verifiable Credentials, RO-Crate, C2PA, and Sigstore offer building blocks for issuer claims, research-object packaging, content provenance, and signed software attestations. Together they make an attestable domain package technically plausible: ordinary files plus a content-addressed manifest, provenance, schemas, licenses, tests, policy constraints, signatures, and issuer credentials. They do not make the package true, current, safe, or “certified” unless an accountable certification scheme defines claims, test methods, issuers, surveillance, revocation, and liability.

The “AI as electricity” analogy is useful only in a narrow economic sense. Both can be pervasive enabling inputs whose productivity contribution depends on complementary capital and organizational redesign; both create large infrastructure and grid effects. The analogy breaks at the point of the delivered service. Within a specified grid product and quality standard, electricity is metered as a largely fungible commodity. A reasoning output is probabilistic, versioned, context-dependent, and entangled with a user’s data, policies, tools, and authority. Electricity does not silently change its semantics after a provider update. AI is therefore closer to an evolving cognitive production system that runs on concentrated digital and electrical infrastructure than to electricity itself.

1. Method and evidence discipline

This memo prioritizes work published or materially updated from August 2024 through August 2026. It emphasizes peer-reviewed studies, official statistics, regulator investigations, and adopted standards. Preprints, working papers, vendor-authored studies, draft guidance, and scenarios are labeled. Older work appears only in a separate foundational-lineage section.

Five analytic labels are used:

Replication note. Peer review is not replication. No empirical result below is treated as independently replicated unless the text says so. Results based on one paper, one intervention, one firm, or one benchmark suite are identified in their limitation paragraphs; where a finding is especially easy to overgeneralize and no independent replication was located, that absence is stated directly.

Benchmark results are not silently generalized to institutions. A model score is not a completed-work outcome; a completed-work count is not quality-adjusted productivity; an energy estimate per query is not a provider’s fleet total; a signed artifact is not verified knowledge; and a regulation is not evidence that compliance has been achieved.

2. The economic object: ownership, execution, and complements

“Locally owned reasoning infrastructure” combines several choices that should be analyzed separately:

  1. Custody: who holds the canonical files, memory, annotations, provenance, policies, and relationship data.
  2. Control: who can inspect, version, authorize, modify, export, or retire those assets.
  3. Execution location: whether a particular inference or tool operation runs on a client device, local server, private cloud, or public service.
  4. Model sourcing: whether model weights are locally controlled, remotely served, or selected dynamically.
  5. Assurance: how behavior, changes, exceptions, and human decisions are evaluated and recorded.
  6. Interoperation: how tools, agents, and knowledge packages exchange capabilities and evidence.

These dimensions are often conflated. Canonical files and accepted ground can remain person-controlled while a remote model performs a difficult inference. Conversely, a local model can still create lock-in if the surrounding memory store, embeddings, schemas, or workflow engine cannot be exported intelligibly. “Local-first” is thus primarily a claim about default custody, inspectability, and continuity; it is not synonymous with offline-only inference or self-sufficiency from industrial compute.

The relevant unit of economic analysis is the accepted task outcome, not the token. Model weights are one productive input alongside context quality, memory, metadata, tools, workflow design, verification, exact integration, and human judgment. A weaker or smaller model embedded in a better system may have a lower cost per accepted outcome than a stronger model requiring repeated calls and rework; the reverse may hold when scaffolding overhead exceeds its gains. A minimal comparison is:

Effective cost per accepted task = compute and service charges + amortized hardware + engineering and operations + network and movement costs + evaluation and human review + expected failure/rework + compliance and continuity costs.

For local execution, the dominant term may be a sunk device at high utilization, or it may be scarce engineering time maintaining model variants across heterogeneous hardware. For remote execution, the token bill may be small while retained-data review, outage exposure, contractual restrictions, or provider-specific workflow dependencies dominate. Any comparison that omits the acceptance criterion and review burden is incomplete.

3. Inference prices, agentic demand, and the rebound question

3.1 What has been measured

Research question and method. How quickly has the posted price of a minimum level of model capability fallen? The Stanford AI Index 2025 uses public API prices and benchmark thresholds to compare the cheapest available model at a fixed capability. It reports that the price of a system scoring at least 64.8 on MMLU fell from $20 per million tokens in November 2022 to $0.07 in October 2024, a decline greater than 280-fold. For a greater-than-50% threshold on GPQA, the minimum observed price fell from $15 per million tokens in May 2024 to $0.12 in December 2024. These are unusually rapid declines.

The underlying Epoch AI analysis combines public prices with benchmark observations, assumes a 3:1 mix of input to output tokens, selects the cheapest model above each task-specific threshold, and fits log-linear trends. Across six benchmark/threshold combinations, it estimates declines from roughly 9-fold to 900-fold per year, with a median around 50-fold over the full observation period and faster estimates after January 2024. Epoch makes code and observations available, which improves inspectability.

Limits and relevance. These series measure posted retail prices at selected benchmark levels. They do not identify marginal provider cost, discounts, capacity reservations, batching, caching, latency, reliability, output length, orchestration, or review. Benchmark coverage is partial and vulnerable to contamination and saturation. Reasoning models were excluded from the initial Epoch reconstruction. Selecting a moving “cheapest adequate model” also combines engineering progress, new entry, and pricing strategy. The series is directly relevant to the price of substitutable model calls, but only indirectly relevant to the cost of an accepted organizational result.

A later preprint, Gundlach et al., “The Price of Progress”, asks how fixed-performance inference price changed and how much of the change can plausibly be separated into competition, hardware, and algorithmic efficiency. It reconstructs performance-adjusted price trends with a larger historical combination of Artificial Analysis pricing and Epoch benchmark data. It estimates price declines of approximately 5–10-fold per year for fixed performance in knowledge, reasoning, mathematics, and software tasks. Restricting parts of the analysis to openly available models and accounting for hardware-price improvement, the authors estimate algorithmic efficiency gains near three-fold per year.

Limits and relevance. This is a preprint. Its decomposition depends on model coverage, benchmark comparability, posted prices, hardware assumptions, and fitted functional forms. It also distinguishes cheap fixed performance from frontier performance: the cost of evaluating or serving the most capable available system need not fall on the same trajectory. The decomposition is relevant to long-run sourcing assumptions, but is too uncertain to serve as a budgeting forecast.

3.2 Cost of a correct answer

Price per token is a poor sufficient statistic when models vary in accuracy and verbosity. The preprint Erol et al., “Cost-of-Pass”, defines the expected monetary cost of obtaining a correct solution and compares models and inference strategies on quantitative and knowledge benchmarks. It finds that model rankings change when accuracy is combined with price, and that majority voting or self-refinement often supplies too little marginal accuracy to justify its added cost in the evaluated settings.

Limits. “Correct” is benchmark-defined and often easier to determine than acceptance in professional work. API prices and model behavior change rapidly. The metric does not include latency, integration, review, or downstream damage from an accepted error. Its main relevance is conceptual: even a narrow task should be priced per accepted outcome rather than per token.

3.3 Agentic token demand

The preprint Kim et al., “The Cost of Dynamic Reasoning”, compares chain-of-thought, ReAct, Reflexion, LATS, and LLMCompiler agent designs using Llama 3.1 8B and 70B models on question-answering and interactive benchmarks. Tool-using agents made 9.2 times as many model calls as the chain-of-thought baseline on average; LATS averaged 71 calls per request. On HotpotQA, selected configurations increased measured GPU energy from 0.32 Wh for a single 8B call to 22.76 Wh for LATS and 41.53 Wh for Reflexion; corresponding 70B figures rose from 2.55 Wh to 158.48 Wh and 348.41 Wh. Accuracy improvements showed diminishing returns.

Method and sample. The authors run five designs on A100 GPUs, measure power with NVIDIA tooling, and use 50-question samples for key design comparisons. They extrapolate several illustrative fleet scenarios.

Limits and relevance. This is a preprint using older model families, selected configurations, no production batching, and GPU-only energy. Its fleet totals are scenarios, not observed consumption. The measured per-workflow result is nevertheless important: architecture can change energy and call counts by two orders of magnitude even before a task leaves the benchmark.

The preprint Bai et al., “How Do AI Agents Spend Your Money?”, analyzes eight frontier models on 500 SWE-bench Verified software-maintenance tasks. It reports that agentic coding uses roughly a thousand times the tokens of code-reasoning or chat-style requests, that repeated runs of the same task can differ by up to 30-fold in token use, and that higher consumption does not monotonically improve success. Models are also poor at predicting their own eventual cost.

Limits and relevance. The result is specific to software maintenance, the selected scaffold, benchmark, and current model behavior. Tokens are not identical to billed tokens when caches and discounts apply, and repository-scale contexts are unusually large. It is direct evidence of high within-task cost dispersion and weak self-budgeting, not an estimate of typical office-agent use.

3.4 Rebound: supported mechanism, unmeasured elasticity

The rebound hypothesis is that cheaper inference encourages longer contexts, more attempts, larger models, more users, more background agents, and more verification, so aggregate resource use grows even as unit cost falls. The IEA’s 2026 assessment reports both rapid improvement in energy per AI task and rising aggregate consumption; it estimates 2025 electricity growth of 17% for data centers overall and about 50% for AI-focused facilities. This coexistence is consistent with rebound.

It is not, by itself, a causal estimate of rebound. Aggregate growth also reflects training, new capacity, conventional cloud workloads, demand growth, model scaling, redundancy, and geographical changes. No source reviewed here estimates a stable economy-wide elasticity linking an exogenous fall in inference price to total agentic tokens or electricity use. The safe statement is:

Falling price and energy per inference are real; agent designs can multiply inference demand; and aggregate data-center energy is rising. The extent to which the second causally offsets the first remains unresolved.

The AIOS system-level thesis introduces a countervailing mechanism. Better context composition, memory selection, provenance, bounded actions, verification, and routing could reduce failed attempts, oversized prompts, repeated retrieval, and unnecessary frontier calls. Those gains could lower tokens and energy per accepted task even when the selected model is unchanged. But the components can also add orchestration steps, review calls, caches, indexes, and background work. No reviewed study estimates the cumulative net effect of this full bundle. It should therefore be represented as a potentially important efficiency complement—and possible rebound counterforce—not assumed savings.

A vendor-authored, bottom-up perspective from Microsoft researchers, Oviedo et al., “Energy Use of AI Inference”, estimates a median of 0.34 Wh for a frontier-model query under an at-scale H100 configuration and about 4.32 Wh in a scenario with 15 times more test-time tokens. It argues that batching, hardware utilization, quantization, model routing, and hardware improvement can reduce energy by 8–20 times in combined scenarios.

Limits. The figures are modeled estimates rather than a provider fleet disclosure. Configuration, query, facility overhead, hardware, and utilization assumptions strongly affect results. The paper usefully demonstrates sensitivity to test-time compute and deployment efficiency; it should not be quoted as a universal “energy per prompt.”

4. On-device economics and the case for hybrid execution

4.1 Technical evidence

The strongest recent peer-reviewed systems evidence is Xu et al., “Fast On-device LLM Inference with NPUs”, published at ASPLOS 2025. The research question is whether commodity smartphone neural processing units can overcome the memory movement and operator constraints that make local language-model inference slow and energy-intensive.

Method and sample. The authors implement an NPU-oriented inference system and test it on two Android phones: a Snapdragon 8 Gen 3 device with 24 GB of memory and a Snapdragon 8 Gen 2 device with 16 GB. They evaluate five models from 1.8B to 7B parameters across language understanding, long-context, question-answering, and mobile task benchmarks, compare five CPU/GPU baselines, repeat runs three times, and sample device power at 100 ms intervals.

Principal finding. Depending on model and baseline, the system reports 7.3–18.4 times CPU speedups, 1.3–43.6 times GPU speedups, 1.9–59.5 times lower energy, less than one percentage point of accuracy loss, and 1.4–32.8 times end-to-end speedups. Prefill can exceed 1,000 tokens per second in favorable cases.

Limitations and relevance. Two Qualcomm phones, five relatively small models, a specialized runtime, and selected benchmarks cannot establish general device economics. Results are strongest for prefill; memory, decoding, thermals, battery aging, hardware fragmentation, and longer production sessions remain constraints. No cloud total-cost comparison is attempted, and no independent replication of the paper’s exact speed and energy claims was identified. The paper demonstrates feasible and potentially efficient local execution for bounded models, not a universal economic advantage.

MLCommons’ MLPerf Client and MLPerf Mobile v6.0 extend standardized measurement to client language models. Client tests include Llama 3.1 8B and Phi-3.5 Mini task suites with accuracy constraints; Mobile v6.0 includes Llama 3.2 1B/3B and Llama 3.1 8B on Android. Power methodology remains less mature than performance measurement, and some devices cannot run the largest or longest components. These are capability and performance benchmarks, not evidence of lower lifetime cost.

4.2 A workload-level break-even model

For a workload class \(w\), a useful break-even condition is:

\[ C_{local,w} = \frac{H_w}{N_w} + E_w + O_w + Q_w + F_w \]

\[ C_{remote,w} = P_w + X_w + L_w + G_w + Q'_w + F'_w \]

where \(H\) is the allocable hardware and deployment cost, \(N\) accepted tasks over its useful life, \(E\) incremental energy, \(O\) local operations and updates, \(P\) remote inference price, \(X\) data movement and egress, \(L\) latency and availability cost, \(G\) governance/contractual burden, \(Q\) evaluation and human review, and \(F\) expected failure and rework. The symbols are not intended to suggest that all terms are easily monetized. They identify omitted variables in simple token-price comparisons.

Interpretation. Local execution is favored by high repeated volume, a stable narrow task, acceptable small-model performance, already-owned hardware, strict latency or offline requirements, sensitive inputs, and expensive network or contractual review. Remote execution is favored by irregular demand, need for frontier capability, very long contexts, shared high-utilization hardware, rapid model upgrades, and organizations that cannot maintain heterogeneous runtimes. A local-first knowledge layer with dynamic routing can preserve option value: ordinary files and evidence remain controlled while execution moves according to task requirements.

4.3 Externalities and distribution

Moving inference to a device can reduce network traffic and data-center demand for the displaced task, but it can also lower utilization, duplicate model storage, accelerate device replacement, or move electricity to less efficient hardware and grids. Conversely, hyperscale facilities may achieve higher accelerator utilization and lower cooling overhead while creating concentrated grid, water, land, and transmission burdens. Comparative environmental claims require a lifecycle boundary, counterfactual utilization, regional electricity mix, embodied hardware, and measured task acceptance. “Local is greener” and “cloud is greener” are both unsafe as general propositions.

5. Capital, energy, and geographic concentration

5.1 Current and retrospective estimates

Research question and method. The IEA asks how much electricity data centers currently use, where that demand is concentrated, and how equipment, facility, and deployment trends could change it. Its global model combines historical and projected server shipments—including accelerated servers—with power, utilization, cooling, facility, and regional assumptions, supplemented by company reporting, project pipelines, geospatial analysis, and consultation. It is a bottom-up accounting model rather than a comprehensive meter census.

The IEA’s 2025 Energy and AI report estimates that data centers used about 415 TWh of electricity in 2024, or roughly 1.5% of global consumption, after growing about 12% annually over the prior five years. The IEA’s 2026 update estimates approximately 485 TWh in 2025 and year-over-year growth of 17%, with AI-focused data centers growing about 50%. These are retrospective modeled global estimates built from equipment, utilization, and facility data; no complete global meter census exists.

The IEA also reports strong geographic concentration. The United States, China, and Europe account for the large majority of current data-center electricity demand, while facilities cluster much more tightly than national electricity totals suggest. Power density has risen quickly, placing connection, generation, transmission, and cooling constraints on particular localities.

Capital concentration is similarly pronounced. The 2026 IEA report states that five large technology firms spent more than $400 billion in capital expenditure in 2025 and announced plans implying a further large increase in 2026. Corporate spending and announced plans are observable, but allocation specifically to AI data centers is estimated and plans can be revised.

Limitations and relevance. Category boundaries between AI-focused and conventional facilities are estimated, provider disclosure is incomplete, and utilization changes rapidly. The estimates are directly relevant to infrastructure concentration and grid planning; they cannot allocate energy reliably to one user response, one agent architecture, or one local-versus-remote counterfactual.

5.2 Projections, not measurements

The IEA’s central scenario reaches about 945–950 TWh of data-center electricity demand in 2030, close to twice 2024–2025 levels, with AI-focused demand roughly tripling. This is conditional on deployment, hardware, utilization, and efficiency assumptions.

The 2024 Lawrence Berkeley National Laboratory US Data Center Energy Usage Report estimates that US data centers accounted for about 4.4% of electricity in 2023 and develops scenarios reaching 6.7–12% in 2028. A 2026 LBNL update uses planned equipment shipments, per-device energy, and cooling simulations to construct 2030 scenarios of 9.5–15.3% of US electricity, with 11.8% as its central case. These are increasingly detailed bottom-up scenarios, not guaranteed demand.

5.3 Institutional consequences

Concentrated capital and energy create at least five institutional effects:

Locally controlled knowledge does not remove these dependencies, but it can reduce the amount of organizational meaning that must be stranded when execution providers change. That is an option-value argument, not proof of lower day-to-day cost.

6. Cloud switching, lock-in, and semantic switching costs

6.1 Regulatory evidence on technical and commercial barriers

The UK Competition and Markets Authority’s cloud-services market investigation, finalized in July 2025, drew on provider submissions, customer evidence, financial analysis, and technical examination. It found highly concentrated infrastructure-as-a-service and platform-as-a-service markets; substantial sustained returns for the largest providers; and adverse effects from egress fees, technical barriers to switching and multicloud use, and Microsoft software-licensing practices. The inquiry examined committed-spend agreements but did not find that their current forms harmed competition. Customers reported that switching or using multiple clouds could require architecture changes, retraining, duplicated operations, and data movement.

Limitations and relevance. The inquiry concerns UK cloud markets and the evidence period preceding the final report. Product boundaries and remedies are jurisdiction-specific. It is authoritative evidence that practical switching is much more than copying data, and that nominal multicloud availability does not imply frequent or economical switching.

The OECD’s 2025 competition paper on cloud computing synthesizes market studies across member jurisdictions. It reports that the two largest providers can account for up to 80% in some major OECD markets, and identifies economies of scale and scope, capital intensity, chip access, proprietary APIs, software dependencies, discounts, and egress costs as reinforcing barriers. The OECD’s 2026 report on AI markets maps similar dependencies across chips, cloud, models, data, energy, talent, and downstream applications.

Limitations. Both reports synthesize heterogeneous jurisdictions and market definitions. They identify mechanisms and recurring findings rather than estimating one universal concentration or switching coefficient.

The FTC’s January 2025 staff report on AI partnerships uses compulsory information requests concerning Alphabet–Anthropic, Amazon–Anthropic, and Microsoft–OpenAI relationships. It documents equity and revenue-sharing interests, consultation and control rights, exclusivity provisions, cloud-spending commitments, and exchanges of information or resources. Staff identifies possible switching costs, input foreclosure, and access to sensitive technical and business information.

Limitations and relevance. This is a competition staff study, not a judicial finding that any arrangement is unlawful or caused a specified price effect. It directly supports the institutional point that model and cloud markets are linked through contracts, not merely technical performance.

The EU Data Act, applicable from September 2025, requires cloud and edge providers to reduce switching obstacles, provide export and interface support, and—subject to staged implementation—eliminate switching charges. It is an important legal response to technical and commercial portability. It does not ensure equivalent model behavior or preserve an organization’s tacit meanings.

6.2 A three-layer switching model

LayerWhat must moveTypical frictionWhat current policy/protocols address
TechnicalFiles, databases, identities, APIs, workflows, compute imagesEgress, proprietary services, incompatible interfaces, downtimeCloud-switching rules, open interfaces, container and tool protocols
BehavioralPrompts, tool descriptions, schemas, thresholds, recovery logic, evaluationsModel-specific sensitivities, version drift, different errors and verbosityEvaluation suites, adapters, prompt migration, conformance tests
Semantic/organizationalCategories, annotations, provenance, exceptions, relationships, memory, decision rightsTacit knowledge, lost context, representation mismatch, changed authorityOnly partially addressed by portable files, explicit metadata, governance and human review

The term semantic switching cost is used here for the third layer and for the portion of the second layer caused by meaning rather than syntax. It is not yet a standardized economic statistic.

A recent peer-reviewed study offers a partial empirical anchor. Jahani et al., “Prompt Adaptation as a Dynamic Complement in Generative AI Systems”, published in Information Systems Research in 2026, asks how users’ prompts adapt when the underlying generative model improves.

Method and sample. Two preregistered experiments assigned 3,750 participants to image-replication and open logo-design tasks, producing about 37,000 prompts. The design compared DALL-E versions, replayed prompts across models, and separated gains due to the model from gains due to changed user behavior.

Principal finding. On the bounded replication task, prompt adaptation accounted for about half of the performance improvement from DALL-E 2 to DALL-E 3. On the open creative task, user adaptation accounted for about 7%, with more than 90% attributed to the model upgrade. Automated prompt rewriting harmed the bounded task and provided only a small, statistically uncertain benefit on the open task.

Limitations and relevance. The study concerns text-to-image generation, one product transition, short sessions, and tasks without durable organizational memory. Some authors have Microsoft/OpenAI relationships, and no independent replication of the estimated adaptation shares was identified. It shows that model performance and user routines can be dynamic complements, especially where an output target is precise. It does not quantify the cost of migrating a file-native knowledge system or establish semantic lock-in as a general market metric.

Emerging preprints on prompt migration and cross-model prompt drift report that structured adaptation can recover performance after model changes, but current results are case-specific and not an adequate foundation for a universal migration-cost claim.

6.3 Why ordinary files help but do not solve semantics

Ordinary local files can lower switching cost when they remain authoritative, documented, versioned, and usable without a particular model. Companion metadata can preserve provenance and relationships without forcing all meaning into a provider’s hidden state. Exact read/write/edit operations can make state transitions inspectable. These are design hypotheses with strong portability logic.

They do not eliminate switching costs. Embeddings must be regenerated; retrieval rankings change; tool descriptions interact differently with different models; long-lived annotations accumulate local conventions; and a new model may obey the same schema with a different distribution of errors. Portability therefore requires behavioral regression tests and semantic stewardship, not just export.

7. Productivity complements and organizational design

7.1 What recent field evidence establishes

Cui et al., “The Effects of Generative AI on High-Skilled Work”, published in Management Science in 2026, asks whether access to an AI coding assistant changes the output of professional software developers.

Method and sample. The authors pool three randomized field experiments at Microsoft, Accenture, and a Fortune 100 firm, covering 4,867 developers. Treatment was access to the assistant; the primary outcome was completed work items in operational repositories rather than a laboratory coding score.

Principal finding. The pooled estimate is a 26.08% increase in completed tasks, with a standard error of 10.3%. Adoption and effects were larger among less-experienced developers. Results varied across the three firms, and individual experiments were noisier than the pooled number.

Limitations and relevance. Completed tasks are not a complete quality-adjusted productivity measure; work-item definitions, code review, task selection, and organizational context differ. The study covers an early tool and developer work at large firms, including the product’s vendor. Its pooled design supplies within-paper, multi-site replication, but no independent replication of the precise effect was identified. It is strong evidence of task-level gains under some conditions, not a general 26% firm-productivity parameter and not evidence about local-first architectures.

Dell’Acqua et al., “The Cybernetic Teammate” and “Navigating the Jagged Technological Frontier” are often discussed together but concern distinct experiments. The peer-reviewed Organization Science “jagged frontier” study asks how GPT-4 affects knowledge workers on tasks inside and outside the system’s capabilities.

Method and sample. A preregistered randomized experiment assigned 758 Boston Consulting Group consultants to no-AI, AI, or AI-plus-prompt-overview conditions. Participants completed realistic consulting tasks selected to lie inside the model’s capability frontier and a task selected outside it.

Principal finding. For inside-frontier work, AI users completed 12.2% more tasks, worked 25.1% faster, and produced higher judged quality. For the outside-frontier task, AI users were 19 percentage points less likely to reach the correct answer. Lower baseline performers received larger inside-frontier gains.

Limitations and relevance. Data were collected in 2023 using an older model. Participants were consultants, the task battery was curated, and the outside-frontier inference rests heavily on one problem type. No independent replication of the precise inside-/outside-frontier effects was identified. The study’s durable contribution is not its precise percentages but the interaction between task fit, reliance, and evaluation.

The separately published “Cybernetic Teammate” study, revised in 2026, asks whether generative AI changes the performance and functional composition of individuals and teams doing product-innovation work.

Method and sample. A preregistered field experiment randomly assigned 791 Procter & Gamble professionals, spanning commercial and research-and-development functions, to work individually or in two-person teams, with or without AI, on real product-development problems. Outputs were evaluated by expert judges.

Principal finding. Individuals using AI produced outputs judged comparable in quality to unaided human teams. AI also reduced the distinction between commercial and technical proposals: individuals generated more balanced solutions across functional domains. Participants reported more positive emotional experience.

Limitations and relevance. The setting is one firm, one workshop-like innovation process, and judged proposals rather than implemented products or long-run outcomes. Several authors were company employees; the article reports research funding from Harvard Business School. No independent replication of the individual-versus-team result was identified. The result supports the possibility that AI can bridge functional vocabularies; it does not show that teams can generally be removed, that knowledge becomes organization-independent, or that outputs remain diverse over time.

A large vendor-affiliated working paper, Dillon et al., “Shifting Work Patterns with Generative AI”, asks whether integrated generative AI changes how office workers allocate time across email, documents, and meetings.

Method and sample. A six-month randomized field experiment across 66 firms offered an integrated assistant to about half of 7,137 knowledge workers. The study uses telemetry and surveys, analyzing both assignment and actual usage.

Principal finding. In the paper’s November 2025 revision, the 80% of treated workers who used the tool in the second half of the experiment spent about two fewer hours per week on email and reduced work outside regular hours. Apart from these individual time savings, the authors did not detect changes in the quantity or composition of tasks from individual-level access.

Limitations and relevance. Several authors are Microsoft researchers and the intervention is a Microsoft product. Access is not use, behavior is not output quality, and early adoption may differ from equilibrium practice. No independent replication of the reported time-allocation effects was identified. The contrast between substantial changes in individual email work and little change in meetings is consistent with coordinated routines requiring organizational redesign rather than a personal tool alone.

Humlum and Vestergaard, “Large Language Models, Small Labor Market Effects”, is a working paper that asks whether rapid chatbot adoption translated into earnings and hours effects in highly exposed occupations.

Method and sample. The authors link representative surveys of adoption and employer encouragement to Danish administrative records for workers and workplaces across 11 exposed occupations, and use difference-in-differences and event-study designs.

Principal finding. Over the first two years after widespread chatbot availability, the estimates for earnings and recorded hours are close to zero and sufficiently precise to rule out effects much larger than about 2%. The study also finds adoption, task restructuring, and occupational movement, particularly where employers encouraged use.

Limitations and relevance. Denmark’s labor institutions and wage setting may not generalize. Earnings and hours are lagging, coarse outcomes; quality, consumer surplus, work intensity, and unrecorded time savings may change first. The study is a useful counterweight to extrapolating task experiments directly into aggregate labor-productivity or displacement claims.

7.2 Complements, not an automatic production function

Taken together, these studies support four propositions:

  1. AI can raise throughput or judged quality on selected, well-scoped tasks.
  2. Effects vary by baseline skill, task position relative to model capability, adoption, and evaluation.
  3. Personal work patterns may change before meetings, approval chains, incentives, and cross-functional coordination.
  4. Short-run task gains do not automatically appear in earnings, hours, firm productivity, or macroeconomic statistics.

The required complements include training, task decomposition, reliable source access, evaluation criteria, workflow integration, review allocation, incident handling, and incentives. Durable files and explicit process state may make these complements easier to build and preserve. They can also add overhead. Whether the balance is positive is an empirical question for each workflow.

7.3 Standardization with local variation

Organizations face a real design tension. Too little standardization produces incomparable records, duplicated integrations, unsafe permissions, and weak assurance. Too much central standardization erases domain distinctions, creates brittle universal prompts, and pushes tacit exceptions into shadow practices.

A defensible architecture separates a common control plane from locally governed semantic content:

Common organizational grammarDomain-owned variation
Identity, authorization, timestamps, versioningVocabulary, rubrics, exceptions, annotations
Provenance and evidence referencesDomain source selection and relationship meaning
Tool-call and file-change recordsTask decomposition and review criteria
Minimum evaluation and incident schemaRisk thresholds beyond the minimum
Export, package, signature, and retention formatsLocal model choice and context composition
Human approval and escalation hooksAllocation of expert authority

This is an institutional design proposal, not a measured optimum. The recent field evidence makes it plausible: AI can cross functional vocabularies, but coordinated practices change slowly and errors depend on task context. The correct object of standardization is therefore often the evidence and control contract, not a single universal way of reasoning.

Local variation also requires governance. Without shared minimum controls, “local autonomy” can become inconsistent risk acceptance, inaccessible records, or model proliferation. Without local authority, “enterprise alignment” can become semantic flattening and dependency on a central team that cannot evaluate every domain. Successful variation is bounded, testable, and revocable.

8. Auditability and regulated-industry assurance

8.1 What authorities currently require or recommend

The EU AI Act, in force since August 2024 with phased application, provides a legal reference point for high-risk systems. Its framework includes risk management, data governance, technical documentation, automatic record-keeping, transparency to deployers, human oversight, accuracy, robustness, cybersecurity, quality management, conformity assessment, and post-market monitoring. Requirements differ by role and use case, and implementation dates and supporting standards continue to develop. The Act is a legal obligation in scope, not an assurance method by itself.

The European Commission’s general-purpose AI obligations, applicable in stages from August 2025, add technical documentation, downstream information, copyright-policy, and training-content-summary duties, with risk evaluation, incident reporting, and security measures for systemic-risk models. Local downstream custody does not remove provider or deployer duties.

NIST AI 800-4, “Challenges to Monitoring Deployed AI Systems”, asks what makes post-deployment monitoring difficult in practice.

Method and sample. NIST combined a literature review with three practitioner workshops held during 2025, then thematically coded challenges across functionality, operations, human factors, security, compliance, and large-scale impacts.

Principal finding. Monitoring is fragmented across disciplines and tools. Organizations struggle to determine what to log, connect technical indicators with human and compliance effects, manage scale, and relate monitoring evidence to audit and governance processes.

Limitations and relevance. This is an authoritative exploratory report, not a controlled trial or a certification standard. Workshop participants are not a representative sample of all deployers. It directly supports a lifecycle view: pre-deployment evaluation alone is inadequate, and a useful record must join technical, operational, and human evidence.

NIST’s 2026 “Building Evaluation Probes into Agentic AI” project is ongoing research, not an adopted standard. Its prototype research-agent pipeline emits machine-readable records and probes citation faithfulness, completeness, and sufficiency. The project is relevant because evaluation is embedded into the workflow rather than appended after the fact. Automated judges and probes remain fallible and require external validation.

The US Food and Drug Administration’s January 2025 draft guidance for AI-enabled medical devices proposes a total-product-lifecycle approach: intended-use specification, data management, bias analysis, validation, transparency, risk management, and post-market performance monitoring. It is explicitly draft, nonbinding guidance. FDA’s subsequent final guidance on predetermined change-control plans addresses how certain planned model changes can be specified, validated, and controlled. These materials are sector-specific; their broader lesson is that change management and monitoring are assurance objects, not administrative afterthoughts.

The Bank for International Settlements’ 2025 report on central-bank AI governance synthesizes practices among member central banks in the Americas. It recommends adaptive governance, clear accountability, human oversight, inventories, tiered risk management, and three lines of defense. It is an institutional practice report rather than causal evidence, but it shows why confidential data, operational resilience, and public legitimacy make governance inseparable from deployment design.

8.2 The minimum evidence chain

For consequential reasoning workflows, an auditable event should be able to identify, at an appropriate level of disclosure:

Full hidden reasoning traces are neither necessary nor sufficient. They may be unavailable, misleading as explanations, sensitive, or too costly to retain. Assurance should prioritize observable inputs, evidence, actions, tests, and accountable decisions.

8.3 Locality: an assurance affordance, not an assurance result

Local custody can improve the chain of evidence by keeping canonical sources, annotations, diffs, and decisions under an accountable institution’s retention and access controls. It can reduce dependence on a provider’s log format or retention policy, and make inspection possible after a service or model is retired.

The same local environment can be poorly secured, silently modified, incompletely logged, or governed by one person without separation of duties. Remote services may supply stronger physical security, immutable logging, validated controls, and professional operations. The correct claim is conditional:

Local-first custody can make evidence continuity and inspection easier when it is paired with identity, authorization, tamper evidence, retention, evaluation, and independent review. Locality alone proves none of those properties.

For regulated sectors, assurance must attach to a defined intended use and deployed configuration. A general-purpose model certificate cannot substitute for workflow validation; a workflow log cannot substitute for clinical, financial, safety, or legal performance evidence; and a management-system certification cannot guarantee each output.

9. Interoperable agents and attestable domain packages

9.1 Protocol layers

The current Model Context Protocol specification, revision 2026-07-28, defines JSON-RPC-based, stateless requests through which model-facing applications can discover and invoke tools, read resources, and use prompts. It also defines an opt-in extension framework, including long-running tasks. Its principal economic value is interface reuse: a capability can be exposed once rather than integrated independently with every model application. Its security value depends on authentication, authorization, consent, validation, and deployment practice outside the core message grammar. The breaking move from a stateful 2025 core to the stateless 2026 revision also illustrates that the protocol layer itself still carries migration cost.

The Agent2Agent protocol, transferred to the Linux Foundation in June 2025, addresses communication between independent agents. Agent Cards advertise capabilities; task, message, artifact, and status structures support longer-running exchanges; authentication is intended to use standard web mechanisms. It complements rather than replaces tool protocols: one addresses agent-to-agent delegation, the other agent-to-resource access.

The Ecma NLIP standards suite, approved in December 2025, and several developing IEEE and ITU work items reflect parallel efforts to standardize agent communication. The NIST AI Agent Standards Initiative, launched in February 2026, explicitly treats interoperability, security, identity, and evaluation as unfinished work.

Protocol specifications establish message and capability contracts. They do not establish:

Open protocols can reduce integration and some switching costs while increasing the number of callable endpoints and trust boundaries. Security, semantic profiles, test suites, and institutional accountability must mature alongside adoption. Claims of a settled universal agent layer are premature.

9.2 Building blocks for portable knowledge packages

Several standards provide reusable components without being designed as one universal “knowledge package” format:

A possible attestable domain package could combine:

  1. ordinary, content-addressed files;
  2. a versioned manifest describing entities and relationships;
  3. source citations, acquisition dates, licenses, and transformation provenance;
  4. schemas and declared semantic profiles;
  5. intended uses, exclusions, policy constraints, and risk classification;
  6. executable conformance tests and benchmark claims with environments;
  7. issuer identity, signature, credential, timestamps, and status/revocation endpoints;
  8. review history and accountable human approval;
  9. privacy partitions separating distributable content from local overlays.

This is a design scenario assembled from adjacent standards, not an adopted cross-domain standard or demonstrated market institution.

9.3 Peer-to-peer distribution and institutional trust

Content addressing permits packages to be cached or exchanged peer to peer while a digest continues to identify the same bytes. Local overlays can preserve private annotations while shared base packages remain verifiable. This could reduce repeated downloads, improve resilience, and let professional communities distribute bounded corpora, methods, or evaluation suites without centralizing every use event.

The hard problems are institutional rather than transport-level:

“Certified knowledge” is therefore unsafe terminology unless a named scheme has an accredited or otherwise accountable issuer, explicit scope, test method, surveillance, appeals, revocation, and liability. A cryptographic signature establishes integrity and a binding to an issuer; a verifiable credential establishes that an issuer made a claim; provenance establishes a history claim. None establishes truth.

10. Evaluating the “AI as electricity” analogy

10.1 What the analogy explains

The OECD’s 2025 assessment, “Is Generative AI a General Purpose Technology?”, evaluates generative AI against economic criteria such as pervasiveness, continuing technical improvement, and innovation complementarities. It concludes that generative AI plausibly displays important general-purpose-technology characteristics, while diffusion, productivity, distribution, and institutional effects remain uncertain.

Method and limitations. This is a structured evidence synthesis, not a causal test that AI has already produced general-purpose productivity growth. The technology and evidence base are changing rapidly, and GPT classification does not specify timing, magnitude, beneficiaries, or governance.

The analogy to electrification is helpful in three respects:

  1. Complementarity and delay. Installing motors did not immediately deliver the productivity of redesigned factories. Similarly, adding a model endpoint does not redesign decisions, evidence flows, incentives, or accountability.
  2. Pervasive input. Reasoning services may enter many activities without being a single final product, making direct sector classification difficult.
  3. Infrastructure externalities. Generation, transmission, siting, and local grid constraints have direct counterparts in the electricity actually used by digital infrastructure.

Acemoglu’s “The Simple Macroeconomics of AI”, published online in August 2024 and in a January 2025 journal issue, asks what current task-exposure and cost-saving evidence implies for aggregate productivity.

Method. It combines a task-based macroeconomic framework with estimates of AI exposure, adoption, cost savings, and the share of tasks likely to be profitably automated.

Principal finding. Under the paper’s assumptions, AI raises total factor productivity by no more than roughly 0.66% over ten years, with smaller baseline estimates; distributional effects may be adverse if capital captures gains or if low-productivity uses expand.

Limitations and relevance. This is a model-based scenario, not a forecast fact. Results depend on exposure measures, profitable task shares, cost savings, complementarity, new-task creation, and the period over which reorganization occurs. It is valuable because it prevents micro-level task gains from being arithmetically treated as macro productivity.

10.2 Where the analogy fails

ElectricityReasoning infrastructure
Metered, largely fungible commodity within specified grid products and quality standardsHeterogeneous output whose usefulness depends on task, evidence, context, and evaluation
Stable physical behavior at the point of useProbabilistic and version-dependent behavior; providers can change models and policies
Commodity normally does not inspect or transform customer meaningService processes confidential data, retrieves sources, invokes tools, and can alter records
Quality is largely independent of local organizational vocabularyPerformance depends on local categories, exceptions, memory, prompts, tools, and authority
Mature metering, safety codes, common-carrier institutions, and grid governanceImmature agent protocols, semantic profiles, assurance practice, and cross-provider accountability
Switching electricity supplier rarely redesigns the appliance’s meaningSwitching models can require prompt, evaluation, workflow, and semantic adaptation
One unit cannot be copied at near-zero marginal costModel software and files are partly non-rival, while chips, latency, energy, and attention remain rival

The “grid versus local generation” metaphor is also incomplete. A locally owned reasoning environment is not merely a generator behind the meter. Its important asset may be canonical context, provenance, memory, and authority—assets for which electricity has no analogue.

10.3 Safe use of the analogy

The analogy should be limited to a claim about diffusion and complements: a broadly applicable technical input can require long periods of reorganization and complementary investment before aggregate productivity appears, while creating concentrated infrastructure requirements. It should not be used to imply that intelligence is a homogeneous utility, that reasoning quality can be metered like energy, that provider switching is semantically neutral, or that regulation can be copied from electricity markets without substantial modification.

For the AIOS thesis, the analogy’s strongest implication is organizational reconfiguration rather than utility equivalence. Frontier and local models are inputs into a larger production system whose productive capital includes purpose, files, memory, relationships, scaffolding, evaluation, and authority. If those complements become portable and locally owned, model services may diffuse widely without owning the durable intelligence layer. That possibility resembles the complement-intensive reorganization associated with electrification; it does not make reasoning output electricity or imply the disappearance of centralized generation-like infrastructure.

11. Institutional scenarios—not forecasts

The following scenarios expose tradeoffs. They are not probability-weighted forecasts.

Scenario A: cloud-dominant execution with portable local state

Canonical files, metadata, evaluations, and approvals stay under user or institutional custody; most inference is remote and routed among providers. This exploits frontier capability and shared utilization while preserving some exit option. Its risks are behavioral drift, provider concentration, confidential-context exposure, and the operational difficulty of maintaining cross-model regression tests.

Scenario B: device and private-edge specialization

Small models handle high-frequency classification, retrieval, drafting, redaction, and file operations on devices or private servers; remote systems handle exceptions and capability-intensive review. The economic case rests on repeat volume, latency, privacy, and already-owned capacity. Its risks are model fragmentation, patching, hardware churn, uneven assurance, and silent quality differences between local and escalated paths.

Scenario C: federated professional knowledge packages

Communities, firms, or public institutions publish versioned, signed packages containing sources, methods, policies, tests, and declared scopes. Peers retain private overlays and choose local or remote execution. The institutional rationale is provenance, reuse, and contestable authority rather than sale of “intelligence.” Its risks are false confidence in signatures, governance capture, revocation failure, liability ambiguity, licensing conflicts, and fragmentation among semantic profiles.

Across all three scenarios, the durable institutional asset is the ability to preserve meaning and evidence independently of a particular execution vendor. Whether that asset is worth its maintenance cost remains workload- and institution-specific.

12. Four-layer examination of the AIOS architectural thesis

This section places the AIOS architectural thesis beside the evidence without collapsing one into the other. AIOS is treated as a complementary reasoning and knowledge architecture: it may call centralized frontier models when their capabilities are needed while moving durable context, expertise, memory, workflow state, and authority into local person- or institution-owned systems. Each subsection distinguishes four layers:

  1. what current evidence establishes;
  2. the AIOS architectural thesis;
  3. first- and second-order implications if the thesis is substantially correct;
  4. conditions, uncertainties, counterforces, and discriminating tests.

The purpose is not to make absence of evidence do the work of counterevidence. It is to preserve the scale of the proposed transition while making its causal chain and empirical burden explicit.

12.1 Bounded semantic judgment, deterministic effects, and human authority

1. Current evidence. Assurance standards and regulatory guidance consistently separate accountable roles, intended use, authorization, validation, logging, human oversight, and post-deployment monitoring. The field experiments reviewed above show that models can improve selected judgments while also inducing confident error outside their capability frontier. Exact state changes, identity checks, and access controls remain conventional deterministic-software strengths. This evidence supports explicit control boundaries; it does not compare complete “pre-AI” and “AI-native” production architectures.

2. AIOS architectural thesis. Language models should make bounded semantic judgments within a declared context and action space. Deterministic code should establish identity and scope, validate schemas and permissions, and guarantee the exact mechanical effect of an authorized operation. Humans should retain purpose, exception authority, and consequential decision rights. Model output is therefore not merely prose for downstream code to scrape; it can be a typed, reviewable judgment inside a controlled state transition.

3. Implications if substantially correct. The first-order implication is a different division of computational labor. Models contribute flexible interpretation where exhaustive rules are costly, while deterministic components make permissions and effects inspectable. This could lower the expected cost of failures, review, and integration, particularly for repeated workflows that mix ambiguous evidence with exact file changes. It could also make model substitution easier because the model is not the canonical store of identity, authority, or state.

The second-order implication is an application architecture organized around local evidence, policies, and state transitions rather than around a remote application’s database and workflow engine. If model judgments can be bounded reliably, more software functions may become compositions of local knowledge, semantic proposals, deterministic constraints, and human ratification. That could reduce dependency on centralized application logic while increasing demand for evaluation packs, policy schemas, local operations, and assurance tooling.

4. Conditions, counterforces, and tests. Deterministic code guarantees only properties correctly captured in its specification and implementation. It cannot guarantee that a semantic judgment is substantively correct, that an identity record is accurate, or that an authorized action is wise. Human approval can be overloaded or ceremonial. Separation adds schemas, interfaces, latency, and maintenance, and observed gains might be attributable to ordinary software modularity rather than to a distinct architectural regime.

A direct test should compare human-only, unrestricted model, free-text-output-plus-parser, typed bounded-judgment, and bounded-judgment-plus-human-approval designs on identical workflows. Outcomes should include accepted-task quality, unauthorized effects blocked, semantic false positives and negatives, review time, recovery, audit reconstruction, and maintenance effort. The comparison must hold model access and inference budgets constant.

12.2 Coordinated intelligence and distinct cognitive modes

1. Current evidence. Outcomes from model-based systems vary with prompts, retrieved evidence, tools, scaffolds, evaluation, and human use; model capability alone does not determine the result. Agent studies also show that added coordination can help or hurt depending on task structure and budget.

Kim et al., “Capable Language Models Can Outgrow the Benefits of Collaboration”, published in Nature Machine Intelligence in July 2026, asks when multi-agent coordination outperforms a strong single agent.

Method and sample. The authors hold prompts, tools, and per-system compute ceilings constant across 260 configurations: a single agent and four coordination structures, three model families, and six benchmarks spanning browsing, financial research, planning, workplace tasks, software maintenance, and terminal work.

Principal finding. Single-agent baseline performance was the most statistically robust predictor of whether coordination helped. The fitted system selected the best architecture in 87% of held-out within-domain configurations. Multi-agent designs incurred 1.6–6.2 times as many realized reasoning turns as the single-agent baseline despite matched compute ceilings. An approximately 45% single-agent-performance threshold helped predict the direction of multi-agent gains in an additional validation, but the underlying interaction did not survive the paper’s strongest cluster-robust correction and is presented as a selection rule rather than a universal scaling law.

Limitations and relevance. The six benchmarks do not cover long-lived file-native domains or consequential institutional workflows; architecture-specific prompts were not tuned; absolute cross-domain prediction was poor; and several tool-overhead patterns did not survive conservative inference. The study was funded by Google, many authors were Google employees, and no independent replication was identified. It is direct evidence that coordination structure matters and that more agents are not inherently better. It does not test sequential cognitive modes operating over stable local files.

2. AIOS architectural thesis. Useful intelligence is a system-level outcome, not merely a property of model weights. A user-visible act of reasoning can emerge from purpose framing, composed context, files, annotations, relational metadata, provenance, multiple memory resolutions, distinct modes such as thinking, writing, planning, editing, and review, bounded model judgments, deterministic operations, verification, post-response integration, and human authority. These functions need not be collapsed into a single persistent agent or a single undifferentiated context. Stable files can carry state between modes, while each mode receives only the evidence and permissions it needs. The central cumulative hypothesis is that many individually modest improvements can interact to reduce inference burden, retries, and hallucination exposure per accepted task.

3. Implications if substantially correct. At the personal level, intelligence infrastructure becomes cumulative: a person’s files, distinctions, annotations, corrections, and methods can improve future work without being locked inside one conversational history or provider. The model remains important but becomes one replaceable participant in a longer-lived cognitive system.

At the organizational level, different functions could share evidence and control contracts while retaining domain-specific reasoning contexts. Work could be distributed across foreground interaction, non-blocking analysis, specialist review, and exact integration without placing one autonomous agent above the whole organization. If role separation reduces context interference and makes review more independent, increasingly capable models could turn mature local scaffolding into a capability multiplier rather than merely a coordination overhead.

Second-order effects could include durable organizational memory that survives staff and vendor changes, new forms of delegation to local systems, and a shift in competitive advantage from exclusive application data silos toward accumulated domain methods and evaluation quality. They could also permit a smaller local model with excellent domain scaffolding to satisfy tasks that would otherwise require repeated frontier calls. If purpose, context, memory, provenance, planning, verification, and integration reduce different sources of failure, their combined economic value could be superadditive: fewer retries reduce tokens, review, latency, and rework together. That cumulative effect is a core implication of the thesis, not a finding established by the component studies.

4. Conditions, counterforces, and tests. The proposition that intelligence “has no single center” is an architectural and philosophical thesis, not a measurement. Multiple modes can discard shared evidence, repeat tokens, propagate inconsistent summaries, or obscure responsibility for integration. A strong single model with a well-structured context may outperform decomposition on short or tightly coupled problems. Any benefit must be distinguished from simply spending more inference and human attention.

A factorial study should vary role separation, context separation, independent review, stable file-mediated state, and total inference budget independently. It should include short and long tasks, parallelizable and sequential tasks, conflicting evidence, and inside- and outside-frontier work. Measures should include final quality, calibration, contradiction, diversity, latency, token and energy use, audit reconstruction, and human integration effort.

12.3 Bounded autonomy through menus, permissions, and escalation

1. Current evidence. Security engineering and regulated assurance support least privilege, authenticated roles, explicit consent, controlled changes, and human oversight. Current agent protocols standardize messages and capabilities but do not themselves establish that a delegation is authorized or safe. There is little direct evidence identifying an optimal action-space size for general-purpose agents.

2. AIOS architectural thesis. Bounded menus and permissions can provide an intermediate design between rigid pipelines and unrestricted autonomy. A model may choose among semantically meaningful actions, ask for clarification, or propose an exception, while deterministic controls constrain resource access and exact effects. The menu is not intended to prescribe every reasoning step; it defines the legitimate action surface and escalation path.

3. Implications if substantially correct. The first-order implication is practical autonomy that is both more adaptable than a fixed workflow and more governable than an unconstrained agent. Institutions could delegate a wider range of ordinary work while preserving visible decision rights and preventing entire classes of unauthorized action.

The second-order implication is a new governance layer for everyday software: permissions would attach not only to users and files but to semantic acts, risk classes, and transitions. Local domain owners could encode their own thresholds within organization-wide minimum controls. This could support decentralized variation without surrendering accountability to either a central platform or a free-ranging agent.

4. Conditions, counterforces, and tests. A menu may omit a safe novel action, force an ill-fitting category, or transfer hidden power to its designer. Formal authorization can still permit harmful sequences; prompt injection, confused-deputy behavior, information leakage, and collusion across tools remain. Menus and policies also require upkeep as work changes.

Identical workflows should be randomized across fixed pipelines, bounded action menus, and open-ended agents with matched tools and budgets. Evaluation should include completion, unsafe attempts, unauthorized effects actually blocked, false refusals, escalation quality, user workarounds, and the long-run cost of changing permissions.

12.4 Local domain systems and selective frontier-model use

1. Current evidence. Recent systems research demonstrates technically viable on-device inference for bounded small-model workloads, and standardized client benchmarks now test language models up to 8B parameters. Public inference prices have fallen rapidly, while frontier capability and shared cloud utilization remain economically attractive for difficult or bursty work. Competition investigations establish that remote cloud and model dependencies can carry technical, contractual, and semantic switching costs.

2. AIOS architectural thesis. Self-contained local domains can hold the canonical files, accepted ground, relationships, memory, evaluation criteria, and authority needed for a large share of ordinary reasoning. Increasingly capable local models and mature scaffolding may handle many repeated personal and organizational tasks; remote frontier models remain available selectively for difficult inference, current external knowledge, or independent review. Model selection can be dynamic by cognitive movement—the kind of step being performed, such as retrieval, classification, drafting, planning, review, or integration—as well as by capability, privacy, latency, cost, sensitivity, and assurance need. Local custody and local execution are separable choices.

3. Implications if substantially correct. The first-order effect would be a large change in the mix of execution rather than the disappearance of remote AI. High-frequency retrieval, classification, drafting, planning, annotation, redaction, and file operations could occur locally; only exceptions or capability-intensive tasks would cross the network. The accepted-task cost could fall where hardware is already owned, workloads repeat, and local context avoids repeated transmission and reconstruction.

The second-order effect would be a change in the boundary of application infrastructure. Remote services could become less responsible for holding every user’s durable state, memory, and workflow. Some centralized application functions might be replaced by local orchestration over ordinary files, while remaining remote services specialize in frontier inference, synchronization, identity, software and model distribution, backup, or cross-party coordination. This would reduce dependence on particular application providers even if total compute demand continued to grow.

A hybrid system also creates option value. When frontier-model prices or capabilities change, canonical context does not have to migrate with them. When a local model improves, more tasks can move onto already-governed state without rebuilding the entire knowledge environment. Dynamic routing can reserve expensive or less-private frontier calls for cognitive movements where their marginal capability is worth the cost and risk, while keeping routine or sensitive movements local. Repeated substitution at the execution layer could make model competition more effective than nominal API compatibility alone.

4. Conditions, counterforces, and tests. “A large share” needs a declared task denominator. Ordinary work often depends on live external facts, other people’s authority, specialized tools, and frontier capability. A local system can appear sufficient because difficult cases are silently deferred to humans or remote services. Hardware heterogeneity, patching, model updates, backup, and evaluation can outweigh saved inference charges. Cheaper private execution can also cause rebound through more background work.

A prospective workload study should inventory tasks before classifying outcomes as local-only, local with external data, hybrid with remote inference, remote-first, or human-only. It should measure acceptance quality, review, latency, privacy exposure, failure, and total cost by risk and task type. Local knowledge custody and local model execution must be varied independently.

12.5 Private local cognition and person-owned intelligence

1. Current evidence. Local execution can keep selected raw inputs off remote inference services, and local custody can give users direct control over files and retention. Existing research does not measure the full privacy outcome of a mature personal reasoning environment. Regulators and assurance reports show that privacy depends on lifecycle data flows, access control, logs, third parties, and incident response—not location alone.

2. AIOS architectural thesis. A person should be able to accumulate durable context, memory, annotations, relationships, drafts, and methods in an inspectable local system that serves the person rather than a remote application’s account model. Private reasoning can remain private by default, with deliberate disclosure to models, collaborators, or institutions. Human purpose and authority remain upstream of the system.

3. Implications if substantially correct. The first-order effect is cognitive continuity and a stronger personal exit right. A person could change models or applications without abandoning the accumulated structure through which they think and work. Sensitive exploration, incomplete drafts, and private reflection could occur without routine transmission to a central service, expanding the practical space for unobserved cognition.

Second-order effects could be institutionally significant. Person-owned cognitive capital may alter bargaining relationships with employers and platforms; make expertise more portable across projects; support long-horizon learning and memory; and create demand for new norms around inheritance, agency, deletion, search, and selective disclosure. Personal systems could collaborate by exchanging bounded packages rather than exposing entire histories. Private local cognition could become a practical civil-liberties affordance, especially where central services are surveilled, censored, unreliable, or unavailable.

4. Conditions, counterforces, and tests. Endpoint compromise, coercive device access, uncontrolled copies, weak key management, and household or workplace surveillance can defeat local privacy. Portability can conflict with confidentiality, employer ownership, professional duties, or another person’s rights. Accumulated memory may entrench errors or become burdensome to curate. Effective ownership requires comprehensible export, deletion, backup, recovery, permissions, and the ability to operate without the original vendor.

Evaluation should trace actual data flows and retention under realistic tasks; test compromise, loss, shared-device use, deletion, selective disclosure, and provider exit; and compare user comprehension and control with well-governed remote alternatives. “Private” should describe measured access and flow properties, not merely physical location.

12.6 Organizational intelligence and regulated use

1. Current evidence. Productivity studies show that AI gains depend on task fit, worker experience, adoption, evaluation, and organizational complements. Coordinated routines change more slowly than individual work. The EU AI Act, NIST, FDA, BIS, and management-system standards emphasize lifecycle governance, evidence, accountability, change control, and monitoring. None validates a specific local-first implementation.

2. AIOS architectural thesis. An organization can standardize identity, authorization, provenance, evaluation, incident handling, and evidence formats while allowing departments and professions to control local vocabulary, sources, annotations, rubrics, exceptions, and authority. Domain systems remain self-contained enough to preserve meaning and confidentiality, but interoperable enough to exchange bounded evidence and requests.

3. Implications if substantially correct. The first-order organizational effect would be distributed intelligence with common controls. Domain experts could build and retain operational knowledge without waiting for one central application or AI team to encode every exception. Institutional memory could reside in inspectable files, relationships, and evaluation packs rather than in scattered conversations or one provider’s hidden state.

For regulated use, local evidence continuity could support reconstruction of what sources, model, policy, permissions, tool calls, file changes, evaluations, and approvals produced a consequential outcome. Local domains could be validated for defined intended uses and updated through controlled change packages. Frontier models could still be used when allowed, but their judgments would enter a retained evidence chain rather than become the whole system of record.

Second-order effects could include new professional roles for domain-package stewardship, model and workflow assurance, semantic migration, and independent review. Organizational boundaries may become more permeable where bounded domain packages can be shared without surrendering all internal context. Regulators may be able to inspect evidence contracts and declared transitions rather than rely only on provider-level model documentation.

4. Conditions, counterforces, and tests. Local autonomy can become inconsistent risk acceptance or inaccessible records; central standardization can flatten domain meaning. Detailed logs create privacy, security, litigation, and surveillance risks. Human authority must be real, trained, and resourced. A local workflow still needs substantive performance validation for its intended use, separation of duties, incident response, and ongoing monitoring.

Multi-institution pilots should compare a common evidence/control plane with centralized semantic standardization and with uncoordinated local practice. Measures should include accepted outcomes, policy exceptions, audit reconstruction, domain-user adaptation, central-team burden, migration effort, incident detection, and time to correct shared and local errors.

12.7 Peer collaboration, knowledge packages, and global accessibility

1. Current evidence. MCP, A2A, and related protocols can reduce interface work; RO-Crate, Verifiable Credentials, C2PA, and Sigstore supply components for portable manifests, issuer claims, provenance, and signatures. They do not establish semantic agreement, truth, fitness, or certification. Client inference benchmarks and on-device results demonstrate some technical basis for operation outside continuous cloud access.

2. AIOS architectural thesis. Local domains could exchange content-addressed, versioned packages containing files, metadata, sources, relationships, methods, tests, policies, signatures, and declared scope. Recipients could preserve private overlays and use their own local or remote models. Peer exchange would distribute bounded knowledge and evidence without centralizing every use event.

3. Implications if substantially correct. The first-order effect would be reusable domain infrastructure. Professional communities, educational networks, public agencies, firms, and individuals could publish inspectable packages that others can run, test, annotate, and fork. Collaboration would center on evidence and declared methods rather than on access to one hosted application.

Second-order effects could include federated knowledge institutions, contestable certification, and faster diffusion of expert practice. A rural clinic, small firm, school, or independent professional could combine an inexpensive local model with a high-quality domain package and selectively obtain remote expertise. Low-bandwidth or intermittently connected communities could continue to reason over local materials. Multiple packages could preserve competing interpretations rather than forcing one global ontology.

This architecture could broaden global accessibility by lowering recurring bandwidth and service dependence, making translation and adaptation local, and allowing communities to retain authority over sensitive knowledge. It could also shift value from exclusive access to hosted data toward the credibility, maintenance, and interoperability of shared packages.

4. Conditions, counterforces, and tests. Hardware cost, electricity, language coverage, technical literacy, accessibility needs, and maintenance can reproduce global inequality. A signed package can be obsolete, biased, malicious, or outside its intended scope. Certification requires accountable issuers, explicit claims, test methods, surveillance, revocation, appeals, and liability. Forks can support pluralism or produce fragmentation and incompatible semantic profiles.

A peer pilot should test package discovery, offline verification, local overlays, correction and revocation, cross-language adaptation, license enforcement, and migration across models. Institutional evaluation should measure who can publish, challenge, and maintain packages; whether small or resource-constrained users achieve accepted outcomes; and whether the system reduces or merely relocates dependence.

12.8 Scaffolding, context, and the usable frontier

1. Current evidence. Scaffold and context design materially change model behavior, but more context and more agent steps do not reliably produce better results. Agent studies show large variation in cost and success across architectures.

Du et al., “Context Length Alone Hurts LLM Performance Despite Perfect Retrieval”, published in Findings of EMNLP 2025, asks whether long-input performance loss is only a retrieval failure. The authors test five open- and closed-source models on mathematics, question answering, and coding while making all required evidence perfectly retrievable and manipulating input length.

Principal finding. Performance declined by 13.9% to 85% as input length increased within advertised context windows, including conditions in which filler was whitespace or masked and evidence appeared immediately before the question. A short-context recitation intervention improved GPT-4o by up to four percentage points on RULER.

Limitations and relevance. These are controlled benchmark manipulations, not end-to-end institutional tasks; the wide range shows strong model/task dependence; a maximum improvement is not an average treatment effect; and no independent replication of the exact estimates was identified. The study does not imply that context should always be short. It establishes that retrieval of all relevant information and nominal context capacity are not sufficient for reliable reasoning.

2. AIOS architectural thesis. Prevailing application scaffolds may underuse frontier-model reasoning when they flatten durable knowledge into one prompt, treat model output as unstructured text, mix incompatible cognitive tasks, or obscure evidence and authority. File-native state, multiple-resolution memory, role-specific context, bounded judgments, and post-process integration may let both frontier and local models operate closer to their useful capability.

3. Implications if substantially correct. Better scaffolding could increase the amount of useful work obtainable from a given model and token budget. The gain may be especially important locally: a smaller model operating over well-curated domain state could satisfy tasks that a generic prompt cannot. As models improve, a stable local architecture could capture those gains without repeatedly rebuilding user knowledge inside new applications. A router that recognizes different cognitive movements could also avoid treating every step as a demand for the same largest model, lowering cost and exposure without imposing one local-only ceiling.

Second-order effects could compound. If scaffolding increases accepted-task capability while local model hardware improves, remote inference may become an exception path for more workflows. If exact integration and evaluation reduce rework, the relevant productivity gain would exceed the reduction in model-call price alone. Mature scaffolding could thus be an economic complement to model progress, not merely overhead around it. At sufficient scale, cumulative reductions in prompts, retries, verification failures, and unnecessary frontier routing could change both application-layer infrastructure demand and the distribution of remote inference, even though frontier training and centralized high-capability services remain essential.

4. Conditions, counterforces, and tests. Current evidence does not identify scaffolding as the sole or dominant cause of agent failure. Failures also arise from base-model limits, excess or irrelevant context, retrieval errors, unreliable tools, ambiguous objectives, bad decompositions, stale knowledge, evaluation mismatch, unsafe delegation, and compounding errors. More elaborate scaffolding can suppress performance, multiply calls, and make debugging harder.

Preregistered ablations should vary context sufficiency, irrelevant-context load, memory resolution, role separation, scaffold, tool reliability, feedback, and token budget independently across frontier and smaller models. Claims that a scaffold unlocks a model’s latent capability require resource-matched improvements across task classes, transparent failure analysis, and independent replication. The full routing benchmark belongs to a dedicated investigation; for this memo, the required economic output is narrower: accepted-task cost and quality by route, including retries, review, latency, energy, privacy exposure, and switching effects.

12.9 Integrated first- and second-order implication map

The table preserves the scale of the thesis while keeping every consequence conditional.

DomainFirst-order implication if the thesis holdsSecond-order implicationKey conditions and counterforces
Personal intelligenceDurable memory, context, methods, and model-assisted work remain person-controlledPortable cognitive capital; stronger exit rights; private space for thoughtEndpoint security, usable ownership, curation, rights in shared data
Organizational intelligenceDomain systems combine local semantics with common evidence and controlsDistributed institutional memory; less dependence on central application teamsGovernance, evaluation, cross-domain coordination, correction processes
Application infrastructureMore state, workflow, and knowledge orchestration occurs locallyCentral services become thinner or more specialized; switching becomes more credibleSynchronization, identity, backup, updates, collaboration still require infrastructure
Remote inferenceRoutine bounded work shifts toward devices or private edge; frontier calls become selectiveLower dependency on remote inference for ordinary work, even as frontier access remains valuableLocal capability, utilization, accepted-task cost, rebound, current external data
System efficiency and routingContext, memory, verification, and model choice reduce avoidable calls and reworkCumulative complementarity may lower infrastructure required per accepted taskComponent gains may not add; routing errors and orchestration can increase cost and failure
Data centers and trainingLittle direct first-order change to frontier training requirementsDemand composition may change; total demand can still rise with expanded useChip and energy concentration, training scale, remote exception work, rebound
Regulated activityLocal evidence chains join model judgments to sources, controls, exact effects, and approvalsWorkflow-level assurance and more portable supervisory evidenceIntended-use validation, tamper evidence, minimization, monitoring, accountable review
Peer collaborationSigned packages and private overlays enable bounded exchangeFederated professional and public knowledge institutionsCertification scope, revocation, liability, semantic profiles, governance capture
Global accessibilityLocal packages and models reduce recurring connectivity dependenceWider access to domain reasoning and locally governed adaptationHardware, electricity, languages, disability access, maintenance, institutional capacity
Energy and resource useSome remote workloads move to heterogeneous local hardwareCentral energy may fall for displaced tasks or rise overall through rebound and new useLifecycle boundaries, device turnover, grid mix, utilization, inference intensity

The most consequential implication is therefore not that centralized models or data centers disappear. It is that the durable intelligence layer—the purpose, context, expertise, memory, relationships, workflow, evidence, verification, and authority that make model calls useful—could migrate toward people and institutions. Dynamic local/frontier models would then serve a more portable system of intelligence rather than own the entire application relationship.

12.10 Minimum empirical program

A compact program capable of testing the architecture and its implications would include:

  1. Architecture trial: human-only, direct-model, conventional model-output-plus-parser, bounded-judgment/deterministic-effect, and role-separated designs on the same real workflows and resource budgets.
  2. Locality trial: local-only, remote-only, and hybrid execution while holding the canonical file representation and accepted ground constant, separating custody effects from model-capability effects.
  3. Representative workload census: prospectively sample personal and organizational tasks, declare acceptance criteria and risk classes, and measure the share handled locally, selectively remotely, or by humans without assuming the answer.
  4. Migration challenge: change the model and execution provider mid-study; measure export completeness, regression failures, retuning time, stranded semantics, and interruption.
  5. Assurance challenge: inject ambiguous authority, prompt attacks, stale sources, package revocation, and log-tampering attempts; test prevention, detection, reconstruction, and human response.
  6. Infrastructure accounting: measure accepted-task cost, endpoint and server energy, network and storage, maintenance labor, review, rework, recovery, and centralized services over months rather than extrapolating from token prices.
  7. Privacy and sovereignty audit: trace data flows, retention, access, deletion, keys, backups, provider dependencies, and operation under disconnection or service withdrawal.
  8. Peer-package pilot: test publication, offline verification, private overlays, cross-language adaptation, correction, revocation, licensing, and accountable certification across more than one institution.
  9. Independent replication: publish tasks, schemas, evaluation packs, failure taxonomies, system boundaries, and analysis code where confidentiality permits; repeat across domains, organizations, model families, and hardware classes.

12.11 Interpretive boundary

Current evidence establishes mechanisms and boundary conditions: local inference is feasible for bounded workloads; scaffolding changes performance and cost; productivity depends on complements; portability is not the same as semantic continuity; auditability requires a whole evidence chain; and centralized infrastructure is concentrated. It does not yet measure AIOS as a complete system.

The appropriate conclusion is not that the larger implications are excluded. It is that they form a coherent conditional research program. If local model capability, domain scaffolding, portable state, assurance, and user governance mature together, substantial movement of ordinary reasoning and durable intelligence toward local systems is economically and institutionally plausible. The scale, distribution, and net benefits depend on the empirical conditions set out above.

13. Foundational lineage: indispensable older work

The following sources predate the main research window and are included only because they frame the complementarity and assurance questions.

14. Conclusions

Strongest supported conclusions

  1. Fixed-capability retail inference became dramatically cheaper through 2024. The direction and rough order are supported by reproducible benchmark-price reconstructions. The series does not measure complete task cost or provider marginal cost.
  2. Agent design can dominate token, call, and energy use. Recent benchmark studies show one- to three-order-of-magnitude variation across workflows and runs. Evidence is strongest for selected coding and research-style tasks, not economy-wide use.
  3. On-device language-model inference is technically viable for bounded models and tasks. Peer-reviewed smartphone results and standardized client benchmarks support feasibility. Total-cost, lifecycle, and cross-workload economic superiority remain unmeasured.
  4. AI infrastructure is capital-, energy-, and geographically concentrated. IEA/LBNL estimates and competition-authority investigations consistently show rapid demand growth, clustered grid effects, hyperscaler scale economies, and contractual/technical dependence.
  5. Cloud portability is not semantic portability. Regulatory evidence establishes material technical and commercial switching barriers. Experimental evidence shows that model upgrades can require user adaptation. Full semantic switching cost is conceptually important but not yet directly measured.
  6. Productivity gains are conditional on task fit and complements. Randomized field studies find meaningful gains on selected tasks and losses or weak effects elsewhere. Administrative labor-market evidence does not yet show corresponding broad earnings or hours effects.
  7. Auditability is a system property. Local custody can improve evidence continuity, but only lifecycle logging, identity, authorization, evaluation, tamper evidence, monitoring, and accountable human review create an assurance basis.
  8. Open protocols reduce plumbing, not trust or meaning. MCP, A2A, and related efforts can lower interface costs. Semantic alignment, safe delegation, identity, conformance, and cross-organizational accountability remain open.
  9. Attestable domain packages are technically plausible; certified truth is not. Existing provenance, credential, research-object, and signing standards can bind files and claims to issuers. They do not validate knowledge without an explicit assurance institution.
  10. The electricity analogy is useful for complementarity and infrastructure, not for the nature of reasoning. AI outputs lack the homogeneity, stability, metering, and mature institutional framework of electricity.
  11. Bounded model judgments, deterministic effect controls, and explicit human authority are directionally consistent with current assurance practice. This supports evaluating the division of labor, but no comparative field evidence reviewed here establishes it as a generally superior architecture.
  12. More agents or roles are not monotonically better. A recent vendor-involved, peer-reviewed controlled study found that the value of coordination depended on task and base-model capability and imposed substantial reasoning-turn overhead. It does not directly test file-centered, role-specific workflows.
  13. Useful intelligence is economically a system-level outcome. Model capability, context, tools, scaffold, evaluation, and human use jointly affect accepted-task cost and quality. Current studies establish several component effects, but not the cumulative effect of the complete AIOS bundle.

Unresolved or contradictory evidence

Possible AIOS connection points for later consideration

These are design implications to evaluate, not claims that the reviewed work validates AIOS.

Claims that would be unsafe to make

The formulations below collapse conditional implications into universal or already-demonstrated results. They do not exclude the more carefully stated AIOS theses examined in Section 12.

15. Source table

Dates identify publication, release, or material revision within the research window. Status labels distinguish peer review and authority from emerging evidence.

DateSource and statusResearch roleDirect link
Aug. 2024EU Artificial Intelligence Act — law/official textHigh-risk lifecycle controls, logging, documentation, oversightRegulation (EU) 2024/1689
Aug. 2024; Jan. 2025 issueAcemoglu, The Simple Macroeconomics of AIpeer-reviewed model-based analysisTask-to-macro productivity scenarios and distributional implicationsDOI
Dec. 2024Lawrence Berkeley National Laboratory, 2024 US Data Center Energy Usage Reportauthoritative modeled estimate/scenariosUS historical estimate and 2028 electricity scenariosReport landing page
Jan. 2025US Federal Trade Commission, AI partnerships 6(b) study — regulator staff reportCloud/model contracts, exclusivity, spending ties, possible switching costsReport PDF
Jan. 2025US FDA, AI-enabled device software guidance — draft, nonbinding guidanceTotal-product-lifecycle assurance, validation, monitoringDraft guidance PDF
Jan. 2025Bank for International Settlements, Governance of AI Adoption in Central Banksauthoritative sector-practice reportAccountability, inventories, three lines of defenseBIS report
Jan. 2025 updateSigstore Bundle Format 0.3 — open software-signing specificationPortable signature, identity, timestamp, and transparency evidenceSpecification
Mar. 2025Xu et al., Fast On-device LLM Inference with NPUspeer reviewed, ASPLOS 2025Smartphone NPU performance and energy measurementsDOI
Mar. 2025Cottier et al., LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasksreproducible research-organization data analysisBenchmark-threshold retail price trends and assumptionsEpoch AI analysis
Apr. 2025Stanford HAI, AI Index Report 2025, ch. 1 — authoritative synthesisFixed-capability API price decline and hardware trendsChapter PDF
Apr. 2025Humlum and Vestergaard, Large Language Models, Small Labor Market Effectsworking paperAdoption linked to Danish administrative labor outcomesBFI page and paper
Apr. 2025Erol et al., Cost-of-PasspreprintExpected price per correct benchmark resultarXiv
Apr. 2025International Energy Agency, Energy and AIauthoritative modeled estimate/scenarios2024 global demand, infrastructure and 2030 scenariosIEA report
May 2025Dillon et al., Shifting Work Patterns with Generative AIvendor-affiliated working paperSix-month randomized workplace experimentNBER working paper
May 2025OECD, Competition in the Provision of Cloud Computing Servicesauthoritative policy synthesisConcentration, technical and commercial switching barriersOECD report PDF
May 2025W3C Verifiable Credentials Data Model 2.0 — W3C RecommendationPortable issuer claims and verification modelRecommendation
May 2025C2PA 2.2 — industry provenance specificationSigned manifests, content history, validationSpecification
June 2025Kim et al., The Cost of Dynamic ReasoningpreprintAgent calls, measured GPU energy, diminishing returnsarXiv
June 2025Agent2Agent project transferred to Linux Foundation — open protocol/projectCapability discovery and agent task exchangeLinux Foundation announcement
June 2025OECD, Is Generative AI a General Purpose Technology?authoritative evidence synthesisCareful assessment of GPT criteria and uncertaintyOECD report
July 2025UK Competition and Markets Authority, cloud market investigation — regulator final reportConcentration, profitability, switching and licensing barriersCase and final report
Aug. 2025 applicationEU AI Act general-purpose AI obligations — official legal guidanceProvider documentation, copyright, systemic-risk, and incident dutiesEuropean Commission fact page
Sept. 2025 applicationEU Data Act — law/official explanationCloud and edge switching, interfaces, export, charge removalEuropean Commission explainer
Sept. 2025; journal 2026Oviedo et al., Energy Use of AI Inferencevendor-authored bottom-up perspective; peer-reviewed journal versionAt-scale inference energy model and test-time-compute scenariosMicrosoft Research page
Nov. 2025Gundlach et al., The Price of ProgresspreprintPerformance-adjusted inference prices and efficiency decompositionarXiv
Nov. 2025Du et al., Context Length Alone Hurts LLM Performance Despite Perfect Retrievalpeer reviewed, Findings of EMNLP; single studyLong-input degradation with controlled retrieval and fillerACL Anthology
Dec. 2025Ecma NLIP standards suite — adopted Ecma standardsAdditional agent-communication layerEcma announcement
Feb. 2026Cui et al., The Effects of Generative AI on High-Skilled Workpeer reviewed, randomized field experimentsDeveloper task completion across three firmsDOI
Feb. 2026NIST AI Agent Standards Initiative — official initiative, early-stageSecure interoperability, identity, evaluation and standards coordinationNIST announcement
Mar. 2026Dell’Acqua et al., Navigating the Jagged Technological Frontierpeer reviewed, preregistered experimentInside/outside-frontier productivity and errorDOI
Mar. 2026NIST AI 800-4, Challenges to Monitoring Deployed AI Systemsauthoritative workshop/literature studyMonitoring gaps across technical and institutional layersNIST report page
Apr. 2026Bai et al., How Do AI Agents Spend Your Money?preprintToken use and cost variability in coding-agent trajectoriesarXiv
Apr. 2026International Energy Agency, Key Questions on Energy and AIauthoritative retrospective estimates and forward scenarios2025 growth, capex, density, and 2030 central scenarioIEA report
Apr. 2026Jahani et al., Prompt Adaptation as a Dynamic Complement in Generative AI Systemspeer reviewed, preregistered experimentsUser adaptation across model versions and task typesDOI
May 2026NIST, Building Evaluation Probes into Agentic AIongoing official researchEmbedded evaluation and machine-readable agent audit trailsNIST project
June 2026Dell’Acqua et al., The Cybernetic Teammatepeer reviewed; firm-involved field experimentIndividuals/teams, functional silos, judged innovation outputDOI
June 2026LBNL, United States Data Center Energy Usage Report, 2025authoritative bottom-up scenariosPlanned hardware and cooling model for 2030 US demandReport landing page
June 2026RO-Crate 1.3 — community recommendationPortable files, entities, provenance, relationshipsSpecification
June 2026MLCommons, MLPerf Mobile v6.0 — industry benchmark releaseStandardized Android LLM performance and accuracyRelease
July 2026Kim et al., Capable Language Models Can Outgrow the Benefits of Collaborationpeer reviewed; vendor-funded/involved controlled benchmark studyMatched-budget single-/multi-agent architecture comparisonNature Machine Intelligence
July 2026OECD, Artificial Intelligence Marketsauthoritative policy synthesisCross-layer concentration and vertical dependenciesFull report
July 2026Model Context Protocol, 2026-07-28 revision — open protocol specificationStateless model/application access to tools, resources, prompts, and extensionsSpecification
Current through Aug. 2026MLCommons, MLPerf Client — industry benchmark specificationClient LLM task, performance, and accuracy methodologyBenchmark documentation

Foundational sources outside the main research window

DateSource and statusResearch roleDirect link
1990David, The Dynamo and the Computerpeer-reviewed historical perspectiveElectrification delay and complementary reorganizationBibliographic record
1999David and Wright, General Purpose Technologies and Surges in Productivityhistorical economic researchFactory electrification and production-system redesignOxford repository
2000Brynjolfsson and Hitt, Beyond Computationpeer-reviewed evidence synthesisIT, organizational practice, and intangible complementsDOI
2021Brynjolfsson, Rock, and Syverson, The Productivity J-Curvepeer-reviewed model and empirical applicationIntangible investment and delayed measured productivityDOI
2023ISO/IEC 42001 — international management-system standardCertifiable organizational AI management processISO record
July 2024NIST AI 600-1, Generative AI Profilevoluntary official risk-management profileGenerative-AI risk, evaluation, provenance, incident practicesNIST PDF