AIOS Proresearch
AIOS Intelligence System · Research Overview

Part VI · Evaluation, Economics, and Research Positioning

Capability Horizons, Experiments, Proof, and Falsification

The architecture remains a proposal until tested against simpler systems over long-running tasks. This chapter sets proof levels, fair comparisons, tests with one component removed, full costs, outcome measures, and results that would disprove its claims.

General intelligence harnessSystem-level capabilityAccepted outcomeCumulative capability amplification

1. Research posture

AIOS should be evaluated neither as a single foundation model nor as a collection of isolated interface features. It is a proposed general intelligence harness whose claims concern the interaction of:

This chapter is the publication’s evaluation companion. Its waves organize progressively stronger proofs of the architecture; they are not a product status report or an implementation roadmap.

The correct unit of proof is therefore a longitudinal human–system journey, supported by component evidence and compared with simpler baselines. The ten reference journeys provide the integrated product traces from which these experiments are constructed.

2. An evidence ladder

Different statements about the system require different kinds of evidence. Keeping them distinct allows the vision to remain ambitious without treating architecture, implementation, and measured effect as the same thing.

LevelMeaningWhat supports it
Architectural propositionWhat AIOS is designed to do and why its parts relateComplete mechanism, explicit boundaries, and internal consistency
Implemented behaviorWhat a particular build actually doesInspectable source, runtime behavior, and reproducible traces
End-to-end demonstrationWhether one complete journey works under defined conditionsFrozen task, acceptance criteria, exact context, result, reload, and review evidence
Replicated effectWhether an outcome survives repeated users, tasks, or modelsPreregistered comparisons, repeated runs, and independent evaluation
Bounded generalizationWhether an effect transfers beyond the original conditionsDiverse tasks, domains, users, model classes, and declared limits
Research hypothesisA causal or economic effect the architecture makes plausibleComparative experiments capable of disconfirming it
Long-range implicationA future possibility that follows if several hypotheses holdScenario analysis, enabling conditions, and visible uncertainty

Architecture is not empirical performance. A capability that is well specified but not built remains a target. A feature that exists but lacks semantic evaluation remains implemented, not demonstrated.

3. Capability waves

Capability waves and their proof thresholds

How evidence advances from governing boundaries and one complete loop toward long-running tests of compounding intelligence and bounded autonomy.

flowchart LR W0["Wave 0: Constitution and evidence boundary"] --> W1["Wave 1: Minimum complete composition"] W1 --> W2["Wave 2: Standard composition workspace"] W2 --> W3["Wave 3: Core reasoning and composition"] W3 --> W4["Wave 4: Workflows and recommendations"] W4 --> W5["Wave 5: Projects, plans, and role teams"] W5 --> W6["Wave 6: Compounding local intelligence"] W6 --> W7["Wave 7: Extended autonomy experiments"]
System-level capability · canonical24-capability-horizons-experiments-proof-and-falsification--m01.mmd
Text equivalent

Wave 0: Constitution and evidence boundary → Wave 1: Minimum complete composition; Wave 1: Minimum complete composition → Wave 2: Standard composition workspace; Wave 2: Standard composition workspace → Wave 3: Core reasoning and composition; Wave 3: Core reasoning and composition → Wave 4: Workflows and recommendations; Wave 4: Workflows and recommendations → Wave 5: Projects, plans, and role teams; Wave 5: Projects, plans, and role teams → Wave 6: Compounding local intelligence; Wave 6: Compounding local intelligence → Wave 7: Extended autonomy experiments.

Each wave earns the next by satisfying an exit condition, not by adding features. The sequence moves from a rebuildable evidence boundary through coherent composition, workflows, projects, and governed memory before extended-autonomy experiments can claim system-level effects.

Wave 0 — constitution, evidence boundary, and rebuild vehicle

Exit: the target architecture, source authority, prohibited constructs, rebuild basis, and first feature contract are settled. No implementation sprint should outrun this boundary.

Wave 1 — minimum complete composition loop

Exit: one provider-backed, file-backed, interface-visible movement passes independent semantic, product, and implementation review from clean state through reload.

Wave 2 — standard composition workspace

Exit: common context, command, landing, failure, metadata-projection, and reconstruction contracts exist; navigation performs no semantic effects.

Wave 3 — core reasoning and composition capabilities

Exit: thinking, writing, editing, structure, planning, review, and integration operate on one semantic situation without duplicating truth or context architecture.

Wave 4 — workflows, stages, recommendations, and completion

Exit: a real workflow survives interruption and returns a completed result into the correct artifact, plan, and interface ground.

Wave 5 — projects, hierarchical planning, and role teams

Exit: a multi-session project integrates bounded roles, compares conflicting returns, changes direction responsibly, and resumes without transcript reconstruction.

Wave 6 — compounding local intelligence and memory governance

Exit: reusable intelligence is accepted, used, challenged, corrected, narrowed, retired, or superseded with complete lineage.

Wave 7 — extended autonomy and contextual-intelligence experiments

Exit: comparative studies show gains in correction, continuity, source fidelity, human authority, or productive reasoning without unacceptable false confidence, disclosure, cost, or control burden.

4. Core hypotheses

H1 — purpose-shaped composition

For complex longitudinal work, a purpose-shaped context composed from durable semantic ground will outperform raw transcript continuation and undifferentiated retrieval on relevance, source fidelity, correction, and person-rated usefulness.

Could be falsified by: no material advantage over simpler retrieval after controlling for token count and model; higher omission rates; excessive authoring burden.

H2 — recursive Why–How–What coherence

Preserving Why–How–What across projects, plans, stages, artifacts, and local edits will reduce objective drift and improve reintegration after interruption.

Could be falsified by: no measurable drift reduction; increased rigidity; users unable to maintain the structure; equivalent outcomes from ordinary outlines.

H3 — file-native reconstruction

Durable artifacts plus companion metadata can reconstruct the operative semantic situation without replaying full conversational history.

Could be falsified by: frequent dependence on hidden session state; low reconstruction agreement; unacceptable manual repair; artifacts too polluted for ordinary use.

H4 — governed promotion

Evidence-linked promotion and revalidation will improve repeated task performance without converting historical practices into rigid universal rules.

Could be falsified by: harmful practice ossification; high false-promotion rates; little later reuse; governance cost greater than repeated reinvention.

H5 — attention-following asynchronous work

Bounded background branches with inspectable returns will reduce human waiting and interruption cost without reducing oversight or integration quality.

Could be falsified by: return-review burden exceeds waiting saved; lost situational awareness; duplicated work; excessive notification and attention fragmentation.

H6 — correction across interruption

Evidence, rationale, supersession, and consumer lineage will enable materially better long-horizon correction than version history or transcript memory alone.

Could be falsified by: missed consumers, false propagation, no change in later reasoning, or greater correction cost than manual expert reconstruction.

H7 — provider portability

Keeping durable knowledge, context sources, workflows, and authority outside a provider can reduce switching cost and preserve continuity across models.

Could be falsified by: behavior depends so heavily on provider-specific features that migration is impractical; adapter and revalidation cost eliminates the benefit.

H8 — task-specific capability amplification

High-quality local context and reusable expertise may allow less expensive or smaller models to approach stronger models on bounded, knowledge-intensive tasks.

Could be falsified by: no task-quality convergence after controlling for latency and cost; smaller models fail on judgment despite excellent context; maintenance costs dominate savings.

This is not a claim of general frontier equivalence.

H9 — interaction-dependent system capability

The complete person–model–artifact system will produce outcome differences that cannot be explained by adding the average effects of its components independently. Some context, memory, movement, verification, and reintegration layers will reinforce one another while others will be redundant or antagonistic.

Could be falsified or narrowed by: performance explained almost entirely by model choice or added tokens; no repeatable interaction effects; simpler harnesses matching the full system after resource control.

H10 — bounded semantic sovereignty

Allowing models to make genuine semantic judgments inside explicit jurisdictions, while deterministic systems retain custody of exact effects, will outperform both keyword-driven decision extraction and broadly autonomous agents on accepted-outcome quality, recoverability, and unauthorized-effect rate.

Could be falsified or narrowed by: semantic decisions proving too variable for useful mediation; deterministic baselines producing equivalent quality; authority boundaries creating prohibitive friction; or hidden action paths bypassing the jurisdiction.

H11 — meaningful human agency

Interfaces that preserve task-relevant cognitive acts while moving mechanical latency into asynchronous work will improve reliance calibration and retained human capability relative to either unaided work or end-to-end automation.

Could be falsified or narrowed by: unchanged error detection; greater rubber-stamping; weaker later unaided performance; review debt exceeding attention saved; or benefits limited to user preference rather than capability.

H12 — policy-gated dynamic routing

Selecting a model–harness configuration for each bounded movement after applying privacy and authority constraints will reduce total cost or latency while maintaining a declared quality threshold and escalating difficult cases appropriately.

Could be falsified or narrowed by: poor routing transfer, privacy leakage, frequent under-escalation, provider-specific instability, or routing overhead exceeding the savings.

H13 — staged reasoning and situation renewal

For multifaceted tasks, movement-specific reasoning with purpose-shaped contexts, landed intermediate products, and recomposition may outperform one extended reasoning trace under a matched total resource budget. The proposed advantage is not that longer reasoning is inherently weak. It is that different cognitive movements can receive different grounds, products can be evaluated before they become later premises, and each accepted delta can renew the situation for the next judgment.

The comparison should vary task horizon, decomposition quality, model class, and whether the task has a strong deterministic verifier. It should measure correct-answer preservation, unsupported claims, error propagation, handoff loss, context and inference cost, latency, human preparation and review, and accepted-outcome quality.

Could be falsified or narrowed by: one extended trace or repeated self-refinement matching or exceeding the staged condition after total resources are controlled; staging overhead dominating any quality gain; poor intermediate judgments compounding error; benefits appearing only in one model or task family; or an expert-designed simple workflow matching the full composition.

H14 — intent engineering and purpose traceability

When a person begins with a tension, partial aim, or internally conflicted direction, an explicit intent-engineering movement may improve later planning, reasoning, and artifact quality by clarifying possible purposes, preserving which purpose received intentional standing, and making foreground work traceable back to it. Human commitment is the experimental boundary: model clarification should improve the decision surface without substituting a model-selected objective for the person's purpose.

Evaluation should compare fixed-request execution, open model-led goal inference, person-only clarification, and joint clarification with explicit commitment and purpose traceability. Measure later objective drift, plan and artifact alignment, responsible purpose revision, correction burden, person understanding, model anchoring, and whether consequential work can be explained upward through commitments, architecture, strategy, and outcome.

Could be falsified or narrowed by: no advantage when initial purpose is incomplete; added ceremony or latency outweighing alignment gains; model suggestions anchoring or displacing the person's own purpose; traceability rewarding superficial conformity; or prior intent becoming harder to reopen despite contrary results.

5. Baselines

Every evaluation should select relevant simpler alternatives:

  1. model with raw conversational transcript;
  2. model with a hand-written summary;
  3. flat document retrieval by lexical search;
  4. embedding-based retrieval over the same corpus;
  5. ordinary project files without companion semantic metadata;
  6. conventional workflow or task tracker plus model chat;
  7. expert human performing manual reconstruction;
  8. provider-native memory or project context;
  9. the same model under a minimal and an AIOS-like harness;
  10. multiple models under the same minimal and strong harnesses;
  11. AIOS with one subsystem ablated.
  12. one extended reasoning trace under a matched inference budget;
  13. repeated self-refinement over one accumulating context;
  14. an expert-designed simple workflow with the same model and review budget;
  15. task-only, full-corpus, and composed-commission context conditions.

The aim is not to defeat every baseline on every measure. It is to identify the conditions under which added architecture earns its complexity.

6. Ablation program

Testing the architecture one boundary at a time

Removing each major layer from the full architecture reveals which relationships improve journey outcomes, add no value, or interact.

flowchart TD F["Full reference architecture"] --> A1["Remove purpose-shaped context"] F --> A2["Remove companion metadata"] F --> A3["Remove explicit authority gates"] F --> A4["Remove source descent"] F --> A5["Remove dissent and counterevidence"] F --> A6["Remove promotion/revalidation"] F --> A7["Remove reload reconstruction"] F --> A8["Replace bounded roles with generic agent"] F --> A9["Collapse distinct cognitive movements"] F --> A10["Remove post-response reintegration"] F --> A11["Disable policy-gated model routing"] A1 --> C["Compare journey outcomes"] A2 --> C A3 --> C A4 --> C A5 --> C A6 --> C A7 --> C A8 --> C A9 --> C A10 --> C A11 --> C
Cumulative capability amplification24-capability-horizons-experiments-proof-and-falsification--m02.mmd
Text equivalent

Full reference architecture → Remove purpose-shaped context; Full reference architecture → Remove companion metadata; Full reference architecture → Remove explicit authority gates; Full reference architecture → Remove source descent; Full reference architecture → Remove dissent and counterevidence; Full reference architecture → Remove promotion/revalidation; Full reference architecture → Remove reload reconstruction; Full reference architecture → Replace bounded roles with generic agent; Full reference architecture → Collapse distinct cognitive movements; Full reference architecture → Remove post-response reintegration; Full reference architecture → Disable policy-gated model routing; Remove purpose-shaped context → Compare journey outcomes; Remove companion metadata → Compare journey outcomes; Remove explicit authority gates → Compare journey outcomes; Remove source descent → Compare journey outcomes; Remove dissent and counterevidence → Compare journey outcomes; Remove promotion/revalidation → Compare journey outcomes; Remove reload reconstruction → Compare journey outcomes; Replace bounded roles with generic agent → Compare journey outcomes; Collapse distinct cognitive movements → Compare journey outcomes; Remove post-response reintegration → Compare journey outcomes; Disable policy-gated model routing → Compare journey outcomes.

Every branch removes one layer but returns to the same end-to-end journey comparison, so the effect of that missing boundary is visible against the complete system. Because layers may reinforce or obstruct one another, active combinations must also be tested together rather than credited independently.

Ablations reveal whether the integrated architecture produces an effect and which elements carry it. A cumulative ladder should be followed by factorial tests over the most active components. Otherwise, gains that depend on interactions can be mistaken for independent additive improvements, while gains bought entirely with more context or inference can be mistaken for architecture.

7. Measurement families

Epistemic performance

Compositional performance

Human performance

System performance

Governance performance

8. The task-specific capability benchmark

The claim that localized structure can produce “frontier-like intelligence at a fraction of the cost” must be decomposed.

For each task class, compare:

ConditionModelContextReusable expertiseCost counted
Afrontierordinary promptnoneinference only and total system cost
BfrontierAIOS-composedAIOS guidanceinference and maintenance
Cmid-sizeordinary promptnoneinference only and total system cost
Dmid-sizeAIOS-composedAIOS guidanceinference and maintenance
Elocal/smallordinary promptnonehardware, energy, latency
Flocal/smallAIOS-composedAIOS guidancehardware, energy, authoring, maintenance

Evaluate accepted-outcome quality, not response fluency. Count human context preparation, review, correction, and failed attempts. A favorable task-specific result cannot be generalized to open-ended reasoning without further evidence.

9. Longitudinal experimental designs

Interrupted project study

Participants perform a multi-week research or writing task with controlled interruptions. Measure purpose retention, reorientation time, source fidelity, and artifact quality.

Contradictory-evidence injection

After an accepted conclusion and delay, introduce credible contrary evidence. Measure source recovery, correction scope, consumer recall, and changed later behavior.

Provider-switch study

Migrate the same project among model providers and a local model. Measure data conversion, prompt changes, behavioral regression, revalidation time, and retained semantic ground.

Practice-compounding study

Repeat a class of tasks. Compare ungoverned prompt reuse, static templates, and promotion with evidence/counterevidence/revalidation.

Asynchronous-attention study

Compare synchronous waiting, unmanaged parallel agents, and bounded asynchronous branches. Measure blocked attention, duplicated work, review load, integration quality, and the person’s control.

Organizational-variation study

Distribute a shared workflow package across teams with local adaptations. Measure standard conformance, legitimate variance, innovation return flow, and policy drift.

Crossed model–harness study

Run local, balanced, and frontier models under both a common minimal harness and a common strong harness. This separates model effect from harness lift and tests whether the gap changes by task family, evidence sufficiency, and verifiability.

Factorial composition study

Vary a selected set of high-value layers—purpose framing, operation-aware context, cognitive-mode separation, relational metadata, source provenance, bounded effects, and reintegration—within a fractional-factorial design. Measure main effects and interactions rather than assuming each layer contributes monotonically.

Human capability-retention study

Measure immediate assisted quality separately from later unaided transfer, independent error detection, fallback performance, and longitudinal decision quality. Compare fully delegated, human-first, AIOS mixed-initiative, and unaided conditions.

10. Cost accounting

Token price alone is an incomplete denominator. Total cost per accepted outcome includes:

model inference
+ context selection and transport
+ human explanation and reconstruction
+ review and verification
+ correction and rework
+ metadata and knowledge maintenance
+ provider migration
+ security and governance
+ failure recovery
+ infrastructure and energy

AIOS may reduce some categories while increasing others. The economic question is whether durable semantic capital lowers repeated reconstruction and failure cost enough to justify composition and governance overhead.

11. Evidence packages

Every capability experiment should preserve:

12. Research threats

ThreatMitigation
builder evaluates own systemindependent raters, preregistered criteria, blind artifact review
stronger underlying model explains resultmodel-controlled comparisons and ablations
curated tasks favor architectureadversarial and externally supplied tasks
novelty effectlongitudinal use and delayed measurement
metadata labor omitted from costfull human and maintenance accounting
provider behavior changesrecord versions and rerun sentinel tasks
private corpus prevents replicationpublish synthetic or consented benchmark packages
subjective “intelligence” scoremultiple operational measures and accepted-outcome criteria
automation hides authority lossnegative-authority tests and action logs

13. Falsification ledger

The research program should maintain an explicit ledger for claims that fail, narrow, or remain unresolved:

hypothesis_id: H8-task-specific-capability-amplification
status: open
task_scope: regulated-document-review
baseline: frontier-model-with-flat-retrieval
predicted_effect: comparable-accepted-outcome-at-lower-total-cost
disconfirming_result: quality-gap-greater-than-accepted-threshold
required_evidence:
  - blinded-domain-review
  - full-cost-accounting
  - repeated-provider-runs
  - failure-analysis
next_review: after-benchmark-v1

This is an illustrative record. Failed hypotheses are valuable architectural evidence, not marketing liabilities to be hidden.

14. What would count as a system-level contribution

Evidence for the integrated thesis would require more than one impressive demonstration. A credible program would show:

  1. repeatable gains on at least several reference journeys;
  2. benefits surviving provider and task variation;
  3. measurable contribution from the architecture beyond model quality;
  4. successful correction after interruption;
  5. preserved person authority under increasingly autonomous work;
  6. tolerable governance and maintenance cost;
  7. honest boundaries where simpler systems perform as well or better.

15. What recent research makes plausible

Current evidence supports the need for this system-level program but does not perform it on AIOS.

The implication is methodological: model, harness, context, task, resource budget, human role, and longitudinal integration must be measured together and then separated analytically. The system capability, context composition, cognitive agency, and AI-native architecture briefs provide the exact evidence and limitations.

16. What only direct AIOS evaluation can establish

Adjacent research cannot establish that the complete AIOS composition works, that its micro-improvements compound, that it protects human agency, or that it changes the economics of accepted outcomes. Those questions require the reference journeys, crossed comparisons, ablations, longitudinal studies, independent review, and full cost accounting described here.

The most important result would not be a single high score. It would be a reproducible map showing where the architecture improves outcomes, where frontier capability remains decisive, where simpler systems are sufficient, where human cognition is strengthened or weakened, and which system relationships create the difference.

This chapter owns experimental method: baselines, ablations, measures, evidence thresholds, and falsification. Open Research Questions and the AIOS Experimental Program owns the unresolved questions, research priorities, and process by which findings should change the architecture.