Part VI · Evaluation, Economics, and Research Positioning
Capability Horizons, Experiments, Proof, and Falsification
General intelligence harnessSystem-level capabilityAccepted outcomeCumulative capability amplification
1. Research posture
AIOS should be evaluated neither as a single foundation model nor as a collection of isolated interface features. It is a proposed general intelligence harness whose claims concern the interaction of:
- person-owned purpose and authority;
- file-native artifacts and memory;
- purpose-shaped context composition;
- bounded model judgments;
- workflows, roles, plans, and asynchronous work;
- recommendation, promotion, correction, and forgetting;
- derived lineage, impact, and registry infrastructure.
This chapter is the publication’s evaluation companion. Its waves organize progressively stronger proofs of the architecture; they are not a product status report or an implementation roadmap.
The correct unit of proof is therefore a longitudinal human–system journey, supported by component evidence and compared with simpler baselines. The ten reference journeys provide the integrated product traces from which these experiments are constructed.
2. An evidence ladder
Different statements about the system require different kinds of evidence. Keeping them distinct allows the vision to remain ambitious without treating architecture, implementation, and measured effect as the same thing.
| Level | Meaning | What supports it |
|---|---|---|
| Architectural proposition | What AIOS is designed to do and why its parts relate | Complete mechanism, explicit boundaries, and internal consistency |
| Implemented behavior | What a particular build actually does | Inspectable source, runtime behavior, and reproducible traces |
| End-to-end demonstration | Whether one complete journey works under defined conditions | Frozen task, acceptance criteria, exact context, result, reload, and review evidence |
| Replicated effect | Whether an outcome survives repeated users, tasks, or models | Preregistered comparisons, repeated runs, and independent evaluation |
| Bounded generalization | Whether an effect transfers beyond the original conditions | Diverse tasks, domains, users, model classes, and declared limits |
| Research hypothesis | A causal or economic effect the architecture makes plausible | Comparative experiments capable of disconfirming it |
| Long-range implication | A future possibility that follows if several hypotheses hold | Scenario analysis, enabling conditions, and visible uncertainty |
Architecture is not empirical performance. A capability that is well specified but not built remains a target. A feature that exists but lacks semantic evaluation remains implemented, not demonstrated.
3. Capability waves
Capability waves and their proof thresholds
How evidence advances from governing boundaries and one complete loop toward long-running tests of compounding intelligence and bounded autonomy.
Text equivalent
Wave 0: Constitution and evidence boundary → Wave 1: Minimum complete composition; Wave 1: Minimum complete composition → Wave 2: Standard composition workspace; Wave 2: Standard composition workspace → Wave 3: Core reasoning and composition; Wave 3: Core reasoning and composition → Wave 4: Workflows and recommendations; Wave 4: Workflows and recommendations → Wave 5: Projects, plans, and role teams; Wave 5: Projects, plans, and role teams → Wave 6: Compounding local intelligence; Wave 6: Compounding local intelligence → Wave 7: Extended autonomy experiments.
Each wave earns the next by satisfying an exit condition, not by adding features. The sequence moves from a rebuildable evidence boundary through coherent composition, workflows, projects, and governed memory before extended-autonomy experiments can claim system-level effects.
Wave 0 — constitution, evidence boundary, and rebuild vehicle
Exit: the target architecture, source authority, prohibited constructs, rebuild basis, and first feature contract are settled. No implementation sprint should outrun this boundary.
Wave 1 — minimum complete composition loop
Exit: one provider-backed, file-backed, interface-visible movement passes independent semantic, product, and implementation review from clean state through reload.
Wave 2 — standard composition workspace
Exit: common context, command, landing, failure, metadata-projection, and reconstruction contracts exist; navigation performs no semantic effects.
Wave 3 — core reasoning and composition capabilities
Exit: thinking, writing, editing, structure, planning, review, and integration operate on one semantic situation without duplicating truth or context architecture.
Wave 4 — workflows, stages, recommendations, and completion
Exit: a real workflow survives interruption and returns a completed result into the correct artifact, plan, and interface ground.
Wave 5 — projects, hierarchical planning, and role teams
Exit: a multi-session project integrates bounded roles, compares conflicting returns, changes direction responsibly, and resumes without transcript reconstruction.
Wave 6 — compounding local intelligence and memory governance
Exit: reusable intelligence is accepted, used, challenged, corrected, narrowed, retired, or superseded with complete lineage.
Wave 7 — extended autonomy and contextual-intelligence experiments
Exit: comparative studies show gains in correction, continuity, source fidelity, human authority, or productive reasoning without unacceptable false confidence, disclosure, cost, or control burden.
4. Core hypotheses
H1 — purpose-shaped composition
For complex longitudinal work, a purpose-shaped context composed from durable semantic ground will outperform raw transcript continuation and undifferentiated retrieval on relevance, source fidelity, correction, and person-rated usefulness.
Could be falsified by: no material advantage over simpler retrieval after controlling for token count and model; higher omission rates; excessive authoring burden.
H2 — recursive Why–How–What coherence
Preserving Why–How–What across projects, plans, stages, artifacts, and local edits will reduce objective drift and improve reintegration after interruption.
Could be falsified by: no measurable drift reduction; increased rigidity; users unable to maintain the structure; equivalent outcomes from ordinary outlines.
H3 — file-native reconstruction
Durable artifacts plus companion metadata can reconstruct the operative semantic situation without replaying full conversational history.
Could be falsified by: frequent dependence on hidden session state; low reconstruction agreement; unacceptable manual repair; artifacts too polluted for ordinary use.
H4 — governed promotion
Evidence-linked promotion and revalidation will improve repeated task performance without converting historical practices into rigid universal rules.
Could be falsified by: harmful practice ossification; high false-promotion rates; little later reuse; governance cost greater than repeated reinvention.
H5 — attention-following asynchronous work
Bounded background branches with inspectable returns will reduce human waiting and interruption cost without reducing oversight or integration quality.
Could be falsified by: return-review burden exceeds waiting saved; lost situational awareness; duplicated work; excessive notification and attention fragmentation.
H6 — correction across interruption
Evidence, rationale, supersession, and consumer lineage will enable materially better long-horizon correction than version history or transcript memory alone.
Could be falsified by: missed consumers, false propagation, no change in later reasoning, or greater correction cost than manual expert reconstruction.
H7 — provider portability
Keeping durable knowledge, context sources, workflows, and authority outside a provider can reduce switching cost and preserve continuity across models.
Could be falsified by: behavior depends so heavily on provider-specific features that migration is impractical; adapter and revalidation cost eliminates the benefit.
H8 — task-specific capability amplification
High-quality local context and reusable expertise may allow less expensive or smaller models to approach stronger models on bounded, knowledge-intensive tasks.
Could be falsified by: no task-quality convergence after controlling for latency and cost; smaller models fail on judgment despite excellent context; maintenance costs dominate savings.
This is not a claim of general frontier equivalence.
H9 — interaction-dependent system capability
The complete person–model–artifact system will produce outcome differences that cannot be explained by adding the average effects of its components independently. Some context, memory, movement, verification, and reintegration layers will reinforce one another while others will be redundant or antagonistic.
Could be falsified or narrowed by: performance explained almost entirely by model choice or added tokens; no repeatable interaction effects; simpler harnesses matching the full system after resource control.
H10 — bounded semantic sovereignty
Allowing models to make genuine semantic judgments inside explicit jurisdictions, while deterministic systems retain custody of exact effects, will outperform both keyword-driven decision extraction and broadly autonomous agents on accepted-outcome quality, recoverability, and unauthorized-effect rate.
Could be falsified or narrowed by: semantic decisions proving too variable for useful mediation; deterministic baselines producing equivalent quality; authority boundaries creating prohibitive friction; or hidden action paths bypassing the jurisdiction.
H11 — meaningful human agency
Interfaces that preserve task-relevant cognitive acts while moving mechanical latency into asynchronous work will improve reliance calibration and retained human capability relative to either unaided work or end-to-end automation.
Could be falsified or narrowed by: unchanged error detection; greater rubber-stamping; weaker later unaided performance; review debt exceeding attention saved; or benefits limited to user preference rather than capability.
H12 — policy-gated dynamic routing
Selecting a model–harness configuration for each bounded movement after applying privacy and authority constraints will reduce total cost or latency while maintaining a declared quality threshold and escalating difficult cases appropriately.
Could be falsified or narrowed by: poor routing transfer, privacy leakage, frequent under-escalation, provider-specific instability, or routing overhead exceeding the savings.
H13 — staged reasoning and situation renewal
For multifaceted tasks, movement-specific reasoning with purpose-shaped contexts, landed intermediate products, and recomposition may outperform one extended reasoning trace under a matched total resource budget. The proposed advantage is not that longer reasoning is inherently weak. It is that different cognitive movements can receive different grounds, products can be evaluated before they become later premises, and each accepted delta can renew the situation for the next judgment.
The comparison should vary task horizon, decomposition quality, model class, and whether the task has a strong deterministic verifier. It should measure correct-answer preservation, unsupported claims, error propagation, handoff loss, context and inference cost, latency, human preparation and review, and accepted-outcome quality.
Could be falsified or narrowed by: one extended trace or repeated self-refinement matching or exceeding the staged condition after total resources are controlled; staging overhead dominating any quality gain; poor intermediate judgments compounding error; benefits appearing only in one model or task family; or an expert-designed simple workflow matching the full composition.
H14 — intent engineering and purpose traceability
When a person begins with a tension, partial aim, or internally conflicted direction, an explicit intent-engineering movement may improve later planning, reasoning, and artifact quality by clarifying possible purposes, preserving which purpose received intentional standing, and making foreground work traceable back to it. Human commitment is the experimental boundary: model clarification should improve the decision surface without substituting a model-selected objective for the person's purpose.
Evaluation should compare fixed-request execution, open model-led goal inference, person-only clarification, and joint clarification with explicit commitment and purpose traceability. Measure later objective drift, plan and artifact alignment, responsible purpose revision, correction burden, person understanding, model anchoring, and whether consequential work can be explained upward through commitments, architecture, strategy, and outcome.
Could be falsified or narrowed by: no advantage when initial purpose is incomplete; added ceremony or latency outweighing alignment gains; model suggestions anchoring or displacing the person's own purpose; traceability rewarding superficial conformity; or prior intent becoming harder to reopen despite contrary results.
5. Baselines
Every evaluation should select relevant simpler alternatives:
- model with raw conversational transcript;
- model with a hand-written summary;
- flat document retrieval by lexical search;
- embedding-based retrieval over the same corpus;
- ordinary project files without companion semantic metadata;
- conventional workflow or task tracker plus model chat;
- expert human performing manual reconstruction;
- provider-native memory or project context;
- the same model under a minimal and an AIOS-like harness;
- multiple models under the same minimal and strong harnesses;
- AIOS with one subsystem ablated.
- one extended reasoning trace under a matched inference budget;
- repeated self-refinement over one accumulating context;
- an expert-designed simple workflow with the same model and review budget;
- task-only, full-corpus, and composed-commission context conditions.
The aim is not to defeat every baseline on every measure. It is to identify the conditions under which added architecture earns its complexity.
6. Ablation program
Testing the architecture one boundary at a time
Removing each major layer from the full architecture reveals which relationships improve journey outcomes, add no value, or interact.
Text equivalent
Full reference architecture → Remove purpose-shaped context; Full reference architecture → Remove companion metadata; Full reference architecture → Remove explicit authority gates; Full reference architecture → Remove source descent; Full reference architecture → Remove dissent and counterevidence; Full reference architecture → Remove promotion/revalidation; Full reference architecture → Remove reload reconstruction; Full reference architecture → Replace bounded roles with generic agent; Full reference architecture → Collapse distinct cognitive movements; Full reference architecture → Remove post-response reintegration; Full reference architecture → Disable policy-gated model routing; Remove purpose-shaped context → Compare journey outcomes; Remove companion metadata → Compare journey outcomes; Remove explicit authority gates → Compare journey outcomes; Remove source descent → Compare journey outcomes; Remove dissent and counterevidence → Compare journey outcomes; Remove promotion/revalidation → Compare journey outcomes; Remove reload reconstruction → Compare journey outcomes; Replace bounded roles with generic agent → Compare journey outcomes; Collapse distinct cognitive movements → Compare journey outcomes; Remove post-response reintegration → Compare journey outcomes; Disable policy-gated model routing → Compare journey outcomes.
Every branch removes one layer but returns to the same end-to-end journey comparison, so the effect of that missing boundary is visible against the complete system. Because layers may reinforce or obstruct one another, active combinations must also be tested together rather than credited independently.
Ablations reveal whether the integrated architecture produces an effect and which elements carry it. A cumulative ladder should be followed by factorial tests over the most active components. Otherwise, gains that depend on interactions can be mistaken for independent additive improvements, while gains bought entirely with more context or inference can be mistaken for architecture.
7. Measurement families
Epistemic performance
- factual and source fidelity;
- unsupported-claim rate;
- uncertainty calibration;
- contradiction detection precision and recall;
- alternative preservation;
- correction uptake;
- provenance completeness.
Compositional performance
- purpose retention;
- cross-section coherence;
- plan-to-task alignment;
- artifact quality;
- integration sufficiency;
- drift after interruption;
- recovery of the next meaningful movement.
Human performance
- time to useful continuation;
- attention blocked by waiting;
- orientation time after interruption;
- correction and review burden;
- perceived control and intelligibility;
- error-detection rate;
- trust calibration rather than undifferentiated trust;
- quality of later unaided performance;
- retained fallback capability when AI is unavailable;
- frequency of meaningful redirection rather than ceremonial approval.
System performance
- context tokens and monetary cost per accepted outcome;
- latency distribution;
- reconstruction equality;
- write and collision integrity;
- provider migration effort;
- derived-index rebuild consistency;
- local/remote disclosure volume;
- failure recovery time.
- accepted-outcome quality at a fixed total resource budget;
- routing accuracy, escalation precision, and abstention quality;
- performance after model or index substitution.
Governance performance
- unauthorized-effect rate;
- promotion precision;
- stale-guidance rate;
- supersession completeness;
- dissent retention;
- audit-trace completeness;
- review effort by risk class.
8. The task-specific capability benchmark
The claim that localized structure can produce “frontier-like intelligence at a fraction of the cost” must be decomposed.
For each task class, compare:
| Condition | Model | Context | Reusable expertise | Cost counted |
|---|---|---|---|---|
| A | frontier | ordinary prompt | none | inference only and total system cost |
| B | frontier | AIOS-composed | AIOS guidance | inference and maintenance |
| C | mid-size | ordinary prompt | none | inference only and total system cost |
| D | mid-size | AIOS-composed | AIOS guidance | inference and maintenance |
| E | local/small | ordinary prompt | none | hardware, energy, latency |
| F | local/small | AIOS-composed | AIOS guidance | hardware, energy, authoring, maintenance |
Evaluate accepted-outcome quality, not response fluency. Count human context preparation, review, correction, and failed attempts. A favorable task-specific result cannot be generalized to open-ended reasoning without further evidence.
9. Longitudinal experimental designs
Interrupted project study
Participants perform a multi-week research or writing task with controlled interruptions. Measure purpose retention, reorientation time, source fidelity, and artifact quality.
Contradictory-evidence injection
After an accepted conclusion and delay, introduce credible contrary evidence. Measure source recovery, correction scope, consumer recall, and changed later behavior.
Provider-switch study
Migrate the same project among model providers and a local model. Measure data conversion, prompt changes, behavioral regression, revalidation time, and retained semantic ground.
Practice-compounding study
Repeat a class of tasks. Compare ungoverned prompt reuse, static templates, and promotion with evidence/counterevidence/revalidation.
Asynchronous-attention study
Compare synchronous waiting, unmanaged parallel agents, and bounded asynchronous branches. Measure blocked attention, duplicated work, review load, integration quality, and the person’s control.
Organizational-variation study
Distribute a shared workflow package across teams with local adaptations. Measure standard conformance, legitimate variance, innovation return flow, and policy drift.
Crossed model–harness study
Run local, balanced, and frontier models under both a common minimal harness and a common strong harness. This separates model effect from harness lift and tests whether the gap changes by task family, evidence sufficiency, and verifiability.
Factorial composition study
Vary a selected set of high-value layers—purpose framing, operation-aware context, cognitive-mode separation, relational metadata, source provenance, bounded effects, and reintegration—within a fractional-factorial design. Measure main effects and interactions rather than assuming each layer contributes monotonically.
Human capability-retention study
Measure immediate assisted quality separately from later unaided transfer, independent error detection, fallback performance, and longitudinal decision quality. Compare fully delegated, human-first, AIOS mixed-initiative, and unaided conditions.
10. Cost accounting
Token price alone is an incomplete denominator. Total cost per accepted outcome includes:
model inference
+ context selection and transport
+ human explanation and reconstruction
+ review and verification
+ correction and rework
+ metadata and knowledge maintenance
+ provider migration
+ security and governance
+ failure recovery
+ infrastructure and energy
AIOS may reduce some categories while increasing others. The economic question is whether durable semantic capital lowers repeated reconstruction and failure cost enough to justify composition and governance overhead.
11. Evidence packages
Every capability experiment should preserve:
- frozen task and acceptance contract;
- source corpus and authority map;
- model/provider/version and settings;
- exact provider-visible context;
- prompt and response contract;
- raw and parsed outputs where policy allows;
- person actions and authority decisions;
- artifact and metadata revisions;
- failures, retries, latency, tokens, and costs;
- semantic, product, and implementation reviews;
- reload or interruption evidence;
- comparison and ablation results;
- limitations and non-generalizable conditions.
12. Research threats
| Threat | Mitigation |
|---|---|
| builder evaluates own system | independent raters, preregistered criteria, blind artifact review |
| stronger underlying model explains result | model-controlled comparisons and ablations |
| curated tasks favor architecture | adversarial and externally supplied tasks |
| novelty effect | longitudinal use and delayed measurement |
| metadata labor omitted from cost | full human and maintenance accounting |
| provider behavior changes | record versions and rerun sentinel tasks |
| private corpus prevents replication | publish synthetic or consented benchmark packages |
| subjective “intelligence” score | multiple operational measures and accepted-outcome criteria |
| automation hides authority loss | negative-authority tests and action logs |
13. Falsification ledger
The research program should maintain an explicit ledger for claims that fail, narrow, or remain unresolved:
hypothesis_id: H8-task-specific-capability-amplification
status: open
task_scope: regulated-document-review
baseline: frontier-model-with-flat-retrieval
predicted_effect: comparable-accepted-outcome-at-lower-total-cost
disconfirming_result: quality-gap-greater-than-accepted-threshold
required_evidence:
- blinded-domain-review
- full-cost-accounting
- repeated-provider-runs
- failure-analysis
next_review: after-benchmark-v1
This is an illustrative record. Failed hypotheses are valuable architectural evidence, not marketing liabilities to be hidden.
14. What would count as a system-level contribution
Evidence for the integrated thesis would require more than one impressive demonstration. A credible program would show:
- repeatable gains on at least several reference journeys;
- benefits surviving provider and task variation;
- measurable contribution from the architecture beyond model quality;
- successful correction after interruption;
- preserved person authority under increasingly autonomous work;
- tolerable governance and maintenance cost;
- honest boundaries where simpler systems perform as well or better.
15. What recent research makes plausible
Current evidence supports the need for this system-level program but does not perform it on AIOS.
- SWE-agent used same-domain interface comparisons to show that action and observation design can materially change model performance.
- NoLiMa showed that the ability to admit a long context is not the same as effective access to the relevant relation within it.
- LongMemEval showed that long-term conversational memory requires temporal and relational reasoning rather than simple history retention.
- τ-bench demonstrated that valid tool use and policy access do not guarantee correct final state or consistent repeated performance.
- A 2024 meta-analysis of human–AI combinations found that human–AI systems do not automatically exceed the better standalone component.
- A 2026 multi-agent coordination study found that decomposition and coordination can produce gains or losses depending on the task and participating models.
- A 2025 randomized study of AI-supported learning found that immediate assisted performance and later unaided capability can diverge.
The implication is methodological: model, harness, context, task, resource budget, human role, and longitudinal integration must be measured together and then separated analytically. The system capability, context composition, cognitive agency, and AI-native architecture briefs provide the exact evidence and limitations.
16. What only direct AIOS evaluation can establish
Adjacent research cannot establish that the complete AIOS composition works, that its micro-improvements compound, that it protects human agency, or that it changes the economics of accepted outcomes. Those questions require the reference journeys, crossed comparisons, ablations, longitudinal studies, independent review, and full cost accounting described here.
The most important result would not be a single high score. It would be a reproducible map showing where the architecture improves outcomes, where frontier capability remains decisive, where simpler systems are sufficient, where human cognition is strengthened or weakened, and which system relationships create the difference.
This chapter owns experimental method: baselines, ablations, measures, evidence thresholds, and falsification. Open Research Questions and the AIOS Experimental Program owns the unresolved questions, research priorities, and process by which findings should change the architecture.