AIOS Proresearch
AIOS Intelligence System · Research Overview

Part VI · Evaluation, Economics, and Research Positioning

Capability Horizons, Experiments, Proof, and Falsification

Organizes the architecture into capability horizons and an evaluation program for the whole reasoning environment. Baselines, ablations, longitudinal studies, and domain packages make system-level contribution observable across real workflows.

THE BAND EVERY CLAIM IS IN

A claim moves up only on a comparison that could have gone the other way, counted against its full cost — and it comes back down when the result says so.

1. A system-level thesis requires system-level proof

AIOS makes claims about a relationship among durable human purpose, file-native memory, authored context, bounded model judgment, cognitive movements, workflows, asynchronous work, authority, correction, and renewal. Evaluating only a model response or an isolated feature cannot establish those claims.

The principal unit of proof is a longitudinal human–system journey. Component tests establish the integrity of its parts. The journey establishes whether purpose survives across those parts and produces the intended durable outcome under the proper authority.

This posture separates architecture from evidence. The architecture defines the system being built. Implementation establishes what exists. Demonstration establishes what the implementation did under declared conditions. General claims require progressively broader replication.

2. The claim ladder controls language

LevelWhat has been establishedProper claim form
ConstitutionalA founder-ratified law of the intended architecture“AIOS is designed so that…”
ImplementedSource and runtime evidence show a current behavior“The implementation currently…”
DemonstratedA declared end-to-end test passed under fixed conditions“The system demonstrated…”
ReplicatedThe effect recurred across relevant tasks, users, or models“The result replicated within…”
GeneralizedEvidence held beyond the original conditions“Evidence supports the effect across…”
HypotheticalA plausible causal or capability proposition awaits proof“We hypothesize…”
SpeculativeA useful direction remains too distant for a near-term claim“A future research direction is…”

A precise architecture can remain unimplemented. A working feature can remain undemonstrated at the level of semantic outcome. A successful demonstration can remain narrow. The ladder preserves ambition by making the next evidentiary obligation explicit.

Chapter 27 supplies the controlled vocabulary for authority, horizon, and lifecycle standing. This ladder supplies the evidentiary dimension. Constitutional corresponds to RATIFIED architectural ground; Implemented is supported by OBSERVED source or runtime evidence; Demonstrated is shared directly; Replicated and Generalized extend the evidence beyond one proof; Hypothetical corresponds to HYPOTHESIS; and Speculative remains within an OPEN research horizon. Chapter 28 uses the controlled standing in its standing column and carries evidentiary maturity in its evidence and transition fields.

3. Capability develops through seven waves

Diagram 1 · §3

flowchart LR W0["Wave 0: Constitution and evidence boundary"] --> W1["Wave 1: Minimum complete composition"] W1 --> W2["Wave 2: Standard composition workspace"] W2 --> W3["Wave 3: Core reasoning and composition"] W3 --> W4["Wave 4: Workflows and recommendations"] W4 --> W5["Wave 5: Projects, plans, and role teams"] W5 --> W6["Wave 6: Compounding local intelligence"] W6 --> W7["Wave 7: Extended autonomy experiments"]
24-capability-horizons-experiments-proof-and-falsification--m01.mmd
Text equivalent

Wave 0: Constitution and evidence boundary → Wave 1: Minimum complete composition; Wave 1: Minimum complete composition → Wave 2: Standard composition workspace; Wave 2: Standard composition workspace → Wave 3: Core reasoning and composition; Wave 3: Core reasoning and composition → Wave 4: Workflows and recommendations; Wave 4: Workflows and recommendations → Wave 5: Projects, plans, and role teams; Wave 5: Projects, plans, and role teams → Wave 6: Compounding local intelligence; Wave 6: Compounding local intelligence → Wave 7: Extended autonomy experiments.

Wave 0 establishes the architecture, source authority, architectural boundaries, evidence discipline, and first proof obligations.

Wave 1 establishes one complete movement from clean files through context, model judgment, durable landing, person-visible consequence, and reload.

Wave 2 establishes shared contracts for context, commands, landing, failure, metadata projection, and reconstruction while keeping navigation free of semantic effects.

Wave 3 carries the product's central value proposition: thinking, writing, editing, structure, planning, review, and integration operating on one semantic substrate without creating parallel truth for each surface.

Wave 4 extends the existing workflow library and recommendation architecture through durable stage products, artifact completion, return to origin, and interruption recovery.

Wave 5 extends project-scale planning, bounded roles, conflicting returns, responsible replanning, and multi-session resumption.

Wave 6 develops the full lifecycle of reusable knowledge: candidate formation, evidence, acceptance, use, challenge, correction, narrowing, dormancy, retirement, and supersession.

Wave 7 investigates extended agency and contextual-intelligence effects after the authority, continuity, and correction substrate can support them.

The waves organize development around usable capability. Thinking, writing, editing, and structure provide immediate value; workflows, projects, compounding knowledge, and external systems extend that value across larger fields of work.

4. Core hypotheses and their falsifiers

H1 — purpose-shaped context composition

For complex longitudinal work, context authored from durable semantic ground will improve relevance, source fidelity, correction, and person-rated usefulness over transcript continuation or undifferentiated retrieval.

The hypothesis narrows or fails if controlled comparisons show no material gain, omission rises, or authoring and maintenance burden outweigh the benefit.

H2 — cross-scale Fractal Seed coherence

Preserving Why → How → What across projects, plans, stages, artifacts, and local contributions will reduce objective drift and improve reintegration after interruption.

The hypothesis narrows or fails if ordinary hierarchy performs equally well, the grammar produces rigidity, or people cannot maintain the relationship across scales.

H3 — file-native reconstruction

Canonical artifacts plus companion memory can reconstruct the operative semantic situation without replaying the complete conversation that produced it.

The hypothesis narrows or fails if hidden session state remains necessary, independent reconstruction agreement stays low, repair burden becomes unacceptable, or operational structure degrades the artifacts.

H4 — governed promotion

Evidence-linked promotion and revalidation will improve repeated work while preventing local practices from hardening into universal rules.

The hypothesis narrows or fails under high false-promotion rates, harmful ossification, little later reuse, or governance cost greater than repeated reinvention.

H5 — attention-following asynchronous work

Bounded background branches with origin-bound returns will reduce waiting and interruption cost while preserving oversight and integration quality.

The hypothesis narrows or fails when review debt exceeds time saved, situational awareness declines, duplication grows, or returns fragment attention.

H6 — correction across interruption

Preserving evidence, rationale, scope, supersession, and consumer lineage will improve long-horizon correction over transcript memory or ordinary version history.

The hypothesis narrows or fails through missed consumers, excessive false propagation, unchanged later reasoning, or correction cost above expert manual reconstruction.

H7 — provider portability

Keeping durable knowledge, context sources, workflows, and authority outside the model provider will preserve more continuity and reduce semantic switching cost.

The hypothesis narrows or fails if provider-specific behavior dominates the system, adaptation and revalidation erase the benefit, or the stable contracts cannot support materially different models.

H8 — task-specific capability amplification

On bounded knowledge-intensive tasks, a less expensive model with mature local context and reusable expertise can approach or exceed a stronger model working from weak ground.

The hypothesis narrows or fails if the quality gap remains, if smaller models cannot perform the needed judgment despite strong context, or if authoring, maintenance, and review dominate any inference savings.

The claim concerns defined tasks and accepted outcomes, not general equivalence among models.

H9 — situation renewal versus continuous trace

On sustained, multi-faceted work without a native verifier, bounded judgments over a situation renewed between movements should outperform one continuous reasoning trace at the same token budget. The advantage should become more consequential as the horizon lengthens.

A matched-budget experiment holds the task, sources, model, acceptance criteria, and total token allowance constant. The continuous-trace condition spends that allowance in one uninterrupted trajectory over one static context. The renewal condition divides the allowance among bounded judgments and integrates accepted intermediate products into the ground composed for each next movement.

The hypothesis narrows or fails if the continuous trace matches or exceeds accepted outcomes, if integration cost consumes the benefit, or if the predicted difference does not grow with horizon length. Verifier-rich tasks remain a separate comparison because extended inference can use a native correctness signal there.

5. Baselines establish whether complexity earns itself

Relevant comparisons include raw transcript continuation, a human-written summary, lexical retrieval, embedding retrieval over the same corpus, ordinary project files without companion memory, a conventional task tracker plus chat, provider-native memory, expert manual reconstruction, one continuous extended trace, and AIOS with one subsystem removed.

Matched-budget conditionToken allocationGround between movements
Continuous extended traceOne uninterrupted trajectory over one static contextNo durable integration before completion
Staged situation renewalSeveral bounded judgments within the same total allowanceAccepted intermediate products renew the ground for the next movement

No architecture needs to defeat every baseline on every measure. The research question is where purpose-shaped composition, durable standing, authority, and correction produce enough value to justify their cost.

6. Ablation isolates contribution

Diagram 2 · §6

flowchart TD F["Full reference architecture"] --> A1["Remove purpose-shaped context"] F --> A2["Remove companion metadata"] F --> A3["Remove explicit authority gates"] F --> A4["Remove source descent"] F --> A5["Remove dissent and counterevidence"] F --> A6["Remove promotion/revalidation"] F --> A7["Remove reload reconstruction"] F --> A8["Replace bounded roles with generic agent"] A1 --> C["Compare journey outcomes"] A2 --> C A3 --> C A4 --> C A5 --> C A6 --> C A7 --> C A8 --> C
24-capability-horizons-experiments-proof-and-falsification--m02.mmd
Text equivalent

Full reference architecture → Remove purpose-shaped context; Full reference architecture → Remove companion metadata; Full reference architecture → Remove explicit authority gates; Full reference architecture → Remove source descent; Full reference architecture → Remove dissent and counterevidence; Full reference architecture → Remove promotion/revalidation; Full reference architecture → Remove reload reconstruction; Full reference architecture → Replace bounded roles with generic agent; Remove purpose-shaped context → Compare journey outcomes; Remove companion metadata → Compare journey outcomes; Remove explicit authority gates → Compare journey outcomes; Remove source descent → Compare journey outcomes; Remove dissent and counterevidence → Compare journey outcomes; Remove promotion/revalidation → Compare journey outcomes; Remove reload reconstruction → Compare journey outcomes; Replace bounded roles with generic agent → Compare journey outcomes.

The ablation program asks which relationships carry the measured effect. It prevents a strong underlying model, unusually expert participant, or polished interface from receiving credit for architectural elements that contributed nothing.

Interaction effects also matter. Context composition may depend on companion memory; correction may depend on lineage and context republication; authority may affect both trust calibration and outcome quality. The integrated system is tested both whole and in disciplined subtraction.

7. Measurement follows five performance families

Epistemic performance includes source fidelity, unsupported claims, confidence calibration, contradiction detection, alternative preservation, correction uptake, and provenance completeness.

Compositional performance includes purpose retention, artifact coherence, plan-to-work alignment, integration sufficiency, drift after interruption, and recovery of the next meaningful movement.

Human performance includes time to useful continuation, attention blocked by waiting, reorientation time, correction and review burden, perceived control, intelligibility, error detection, and calibrated trust.

System performance includes total context and inference cost, latency, reconstruction equality, write integrity, collision behavior, provider migration effort, index rebuild equality, disclosure volume, and failure recovery.

Governance performance includes unauthorized effects, promotion precision, stale guidance, supersession completeness, dissent retention, trace completeness, and review effort by risk class.

8. The capability-amplification benchmark counts the whole system

ConditionModel classGroundReusable expertiseCost boundary
AStrongest eligibleOrdinary requestNoneInference and total workflow cost
BStrongest eligibleAIOS-composedAIOS guidanceInference, composition, and maintenance
CMid-rangeOrdinary requestNoneInference and total workflow cost
DMid-rangeAIOS-composedAIOS guidanceInference, composition, and maintenance
ELocal or smallOrdinary requestNoneHardware, energy, latency, and review
FLocal or smallAIOS-composedAIOS guidanceHardware, authoring, maintenance, and review

The benchmark evaluates accepted outcomes rather than verbal fluency. It includes failed attempts, human preparation, verification, correction, and maintenance. A favorable result remains confined to its task class until broader evidence exists.

9. Longitudinal experiments test continuity

An interrupted-project study measures purpose retention, reorientation, source fidelity, and artifact quality across a multi-week writing or research project.

A contradictory-evidence study introduces credible conflict after an accepted conclusion and delay, then measures source recovery, correction scope, consumer precision, authority, and later behavior.

A provider-switch study moves one mature project among eligible models and measures data conversion, adapter change, behavioral regression, revalidation, and preserved semantic ground.

A practice-compounding study compares no reusable memory, static templates, ungoverned prompt reuse, and governed promotion across repeated work.

An asynchronous-attention study compares synchronous waiting, unmanaged parallel work, and origin-bound branches on blocked attention, duplication, review load, integration quality, and control.

An organizational-variation study distributes shared operational knowledge across teams and measures required consistency, legitimate variance, upward learning, and governance drift.

10. Total cost per accepted outcome is the economic denominator

inference: model and tool use
+ composition: context selection, authoring, and transport
+ reconstruction: human explanation and reorientation
+ review: verification and integration
+ failure: correction and rework
+ governance: metadata and knowledge maintenance, lineage, access, retention, change control, and security
+ switching: provider migration and revalidation
+ infrastructure: devices, networking, storage, backup, energy, operations, and recovery

AIOS intentionally adds composition and governance work. The economic hypothesis is that durable semantic capital reduces recurring reconstruction, error, migration, and coordination costs enough to earn that investment.

11. Every experiment produces an evidence package

The package preserves the frozen task and acceptance contract, source corpus and authority map, model identity and settings, exact provider-visible context, prompt and response contract, outputs permitted by policy, person actions and authority decisions, artifact and metadata revisions, failures and retries, latency and cost, independent reviews, reload or interruption evidence, baselines, ablations, and non-generalizable conditions.

The package makes both positive and negative results inspectable. Evidence remains bound to the system state and task that produced it.

12. Threats become part of the design

Builder evaluation calls for independent raters and blind artifact review. Model-quality confounding calls for controlled comparisons. Curated tasks call for adversarial and externally supplied fixtures. Novelty effects call for delayed longitudinal measures. Hidden metadata labor calls for complete time accounting. Provider drift calls for recorded versions and sentinel reruns. Private corpora call for synthetic or consented replication packages. Broad intelligence scores call for operational measures. Authority loss calls for negative-effect tests.

The research program assumes that simpler explanations may win and designs comparisons capable of showing it.

13. Falsification remains a durable record

One illustrative falsification record can take this form:

hypothesis:
  identity: H8-task-capability-amplification
  standing: open
  task_scope: bounded-document-review
  comparison: strong-model-with-flat-ground
  predicted_result: comparable-accepted-outcome-at-lower-total-cost
  disconfirming_condition: quality-gap-exceeds-accepted-threshold
  required_evidence:
    - blind-domain-review
    - complete-cost-accounting
    - repeated-model-runs
    - failure-analysis
  review_after: benchmark-one

A failed hypothesis can narrow a claim, change a capability wave, or reveal that an architectural cost is not justified. That evidence belongs in the system's lineage.

Boundary: proof remains proportional to the tested claim

A journey shows how the architecture operates in one declared setting. Broader research can then compare results across tasks, models, interruptions, authority conditions, and governance cost to identify where the system-level contribution appears.

14. Research completion criteria

The evaluation program succeeds when it can show where AIOS improves accepted outcomes, where it merely adds machinery, which mechanisms carry the difference, how much the difference costs, and which conditions reverse the result.

The architecture earns its strongest claims through falsifiable experiments. Every open hypothesis is an invitation to discover the actual boundary of the system rather than a placeholder for future certainty.