AIOS Proresearch
AIOS Intelligence System · Research Overview

Appendix B

Open Research Questions and the AIOS Experimental Program

A developed architecture creates a more precise research agenda

AIOS does not depend on external research for the origin or internal coherence of its architecture. The Fractal Seed, semantic situation, cognitive movements, context-composition model, file-native memory, authority boundaries, asynchronous engagement model, and knowledge-evolution process have been developed as one system.

What remains open is not whether those concepts have been articulated. It is how their integrated effects behave across people, models, domains, risks, and time.

This distinction allows the publication to hold two positions at once:

  1. AIOS is a complete and ambitious paradigm for locally governed general intelligence infrastructure.
  2. Its capability, human, security, and economic hypotheses should be tested directly rather than inferred from architectural coherence or adjacent research.

The research program therefore asks where the architecture changes accepted outcomes, which relationships create the change, what they cost, when they fail, and how the system should evolve in response.

The integrated hypothesis map

The major propositions are related rather than independent.

How the system hypotheses interact

How durable ground, composed context, bounded movements, model selection, effects, and reintegration jointly shape accepted outcomes.

flowchart TB A["Durable local semantic ground"] --> B["Purpose- and operation-aware context"] B --> C["Bounded cognitive movement"] C --> D["Selected model and method"] D --> E["Semantic judgment"] E --> F["Validated and authorized effect"] F --> G["Review and reintegration"] G --> A B --> H["Lower reconstruction and inference burden"] C --> I["Less cognitive interference"] D --> J["Dynamic capability allocation"] F --> K["Failure containment and auditability"] G --> L["Semantic capital or semantic debt"] H --> M["Accepted-outcome quality, cost, and latency"] I --> M J --> M K --> M L --> M P["Human purpose, attention, and authority"] --> B P --> C P --> F M --> P
System-level capability29-open-research-questions-and-the-aios-experimental-program--m01.mmd
Text equivalent

Durable local semantic ground → Purpose- and operation-aware context; Purpose- and operation-aware context → Bounded cognitive movement; Bounded cognitive movement → Selected model and method; Selected model and method → Semantic judgment; Semantic judgment → Validated and authorized effect; Validated and authorized effect → Review and reintegration; Review and reintegration → Durable local semantic ground; Purpose- and operation-aware context → Lower reconstruction and inference burden; Bounded cognitive movement → Less cognitive interference; Selected model and method → Dynamic capability allocation; Validated and authorized effect → Failure containment and auditability; Review and reintegration → Semantic capital or semantic debt; Lower reconstruction and inference burden → Accepted-outcome quality, cost, and latency; Less cognitive interference → Accepted-outcome quality, cost, and latency; Dynamic capability allocation → Accepted-outcome quality, cost, and latency; Failure containment and auditability → Accepted-outcome quality, cost, and latency; Semantic capital or semantic debt → Accepted-outcome quality, cost, and latency; Human purpose, attention, and authority → Purpose- and operation-aware context; Human purpose, attention, and authority → Bounded cognitive movement; Human purpose, attention, and authority → Validated and authorized effect; Accepted-outcome quality, cost, and latency → Human purpose, attention, and authority.

The central loop is the architecture under test; the side branches name proposed mechanisms, not measured gains. All five mechanisms converge on outcome quality, cost, and latency, which then return to human purpose and authority rather than closing into autonomous optimization.

The architecture predicts that several individually modest changes can interact. Better metadata may not improve a result unless it changes recall. Better recall may not help unless context is composed for the operation. Better context may not improve accepted outcomes unless the movement, model, verification, and reintegration path are also appropriate. A safe action boundary may contain damage without improving the quality of the underlying judgment. A strong result may add no long-term value if it never returns to durable ground.

The central empirical question is therefore not whether each feature is useful in isolation. It is whether the relationships among them create a repeatable system-level effect.

Ten central research propositions

1. Inference-burden transfer

Proposition: Durable purpose, terminology, decisions, relationships, provenance, and artifact state can move recurring work out of model inference and into the environment.

Evidence needed: Same-model comparisons using identical tasks and controlled resource budgets, including immediate-frame-only, flattened long-context, and cross-resolution conditions; analysis of tokens spent reconstructing versus solving; unsupported-claim rates; source use; accepted quality. A measurable cross-resolution context lock should predict outcomes after controlling for model, total context volume, and authoring effort.

What would narrow it: Little outcome difference after controlling for token volume and expert context authoring; no relationship between context-lock measures and accepted quality; or maintenance cost greater than the repeated inference burden avoided.

2. Cumulative capability amplification

Proposition: Purpose framing, operation-aware context, cognitive-mode separation, multi-resolution memory, relational metadata, bounded judgment, verification, and reintegration may produce a combined improvement larger than their isolated average effects.

Evidence needed: Component ablations followed by factorial experiments, held-out tasks, fixed and variable model conditions, and total resource accounting. Under matched budgets, compare one extended reasoning trace and repeated self-refinement with movement-specific reasoning whose intermediate products land, receive standing, and recompose the next situation.

What would narrow it: Gains explained almost entirely by added inference, one dominant component, benchmark-specific curation, or the underlying model; or staging overhead, handoff loss, and error propagation matching or exceeding any benefit from movement separation and situation renewal.

3. Capability substitution

Proposition: On bounded, evidenced, and verifiable work, a smaller or local model inside a mature semantic environment may meet or exceed the accepted-outcome threshold of a stronger model operating through a weak harness.

Evidence needed: Crossed model–harness trials across task families, blinded domain review, explicit quality thresholds, and total cost per accepted outcome.

What would narrow it: Persistent model-quality gaps, weak transfer beyond narrow tasks, routing errors, or context and maintenance costs that dominate the savings.

This proposition does not imply general equivalence among models.

4. Frontier-model complementarity

Proposition: The same architecture can increase the usefulness of frontier models by giving them better ground and reserving their marginal capability for novel, ambiguous, high-consequence, or poorly externalizable movements.

Evidence needed: Frontier-model comparisons under minimal and AIOS-like harnesses; escalation studies; tasks with conflicting or incomplete evidence; measures of corrected and newly introduced error.

What would narrow it: No meaningful harness lift at the frontier or gains arising only through much larger token and inference budgets.

5. Hallucination-surface reduction

Proposition: Explicit sources, standing, narrower jurisdictions, meaningful absence, bounded outputs, and post-response verification can reduce the occasions on which a model must invent unsupported ground.

Evidence needed: Source-fidelity evaluation, missing-evidence cases, contradiction injection, attribution audits, and comparison with retrieval-only and long-context baselines.

What would narrow it: False confidence transferred into metadata or summaries, omission of decisive evidence, or polished unsupported claims surviving all downstream layers.

The architecture may reduce a failure surface; it cannot make every contextual statement true.

6. Longitudinal semantic capital

Proposition: Well-governed files, relationships, decisions, methods, and correction history can lower future reconstruction, review, and error costs across repeated work.

Evidence needed: Multi-month repeated-task studies, measures of reuse and staleness, maintenance labor, correction propagation, and comparison with static templates or flat archives.

What would narrow it: Low reuse, accumulated contradictions, poisoned memory, rising curation burden, or semantic debt exceeding the value of retained structure.

7. Correction across interruption

Proposition: Preserving evidence, rationale, standing, consumers, and supersession paths will allow the system to correct accepted ground after a delay and change later behavior more reliably than transcript memory or version history alone.

Evidence needed: Delayed contradictory-evidence studies, current-state and historical queries, consumer discovery, proportional correction, and later-context verification.

What would narrow it: Records change without future reasoning changing, corrections propagate beyond their valid scope, or human experts must reconstruct the entire history manually.

8. Bounded semantic sovereignty

Proposition: Real model judgment inside explicit semantic jurisdictions, combined with deterministic custody of exact effects, can preserve more intelligence and contain more failure than either keyword-driven decision extraction or broad agent autonomy.

Evidence needed: Matched tasks under all three architectures; independent semantic review; injection and stale-state tests; unauthorized-effect rate; recovery; human authority and burden.

What would narrow it: Semantic judgments too variable to mediate safely, hidden bypasses around the action boundary, or deterministic alternatives achieving equivalent quality with materially lower complexity.

9. Human cognitive agency

Proposition: Fast foreground interaction, meaningful cognitive acts, inspectable asynchronous work, and consequential human authority can increase assisted quality while supporting a co-adaptive relationship that does not require cognitive surrender.

Evidence needed: Separate measures of immediate assisted performance, reliance calibration, error detection, later unaided transfer, retained fallback capability, longitudinal decision quality, attention, review debt, and the person's ability to frame, question, correct, and act as the durable system evolves.

What would narrow it: Rubber-stamping, weaker unaided performance, lost situational awareness, sunk-cost acceptance of completed returns, or attention fragmentation exceeding waiting saved.

10. Local and distributed intelligence economics

Proposition: If durable semantic capital remains locally owned and models become dynamically selectable resources, useful intelligence can become less provider-dependent, more private, more portable, and less costly for a meaningful share of ordinary work. Shared evidence contracts may also allow collaboration without centralizing every participant’s private knowledge.

Evidence needed: Real workload distributions, local task coverage, privacy and data-egress traces, provider-switch trials, full cost and energy accounting, organizational pilots, and peer exchange with provenance and revocation.

What would narrow it: Persistent remote dependence, high device or security burden, weak portability, rebound inference demand, standards capture, or collaboration that requires recreating centralized custody.

This is the empirical bridge from the architecture to the larger Personal AGI and decentralized-intelligence thesis.

Proposition ownership and evidence map

Each proposition descends in three directions: to the chapters that own the architectural mechanism, to the normalized research briefs that summarize adjacent evidence and limits, and to the experimental designs that can test the integrated AIOS claim. The evidence links do not transfer authorship of the proposition to the literature.

PropositionOwning architecture chaptersAdjacent evidence briefsDirect experimental path
1. Inference-burden transferSystem Intelligence, Context Composition, Continuous Context OptimizationContext Composition, System Capability and RoutingAblation program, cost accounting
2. Cumulative capability amplificationCognitive Movements, Context Composition, Memory EcologyContext Composition, System Capability and RoutingFactorial composition study, system-level contribution
3. Capability substitutionCapability Horizons, Research PositioningSystem Capability and Routing, Local and Sovereign AICrossed model–harness study, task-specific capability benchmark
4. Frontier-model complementarityCapability Horizons, Research PositioningSystem Capability and RoutingCrossed model–harness study, system-performance measures
5. Hallucination-surface reductionContext Composition, Memory Ecology, Self-CorrectionContext Composition, Relational MemoryEpistemic-performance measures, contradictory-evidence injection
6. Longitudinal semantic capitalMemory Ecology, Files and Metadata, Registration and PromotionRelational Memory, Economics and InstitutionsPractice-compounding study, longitudinal experimental designs
7. Correction across interruptionMemory Ecology, Self-CorrectionRelational MemoryInterrupted-project study, contradictory-evidence injection
8. Bounded semantic sovereigntyBounded Semantic Sovereignty, Product Mechanics, Reference JourneysAI-Native ArchitectureBounded-sovereignty hypothesis, governance-performance measures
9. Human cognitive agencyAsynchronous Agency, Distributed CognitionCognitive AgencyHuman capability-retention study, asynchronous-attention study
10. Local and distributed intelligence economicsEconomic Architecture, Research PositioningLocal and Sovereign AI, Economics and InstitutionsProvider-switch study, cost accounting

Research frontiers by system layer

Context and cognition

  1. What operational criteria distinguish the smallest complete context from merely short context?
  2. How do lookup, comparison, aggregation, contradiction resolution, temporal update, structural reasoning, creation, and verification change the optimal evidence arrangement?
  3. When does full-source reading outperform retrieval or summarization?
  4. How should relevant-information volume, distribution, and position enter an evidence budget?
  5. Which cognitive movements require separate contexts or model calls, and which separations create only ceremony?
  6. Does the Fractal Seed add measurable coherence beyond a strong ordinary outline or expert-authored brief?
  7. How should compression disclose loss without making ordinary interaction cognitively expensive?
  8. Can cross-resolution context lock be measured independently of context length, and does it predict accepted-outcome quality across models and task families?
  9. Under a fixed inference and review budget, when do concurrent bounded branch frames outperform one accumulating context, and when does reintegration cost erase the gain?
  10. When is a person's purpose already sufficiently formed, and when does an explicit intent-engineering movement improve later work rather than create unnecessary ceremony or model anchoring?
  11. Can purpose traceability be measured from foreground work through commitments, architecture, strategy, and outcome, and does stronger traceability improve responsible replanning without making prior intent rigid?

Relational memory and knowledge evolution

  1. Which memory layers are operationally necessary, and which are redundant descriptions of the same state?
  2. How should current-state questions differ from historical questions?
  3. When should temporal records use valid time, transaction time, or both?
  4. Which relationships should be authored, which may be inferred provisionally, and which should never be promoted without review?
  5. How can delayed memory poisoning be detected before a stored lesson influences later decisions?
  6. Which authority may accept a proposed memory, method, or relationship at different consequence levels?
  7. When should accepted knowledge be revalidated, narrowed, made dormant, retired, or deleted?
  8. How can direct-source order and missing relationships remain visible in graph or summary views?

Human authority and engagement

  1. Which cognitive acts preserve understanding without imposing generic friction?
  2. When should a person form an initial view before receiving model advice?
  3. What makes a consequential choice genuinely understandable and overridable?
  4. Which low-risk, reversible actions can proceed under standing grants?
  5. How should confidence change checking, evidence gathering, abstention, escalation, or review?
  6. When should asynchronous work notify, return quietly, wait, expire, or request a new grant?
  7. How can review debt and automation-induced anchoring be measured over time?
  8. What does recovery look like when the model or application is unavailable?
  9. Which longitudinal signals distinguish co-adaptation that expands human capability from dependence that weakens unaided judgment or meaningful control?

Safe judgment and action

  1. Which semantic judgments can be delegated reliably within explicit jurisdictions?
  2. When should a constrained menu be supplied before reasoning, and when should packaging occur only after open reasoning?
  3. How can syntax, semantics, identity, authorization, and effects be evaluated independently?
  4. What constitutes independence in model review: a different prompt, model, evidence source, method, or human judgment?
  5. How should prompt injection, compromised sources, stale state, and legitimately granted but harmful actions be tested?
  6. Can a trusted action display prove that the approved boundary event is the event that executed?
  7. How should systems record useful partial results when a structured control output fails?

Files, metadata, and the Flow Atlas

  1. What is the minimum companion metadata required for recognition, reconstruction, correction, and model switching?
  2. How can companion metadata records avoid drift, orphaning, duplication, and invisible authority?
  3. Which exact concurrent or atomic operations justify a narrow transactional store?
  4. Can every embedding, index, summary, and graph projection be deleted and rebuilt without loss of the meaning carried by canonical sources, artifacts, and accepted ground?
  5. How should impact analysis distinguish exact consumers from semantic review candidates?
  6. Which visual forms help people recognize relationships before opening complete files?
  7. How should long-lived schemas evolve without invalidating personal domains?
  8. What admission contract best preserves identity, provenance, custody, original files, batch consequences, inactive-by-default state, and later relevance selection across copy and bind policies?
  9. Which mechanical observations are needed for freshness and reconstruction, and how can the system prove that observation alone never initiated semantic work or consequence?

Models, routing, and local execution

  1. What is the right routing unit for real work, and how stable are task-family classifications?
  2. Which policy and privacy gates must precede cost–capability optimization?
  3. How should outcome histories be keyed to model version, endpoint, context policy, and task family?
  4. What signal should trigger escalation from local to specialized or frontier inference?
  5. How do model updates alter previously calibrated routes?
  6. How much energy, latency, and security maintenance shifts to the edge?
  7. Which ordinary workloads can remain local without unacceptable quality loss?

Economics and institutions

  1. What share of knowledge-work cost comes from repeated context reconstruction, correction, and provider-specific state?
  2. When does semantic capital create measurable option value or switching leverage?
  3. Which forms of central application infrastructure are displaced, which are merely moved, and which grow through rebound demand?
  4. How should the labor of authoring, maintaining, and governing a knowledge system be valued?
  5. Can attestable domain packages support regulated use without creating false certification or central standards capture?
  6. What update, rollback, revocation, and liability structures would those packages require?
  7. Can teams share accepted methods and evidence contracts while preserving local exceptions and private ground?
  8. Which business, educational, scientific, or public-sector settings benefit most from locally governed reasoning infrastructure?

The reference journeys

Individual component tests are necessary, but the architecture becomes meaningful through complete journeys. Each journey should begin from a human purpose, traverse the relevant system layers, survive interruption or error, and return a legible semantic delta.

JourneyWhat it testsDecisive completion evidence
Thinking into a durable documentMinimum Fractal Seed and file-continuity loopUseful judgment becomes an authorized artifact, reloads from files, and renews context without hidden session state.
Correction across interruptionEpistemic maintenance over timeContrary evidence reopens the right claim, preserves lineage, reaches affected consumers, and changes later reasoning.
Workflow through completionDistinct cognitive movements and dependent stagesStage products remain coherent, the whole artifact meets criteria, and completion returns to the correct origin.
Structural editingMovement across document altitudesA high-level plan produces bounded local edits and a whole-document review without uncontrolled regeneration.
Hierarchical planning and replanningPurpose continuity across branchesA changed assumption updates strategy, architecture, commitments, and foreground proportionally.
Delegated research and reintegrationDistributed cognition without a central agentIndependent returns disclose their ground and differences, pass sufficiency review, and alter the named receiving object.
Recommendations and ExploreSemantic action fields and human authorityThe person understands the meaningful alternatives, retains an open path, and explicitly grants any consequence.
Practice promotion and revalidationBottom-up learning and top-down reuseA pattern moves through evidence and acceptance, improves later work, and can be narrowed or retired with lineage.
Local domain with external models or agentsSovereign custody and bounded exchangeMinimum necessary context leaves the domain, the return is validated locally, and no external system mutates accepted ground directly.
Team standardization with local discretionShared coherence without total centralizationA common package remains attestable while local variation and upward learning remain visible.
Admission of existing artifacts and sourcesLocal custody without automatic trust or activationPreview, authorization, provenance, stable identity, inactive-by-default reconstruction, and later relevance-based selection remain separate and verifiable.

The detailed state traces appear in Reference Journeys and Semantic State Traces. The experimental comparisons and measures appear in Capability Horizons, Experiments, Proof, and Falsification.

A prioritized experimental sequence

Program 1 — Minimum complete composition

Demonstrate one full movement from purpose and exact source ground through model-visible context, semantic judgment, authorized file effect, companion metadata, reintegration, and reconstruction after restart.

This is first because every later hypothesis depends on a trustworthy semantic and mechanical substrate.

Program 2 — Context and cognitive-movement comparison

Compare raw transcript, expert summary, lexical search, embedding retrieval, full-source context, and operation-aware AIOS composition under matched models and budgets. Within the composition condition, compare immediate-frame-only, flattened long-context, and cross-resolution alignment among the immediate frame, active semantic situation, and durable domain ground. Cross these context conditions with single-call and movement-specific cognitive conditions.

The cognitive comparison should explicitly include one extended reasoning trace, repeated self-refinement over one accumulating context, an expert-designed simple workflow, and staged movements with landed intermediate products and fresh recomposition. A second context comparison should test task-only, full-corpus, and composed-commission conditions in which stable governing ground and relevant patterns are present while deeper sources remain available through question-led descent. Total inference, context-authoring, and human review resources should be matched or reported separately.

For cases that begin with a tension or incomplete aim, add an intent condition comparing direct execution, open model-led goal inference, person-only clarification, and joint clarification followed by explicit human commitment. Track whether later work remains purpose-traceable and whether contrary results can reopen the purpose without erasing its history.

Measure accepted artifact quality, source fidelity, omission, unsupported claims, correct-answer preservation, error propagation, handoff loss, latency, tokens, and human preparation or review cost. Separate verifier-rich from verifier-poor work and test whether any staged advantage changes with task horizon, decomposition quality, or model class.

Program 3 — Component and interaction ablation

Begin with the strongest complete configuration, remove one layer at a time, and identify the most active components. Follow with a fractional-factorial study over those components to estimate reinforcement, redundancy, and antagonism.

Program 4 — Correction and memory security

After a meaningful delay, introduce credible contrary evidence and a separate adversarial memory candidate. Measure current-state resolution, historical recovery, consumer discovery, promotion rejection, correction propagation, and later behavior.

Program 5 — Human agency and attention-following work

Compare unaided, fully delegated, synchronous AI-assisted, and AIOS mixed-initiative conditions. Measure immediate quality, later unaided transfer, error detection, reliance calibration, blocked attention, return review, redirection, fallback capability, and whether repeated use improves or weakens the person's capacity to frame and govern subsequent work.

Program 6 — Model–harness and routing frontier

Run local, balanced, specialized, and frontier models under both minimal and strong harnesses. Evaluate policy eligibility, routing, escalation, cost, latency, energy, source fidelity, and accepted quality across several task families.

Program 7 — Longitudinal semantic capital

Repeat a class of real work over months under no durable memory, static templates, generic memory, and governed AIOS promotion. Measure reuse, reconstruction, error, staleness, semantic debt, maintenance labor, and correction.

Program 8 — Organization and network pilots

Only after the local substrate is stable, test shared workflow or domain packages, local adaptations, attestable evidence, signed updates, rollback, revocation, peer exchange, and return of candidate improvements.

Evidence and reproducibility

Where privacy permits, each experiment should preserve:

Private domains need not be exposed to make the method reproducible. Synthetic, consented, or independently held task packages can preserve the system contracts and evaluation logic while protecting personal ground.

How the architecture should respond to evidence

A mature research program must allow several outcomes:

Negative and null results are not failures of publication. They are inputs to the self-correcting knowledge system. The relevant chapter, glossary definition, model-routing policy, workflow, or promoted practice should be narrowed or revised; the prior rationale and evidence should remain visible; later contexts should use the corrected form.

This is how the research overview can become a living instance of the architecture rather than a frozen defense of its first articulation.

Closing research thesis

The conventional research question asks how much intelligence is contained in a model. AIOS adds a second question:

How much useful intelligence can be created when purpose, context, memory, artifacts, semantic judgment, exact operations, human authority, and reintegration are designed as one evolving system?

If the architecture works as proposed, its significance will not be that a local knowledge system has replaced frontier models or eliminated centralized infrastructure. It will be that models become plural, adaptable, and selectively invoked cognitive resources—subject to compatible adapters, comparative evaluation, and revalidation—inside a durable environment owned by the people and organizations whose judgments gave it value.

That would change the unit of intelligence, the economics of inference, the locus of memory, the meaning of human participation, and the practical architecture through which AI becomes part of ordinary life. The purpose of the experimental program is to determine precisely where that future is already possible, where frontier capability remains essential, and which relationships make the difference.

The Glossary and Source Guide provides the common vocabulary and exact descent into the eight research programs. Together, the research library and the evaluation architecture allow every ambitious claim in this overview to remain visible, investigable, and open to correction.