AIOS Proresearch
AIOS Intelligence System · Research Overview

Full research memo: System-Level Capability Amplification and Dynamic Model Routing

1. Domain question

To what extent is useful AI performance a property of the complete reasoning system rather than the reasoning model alone, and can AIOS exploit that difference by routing bounded cognitive movements among frontier, balanced, small, and on-device models? The research tests whether purpose, context, decomposition, working state, memory, relational metadata, provenance, tools, verification, and integration can reduce inference burden and failure risk; when their gains interact or backfire; and how cumulative system effects, selective frontier use, and reusable semantic structure should be evaluated without mistaking harness intelligence for model intelligence.

2. Executive answer

The evidence supports a bounded but consequential answer: useful performance is jointly produced by the model, harness, task distribution, environment, and resource budget. Model capability remains important, especially for novel, ambiguous, open-ended, or high-consequence reasoning. Yet same-model ablations and within-study comparisons show that task framing, interfaces, exact tools, context policy, workflow structure, and verification can materially change outcomes. In bounded settings with a good external signal, a smaller model plus a specialized tool or optimized workflow can match or exceed a larger model operating without that support.

This is not a universal scale-substitution result. The strongest demonstrations occur where a weak operation can be externalized and checked: calculation, code execution, repository repair, retrieval, structured output, or verifier-guided mathematical search. Scaffolding cannot reliably supply missing world knowledge, invent a valid acceptance function, or resolve genuinely ambiguous judgment. Longer reasoning is not automatically better, and additional context, memory, critique, agents, or workflow stages can introduce interference, correlated error, security exposure, and coordination cost.

Current model gaps are also benchmark-sensitive. Rankings compress on some agent and mathematical evaluations, widen sharply on specialist or broad-knowledge tests, and can reverse for particular tasks. The same nominal model can perform differently across serving endpoints. AIOS should therefore route a bounded cognitive movement—not a whole project—using task type, evidence sufficiency, sensitivity, consequence, tool state, calibrated performance history, and verification results. The router should first enforce privacy and authority constraints, then select the least costly admissible model-harness configuration expected to meet the quality threshold, with explicit escalation to frontier capability when evidence, calibration, or checks are insufficient.

The ambitious implication is that accepted, versioned claims, relations, tests, decisions, and correction dependencies could become reusable semantic capital, shifting repeated work from fresh inference into inspectable structure. That implication is plausible but not yet directly established. Long-term memory studies show both retrieval benefits and large losses from staleness, compression, poisoning, and privacy extraction.

The appropriate conclusion is therefore architectural and experimental. AIOS has a credible basis for selective frontier use within a local, model-plural system, but should treat every amplification layer as a factor to test rather than a presumed improvement. The thesis should be accepted only if factorial ablations, distribution-shifted routing trials, and longitudinal reuse studies improve the accepted-outcome frontier after accounting for tokens, latency, energy, retries, human review, privacy, maintenance, correction propagation, and failure containment.

3. Essential findings

Finding 1 — Observed capability belongs to a model–harness system

Finding 2 — Scaffolding can substitute for scale on bounded, verifiable operations

Finding 3 — Context quality and placement matter more than maximum context volume

Finding 4 — Provenance improves auditability but does not establish truth

Finding 5 — Amplification layers interact and can be harmful

Finding 6 — Dynamic routing is feasible but remains fragile under change

Finding 7 — Tools and memory amplify capability and failure persistence together

Finding 8 — Semantic capital is a plausible longitudinal hypothesis, not a demonstrated result

4. How the evidence refines the AIOS account

  1. It supplies a bounded empirical formulation of system capability: model, harness, environment, task distribution, and budget interact rather than contribute independent fixed gains.
  2. It defines the bounded cognitive movement as the unit for model selection, verification, privacy control, and outcome measurement.
  3. It turns model plurality into a constrained routing problem with admissibility, capability allocation, and evidence-triggered escalation.
  4. It distinguishes factuality, faithfulness, attribution, and source quality, clarifying which provenance mechanisms support audit rather than truth itself.
  5. It provides a falsifiable program—crossed model/harness tests, fractional-factorial ablations, shifted routing trials, and longitudinal reuse—to test cumulative amplification and semantic capital.

5. Architectural boundaries preserved

  1. Routing remains subordinate to the established AIOS architecture of purpose, context, memory, tools, and integration rather than becoming the center of the account.
  2. AIOS remains complementary to frontier models, invoking frontier capability selectively where it changes the accepted outcome.
  3. Architectural availability remains distinct from task-level activation; not every layer is active for every task.
  4. Local execution remains distinct from complete privacy, just as retrieved citations remain distinct from factual correctness.
  5. Benchmark cost savings remain bounded results until accepted outcomes, review, maintenance, security, and model churn are measured.

6. Where the evidence connects to AIOS

Research contributionAIOS connectionOwning chapterWhy it matters
Capability is a joint model–harness propertyCore architectureFrom Model Intelligence to System IntelligenceExplains why AIOS layers can affect outcomes without diminishing model intelligence.
Bounded smaller-model substitutionResearch boundaryCapability Horizons, Experiments, Proof, and FalsificationEvidence is strong in verifiable domains but too conditional for an unrestricted claim.
Cognitive movement as the routing unitCore architectureCognitive Movements, Roles, and CarriersConnects decomposition to model choice, privacy, verification, and cost control.
Constrained router with selective frontier escalationCore architectureFrom Model Intelligence to System IntelligencePreserves frontier use while making privacy, capability, and consequence explicit.
Non-additive layer interactionsExperimental boundaryCapability Horizons, Experiments, Proof, and FalsificationPrevents cumulative demonstrations without ablation.
Provenance as auditability rather than truthCore distinctionFiles, Artifacts, Metadata, and Semantic StandingPrevents citation and grounding mechanisms from being overstated.
Tools and memory as capability–risk multipliersConditional implicationProduct Mechanics: Commands, Files, Records, and ReconstructionConnects amplification to persistent and state-changing failure modes.
Semantic capital hypothesisConditional implicationEconomic Architecture, Sovereignty, and Decentralized IntelligencePreserves the economic and epistemic implication while keeping it explicitly testable.
Factorial system benchmarkExperimental designCapability Horizons, Experiments, Proof, and FalsificationDistinguishes model effects, harness lift, interactions, and total system frontiers.
Exact current benchmark rankingsSource-level depthResearch Positioning, Novelty, and System-Level ContributionFast-changing values support rather than anchor the durable account.
Broad autonomous workflow searchOpen research directionOpen Research Questions and the AIOS Experimental ProgramSearch cost, overfitting, and transfer require prospective evidence before architectural integration.

7. Relationships across research programs

  1. Context composition × provenance: retrieval determines what evidence enters working context; provenance determines whether its use can be audited, refreshed, or reversed.
  2. Memory × security: durable state creates continuity and reuse, but also extends the lifetime of poisoning, privacy leakage, and obsolete instructions.
  3. Routing × privacy: model selection begins with endpoint admissibility and data minimization, not a pure accuracy-versus-cost ranking.
  4. Tools × verification: exact tools reduce computational errors only when argument selection, state changes, and postconditions are independently checked.
  5. Integration × learning economics: accepted outputs become useful reusable structure only if correction dependencies and invalidation propagate across later work.

8. Priority source set

DateSource typeExact supportable claimDirect link
2024-09Peer reviewed, EMNLP IndustryFully resourced long context outperformed tested RAG configurations on public QA, while self-routing retained comparable performance with lower compute.Li et al.
2024-12Peer reviewed, NeurIPSA purpose-built agent-computer interface materially improved repository-repair performance, showing that action and observation design affect measured capability.SWE-agent
2024-12Peer reviewed, NeurIPS benchmarkAcross 97 tasks and 629 attacks, tool-agent task completion and prompt-injection defense were both difficult.AgentDojo
2025Peer reviewed, ICLRAdaptive test-time compute was reported as more than four times as efficient as best-of-N and enabled bounded smaller-model scale substitution in mathematical reasoning.Snell et al.
2025Peer reviewed, ICLRLong-history performance fell 30–60% relative to oracle evidence, while targeted indexing helped and fact condensation created ability tradeoffs.LongMemEval
2025Peer reviewed, ICLRPreference-trained routing reduced cost by more than twofold at matched measured quality and transferred across tested model pairs.RouteLLM
2025Peer reviewed, ICMLWorkflow search improved six-benchmark performance by 5.7% on average and produced selected cases where smaller models exceeded GPT-4o at much lower reported inference cost.AFlow
2025Peer reviewed, npj Digital MedicineA specialized calculator raised Llama 3.1 70B clinical-calculation accuracy from 11% to 84%, exceeding unassisted GPT-4o at 36%.Pillai et al.
2025Peer reviewed, ACLSelf-correction decomposes into critique and preservation; interventions can repair wrong answers while increasing regressions on correct ones.Yang et al.
2025Peer reviewed, Findings EMNLPAcross five models and three task families, increasing context caused 13.9–85% degradation despite perfect evidence retrievability.Context Length Alone Hurts
2025Peer reviewed, NeurIPSQuery-only interaction can plant malicious records in agent memory that are retrieved in later work.MINJA
2026Peer reviewed, Findings ACLAcross more than 400,000 instances, many routers performed similarly to simple baselines and remained far below oracle selection; larger portfolios showed diminishing returns.LLMRouterBench
2026Independent reportFrontier progress, compressed leading ranks, benchmark defects, and large residual failures coexist, demonstrating task-sensitive capability gaps.Stanford AI Index 2026
2026-08-04Independent benchmarkQuantization, output constraints, serving configuration, and tool handling can materially change accuracy for the same nominal model.Endpoint Accuracy Index

9. Open research and design questions

Open questionWhat would resolve it
Is the bounded cognitive movement the most stable routing and evaluation unit?Cross-domain comparison of workflow boundaries, privacy constraints, verification contracts, and accepted outcomes.
What model portfolio is sufficient, and what privacy, tool-authorization, and consequence constraints make an endpoint admissible?A policy-gated routing specification tested against real task families and failure cases.
Which escalation thresholds and checks count as independent verification for each task family?Domain-specific calibration against consequence, uncertainty, correlated error, and human-review cost.
What is the minimum viable factorial design, task panel, accepted-outcome rubric, and resource-accounting method?A preregistered benchmark that crosses models, harness layers, budgets, and task distributions.
Which accepted artifacts may enter durable memory, and under what update, deletion, provenance, rollback, and correction-propagation rules?Longitudinal memory experiments spanning contradiction, poisoning, correction, and reuse.
When does durable semantic structure become net semantic capital rather than accumulated semantic debt?Longitudinal measurement of reuse, correction cost, invalidation, model substitution, and accepted-outcome quality.