Research library · Normalized brief
System Capability and Dynamic Model Routing — Research Brief
AIOS connections: From Model Intelligence to System Intelligence; Capability Horizons, Experiments, Proof, and Falsification; Economic Architecture, Sovereignty, and Decentralized Intelligence; Research Positioning, Novelty, and System-Level Contribution
Full research memo: System-Level Capability Amplification and Dynamic Model Routing
1. Domain question
To what extent is useful AI performance a property of the complete reasoning system rather than the reasoning model alone, and can AIOS exploit that difference by routing bounded cognitive movements among frontier, balanced, small, and on-device models? The research tests whether purpose, context, decomposition, working state, memory, relational metadata, provenance, tools, verification, and integration can reduce inference burden and failure risk; when their gains interact or backfire; and how cumulative system effects, selective frontier use, and reusable semantic structure should be evaluated without mistaking harness intelligence for model intelligence.
2. Executive answer
The evidence supports a bounded but consequential answer: useful performance is jointly produced by the model, harness, task distribution, environment, and resource budget. Model capability remains important, especially for novel, ambiguous, open-ended, or high-consequence reasoning. Yet same-model ablations and within-study comparisons show that task framing, interfaces, exact tools, context policy, workflow structure, and verification can materially change outcomes. In bounded settings with a good external signal, a smaller model plus a specialized tool or optimized workflow can match or exceed a larger model operating without that support.
This is not a universal scale-substitution result. The strongest demonstrations occur where a weak operation can be externalized and checked: calculation, code execution, repository repair, retrieval, structured output, or verifier-guided mathematical search. Scaffolding cannot reliably supply missing world knowledge, invent a valid acceptance function, or resolve genuinely ambiguous judgment. Longer reasoning is not automatically better, and additional context, memory, critique, agents, or workflow stages can introduce interference, correlated error, security exposure, and coordination cost.
Current model gaps are also benchmark-sensitive. Rankings compress on some agent and mathematical evaluations, widen sharply on specialist or broad-knowledge tests, and can reverse for particular tasks. The same nominal model can perform differently across serving endpoints. AIOS should therefore route a bounded cognitive movement—not a whole project—using task type, evidence sufficiency, sensitivity, consequence, tool state, calibrated performance history, and verification results. The router should first enforce privacy and authority constraints, then select the least costly admissible model-harness configuration expected to meet the quality threshold, with explicit escalation to frontier capability when evidence, calibration, or checks are insufficient.
The ambitious implication is that accepted, versioned claims, relations, tests, decisions, and correction dependencies could become reusable semantic capital, shifting repeated work from fresh inference into inspectable structure. That implication is plausible but not yet directly established. Long-term memory studies show both retrieval benefits and large losses from staleness, compression, poisoning, and privacy extraction.
The appropriate conclusion is therefore architectural and experimental. AIOS has a credible basis for selective frontier use within a local, model-plural system, but should treat every amplification layer as a factor to test rather than a presumed improvement. The thesis should be accepted only if factorial ablations, distribution-shifted routing trials, and longitudinal reuse studies improve the accepted-outcome frontier after accounting for tokens, latency, energy, retries, human review, privacy, maintenance, correction propagation, and failure containment.
3. Essential findings
Finding 1 — Observed capability belongs to a model–harness system
- Finding: Model scores do not isolate intrinsic intelligence because task mixture, endpoint configuration, context, interface, tools, and evaluation protocol materially affect observed performance.
- Evidence: Stanford’s 2026 synthesis finds rapid frontier gains alongside jagged residual failures and benchmark defects; endpoint testing shows materially different accuracy for the same nominal model under different quantization, output, and tool configurations.
- Relationship to AIOS: Direct mechanism: AIOS already organizes reasoning through context, tools, state, memory, and workflow, so its correct evaluation unit is the complete movement configuration.
- Implication: Publish separate model-effect, harness-lift, and system-frontier results instead of one undifferentiated capability score.
- Limits or counterevidence: Stronger models still tend to dominate difficult broad evaluations; harness sensitivity does not make model weights interchangeable.
- Exact descent: Full memo headings “1.1 There is no single ‘model gap’” and “5.1 System-level benchmark design”. Sources: Stanford AI Index 2026, Artificial Analysis Endpoint Accuracy Index.
Finding 2 — Scaffolding can substitute for scale on bounded, verifiable operations
- Finding: A smaller model can outperform a larger weakly scaffolded model when the fragile operation is externalized through an exact tool, constrained interface, optimized workflow, or reliable verifier.
- Evidence: In 10,000 clinical-calculation trials, Llama 3.1 70B rose from 11% unassisted accuracy to 84% with OpenMedCalc, exceeding unassisted GPT-4o at 36%; adaptive test-time compute also let a smaller mathematical model exceed a roughly 14-times-larger model in solvable regions.
- Relationship to AIOS: Direct mechanism: route calculation, search, validation, and structured transformation through purpose-built movements and reserve general reasoning for interpretation.
- Implication: Compare a targeted system intervention with a model-tier upgrade before defaulting every operation to the frontier model.
- Limits or counterevidence: These results are strongest in mathematics, code, and tool-bounded work; even the best clinical condition retained interpretation errors.
- Exact descent: Full memo headings “1.2 Scaffolding can substitute for scale when the bottleneck is externalizable” and “3.2 What present evidence actually supports”. Sources: clinical-calculation study, Snell et al., AFlow.
Finding 3 — Context quality and placement matter more than maximum context volume
- Finding: Retrieval, long context, and memory help only when they increase usable evidence density without overwhelming the model with irrelevant, stale, or badly positioned material.
- Evidence: LongMemEval reports 30–60% declines from oracle-evidence conditions over long histories; a controlled Findings of EMNLP study reports 13.9–85% degradation as context length increases even when the relevant evidence remains perfectly retrievable.
- Relationship to AIOS: Direct mechanism and boundary condition: context composition and memory retrieval should operate under an evidence budget with freshness, placement, and sufficiency checks.
- Implication: Store broadly if justified, but generate from a deliberately constructed working set; preserve fallbacks to source text and simpler retrieval.
- Limits or counterevidence: Long context can outperform RAG when fully resourced, while graph retrieval can improve associative reasoning; no single context policy dominates every task.
- Exact descent: Full memo headings “1.3 Context quality is not context quantity” and “3.5 Durable structure as semantic capital”. Sources: LongMemEval, Context Length Alone Hurts, RAG versus long context, HippoRAG 2.
Finding 4 — Provenance improves auditability but does not establish truth
- Finding: Claim-source structure can improve attribution and review while leaving factuality, source authority, and evidence completeness unresolved.
- Evidence: ReClaim reports roughly 90% citation accuracy with interleaved references and claims, while GaRAGe finds leading systems still weak on relevance-aware factuality, deflection, and attribution across 2,366 questions and more than 35,000 annotated passages.
- Relationship to AIOS: Supporting and complicating evidence: relational metadata can make claims inspectable and corrections traceable, but the generator must not treat metadata as proof.
- Implication: Evaluate factuality, faithfulness, attribution, source quality, and freshness separately; expose what was actually checked.
- Limits or counterevidence: Authorship labels alone can change attribution quality by 3–18%, showing that provenance metadata can create authority bias.
- Exact descent: Full memo headings “1.4 Provenance improves auditability, not truth by itself” and “4.3 Hallucination control becomes failure-mode-specific”. Sources: ReClaim, GaRAGe, authorship-metadata study.
Finding 5 — Amplification layers interact and can be harmful
- Finding: Context, memory, critique, decomposition, and multiple agents do not contribute fixed additive gains; they can be synergistic, redundant, bottleneck-shifting, or antagonistic.
- Evidence: ACL 2025 separates critique from preservation and shows that interventions improving error repair can damage initially correct answers; a 180-configuration multi-agent preprint reports large gains on parallel work but 39–70% degradation on sequential tasks.
- Relationship to AIOS: Boundary condition: each established layer should remain available as an architectural capability, but its activation policy should be task-conditional and empirically ablated.
- Implication: Estimate pairwise interactions, coordination cost, and regression rates rather than validating layers only through cumulative demonstrations.
- Limits or counterevidence: The broad multi-agent scaling study is a preprint, and interaction magnitudes may change with models and tasks; positive cases for centralized coordination remain substantial.
- Exact descent: Full memo headings “1.5 Critique, verification, and multiple agents have strong interaction effects” and “3.3 Gains are not generally additive”. Sources: Yang et al., Scaling Agent Systems, corrective in-context learning.
Finding 6 — Dynamic routing is feasible but remains fragile under change
- Finding: Model routers can lower inference cost at matched benchmark quality, but their advantage depends on calibration, portfolio composition, evaluator quality, and distribution stability.
- Evidence: RouteLLM reports more than a twofold cost reduction at matched measured quality; BEST-Route reports up to 60% savings with less than 1% performance loss; LLMRouterBench finds many sophisticated routers statistically similar to simple baselines and far below an oracle across more than 400,000 instances.
- Relationship to AIOS: Direct mechanism: select a model-harness-budget combination for each bounded movement after privacy and authority filtering, then escalate on failed checks or poor coverage.
- Implication: Begin with a small complementary portfolio and a transparent baseline router; justify learned complexity prospectively under endpoint and domain shift.
- Limits or counterevidence: Preference and benchmark labels are imperfect proxies for accepted outcomes, while self-confidence can be badly calibrated outside the training distribution.
- Exact descent: Full memo headings “1.7 Dynamic model routing works, but robust routing remains unsolved” and “3.4 Dynamic model selection as constrained control”. Sources: RouteLLM, BEST-Route, LLMRouterBench.
Finding 7 — Tools and memory amplify capability and failure persistence together
- Finding: Tool access and durable memory improve action and continuity while increasing prompt-injection, unauthorized action, privacy-extraction, poisoning, and rollback risks.
- Evidence: AgentDojo evaluates 97 tasks and 629 injection cases and finds both useful task completion and defense difficult; MINJA plants records through ordinary queries for later retrieval; MEXTRA extracts stored private interactions in black-box settings.
- Relationship to AIOS: Complicating evidence and direct mechanism: local governance, typed actions, scoped memory, provenance, previews, postconditions, deletion, and rollback are part of capability—not optional security wrappers.
- Implication: Measure privacy leakage, unauthorized state changes, reversibility, and blast radius alongside task success.
- Limits or counterevidence: These benchmarks use bounded threat models and representative agents, so attack rates do not directly predict a specific deployment; they establish credible failure classes.
- Exact descent: Full memo headings “1.6 Tool use expands both capability and attack surface” and “4.2 Privacy becomes a routing dimension rather than a universal model property”. Sources: AgentDojo, MINJA, MEXTRA.
Finding 8 — Semantic capital is a plausible longitudinal hypothesis, not a demonstrated result
- Finding: Accepted, versioned claims, relations, tests, decisions, and correction dependencies may reduce future inference and review, but persistence creates capital only when reuse value exceeds maintenance and failure costs.
- Evidence: LongMemEval and HippoRAG 2 show gains from structured indexing and relational retrieval alongside losses from compression and retrieval tradeoffs; no reviewed study directly establishes a declining marginal cost of heterogeneous knowledge work over time.
- Relationship to AIOS: Supporting evidence plus boundary condition: established post-response integration provides the organizing ground, while the research specifies what would count as net semantic capital.
- Implication: Run a 12–24-week comparison of no reuse, raw transcript reuse, and structured accepted-artifact reuse with controlled updates, deletions, poisoning, and store-removal counterfactuals.
- Limits or counterevidence: A growing store can become semantic debt through staleness, contradiction, leakage, and correction burden.
- Exact descent: Full memo headings “3.5 Durable structure as semantic capital”, “4.4 Post-response integration can change the economics of inference”, and “5.6 Longitudinal semantic-capital experiment”. Sources: LongMemEval, HippoRAG 2, MINJA.
4. How the evidence refines the AIOS account
- It supplies a bounded empirical formulation of system capability: model, harness, environment, task distribution, and budget interact rather than contribute independent fixed gains.
- It defines the bounded cognitive movement as the unit for model selection, verification, privacy control, and outcome measurement.
- It turns model plurality into a constrained routing problem with admissibility, capability allocation, and evidence-triggered escalation.
- It distinguishes factuality, faithfulness, attribution, and source quality, clarifying which provenance mechanisms support audit rather than truth itself.
- It provides a falsifiable program—crossed model/harness tests, fractional-factorial ablations, shifted routing trials, and longitudinal reuse—to test cumulative amplification and semantic capital.
5. Architectural boundaries preserved
- Routing remains subordinate to the established AIOS architecture of purpose, context, memory, tools, and integration rather than becoming the center of the account.
- AIOS remains complementary to frontier models, invoking frontier capability selectively where it changes the accepted outcome.
- Architectural availability remains distinct from task-level activation; not every layer is active for every task.
- Local execution remains distinct from complete privacy, just as retrieved citations remain distinct from factual correctness.
- Benchmark cost savings remain bounded results until accepted outcomes, review, maintenance, security, and model churn are measured.
6. Where the evidence connects to AIOS
| Research contribution | AIOS connection | Owning chapter | Why it matters |
|---|---|---|---|
| Capability is a joint model–harness property | Core architecture | From Model Intelligence to System Intelligence | Explains why AIOS layers can affect outcomes without diminishing model intelligence. |
| Bounded smaller-model substitution | Research boundary | Capability Horizons, Experiments, Proof, and Falsification | Evidence is strong in verifiable domains but too conditional for an unrestricted claim. |
| Cognitive movement as the routing unit | Core architecture | Cognitive Movements, Roles, and Carriers | Connects decomposition to model choice, privacy, verification, and cost control. |
| Constrained router with selective frontier escalation | Core architecture | From Model Intelligence to System Intelligence | Preserves frontier use while making privacy, capability, and consequence explicit. |
| Non-additive layer interactions | Experimental boundary | Capability Horizons, Experiments, Proof, and Falsification | Prevents cumulative demonstrations without ablation. |
| Provenance as auditability rather than truth | Core distinction | Files, Artifacts, Metadata, and Semantic Standing | Prevents citation and grounding mechanisms from being overstated. |
| Tools and memory as capability–risk multipliers | Conditional implication | Product Mechanics: Commands, Files, Records, and Reconstruction | Connects amplification to persistent and state-changing failure modes. |
| Semantic capital hypothesis | Conditional implication | Economic Architecture, Sovereignty, and Decentralized Intelligence | Preserves the economic and epistemic implication while keeping it explicitly testable. |
| Factorial system benchmark | Experimental design | Capability Horizons, Experiments, Proof, and Falsification | Distinguishes model effects, harness lift, interactions, and total system frontiers. |
| Exact current benchmark rankings | Source-level depth | Research Positioning, Novelty, and System-Level Contribution | Fast-changing values support rather than anchor the durable account. |
| Broad autonomous workflow search | Open research direction | Open Research Questions and the AIOS Experimental Program | Search cost, overfitting, and transfer require prospective evidence before architectural integration. |
7. Relationships across research programs
- Context composition × provenance: retrieval determines what evidence enters working context; provenance determines whether its use can be audited, refreshed, or reversed.
- Memory × security: durable state creates continuity and reuse, but also extends the lifetime of poisoning, privacy leakage, and obsolete instructions.
- Routing × privacy: model selection begins with endpoint admissibility and data minimization, not a pure accuracy-versus-cost ranking.
- Tools × verification: exact tools reduce computational errors only when argument selection, state changes, and postconditions are independently checked.
- Integration × learning economics: accepted outputs become useful reusable structure only if correction dependencies and invalidation propagate across later work.
8. Priority source set
| Date | Source type | Exact supportable claim | Direct link |
|---|---|---|---|
| 2024-09 | Peer reviewed, EMNLP Industry | Fully resourced long context outperformed tested RAG configurations on public QA, while self-routing retained comparable performance with lower compute. | Li et al. |
| 2024-12 | Peer reviewed, NeurIPS | A purpose-built agent-computer interface materially improved repository-repair performance, showing that action and observation design affect measured capability. | SWE-agent |
| 2024-12 | Peer reviewed, NeurIPS benchmark | Across 97 tasks and 629 attacks, tool-agent task completion and prompt-injection defense were both difficult. | AgentDojo |
| 2025 | Peer reviewed, ICLR | Adaptive test-time compute was reported as more than four times as efficient as best-of-N and enabled bounded smaller-model scale substitution in mathematical reasoning. | Snell et al. |
| 2025 | Peer reviewed, ICLR | Long-history performance fell 30–60% relative to oracle evidence, while targeted indexing helped and fact condensation created ability tradeoffs. | LongMemEval |
| 2025 | Peer reviewed, ICLR | Preference-trained routing reduced cost by more than twofold at matched measured quality and transferred across tested model pairs. | RouteLLM |
| 2025 | Peer reviewed, ICML | Workflow search improved six-benchmark performance by 5.7% on average and produced selected cases where smaller models exceeded GPT-4o at much lower reported inference cost. | AFlow |
| 2025 | Peer reviewed, npj Digital Medicine | A specialized calculator raised Llama 3.1 70B clinical-calculation accuracy from 11% to 84%, exceeding unassisted GPT-4o at 36%. | Pillai et al. |
| 2025 | Peer reviewed, ACL | Self-correction decomposes into critique and preservation; interventions can repair wrong answers while increasing regressions on correct ones. | Yang et al. |
| 2025 | Peer reviewed, Findings EMNLP | Across five models and three task families, increasing context caused 13.9–85% degradation despite perfect evidence retrievability. | Context Length Alone Hurts |
| 2025 | Peer reviewed, NeurIPS | Query-only interaction can plant malicious records in agent memory that are retrieved in later work. | MINJA |
| 2026 | Peer reviewed, Findings ACL | Across more than 400,000 instances, many routers performed similarly to simple baselines and remained far below oracle selection; larger portfolios showed diminishing returns. | LLMRouterBench |
| 2026 | Independent report | Frontier progress, compressed leading ranks, benchmark defects, and large residual failures coexist, demonstrating task-sensitive capability gaps. | Stanford AI Index 2026 |
| 2026-08-04 | Independent benchmark | Quantization, output constraints, serving configuration, and tool handling can materially change accuracy for the same nominal model. | Endpoint Accuracy Index |
9. Open research and design questions
| Open question | What would resolve it |
|---|---|
| Is the bounded cognitive movement the most stable routing and evaluation unit? | Cross-domain comparison of workflow boundaries, privacy constraints, verification contracts, and accepted outcomes. |
| What model portfolio is sufficient, and what privacy, tool-authorization, and consequence constraints make an endpoint admissible? | A policy-gated routing specification tested against real task families and failure cases. |
| Which escalation thresholds and checks count as independent verification for each task family? | Domain-specific calibration against consequence, uncertainty, correlated error, and human-review cost. |
| What is the minimum viable factorial design, task panel, accepted-outcome rubric, and resource-accounting method? | A preregistered benchmark that crosses models, harness layers, budgets, and task distributions. |
| Which accepted artifacts may enter durable memory, and under what update, deletion, provenance, rollback, and correction-propagation rules? | Longitudinal memory experiments spanning contradiction, poisoning, correction, and reuse. |
| When does durable semantic structure become net semantic capital rather than accumulated semantic debt? | Longitudinal measurement of reuse, correction cost, invalidation, model substitution, and accepted-outcome quality. |