AIOS Proresearch
AIOS Intelligence System · Research Overview

Full research memo: Local-First and Sovereign AI

1. Domain question

Under what technical, security, performance, and workload conditions can laptops, phones, tablets, and edge devices support person- or organization-controlled AI reasoning, with selective remote escalation, while ordinary files remain canonical and indexes or databases remain contingent infrastructure? The research also asks whether current evidence supports AIOS’s division of labor among bounded model judgment, exact deterministic operations, and human authority, and what would follow if coordinated local context, memory, metadata, workflow, and model routing substantially reduce dependence on remote inference and centralized application custody.

2. Executive answer

Current evidence supports a meaningful but bounded local-AI capability. Quantized small models can run interactively on recent phones, laptops, and edge accelerators, particularly for constrained document, classification, summarization, extraction, and specialized language tasks. Capability cannot be inferred from parameter count or advertised context alone: memory bandwidth, prefill cost, KV-cache growth, sustained thermals, runtime maturity, energy, and model–task fit determine the usable envelope. Larger or more difficult tasks still benefit from frontier models, and the evidence favors selective complementarity rather than an all-local rule.

Local execution materially reduces routine data transfer, but privacy and sovereignty are end-to-end system properties. Local files, logs, memories, tools, backups, synchronization, outputs, software supply chains, and device compromise all remain in the threat model. Confidential-cloud systems can reduce provider access through isolation, attestation, stateless processing, and constrained egress, but they remain dependent on hardware, firmware, attested code, client integrity, and governance. A privacy-respecting hybrid design therefore requires an authorization and disclosure gate before any learned capability router. Only eligible routes should be compared for quality, latency, energy, and cost.

The file-native proposition is defensible as a contingent architecture. Bounded, curated, human-navigable domains can often use ordinary files, documented companion metadata, exact path selection, and lexical search while keeping knowledge portable and auditable. Research does not identify a universal corpus-size threshold or a universally best retrieval architecture. Repeated queries, vocabulary mismatch, heterogeneous documents, multi-hop relationships, latency limits, concurrent writers, atomic invariants, or high update rates can justify lexical, vector, relational, graph, or transactional infrastructure. Except where transaction semantics require otherwise, those structures can remain derived, versioned, rebuildable projections rather than the only authoritative representation of knowledge.

Evidence also establishes that context selection, tool menus, retrieval, validation, and routing affect observed model performance. It does not yet establish the cumulative AIOS claim that many modest system improvements will reliably reduce total inference, retries, and unsupported claims. If that thesis holds, a stable local domain could make smaller models more useful for recurring work, retain durable cognition across provider changes, and invoke frontier models only for movements that warrant them. Determining conditions include workload closure, model capability, provenance quality, memory integrity, routing error, human review burden, device economics, and the full cost of synchronization and support.

3. Essential findings

Finding 1 — Useful on-device AI is real but workload-bounded

Finding 2 — Privacy is a lifecycle property, not a deployment location

Finding 3 — Hybrid routing needs a policy gate before optimization

Finding 4 — File-native knowledge is defensible only as a contingent choice

Finding 5 — Bounded judgment, exact effects, and human authority solve different problems

Finding 6 — Scaffolding changes capability, but cumulative gains remain unproved

Finding 7 — Stable local domains could reduce centralized dependence without replacing the frontier

4. How the evidence refines the AIOS account

  1. An evidence-calibrated local capability envelope. The research replaces device-level generalities with the coupled constraints that determine useful local performance: task fit, quantization, memory, prefill and decode behavior, context, runtime, sustained thermals, and energy.
  2. A two-stage hybrid-routing rule. It sharpens selective frontier use into deterministic disclosure eligibility first and empirical capability optimization second, preventing cost–quality routers from becoming de facto privacy authorities.
  3. A measurable escalation ladder for knowledge infrastructure. It identifies workload triggers for moving from exact files to lexical, semantic, graph, or transactional layers while preserving the distinction between canonical files and accepted ground, on one side, and rebuildable acceleration on the other.
  4. A system-level threat boundary. It extends “local privacy” to memory poisoning, untrusted files, tools, logs, backup, synchronization, update supply chains, and confidential-compute assumptions.
  5. A cumulative-system evaluation requirement. It specifies that AIOS’s many small improvements must be tested for individual, interaction, and total effects on retries, unsupported claims, human correction, tokens, energy, disclosure, and integration—not credited from adjacent component studies.

5. Architectural boundaries preserved

  1. The model–code–human division of responsibility. External evidence exposes why each layer is needed without justifying a collapse of semantic judgment, exact enforcement, and consequential authority into one agent.
  2. Ordinary files and companion metadata as the preferred durable knowledge ground. Retrieval evidence justifies optional derived indexes and a transactional exception without replacing the file-native default across bounded domains.
  3. Distributed intelligence rather than a single cognitive center. No reviewed result requires AIOS to recast intelligence as one model, one database, or one universal agent.
  4. The Why–How–What grammar and distinct cognitive movements. The research neither validates these structures nor supplies evidence for redesigning them; they remain AIOS hypotheses requiring direct comparative tests.
  5. Frontier models as selective complements. Local evidence does not justify an all-local doctrine, and remote-model evidence does not justify moving canonical files, accepted ground, or authority back into provider custody.

6. Where the evidence connects to AIOS

Research contributionAIOS connectionOwning chapterWhy it matters
Intelligence is evaluated as a longitudinal system outcome, not only a model scoreCore architectureFrom Model Intelligence to System IntelligenceStates the central thesis accurately while keeping cumulative performance conditional.
Canonical files can coexist with rebuildable lexical, vector, relational, or graph projectionsCore architectureFiles, Artifacts, Metadata, and Semantic StandingPrevents the file-native claim from being mistaken for a prohibition on indexing and preserves exact descent.
On-device capability is useful but bounded by task, runtime, context, thermals, and energyResearch boundaryEconomic Architecture, Sovereignty, and Decentralized IntelligenceEstablishes a credible boundary while retaining device-level measurements in the full memo.
Disclosure policy precedes learned local/cloud routingMechanism refinementProduct Mechanics: Commands, Files, Records, and ReconstructionSharpens selective frontier use by separating admissibility from optimization.
Local inference does not establish end-to-end privacyResearch boundaryEconomic Architecture, Sovereignty, and Decentralized IntelligenceQualifies claims about private local cognition with lifecycle conditions.
Stable local domains may reduce remote inference and centralized application custodyConditional implicationThe Top-Down Operational Knowledge SystemPreserves a consequential thesis whose aggregate magnitude remains unmeasured.
Regulated use, peer-domain federation, and low-connectivity accessConditional implicationEconomic Architecture, Sovereignty, and Decentralized IntelligenceKeeps identity, provenance, affordability, language quality, and oversight visible as determining conditions.
Confidential-compute architectures and specific hardware attacksSource-level depthOpen Research Questions and the AIOS Experimental ProgramBounds governed frontier fallback without overloading the core architecture with implementation detail.
Compression algorithms, NPU measurements, cryptographic inference, and retrieval benchmark variantsSource-level depthCapability Horizons, Experiments, Proof, and FalsificationPreserves exact support and counterevidence through descent.
Any universal local-coverage or infrastructure-reduction magnitudeOpen testOpen Research Questions and the AIOS Experimental ProgramNo population-level or reference-deployment evidence yet supports a universal magnitude.

7. Relationships across research programs

  1. Benchmarking and dynamic routing. Local–frontier routing connects system-level and cognitive-movement-level measures of quality, retries, disclosure, energy, latency, and cost to task taxonomy, baselines, and routing objectives.
  2. Context, memory, and provenance. Local capability depends on composed context and multiple memory layers, while persistent context extends the lifetime of poisoning and stale inference; source descent, invalidation, and contamination controls therefore belong to the same system boundary.
  3. Cognitive modes and agent architecture. Bounded menus and phase-specific contexts may make smaller models more effective, but equal-budget comparisons must separate cognitive specialization from additional samples or compute.
  4. Metadata, relationships, and retrieval. The file-native account admits derived indexes when workloads require them; provenance and rebuild guarantees connect the metadata-emergent relationship graph to lexical, semantic, graph, and transactional needs.
  5. Human authority, regulated use, and peer collaboration. Local control creates opportunities for data minimization and federated domains without settling authorization, review, identity, signatures, revocation, conflict resolution, legal accountability, accessibility, or language equity.

8. Priority source set

9. Open research and design questions

  1. Which task distribution and interaction effects would establish system-level intelligence as a cumulative outcome rather than a restatement of model capability?
  2. Which measured workload triggers justify escalation from exact files to lexical, semantic, graph, or transactional infrastructure?
  3. Which task population, stakes, hardware and language strata, success threshold, and frontier-escalation policy define a defensible claim of broad local coverage?
  4. How should ordinary cloud, confidential cloud, and cryptographic private inference be compared without implying that any route is trust-free?
  5. Under what conditions do personal and organizational intelligence, regulated domains, peer federation, and global accessibility follow from local-first architecture?