AIOS Proresearch
AIOS Intelligence System · Research Overview

Part II · Context, Memory, Files, and Composition

Files, Artifacts, Metadata, and Semantic Standing

Files are the durable home of the work, while nearby records explain what each file is, how it relates to other work, and what status it has. Search indexes, maps, and visual views remain disposable ways into that ground.

Canonical artifactCompanion metadataSemantic standingFile-native knowledge graph

1. Files are the durable body

The file-native claim is stronger than “the system saves outputs as files.”

In AIOS, files provide:

The model is temporary. The files preserve the work.

This does not make a file—or a database—“the intelligence.” Canonical files give the distributed system owner-controlled ground to which people and different models can return. Metadata, indexes, maps, and retrieval systems help compose that ground for a bounded movement. Exact operations determine what is stored durably; the appropriate person or institutional authority determines what receives accepted standing.

2. Canonical, companion, record, and projection classes

Canonical ground and rebuildable projections

How durable canonical classes support projections that can be deleted and rebuilt without becoming authority.

flowchart TB subgraph C["Canonical content"] S["Exact sources"] A["Canonical artifacts"] W["Workflow and expertise definitions"] end subgraph M["Canonical companion metadata records"] H["Document Home"] P["Project / Session Home"] D["Decisions · annotations · relations"] end subgraph R["Durable operation evidence"] O["Assignments · returns · receipts"] T["Traces · syntheses · completion"] end subgraph X["Derived projections"] DM["Document Map"] IX["Indexes and search facets"] AT["Atlas and visualizations"] WC["Working contexts"] end C --> M C --> R M --> X R --> X X -. "deleted and rebuilt" .-> C X -. "deleted and rebuilt" .-> M X -. "deleted and rebuilt" .-> R
Derived access structure · canonical09-files-artifacts-metadata-and-semantic-standing--m01.mmd
Text equivalent

Canonical content → Canonical companion metadata records; Canonical content → Durable operation evidence; Canonical companion metadata records → Derived projections; Durable operation evidence → Derived projections.

The asymmetry is the point: canonical content, companion records, and operation evidence feed the projection layer. The dotted return is reconstruction, not promotion—maps, indexes, Atlas views, and working contexts may disappear while their owner ground remains.

The architecture depends on preserving these differences.

3. Core file classes

File classStandingIntelligence function
Thinking MarkdownPrimary exploratory workLets ideas form before they are forced into document structure
Canonical document MarkdownPrimary artifactLanguage the person is creating and delivering
Document Home / companion YAMLCanonical companion metadata recordIdentity, accepted purpose and spine, state, annotations, lineage, relations, completion
Document Map MarkdownDisposable projectionReadable structure and movement contributions
Project or Session Home MarkdownCanonical navigational bearingPurpose, plan, outputs, documents, relationships
Workflow YAML and stage MarkdownReusable method groundIdentity, role, stages, production guidance, quality, completion behavior
Expertise resourcesAccepted reusable knowledgePrinciples, roles, workflows, templates, exemplars, conventions
Materialized context MarkdownRewritable capsule sourceStable reading notes plus fresh bounded context
Decision and relationship recordsDurable semantic stateRationale, authority, dependency, challenge, supersession
Assignment and return recordsDurable delegated-work stateCharge, grant, progress, findings, limitations, delta
Operation and lifecycle receiptsDurable evidenceAuthorization, source state, exact effect, result, failure, continuation
Registry and Atlas outputsDerived infrastructureResolved identity, validation, impact, inspection, visualization

Thinking can be retained physically without becoming future-governing ground. An exploratory trace may remain available as inert history or evidence, while branching, composing, planning, or accepted integration gives selected material the intentional standing required for deliberate reuse. Physical persistence and semantic participation are different properties of the same system.

Admission of existing artifacts and sources

Existing work enters the managed domain through admission, not ingestion. Artifact or source admission is the explicit act by which an external or previously unmanaged file receives a stable local identity, an object class, provenance, and a declared custody or trust state. It makes the object addressable inside the knowledge system without silently changing what the object means.

external or unmanaged file
  -> admission preview
    -> person-authorized custody operation
      -> verified canonical identity + companion metadata + provenance
        -> available for relevance-based selection

Admission does not:

These remain separate judgments and state transitions. A person's existing draft may become a canonical person-owned artifact while some of its claims remain unreviewed. An imported paper may become a locally held source while retaining external provenance and a qualified trust boundary. Both can receive ordinary work affordances—open, read, search, annotate, branch, and compose from—without becoming epistemically equivalent.

Identity, object class, provenance, custody, epistemic standing, intentional standing, availability, and activation therefore remain independent dimensions. Admission creates addressable local custody; later context composition determines whether the admitted object is relevant to a particular movement.

Product Mechanics defines admission as a closed, previewed, person-authorized custody operation. Reference Journey K tests the complete transition through restart and later relevance selection.

4. The canonical artifact remains primary

The document is the artifact the person is making. Metadata exists to help understand and operate on it, not to replace it.

The system may:

It should not contaminate the artifact with every machine-state field required by the runtime.

5. The Document Home

The Document Home is the adjacent companion metadata record for one managed document. Historical sources may call the same class a shadow document or task metadata document. These are lineage names, not separate systems.

A target Document Home can hold:

The compact Why–How–What account inside it is the Document Spine. The spine supports navigation and context. It is not a template imposed on visible prose.

6. Identity, title, path, and revision are distinct

These concepts must not collapse:

A rename should not create a new document. A move should not erase lineage. A stale revision should not receive an exact edit merely because its title still matches.

7. The Document Map

The Document Map is a disposable, natural-language projection compiled from canonical document and companion ground.

It can show:

It is more than a table of contents. It explains how the argument or artifact moves.

Because it is derived, it can be rebuilt when stale and must not outrank the artifact or Document Home.

8. The Project or Session Home

The Project Home is the project at navigational grain.

Its core zones are:

It is a compact launchpad, not a log of everything that happened. From it, the system can descend to a document's companion metadata record, maps, movements, exact passages, decisions, and evidence.

9. Metadata has four jobs

Identity

Resolve stable objects across title, path, revision, and interface changes.

Semantic standing

Record whether material is source, interpretation, proposal, accepted, contested, stale, or superseded.

Relationship and lineage

Show how artifacts, claims, decisions, operations, and sources bear on one another.

Operational reconstruction

Recover what was authorized, in flight, applied, returned, or awaiting attention.

Metadata should not decide semantic truth merely because a field exists.

Provenance is especially important and especially easy to overstate. It records lineage, identity, revisions, and trust boundaries so a claim can be attributed, rechecked, corrected, or recovered. It does not make the claim true.

Annotations preserve situated judgment

An annotation is an anchored, provenance-bearing judgment about how an artifact, source, passage, or cognitive movement should be understood later. It can preserve an interpretation, qualification, critique, relationship, contradiction, or proposed supersession without rewriting the underlying artifact or its accepted Document Spine.

A durable annotation carries a stable identity, an exact target anchor and source revision, its author or originating model and principal, time, provenance, scope, and semantic standing. Companion metadata can surface that compact account during recognition and relational recall before the system opens the complete annotated material. The annotation may then become candidate ground for a fitting context, with exact descent to both its target and its own source.

An annotation does not change the target's standing, grant authority, or trigger an effect merely because it exists. Those consequences require their own review, authorization, verified operation, and—where acceptance is intended—settlement and reintegration.

10. Structured fields and natural language cooperate

Use structured fields for:

Use natural language for:

This lets mechanisms resolve what must be exact without asking schemas to carry the full meaning of the work.

11. Records and receipts

Consequential seams deserve durable records:

Not every streaming fragment or interface refresh needs a receipt. Records exist where authority, identity, source standing, result, or continuation must be reconstructed.

12. Bounded file operations

The model can declare meaning; the host applies exact effects.

Typical operation classes include:

Each effect should identify:

13. File-bound authority

Different semantic states imply different effect rights.

InputStandingNormal consequence
“Add this to the plan.”Direct instructionMay authorize bounded plan change
“This may belong in the plan.”Observed possibilityCandidate only
Repeated themeForming patternRecord as candidate; do not install as principle
Delegated research returnEvidence-shaped returnIntegrate as evidence or proposal, not automatic decision
Approved reusable methodPromoted expertiseAvailable to context composition, not mandatory
Stale mapDerived projectionRebuild; never overwrite canonical artifacts or accepted ground from it

The architecture prevents semantic implication from silently becoming mutation authority.

14. Relationships create a file-native knowledge graph

Relationships as a file-native graph

How explicit relations turn separate files into a traversable field of evidence, decisions, production, and learning.

flowchart LR S["Source"] -->|"supports"| C["Claim"] C -->|"informs"| D["Decision"] D -->|"constrains"| P["Plan"] P -->|"commissions"| A["Artifact"] A -->|"reveals tension in"| C A -->|"produced by"| W["Workflow run"] W -->|"provides evidence for"| K["Candidate practice"]
File-native knowledge graph · canonical09-files-artifacts-metadata-and-semantic-standing--m02.mmd

The path is not merely forward. An artifact can reveal tension in the claim that helped commission it, while its workflow run can become evidence for a candidate practice. That makes consequence and correction traversable without replacing the files with a graph database.

The graph emerges from explicit records and exact handles. Search, source descent, and derived projections recover it. A database or vector store may assist retrieval when the workload requires it but does not become canonical authority.

Relationship traversal is one member of a plural retrieval system. Exact paths answer identity-sensitive questions; lexical search finds known language; metadata filters use standing and scope; vector indexes address vocabulary mismatch; relationship graphs support explicit multi-hop bearings; rerankers reorder candidates; and direct reading preserves original evidence and order. The operation and workload determine the useful combination.

Escalation should respond to measured conditions such as query ambiguity, vocabulary mismatch, document heterogeneity, multi-hop depth, latency, update rate, and concurrency—not to an arbitrary number of files. Every derived candidate should preserve a route to the exact owner-controlled source.

15. Reconstruction after restart

A robust file-native system should be able to rebuild:

If this state exists only in a running process, the system is not durably file-native.

16. Files and provider independence

Because prompts, artifacts, memory, and records remain outside any one model provider:

Provider flexibility still requires adapter and capability proof. It should not be overstated as universal interchangeability.

17. Derived views remain removable; transactions are a narrow exception

The registry, Atlas, maps, search facets, and working contexts are valuable because they make the system legible. They remain safe only if they can be deleted and rebuilt without loss of canonical artifacts, sources, or accepted standing.

The canonicality test is:

If this projection disappears, can the system reconstruct it from exact owner files and accepted records?

If not, the projection has become an undeclared canonical source.

This test applies to lexical indexes, embeddings, extracted facts, relationship graphs, reranked result sets, summaries, and visual maps. They can be versioned, evaluated, discarded, and rebuilt from canonical files and accepted records.

Transactional state is the narrow exception. When concurrent writers, high update rates, or rules requiring several related changes to succeed or fail together make a plain-file update unsafe, a transactional system may be operationally canonical for that exact state. The exception should remain bounded: it does not justify moving documents, purposes, evidence, and semantic relationships into an opaque application store by default. Transactional records need stable links to owner files, visible effects and receipts, recovery paths, and exportable state.

18. Failure modes

Metadata monoculture

One giant schema attempts to represent all meaning, workflows, projects, and artifacts.

Parallel companions

Different features create competing metadata files for the same document.

Projection authority

A map, index, or Atlas silently outranks its sources.

Transactional creep

A database introduced for one atomic invariant gradually becomes the hidden home of semantic knowledge and authority.

Title identity

Renaming or moving a file breaks history and references.

Passive semantic initiation

A file watcher interprets content changes and starts new intelligence work without an admitted action.

Hidden state

Important progress or decisions exist only in memory of a running process.

Receipt excess

The system records every trivial event, producing more operational noise than reconstructive value.

19. Research evaluation

Evaluate whether the file architecture can:

  1. reconstruct current state after restart;
  2. preserve stable identity across rename and move;
  3. prevent stale exact edits;
  4. distinguish canonical source from derived projection;
  5. preserve proposal, acceptance, and execution as different states;
  6. recover source and decision lineage;
  7. isolate malformed or conflicting records locally;
  8. rebuild indexes and maps deterministically;
  9. route among exact, lexical, semantic, and relational retrieval without losing source descent;
  10. confine transactional authority to declared concurrent or atomic state;
  11. support multi-model evaluation without loss of work;
  12. remain understandable to the person without exposing all machinery.

20. Research connection: plural access over canonical files

AIOS begins with ordinary files because the product must preserve inspectable, portable, owner-controlled knowledge while models and interfaces change. Research sharpens the access architecture around that commitment. BRIGHT shows that lexical and dense retrievers can both struggle when relevance itself requires reasoning. HybGRAG finds value in combining textual and relational retrieval for questions that genuinely require both. Laitenberger, Manning, and Liu show that preserving original document structure and order can match or outperform more elaborate multistage approaches under matched budgets in their long-document evaluations.

Together, these findings support plural retrieval and exact source descent rather than one universal index. They do not establish that files alone are sufficient for every workload, that relationships are always useful, or that any derived projection is complete. AIOS retains the file-native default while allowing rebuildable access structures and the declared transactional exception.

Read deeper in the normalized relational-memory brief, full relational-memory memo, normalized local-sovereign AI brief, and full local-sovereign AI memo. Chapter 10 shows how the same ground becomes navigable across semantic resolutions.