Part V · Knowledge Evolution and Product Mechanics
Self-Correction, Contradiction, and Epistemic Maintenance
Temporal state transitionAuthorized departureRevalidationAccepted ground
1. Self-correction is a continuity problem
A system does not become self-correcting merely because a model can revise an answer. Long-horizon correction requires the system to preserve enough structure to answer seven questions:
- What was previously believed or decided?
- Why was it believed?
- Which evidence, assumptions, and scope supported it?
- What new observation or argument challenges it?
- How far does the correction propagate?
- Who has authority to accept the changed ground?
- Does later reasoning actually use the correction?
This chapter develops the corrective half of the knowledge-evolution loop. Registration and promotion explain how experience can acquire reusable standing. The Flow Atlas makes lineage and affected relationships visible. Epistemic maintenance determines how accepted ground is challenged, narrowed, superseded, retired, and demonstrably changed in later use.
The governing principle is: correction changes the future without falsifying the past. Historical records should continue to show what was believed or decided, by whom, and from which evidence. Their present authority can change.
2. The epistemic object is larger than a claim
An important conclusion should be recoverable as a structured relationship among:
- claim or decision;
- source observations;
- interpretation;
- assumptions;
- scope and boundary conditions;
- supporting evidence;
- counterevidence;
- alternatives considered;
- confidence and its basis;
- unknowns and contradictions;
- author, reviewer, and acceptance standing;
- downstream consumers;
- revision and supersession lineage.
This does not require every sentence to carry a heavy schema. Consequential claims need deeper memory than ordinary prose, and the depth should be proportional to risk and future reuse.
3. The correction journey
Correction crosses the same standing boundaries as promotion:
observed contradiction or new evidence
→ proposed correction
→ accepted revised knowledge
→ canonical supersession, qualification, or retirement
→ invalidation and recompilation of affected contexts and projections
→ behavioral proof in a later movement
The evidence and the proposed correction remain distinct. A model can identify conflict and draft a revision; it cannot make that revision governing merely by writing persuasively. An acceptance decision and authorization to canonicalize follow the authority of the affected scope; accepted standing settles only after the exact file effect is verified and reconstructable.
Correction proved by changed later behavior
A credible contradiction triggers source recovery, scoped correction, authorized supersession, recompilation, and a later behavioral test.
Text equivalent
Accepted understanding → Interruption or elapsed time; Interruption or elapsed time → New evidence or credible contradiction; New evidence or credible contradiction → Recover original sources and rationale; Recover original sources and rationale → Challenge assumptions, scope, and alternatives; Challenge assumptions, scope, and alternatives → Determine correction and affected consumers; Determine correction and affected consumers → Acceptance required?; Person or designated authority reviews → Record correction; Record correction → Supersede obsolete scope without erasure; Supersede obsolete scope without erasure → Update canonical artifacts; invalidate and recompile affected views; Update canonical artifacts; invalidate and recompile affected views → Test a later invocation; Test a later invocation → Behavior changed appropriately?; Trace missed consumer or failed activation → Determine correction and affected consumers.
The loop at the end is decisive: if a later invocation still follows obsolete ground, the system traces a missed consumer or failed activation and returns to impact analysis. Correction succeeds only when changed canonical ground produces appropriately changed behavior.
The final test is behavioral, not clerical. A status update, revision record, or green validation result is insufficient if the obsolete conclusion still shapes future work.
A correction that proceeds under delegated authority still requires a recoverable grant naming the scope, actor, permitted operation, and revocation path. “No new acceptance required” does not mean “no authority record.”
4. Correction is not overwrite
The system distinguishes:
| Operation | Meaning |
|---|---|
| edit | Change a working expression that has not acquired durable historical standing |
| revise | Produce a new version while retaining prior versions and their relationship |
| narrow | Preserve a claim but reduce its valid scope |
| qualify | Add conditions, uncertainty, or exceptions |
| retract | Remove present support while preserving the historical record |
| supersede | Establish a new governing object for a declared scope |
| retire | Remove an object from ordinary activation without denying its history |
| revoke | Withdraw present authority or eligibility while preserving why and when it was withdrawn |
| roll back | Restore a prior operational state while retaining the intervening record and reason |
| restore | Re-activate an earlier object with a new justification |
This vocabulary prevents two opposite errors: treating all change as deletion, and treating every historical statement as eternally active.
5. Contradictions are first-class signals
A contradiction can occur between:
- two sources;
- a source and an interpretation;
- two accepted conclusions;
- a plan and later evidence;
- an artifact and its declared purpose;
- a workflow rule and observed practice;
- a current recommendation and a prior decision;
- a provider response and authoritative ground;
- two scopes that were mistakenly treated as identical;
- a stated state and reconstructed file evidence.
Not every difference is a contradiction. Some are distinctions of time, scope, abstraction, role, or vocabulary. Contradiction review therefore begins by asking whether the propositions are actually co-referential and mutually exclusive.
Authorized departures are evidence about standards
A legitimate departure from applicable guidance is not necessarily a failure of either the person or the standard. It can reveal a boundary condition. A first-class departure record preserves the guidance, current scope, actor, authority, circumstance, rationale, consequence, and observed result. That record keeps the exception legible without silently changing the guidance that governed the original scope.
Aggregated departures can later surface a practice for review. Repeated exceptions may show that a standard is too broad, stale, poorly scoped, or missing a known alternative. They remain evidence rather than votes: recurrence can open a review, but only the relevant governing authority can reaffirm, narrow, qualify, split, supersede, or retire the practice. In this way, rules can learn from their exceptions through the same promotion and correction boundaries that govern every other change to accepted ground.
6. A family of review movements
AIOS treats review as a set of distinct cognitive movements rather than a generic request to “check this.”
Contradiction review
Find propositions that cannot jointly govern the same scope. Distinguish genuine conflict from temporal or contextual difference.
Evidence review
Ask whether cited evidence supports the exact claim, whether contrary evidence exists, and whether source quality is proportionate to consequence.
Assumption review
Make hidden premises visible. Separate required assumptions from convenient habits inherited from prior work.
Alternative-explanation review
Generate plausible accounts that fit the same observations and identify discriminating evidence.
Decision-reconsideration review
Reconstruct what was known when the decision was made, then evaluate whether changed evidence, objectives, or constraints warrant a new decision.
Artifact critique
Judge the produced object against purpose, audience, contract, and material quality—not merely against formal completion.
Transition and coherence review
Check whether local sections remain intelligible in the whole composition and whether important conceptual transitions are earned.
Source and quotation review
Verify attribution, source descent, context, quotation fidelity, and the boundary between sourced fact and synthesis.
Confidence review
Ask whether the expressed confidence matches evidence, uncertainty, model limitations, and the cost of error.
Semantic-drift review
Compare a concept’s current use with its governing definition and revision lineage. Drift may signal either corruption or legitimate conceptual development.
Dissent-preservation review
Confirm that a minority position, counterexample, or unresolved objection has not disappeared merely because a prevailing view was promoted.
7. Independent review has conditions
Multiple model outputs are not automatically independent evidence. They may share:
- training distributions;
- sources supplied in context;
- prompts or framing assumptions;
- provider infrastructure;
- social and linguistic priors;
- incentives to agree with the person.
Independent review becomes stronger when it introduces a materially different source set, method, role, adversarial objective, empirical test, or human expertise. Agreement without methodological independence is convergence, not corroboration.
8. Risk-proportionate epistemic maintenance
The depth and independence of review should rise with:
- irreversibility of the effect;
- regulatory or safety consequence;
- number of downstream consumers;
- expected lifetime of the conclusion;
- degree of uncertainty;
- novelty of the domain;
- likelihood that social agreement masks shared error;
- cost of later correction.
Review effort follows consequence and reuse
Claims and changes receive progressively stronger review as their consequences, reuse, and correction costs rise.
Text equivalent
Proposed claim, decision, or change → Consequence and reuse; Local source check and ordinary revision → Record proportionate lineage; Named review plus counterevidence search → Record proportionate lineage; Independent method, domain review, and explicit acceptance → Record proportionate lineage; Record proportionate lineage → Activate only within accepted scope.
The branches are proportional, not absolute: low-risk work may need a local source check, moderate work a named review and counterevidence search, and high-consequence work an independent method, domain review, and explicit acceptance. Every path records lineage and activates only within accepted scope.
The purpose is not bureaucratic maximalism. It is to spend epistemic effort where a mistaken conclusion can compound.
9. The role of metadata
Companion metadata can hold the relationships that would burden person-owned prose:
claim_id: claim-example-014
standing: accepted
scope: project-and-domain-specific
source_refs:
- source-a
- source-b
assumptions:
- assumption-03
counterevidence_refs:
- objection-07
confidence:
level: moderate
basis: convergent-sources-with-one-unresolved-boundary
consumed_by:
- plan-04
- workflow-guide-02
supersedes: claim-example-009
superseded_by: null
review_trigger: source-standard-change-or-material-counterexample
This is an illustrative shape, not a universal schema. The principle is stable: operational structure belongs beside the artifact when putting it inside the artifact would make the artifact worse for the person.
Metadata and provenance support attribution, temporal resolution, impact analysis, recovery, and exact source descent. They do not establish truth. A fully traceable conclusion can still rest on weak evidence or a mistaken interpretation; correction therefore returns from metadata to the underlying sources and reasoning.
10. Correction scope and impact analysis
A corrected conclusion may affect:
- contexts that quote or summarize it;
- plans whose assumptions depend on it;
- workflows or guidance that operationalize it;
- recommendations whose rationale uses it;
- artifacts that present it as accepted ground;
- tests that encode its earlier behavior;
- assignments currently operating under it;
- organizational or domain packages that redistribute it.
The Flow Atlas can identify exact registered consumers and offer semantic-neighbor candidates for human review. It must not claim exact impact from vector similarity or textual resemblance alone.
11. Preserving dissent and alternative paths
A knowledge system becomes brittle when promotion collapses disagreement into a single apparently unanimous truth. Durable dissent should record:
- what is disputed;
- who or what method produced the objection;
- the evidence or principle behind it;
- the scope in which it matters;
- what observation would strengthen or weaken it;
- whether it blocks action or remains a monitored minority position.
The goal is neither permanent indecision nor artificial balance. It is recoverable disagreement: future reasoning can see why consensus was limited and what would justify reopening it.
12. Self-correction across interruption
Transcript continuity is not enough. After a long interruption, the system should reconstruct:
- the governing purpose;
- the accepted conclusion and its exact scope;
- the evidence and rationale available at acceptance time;
- unresolved objections;
- later evidence and its provenance;
- affected artifacts and consumers;
- the authority required for correction;
- the next meaningful test.
This is why the strongest self-correction experiment is multi-session. A correction performed inside one conversational window may only demonstrate short-term attention.
13. Worked semantic trace
Suppose a team accepts: “All high-quality workflow stages should require a fixed three-step model topology.” Later, the Fractal Seed’s reasoning grammar establishes that Why–How–What is a compositional pattern, not a mandatory number of calls.
The correction should not silently rewrite the original decision. It should:
- recover the earlier evidence and the problem the three-step design solved;
- distinguish semantic completeness from call topology;
- identify workflows, prompts, tests, and documentation that treat three calls as mandatory;
- propose the narrower rule: every consequential composition preserves Why, How, and What, while implementation topology remains task-shaped;
- obtain acceptance for the changed architectural rule;
- supersede the broader rule while preserving why it once seemed useful;
- update canonical guidance and invalidate affected derived contexts and projections;
- demonstrate that a later workflow can use two, three, or more calls without losing semantic completeness.
The system has corrected itself only at step 8.
14. Failure modes
| Failure | Consequence |
|---|---|
| silent overwrite | the past becomes unintelligible and audit claims weaken |
| perpetual accumulation | obsolete ground continues contaminating context |
| consensus-as-truth | correlated model agreement masquerades as independent evidence |
| source-free correction | a new assertion replaces an old assertion without stronger ground |
| unbounded propagation | a local exception destabilizes unrelated scopes |
| under-propagation | downstream plans and contexts retain the obsolete conclusion |
| correction theater | records change but later reasoning does not |
| dissent erasure | future reviewers cannot recover unresolved objections |
| review maximalism | epistemic maintenance consumes attention without proportional value |
| authority bypass | the system changes consequential ground without valid acceptance |
15. Evaluation program
Self-correction should be measured through adversarial longitudinal tasks:
| Measure | Question |
|---|---|
| source recovery | Can the original evidence and rationale be recovered after interruption? |
| contradiction precision | Are genuine contradictions separated from changes of time or scope? |
| correction recall | Are the material downstream consumers found? |
| correction precision | Are unrelated consumers left unchanged? |
| lineage integrity | Can a reviewer reconstruct what changed and why? |
| dissent retention | Do unresolved objections survive promotion and correction? |
| authority fidelity | Are consequential changes accepted by the right authority? |
| behavioral uptake | Does later reasoning use the corrected ground? |
| false-notice burden | How often do proactive drift or contradiction notices waste attention? |
| recovery cost | How much person effort is required to re-establish the situation? |
Useful baselines include transcript-only continuity, flat retrieval over documents, ordinary version history without semantic lineage, and an expert manually reconstructed control.
16. Research hypotheses
- Explicit epistemic relationships improve long-horizon correction relative to transcript replay alone.
- Purpose-shaped source descent reduces confident but unsupported correction.
- Preserving dissent improves later adaptation when environmental conditions change.
- Risk-proportionate review produces better assurance-to-attention ratios than uniform review.
- A correction graph improves downstream update recall but requires strict confidence labels to avoid false exactness.
- Behavioral proof—changed later reasoning—is a stronger acceptance criterion than successful record mutation.
These are testable architectural hypotheses, not established performance claims. Thresholds, review boundaries, schemas, revalidation intervals, and comparative performance against simpler baselines remain empirical questions.
17. Research connection
Current long-horizon memory evaluations support treating correction as temporal state resolution rather than overwrite. LongMemEval found that incorrect temporal pruning can remove evidence needed for later answers. MemoryAgentBench reports much stronger performance on single-hop consolidation than on multi-hop updates, where evaluated memory systems remained substantially weaker. These studies do not test the AIOS correction lifecycle, but they support explicit operations such as qualify, narrow, supersede, revoke, and preserve for history—followed by a test that later reasoning uses the changed state.
Research also limits what can be inferred from additional reviewers. ContextualJudgeBench found that even its strongest tested judge was only about 55% consistently accurate across difficult contextual comparisons. A 2026 matched-compute study found that multi-agent coordination produced large gains on some tasks and large losses on others. Review becomes more independent through different evidence, methods, adversarial objectives, empirical tests, or qualified human expertise—not through model or agent count alone.
The full evidentiary account is available in the relational-memory and AI-native architecture briefs.
18. What this page contributes to the whole
Self-correction prevents the self-evolving knowledge system from becoming a one-way accumulation mechanism. It preserves historical intelligibility while changing present authority, uses the Atlas to locate consequences without claiming omniscience, and completes correction only when renewed context produces appropriately changed behavior. Chapter 22 now descends from this semantic architecture into the commands, files, validations, receipts, and reconstruction that make those transitions real.