AIOS Proresearch
AIOS Intelligence System · Research Overview

Research library · Full memo

Cognitive Agency in AI-Mediated Reasoning

Evidence on cognitive surrender, offloading, overreliance, deskilling, metacognition, and interaction design

This memo reviews evidence published or materially updated through 10 August 2026. It distinguishes assisted performance, learning, durable skill, and agency while preserving direct descent to primary studies.

Executive findings

The strongest available evidence does not support either of the simple positions that generative AI inevitably erodes cognition or that retaining a human “in the loop” reliably prevents harm.

It supports a conditional account:

  1. AI advice can move human performance sharply in either direction. In short reasoning and decision tasks, correct AI advice often improves answers, while confident but wrong advice can reduce accuracy below an unaided baseline. People do not reliably discriminate between the two, and explanations can increase reliance on incorrect as well as correct outputs.
  1. Assisted performance is not the same outcome as learning, skill, or agency. Several experiments show better work while AI is available but no later unaided advantage; one large field RCT and one smaller, high-attrition RCT report later unaided penalties after access to general-purpose ChatGPT, with the clearest answer-delivery mechanism shown by Bastani et al. A small number of studies find positive unaided learning or transfer under structured tutoring or regulated constructive use.
  1. Evidence of general or permanent cognitive decline is not yet available. The literature measures task accuracy, confidence, effort, immediate or short-delay learning, and—in one clinical observational study—changes in unaided professional performance over several months. It does not establish irreversible loss of critical thinking, a population-wide collapse of human agency, or a stable new cognitive architecture.
  1. Cognitive offloading is not inherently harmful. Offloading can free scarce attention and working memory, distribute expertise, accelerate work, and improve learning when the external system supports rather than replaces the cognitive operations the user is meant to acquire or retain. Some studies find positive unaided learning or transfer: structured tutoring improved immediate learning, while regulated constructive use improved later unaided performance in one study. The central design question is not whether cognition is offloaded, but which operations are delegated, which remain visible and practiced, and whether the person can detect error, contest output, and perform when assistance is absent.
  1. The best-supported candidate mechanisms for testing elicit a meaningful cognitive act. Evidence is most favorable for forming an initial view, generating ideas before AI exposure, stating confidence or decision criteria, taking constructive notes, working through hints, checking sources, comparing alternatives, receiving diagnostic feedback, and retaining real authority to override or stop a process. Most evidence concerns immediate reliance or task performance rather than durable cognition or agency, and several interventions increase underreliance, time, or cognitive load. A generic approval step or arbitrary delay is not enough.
  1. Short foreground responses are a plausible attention-management technique, not an established safeguard. Progressive disclosure can increase exploratory engagement, but terse answers can also conceal assumptions and amplify the authority of a fluent conclusion. The best current inference is to keep the foreground compact while preserving on-demand evidence, uncertainty, alternatives, and provenance. “Shorter is safer” is unsupported.
  1. There is no direct evidence that non-blocking downstream work protects cognition. It may preserve flow and allow the user to keep thinking, or it may move consequential reasoning out of view and accumulate review debt. Its cognitive effect depends on whether background work is bounded, reversible, inspectable, and returned to a real judgment point before consequential integration.
  1. These findings identify testable connection points for AIOS; they do not validate AIOS. Local files, canonical person-controlled knowledge, exact operations, visible provenance, multiple-resolution memory, the Fractal Seed, and explicit human authority could support inspectability and contestability. Their cognitive effects have not been directly evaluated.
  1. The larger AIOS thesis deserves examination beyond the evidence already available for its parts. AIOS proposes a complementary local reasoning and knowledge architecture, not the end of frontier models, centralized training, or every remote service. If increasingly capable local models, durable person-controlled context, bounded semantic judgment, exact software operations, and meaningful human authority work together as proposed, the first-order consequence could be a larger share of ordinary reasoning completed privately and locally; second-order consequences could include personal and organizational intelligence systems, less dependence on centralized application layers and routine remote inference, new peer-to-peer collaboration patterns, and wider access to durable domain expertise. Current studies establish feasibility and risks for components, not those combined outcomes. The architecture-hypothesis section therefore separates evidence, thesis, implications, and conditions/tests.

Scope, method, and evidentiary discipline

This is a structured critical review, not a formal systematic review or meta-analysis. It prioritizes primary peer-reviewed studies and authoritative reports published or materially updated from August 2024 through August 2026. Papers were selected for methodological relevance to at least one of six questions:

Preprints, working papers, vendor-authored research, perspectives, and non-replicated findings are labeled. Older sources appear only in the foundational-lineage section. Evidence from education, consumer choice, knowledge work, and clinical decision support is not treated as interchangeable. Each can establish a mechanism or boundary condition without automatically generalizing to AIOS.

Unless an independent replication is explicitly cited, each primary finding below should be treated as unreplicated for its specific task, interface, population, and outcome. Peer review and multiple experiments within one paper are not substitutes for independent replication.

The review uses four outcome levels that should not be conflated:

Outcome levelWhat it can establishWhat it cannot establish
Assisted task performanceWhether a human-AI configuration improves the current outputLearning, durable skill, or independent capability
Reliance and calibrationWhether people accept correct and incorrect advice appropriatelyAbsence of internal thought or transfer of moral agency
Later unaided performanceShort-term retention, transfer, or possible skill substitutionPermanent deskilling unless observed over an adequate duration
Longitudinal independent performanceChange in capability or behavior after repeated exposureA causal AI effect without an adequate counterfactual and control of temporal confounding

Concepts and their evidentiary status

Cognitive offloading is the use of an external resource to reduce internal cognitive demand. Notes, calculators, search engines, checklists, collaborators, and AI systems can all serve this function. Offloading may be adaptive or maladaptive depending on the task objective, future need for unaided performance, accuracy of the external resource, and the user’s metacognitive calibration.

Automation bias and overreliance refer to inappropriate acceptance of automated advice or failure to detect its errors. These constructs are behaviorally measurable. A system can simultaneously produce appropriate reliance on correct answers, overreliance on wrong answers, and underreliance on useful advice.

Cognitive surrender, as used by Shaw and Nave, denotes acceptance of an AI answer with minimal or absent internal deliberation. It is a provocative hypothesis, but following an AI answer does not by itself show that thought stopped. Establishing surrender would require process measures such as an initial judgment, information-search behavior, deliberation traces, verification attempts, interviews, or convergent physiological measures.

Deskilling is a decline in a capability that was previously present. It requires a longitudinal or at least pre/post unaided measure. Lower effort during assisted work, weaker output on one later task, or neural differences during tool use are not by themselves proof of durable deskilling.

Metacognition includes monitoring one’s knowledge and performance and controlling strategy accordingly. Relevant measures include confidence calibration, error discrimination, knowing when to seek or reject assistance, and selecting a task-allocation strategy.

Human agency is broader than final decision authority. Operationally, it includes the ability to set goals, define criteria, understand relevant system limits, inspect and contest outputs, choose whether and how to delegate, reverse actions, and remain accountable for consequential commitments.

1. Assessment of “Thinking—Fast, Slow, and Artificial”

Status and contribution

Steven D. Shaw and Gideon Nave’s “Thinking—Fast, Slow, and Artificial: How AI Is Reshaping Human Reasoning and the Rise of Cognitive Surrender” is a 2026 preprint/Wharton working paper, not peer reviewed. A version is also registered at PsyArXiv. The paper reports three experiments, described as preregistered, totaling 1,372 participants and 9,593 trials.

Its theoretical proposal adds “System 3”—external, automated AI cognition—to the familiar intuitive/deliberative System 1/System 2 distinction. It differentiates beneficial cognitive offloading, in which internal evaluation remains active, from cognitive surrender, in which an AI answer is accepted without verification. The language is useful for generating hypotheses about human-AI interaction. The experiments, however, test advice influence on short Cognitive Reflection Test problems; they do not establish a distinct cognitive system.

The architectural idea also has a close recent precedent. Chiriatti and colleagues proposed AI-mediated “System 0” thinking in a two-page Nature Human Behaviour correspondence published in October 2024. That piece is a perspective, not empirical evidence, but it weakens any implied novelty in simply adding an external AI system to a dual-process model.

Research question

When people can consult an AI assistant on reasoning problems, how strongly do their final answers follow correct or deliberately faulty AI advice, and do time pressure, incentives, feedback, trust, need for cognition, or fluid intelligence change that reliance?

Method and sample

The manuscript links an OSF project for preregistrations and materials. At the time of this review, its unauthenticated endpoints were not accessible, so the preregistrations, code, and exclusions could not be independently audited.

Principal findings

In Study 1, the assistant was consulted on approximately 53–54% of eligible trials. Conditional on consultation, participants followed 92.7% of correct AI recommendations and 79.8% of deliberately faulty recommendations. Accuracy was 45.8% without AI, 71.0% with correct AI, and 31.5% with faulty AI. Correct advice therefore raised performance by 25.2 percentage points relative to the no-AI baseline, while faulty advice lowered it by 14.3 points. Global confidence was also higher with AI.

Study 2 found high faulty-advice following both with and without time pressure. The direct interaction between time pressure and selectivity in following correct versus wrong advice was not significant; a pooled categorical “surrender” result was borderline and imprecise.

In Study 3, the incentive-plus-feedback bundle increased override of faulty advice from 20.0% to 42.3% and improved accuracy on both faulty- and correct-AI trials. Per-item confidence remained higher with AI and did not reliably distinguish correct from faulty AI.

In the pooled analysis, the authors classified 73.2% of AI-consulted, AI-incorrect trials as cognitive surrender, 19.7% as offloading, and 7.1% as failed overrides. With incentives and feedback, the corresponding surrender share was 57.9% and offloading 37.1%.

What the experiments establish

What remains interpretation or hypothesis

Evidentiary verdict

The paper offers strong, narrow evidence of context-bound overreliance; preliminary evidence for a broader construct of cognitive surrender; and weak evidence for a new cognitive architecture. It is valuable precisely because it shows both directions of AI influence and because its intervention changes behavior. Its strongest claim is not that AI makes people stop thinking; it is that, on a class of deceptively simple reasoning tasks, voluntary consultation of a fluent assistant creates poor error discrimination, and that the consequences depend strongly on whether the assistant is right.

2. Recent empirical evidence

2.1 Acute reliance, confidence, and critical-thinking behavior

Stadler, Bannert, and Sailer, “Cognitive ease at a cost” — peer-reviewed primary study, 2024

Question. Does using ChatGPT rather than web search reduce cognitive load, and does that ease affect the quality of reasoning? Method. Randomized experiment with 91 university students asked to research nanoparticles in sunscreen and make and justify a recommendation, using either ChatGPT-3.5 or Google. Finding. ChatGPT users reported lower cognitive load on all measured dimensions but produced fewer relevant arguments and weaker justifications. Limitations. One 20-minute task, student sample, older model, self-reported load, and no later unaided test. Relevance. The experiment shows that lower effort can coincide with weaker synthesis. It does not establish cognitive decline or show that effort should always be maximized.

Lee et al., “The Impact of Generative AI on Critical Thinking” — peer-reviewed CHI study; vendor-authored, 2025

Question. How do knowledge workers describe critical thinking when using generative AI, and how is effort related to confidence? Method. Survey of 319 frequent generative-AI users reporting on 936 real work episodes. Finding. Greater confidence in the AI was associated with less self-reported critical-thinking effort; greater confidence in one’s own task ability was associated with more. Workers described a shift from direct execution toward goal and prompt formulation, output verification, and integration or stewardship. Limitations. Retrospective self-report, selected frequent users, no objective performance measure, no causal inference, and Microsoft co-authorship. Relevance. The study suggests redistribution rather than simple disappearance of critical thinking. It also warns that identical output reuse can follow either careful review or uncritical acceptance; artifacts alone do not reveal the cognitive process.

Wu et al., “Human–generative AI collaboration enhances task performance but undermines human’s intrinsic motivation” — peer-reviewed primary studies, 2025

Question. What happens to output, later solo performance, perceived control, motivation, and boredom when people move between AI-assisted and solo work? Method. Four preregistered online experiments with 3,562 UK participants using two-task sequences on promotional writing, performance reviews, email, and brainstorming. Finding. AI usually improved selected attributes of the assisted output, but gains generally did not persist to the next solo task. AI-to-solo transitions consistently increased perceived control. Motivation declined and boredom increased across task sequences; some comparisons showed larger motivation declines after AI-to-solo transitions, but these effects were not unique to that transition in every study. In the largest study, moving from solo to AI sharply reduced perceived control. Limitations. Two brief tasks, immediate self-reports, some automated text proxies for quality, substantial instruction-based exclusions in the largest experiment, and no durable skill measure. Relevance. The experiments show a short-term motivational and perceived-agency tradeoff. They do not show that users permanently lose agency.

Kim et al., “Fostering Appropriate Reliance on Large Language Models” — peer-reviewed CHI study; partly vendor-authored, 2025

Question. How do explanations, sources, and internal inconsistencies affect reliance on correct and incorrect LLM answers? Method. A think-aloud study informed a preregistered controlled experiment with 308 participants answering objective questions in an LLM-infused application. Finding. Explanations increased reliance on both correct and incorrect answers. Clickable sources causally reduced reliance on incorrect answers. Naturally occurring inconsistencies in erroneous explanations were associated with lower reliance, but inconsistency was not experimentally manipulated. Limitations. Short, low-stakes tasks; reliance rather than durable learning; availability of a source does not ensure it will be opened. Relevance. This study directly rejects “more explanation” as a universal remedy. Verifiability and diagnostic cues matter more than persuasive fluency.

Spatharioti et al., “Effects of LLM-based Search on Decision Making” — peer-reviewed CHI study; vendor-authored, 2025

Question. How does LLM-based search affect speed, accuracy, and overreliance, and can uncertainty highlighting improve calibration? Method. Two controlled experiments, N=90 and N=120, on decision tasks using LLM-based versus conventional search and token-probability-based uncertainty highlighting. Finding. LLM search roughly halved completion time, but users overrelied on an erroneous model response on a deliberately selected difficult item. Low token-probability highlighting increased accuracy on difficult items from about 26% to 53–58% and encouraged follow-up checking. Limitations. Small, short experiments, particular search interfaces, and Microsoft affiliation. Relevance. The paper provides direct evidence that a compact, accurate error-likelihood cue can provoke verification. It does not show that generic model self-confidence will be calibrated enough to serve this role.

De Jong et al., “Cognitive Forcing for Better Decision-Making” — peer-reviewed CSCW study, 2025

Question. Can showing only part of an AI explanation reduce overreliance? Method. Two experiments: N=264 on weighted shortest-path problems and N=210 on spelling and grammar, comparing no, full, and partial explanations in the presence of induced AI errors. Finding. Partial explanations reduced acceptance of wrong suggestions relative to a baseline, but also reduced acceptance of correct suggestions. Full explanations produced the best total performance in these tasks. Difficulty and need for cognition changed the effects. Limitations. Prolific participants, synthetic low-stakes tasks, no delayed transfer, and no demonstration that partial explanation improves understanding. Relevance. The experiments show an overreliance–underreliance tradeoff. Reducing reliance is not the same as improving calibrated reliance, and less information is not categorically better.

Spitzer et al., “The effect of medical explanations from large language models on diagnostic accuracy in radiology” — peer-reviewed randomized study, 2026

Question. How do a short diagnosis, a differential list, and a detailed explanation affect expert diagnostic accuracy and reliance? Method. Preregistered randomized experiment with 101 US radiologists, each assessing 20 image-based cases, for 2,020 assessments. Conditions included no AI and three LLM-output formats. Finding. Detailed explanations improved accuracy by 12.2 percentage points over no AI, 7.2 points over a short diagnosis, and 9.7 points over a differential list. When AI was wrong, adherence was substantially lower with a detailed explanation than in an erroneous differential-list condition; when AI was right, appropriate adherence was highest with detailed support. Limitations. Controlled vignettes, one medical specialty and model, no patient or longitudinal skill outcome, and output correctness varied somewhat by format. Generated reasoning need not faithfully reveal a model’s internal process. Relevance. The study is strong counterevidence to a universal brevity principle. Experts may need enough case-specific reasoning to audit an answer; a compact foreground should therefore retain access to inspectable evidence and rationale.

Tian, Amin, and Yin, “AI-Assisted Critical Thinking” — peer-reviewed CHI study, 2026

Question. Can an interface that elicits a user’s initial estimate, rationale, confidence, and reflection on conflicting evidence reduce overreliance in AI-assisted decisions? Method. Controlled experiment with 402 Prolific participants completing 20 house-price decisions across human-only, recommendation, hypothesis-driven analyzer, and critical-thinking-assistant conditions. Finding. The most reflective interface reduced switching to AI compared with a recommender (24.3% versus 40.9%) and reduced overreliance when AI was wrong. It also increased underreliance when AI was correct, time, and cognitive load, and did not improve overall accuracy for the full sample. Benefits were larger in some more knowledgeable groups. Limitations. Synthetic task, lay participants, possible information overload, exploratory subgroup effects, and no long-term outcome. Relevance. The study is strong direct evidence that substantive judgment points change behavior. It also shows that reflection aids can overshoot and should be adapted to expertise and stakes.

Keke, Eisenhardt, and Meske, “Thinking Twice” — peer-reviewed conference experiment, 2025

Question. Does requiring a decision before consultation improve immediate human-AI problem solving relative to simultaneous access? Method. Online experiment with 130 analyzed UK participants assigned to no AI, concurrent AI, or sequential “decide first, consult later” assistance on two brain teasers using prerecorded, correct ChatGPT responses. Finding. Accuracy was 33.0% without AI, 48.8% with concurrent AI, and 66.7% with sequential AI; the sequential–concurrent difference was 17.9 percentage points. Sequential participants used fewer prompts. Limitations. Two familiar puzzles, small sample, correct and preconstructed advice, optional initial response in implementation, and no erroneous-AI, transfer, or durable-skill test. Relevance. The paper directly supports a pre-AI judgment point for immediate performance, but not as a complete solution to automation bias.

Fernandes et al., “AI Makes You Smarter but None the Wiser” — peer-reviewed primary study, 2026

Question. How does AI assistance affect logical-reasoning performance and metacognitive accuracy? Method. Study 1 compared 246 AI-assisted Prolific participants on 20 LSAT-style items with an older unaided benchmark; Study 2 randomized 452 participants to AI or no AI. Participants predicted performance and rated confidence. Finding. AI substantially improved task performance. Participants in both conditions overestimated their scores, and trial-level confidence was poorly diagnostic. In the randomized study, AI did not clearly worsen absolute overestimation, though it altered the relation between ability and self-assessment and produced no human-AI synergy over the model alone. Limitations. Study 1’s historical comparison is noncausal; one logical-reasoning task; immediate metacognition only; ceiling effects; mandatory AI interaction in the AI condition. Relevance. The paper complicates claims that AI uniquely destroys metacognitive accuracy. Poor calibration can pre-exist AI, while AI still changes the performance–self-knowledge relationship.

2.2 Learning, transfer, and possible deskilling

Bastani et al., “Generative AI without guardrails can harm learning” — peer-reviewed PNAS field RCT, 2025

Question. Does unrestricted GPT access improve practice while harming later unaided mathematics performance, and can a tutor design prevent the harm? Method. Preregistered classroom-randomized trial with nearly 1,000 Turkish high-school students over four 90-minute sessions. Conditions were no AI, a general GPT-4 interface, and a tutor grounded in teacher solutions and common errors and instructed to give hints rather than final answers. Finding. During assisted practice, general GPT raised performance by 48% and the tutor by 127% relative to control. On the subsequent unassisted exam, general-GPT users scored 17% below control. Tutor users were statistically indistinguishable from control: the tutor removed the measured penalty but did not create an unaided advantage. Logs suggested more answer requests and copying in the unrestricted condition and more attempts and help-seeking in the tutor condition. Students were overly optimistic about learning. Limitations. One school and subject, short intervention, limited delayed retention, classroom-level allocation, and a bundled tutor treatment whose components cannot be isolated. Relevance. This study is the strongest direct evidence that answer delivery can substitute for learning and that interaction design can change the outcome.

Melumad and Yun, “Experimental evidence of the effects of large language models versus web search on depth of learning” — peer-reviewed PNAS Nexus, 2025

Question. Does learning from an LLM synthesis rather than web links change effort, perceived knowledge depth, and the quality of later advice? Method. Seven online and laboratory experiments totaling N=10,462. Conditions included live ChatGPT versus Google, identical facts presented as a summary versus linked pages, and Google AI Overviews versus conventional results. Participants researched a topic and then advised someone else; independent recipients evaluated some advice. Finding. LLM-summary users generally spent less effort, reported shallower knowledge, and produced shorter, less factual, less original advice that recipients found less informative and were less likely to adopt. Effects persisted when facts were held constant and when links accompanied the summary. Limitations. “Depth” partly relies on self-report and downstream prose proxies; browsing creates repeated exposure and source-navigation activity in addition to synthesis; the experiments do not test delayed factual retention or durable skill. Relevance. The study suggests that pre-synthesized answers can remove knowledge-construction operations. It does not imply that every compact synthesis is cognitively harmful; task purpose and required follow-on work matter.

Kreijkes et al., “Effects of LLM use and note-taking on reading comprehension and memory” — peer-reviewed, online 2025/issue 2026

Question. How do LLM use and note-taking affect delayed comprehension and memory? Method. Preregistered randomized experiment with 344 analyzed students aged 14–15 across seven English schools. Students read history passages using LLM-only, notes-only, or LLM-plus-notes support; outcomes were tested three days later. Finding. Notes-only outperformed LLM-only on retention (d=.44), comprehension (d=.38), and recall (d=.21). Adding notes to LLM use produced smaller retention and comprehension gains over LLM-only. Students nevertheless preferred the LLM and judged it easier and more helpful. Limitations. GPT-3.5, one reading activity, no reading-only group, and some combined-condition notes copied AI text rather than reconstructing it. Relevance. The paper supports pairing assistance with a user-generated artifact and shows a gap between perceived and measured learning.

Barcaui, “ChatGPT as a cognitive crutch” — peer-reviewed RCT, 2025

Question. Does access to ChatGPT during instruction affect longer-term retention? Method. Randomized study with 120 undergraduates learning about AI; 85 remained for a surprise test 45 days later. Finding. The unrestricted ChatGPT group retained less knowledge than the traditional condition (reported means 57.5% versus 68.5%, d=.68). Limitations. High attrition, one Brazilian business-school course, modest retained sample, one intervention, and limited information about participants’ actual AI strategies. Differential attrition and context constrain causal confidence in the delayed estimate. Relevance. The study is unusual in measuring a 45-day outcome, but it is not sufficient on its own to establish general deskilling.

Budzyń et al., “Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy” — peer-reviewed observational study, 2025

Question. Did experienced clinicians’ unaided adenoma-detection performance change after regular exposure to computer-aided detection? Method. Retrospective multicenter study at four Polish endoscopy centers participating in a trial. Nineteen experienced endoscopists performed 1,443 non-AI colonoscopies: 795 before and 648 after AI introduction. Finding. Unaided adenoma detection fell from 28.4% to 22.4%, an absolute difference of −6.0 percentage points (95% CI −10.5 to −1.6); adjusted odds ratio for the post-exposure period was .69. Limitations. Before/after observational design, three-month windows, possible workload, case-mix, temporal, or behavior changes, and no direct measurement of visual-search skill. AI exposure is a plausible explanation, not a uniquely identified cause. Relevance. The study00133-5) is the strongest recent real-world signal of possible professional deskilling, but its design requires cautious causal language.

Lower-weight, high-publicity preprints

These studies are useful hypothesis generators. Neither justifies claims of generalized or permanent atrophy.

2.3 Evidence that AI can extend or improve cognition

Vaccaro, Almaatouq, and Malone, human-AI meta-analysis — peer-reviewed systematic review, 2024

Question. When do human-AI combinations outperform humans alone, AI alone, or the better of the two? Method. Preregistered meta-analysis of 106 experiments and 370 effect sizes from studies available through June 2023. Finding. Human-AI configurations outperformed humans alone on average (Hedges’ g=.64) but underperformed the better of the human or AI (g=−.23). Decision tasks showed negative synergy; creation tasks were more promising. When the human was stronger than the AI, combination tended to help; when AI was stronger, combination often degraded AI performance. Limitations. Extremely high heterogeneity, mostly pre-generative-AI studies, few creation tasks, varying definitions and interfaces, and evidence no later than June 2023. Confidence and explanation displays were not significant moderators in this broad sample. Relevance. The synthesis makes a vital distinction: improvement over unaided humans is augmentation, not necessarily synergy. More than 95% of studies left the final decision to a human, showing that nominal final authority is insufficient.

Kestin et al., “AI tutoring outperforms in-class active learning” — peer-reviewed RCT, 2025

Question. Can a carefully structured AI tutor improve physics learning relative to an active-learning class? Method. Randomized crossover study with 194 Harvard introductory-physics students across two lessons. A custom GPT-4 tutor used expert-authored solutions, sequential scaffolding, targeted feedback, active-learning principles, and self-pacing. Finding. Median learning gains were more than twice those of the classroom condition. The linear model estimated an effect of .63 SD; quantile estimates correcting for ceiling effects ranged from .73 to 1.3 SD. Students spent less time and reported higher engagement and motivation. Limitations. Immediate post-tests, two topics, selective institution, treatment bundling, and an engineered tutor backed by expert-written material. It was not a test of a generic chatbot or teacher replacement. Relevance. The study demonstrates that AI can improve learning when accuracy, sequence, feedback, and pacing are deliberately designed.

Brynjolfsson, Li, and Raymond, “Generative AI at Work” — peer-reviewed field study, 2025

Question. How does an AI conversational assistant affect productivity and skill distribution in customer support? Method. Staggered deployment to 5,172 agents at a Fortune 500 software firm. The assistant generated real-time suggested replies that agents could ignore or edit. Finding. Resolutions per hour increased by about 15% overall and by roughly 30% for novice and lower-skill workers. Newer workers reached the performance of untreated workers with six months’ experience after about two months. During rare outages, previously exposed agents retained some performance advantage, offering suggestive—though noisy—evidence of learning. Top workers gained little speed and showed small quality declines. Limitations. Nonrandom deployment, one firm and occupation, stable recurring problem types, noisy outage inference, and job performance rather than general cognition. Relevance. The study shows that AI can distribute tacit expertise and compress a workplace learning curve. Effects are heterogeneous and can include downward imitation or over-adherence among experts.

Dell’Acqua et al., “The Cybernetic Teammate” — peer-reviewed field experiment, 2026

Question. Can generative AI substitute for or complement teamwork on realistic product innovation? Method. Preregistered 2×2 field experiment at Procter & Gamble. The workshop involved 826 participants; randomized analyses focused on 791 assigned to individual or team work, with or without AI, on product-development challenges. Finding. AI increased solution quality by about .37 SD for individuals and .39 SD for teams; individuals with AI matched teams without AI. AI broadened contributions across functional boundaries and improved affect. Human evaluation still mattered: AI-supported participants were slightly less effective at identifying their own strongest ideas. Limitations. One company and domain, one-day flash teams, one model, expert-rated rather than commercial outcomes, no longitudinal skill measure, and disclosed company support or employment for some authors. Relevance. The experiment is strong evidence of cognitive extension in ideation and cross-functional synthesis, alongside the need to preserve human evaluation and selection.

Dell’Acqua et al., “Navigating the Jagged Technological Frontier” — peer-reviewed field experiment, 2026

Question. How does GPT-4 affect professional performance on tasks inside and outside its capability frontier? Method. Preregistered field experiment with 758 Boston Consulting Group consultants assigned to no AI, GPT-4, or GPT-4 plus prompt guidance on realistic consulting tasks. Finding. On tasks selected to be within the model’s frontier, AI users completed 12.2% more work, were 25.1% faster, and improved quality by more than 30%. On one deliberately selected outside-frontier task, AI users were 19 percentage points less likely to reach the correct answer. Limitations. One consultancy, GPT-4, constructed tasks, only one outside-frontier task, short-term performance, no skill-transfer measure, and BCG coauthors. Relevance. The experiment shows why capability and task fit must be diagnosed before delegation. Fluency can extend expert work inside a frontier and mislead outside it.

Wong and Qiu, “Think first, ChatGPT later” — peer-reviewed experiment, 2026

Question. Does independent ideation before AI use change later unaided creativity? Method. N=196 university students assigned to human-only, unrestricted AI, or regulated “think first” collaboration, followed by a novel invention task without AI. Finding. Unrestricted AI produced the strongest initial products but no later unaided advantage. The think-first group outperformed both alternatives on the unaided transfer task and used AI more often to elaborate and evaluate self-generated ideas rather than delegate generation. Limitations. One creativity paradigm, student sample, and immediate rather than longitudinal transfer. Relevance. The study is the clearest recent evidence that a meaningful pre-AI contribution can preserve or improve independent performance.

Additional positive but bounded evidence

2.4 Authoritative standards and evidence syntheses

These sources establish governance expectations or synthesize evidence; they are not intervention trials and should not be cited as proof that a particular interface works.

3. Do the proposed interaction patterns mitigate the risks?

Evidence-grade summary

Proposed patternCurrent evidenceMost defensible interpretation
Short foreground responsesWeak and conditionalPotentially useful as progressive disclosure; unsafe if brevity hides evidence, uncertainty, or alternatives
Explicit choicesIndirect and incompleteCan preserve initiative when choices expose real alternatives and tradeoffs; mere optionality does not prevent overreliance
Human judgment pointsModerate, design-dependent supportMost promising when the person contributes an initial view, criteria, confidence, rationale, or evaluation before commitment
Non-blocking downstream workNo direct cognitive-protection evidenceTreat as a workflow hypothesis; make work provisional, bounded, inspectable, reversible, and subject to consequential review

Short foreground responses

The intuition is plausible: a compact response can reduce attentional capture, keep the user oriented to the current goal, and leave room for active inquiry. A 2026 within-subject experiment with 36 participants found that progressive disclosure and higher user prompt initiative increased perceived knowledge change, engagement, germane cognitive load, and exploratory behavior. This is encouraging but small and reliant largely on subjective outcomes. Kim and Ji.

The contrary evidence is equally important:

Therefore the defensible pattern is not “short answers preserve cognition.” It is layered disclosure: a compact foreground that communicates the decision-relevant claim, uncertainty, major assumption, and next choice, backed by accessible evidence, alternatives, provenance, and detailed work. For consequential judgments, terseness should not remove the information needed to contest the answer.

Explicit choices

Voluntary consultation and freedom to override were already present in the Shaw and Nave studies, yet faulty-advice following remained high. An explicit choice is meaningful only if the option set does not silently frame one answer as the default and if the user can understand consequences.

Choice is more likely to preserve agency when it includes:

The empirical case remains indirect. Studies of critical-thinking aids, decision-relevant uncertainty, and hypothesis-driven interfaces show that structured alternatives and reflection can change reliance, but no strong study isolates “explicit choice” as a general cognitive safeguard. Choice architecture can also anchor and narrow thought.

Human judgment points

This has the strongest empirical support of the four patterns, with an essential qualification: the judgment must be substantive.

Supported forms include:

Bastani et al.’s tutor replaced direct solutions with hints and attempts. Tian et al. elicited an estimate, rationale, and confidence before reflective prompts. Keke et al. required a decision before consultation. Wong and Qiu required ideation before AI. Kreijkes et al. required a user-produced note artifact. These mechanisms differ, but each makes the human contribution cognitively meaningful. Separately, Shaw and Nave’s incentives-plus-feedback intervention more than doubled override of faulty advice, supporting diagnostic feedback and salient stakes rather than an initial judgment checkpoint.

A final approval click is weak protection. The human-AI meta-analysis found no general synergy despite final decisions usually remaining with humans. The EU AI Act’s human-oversight requirements likewise include the competence and authority to understand limitations, interpret outputs, disregard or reverse recommendations, and stop a system; they do not treat nominal presence as sufficient.

Judgment points also impose costs. They can increase time and cognitive load, annoy users, and create underreliance. They should be concentrated at high-value boundaries: goal selection, acceptance criteria, disputed evidence, high uncertainty, irreversible operations, canonical integration, and final normative commitments.

Non-blocking downstream work

No located study directly tests whether continuing bounded work in the background preserves metacognition, critical thinking, or agency. It should therefore be treated as a systems-design hypothesis, not an evidence-backed mitigation.

Two causal pathways are plausible:

The cognitively safer form is bounded delegation:

This pattern could reduce interface friction while preserving authority, but it requires direct evaluation.

The combined interaction pattern

Short foreground responses, explicit choices, substantive judgment points, and non-blocking downstream work form a coherent design hypothesis when combined: keep the person’s immediate attention on purpose and choice; let the system perform bounded, reversible work; then return evidence, uncertainty, and unresolved decisions before integration. No located experiment evaluates that bundle. Its effect cannot be inferred by adding together findings from tutoring, decision aids, search interfaces, and workplace systems.

The bundle could fail in two opposite ways. If the foreground is too thin, the user may accept a system-framed direction without understanding it. If background work becomes too complete or costly to reverse, the later judgment point may become a rubber stamp. The key empirical test is therefore not whether a user was shown a choice, but whether the architecture improves calibrated error detection, rationale quality, revision behavior, delayed independent performance, and the practical exercise of override.

Architecture-hypothesis stress test

This section evaluates the parts of the AIOS architectural thesis that intersect with cognitive engagement, calibrated delegation, human authority, context integrity, and their direct architectural implications. It is an interpretation of adjacent evidence, not an additional research finding and not validation of AIOS. The analysis preserves large implications while distinguishing them from observations already established in component studies.

Evidence from adjacent architecture research

Hsieh et al., RULER — peer-reviewed COLM benchmark, 2024

Question. How much of an advertised long context can a model use effectively on tasks beyond simple needle retrieval? Method. Synthetic benchmark with 13 configurable tasks in four categories, evaluated from 4K to 128K tokens. The original paper evaluated ten long-context models; the maintained benchmark has since reported additional models. Finding. Models that performed nearly perfectly on simple needle tests often degraded substantially as context length and task complexity increased. In the original evaluation, only about half of models claiming at least 32K context maintained the paper’s performance threshold at 32K. Limitations. Synthetic tasks, a qualitative threshold, rapidly superseded model versions, and no direct measure of real organizational reasoning. Relevance. RULER supports active context selection and layered memory. It contradicts both “more context is always better” and “insufficient context is the sole cause of agent failure.”

Du et al., “Context Length Alone Hurts LLM Performance” — peer-reviewed Findings of EMNLP study, 2025

Question. Does performance decline with longer inputs only because retrieval becomes harder, or can length itself impair reasoning after relevant evidence is controlled? Method. Five open and closed models on mathematics, question answering, and code, with experiments that fixed evidence position, replaced irrelevant content with whitespace, or masked irrelevant tokens. Finding. Performance declined by 13.9% to 85% as inputs lengthened within supported windows, including conditions designed to remove ordinary retrieval confounds. Having the model recite relevant evidence before solving recovered some performance on one benchmark. Limitations. Benchmark tasks and artificial controls; the mechanism is not fully identified; results do not imply that every longer context is worse. Relevance. The study strengthens the case for staged, salient context. It directly rules out “just provide all available context” as a universal remedy.

Yang et al., SWE-agent — peer-reviewed NeurIPS system benchmark, 2024

Question. Can a model-specific computer interface improve an LLM agent’s ability to repair software? Method. A custom agent–computer interface provided constrained repository navigation, editing, and testing operations; the system was evaluated on SWE-bench and HumanEvalFix. Finding. Interface design materially changed agent behavior and raised performance, reaching 12.5% pass@1 on SWE-bench and 87.7% on HumanEvalFix at the time. Limitations. Software engineering only, older models and benchmarks, no human-agency outcome, and comparison conditions differ in more than “scaffolding” alone. Relevance. The study is direct counterevidence to the broad claim that scaffolding suppresses frontier reasoning. A task-appropriate interface can expose capability that a generic text interface does not.

Tam et al., structured generation and reasoning — peer-reviewed EMNLP Industry study, 2024

Question. How do JSON mode, format instructions, and staged conversion affect reasoning and classification? Method. Five models across reasoning and classification benchmarks, comparing unrestricted natural language, direct strict-format generation, format instructions, and a two-stage natural-language-to-structured conversion. Finding. Strict formats generally degraded reasoning performance, though they sometimes improved classification. Separating semantic work from later conversion often avoided part of the penalty. Limitations. Older model generation, benchmark prompts, and no tool execution or security outcome. Relevance. The study supports a model-proposal/code-enforcement boundary, but warns against forcing every intermediate cognitive operation into a rigid schema.

Xia et al., Agentless — peer-reviewed FSE study, 2025

Question. Are complex autonomous tool-using agents necessary for repository issue resolution? Method. A fixed three-stage workflow—localization, repair, and patch validation—without model-directed future actions or a large interactive tool space, evaluated on SWE-bench Lite. Finding. The simple system resolved 32.67% of the benchmark at a reported average cost of $0.68, outperforming the compared open-source agents at the time. The authors also identified benchmark problems and produced a filtered set. Limitations. One software benchmark, fast-moving baselines, system components and inference budgets not reducible to a single complexity variable, and no evidence about creative or ambiguous tasks. Relevance. The paper supports bounded decomposition and deterministic validation. It also shows that extra autonomy and orchestration are not automatically beneficial.

Ben Sghaier et al., agent-harness evolution — 2026 preprint; not peer reviewed

Question. How does an agent harness change effectiveness and efficiency when the underlying model is held constant across releases? Method. The authors first characterized five rapidly evolving open-source coding harnesses, then evaluated 35 sequential Qwen Code releases on 50 stratified SWE-bench Verified tasks with a fixed model. They measured resolve rate, token consumption, and tool calls and associated changes with development and architecture components. Finding. Resolve rate fluctuated but showed no statistically significant improvement across releases, while later releases consumed nearly twice the tokens and tool calls without corresponding average quality gains. Context-management changes were a high-risk area and expanded context management was associated with lower token efficiency; the study did not establish a general fall in correctness from greater scaffold complexity. Limitations. The paper is a July 2026 preprint, uses one coding harness, one fixed model configuration, and 50 benchmark tasks; component associations are not all randomized causal effects. Relevance. This is unusually direct evidence that a harness is a quality-critical variable independent of the model. It supports matched-model regression testing, not the universal claim that current scaffolds suppress frontier reasoning.

Kamoi et al., LLM self-correction — peer-reviewed TACL critical survey, 2024

Question. Under what conditions do inference-time self-correction methods actually improve an LLM’s initial response? Method. Critical review and methodological categorization of the self-correction literature, including the source of feedback, task suitability, training status, and fairness of evaluation baselines. Finding. The review found no convincing general evidence that an ordinarily prompted model corrects itself using only prompted self-feedback, apart from unusually self-correctable tasks. Reliable external feedback and large-scale training for correction were the more successful conditions. Limitations. A synthesis of heterogeneous prior studies rather than a new controlled benchmark; the field and model capabilities change quickly, and conclusions are conditional on the included literature. Relevance. The survey distinguishes a nominal “review role” from an operationally independent check. A separate label or fresh prompt is not equivalent to external evidence, deterministic validation, a trained critic, or human review.

Luz de Araujo et al., “Principled Personas” — peer-reviewed EMNLP benchmark, 2025

Question. Do expert-role prompts improve performance robustly and for the intended reason? Method. Nine state-of-the-art models across 27 tasks, evaluating expert advantage, sensitivity to irrelevant persona details, and fidelity to expertise attributes. Finding. Expert personas usually produced positive or nonsignificant changes, but irrelevant persona details could reduce performance by almost 30 percentage points. Education, specialization, and domain fit had inconsistent or negligible effects across tasks; proposed robustness mitigations worked mainly for the largest models. Limitations. Benchmark tasks rather than long workflows; persona labels are not equivalent to separate context stores, tools, or system roles. Relevance. The study supports careful task-specific role design but rejects a general assumption that assigning distinct expert identities reliably improves reasoning.

“Capable language models can outgrow the benefits of collaboration” — peer-reviewed Nature Machine Intelligence benchmark, 2026

Question. When do multi-agent architectures outperform a single-agent system as model capability rises? Method. Five canonical single- and multi-agent architectures, three model families, six agentic benchmarks, and 260 controlled configurations with matched compute. Finding. Multi-agent benefit depended on task, model family, coordination efficiency, error amplification, and message overhead. Some tasks showed large gains; on others coordination cost dominated, and stronger models could outgrow the benefit of collaboration. Limitations. Benchmark agents rather than human-AI teams; architecture implementations and capability index choices may shape results; publication is too recent for independent replication. Relevance. The study is strong counterevidence to universal claims for multiple roles, agents, or centers of reasoning. Decomposition must earn its overhead empirically.

Köbis et al., “Delegation to artificial intelligence can increase dishonest behaviour” — peer-reviewed Nature experiments, 2025

Question. Does delegating an incentivized task to machines change unethical behavior, responsibility, or agent compliance? Method. A series of incentivized die-roll and tax-reporting experiments using human and machine delegation. In the natural-language implementation study, a broadly US-representative Prolific sample of 975 acted on delegated instructions; follow-on tests used GPT-4, GPT-4o, Claude 3.5 Sonnet, and Llama 3.3 and varied prompt guardrails. Finding. Natural-language LLM agents complied with fully unethical requests at very high rates—79% for Llama and 98% for the other tested models in the reported condition. Prompt guardrails reduced compliance, but the strongest task-specific prohibition was the least scalable. Delegation could also distance principals from direct implementation. Limitations. Artificial, low-stakes cheating tasks; specific model versions; ethical compliance rather than operational authorization; no deterministic policy layer. Relevance. The experiments support retaining consequential authority and enforcing prohibitions outside natural-language persuasion. They also show that delegation architecture changes responsibility-related behavior.

Hemken et al., “Can a Large Language Model Keep My Secrets?” — peer-reviewed ACL Student Research Workshop, 2025

Question. Can LLM agents make confidentiality-aware access decisions from natural-language organizational scenarios? Method. Synthetic planning and deduction benchmark evaluated with Llama 3 and GPT-4o-mini, comparing prompting and fine-tuning; a 23-person human study provided a baseline. Finding. Humans reached up to 79% accuracy. Chain-of-thought and few-shot prompting remained below the human baseline and below real-world applicability, while task-specific fine-tuning reached up to 98% on the synthetic benchmark. Limitations. Small human baseline, synthetic scenarios, student-workshop publication, and no adversarial production environment. High benchmark accuracy is not a security guarantee. Relevance. The study supports using models to interpret ambiguous policy language but not trusting a prompted model as the authorization boundary. Semantic policy translation and policy enforcement are different functions.

Debenedetti et al., AgentDojo — peer-reviewed NeurIPS security benchmark, 2024

Question. How well do tool-using agents complete benign tasks and resist indirect prompt injection, and does restricting the tool surface help? Method. A dynamic environment with 97 realistic tasks and 629 security test cases. A defense filtered tools before untrusted data entered the context, limiting which actions could be exposed to the agent. Finding. The tool filter reduced targeted attack success from 57.69% to 6.84% while slightly improving benign utility. It was weaker when the same tool could serve benign and malicious goals or when required tools could not be known in advance. Limitations. Simulated benchmark, modeled attack patterns, and a defense that assumes an externally selected action set. It does not cover every execution path or policy error. Relevance. The study is direct support for task-scoped capabilities and separation of untrusted content from executable authority.

Yao et al., τ-bench — peer-reviewed ICLR benchmark, 2025

Question. Can conversational agents follow domain policy, interact with users, call APIs, and reach the exact correct database state reliably? Method. Stateful retail and airline scenarios with policy documents, user simulation, tool APIs, and evaluation of final database state and repeated-run reliability. Finding. Then-current function-calling agents completed fewer than half of tasks; repeated-run reliability was lower still, with retail pass^8 below 25% in the reported results. Valid tool calls did not imply correct policy-compliant outcomes. Limitations. Two simulated domains, model and prompt versions age quickly, and exact-state success does not measure open-ended output quality. Relevance. The benchmark demonstrates why deterministic state evaluation can reveal model planning and semantic errors, but cannot prevent them automatically.

Faghih et al., tool-description sensitivity — peer-reviewed EMNLP study, 2025

Question. How robust is tool selection to the wording of descriptions in a bounded tool menu? Method. Seventeen models evaluated under systematically varied tool descriptions and task conditions. Finding. Irrelevant or strategically altered descriptions substantially shifted selection; for some models, an edited tool became more than ten times as likely to be chosen. Limitations. Benchmark tool choice, not full end-to-end execution; wording effects vary by model; a malicious description is not an ordinary menu. Relevance. The study shows that bounded menus reduce the action space without eliminating semantic manipulation or choice architecture effects.

Li et al., RAG versus long context — peer-reviewed EMNLP Industry study, 2024

Question. When should a system retrieve selected passages rather than place the full corpus in a long model context? Method. Comparison of retrieval-augmented generation and long-context prompting across several public datasets and three then-current LLMs; the authors also tested a routing hybrid. Finding. With sufficient resources, long-context prompting had higher average performance, while RAG was substantially cheaper. A self-routing hybrid preserved performance close to long context at lower cost. Limitations. Benchmark QA, three models, model-self-reflection used for routing, no private organizational deployment, and no security or user-agency outcome. Relevance. The study supports selective context composition for efficiency, but not the superiority of local file retrieval in every task.

Zou et al., PoisonedRAG — peer-reviewed USENIX Security study, 2025

Question. Can a small number of malicious documents corrupt answers from a retrieval-grounded system? Method. Black- and white-box knowledge-corruption attacks against RAG systems with very large document collections. Finding. Injecting five malicious texts per target question achieved a reported 90% targeted attack success rate in a knowledge base with millions of documents. Limitations. Adversarial benchmark rather than ordinary accidental error; attack success depends on retrieval and generation conditions; it does not show that all local repositories are insecure. Relevance. The paper shows that person-controlled files and local retrieval create a governable attack surface, not an automatically trustworthy one. Provenance, ingestion authorization, and integrity checking remain necessary.

PALMBench — peer-reviewed ICLR benchmark, 2025

Question. What capability, latency, energy, and safety tradeoffs arise when compressed language models run on mobile devices? Method. Automated evaluation of widely used 1.1-billion- to 8-billion-parameter models, 2- to 8-bit quantization, and two inference frameworks across Pixel, Samsung, and iPhone devices plus Orange Pi and Jetson hardware. The benchmark measured memory, latency, throughput, power, question answering, MTBench, hallucination, and toxicity. Finding. Four-bit quantization often retained useful performance while reducing model size substantially, but platforms differed materially in energy and throughput and aggressive compression could produce large quality and safety penalties. In one reported Llama-3-8B configuration, measured hallucination was 32.7% at 2-bit and 37.5% at 3-bit, versus 9.1% at GPTQ 4-bit and 8.1% at 8-bit. Models above roughly 7–8 billion parameters were impractical on the tested phones under the evaluated configurations. Limitations. Mobile benchmark rather than representative personal or organizational workflows; automated and model-judge safety metrics are imperfect; hardware and models change quickly; privacy benefit was inferred from avoiding transmission rather than measured as end-to-end leakage. Relevance. The benchmark supports local execution as a feasible option with real tradeoffs. It does not establish that local models can perform most ordinary reasoning or match frontier services.

SlimLM — peer-reviewed ACL system demonstration, 2025

Question. Can compact language models provide useful document assistance directly on a current smartphone? Method. Models from 125 million to 8 billion parameters were trained and evaluated for summarization, question answering, and suggestion tasks using the authors’ DocAssist dataset; device execution was demonstrated on a Samsung Galaxy S24. Finding. The system completed the bounded document-assistance tasks on-device and exposed explicit tradeoffs among model size, context, quality, and inference time. Limitations. A demonstration-track paper using a self-constructed dataset, one phone, and narrow tasks; it does not establish general reasoning coverage or longitudinal utility. Relevance. SlimLM supplies a concrete local document-work example, while delimiting what current on-device evidence can support.

BRIGHT — peer-reviewed ICLR retrieval benchmark, 2025

Question. How well do retrieval systems find evidence for realistic queries that require reasoning rather than surface similarity? Method. The authors evaluated retrieval on 1,398 queries spanning economics, psychology, mathematics, coding, and other domains and tested query-reasoning interventions and downstream question answering. Finding. A retriever scoring 59.0 nDCG@10 on the conventional MTEB benchmark scored 18.0 on BRIGHT. Explicit query reasoning improved retrieval by as much as 12.2 points, and higher-ranked evidence improved downstream answering by more than 6.6 points. Limitations. Benchmark queries and relevance judgments are not a longitudinal personal knowledge system; results depend on corpus and retrieval setup. Relevance. BRIGHT shows both the value of relevant domain evidence and the difficulty of finding it. File ownership does not solve retrieval.

Fensore et al., domain RAG in consumer health — peer-reviewed Machine Learning for Healthcare study, 2025

Question. Does retrieval from a curated medical corpus improve open-source models on consumer-health questions? Method. Four open-source models answered 828 unique questions with 1,192 reference answers, with and without retrieval from 151 curated NIDDK documents; evaluation combined automatic metrics, model-based assessment, and clinical validation. Finding. The unaugmented models outperformed the RAG variants overall. Retrieval precision at five was 0.15, although RAG remained competitive in narrower areas such as scientific consensus and harm reduction. Limitations. One medical subdomain and one family of RAG designs; the result does not show that well-engineered domain retrieval is generally harmful. Relevance. The study is direct counterevidence to the assumption that a curated, self-contained corpus automatically improves model judgment.

Consistent local-first software and Janus — peer-reviewed distributed-systems studies, 2024

Question. What coordination is required for local replicas to preserve application invariants, stronger consistency, and Byzantine tolerance? Method. ConLoc formalized application invariants and added synchronization only where required, evaluating research and production CRDTs. Janus implemented middleware combining CRDT operation with asynchronous Byzantine-fault-tolerant consensus and evaluated throughput against a naïve strongly consistent design. Finding. ConLoc showed that explicit invariant analysis can retain much of the performance of local execution while preserving specified correctness. Janus achieved markedly higher throughput than naïvely applying HotStuff while supporting stronger guarantees. Both studies also identify cases in which eventual local convergence is insufficient. Limitations. Programming-system and database evaluations, not studies of AI, organizational adoption, privacy, or total infrastructure. Their guarantees depend on specified invariants, threat models, and middleware. Relevance. These studies support technically serious local-first operation while showing that collaboration, malicious replicas, freshness, and transactions can reintroduce identity, communication, and consensus services.

Edge and centralized inference efficiency — peer-reviewed systems evidence, 2025

Question. How much resource use can edge execution save in selected deployments, and how much can centralized utilization improve? Method. A short cloud-versus-edge case study compared the same deployable generative models on smartphone hardware and cloud GPUs for energy, carbon, and water. The industry-affiliated Aegaeon study evaluated token-level GPU pooling in benchmarks and a multi-month Alibaba deployment. Finding. The edge case reported more than 90% inference-energy savings and more than 80% carbon and water reductions in its tested configurations while setting accuracy and latency aside. Aegaeon reported 2–2.5 times higher sustainable arrival rates or 1.5–9 times greater goodput and a production reduction from 1,192 to 213 allocated GPUs for the reported model-market workload. Limitations. The edge paper is a six-page case study rather than lifecycle accounting; the cloud comparison was not a fully optimized production service. Aegaeon is vendor-affiliated and workload-specific. Neither measures rebound effects, device manufacture, or an AIOS-like application architecture. Relevance. The pair establishes competing possibilities: local execution can reduce the remote load for tasks completed locally, while centralized pooling can reduce the infrastructure required per remote task. Infrastructure implications require measurement of both.

Claim-by-claim assessment in four layers

The four layers below prevent two opposite errors: presenting a prospective architecture as already validated, or treating its first- and second-order implications as unworthy of analysis because the integrated system has not yet been tested. “If the thesis is substantially correct” denotes a conditional implication, not a forecast.

0. The transition from prose-producing AI features to AI-native intelligence architecture

1. What current evidence establishes. Systems that incorporate probabilistic semantic interpretation, context construction, natural-language tool selection, and generative planning require explicit boundaries between uncertain judgments and exact operations. SWE-agent shows that the interface through which a model acts can materially change performance; Agentless shows that a deliberately bounded workflow can outperform more elaborate agency; structured-generation research shows that placing every intermediate thought inside a rigid grammar can impair reasoning. Treating model output as an untrusted proposal is more robust than having deterministic code scrape arbitrary prose for consequential commands. These findings concern mechanisms, not a historical law: earlier software already combined uncertain classifiers, databases, reference monitors, workflow engines, and human approval.

2. The AIOS architectural thesis. “AI-native” means that semantic judgment is a first-class bounded component rather than unstructured text pasted into a conventional application. Purpose, context composition, domain files, companion metadata, model proposals, policy, exact operations, post-process integration, memory, and human authority have distinct responsibilities and explicit interfaces. The thesis is not that all deterministic pipelines disappear; it is that software for open-ended reasoning should be organized around the coordination of semantic and exact components instead of treating model prose as the center of intelligence.

3. What follows if the thesis is substantially correct. First-order, systems could adapt to new semantic tasks without replacing identity, authorization, validation, and state-management machinery, while model upgrades would not require rebuilding the canonical-file and accepted-ground layer. Models could be selected by task and sensitivity rather than being the product’s single center. Second-order, intelligence products could become less application-bound: durable context, workflow, and memory could remain with a person or organization while models and specialized services become substitutable. The competitive and governance locus could shift partly from owning an application silo to stewarding an interoperable knowledge-and-authority environment.

4. Conditions, counterforces, and tests. “AI-native” needs observable properties and comparison against credible alternatives. Some tasks remain better served by deterministic pipelines; gains may come from a better interface, a stronger model, retrieval, extra samples, or test execution rather than the architecture. Technical feasibility does not ensure adoption: network effects, standards, migration cost, provider bundling, procurement, institutional policy, and local maintenance can preserve application silos. Compare a rule pipeline, prose-plus-parser, typed model proposal plus deterministic executor, and human-reviewed hybrid on matched tasks and budgets. Measure semantic accuracy, exact-state success, security violations, recovery, auditability, interoperability, migration, model substitutability, human comprehension, latency, and total operating cost. The thesis is strengthened only if the integrated division of labor adds value beyond its individual components.

1. Models make bounded semantic judgments; deterministic mechanisms guarantee exact properties; humans retain purpose and consequential authority

1. What current evidence establishes. Language models can usefully interpret ambiguous language, rank, summarize, classify, propose relationships, and draft plans. Deterministic mechanisms are better suited to enforcing typed schemas, resource identifiers, access-control rules, scope limits, transaction boundaries, idempotency, and exact file transformations. AgentDojo shows that preselecting task-relevant tools can sharply reduce prompt-injection success; the confidentiality study shows that prompted models are not reliable authorization boundaries; τ-bench shows that valid tool calls can still end in the wrong state. The human-AI literature supports substantive human control over goals and consequential commitments, while also showing that a nominal approval step can become automation-biased rubber-stamping.

2. The AIOS architectural thesis. Models should make bounded, inspectable semantic proposals; deterministic code should mediate identity, scope, validation, authorization, structure, and controlled state change; humans should retain purpose, acceptance criteria, consequential commitment, and the ability to stop or redirect work. This is a layered trust model: semantic fluency is not execution authority, valid syntax is not correct judgment, and a person’s presence is not meaningful control unless the person has information and practical alternatives.

3. What follows if the thesis is substantially correct. First-order, the architecture could combine flexible natural-language reasoning with reproducible and auditable effects. Models could be upgraded or switched without granting them ambient authority; regulated deployments could attach policy, provenance, and review requirements to operations rather than rely on prompt wording. Second-order, organizations could delegate more interpretive work without turning every model error into an action, and individuals could use powerful models while keeping durable authority in a person-controlled system. This could support accountable professional use in which model judgments remain contestable and exact changes remain attributable.

4. Conditions, counterforces, and tests. Deterministic code guarantees only correctly specified invariants inside its trusted boundary. It cannot make a wrong model judgment true, infer legitimate purpose from authentication, or guarantee uncontrolled real-world consequences. Human purposes may also be mistaken, discriminatory, or unlawful; authority must remain accountable to external rights and policy, not become unlimited personal sovereignty. Strict schemas can impair reasoning, generic tools can bypass confinement, and repeated approvals can create fatigue. Compare model-only action, prompted guardrails, typed proposals with deterministic enforcement, and the same system with substantive human review under ambiguous, adversarial, stale-identity, partial-failure, and irreversible-action conditions. Measure semantic error, unauthorized effects, invariant violations, recovery, approval quality, contestability, and under-action as well as over-action.

PropertyDefensible deterministic claimUnsafe extension
IdentityBind an action to an authenticated credentialProve the credential represents the intended real-world person or legitimate purpose
ScopeEnforce paths, resources, operation types, capabilities, and time limitsGuarantee confinement if an unmediated path or generic execution capability remains
ValidationEnforce syntax, types, schemas, preconditions, and machine-checkable invariantsEstablish truth, quality, fairness, or correct intent
AuthorizationEnforce a formal policy at a reference monitorEstablish that the policy or delegation is appropriate
StructureGuarantee parseability or grammar membershipGuarantee sound reasoning or meaningful content
Exact effectsVerify controlled pre/post-state changes, hashes, diffs, or transactionsGuarantee external, concurrent, physical, or human downstream consequences

2. Intelligence emerges from coordinated components rather than one central agent

1. What current evidence establishes. System outcomes depend on context selection, retrieval, model capability, interface, tools, validation, execution, memory, and human decisions; model benchmark scores alone do not predict the performance of a complete human-AI system. SWE-agent, Agentless, RAG/long-context comparisons, and the human-AI experiments all demonstrate configuration effects. The evidence does not establish that a distributed configuration is always superior: recent matched-compute multi-agent results show coordination overhead, error amplification, and benefits that can shrink as base models improve.

2. The AIOS architectural thesis. Intelligence has no necessary single operational center. A user-visible result can arise from coordination among purpose, canonical files, metadata, annotations, relationships, mode-specific contexts, bounded model judgments, exact operations, memory, post-processing, and human authority. “Emergence” here is a testable system claim: changing one component or relationship can change the result even with the same base model.

3. What follows if the thesis is substantially correct. First-order, evaluation and debugging would shift from asking whether “the model” is intelligent to locating where evidence was selected, transformed, authorized, or lost. Components could be improved independently and multiple models could contribute without becoming authorities over the whole system. Second-order, durable expertise could become an evolving property of a person- or organization-controlled environment rather than a transient property of one chat or vendor. Peer collaboration could exchange inspectable files, annotations, methods, and bounded judgments while preserving separate canonical authority.

4. Conditions, counterforces, and tests. The thesis needs causal traces rather than metaphysical language. A model, orchestrator, retrieval policy, or owner can still dominate; more components can diffuse accountability and create hidden coupling. Perform component ablations and matched-budget comparisons, tracing which component originates, catches, amplifies, or corrects each error. Compare centralized and distributed orchestration with the same compute, information, and action budget, and assess not only output quality but provenance completeness, debugging time, error correlation, accountability, and recovery.

3. Current scaffolding suppresses frontier-model reasoning

1. What current evidence establishes. Scaffolding is an active causal variable. Irrelevant persona details, excessive or poorly placed context, unnecessary coordination, and rigid output formats can impair performance; specialized interfaces, selective context, staged conversion, and bounded workflows can improve it. RULER and Du et al. show that context length itself can degrade performance within supported windows. Agent failure also arises from limited capability, retrieval error, planning and tool errors, weak feedback, environment change, security attacks, and long-horizon error accumulation. Insufficient context is therefore one cause, not the sole cause.

2. The AIOS architectural thesis. Much current scaffolding was designed around generic chat, prompt templates, or earlier models and may obscure the reasoning capabilities of newer models. AIOS proposes leaner, task-specific scaffolds: give a model the context, cognitive mode, tools, and output latitude needed for a bounded semantic judgment, then move structural conversion and exact execution into later deterministic stages.

3. What follows if the thesis is substantially correct. First-order, the same frontier model could deliver higher-quality judgments with less prompt ritual, less irrelevant context, and fewer coordination tokens, while audit and safety remain at system boundaries. Model capability gains could be realized by changing context composition and interfaces rather than continually adding agent layers. Second-order, a local durable architecture could improve over time by revising domain scaffolds independently of model vendors, making model advances more transferable to person-owned knowledge and workflows.

4. Conditions, counterforces, and tests. “Current scaffolding” and “frontier reasoning” must be operationalized per model, task, and objective. Some restrictions intentionally trade benchmark quality for lower risk; a lean scaffold can remove needed context or verification. Evaluate no scaffold, minimal scaffold, proposed scaffold, and representative alternatives across reasoning, writing, editing, planning, research, and exact operations, matching model, tools, samples, tokens, latency, and retries. Score judgment quality, groundedness, invariant violations, security, cost, user calibration, and recovery. Model-version crossover studies should test whether a scaffold helps weaker models but constrains stronger ones.

4. Thinking, writing, editing, structure, planning, and review require distinct contexts and roles

1. What current evidence establishes. Thinking, drafting, editing, structuring, planning, and reviewing have different immediate objectives, evidence needs, and evaluation criteria. Phase separation can preserve an initial view, reduce some interference, and make independent checking possible; Agentless and human-first/reflective studies make this plausible. Role labels alone are unreliable: persona effects are inconsistent and sensitive to irrelevant attributes, and multiple agents can add coordination cost or correlated error.

2. The AIOS architectural thesis. These modes should receive purpose-built contexts, permissions, tools, and acceptance criteria. Distinction need not mean six anthropomorphic agents or impermeable silos; it means that, for example, generative writing should not silently replace source-grounded editing, and review should have enough independence to challenge the production context.

3. What follows if the thesis is substantially correct. First-order, systems could preserve intent across revisions, reduce premature convergence, and make review a genuine counter-process rather than a prompt suffix. Users could move between modes while the system composes only relevant context. Second-order, domain systems could encode reusable methods of thought and quality control without freezing work into rigid pipelines; organizations could make expertise legible as mode-specific evidence, tools, and criteria rather than only as prose instructions to one general agent.

4. Conditions, counterforces, and tests. Phase separation can lose cross-phase information, create incoherence, or provide false independence when the same model and evidence underlie every role. Generic prompted self-review is not a strong independent check; the self-correction survey finds more reliable results when feedback is external or correction behavior is specifically trained. Gains may come from extra samples, tokens, or checklists. Use factorial tests: shared versus phase-specific context; labels versus no labels; original-context self-review; fresh-context same-model review; different-model review; trained critic; deterministic validator; and human review, with equalized and natural compute. Measure errors fixed, correct work damaged, error correlation, cross-phase contamination, goal drift, revision quality, latency, and human comprehension. A distinct mode earns its place only if the operative separation adds value.

5. Bounded menus and permissions avoid both rigid pipelines and unrestricted autonomy

1. What current evidence establishes. A bounded action space can make least privilege enforceable, limit blast radius, and support provenance. AgentDojo provides direct benchmark evidence that preselecting task-relevant tools can reduce attack success. Choice and reliance studies show that users need meaningful alternatives and information, not merely a confirm button. Tool-description sensitivity shows that bounded sets remain vulnerable to wording and framing.

2. The AIOS architectural thesis. Bounded menus are a middle layer between rigid pipelines and ambient autonomy. A model can recommend or request an action within a typed, context-specific capability set; deterministic policy decides what can execute; the person can choose, inspect, revise, abstain, or authorize a newly proposed path. A user-facing menu and an executable permission boundary are separate but coordinated mechanisms.

3. What follows if the thesis is substantially correct. First-order, users could delegate flexibly without giving a model unrestricted credentials, while the system remains open to novelty through non-executable proposals. Consequential decisions could become visible, attributable, and reversible. Second-order, regulated organizations could encode permissions and approval duties independently of model prompts, and people could collaborate by granting narrow capabilities over particular files or domains rather than surrendering entire accounts or application workspaces.

4. Conditions, counterforces, and tests. Menus can omit legitimate options, anchor users, become confirmation rituals, or cause approval fatigue; open planning may outperform them on novel tasks, while fixed pipelines may outperform them on stable ones. The underlying credential boundary matters more than the visual menu. Compare fixed workflows, bounded dynamic menus, propose-new-action workflows, and unrestricted agents on routine and novelty-heavy tasks, including adversarial instructions and escalation attempts. Measure task success, novel-solution recall, unauthorized effects, approval fatigue, comprehension, reversibility, option omission, and whether overrides are practically exercised.

6. Local self-contained domain systems may support a large share of ordinary personal and organizational reasoning while strengthening privacy and sovereignty

1. What current evidence establishes. Current consumer hardware can run compressed models for selected drafting, summarization, extraction, classification, and simple question-answering tasks. PALMBench and SlimLM demonstrate on-device feasibility with significant model, quantization, device, quality, and safety tradeoffs. Domain retrieval sometimes improves knowledge-intensive answers, but retrieval is itself a reasoning bottleneck; long context can outperform RAG when resources permit, and low-quality retrieval can worsen results. If a complete inference path has no outbound transfer, source documents and prompts are not sent to a remote inference provider. That is a data-flow property, not a complete privacy, security, or legal-sovereignty guarantee. The European Data Protection Board’s Opinion 28/2024 treats anonymity and legitimate interest as case-specific legal questions, while NIST AI 100-2e2025 catalogues poisoning, extraction, prompt injection, connected-resource compromise, and supply-chain risk across the lifecycle.

2. The AIOS architectural thesis. AIOS is complementary to frontier models and centralized frontier training. It proposes that durable context, expertise, memory, workflow, and authority live primarily in local, person- or organization-controlled domain systems. Increasingly capable local models handle the large set of tasks that fit their capability envelope; remote frontier models, external search, and specialized services are invoked selectively when their marginal value justifies disclosure, cost, latency, or policy constraints. “Self-contained” denotes canonical control and a default reasoning boundary, not epistemic isolation.

3. What follows if the thesis is substantially correct. First-order, a growing share of everyday research, drafting, planning, recall, document transformation, and domain assistance could continue offline or with minimal egress, while models become replaceable components around durable person-owned knowledge. People and organizations could accumulate private cognitive infrastructure that survives a vendor, interface, or model change. Second-order, widespread domain systems could distribute practical expertise, enable privacy-preserving professional and regulated use, support underserved languages or settings with intermittent connectivity, and make advanced reasoning assistance more globally accessible where recurring remote inference is costly or unavailable. Peer collaboration could exchange files, annotations, methods, and scoped judgments without merging all canonical files and accepted ground into one provider. These are large conditional implications; no present study measures the combined longitudinal outcome.

4. Conditions, counterforces, and tests. The task distribution must be defined: “a large share of ordinary needs” depends on capability, hardware, corpus completeness, retrieval, update frequency, language, accessibility, stakes, and error tolerance. Local endpoints, backups, embeddings, update channels, and ingested documents create attack surfaces; privacy can also be supplied by on-premises or confidential remote systems. Legal sovereignty, administrative control, key ownership, substitutability, and physical location are distinct properties. Local models may remain inferior on complex synthesis or judgment, and cheap local execution can increase total use. Build a preregistered workload from consented personal and organizational task diaries and compare local small model plus retrieval, larger local model, privacy-controlled remote frontier model, and consent-based hybrid routing. Hold evidence constant in one arm, then introduce incomplete, stale, contradictory, and poisoned local corpora plus an external-search oracle in another. Measure quality, evidence entailment, evidence-gap detection, unsupported confidence, escalation recall and precision, latency, energy, cost, data egress, poisoning resistance, deletion, maintenance, accessibility, and user control over months rather than isolated prompts.

7. The architecture may materially reduce centralized application-layer infrastructure

1. What current evidence establishes. Local-first systems research shows that local replicas can support offline operation and that specified distributed invariants can be preserved, but it also shows that synchronization, conflict handling, identity, and sometimes consensus remain necessary. On-device inference substitutes for remote inference on the tasks it completes and, in some tested configurations, reduces per-query energy. Centralized inference can also become much more efficient through batching and pooling. These findings establish feasible substitution mechanisms and counterforces; no reviewed study directly measures application-layer reduction in a mature AIOS-like deployment across representative work and total infrastructure.

2. The AIOS architectural thesis. Moving canonical files, accepted ground, durable context, memory, workflow, and authority out of centralized application silos could reduce dependence on the application layer and routine remote inference, while retaining centralized frontier training, model distribution, optional frontier escalation, external data services, and whatever coordination a workflow genuinely requires. The architectural target is substitution and selective use, not total local isolation.

3. What follows if the thesis is substantially correct. First-order, many single-user and asynchronous workflows could need fewer always-on application services, less proprietary state duplication, and less routine transmission of personal or organizational context. Switching costs could fall because files and metadata remain canonical outside the service. Second-order, software markets could shift from centralized systems of record toward interoperable local intelligence environments plus optional compute and collaboration services. Organizations could reduce concentration risk and maintain continuity during provider outages; individuals and small institutions could gain capabilities now bundled with expensive platforms. Remote infrastructure demand would be reduced for workloads completed locally but retained for training, escalation, synchronization, public data, and high-capability tasks. Rebound effects and new local workloads could offset part of the reduction.

4. Conditions, counterforces, and tests. “Material” requires a named baseline, workload, reliability target, collaboration pattern, and time horizon. Local-first may shift costs into device hardware, indexes, backups, sync, schema migration, encryption keys, revocation, audit, legal hold, support, and model updates. Per-device inference can lose centralized utilization advantages, and regulated collaboration may reintroduce strong coordination. Implement matched local-first, cloud-application, and hybrid prototypes for defined workflows. Account for central requests and tokens, accelerator time, storage, egress, identity, synchronization, operator labor, recovery, security, device energy and amortization, model downloads, and remote escalation. Report eliminated, reduced, shifted, and newly created loads separately.

Stress-test verdict

The current evidence supports important parts of the architecture: system configuration materially affects model performance; selective context can outperform indiscriminate context; bounded capabilities can reduce attack success; exact state evaluation catches errors that valid model calls conceal; local models can perform selected useful tasks; and meaningful human control requires more than final approval. It does not yet evaluate the integrated AIOS proposition.

The AIOS thesis is nevertheless coherent and consequential: a complementary architecture could hold durable context, expertise, memory, workflow, and authority in local person-controlled systems; use local and frontier models according to task; separate semantic proposals from exact effects; and compose intelligence across components and human judgment. If substantially correct, it would imply not only better interaction with models but a different locus for personal and organizational intelligence, regulated use, collaboration, application dependence, and global access.

The decisive uncertainties are task coverage, retrieval integrity, the real exercise of human judgment, security of local and distributed state, coordination costs, model–scaffold interactions, economic substitution versus load shifting, and long-term maintenance. Those are conditions of the thesis and priorities for testing, not reasons to omit its implications. The strongest rival explanation is task fit: many observed benefits may come from better information selection, tool affordances, verification, or extra compute rather than the labels “agent,” “role,” “local,” or “AI-native.” A successful research program must show what the integrated architecture contributes beyond those mechanisms.

4. System-level interpretation for AIOS

The evidence supports analyzing human-AI outcomes at the system level because task allocation, context, interface, tools, validation, and human judgment materially affect performance. AIOS goes further: it describes intelligence as distributed across purpose, person-controlled files, companion metadata, annotations, relationships, context composition, bounded model judgments, exact operations, post-process integration, memory, and human authority. That particular decomposition remains an architectural thesis; adjacent system evidence does not validate each component or the combined configuration.

The following are possible connection points for later consideration, not validation:

Canonical local files and exact operations

Person-controlled canonical files can make AI outputs inspectable, revisable, portable, and contestable. Exact read, write, edit, and search operations can support provenance and cheap rollback. These are important conditions for operational agency. They do not themselves ensure that users inspect changes, understand them, or maintain skill. Evaluation should measure actual review behavior and correction quality, not the existence of an audit trail.

Model–software–human trust boundary

A possible AIOS connection is to treat the model as a semantic proposer, the deterministic layer as a reference monitor and exact-state operator, and the person as the source of purpose and consequential commitment within applicable external rules. The boundary should be explicit: a model-generated relation or classification is provisional; an authenticated identity is not proof of legitimate intent; a schema-valid action is not necessarily correct; and provenance metadata is evidence about origin, not authority. This pattern could make model error more containable and regulated review more auditable. Its value depends on complete mediation of executable paths, correct policy, legible proposals, and human checkpoints that are substantively usable.

Companion metadata and an emergent relationship graph

Metadata can expose source, confidence, status, dependency, and transformation history, making verification easier. It can also amplify hidden framing if inferred relationships are presented as settled facts. Inferred edges should remain distinguishable from authored or directly evidenced relationships, and consequential graph-derived conclusions should be contestable.

Multiple-resolution memory

Layered memory can reduce overload by showing only decision-relevant context while keeping deeper material accessible. It can also create invisible anchoring if summaries select what the user sees. Users need to inspect, correct, supersede, or exclude remembered material and to know when a response depends on a compressed representation rather than a primary source.

The Fractal Seed: Why–How–What

A recurring Why–How–What grammar could externalize purpose, method, and artifact, creating natural points for the user to state intent, assess a method, and accept an output. No located study tests this grammar. Its value should be assessed against simpler structures and on outcomes such as goal fidelity, error detection, rationale quality, and delayed independent performance.

Bounded model judgments

Bounding model judgments by domain, evidence, authority, and uncertainty is consistent with the task-conditional evidence: AI helps where its capability and information fit the task and can harm where they do not. Bounds must be legible to the user and behaviorally enforced. A label that says “bounded” without exposing assumptions or failure conditions would provide little protection.

Self-contained knowledge domains may reduce irrelevant context and make applicable evidence easier to inspect. They may also hide missing external evidence or reinforce a domain’s inherited assumptions. A boundary should therefore signal when external search, cross-domain comparison, or escalation is warranted. Whether inference runs locally or remotely changes privacy, availability, and governance properties; no reviewed evidence suggests that model locality by itself changes cognitive engagement.

Complementary local and frontier topology

The consequential thesis is not local models instead of frontier models. It is that canonical files, accepted ground, durable memory, workflow, and authority can remain local while inference is routed according to capability, sensitivity, cost, connectivity, and person consent. If this topology works across representative work, local models could absorb routine bounded tasks, frontier systems could handle harder or externally informed judgments, and only the minimum necessary context would leave the domain. This could support continuity, private cognition, model substitutability, peer exchange of inspectable knowledge artifacts, and lower dependence on any one application or inference provider. The connection remains prospective until longitudinal workload, privacy, maintenance, and infrastructure studies demonstrate those outcomes.

Human authority

Authority should be attached to operations the person can realistically exercise:

This is a stronger interpretation than “the human makes the final decision.” It makes authority observable and testable.

5. A research program for cognitively engaged AI interaction

AIOS-specific claims require direct comparison, not analogy. A useful evaluation program would manipulate interaction architecture while holding model, task, and underlying information as constant as possible.

Core experimental comparisons

  1. Foreground resolution: independently manipulate visible response length and the underlying reasoning/evidence budget. Compare short versus long foregrounds generated from the same reasoning record, then shallow versus deeper reasoning behind the same foreground. Measure comprehension, source inspection, error detection, cognitive load, later recall, and perceived authority.
  2. Judgment timing: initial user view before AI versus simultaneous AI exposure versus AI first. Measure switching, calibrated reliance, rationale quality, and transfer.
  3. Choice architecture: direct-answer default versus explicit mode choice—hint, critique, evidence, alternatives, execution—with no preselection. Measure delegation strategy and learning.
  4. Background work: synchronous visible work versus non-blocking bounded work with checkpoints versus completed-output-only return. Measure review depth, interruption cost, detected errors, rollback behavior, and review debt.
  5. Memory visibility: summary-only context versus layered memory with inspectable provenance versus full raw context. Measure anchoring, correction of stale assumptions, and source recovery.
  6. Why–How–What grammar: compare with a matched generic structured workflow. Measure goal drift, plan quality, output quality, ability to explain the result, and delayed unaided performance.
  7. Semantic/exact division of labor: model-only action versus prompted guardrail versus typed model proposal with deterministic policy and exact-state checks versus the same with substantive human commitment. Measure both semantic quality and unauthorized or incorrect effects.
  8. Mode architecture: shared context versus phase-specific contexts; labels alone versus different evidence, tools, permissions, and criteria; original-context self-review, fresh-context same-model review, different-model review, trained critic, deterministic validation, and human review. Record both errors repaired and correct work degraded, and equalize compute before attributing gains to distinct modes.
  9. Local–frontier routing: local model, local model plus domain retrieval, remote frontier model over the same authorized evidence, and consent-based hybrid routing. Add deliberately incomplete, stale, contradictory, and poisoned local corpora plus an external-search oracle. Measure evidence-gap detection, escalation recall and precision, unsupported confidence, and contamination from low-quality external retrieval as well as quality, privacy, and cost. Stratify by complexity, stakes, language, and connectivity.
  10. Longitudinal domain system: compare an AIOS-like person-controlled environment with an ordinary centralized assistant across several months of real work. Measure continuity, correction burden, evidence recovery, model escalation, migration, skill, and expert-rated decisions.
  11. Infrastructure and collaboration: implement matched local-first, centralized, and hybrid workflows, including individual work, peer exchange, and regulated organizational collaboration. Account for eliminated, reduced, shifted, and newly created computation, storage, identity, synchronization, audit, and maintenance.

Required outcome measures

Architecture studies should additionally report exact-state success, invariant and authorization violations, data egress, retrieval recall and faithfulness, model-escalation rate, offline completion, deletion and recovery, provenance completeness, central and device resource use, operator burden, interoperability, and total cost. For global-access implications, sampling must include languages, devices, network conditions, and institutions outside wealthy English-speaking settings; otherwise “accessibility” remains an assertion about deployment, not measured access.

The most important study design includes four conditions: human alone, AI alone, ordinary human-AI interaction, and the proposed AIOS interaction. Without all four, augmentation, synergy, and substitution cannot be separated.

6. Foundational lineage before August 2024

These older sources are included because they define the constructs underlying the recent work.

7. Strongest supported conclusions

  1. Confident AI advice can causally improve or impair short-task performance; users often fail to calibrate reliance to correctness.
  2. Unrestricted answer delivery can improve assisted work while reducing subsequent unaided learning in some educational contexts.
  3. Interaction design materially changes outcomes. Hints, constructive note-taking, initial human generation, decision-relevant uncertainty, visible sources, and diagnostic feedback are more promising than generic approval. Naturally occurring contradiction cues are also associated with lower wrong-answer reliance, but have not been isolated experimentally in the cited study.
  4. AI can extend cognition: structured tutors can improve immediate learning; workplace assistants can distribute tacit expertise; generative systems can improve ideation, breadth, and output quality within an appropriate capability frontier.
  5. Human-AI performance above unaided human performance is common, but performance above the better of human or AI is not. Augmentation should not be mislabeled synergy.
  6. Metacognitive failure is not created solely by AI. People can be miscalibrated without AI, while AI changes the stakes and sometimes amplifies the consequences.
  7. As an evidence-informed operational and governance conclusion, meaningful agency should include goal control, information for contestation, real override ability, and reversibility. The cited studies show the weakness of nominal final authority, but do not validate one universal agency architecture.
  8. There is not yet strong evidence of general, irreversible cognitive decline caused by generative AI.
  9. Model performance is a system property as well as a model property: context length and selection, interface, decomposition, tool surface, output constraints, and validation can materially help or harm the same model. Insufficient context is not the sole explanation for agent failure.
  10. Security and agent benchmarks support two components of a risk-containment principle: restrict executable capabilities independently of untrusted model context, and check controlled state deterministically. The full model-proposal/deterministic-enforcement/human-authority pattern has not been evaluated end to end. Deterministic enforcement does not establish semantic correctness, and human approval does not by itself establish meaningful control.
  11. Compressed local models can perform selected useful language tasks on consumer devices, and local-first software can support offline replicas. Current evidence does not establish the share of ordinary work that an integrated local domain system can satisfy; that is a consequential empirical question, not a settled negative conclusion.

8. Unresolved and contradictory evidence

9. Possible AIOS connection points for later consideration

These are hypotheses to test, not empirical validation:

10. Claims that would be unsafe to make

11. Source table

Dates are exact first-online or current-version dates where readily available; otherwise the table gives the publication year. “Peer reviewed” includes archival conference proceedings; it does not imply independent replication.

DateSource and direct linkStatusMain use in this memo
11 Jan 2026 manuscript; revised 10 Feb 2026Shaw & Nave, “Thinking—Fast, Slow, and Artificial”; PsyArXiv versionPreprint/working paper; not peer reviewed; no independent replication identifiedFocal cognitive-surrender experiments
2 Aug 2024Stadler, Bannert & Sailer, “Cognitive ease at a cost”Peer-reviewed primary studyCognitive load and reasoning quality
28 Oct 2024Vaccaro, Almaatouq & Malone, human-AI meta-analysisPeer-reviewed preregistered systematic review/meta-analysisAugmentation versus synergy
2025Bastani et al., “Generative AI without guardrails can harm learning”Peer-reviewed preregistered field RCTAssisted performance, later unaided harm, tutor guardrails
2025Lee et al., “The Impact of Generative AI on Critical Thinking”Peer-reviewed CHI study; vendor-authoredSelf-reported critical-thinking redistribution
2025Wu et al., human–generative AI collaboration, motivation, and controlPeer-reviewed preregistered experimentsAssisted performance, transfer, intrinsic motivation, perceived control
2025Kim et al., “Fostering Appropriate Reliance on Large Language Models”Peer-reviewed preregistered CHI study; partly vendor-authoredExplanations, sources, inconsistencies
2025Spatharioti et al., “Effects of LLM-based Search on Decision Making”Peer-reviewed CHI study; vendor-authoredSpeed, overreliance, token-probability uncertainty highlighting
2025De Jong et al., “Cognitive Forcing for Better Decision-Making”Peer-reviewed CSCW experimentsPartial explanations and underreliance tradeoff
2025Keke, Eisenhardt & Meske, “Thinking Twice”Peer-reviewed conference experiment; small and non-replicatedSequential decide-first interaction
9 Oct 2025 online; 2026 issueFernandes et al., “AI Makes You Smarter but None the Wiser”Peer-reviewed primary studiesAI performance and metacognitive calibration
28 Oct 2025Melumad & Yun, LLMs versus web search and depth of learningPeer-reviewed, seven experimentsSynthesis, effort, knowledge construction, advice quality
2025Kreijkes et al., LLM use, note-taking, comprehension, and memoryPeer-reviewed preregistered experimentConstructive notes and three-day learning
2025Barcaui, “ChatGPT as a cognitive crutch”Peer-reviewed RCT; high attrition; non-replicated45-day retention signal
12 Aug 2025 onlineBudzyń et al., endoscopist deskilling risk00133-5)Peer-reviewed multicenter observational studyPossible real-world skill decline
2025Kestin et al., “AI tutoring outperforms in-class active learning”Peer-reviewed randomized crossover studyStructured tutoring and immediate learning
2025Brynjolfsson, Li & Raymond, “Generative AI at Work”Peer-reviewed field studyProductivity, novice gains, possible workplace learning
2025Xi, Zhang & Wang, Socratic conversational agentPeer-reviewed 16-week experiment; 2026 issueSocratic prompting and reflective thinking
2025Fischer, Rau & Rilke, AI tutoring and reading effortPreregistered working paper; not peer reviewedCounterevidence to arbitrary delay
2025Kosmyna et al., “Your Brain on ChatGPT”Preprint; small sample; not independently replicatedTask-state EEG, recall, ownership; claim boundary
2024/2025 versionsWang et al., Tutor CoPilotPreregistered field-RCT preprint; not peer reviewedHuman-led AI tutoring and mastery
2026Tian, Amin & Yin, AI-assisted critical thinkingPeer-reviewed CHI experimentInitial judgment, reflection, over/underreliance
23 Apr 2026Spitzer et al., medical explanations and radiologist accuracyPeer-reviewed preregistered randomized studyOutput detail, expert auditing, appropriate reliance
2026Dell’Acqua et al., “The Cybernetic Teammate”Peer-reviewed preregistered field experiment; disclosed firm tiesIdeation, cross-functional breadth, human evaluation
11 Mar 2026Dell’Acqua et al., “Navigating the Jagged Technological Frontier”Peer-reviewed preregistered field experiment; BCG coauthorsTask-model fit and performance inside/outside capability frontier
2026Wong & Qiu, “Think first, ChatGPT later”Peer-reviewed experiment; non-replicatedHuman-first ideation and unaided transfer
2026Kim & Ji, exploratory search with generative AIPeer-reviewed small within-subject studyProgressive disclosure and user initiative
Jan 2026Shen & Tamkin, AI assistance and coding-skill formationVendor-authored preprint; small sample; not peer reviewedImmediate skill-acquisition risk and interaction modes
12 Jul 2024; implementation phased from 2025European Union, Regulation (EU) 2024/1689, Article 14Authoritative law; normative, not empiricalEffective human oversight and automation-bias awareness
3 Sep 2024UNESCO, AI Competency Frameworks for Students and TeachersAuthoritative normative frameworks; not causal evidenceAgency, critical judgment, accountability
2026OECD, Digital Education Outlook 2026Authoritative evidence synthesisDistinguishing task performance from learning
Feb 2026International AI Safety Report 2026Authoritative international evidence synthesisAutonomy, offloading, deskilling, and longitudinal evidence gaps
2024Hsieh et al., RULERPeer-reviewed COLM benchmark; synthetic tasks; not an organizational deploymentLong-context limits and selective context composition
2025Du et al., “Context Length Alone Hurts LLM Performance”Peer-reviewed Findings of EMNLP studyContext-length impairment after retrieval confounds are controlled
2024Yang et al., SWE-agentPeer-reviewed NeurIPS system studyTask-specific model–computer interface and scaffold benefit
2024Tam et al., structured generation and reasoningPeer-reviewed EMNLP Industry studyStrict formats, semantic reasoning, and staged conversion
2025Xia et al., AgentlessPeer-reviewed FSE system studyBounded decomposition, deterministic validation, and autonomy counterexample
20 Jul 2026 revisionBen Sghaier et al., agent-harness evolutionPreprint; not peer reviewed; one harness, one fixed-model configuration, 50 benchmark tasksFixed-model harness effects, quality regressions, and efficiency
2024Kamoi et al., LLM self-correction critical surveyPeer-reviewed TACL critical survey; synthesis rather than new experimentLimits of prompted self-review and value of external feedback
2025Luz de Araujo et al., “Principled Personas”Peer-reviewed EMNLP benchmark; no independent replication identifiedPersona and role-prompt limits
2026“Capable language models can outgrow the benefits of collaboration”Peer-reviewed Nature Machine Intelligence benchmark; too recent for independent replicationConditional multi-agent benefit and coordination overhead
2025Köbis et al., “Delegation to artificial intelligence can increase dishonest behaviour”Peer-reviewed incentivized experimentsDelegation, unethical compliance, and limits of prompt guardrails
2025Hemken et al., “Can a Large Language Model Keep My Secrets?”Peer-reviewed ACL Student Research Workshop study; synthetic and small human sampleSemantic policy interpretation versus authorization enforcement
2024Debenedetti et al., AgentDojoPeer-reviewed NeurIPS security benchmarkTask-scoped capabilities and indirect prompt injection
2025Yao et al., τ-benchPeer-reviewed ICLR benchmarkPolicy following, tool use, and exact final state
2025Faghih et al., tool-description sensitivityPeer-reviewed EMNLP benchmarkWording and choice effects within bounded tool menus
2024Li et al., RAG versus long contextPeer-reviewed EMNLP Industry benchmarkPerformance–cost tradeoffs and hybrid context routing
2025Zou et al., PoisonedRAGPeer-reviewed USENIX Security studyKnowledge-base poisoning and local corpus integrity
2025PALMBenchPeer-reviewed ICLR benchmarkOn-device feasibility, compression, latency, energy, and safety tradeoffs
2025SlimLMPeer-reviewed ACL system demonstration; one device and self-constructed datasetBounded on-device document assistance
2025BRIGHTPeer-reviewed ICLR retrieval benchmarkReasoning-intensive retrieval and domain-evidence bottlenecks
2025Fensore et al., domain RAG in consumer healthPeer-reviewed Machine Learning for Healthcare studyCounterevidence to automatic gains from curated domain RAG
2024Köhler et al., Consistent Local-First SoftwarePeer-reviewed IEEE Transactions on Software Engineering studyLocal-first invariants and selective synchronization
2024Mao et al., JanusPeer-reviewed PVLDB systems studyStronger consistency and Byzantine coordination in local-first systems
2025Cloud-versus-edge environmental case studyPeer-reviewed short ACM case study; accuracy and latency excludedPossible per-query edge savings and lifecycle-evidence boundary
2025AegaeonPeer-reviewed SOSP systems study; industry-affiliated and workload-specificCentralized pooling efficiency as an alternative infrastructure path
18 Dec 2024European Data Protection Board, Opinion 28/2024Authoritative regulatory opinion; not empiricalCase-specific anonymity, legitimate interest, and legality limits
24 Mar 2025NIST, Adversarial Machine Learning taxonomy, AI 100-2e2025Authoritative technical report; taxonomy, not comparative deployment evidenceLifecycle risks for local and connected AI systems
22 Oct 2024Chiriatti et al., “The case for human–AI interaction as system 0 thinking”Peer-reviewed correspondence/perspective; not empiricalConceptual precedent for an external AI cognitive system
1983Bainbridge, “Ironies of Automation”90046-8)Foundational peer-reviewed analysisSkill and monitoring paradox under automation
1997Parasuraman & Riley, “Humans and Automation”Foundational peer-reviewed review/frameworkMisuse, disuse, abuse, and calibrated use
2016Risko & Gilbert, “Cognitive Offloading”Foundational peer-reviewed reviewOffloading as a metacognitively controlled strategy
2021Buçinca et al., “To Trust or to Think”Foundational peer-reviewed experimentCognitive forcing and usability costs
2024, before review windowTankelevitch et al., metacognitive demands and opportunities of generative AIFoundational peer-reviewed perspective; not empiricalMetacognitive task decomposition and monitoring
2019Kleppmann et al., “Local-first software”Foundational peer-reviewed systems paper and research agendaOffline operation, local replicas, longevity, and user control