{"id":"5883440b-e3cb-4978-82ac-985409752ab9","arxiv_id":"2412.14741","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper argues that Active Inference can serve as a unifying quantitative theory for human-computer interaction, presenting a tutorial, a taxonomy of configurations, and three illustrative scenarios, but no empirical validation.","lead":"This paper proposes using Active Inference, a neuroscience-inspired theory in which agents act to minimize surprise, as a unifying framework for human-computer interaction. It argues that modeling both users and systems as predictive agents could improve interface design and analysis.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The programmatic claims rest on an untested assumption that human users' interaction behavior is generated by expected free energy minimization; a behavioral prediction test against a simpler baseline is needed.","rationale":"The reader correctly identifies the human-AIF premise as the weakest link; I agree. The paper is internally consistent and mathematically standard, and it gains credit for a clear exposition, a useful taxonomy of interaction configurations, and concrete vignettes, but those do not substitute for validation. The central claim is programmatic, not demonstrated. A conditional verdict is appropriate: accept as a conceptual proposal contingent on empirical validation. No ad hominem; the issue is the evidence gap. One could raise the mutual-recursion issue as a technical concern, but the human-model validity is more fundamental and is the natural first checkpoint. The proposed experiment is feasible with existing tooling (pymdp/RxInfer) and directly targets Eq. (4)–(6). If the check passes, the CONDITIONAL verdict can be upgraded; if it fails, the framework's empirical relevance is in doubt.","tokens_in":36578,"tokens_out":3980,"duration_ms":31183,"concrete_test":"Conduct a controlled behavioral experiment in a canonical HCI task—e.g., noisy 1-of-N ordinal selection (cf. [135])—with at least 30 participants. Pre-register an AIF user model (generative model with observation noise; choose action distribution via Eq. (6)) and a parsimonious bounded-rationality baseline (softmax utility maximization with the same noise and task costs). Fit both models to action sequences and choice outcomes using k-fold cross-validation, and evaluate predictive log-likelihood and calibration of choice probabilities. Also probe an epistemic-action prediction: when state uncertainty is high, the AIF model predicts users should choose information-gaining actions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that AIF gives a coherent model-based theory of interaction and usable design tools—stands or falls on the empirical premise introduced in §2.4 ('AIF Human model'): a user's actions, perceptions, and preferences are governed by expected free energy (EFE) minimization under a known generative model. This premise is load-bearing for every downstream contribution: the offline simulations ((U')S and (U')(S')), the online constructions (U(S), U(I)S), the reflective models U(S(U')), and the proposed agency/engagement measures. The paper presents no test of this premise; §1.1 explicitly says it is a theory paper without implementations, evaluations, or results, and §5.2 lists modeling and preference elicitation as open problems. Without evidence that Eq. (4)–(6) describe actual user choice in an interactive task, the framework could be perfectly coherent yet empirically inert. A second, related gap is the mutual-recursion configuration: in §2.6.3 each agent's generative model nests the partner's model, and the paper only gestures at truncating prediction horizons, leaving the well-definedness of mutual AIF models unresolved. The primary concern, however, is the human-model validity: if human behavior deviates systematically from EFE-optimal action selection, the proposed user models and measures lose their foundation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a theory/position paper that proposes Active Inference (AIF) as a unifying computational framework for human–computer interaction. It reviews AIF for an HCI audience, presents a taxonomy of configurations for embedding AIF agents in the interaction loop (offline user simulation, mutual simulation, online construction, transduction, and reflective models), and discusses core elements such as Markov blankets, forward models, expected free energy, and preference priors. Three appendix vignettes (semi-autonomous driving, a soft companion robot, and an intelligent music speaker) illustrate how AIF could be applied. The central claim is that AIF provides a coherent model-based theory of interaction, supports offline design and online adaptation, and yields new quantitative measures of agency, engagement, and interaction freedom. The paper explicitly states in §1.1 that it is a theory paper without implementations, evaluations, or results.","tokens_in":36755,"tokens_out":7048,"duration_ms":52918,"significance":"If the framework were established, it would offer HCI a unified, quantitative, model-based account of the entire interaction loop, with potential benefits for simulation, adaptive interfaces, and formal analysis of agency and boundaries. The paper's strengths are its clear tutorial structure, the standard and apparently correct mathematical formulation in Appendix C, a useful taxonomy of interaction configurations in §2.6, a broad and relevant literature review, and an unusually explicit acknowledgment of its own open problems in §5. However, the load-bearing empirical premise—that human users can validly be modeled as AIF agents minimizing expected free energy—is asserted rather than tested. The paper is honest about this, but the honesty does not remove the gap between the programmatic claims and the evidence. As a conceptual proposal, the paper is valuable; as a demonstration of the claimed new measures, resilient systems, or predictive power, it is incomplete.","major_comments":[{"comment":"The central claim that AIF gives a coherent model-based theory of interaction rests on the empirical premise stated in §2.4 ('AIF Human model') that a user's actions, perceptions and preferences are governed by expected free energy minimization under a known generative model. The paper states in §1.1 that it is 'a theory paper without implementations, evaluations or results,' and §5.2 lists preference elicitation and forward-model construction as open problems. Every downstream configuration—(U')S, (U')(S'), U(S), U(I)S, and U(S(U'))—inherits this premise, but no behavioral evidence or falsifiable prediction is provided; the vignettes in Appendix D are illustrative constructions rather than tests. I ask the authors to reframe this premise as an explicit hypothesis and to specify a minimal experiment that could distinguish an EFE-based user model from a bounded-rationality or utility-maximization baseline of the kind reviewed in §A.3.3.","section":"§2.4 and §1.1"},{"comment":"The reflective and mutual configurations nest each agent's generative model inside the other's. The paper does not specify how this recursion is defined: §3.4.4 mentions that prediction horizons help terminate mutual theories of mind, but it gives no termination rule, no proof that finite-depth truncation yields a consistent AIF construction, and no statement about what happens when the two agents' models disagree. Since §2.6.3 presents reflection as a distinctive contribution, this is a load-bearing gap. Please provide a formal recursion semantics (for example, define U(S(U')) as a finite-depth construction with a fixed depth parameter) or explicitly state that reflective AIF is currently only a conceptual schema.","section":"§2.6.3 and §3.4.4"},{"comment":"The abstract and §4.3 promise 'new tools to measure important concepts such as agency and engagement,' yet no operational definition is given. The statement that AIF agents 'are directly computing their freedom to act' is not a measurement procedure; no formula for an agency or engagement metric appears in the paper, and no validation against existing instruments (e.g., subjective sense-of-agency questionnaires) is proposed. Please provide formal definitions—for example, in terms of the entropy of the policy distribution or the divergence between actual and preferred state occupancy—and describe how these metrics would be validated.","section":"§4.3"},{"comment":"Section 2.5 asserts that AIF-powered interfaces will have 'superior qualities in remaining stable and controllable,' and §3.4.3 asserts that curiosity-driven exploration 'makes a system more resilient to inter-user variability, context or varying preferences.' These are empirical claims, but no implemented system or simulation is presented to support them. The authors should either mark these as hypotheses or cite existing empirical evidence; as written, the assertions outrun the evidence supplied in the paper.","section":"§2.5 and §3.4.3"}],"minor_comments":[{"comment":"The second bullet in §2.6.1 is labelled '((U') S) Mutual interaction', which duplicates the first bullet's label; Figure 4 labels this configuration (U')(S'). Please correct the notation.","section":"§2.6.1 and Figure 4"},{"comment":"The sentence 'agents will act to change the environment in future-oriented ways which maximise Expected Free Energy' should read 'minimise', since AIF agents select actions to minimize expected free energy.","section":"§3.1.3"},{"comment":"The phrase 'an agent being a entity distinct from its environment' contains an article error; it should be 'an entity'.","section":"§2.2"},{"comment":"References [65] and [67] both cite Hornbæk and Oulasvirta's CHI 2017 paper 'What is Interaction?'; one duplicate should be removed.","section":"References"},{"comment":"The sentence 'a one-page visual summary of the core computational elements of AIF is presented in Figure 2' would read more clearly as 'Figure 2 presents a one-page visual summary...'.","section":"§2.3"},{"comment":"In the sentence 'Markov blankets could bedetected;', there is a missing space between 'be' and 'detected'.","section":"§3.2.5"}],"recommendation":"major_revision","confidential_remarks":"This is a position and tutorial paper rather than a report of novel results. The equations in Appendix C are standard, and the taxonomy in §2.6 is a useful organizing contribution. My main concern is that the paper's advertised contributions—quantitative measures of agency and engagement, resilient systems, and a validated basis for interaction theory—promise more than the programmatic content delivers; the major comments address these specific gaps. I do not see grounds for rejection, because the paper is transparent about its status and the conceptual framework is internally coherent. A major revision that explicitly repositions the empirical claims as hypotheses, adds a validation strategy, and resolves the recursion semantics would bring it to publishable standard. The paper is broad and may be better suited to a venue that explicitly welcomes theoretical or vision contributions, but that is an editorial judgment."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a theory/review paper, and it is a good one of its kind. The authors introduce Active Inference to an HCI audience with unusual clarity, and propose a taxonomy of five configurations — offline simulation, mutual interaction, construction, transduction, reflection — that should help structure future work. The math in Appendix C is standard and correctly stated, and the vignettes in Appendix D make the abstractions concrete without overclaiming. The paper is honest from the start: it says it is a theory paper without implementations, evaluations, or results.\n\nThe soft spot is exactly what the authors admit. Everything downstream — agency measures, resilience claims, design tools — rests on the premise that human users can be modeled as active inference agents, with choices governed by expected free energy minimization under a known generative model. That premise is stated in Section 2.4 and never tested. Section 5.2 correctly lists preference modeling and forward model construction as open problems, so the paper is not hiding anything. Still, the central claims are programmatic, not demonstrated. That is fine for a research agenda; it would be a problem if anyone treated the proposed agency or engagement measures as validated instruments. The reflective configurations also get short shrift formally — truncating prediction horizons is mentioned, but well-definedness of nested mutual models is left open.\n\nThere is some circularity, in that the vignettes are built to illustrate the framework rather than test it. For a tutorial that is acceptable, but it means the paper cannot independently ground the AIF-HCI approach. The literature coverage is broad and fair, and the authors are appropriately critical of the computational and modeling challenges.\n\nOverall, this deserves a serious referee. It is a well-written, well-scoped contribution that will be useful to HCI researchers who want to understand what Active Inference might offer and where the open problems are. I would not cite it as evidence that AIF works in HCI; I would cite it as the clearest map of the research space so far. Bring it to reading group if you want a lively discussion about whether AIF is a genuinely unifying theory or a repackaging of control theory with different vocabulary.","headline":"A clear and honest programmatic review that maps Active Inference onto HCI with a genuinely useful taxonomy; the untested human-model premise keeps it from being more than a research agenda, but it deserves serious engagement.","tokens_in":37335,"tokens_out":1883,"would_cite":true,"duration_ms":16816,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T05","68U35"],"pacs":[],"model":"deepseek-v4-flash","headline":"Active Inference offers HCI a unified, quantitative theory of the whole interaction loop.","keywords":["Active Inference","Human-Computer Interaction","Computational Interaction","Free Energy Principle","Expected Free Energy","Markov Blanket","Generative Models","Agency"],"falsifier":"A controlled experiment in which a real user's interaction trajectories deviate systematically and reproducibly from the predictions of a calibrated Active Inference user model—for example, choices that consistently prefer higher expected surprise over lower expected surprise under the measured preference prior—would weaken the central premise.","tokens_in":36321,"feed_emoji":"🖥️","tokens_out":1270,"duration_ms":14473,"temperature":0.7,"pith_summary":"This paper proposes that Active Inference, a computational theory in which agents act to minimize expected surprise using internal generative models, can serve as a coherent framework for human–computer interaction. The authors argue that modeling both the human user and the computer system as Active Inference agents makes the entire interaction loop—perception, action, adaptation, and even concepts like agency and engagement—explicit, probabilistic, and simulable. They present this as a response to HCI's long-standing need for theories with both high determinacy and broad scope, and they illustrate the proposal with worked vignettes in semi-autonomous driving, companion robots, and intelligent music playback. If the proposal is right, HCI would gain a single mathematical language for offline simulation, online real-time adaptation, and quantitative measures of concepts that have resisted formal treatment.","feed_headline":"One theory could unite the whole HCI loop","feed_subtitle":"Active Inference models users and systems as surprise-minimizing agents, giving HCI simulation and measures of agency.","key_machinery":"The central object is the Active Inference agent, defined by three components: a preference prior encoding goals as a probability distribution rather than a single target, a forward model predicting how the environment evolves under candidate actions, and an observation model predicting sensations from states. The mechanism that carries the argument is expected free energy (EFE), a single scalar that bounds future surprise and decomposes into a pragmatic term (agreement with preferences) and an information-gain term (expected learning), so action selection balances exploitation and exploration in one unified objective.","core_discovery":"The paper's central claim is that the human–computer interaction loop can be productively reconceived as a dyad of mutually embedded Active Inference agents, each maintaining probabilistic beliefs about hidden states, each acting to minimize expected free energy, and each treating the other as part of its environment. From this move, the authors derive a family of concrete configurations: offline simulation of user behavior, offline simulation of mutual interaction, online construction of an Active Inference system, transduction by a mediating interface agent, and reflective systems that embed a model of the user within the system's forward model. They argue that this framework gives HCI predictive power, explanatory power for boundary phenomena via Markov blankets, and evaluation tools for measuring freedom, agency, and engagement.","pith_inferences":["A testable near-term extension is the transduction configuration, where a mediating Active Inference agent sits between an existing user and system; the paper notes initial simulation results already exist, and this seems the most direct route to empirical validation.","If user behavior in interaction is well described by Active Inference, then interface designs could be ranked by the expected free energy they induce in a simulated user, turning design optimization into a search over models rather than a search over heuristics.","The framework implies that some seemingly irrational user behaviors—exploration, checking, hesitation—might be reinterpreted as rational information-gathering driven by the information-gain term, which could change how such behaviors are evaluated in usability studies."],"forward_implications":["HCI research could simulate user behavior and joint user–system behavior before building systems, testing designs in silico across diverse users and contexts.","Interactive systems could adapt in real time by reasoning over predictive models, absorbing latency and uncertainty rather than reacting to raw sensor events.","Concepts such as agency, engagement, autonomy, and freedom could receive quantitative, counterfactual measures derived from the agent's distribution over future actions and its control over the interaction loop.","Machine-learned perception models could be integrated into interaction design through a principled Bayesian structure, replacing brittle input-specific heuristics.","The Markov blanket formalism could provide an objective way to analyze where the human–computer boundary lies and how it shifts with assistive or autonomous technology."],"supporting_citations":[{"why":"Establishes the free energy principle that grounds Active Inference as the foundation for agent behavior.","marker":"[47]"},{"why":"Provides the textbook formulation of Active Inference agents, policies, and expected free energy that the paper adapts to HCI.","marker":"[105]"},{"why":"Supplies the critique of interaction theories lacking determinacy and scope that motivates the paper's proposal.","marker":"[67]"},{"why":"Supports the account of Markov blankets as flexible, dynamically reconfigurable boundaries that can extend beyond the body.","marker":"[28]"},{"why":"Offers the paper's own prior example of an Active Inference transducer for noisy ordinal selection, grounding the transduction configuration.","marker":"[135]"},{"why":"Supplies the claim that Active Inference dispenses with inverse models, a key contrast with optimal control that the paper adopts for sensing.","marker":"[111]"},{"why":"Links Active Inference to a formal sense of agency via replacing cost functions with the free energy principle.","marker":"[49]"},{"why":"Provides the levels of mutual modeling that structure the reflective and recursive Active Inference configurations.","marker":"[77]"}],"fun_headline_variants":["Active Inference unifies the HCI loop","Modeling HCI as mutual surprise minimization","A new theory for human-computer interaction","Rethinking HCI through predictive agents"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole proposal rests on treating real human users as if their perception, action, and preferences in interactive settings are governed by expected free energy minimization under a known internal generative model.","fun_headline_variants_meta":{"raw":{"variants":["Active Inference unifies the HCI loop","Modeling HCI as mutual surprise minimization","A new theory for human-computer interaction","Rethinking HCI through predictive agents"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000167,"raw_usage":{"total_tokens":1196,"prompt_tokens":823,"completion_tokens":373,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":439,"completion_tokens_details":{"reasoning_tokens":319}},"tokens_in":439,"tokens_out":373,"duration_ms":3752,"temperature":1.0,"reasoning_tokens":319,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:56:23.020556+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled experiment in which a real user's interaction trajectories deviate systematically and reproducibly from the predictions of a calibrated Active Inference user model—for example, choices that consistently prefer higher expected surprise over lower expected surprise under the measured preference prior—would weaken the central premise.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the textbook formulation of Active Inference agents, policies, and expected free energy that the paper adapts to HCI."},{"cited_title":"Williamson, and Roderick Murray-Smith","cited_arxiv_id":null,"evidence_quote":"Offers the paper's own prior example of an Active Inference transducer for noisy ordinal selection, grounding the transduction configuration."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the claim that Active Inference dispenses with inverse models, a key contrast with optimal control that the paper adopts for sensing."},{"cited_title":"The Role of Higher-Order Cognitive Models in Active Learning","cited_arxiv_id":"2401.04397","evidence_quote":"Provides the levels of mutual modeling that structure the reflective and recursive Active Inference configurations."}],"review_version":1}