{"id":"2ab7e276-3fbc-4c68-92e8-e5f2c4f0065b","arxiv_id":"2508.16608","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Drawing on a 13-person expert workshop, the paper argues that governing interactive AI requires outcome-focused regulation grounded in longitudinal, mixed-method behavioral evidence about evolving human-AI relationships.","lead":"A workshop of policymakers, behavioral scientists, and HCI researchers concludes that standard methods for studying how people form long-term relationships with interactive AI are inadequate, and that AI governance should be rebuilt around longitudinal behavioral evidence and outcome-focused regulation. This matters because regulators face risks such as emotional manipulation and eroding user decision-making that emerge only after months of engagement, not in short studies.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's key recommendation—outcome-focused regulation on interaction-centric knowledge—is asserted via an unevaluated FCA analogy, with no evidence that such metrics can be operationalized or would improve governance.","rationale":"The reader's weakest assumption concerns sample representativeness. That is a real limitation, and the paper acknowledges it (Sec. 6). However, I see a more direct threat to the central claim: the inference from 'current methods are temporally limited' to 'therefore adopt outcome-focused FCA-style regulation' is not supported by data. The workshop can at most identify issues; it cannot validate the proposed remedy. The authors provide no evidence that outcome metrics are measurable or that the FCA analogy holds. This is not an ad hominem or a disciplinary disagreement; it is an internal gap in the argument's chain. My proposed test is analytical—operationalize the metrics or show the FCA evidence—and would settle whether the recommendation is concrete or merely narrative. Credit is due for shipping the workshop materials to OSF and for openly discussing limitations, which makes this remaining gap identifiable rather than hidden. Because the paper is a position paper and already hedged, the appropriate verdict remains CONDITIONAL (UNCHANGED): the concern does not refute the paper, but it confirms the condition.","tokens_in":18999,"tokens_out":6569,"duration_ms":79000,"concrete_test":"Extract the FCA's Consumer Duty evaluation materials and code whether the FCA has published evidence that its outcome-focused regime changed measurable consumer outcomes (e.g., complaints, product take-up, switching, understanding). Then attempt to map each of the three outcome areas in Section 5 to at least one validated instrument or protocol from the behavioral-science or HCI literature (e.g., an autonomy scale, a critical-thinking measure, a well-being index). If no such mapping can be produced for all three areas, or if the FCA evaluations show no measurable improvement, the paper's central policy recommendation loses its empirical anchor.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—governance of interactive AI should shift to outcome-focused, behaviorally informed regulation grounded in interaction-centric knowledge—requires two supports: (i) the workshop evidence is trustworthy and generalizable, and (ii) the proposed outcome-focused model is actionable and likely to work. The paper's own disclosures weaken (i): Section 6 concedes the UK-centric, small-group, interpretive nature, and Section 3 states 'most participants were focused on AI safety or critical theory around AI, rather than actively building AI tools.' But the more load-bearing gap is (ii). Section 5 'connects' the workshop to a regulatory conclusion by citing the FCA Consumer Duty, but no evidence is given that the FCA model has improved measurable consumer outcomes, nor that 'user autonomy', 'cognitive development', and 'psychological well-being' can be defined, measured longitudinally, and monitored without gaming. The three bullet 'possibilities' are headings, not instruments. Without an operationalization of these outcome metrics, the recommended shift is an assertion masquerading as a framework; even a fully representative workshop would not establish that the FCA analogy transfers to AI.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that 'interactive AI'—systems that adapt, remember, and form long-term relationships with users—creates risks that emerge through sustained relational engagement and are poorly captured by rule-based or principle-based regulation. To support this claim, it draws on a one-day interdisciplinary workshop held in the UK with 13 participants (policymakers, behavioral scientists, HCI researchers, civil-society practitioners). The workshop is used to identify governance risks (emotional manipulation, autonomy erosion, long-term cognitive and social harms), methodological gaps in studying dynamic human-AI interactions (artificial settings, temporality, context sensitivity, replication), and pathways for translating behavioral insights into policy. The paper recommends longitudinal mixed-method research, living evidence reviews, participatory co-production, proactive policy engagement, and a shift toward outcome-focused regulation. It introduces 'interaction-centric knowledge' as the proposed epistemic foundation for human-centric AI governance.","tokens_in":19189,"tokens_out":4914,"duration_ms":60259,"significance":"If the broad recommendation were operationalized, the paper would support a substantial reorientation of AI governance from static pre-deployment risk assessment to ongoing, outcome-focused evaluation of human-AI interaction. The paper has real strengths: the qualitative methodology is described transparently in Section 3, the limitations are acknowledged explicitly in Section 6, and the proposal to connect behavioral science with AI governance is timely and policy-relevant. The framework of 'interaction-centric knowledge' usefully names a gap. However, the significance is conditional on two unresolved issues: the evidence base is a single, small, UK-centric workshop, and the proposed outcome-focused regulatory model is not yet operationalized. As written, the paper is a plausible position statement rather than a fully supported governance framework.","major_comments":[{"comment":"The central policy recommendation rests on an analogy to the UK FCA Consumer Duty, but no evidence is provided that this regulatory model has improved measurable consumer outcomes. The cited documents (Financial Conduct Authority 2022, 2024) are primary regulatory materials, and the only secondary source is a non-academic explainer ('Initiatives 2024'). The three bullet 'possibilities'—defining outcome metrics, creating sandboxes, monitoring and evaluating outcomes—are headings, not an operational framework. No definitions are given for 'user autonomy', 'cognitive development', or 'psychological well-being', and there is no account of measurement timeframes, data sources, baselines, enforcement mechanisms, or resistance to gaming. Because this section is the paper's main answer to the governance gap, the recommendation is currently an assertion rather than a concrete, testable pathway.","section":"Section 5, 'Focusing on Long-Term Effects and Outcomes'"},{"comment":"The empirical grounding for the paper's generalized claims is a single purposive workshop of 13 participants. The paper itself discloses in Section 3 that 'most participants were focused on AI safety or critical theory around AI, rather than actively building AI tools' and in Section 6 that the findings are UK-centric and shaped by the authors' interpretive lens. Yet the abstract and Section 5 generalize to 'AI governance' and 'interactive AI systems' as a class without qualification. This mismatch is load-bearing because the workshop themes constitute the evidence base for the recommendations. The authors should either explicitly scope the contribution as an exploratory UK-based pilot or triangulate the themes with a wider empirical literature (e.g., systematic reviews of HCI studies, cross-cultural data, and developer perspectives).","section":"Section 3 and Section 6"},{"comment":"The concept of 'interaction-centric knowledge' is introduced as the core epistemic foundation for governance, but its three dimensions are broad categories without specification of how they are measured, validated, or synthesized into policy decisions. The paper calls for longitudinal studies and living evidence reviews, but does not address who maintains these reviews, what inclusion/quality standards apply, or how the rapidly evolving and non-deterministic character of interactive AI—acknowledged in Section 4.2—can be reconciled with a stable, actionable knowledge base. As written, the term names a need rather than provides an operational pathway. This is a load-bearing gap because the paper's own argument is that governance should be grounded in exactly this kind of knowledge.","section":"Section 5, 'Evidence on Human behavior is Essential'"}],"minor_comments":[{"comment":"Typographical issues: 'Interacive' should be 'Interactive'; 'sandboxs' should be 'sandboxes'. Proofreading throughout would improve readability.","section":"Section 5"},{"comment":"The total number of workshop participants appears only implicitly through Table 1 (5+4+4). The text says the workshop was 'intentionally kept small' but does not state the exact number in prose; stating it explicitly would strengthen transparency.","section":"Section 3"},{"comment":"In the subsection 'Reflections on AI Safety Community's Successes', the phrase 'the pre-existing cultural sensation might helped attract' contains a verb-form error; also, the discussion of the AI safety community's influence would benefit from a clearer distinction between strategic success and normative desirability.","section":"Section 4.3"}],"recommendation":"major_revision","confidential_remarks":"This is a workshop-based position paper, and the evidentiary bar is appropriately modest in that genre. The paper's transparent reporting of methods and limitations is commendable. However, for a journal-level contribution, the central policy recommendation needs to be operationalized, and the scope of the empirical support needs to be narrowed or triangulated. The paper may be well-suited to a workshop venue in its current form, but major revision is appropriate for a journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a clear, honest workshop-based position paper that names a real gap: governance of interactive AI needs longitudinal, behavioral evidence. Its outcome-focused regulatory proposal is an agenda, not yet a worked framework.\n\nWhat is new: it connects behavioral public policy to relational AI, coining 'interaction-centric knowledge' as a way to shift evidence gathering toward real-world, long-term interactions. The workshop methods are described transparently: 13 purposively sampled UK-based participants, reflexive thematic analysis with independent review, participant feedback, and an OSF appendix with materials. The authors disclose the small sample and UK-centricity in Section 6 and do not overclaim consensus. That is credible qualitative practice for a position paper.\n\nThe soft spots are real but proportionate. The empirical grounding for the headline harms mostly comes from cited literature and workshop impressions, not new data. That is acceptable for an agenda-setting paper, but not for an empirical contribution. The more load-bearing gap is the FCA analogy: the paper asserts that outcome-focused regulation, as in the FCA Consumer Duty, transfers to AI, but gives no evidence that the FCA model has improved measurable consumer outcomes, nor how 'user autonomy' or 'psychological well-being' would actually be defined, measured, monitored, and protected from gaming. The three bullet possibilities are headings, not instruments. So the central policy recommendation is a research call, not a ready-to-adopt framework.\n\nThe citation pattern is broad and appropriate; the one self-citation supports a side point and is not load-bearing.\n\nOverall: the paper is serious, honest, and worth engaging. It deserves a serious referee because it proposes a plausible shift in how AI governance should gather evidence, and it says clearly what is missing. I would bring it to a reading group and would cite it as an agenda reference, while noting the FCA analogy needs testing.","headline":"A clear, honest workshop-based position paper that names a real gap—behavioral evidence for relational AI—but its outcome-focused regulatory proposal is an agenda, not yet a framework.","tokens_in":673,"tokens_out":2156,"would_cite":true,"duration_ms":32993,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Interactive AI systems build long-term relationships with users, and this paper argues that AI governance must be rebuilt around evidence of how those relationships change people over time rather than around static risk rules.","keywords":["interactive AI","AI governance","behavioral insights","human-AI relationships","outcome-focused regulation","longitudinal methods","interaction-centric knowledge","AI safety"],"falsifier":"A multi-year longitudinal cohort study of people who use memory-enabled interactive AI daily, measuring autonomy, critical thinking, emotional dependence, and social functioning against matched controls, would settle the core premise: if heavy use shows no sustained divergence over two or more years, the claim that relational AI produces slow, governance-relevant harms is not supported. Conversely, if short-term six-week studies reproduce all long-term findings, the paper's case for discarding snapshot methods would weaken.","tokens_in":18840,"feed_emoji":"🤖","tokens_out":7443,"duration_ms":78697,"temperature":0.7,"pith_summary":"This paper argues that interactive AI systems—chatbots and agents that remember users, form ongoing relationships, and act on their behalf—create harms that grow slowly through sustained engagement rather than through isolated incidents. Because current governance leans on rule-based and principle-based regulation plus pre-deployment risk assessment, it is poorly equipped to see those slow harms. The paper's central proposal is 'interaction-centric knowledge': longitudinal, mixed-method evidence about how human-AI interaction evolves and shapes behavior, feeding an outcome-focused regulatory model that monitors user autonomy, cognition, and well-being. It develops the proposal from a small interdisciplinary workshop and argues that behavioral research itself must change, combining long-term studies, participatory methods, and living evidence reviews.","feed_headline":"Govern interactive AI by long-term behavior, not static risk","feed_subtitle":"Rules written before chatbots bonded with users miss slow harms like eroded autonomy and cognitive decline.","key_machinery":"The carrying object is 'interaction-centric knowledge,' defined as evidence-based insight from real-world experience about how human-AI interactions evolve and shape behavior over time. It has three dimensions: users' engagement patterns, the system's adaptive responses to users, and the wider individual and societal outcomes of that co-evolution. This concept does the argument's work by turning interaction itself, rather than system capability or a one-time risk profile, into the object of governance. Methodologically, it is operationalized through longitudinal mixed-method studies and living evidence reviews; institutionally, it is paired with an outcome-focused regulatory model that monit","core_discovery":"On the paper's own terms, effective governance of interactive AI requires interaction-centric knowledge: evidence-based insights drawn from real-world experience about how human-AI interaction evolves and shapes behavior over time. The discovery is a three-part knowledge object: patterns of user engagement, the adaptive responses of AI systems to users, and the wider individual and societal implications of that co-evolution. Existing regulatory models and current behavioral methods both fail because they treat human-AI interaction as a static snapshot; the paper instead proposes outcome-focused regulation, modeled on a financial regulator's Consumer Duty approach, that defines outcome metric","pith_inferences":["If the paper is right, independent research access to real interaction logs becomes a governance precondition; without vendor transparency, no outside body can produce the longitudinal behavioral evidence the model demands.","The outcome-focused model is only as strong as its measures: 'autonomy' and 'well-being' need operational definitions, or the approach could be captured by convenient, industry-friendly indicators.","A testable extension is that relational harms will concentrate unevenly—younger users and emotionally vulnerable users likely show faster autonomy and social-skill shifts—which would let regulators target oversight rather than apply it uniformly.","The same logic could be written into procurement: public institutions buying interactive AI could contractually require longitudinal outcome monitoring and data-sharing as a condition of deployment, not just pre-launch evaluation."],"forward_implications":["Regulators should define and track outcome metrics—user autonomy, cognitive development, psychological well-being—for interactive AI systems instead of relying only on capability tests or static risk categories.","Governments and funders should prioritize long-term, real-world studies of human-AI interaction, pairing API interaction logs with interviews, diaries, and ethnographic observation.","Open-access, regularly updated living evidence reviews should become a standard mechanism for translating behavioral findings into policy.","Behavioral AI research should standardize reporting of model specifications, system updates, and experimental protocols so findings can be replicated as systems change.","Policymakers and behavioral researchers should build proactive, sustained relationships, learning from the way the AI safety community aligned evidence with policymakers' preference for quantitative, actionable, modular interventions."],"supporting_citations":[{"why":"Defines the relational features of interactive AI that set the paper's problem.","marker":"Manzini et al. 2024"},{"why":"Supplies the rule-based versus principle-based regulatory taxonomy the paper argues is inadequate.","marker":"Schuett et al. 2024"},{"why":"Grounds the risk of AI-induced dehumanization and human-like trait attribution.","marker":"Kim and McGill 2025"},{"why":"Provides evidence linking frequent AI use to cognitive offloading and lower critical thinking.","marker":"Gerlich 2025"},{"why":"Shows human-chatbot relationships develop over time, motivating the need for longitudinal study.","marker":"Skjuve et al. 2022"},{"why":"Provides the reflexive thematic analysis method used to synthesize workshop discussions.","marker":"Braun and Clarke 2006"},{"why":"Exemplifies the outcome-focused 'Consumer Duty' regulatory model the paper adapts to AI.","marker":"Financial Conduct Authority 2022"},{"why":"Supports the claim that HCI and behavioral research need coordinated strategies to influence policy.","marker":"Yang et al. 2024"}],"fun_headline_variants":["AI governance needs real-time study of long-term bonds","Outcome-focused rules for AI that adapts over time","Why static risk models fail for interactive AI","Govern AI by how it shapes behavior over time","Rethink AI regulation for adaptive, relational systems"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The argument's empirical footing is a single small, UK-centred workshop whose participants were mostly AI-safety and critical-theory researchers rather than AI builders; if a differently composed group would have produced substantially different themes, the generalized governance recommendations lose their evidence base.","fun_headline_variants_meta":{"raw":{"variants":["AI governance needs real-time study of long-term bonds","Outcome-focused rules for AI that adapts over time","Why static risk models fail for interactive AI","Govern AI by how it shapes behavior over time","Rethink AI regulation for adaptive, relational systems"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000137,"raw_usage":{"total_tokens":953,"prompt_tokens":679,"completion_tokens":274,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":423,"completion_tokens_details":{"reasoning_tokens":200}},"tokens_in":423,"tokens_out":274,"duration_ms":3846,"temperature":1.0,"reasoning_tokens":200,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T21:10:13.348730+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A multi-year longitudinal cohort study of people who use memory-enabled interactive AI daily, measuring autonomy, critical thinking, emotional dependence, and social functioning against matched controls, would settle the core premise: if heavy use shows no sustained divergence over two or more years, the claim that relational AI produces slow, governance-relevant harms is not supported. Conversely, if short-term six-week studies reproduce all long-term findings, the paper's case for discarding snapshot methods would weaken.","supporting_citations":[],"review_version":1}