{"id":"c67947aa-988c-4926-aaf1-39025e37e8b6","arxiv_id":"2501.00867","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Higher education should be redesigned around dialogical interaction with LLM agents, with the aim of developing a new 'interactional intelligence' skill set.","lead":"This paper proposes a new educational framework, Interactionalism, in which students learn by conversing with AI agents rather than by passive lectures, and it introduces 'interactional intelligence' as a teachable skill set. The authors offer a blueprint for redesigning curricula, selection, and assessment, but provide no experimental evidence.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Transfer assumption is load-bearing and untested: Tables 1-3 map LLA-design tasks onto skill labels without psychometric evidence, and the proposed LLA-transcript evaluation risks circularity.","rationale":"Read in good faith, the paper is a coherent conceptual blueprint rather than an empirical theory; it explicitly disclaims being a theory of learning and acknowledges the absence of interactional skill ontologies and measures. That transparency is commendable, but it does not supply evidence for the causal claim at the center of the proposal. The single most load-bearing assumption is that the metacognitive and meta-emotional operations engaged by LLA design and prompting are the same operations required for competent human-human and human-AI interaction, and that practicing them in the LLA modality transfers to those settings. Tables 1-3 are rational reconstructions, not validated psychometric mappings. The paper's own evaluation proposal—scoring transcripts of interaction with a dialogical agent—makes the intervention and the outcome same-modality, so observed improvements could be specific to interacting with LLMs rather than evidence of generalizable metaskills. This is not an internal inconsistency; the argument is structurally sound but empirically underdetermined. The reader's UNVERDICTED verdict is the honest status, so no verdict change is needed. The concern is worth flagging because any adopter of the framework should require transfer evidence before redesigning curricula around it.","tokens_in":12919,"tokens_out":3903,"duration_ms":39916,"concrete_test":"Run a preregistered within-institution experiment (N ≥ 120 per arm): Intervention arm completes a semester-long LLA design/prompting curriculum; control arm completes a matched content curriculum without agent design. Before and after, measure (a) metacognitive self-regulation with an established instrument, (b) emotional-regulation and perspective-taking scales, and (c) blind-rated performance on human-human collaborative tasks (e.g., a dyadic negotiation and a peer-teaching episode) and on a novel human-AI task. The central claim is supported only if the intervention arm shows significantly larger gains than control on the human-human measures and on the novel human-AI task after controlling for baseline scores. If gains appear only on LLA-specific tasks, the transfer assumption is refuted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central causal claim is that the metacognitive and meta-emotional demands of prompting, testing, and architecting LLM agents are developmental opportunities that build 'interactional intelligence,' and that this skill transfers to human-human and human-AI work. The load-bearing assumption is that the latent skill exercised when, e.g., 'specifying the objective function for an LLA' (Table 2) is the same latent skill needed to manage a human collaborator's goals, and that practice on the former improves the latter. This is asserted by mapping tasks onto skill-class labels, but no psychometric or empirical evidence is offered for the mapping's validity. The paper itself concedes: 'We do not have good interactional and dialogical skill ontologies – let alone measures we can reliably gauge learner progress against' (Section 'But was not Cognition Always-Already Interactional?'). Moreover, because the proposed evaluation is a transcript of interaction with a dialogical agent, the intervention and the outcome are measured in the same modality, risking circularity: improved LLA-interaction transcripts may reflect AI-specific strategies (e.g., prompt-formatting heuristics) rather than generalizable metaskills. If the transfer assumption fails, the architectural redesign recommendation—making the learning process 'a model of the desired skills'—has no empirical footing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces 'Interactionalism' as a set of design principles for higher education in the era of large language models (LLMs), arguing that the tasks involved in prompting, testing, and architecting LLM-based agents can develop an 'interactional intelligence' composed of metacognitive and meta-emotional skills. The authors propose restructuring learner selection, learning experiences, and evaluation around dialogical interactions with these agents, so that 'the learning experience embodies the learning objectives.' The paper is explicitly positioned not as a theory of learning but as a practical blueprint and research agenda, with conceptual decompositions of the target skills and mappings from LLM-agent design tasks to those skills.","tokens_in":13141,"tokens_out":3739,"duration_ms":36092,"significance":"If the central claim were established, this framework would be significant for both educational practice and human-AI interaction research: it offers a coherent vocabulary for a class of skills that many argue are increasingly valuable, and it proposes a concrete architectural response to the assessment-integrity crisis that GenAI poses to traditional education. The paper's strengths include its explicit recognition of the current absence of interactional skill ontologies and measures, its design-oriented focus on existing technologies, and its clear articulation of a research agenda. However, the manuscript currently asserts, rather than demonstrates, the load-bearing causal claim that practicing with LLM agents builds generalizable interactional skills; this is a proposal in need of empirical validation, not yet a finding.","major_comments":[{"comment":"The central claim that 'working with Large Language Model (LLM)-based agents can be proactively used to help develop learners' is asserted without empirical evidence or a comparison baseline. The mapping from LLA-design tasks to skill classes in Tables 2 and 3 is a plausible taxonomy but not a demonstration of construct validity or of transfer to human-human or human-AI contexts. The paper itself concedes, in the section 'But was not Cognition Always-Already Interactional?', that 'We do not have good interactional and dialogical skill ontologies – let alone measures we can reliably gauge learner progress against.' This concession undercuts the strength of the developmental claim. Please either provide evidence (even indirect) for the transfer assumption, or explicitly reframe the manuscript as a proposal with testable hypotheses and a validation plan.","section":"Abstract and 'Interactional Intelligence'"},{"comment":"The proposed evaluation method, which replaces the individually produced artifact with 'a transcript of the interaction with the dialogical agent' as the basis for evaluation, risks circularity: the intervention (interacting with an LLA) and the outcome measure (a transcript of exactly that interaction) share the same modality. Improved transcripts may reflect AI-specific strategies (e.g., prompt-formatting heuristics) rather than generalizable interactional intelligence. To support the transfer claim, the paper should specify outcome measures that are external to the LLA-interaction itself, such as human-human collaboration tasks, standardized negotiation or teaching scenarios, or pre/post psychometric instruments.","section":"Dialogical Learning for Interactional Skill development"},{"comment":"The construct 'interactional intelligence' is not defined operationally enough to be falsifiable. The paper states that meta-cognitive and meta-emotional components 'can be understood, quantified, measured and developed' but does not provide a measurement model, item-level definitions, or reliability/validity evidence for the skill classes in Table 1 or the mappings in Tables 2 and 3. As a result, the central claim that these skills are 'articulable and measurable' (Conclusion) is circular: the authors define the construct through the tasks that allegedly exercise it, and then use those same tasks as evidence of the construct's existence. A clearer separation between the construct definition, the measurement instrument, and the intervention tasks is needed.","section":"The Metacognitive and Meta-Emotional Demands... and Tables 1-3"},{"comment":"The manuscript oscillates between modest disclaimers and strong causal assertions. For example, it states that Interactionalism 'is not advanced as a theory of learning' but also that the learning method 'embodies the learning objectives' and that designers of LLA tasks 'require' a 'dense and proactive exercise of meta-cognitive functions.' These claims are not inherently contradictory, but the paper does not reconcile them: if this is a blueprint, then the developmental claims should be framed as hypotheses to be tested, not as established outcomes. Please clarify the epistemic status of each major claim, especially the developmental and transfer claims.","section":"Entire manuscript"}],"minor_comments":[{"comment":"There are several typographical errors that should be corrected: 'platrforms' (Abstract), 'abnd' (Section 'GenAI in Education'), 'expeirnce' (Section 'Interactional Intelligence'), 'escalatre' (Section 'But was not Cognition Always-Already Interactional?'), and 'which This closely mimics' (Section 'Learner Evaluation').","section":"Abstract and throughout"},{"comment":"The reference list is inconsistent with in-text citations. For instance, 'Johnson, Manyika and Yee, 2007' appears in text but the reference is dated 2005; 'Darvishi, et al, 2024' is cited but missing from the reference list; 'Moldoveanu and Djikic, 2017' and 'Deming, 2017' are cited but not listed; and 'Mercier and Sperber, 2012' is cited in text while the reference list gives 2011.","section":"References"},{"comment":"Table 2's header 'Sample instance from large Language Agent Design' has inconsistent capitalization; consider harmonizing table formatting and ensuring all table entries are complete sentences.","section":"Tables"},{"comment":"The sentence beginning 'A key component of an interactionalist approach to learning is a recognition and sharp definition of dialogical agents (DA's) that capture with fidelity and nuance the conversational and interactional structures of teaching and learning' could be split for clarity, as it currently conflates the definition of dialogical agents with their role in learning.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is more of a white paper or programmatic essay than a conventional research article. Its central contribution is conceptual and would be valuable if explicitly framed as a research agenda with falsifiable hypotheses. However, in its current form the causal and transfer claims are insufficiently supported, and the evaluation proposal is circular. The heavy reliance on the authors' own prior publications for key constructs (e.g., 'interactional intelligence', 'dialogical grammar') is acceptable but worth checking for novelty relative to those works. The paper also shows signs of incomplete editing (typos, missing references), which suggests it may be a preprint not yet ready for formal publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a clearly argued, honestly scoped position paper. It identifies a real problem—GenAI breaks the assessment model of individual, monological skill demonstration—and proposes a coherent redesign built around dialogical learning with LLM agents. It is not an empirical paper, and the authors do not pretend otherwise: they state explicitly that we lack good interactional and dialogical skill ontologies and reliable measures. That admission is the paper's most useful honesty and also the hinge on which everything turns.\n\nWhat is actually new is small but real. The named framework 'interactionalism' and the specific decomposition in Tables 1–3, mapping meta-cognitive and meta-emotional skill classes onto LLA design tasks such as prompt refinement, objective specification, and evaluation-rubric design, are not present in the cited literature. The underlying dialogical tradition (Vygotsky, Mercier and Sperber) is old, but the translation into concrete instructional moves for agent-work is a useful contribution, and the tables give designers something to argue with.\n\nThe soft spots are proportionate to the genre. The load-bearing claim is that practicing LLA design develops transferable 'interactional intelligence' that carries to human-human and human-AI work. That claim is asserted through the mapping tables, not tested. There is no psychometric evidence that the latent skill behind 'specifying an objective function for an LLA' is the same as the skill behind managing a human collaborator's goals. The proposed evaluation—transcripts of interaction with a dialogical agent—risks circularity: the intervention and the outcome live in the same modality, so improved transcripts could just be better prompt-formatting heuristics. The authors also cite their own prior work heavily, which is not disqualifying when the cited work is substantively relevant, but the construct 'interactional intelligence' is anchored mainly in those self-citations rather than in independent measures.\n\nThis paper deserves a serious referee. It would be a waste to desk-reject it: the argument is internally coherent, the limitations are stated in the text itself, and the design tables are concrete enough to generate testable hypotheses for learning-science and HCI researchers. A referee should require the authors to separate the programmatic vision from the empirical claims and to specify what evidence would falsify the transfer assumption. I would not cite it as evidence of transfer; I would cite it, if at all, as a representative position piece.","headline":"A clearly argued, self-aware position paper that names a real problem and offers a concrete blueprint, but the transfer claim at its core is asserted, not demonstrated.","tokens_in":13683,"tokens_out":2830,"would_cite":false,"duration_ms":28798,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Working with large language agents can be deliberately designed to develop metacognitive and meta-emotional skills, and higher education should reorganize around that insight.","keywords":["interactionalism","interactional intelligence","metacognition","meta-emotional skills","large language model agents","dialogical learning","higher education redesign","generative AI in education"],"falsifier":"Train one group to design and refine LLM agents and another on equivalent content through traditional methods, then measure both groups on established metacognitive and social-emotional instruments in human interaction tasks; if the agent-design group shows no advantage outside AI-specific tasks, the central recommendation collapses.","tokens_in":12695,"feed_emoji":"🎓","tokens_out":8447,"duration_ms":71403,"temperature":0.7,"pith_summary":"Interactionalism argues that the rise of generative AI, especially dialogical agents built on large language models, turns the act of learning itself into a new kind of skill: interactional intelligence. The paper tries to establish that designing, prompting, testing, and refining LLM agents is not just a way to get work done but a deliberate exercise of metacognitive and meta-emotional abilities, from task decomposition to monitoring one's own objectives to inferring a user's emotional state. If that holds, universities could redesign the entire learning architecture—selection, instruction, feedback, assessment—so that the method of learning models the interactional skills graduates will need in AI-mediated workplaces. The value would be a higher education system whose process, not just its content, produces the 'meta-human' skills the labor market increasingly rewards.","feed_headline":"Prompting AI agents could become a core college skill","feed_subtitle":"A proposed redesign makes higher learning itself a practice of interactional intelligence","key_machinery":"The load-bearing mechanism is the mapping, in the paper's tables, between a decomposition of metacognitive skills (monitoring, registering, and controlling one's own mental events) and meta-emotional and meta-relational skills (inferring and measuring emotional states, interactive reasoning), on one side, and the concrete tasks of designing LLM agents on the other. Writing and refining prompts, specifying objective functions and personas, chaining tasks, setting evaluation rubrics, and testing a user's likely responses are each claimed to require rapid switching between planning, monitoring, and control, and between cognitive and meta-cognitive frames. The complementary mechanism is the design principle that the learning experience should embody the learning objectives: interactional and dialogical skills are learned by exercising them in the fabric of the course itself.","core_discovery":"The paper's central claim is that metacognitive and meta-emotional skills—not raw recall or calculation—are the bottleneck for productive human-AI collaboration, and that these skills can be developed by engaging learners in the design and use of LLM agents. It breaks interactional intelligence into named components such as self-explicitation, task specification, task decomposition, sub-task energization, task-performance evaluation, task switching, partial-credit assignment, and objective and task refinement, then maps each onto concrete agent-design tasks, for instance specifying an agent's objective function, chaining sub-tasks, designing evaluation rubrics, and testing a user's likely responses. The same mapping is argued to hold for meta-emotional and meta-relational skills such as inferring intentionality and designing for emotional connectedness. The paper then proposes replacing individual one-shot productions like essays, calculations, models, exegeses, and verifications with evaluated transcripts of interactions with dialogical agents, and making the learning environment an always-on, agent-mediated dialogue. It explicitly does not claim to offer a theory of learning; it offers a blueprint for practice.","pith_inferences":["If the transfer premise holds, the same agent-design tasks could be used inside organizations as low-cost developmental exercises for the soft skills that current training programs struggle to build, turning routine AI use into upskilling.","The framework implies a testable psychometric program: build an interactional skill ontology and check whether performance on agent-design tasks predicts performance in human team settings; the paper stops at mapping tables rather than delivering such measures.","The paper's de-emphasis on 'know-what' leaves an open tension: interactional skill may depend on a base of domain knowledge, so a working redesign may need to preserve some monological content, a question the paper does not address.","A natural next experiment is to compare interaction-transcript evaluation with expert human judgment of the same artifacts to see whether transcripts carry incremental information about a learner's future workplace performance."],"forward_implications":["Higher education could shift from one-shot individual assessment to evaluation of interaction transcripts, scoring learners on the quality of their questions, challenges, and refinements rather than only on final artifacts.","Course architecture could become a network of always-on dialogical agents serving as tutor, teaching assistant, evaluator, guide, and mentor, making the large one-to-one tutoring advantage seen in classic studies achievable at scale.","Admissions and learner evaluation could include chained, multi-shot, meta-cognitive questioning that probes how a learner knows, not just what she knows.","The design goal for educational interfaces would invert: instead of reducing the metacognitive load of generative AI, educators would deliberately compose tasks that raise it as a developmental opportunity.","The skill that a degree certifies would shift from individual know-how to interactional know-how, changing what credentials signal to employers."],"supporting_citations":[{"why":"Documents the metacognitive demands and opportunities of generative AI, which the paper reframes as developmental rather than as load to be reduced.","marker":"[Tankelevitch et al., 2024]"},{"why":"Provides field evidence that AI creates a jagged and volatile capability frontier, motivating the need for meta-level human skills.","marker":"[Dell'Acqua et al., 2023]"},{"why":"Supplies the large one-to-one tutoring effect that dialogical agent-mediated learning is meant to make affordable at scale.","marker":"[Bloom, 1984]"},{"why":"Establishes the rising labor market value of social and interactional skills, which the paper's proposed skill set is designed to meet.","marker":"[Deming, 2017]"},{"why":"Provides the author's prior framework for interactional and soft skills that Interactionalism extends into learning design.","marker":"[Moldoveanu, 2024]"},{"why":"Offers evidence that programming with AI involves interactional skills that extend beyond collaboration with human partners.","marker":"[Sarkar et al., 2022]"},{"why":"Grounds the claim that thinking is internalized dialogue, which justifies treating dialogical interaction as a site of cognitive development.","marker":"[Vygotsky, 1981]"},{"why":"Supplies the definition of metacognitive activities as monitoring and controlling one's own mental states that the skill decomposition builds on.","marker":"[Stuss, 2011]"}],"fun_headline_variants":["Metaskills over memorization: a blueprint for AI-era learning","Teaching interactional intelligence: new design for higher ed","AI agents in class: a new blueprint for college learning","Rethink higher learning: prioritize meta-cognitive skills with AI","New framework: Interactionalism for learning with LLM agents"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the metacognitive and meta-emotional moves practiced while designing and prompting LLM agents are the same skills that make people effective in human-human and human-AI work, and that practice with AI transfers to those settings; the mapping tables assert this correspondence but offer no empirical or psychometric validation.","fun_headline_variants_meta":{"raw":{"variants":["Metaskills over memorization: a blueprint for AI-era learning","Teaching interactional intelligence: new design for higher ed","AI agents in class: a new blueprint for college learning","Rethink higher learning: prioritize meta-cognitive skills with AI","New framework: Interactionalism for learning with LLM agents"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001092,"raw_usage":{"total_tokens":4517,"prompt_tokens":861,"completion_tokens":3656,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":3572}},"tokens_in":477,"tokens_out":3656,"duration_ms":23398,"temperature":1.0,"reasoning_tokens":3572,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:40:28.416953+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train one group to design and refine LLM agents and another on equivalent content through traditional methods, then measure both groups on established metacognitive and social-emotional instruments in human interaction tasks; if the agent-design group shows no advantage outside AI-specific tasks, the central recommendation collapses.","supporting_citations":[{"cited_title":"The learning experience embodies the learning objectives","cited_arxiv_id":null,"evidence_quote":"Provides the author's prior framework for interactional and soft skills that Interactionalism extends into learning design."}],"review_version":1}