{"id":"1ced5fb0-9c05-43e0-9755-b1a869ce4b9a","arxiv_id":"2507.07550","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A five-dimensional pluri-perspectivist model is introduced to guide context-sensitive, co-creative human-robot interaction, grounded in theory and interviews with artists and art teachers.","lead":"A position paper proposes a five-dimensional framework for how robots and humans can share viewpoints during co-creative art making, from social roles to spatial position. It grounds the framework in interviews with artists and art teachers and sketches how vision-language models could implement it.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The five-dimensional model's completeness and its benefit as a VLM prompting schema are asserted, not established: the interview coding uses the model's own categories as primary codes, and the target HRI application is explicitly untested.","rationale":"The reader's weakest assumption identifies the same load-bearing issue: the model's completeness is validated circularly by coding interview data with the model's own categories. My concern adds that the central claim about improving robot context sensitivity is also untested, since the paper explicitly states no HRI experiment was conducted. These concerns do not warrant rejection because the paper is framed as a position paper and a foundation for future work; the conditional accept already reflects the appropriate level of confidence. I therefore see no reason to change the reader's verdict, provided the final version continues to present the model as a proposal rather than an empirically validated design schema.","tokens_in":11682,"tokens_out":4172,"duration_ms":51067,"concrete_test":"Re-analyze a random sample of interview transcripts using open, inductive thematic coding without reference to the five dimensions. If coders, blind to the model, identify main themes that cannot be mapped into the five dimensions and two layers, the sufficiency claim fails, and any schema built from the model would be incomplete.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim needs the five dimensions to be a sufficient map of co-creative context and the layered structure to be usable for VLM-guided robot behavior. The interview study in Section 5 cannot establish sufficiency: the coding scheme was developed with the perspectival dimensions as primary categories, so the statement in Section 5.1 that 'no additional main categories were found' is an artifact of the coding scheme, not independent confirmation of completeness. The reported intercoder reliability (Krippendorff's alpha = 0.549) is moderate, and the explanation that categories are non-discrete does not rescue the confirmatory use of the coding. In addition, Section 7 explicitly says the model 'has not yet been tested in human-robot interactions'; no VLM or robot experiment supports the abstract's claim that using the model as an internal schema improves robot context sensitivity. The speculative scenario in Figure 3 is illustrative, not evidence. The argument is internally coherent as a design proposal, but its load-bearing assumptions are unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper introduces a layered five-dimensional model of 'pluri-perspectivism' for human-robot co-creativity with older adults. The five dimensions—social, material, temporal, semiotic, and spatial—extend Glăveanu and Gillespie's three creatogenetic differences, and the model separates a creative-task layer from a collaboration layer. The authors present a qualitative interview study with 10 older visual artists and 8 art teachers, report that the five dimensions could be used to code their practices and that no additional main categories emerged, and then propose using the model as a schema for prompting VLM-enhanced robots. A speculative human-robot drawing scenario illustrates the idea. The paper concludes that the model can guide future design and evaluation, while acknowledging that it has not yet been tested in human-robot interactions.","tokens_in":11868,"tokens_out":5943,"duration_ms":65590,"significance":"The paper's contribution is a conceptual framework that could give HRI designers a shared vocabulary for analyzing and prompting co-creative interactions. Its grounding in the creativity literature—Glăveanu and Gillespie's creatogenetic differences, Clark's joint action, and De Jaegher and Di Paolo's participatory sense-making—is independent and gives the proposal intellectual coherence. The interview material with older artists and art teachers is rich and lends plausibility to the claim that the five dimensions are present in artistic practice. The paper is also commendably explicit in Section 7 that the model has not yet been tested in HRI. However, the manuscript currently overstates the empirical support: the interview study cannot validate completeness because its coding frame is the model itself, and the VLM-prompting benefit is asserted rather than demonstrated. If the authors reframe the contribution as a design proposal with a validation agenda, the paper could be an acceptable conceptual contribution; as submitted, the abstract's claims exceed the evidence.","major_comments":[{"comment":"The statement that 'no additional main categories were found' cannot serve as evidence for the completeness of the five-dimensional model, because the coding scheme used those five dimensions as primary categories (Section 5). The absence of new main categories is therefore built into the coding protocol, not an independent result. Moreover, the two-layer task-versus-collaboration structure is not evaluated by this coding at all, since the primary categories are the dimensions rather than the layers. To substantiate the model's sufficiency, the authors should report an open-coding pass or a separate elicitation task that would allow new dimensions and layers to emerge.","section":"Section 5.1 and Table 1"},{"comment":"The paper claims the pluri-perspectivist model can be used as an internal schema to guide a VLM-enhanced robot, 'improving robot context sensitivity' (Section 1), but Section 7 states the model 'has not yet been tested in human-robot interactions' and Section 6 characterizes its application as exploratory. No VLM or robot experiment is reported; Figure 3 is an illustrative scenario. The causal benefit of the schema is therefore an untested hypothesis. The authors should either add empirical evidence (for example, a small technical probe comparing prompted versus unprompted VLM behavior) or explicitly reframe the abstract and discussion as a design proposal rather than a demonstrated improvement. The relevance to the target HRI population also remains an assumption, since none of the interview participants interacted with a robot.","section":"Abstract, Section 1, and Section 7"},{"comment":"The reported Krippendorff's alpha of 0.549 is moderate, and the explanation that the model 'does not consist of discrete or exclusive categories' addresses segment granularity but not the confirmatory use of the coding: because the same predefined categories were used to code the data, the reliability estimate cannot speak to whether the model captures the full range of relevant contextual dimensions. The paper should report how the absence of new categories could have been detected in the analysis and should temper the validation language in Sections 5.1 and 6 accordingly.","section":"Section 5, intercoder reliability"}],"minor_comments":[{"comment":"The phrase 'Kippendorf's c-a-binary' appears to be a misspelling of 'Krippendorff's alpha with binary coding'; please use the standard name and cite the exact software or formula used.","section":"Section 5"},{"comment":"The first sentence of the Proposition section reads 'a VLM-enhanced robots'; this should be 'a VLM-enhanced robot'.","section":"Section 3"},{"comment":"The model's two-layer structure is described in prose but is not included in Table 1 or made explicit in the figure's legend; adding the layer labels directly to the table would help readers connect the dimensions to the task-versus-collaboration distinction.","section":"Figure 2 and Table 1"},{"comment":"The paper explains how the model could be used for schema-guided prompting but does not include a concrete example of a prompt built from the five dimensions; a short prompt template would make the implementation proposal easier to evaluate.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The paper is clearly positioned as a position paper, so the lack of an HRI experiment is not by itself disqualifying. The larger concern is framing: the abstract and Section 1 present the model as a validated map and the VLM benefit as a demonstrated improvement, whereas the evidence supports only a plausible conceptual framework. If the authors revise to match the evidence, the paper could fit a design/position venue. The relation to the authors' own prior study in reference [8] should also be clarified so that the novel contribution of the present model is clear."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper offers something useful: a structured vocabulary for thinking about co-creativity with robots. The five-dimensional, two-layer model—social, material, temporal, semiotic, spatial dimensions split into creative-task and collaboration layers—is a new synthesis built from Glaveanu and Gillespie, De Jaegher and Di Paolo, and Clark. That synthesis is the contribution, and it makes sense. The bridge to VLM-based schema-guided prompting is also a practical step that could inform implementation.\n\nThe authors ground the model well. The speculative scenario in Figure 3 is a helpful illustration of how dimensions interleave in dialogue. And the interview data, while not validating the model, are rich enough to inspire concrete few-shot prompts.\n\nThe soft spots are real but not fatal, provided the paper is read as a position piece. The 'validation' via interviews is circular: the coding scheme used the five dimensions as primary categories, so the finding that no new main categories emerged carries little weight. The reported intercoder reliability (alpha=0.549) is moderate, and the explanation about non-discrete categories doesn't rescue the confirmatory use. The sample was artists and teachers, not older adults interacting with a robot, so relevance to the target HRI setting is assumed. And the central applied claim—that the model improves VLM context sensitivity—is untested; the paper itself says so in Section 7. That honesty is to its credit, though the abstract's 'demonstrating the potential' overstates what the interviews show.\n\nI'd take this as a conceptual proposal, not an empirical demonstration. For the HRI/co-creativity community, it's a useful map. It deserves serious peer review, with revisions that soften the validation language, address the missing supplementary materials, and make the position-paper status explicit. If I were doing work on co-creative robots I'd cite it for the framework.\n\nRecommendation: engage with it.","headline":"A useful conceptual framework for co-creative HRI, but the empirical validation is circular and the VLM benefit is asserted rather than shown.","tokens_in":12410,"tokens_out":2258,"would_cite":true,"duration_ms":24854,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A five-dimensional map tells co-creative robots what to notice","keywords":["human-robot co-creativity","pluri-perspectivism","five-dimensional contextual model","creative aging","vision-language models","schema-guided prompting","co-creative interaction model"],"falsifier":"Run a co-creative drawing session in which a VLM-enhanced robot is prompted with the pluri-perspectivist schema and compare it with a no-schema baseline; if older adults do not rate the schema-condition robot as more context-sensitive and if coded interactions show no increase in perspective taking or offering across the five dimensions, the practical claim fails. Alternatively, a fresh interview study using open coding that yields a main category outside the five would refute the model's completeness.","tokens_in":11448,"feed_emoji":"🎨","tokens_out":4212,"duration_ms":41757,"temperature":0.7,"pith_summary":"This position paper argues that pluri-perspectivism, the active exploration and integration of multiple viewpoints, is central to creative experience and should organize how robots collaborate with humans in creative tasks. It proposes a layered model with five dimensions (social, material, temporal, semiotic, and spatial) split into a collaboration layer and a creative-task layer, grounded in creativity theory and interview data from older visual artists and art educators. If the model holds, it gives robot designers and researchers a common map for designing co-creative behaviors and for prompting vision-language-model-enhanced robots to be sensitive to the creative context. The paper positions this model as a foundation for future empirical testing with older adults.","feed_headline":"Five dimensions tell co-creative robots what to notice","feed_subtitle":"A new model grounds robot context-sensitivity in social, material, temporal, semiotic, and spatial perspectives.","key_machinery":"The central object is the layered five-dimensional pluri-perspectivist model, a contextual map that defines a multidimensional co-creative space. The model works by pairing Glăveanu and Gillespie's three creatogenetic differences (social, semiotic, temporal) with two added dimensions (material and spatial), and by splitting the space into collaboration-level and task-level layers derived from Clark's theory of joint action. This map carries the argument because it simultaneously serves as an analysis scheme for coding interaction dynamics and as a prompting schema for VLM-enhanced robots.","core_discovery":"The paper's central claim is that the space of co-creative interaction can be described by five perspectival dimensions, social (self versus others), material (goals versus affordances), temporal (past versus future), semiotic (sign versus object), and spatial (here versus there), arranged in two layers: an outer collaboration layer that organizes, negotiates, and guides the process, and an inner creative-task layer that directly advances the artwork. The claim is that a VLM-enhanced robot can use this model as an internal schema for structured prompting and few-shot examples, so that it knows what contextual information to attend to, when to take a perspective, and when to offer one. Evidence comes from interviews with 10 visual artists aged 65 and older and 8 art educators, whose practices exhibited all five dimensions and no additional main categories, and from a review of robot behaviors that map onto the dimensions. The authors state the model is conceptual and has not yet been tested in human-robot interaction.","pith_inferences":["Editorial: The model also offers a concrete way to operationalize 'context sensitivity' itself, turning a vague quality into measurable categories of perspective taking and offering across five dimensions.","Editorial: Because several artists preferred solitude in early creative stages, a fully context-sensitive robot might need to modulate social engagement by creative phase, a nuance the paper hints at but does not develop.","Editorial: The two-layer structure could generalize beyond drawing to other joint creative domains such as music, craft, or writing, where coordination actions likewise tend to precede task actions.","Editorial: A direct test would compare older adults' perceived co-creativity and willingness to continue sessions with a VLM-enhanced robot prompted with the model versus one prompted with generic creative instructions."],"forward_implications":["The model can serve as a coding scheme for analyzing co-creative human-robot interactions, enabling researchers to track which perspectival dimensions are engaged and whether perspectives are taken or offered.","Schema-guided prompting using the five dimensions can improve a VLM-enhanced robot's context sensitivity in co-creative tasks.","Social and semiotic dimensions align well with text-and-image prompting, while temporal and spatial dimensions likely require additional sensor modules such as action recognition, drawing tracking, and pacing data.","The model can guide robot behavior choices, such as when to ask about preferences, reinterpret an artwork, suggest a material, or adjust turn-taking rhythm.","The speculative scenario suggests that collaboration-level perspectives often precede task-level perspectives, implying a staged or layered unfolding of co-creative interaction."],"supporting_citations":[{"why":"Supplies the three creatogenetic differences (social, semiotic, temporal) that form the model's foundation.","marker":"[26]"},{"why":"Justifies adding material and spatial dimensions through participatory sense-making and embodied interaction.","marker":"[15]"},{"why":"Provides the task versus coordination distinction that structures the model's two layers.","marker":"[14]"},{"why":"Defines creative experience with pluri-perspectivism as a key principle.","marker":"[24]"},{"why":"Reports the prior finding that a VLM-enhanced robot lacked sensitivity to the creative context, motivating the model.","marker":"[8]"},{"why":"Provides the schema-driven prompting mechanism for using the model as an internal schema.","marker":"[32]"},{"why":"Grounds the few-shot learning strategy for conditioning a VLM on structured examples.","marker":"[11]"},{"why":"Supports the integration of large language models and VLMs into robots for contextual understanding.","marker":"[57]"}],"fun_headline_variants":["Five perspectives give robots a co-creation compass","A five-axis model for context-aware co-creative robots","From elder artists to robot co-creation: a five-dimension map","Co-creative robots get a five-lens perspective model","New model maps co-creativity for robots across five dimensions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model's five dimensions and two-layer structure are assumed to be the correct and complete map of co-creative context, yet the interviews used a coding scheme built from those exact dimensions, and the participants were artists and art teachers rather than older adults actually collaborating with a robot.","fun_headline_variants_meta":{"raw":{"variants":["Five perspectives give robots a co-creation compass","A five-axis model for context-aware co-creative robots","From elder artists to robot co-creation: a five-dimension map","Co-creative robots get a five-lens perspective model","New model maps co-creativity for robots across five dimensions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000205,"raw_usage":{"total_tokens":1344,"prompt_tokens":848,"completion_tokens":496,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":414}},"tokens_in":464,"tokens_out":496,"duration_ms":6608,"temperature":1.0,"reasoning_tokens":414,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:37:27.830516+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a co-creative drawing session in which a VLM-enhanced robot is prompted with the pluri-perspectivist schema and compare it with a no-schema baseline; if older adults do not rate the schema-condition robot as more context-sensitive and if coded interactions show no increase in perspective taking or offering across the five dimensions, the practical claim fails. Alternatively, a fresh interview study using open coding that yields a main category outside the five would refute the model's completeness.","supporting_citations":[{"cited_title":"In: Rethinking creativity, pp","cited_arxiv_id":null,"evidence_quote":"Supplies the three creatogenetic differences (social, semiotic, temporal) that form the model's foundation."},{"cited_title":"Phenomenology and the cognitive sciences6, 485–507 (2007)","cited_arxiv_id":null,"evidence_quote":"Justifies adding material and spatial dimensions through participatory sense-making and embodied interaction."},{"cited_title":"Cambridge university press (1996)","cited_arxiv_id":null,"evidence_quote":"Provides the task versus coordination distinction that structures the model's two layers."},{"cited_title":"Creativity Research Journal33(2), 75–80 (2021)","cited_arxiv_id":null,"evidence_quote":"Defines creative experience with pluri-perspectivism as a key principle."},{"cited_title":"LLM-enhanced Interactions in Human-Robot Collaborative Drawing with Older Adults","cited_arxiv_id":"2506.18711","evidence_quote":"Reports the prior finding that a VLM-enhanced robot lacked sensitivity to the creative context, motivating the model."},{"cited_title":"Dialogue State Tracking with a Language Model using Schema-Driven Prompting","cited_arxiv_id":"2109.07506","evidence_quote":"Provides the schema-driven prompting mechanism for using the model as an internal schema."},{"cited_title":"Biomimetic Intelligence and Robotics 3(4), 100131 (2023)","cited_arxiv_id":null,"evidence_quote":"Supports the integration of large language models and VLMs into robots for contextual understanding."}],"review_version":1}