{"id":"4939ba75-7a76-4ea3-b289-be52e0b7a9b9","arxiv_id":"2607.05110","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Clinical social-robot personalization for children requires four data principles—integrated profiles, effectiveness signals, linkable coverage, and exposure records—because existing observational session scores cannot train recommender-style models.","lead":"Personalizing hospital social robots for children is blocked by how data is collected, not by models; the paper maps four recommender-system data problems onto clinical HRI and proposes matching collection principles. Generalists should care because robots already screen and comfort stressed kids, yet current logs cannot support real adaptation.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The strongest claim is a requirements diagnosis, not a performance guarantee. The four challenges are accurately diagnosed from the cited hospital studies (session-level scores, single-visit anonymized designs, non-randomized care). The four principles are the standard RS remedies for non-stationarity, implicit feedback, cold-start/unlinkable histories, and off-policy bias, and Table II correctly shows they cut across UP/R/RC rather than mapping one-to-one. The transferability of CF/IPW to pediatric clinical data is left open by the authors (§VI) and is therefore not a hidden premise that, if false, collapses the paper. Because the contribution is a well-supported data-collection agenda rather than a measured result, the reader’s ACCEPT / low correctness_risk / HIGH confidence verdict is appropriate and needs no adjustment.","tokens_in":9711,"tokens_out":470,"duration_ms":4572,"concrete_test":"Independently re-map each row of Table I against Principles 1–4 (integrated multi-dimensional profile, per-action effectiveness signals, linkable identity, exposure/rationale log). If any listed study already satisfies all four at action granularity, the “blocked by data” premise weakens; otherwise the diagnosis stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper is a Late Breaking Report position piece whose central claim is conceptual: existing clinical social-robot data (Table I) cannot instantiate the RS framework of [13] because of four familiar challenges, and four corresponding data principles would supply what user profiling, ranking, and responsible computing need. That claim is supported by the literature mapping and standard RS arguments; no empirical transfer result or formal derivation is asserted. The reader’s weakest_assumption (that collaborative filtering / IPW will transfer to sparse pediatric interactions) is a genuine open empirical question the authors themselves flag in §VI (“how much is enough \times scale / diversity / temporal depth”), but it is not load-bearing for the claim as stated—the claim is that the data must be designed this way, not that the models are already known to work once the data exist. No internal inconsistency, missing premise, or over-claim that would overturn the argument was found.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This Late Breaking Report argues that personalizing social robots for child well-being in clinical settings—framed as a recommendation problem via the authors’ prior RS-for-robots framework—is blocked by data design rather than by models. From a literature survey of robot-administered assessment studies (Table I), it diagnoses four challenges: non-stationary multi-dimensional child state, weak/indirect per-action feedback, sparse and unlinkable cross-session histories under anonymization, and observational (non-randomized) exposure. It proposes four corresponding data principles—integrated multi-layer profile, effectiveness signals aligned to actions, linkable representative coverage, and exposure logging at collection time—maps them to user profiling, ranking, and responsible computing (Table II), and states which subset each capability (within-session, cross-session, cold-start, cross-domain) requires. The discussion flags open practical questions of consent-constrained collection and empirical scale/diversity/depth.","tokens_in":9935,"tokens_out":1130,"duration_ms":23569,"significance":"If the diagnosis and principles hold, the paper supplies a clear, reusable checklist for data collection in personalized clinical HRI, connecting standard recommender-system problems (implicit feedback, cold-start, off-policy bias) to a high-stakes pediatric setting where existing practice yields only session-level anonymized scores. The literature mapping in Table I and the principle-to-component mapping in Table II are concrete contributions that can orient future datasets and deployments (e.g., Haru-style oncology screening). As a conceptual LBR it does not claim empirical transfer of CF/IPW to sparse pediatric interactions; that open question is appropriately deferred to §VI. The work is therefore significant as an agenda-setting data-requirements paper rather than as a validated system result.","major_comments":[{"comment":"§IV Principle 3 and §VI: Linkable coverage is presented as essential for collaborative filtering and returning-child personalization, yet the same sections (and Challenge iii) note that consent and anonymization routinely break per-child identity. The paper gestures at “privacy-preserving or federated collection” but does not specify even a minimal mechanism (e.g., site-local stable pseudonyms, cross-site federated profiles, or what may be recorded without re-identification). Because P3 is load-bearing for cross-session and cold-start claims, a short operational sketch of how linkability can be achieved under the constraints the paper itself identifies would strengthen implementability.","section":"§IV Principle 3 / §VI"},{"comment":"§V (Cold start): The suggested route of “a few probing actions, such as briefly asking the user” to obtain Principle-2 signals is in tension with the motivating setting—anxious children in oncology or procedural care—where extra questioning can itself raise burden or distress. The section should either qualify when direct elicitation is appropriate or emphasize passive implicit signals (gaze, silence, facial expression) already listed under Principle 2 as the primary cold-start path in clinical use.","section":"§V Cold start"}],"minor_comments":[{"comment":"Abstract and §I claim the principles are framed as “concrete guidelines for data collection.” The body gives clear principles and examples but not protocol-level guidelines (what fields to log, sampling rates, consent language). Softening “concrete guidelines” to “design principles / requirements” would better match the content of an LBR.","section":"Abstract / §I"},{"comment":"Table I is valuable; adding a brief note on whether any listed study logs per-action robot content (even if not used for personalization) would make the gap to Principle 2 even sharper.","section":"Table I"},{"comment":"Fig. 1 caption and body: “P1–P4 refer to the four data principles of Section IV” is clear; ensure the figure itself labels P1–P4 consistently with the principle names (integrated profile, effectiveness signals, linkable coverage, exposure record) for readers who land on the figure first.","section":"Fig. 1"},{"comment":"Author-name encoding artifacts appear in the text (e.g., “Do ˘gan”, “Do˘gan”); clean for the camera-ready version.","section":"Author block / citations"},{"comment":"§II.B: A one-sentence reminder of what “ranking” and “responsible computing” consume as inputs would help readers who have not read [13], given how heavily Table II depends on those components.","section":"§II.B"}],"recommendation":"minor_revision","confidential_remarks":"This is an LBR position piece that builds directly on the authors’ own HRI 2026 framework [13]. That is appropriate for a follow-on data-requirements note; novelty is in the clinical-child diagnosis and the four principles, not in a new model. No circularity problem: the principles are motivated by Table I and standard RS challenges, not solely by the prior framework. Fit for RO-MAN LBR is good; I would not hold it to full empirical-journal standards."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a short, clear position piece. The punchline is right: personalization for pediatric social robots is blocked by data design, not by models. Existing hospital studies (their Table I) mostly produce one anonymized session score from a standardized instrument; that cannot support per-action preference learning, cross-visit profiles, or debiasing. Mapping those four familiar RS problems onto clinical HRI and turning them into four concrete principles—integrated multi-layer profile, action-aligned effectiveness signals, linkable coverage, and exposure logged at collection time—is the real contribution. The capability-to-principle table is useful and not just restating the earlier HRI framework paper.\n\nWhat it does well: the diagnosis is grounded in the literature table rather than hand-waving, the principles are stated as collection guidelines rather than vague desiderata, and the authors are honest in the discussion that scale, diversity, and temporal depth remain open empirical questions. No over-claim that collaborative filtering or IPW already work in this domain; they only claim the data must be designed this way before those tools can be tried. Citations look solid and the self-reference to their RS framework is load-bearing but not circular—the principles are motivated independently from Table I and standard RS issues.\n\nSoft spots are minor and proportionate to an LBR. There is no empirical pilot of the logging scheme, no privacy-preserving collection protocol, and no estimate of how much data would be enough. The transfer assumption (that sparse pediatric histories can be served by similar users) is left open, which is fine for a requirements paper. Nothing load-bearing is broken.\n\nThis is for people building or evaluating clinical social robots who need to redesign data pipelines before chasing new models. Worth a serious referee at RO-MAN LBR level; I would bring it to reading group if anyone is doing hospital HRI data collection. Engage with it if that is your area; cite the principles when arguing for action-level logging.","headline":"Clean LBR that correctly diagnoses why clinical robot data cannot train personalization and gives four usable collection principles.","tokens_in":10520,"tokens_out":484,"would_cite":true,"duration_ms":5320,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Personalization for social robots that support children in clinics is blocked by how data is collected, not by the models; four data principles can supply what is missing.","keywords":["social robots","personalization","recommender systems","child well-being","clinical HRI","data collection principles","user profiling","exposure bias"],"falsifier":"Instrument a multi-site pediatric robot deployment with the four principles, train collaborative-filtering and inverse-propensity models on the resulting action-level logs, and test whether the ranked actions improve measured child anxiety or engagement relative to existing scripted or single-score baselines; clear failure of transfer would falsify the claim that the data principles alone unblock personalization.","tokens_in":10625,"feed_emoji":"🤖","tokens_out":1043,"duration_ms":16204,"temperature":0.7,"pith_summary":"Social robots are already used in hospitals to comfort and screen children, but effective support must adapt to each child and to how that child changes from moment to moment and visit to visit. The paper treats choosing the best robot action as a recommendation problem and shows that a recent recommender-system framework—user profiling, ranking, and responsible computing—could in principle deliver that personalization. What blocks it is not the model but the data: existing studies reduce a whole session to one anonymized score on a single construct, so they cannot track a shifting state, recover which actions helped, link sparse visits, or correct for the fact that care is never randomized. The authors answer with four collection principles: an integrated multi-layer profile, per-action effectiveness signals (both explicit and implicit), linkable coverage across children and visits, and an exposure record of why each action was chosen, logged at the moment of selection. If these guidelines are followed, the same framework can begin to personalize within a session, across return visits, and even for children seen only once.","feed_headline":"Robot personalization is blocked by data, not models","feed_subtitle":"Clinical child-wellbeing robots need action-level logs, not session scores, to learn what works for each child","key_machinery":"The four data principles (integrated profile, effectiveness signals, linkable coverage, exposure record). They convert the four familiar recommender problems—non-stationary user state, weak implicit feedback, cold-start and unlinkable histories, and off-policy bias—into concrete collection guidelines that supply what user profiling, ranking, and responsible computing require.","core_discovery":"Instantiating a recommender-system framework for personalizing social-robot actions in child well-being is blocked not by the model but by the data. Existing hospital studies yield only fixed, single-construct, end-of-session scores that cannot track a shifting child state, recover per-action feedback, link sparse visits, or correct for non-random assignment. Four data principles—an integrated profile, effectiveness signals, linkable representative coverage, and an exposure record logged at collection time—directly answer these four challenges and map onto the framework’s profiling, ranking, and responsible-computing components.","pith_inferences":["Hospitals that adopt action-level logging with exposure records would create datasets usable for offline policy evaluation, reducing the need for new randomized trials each time a robot behavior is proposed.","The same four principles could transfer to other high-stakes, sparse-interaction domains such as elderly care or special education where randomization is ethically limited.","Privacy-preserving or federated collection will almost certainly be required to reconcile the demand for linkable profiles with the anonymization constraints the paper itself flags as open.","Empirical thresholds for scale, diversity, and temporal depth remain untested; small pilots that instrument only integrated profiles and effectiveness signals could already enable basic within-session ranking."],"forward_implications":["Data collection for clinical social robots must log one record per robot action rather than one score per session.","Within-session personalization needs per-action signals plus population coverage; cross-session personalization additionally needs stable linkable identities.","Exposure (why an action was chosen) must be recorded at the moment of selection or it cannot be recovered later for debiasing.","Capabilities such as cold-start and cross-domain transfer can run on subsets of the four principles rather than requiring all of them at once.","Whether social robots can personalize therefore turns less on new models than on collecting the right data from the start."],"fun_headline_variants":["Robot personalization blocked by data not models","Four data principles unlock child social-robot personalization","Shifting child states demand integrated profiles and action signals","Sparse visits and observational bias stall robot recommenders","Effectiveness signals and exposure logs needed for child robots"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The claim rests on the premise that once multi-dimensional, linkable profiles and per-action signals exist, ordinary recommender techniques will usefully transfer preferences across sparse, high-stakes pediatric clinical interactions the way they do in everyday recommendation settings.","fun_headline_variants_meta":{"raw":{"variants":["Robot personalization blocked by data not models","Four data principles unlock child social-robot personalization","Shifting child states demand integrated profiles and action signals","Sparse visits and observational bias stall robot recommenders","Effectiveness signals and exposure logs needed for child robots"]},"model":"grok-4.5","effort":"low","cost_usd":0.003294,"raw_usage":{"total_tokens":1123,"prompt_tokens":813,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":32940000,"prompt_tokens_details":{"text_tokens":813,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":236,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":813,"tokens_out":74,"duration_ms":2775,"temperature":1.0,"reasoning_tokens":236,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T08:40:32.358627+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Instrument a multi-site pediatric robot deployment with the four principles, train collaborative-filtering and inverse-propensity models on the resulting action-level logs, and test whether the ranked actions improve measured child anxiety or engagement relative to existing scripted or single-score baselines; clear failure of transfer would falsify the claim that the data principles alone unblock personalization.","supporting_citations":[],"review_version":1}