{"id":"65849ccd-146f-4e01-92d6-ca2bed66f4a2","arxiv_id":"2506.14968","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"FEAST is a mealtime assistance robot that uses LLM-editable behavior trees and modular tools to let care recipients personalize feeding, drinking, and mouth wiping in real home settings.","lead":"FEAST is a robot system that helps people with disabilities eat, drink, and wipe their mouths at home, and lets users personalize its behavior using plain language. Researchers tested it with two co-designer care recipients over five days and with an occupational therapist, finding that users could adapt the system to their needs and preferences.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"In-the-wild personalization evidence comes only from two co-author care recipients; naive end-user success is unverified.","rationale":"The reader identified LLM reliability as the weakest assumption, and that is a genuine concern. However, the paper's own data show the LLM error was caught and corrected through transparency, and no unsafe behavior resulted; the system is explicitly designed around fallible LLMs. The more load-bearing gap is the participant sample. The central claim is about \"individual care recipients\" and \"in-the-wild\" personalization, yet the only in-the-wild users are co-authors. Insider status can confound all reported successes: they know the interface, the failure modes, and how to phrase requests. The paper is commendably explicit about this limitation, but it remains a threat to the abstract's generalization. The 18 interventions and 20 explanations further show that fully autonomous user-driven personalization was not actually observed; some interventions were manual parameter patches. I therefore agree with the conditional verdict but for a different primary reason. A follow-up with naive care recipients is the decisive test.","tokens_in":36348,"tokens_out":4336,"duration_ms":40372,"concrete_test":"Recruit at least three care recipients who were not involved in FEAST's development, and have each independently complete at least two full meals in their own homes, following only the provided instruction manual and with experimenters intervening only for safety-critical issues. Log (a) adaptability request success rate on first attempt, (b) number of requests requiring experimenter repair, (c) meal completion success, and (d) safety incidents. Compare these to Table II for CR1 and CR2. If naive users show materially lower request success or need frequent engineer repairs, the abstract's blanket personalization claim should be restricted to co-designed deployments.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that end users can personalize FEAST in-the-wild without engineer involvement. The only in-the-wild evidence is from CR1 and CR2, who are community researchers and co-authors deeply involved in the system since November 2022. Their familiarity likely reduces the difficulty of issuing and debugging personalization requests; the paper itself acknowledges this in Section VIII. Objective logs weaken the \"without engineer involvement\" reading: across six meals there were 18 experimenter interventions and 20 explanations, and some interventions were direct engineering fixes to personalization state, e.g., Meal ID 6 \"Bug in behavior tree param... manually updated to 100\" and Meal ID 4 \"Had to code back continuous transfer in (experimenter fault).\" The OT evaluation uses a non-developer, but it is a controlled, scenario-based assessment, not an in-the-wild meal, and the OT is not a care recipient. Therefore the evidence base for \"individual care recipients\" personalizing in-the-wild is a sample of two insiders. If success depends on insider familiarity, the central claim fails to transfer to the target population. The LLM misinterpretation in Section VI-D is real but secondary: transparency and safety checks allowed recovery, so it does not by itself falsify the personalization claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"FEAST is a mealtime-assistance robot system that combines modular hardware (feeding, drinking, mouth-wiping tools), a web-based interface, head-gesture and button inputs, and parameterized behavior trees that can be adapted through natural language via a large language model. The authors conduct a formative study with 21 care recipients to identify personalization needs, then evaluate FEAST in a five-day in-home study with two care recipients (who are also community researchers and co-authors), spanning six meals in personal, TV-watching, and social contexts. They also report an evaluation with an occupational therapist unfamiliar with the system. The paper claims that FEAST can be personalized in the wild to meet the unique needs of individual care recipients, with low workload, high user acceptance, and broad coverage of the adaptability requests identified in the formative study.","tokens_in":1794,"tokens_out":1955,"duration_ms":78823,"significance":"If the claims hold, FEAST is a meaningful step towards deployable, personalized mealtime assistance: it is one of the first systems to integrate feeding, drinking, and mouth wiping with open-ended, natural-language personalization, and it reports real in-home meals across varied contexts. The paper is unusually transparent in its appendices, which document per-meal experimenter interventions, personalization requests, LLM responses, and OT scenarios; this level of detail is a genuine strength for reproducibility and for understanding failure modes. The community-based participatory research methodology is appropriate for assistive robotics, and the authors are candid about the insider status of the two care recipients. However, the evidence base for the central 'personalized in-the-wild' claim is narrow: only two co-designer participants, an acknowledged limitation in Section VIII, and the abstract's wording overstates what the data can support. The 'outperforms a state-of-the-art baseline' claim is also based on a coverage count rather than a direct comparative evaluation.","major_comments":[{"comment":"The abstract claims that FEAST 'outperforms a state-of-the-art baseline limited to fixed customizations,' but the only support is the coverage count in Section V-A (36 of 46 request types for FEAST vs. 17 for Nanavati et al. and 15 for FLAIR). This is a self-scored capability comparison based on the authors' categorization in Table III, not a head-to-head user evaluation. The OT study in Section VII compares personalized FEAST against the default no-personalization FEAST, not against another system. The abstract and Section V-A should be reworded to say 'covers more of the personalization request types identified in our formative study' rather than 'outperforms,' unless a direct comparative study is added.","section":"Section V-A, Figure 5, and abstract"},{"comment":"The statement that restricting node additions to 'pause,' 'wait for gesture,' and 'retract,' and confining parameter changes to predefined domains, 'guarantees safety' is not established. The static checks verify syntactic conformance to a whitelist and parameter bounds, but they do not prove that every accepted combination of nodes and parameters is physically safe across all user states, food configurations, and environmental contexts. The safety hardware and monitoring features in Section V-C are reasonable mitigations, but 'guarantee' is too strong. I recommend replacing 'guarantee' with 'mitigate' and adding a hazard analysis or a systematic failure-mode evaluation to support the safety claim.","section":"Section IV-B (Safety Checks) and Section V-C"},{"comment":"The claim that users can personalize FEAST 'without engineer involvement' is not fully supported by the reported study. There were 18 experimenter interventions and 20 explanations across six meals, and some interventions directly modified personalization state: Meal ID 6 reports 'Bug in behavior tree param... manually updated to 100' and Meal ID 4 reports 'Had to code back continuous transfer in (experimenter fault).' The presence of experimenters who can edit behavior trees is reasonable for a prototype study, but the paper should explicitly report which personalization operations were completed end-to-end by the participant alone and which required experimenter action, and it should temper the 'without engineer involvement' phrasing accordingly.","section":"Section VI-C, Table II, and Appendix E"},{"comment":"No aggregate success rate is reported for the central natural-language-to-behavior-tree personalization mechanism. The appendix shows multiple requests that were partially applied, misapplied to the wrong set of tools, or rejected as invalid (e.g., the 'Use button when completing a transfer when taking a sip' request initially changed all tools in Meal ID 1, and several requests in Meal IDs 3 and 6 required iterative rephrasing). Since reliable LLM translation is load-bearing for the personalization claim, the paper should quantify, from the existing logs, the first-attempt success rate, the fraction of requests requiring user correction, and the final success rate after transparency-based iteration.","section":"Section VI-D, Lesson 2, and Appendix E"},{"comment":"The only in-the-wild personalization evidence comes from CR1 and CR2, who are community researchers and co-authors deeply involved in the system's design since 2022; the occupational therapist is not a care recipient and was evaluated in a controlled scenario, not in the home. The paper acknowledges this in Section VIII, but the abstract's unhedged claim that FEAST 'can be personalized in-the-wild to meet the unique needs of individual care recipients' goes beyond the evidence. I recommend softening the central claim to describe a feasibility demonstration with two co-designer care recipients, and to state more prominently that transfer to non-expert end users remains untested.","section":"Section VI-A and Section VIII"}],"minor_comments":[{"comment":"The safety standard number is inconsistent: Table I references ISO 13482, while Section V-C and Appendix C repeatedly use 'ISO 13842.' The correct designation for 'Robots and robotic devices — Safety requirements for personal care robots' is ISO 13482; please correct all occurrences.","section":"Section V-C and Appendix C vs. Table I"},{"comment":"The NASA-TLX description in Section VI-B says participants completed a 7-point Likert scale, while Section VI-C reports workload on a 0-100 scale. Please clarify the mapping or make the scales consistent.","section":"Section VI-B and Section VI-C"},{"comment":"The statement that drink acquisition averaged 100.0 ± 0.0% across Meal IDs 2-5 is slightly confusing because Meal ID 6 also had successful drink acquisition (2/2 in Table II). Consider reporting across Meal IDs 2-6 to avoid an apparent omission.","section":"Section VI-C, Figure 7"}],"recommendation":"major_revision","confidential_remarks":"This is a solid systems paper with unusually honest and detailed appendices, and the community-based participatory research framing is appropriate for HRI. The core concern is that several abstract-level claims ('outperforms a state-of-the-art baseline,' 'guarantees safety,' 'without engineer involvement') are stronger than what the evidence supports. These are fixable by recalibrating the language and by extracting quantitative success metrics from the existing logs, so I do not think rejection is warranted. For the journal's scope, the paper would be stronger if the authors reframe the contribution as a feasibility demonstration with expert co-designers rather than a generalizable end-user evaluation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"FEAST is a genuine step forward for assistive feeding. It is the first system I know of that couples LLM-editable parameterized behavior trees with modular tools for feeding, drinking, and mouth-wiping, and it adds user-synthesized gestures on top of that. The formative study with 21 care recipients is a real contribution, and the five-day in-home deployment—six meals across personal, TV, and social contexts—is more than most systems papers in this area attempt. The paper is also unusually honest: it documents LLM misinterpretations, a robot fall, aborted meals, and experimenter interventions in the appendices, and it has an explicit limitations section. Open-sourcing the hardware and software is the right move.\n\nThe main soft spot is exactly what the stress-test flags: the in-the-wild evidence for 'personalization by care recipients' comes from two people who are co-authors and have been co-designing the system since November 2022. Their familiarity likely lowers the cognitive workload and makes it easier to recover from bad personalization requests. The paper acknowledges this in Section VIII, but the abstract still says 'care recipients successfully personalize FEAST,' which overstates the generality. Relatedly, the claim that FEAST 'outperforms a state-of-the-art baseline' is supported by a coverage analysis against Nanavati et al. and FLAIR and by a gesture-detection ablation with emulated users—not by a controlled comparison with care recipients. That is a coverage argument, not an empirical win.\n\nThe 18 experimenter interventions and 20 explanations across six meals, including one manual behavior-tree parameter fix and one 'experimenter fault' note, further weaken the 'without engineer involvement' reading. None of this kills the paper: FEAST is a systems contribution, and the engineering is credible. The LLM errors are not fatal because the transparency loop let users recover; that is actually the paper's best lesson. The OT evaluation is scenario-based rather than in-the-wild, but it does show a non-developer can drive the personalization pipeline, which is useful evidence.\n\nWho should read it: anyone building physically assistive robots with user personalization, and HRI reviewers interested in how LLM-based adaptation can be constrained for safety. It deserves peer review, with the expectation that the authors temper the generalization claims and ideally add at least one non-co-designer care recipient in a future study.","headline":"A credible systems contribution with real in-home deployment, but the personalization evidence is thinner than the abstract suggests; worth refereeing with expected revisions.","tokens_in":37118,"tokens_out":2840,"would_cite":true,"duration_ms":27776,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A feeding robot that care recipients can personalize in plain language, through mid-meal requests, custom gestures, and buttons, completed all six in-home test meals across three dining contexts.","keywords":["mealtime assistance","assistive robotics","in-the-wild personalization","behavior trees","large language models","human-robot interaction","in-home evaluation","community-based participatory research"],"falsifier":"Run the paper's 46 formative request types through the LLM personalization pipeline and count the updates that pass static validation yet change the wrong nodes or parameters; if that misapplication rate is high enough that users cannot reliably detect and correct errors within a meal, the in-the-wild claim fails. The documented 'use the button only when sipping' incident, where the LLM changed every tool, is the observed instance this test generalizes.","tokens_in":36136,"feed_emoji":"🤖","tokens_out":11807,"duration_ms":108287,"temperature":0.7,"pith_summary":"The paper sets out to establish that in-the-wild personalization of mealtime assistance is practical: care recipients themselves, not engineers, can adapt a feeding robot to their individual needs and preferences during real meals. FEAST combines modular hardware that swaps between feeding, drinking, and mouth-wiping tools with a web interface, head gestures, and physical buttons, and lets users change how the robot behaves through natural language. The adaptation is carried by parameterized behavior trees, where an LLM turns a request into structured updates that are statically checked for safety and summarized back to the user for transparency. Evidence includes a formative study with 21 care recipients, a five-day in-home evaluation in which two care recipients finished six meals across personal, TV, and social contexts with low reported workload, and an evaluation by an occupational therapist unfamiliar with the system. A sympathetic reading is that FEAST shows end users can tailor a multi-task care robot to their own bodies, habits, and settings.","feed_headline":"Feeding robot adapts to plain-language requests in home meals","feed_subtitle":"Six in-home meals, three settings: users tuned speed, gestures, and bite transfer by talking to the robot.","key_machinery":"The load-bearing mechanism is the parameterized behavior tree: each skill (pick up tool, acquire bite, transfer, retract) is a behavior tree whose nodes and parameters carry human-readable names and bounded domains, such as $Speed \\in \\{low, medium, high\\}$ or $TimeToWaitBeforeAutocontinue \\in [5, 100]$. This representation makes LLM-driven personalization safe and inspectable: a natural-language request is translated by an LLM into structured tree updates, which are statically validated for valid names and in-domain values, with new nodes restricted to three empirically tested types: 'pause,' 'wait for gesture,' and 'retract.' A second LLM pass summarizes each change in plain language, and transparency queries are answered from the same tree encodings plus logs of node execution, perception, and safety checks. Around this, an automated-planning (PDDL) planner sequences skills to reach user goals, and custom gesture detectors are synthesized by LLM code generation, validated against a few user-recorded positive and negative examples.","core_discovery":"On its own terms, the paper's discovery is that in-the-wild personalization can be achieved by making every behavior change a small, checkable edit to a structured representation. The robot feeds, serves drinks, and wipes mouths with a single arm that swaps tools, and users steer it with requests such as 'feed me as fast as you can,' 'dip the strawberry deeper into the whipped cream,' 'do not show continue pages,' and 'move to retract position after every bite.' The paper's central empirical claims are that both care recipients completed three meals each in diverse in-the-wild contexts with few experimenter interventions, reported low cognitive workload on NASA-TLX, rated FEAST at or above 4 out of 5 on the Technology Acceptance Model, and said FEAST gave them more control and a stronger sense of independence than their human caregivers. It also claims that when the LLM misapplied a request, changing button-use for transfer completion across all tools instead of only sips, the transparency features let the user detect and iteratively correct the error.","pith_inferences":["Editorial inference: the safety design, restricting what an LLM may change at the representation level rather than trusting its output, transfers to any assistive or domestic robot whose policy is edited by non-experts; the open question is whether a three-node whitelist stays expressive enough as user requests diversify.","Editorial inference: the paper reports LLM misapplications qualitatively but not their rate; a precision measurement over the 46 request types (updates that pass static checks yet change the wrong nodes or parameters) would set a concrete reliability target for the pipeline.","Editorial inference: because the two in-home participants were co-developers of the system, the occupational-therapist evaluation is the strongest current evidence for unfamiliar-user usability; a longitudinal study with naive care recipients would test whether the low workload scores survive outside co-design.","Editorial inference: the gesture-synthesis ablation (F1 score of about 0.9 with personalized parameters versus 0.63 without) suggests a broader design lesson: for users with limited mobility, interaction modalities are better synthesized from a few personal examples than chosen from a fixed library."],"forward_implications":["If the central claim holds, care recipients can reconfigure a feeding robot's speed, transfer style, confirmation behavior, and interaction modality from meal to meal, without an engineer present.","Personalization would extend beyond feeding to drinking, mouth-wiping, retracting between bites, and custom gesture control, capabilities the paper says prior systems with fixed customizations lack.","The static safety net (whitelisted node types plus bounded parameter domains) means safe adaptation does not require flawless language understanding, only errors rare enough and transparent enough for users to catch.","The reported workload figures (mean NASA-TLX around 7–22 versus a literature baseline of 37) and Technology Acceptance Model scores at or above 4 out of 5 are the paper's evidence that in-the-wild adaptation does not impose an unusable cognitive burden.","A corollary the paper draws is that transparency must be reachable throughout the meal: in the study, users asked the experimenters questions that the transparency page could have answered had it been accessible from every screen."],"supporting_citations":[{"why":"FEAST's bite acquisition skills, including skewering, scooping, twirling, dipping, and bite-order prediction, build directly on the FLAIR framework.","marker":"[20]"},{"why":"Supplies the inside-mouth bite transfer method and head-pose perception pipeline for users who cannot lean forward.","marker":"[11]"},{"why":"Supplies the default outside-mouth bite transfer strategy and the analysis that transfer depends on acquisition.","marker":"[23]"},{"why":"The out-of-lab feeding robot system whose fixed-customization approach serves as the comparison baseline for adaptability coverage.","marker":"[36]"},{"why":"The NASA-TLX workload instrument that produced the reported low-workload scores.","marker":"[37]"},{"why":"The Technology Acceptance Model survey behind the high usability and acceptance ratings.","marker":"[38]"},{"why":"The ISO 13482 personal-care-robot safety standard whose principles the safety layer claims to follow.","marker":"[39]"},{"why":"The IEEE 7001 transparency standard whose five levels structure the transparency evaluation.","marker":"[40]"},{"why":"GPT-4o, the LLM that performs request translation, transparency summarization, and gesture-program synthesis.","marker":"[98]"}],"fun_headline_variants":["Robot helps eat, drink, and wipe mouth via plain-language requests","Home robot personalizes feeding, drinking, and wiping via your words","LLM-powered robot serves meals, drinks, and wipes mouths on command","In-the-wild mealtime assistant personalizes via spoken requests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the LLM will usually turn a user's natural-language request into the behavior-tree change the user intended; the study itself shows this failing when a request to use the button 'only when taking a sip' was applied to all tools, so the in-the-wild claim depends on such errors being rare enough, and visible enough through transparency summaries, that users catch and correct them.","fun_headline_variants_meta":{"raw":{"variants":["Robot helps eat, drink, and wipe mouth via plain-language requests","Home robot personalizes feeding, drinking, and wiping via your words","LLM-powered robot serves meals, drinks, and wipes mouths on command","In-the-wild mealtime assistant personalizes via spoken requests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001217,"raw_usage":{"total_tokens":5079,"prompt_tokens":1092,"completion_tokens":3987,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":708,"completion_tokens_details":{"reasoning_tokens":3911}},"tokens_in":708,"tokens_out":3987,"duration_ms":30134,"temperature":1.0,"reasoning_tokens":3911,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:09:02.520595+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's 46 formative request types through the LLM personalization pipeline and count the updates that pass static validation yet change the wrong nodes or parameters; if that misapplication rate is high enough that users cannot reliably detect and correct errors within a meal, the in-the-wild claim fails. The documented 'use the button only when sipping' incident, where the LLM changed every tool, is the observed instance this test generalizes.","supporting_citations":[{"cited_title":"User acceptance of computer technology: A comparison of two theoretical models,","cited_arxiv_id":null,"evidence_quote":"The Technology Acceptance Model survey behind the high usability and acceptance ratings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The ISO 13482 personal-care-robot safety standard whose principles the safety layer claims to follow."},{"cited_title":"Ieee standard for trans- parency of autonomous systems,","cited_arxiv_id":null,"evidence_quote":"The IEEE 7001 transparency standard whose five levels structure the transparency evaluation."}],"review_version":1}