Pith. sign in

REVIEW 4 major objections 5 minor 26 references

Bot App\'etit! Exploring how Robot Morphology Shapes Perceived Affordances via a Mise en Place Scenario in a VR Kitchen

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Robot body shape changes which kitchen tasks people trust them with, a virtual-reality study suggests.

desk verdict An honest, exploratory HRI study that generates three testable hypotheses from a small but openly shared VR dataset; the quantitative grounding is thin, but the paper never overclaims and the qualitative observations carry the weight. read the letter →

arxiv 2507.19082 v1 pith:Q5TGYGLV submitted 2025-07-25 cs.RO

classification cs.RO
keywords human-robotinteractionrobotmorphologyvirtualrealityperceivedaffordancestaskdelegationkitchencollaborationthink-aloudprotocolsharedspaces
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This exploratory study asks whether a robot's visible body shape changes how people would work beside it in a kitchen. Twenty-two participants used a virtual-reality kitchen to arrange ingredients, tools, and one of eleven differently shaped robots for a mise en place task, narrating their reasoning as they went. From the resulting placements, think-aloud transcripts, and interviews, the authors put forward three testable hypotheses: people prefer to collaborate with biomorphic robots; people infer a robot's sensory abilities less from its shape than its action abilities; and people use fewer avoidance strategies around slender, less bulky robots. If these hypotheses hold in larger studies, robot appearance could be used to encourage task-sharing and comfortable co-location without changing a robot's underlying software.

What carries the argument

The central instrument is the mise en place scenario: a pre-cooking kitchen-arrangement task conducted in virtual reality, in which participants place tools, ingredients, and a static robot model that they are told to imagine as fully capable. Each robot's morphology was categorized along axes such as silhouette (humanoid-hybrid, zoomorphic, technomorphic), size, and bulkiness, and the arrangements, think-aloud protocols, and post-task interviews were annotated for factors like safety concerns, willingness to share proximity, and who was assigned each recipe subtask. This setup lets the paper measure perceived affordances through spatial decisions rather than through rating scales alone.

What would settle it

A larger follow-up in a physical kitchen with a moving robot, measuring the distances participants keep and which subtasks they delegate to robots of different body shapes, would settle the hypotheses: if bulky or non-biomorphic robots are not given fewer shared tasks, and if slender robots do not draw people closer, the claims would be contradicted.

Watch

Extended reading notes

Core claim

The paper claims that robot morphology acts as a surface for perceived affordances: humans infer from a robot's body what it can do and how it will behave, and they use those inferences when deciding whether to delegate tasks and whether to allow the robot into their personal workspace. Specifically, the authors hypothesize that biomorphic silhouettes invite task-sharing, that action-relevant body parts (grippers, arms, height) shape beliefs about motor competence more than visible sensors shape beliefs about perception, and that 'gracile' robots (slender rather than necessarily small) are avoided less than physically imposing ones. These hypotheses are grounded in observed patterns, such as robots with humanoid-hybrid or animal-like silhouettes receiving more offers to collaborate while being given fewer solo tasks, and participants arranging separate workstations to keep a buffer from bulkier robots.

Load-bearing premise

The study assumes that how people arrange a kitchen and delegate tasks in virtual reality, with a robot they are told to imagine as competent, reflects how they would actually behave in a real kitchen with a moving robot.

Editorial extensions

If this is right

  • If H1 is correct, giving a robot a biomorphic silhouette is a design lever for increasing a person's willingness to collaborate on a task.
  • If H2 is correct, visible camera and sensor placement may matter less for trust than the shape of arms, grippers, and body size, because people assume hidden sensing.
  • If H3 is correct, making a robot less bulky, even without making it smaller, should reduce avoidance behaviors and make shared workspaces more acceptable.
  • Because perceived competence and willingness to collaborate appeared to split apart, evaluating a robot's capability alone is not enough; studies should measure spatial and task-sharing choices separately.
  • The same protocol can be scaled to a larger participant pool and to moving robots to test whether the hypotheses survive outside imagination.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's account of H2 implies a possible over-trust effect: people may assume sensing abilities a robot does not actually have, so designers may need to make sensing limits visible rather than rely on morphology.
  • The avoidance differences attributed to bulky versus gracile bodies could really be driven by perceived predictability or maneuverability; a follow-up that controls for silhouette while varying only bulk, or that animates a bulky robot with graceful motion, could separate these.
  • The mise en place protocol could be recycled to test whether the type of tool (sharp versus fragile) interacts with morphology, since the paper's related work suggests knives and fragile objects produce different danger perceptions when handled by a robot.
  • The study's small sample and static robots leave open whether the biomorphic preference is specific to kitchen collaboration or generalizes to other shared manual tasks, such as assembly or cleaning.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents an exploratory VR study in which 22 participants arranged a virtual kitchen for collaborative cooking with one of 11 robots (two recipes per participant, random robot assignment), producing 3D placements, think-aloud protocols, and interview transcripts. The authors analyze task assignment frequencies per robot (Figure 5) and qualitative annotations of the think-alouds. Based on these observations, they formulate three hypotheses: H1 (more biomorphic morphology encourages task sharing), H2 (beliefs about sensory capabilities are inferred from action capabilities suggested by the body, e.g., 'if it has a hand it can probably feel by touch'), and H3 (gracile robots are avoided less than imposing ones). The paper explicitly frames the contribution as hypothesis generation, not hypothesis testing, and states that follow-up studies are needed to verify the hypotheses.

Significance. If taken as an exploratory hypothesis-generating contribution, the paper has real value: the mise en place scenario in VR is a relatively novel and ecologically plausible method for studying perceived affordances from robot morphology, the dataset is openly available, and the hypotheses H1-H3 are concrete, falsifiable, and clearly separated from confirmatory findings. The authors are commendably transparent about the exploratory nature of the work and about limitations such as static robot models and small sample size. The use of a structured annotation scheme with reported inter-coder reliability (albeit limited) and the explicit linkage to the MetaMorph taxonomy are strengths. However, because the central claim is that the hypotheses are 'grounded in observed participant behavior,' the quantitative and qualitative evidence base must be robust enough to support that grounding; the current evidence is fragile in several load-bearing places.

major comments (4)
  1. [§IV-A, Figure 5] The rank ordering in Figure 5 that motivates H1 and H3 is based on only 3-5 participants per robot (22 participants, 11 robots, 2 recipes each, no pair repeated). The paper itself says the data are 'too small to extract statistically robust conclusions' (Section IV-A), yet the hypotheses are presented as grounded in the observed patterns. A single participant's 'collaborate or not' decision changes the collaborator ratio in Figure 5c by about 0.25 for a robot seen by 4 participants. No confidence intervals, bootstrapped errors, or inferential tests are reported, and no account is taken of robot-specific confounds such as prior familiarity (e.g., Spot is a widely known commercial robot), color, or brand. To make the claimed grounding credible, the authors should either report per-robot n, raw counts, and some measure of uncertainty (e.g., bootstrap CIs), or explicitly re-label the patterns as anecdotal rather than hypothesis-grounding.
  2. [§V] The hypotheses are post-hoc interpretations of the same dataset from which they are derived. The authors state that 'we had to look at other features of the popular robots to identify what they may have in common' (Section V), which is a direct admission that H1 and H3 were selected after seeing the data. This is acceptable for an explicitly exploratory paper, but the current wording ('formulated our hypotheses') could be read as implying more independence than exists. The manuscript should clearly state, for each hypothesis, that it was generated from the observed patterns rather than predicted a priori, and should describe a specific preregistered confirmatory design that would test H1-H3 with a new sample. Without that, the hypotheses are restatements of the ranking in Figure 5 rather than generalizable claims.
  3. [§IV-B] The central support for H3 is the claim about avoidance strategies (e.g., 'participants rarely envisioned sharing proximity with the robot and often took measures to avoid it'), but the annotation data behind this claim is explicitly not shown: 'additional annotations of the TAPs (not shown here due to space constraints).' The numbers given (7 of 22 participants for the first recipe, 3 of 22 for the second) are not linked to robot identity, so the reader cannot verify the proposed association between gracile morphology and fewer avoidance strategies. Since H3 is one of the three central hypotheses, the supporting per-robot avoidance data should be presented (e.g., as a table or additional panel in Figure 5) or the hypothesis should be explicitly downgraded to an impression from the transcripts.
  4. [§III-D] The inter-coder reliability check is based on only 7 participants (participants 1-8 excluding 4), and the Cohen's kappa for the category that directly supports H2 and H3 (agent placement) is 0.62, which is generally considered moderate agreement. After this check, a single annotator coded all remaining participants. Given that the qualitative annotations carry much of the interpretive weight for H2 and H3 (e.g., judgments about sensoric vs. motoric capabilities and avoidance strategies), the reliability evidence is thin. The authors should either (a) report kappa by subcategory (e.g., safety, obstruction, sensory vs. action inferences) to show which specific codes were reliable, or (b) present a second-annotator pass on a random subset of the remaining participants. As it stands, the qualitative foundations of H2 and H3 are not sufficiently demonstrated.
minor comments (5)
  1. [Abstract] The abstract statements 'humans prefer to collaborate with biomorphic robots...' are written as assertions in the first reading; while the following sentence clarifies these are hypotheses, the phrasing should be adjusted to make the hypothetical status explicit (e.g., 'we hypothesize that humans prefer...').
  2. [§III-C] The word 'ifluenced' is a typo for 'influenced' in the sentence 'asked if and how the robot's appearance and size ifluenced its placement.'
  3. [Figure 5] Figure 5 is difficult to parse because panels (a), (b), and (c) share the same x-axis ordering (sorted by average task assignments in panel a), but the reader must infer this from the caption text. Please state explicitly in the caption that all three panels use the same x-axis ordering, or sort each panel independently for readability.
  4. [§III-B] The participant section reports 23 recruited, data from 22 after technical loss; it would be helpful to state the age range and gender split for the 22 participants whose data were analyzed (the current numbers appear to mix the full 23).
  5. [§IV-B] The observation that participants gave tasks to robots 'with no clear indication of visual sensors' (e.g., Spot) is interesting, but the term 'no clearly visible cameras' is potentially contested; consider adding a supplementary image or a MetaMorph-based sensor-feature table to make this claim verifiable.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: H1–H3 are explicitly post-hoc hypotheses, not verified predictions, and the MetaMorph self-citation is descriptive rather than load-bearing.

full rationale

The paper makes no claim to derive or verify H1–H3; it explicitly labels the study exploratory and states in the Discussion that 'we do not present findings but rather hypotheses suggested by our observations, which we intend to follow up on in larger, subsequent studies.' Section IV presents the quantitative and qualitative results as descriptive patterns ('suggestive', 'may suggest'), not as confirmatory predictions. The self-citation to the MetaMorph model [6] is used as a descriptive vocabulary for robot morphology and silhouette labels; it does not supply the content of the behavioral hypotheses, and no uniqueness or exclusion argument is imported from it. The small per-robot samples and single-annotator coding are validity limitations, not circular reductions: the hypotheses are not constructed so that the data are true by definition, and the paper openly frames them as requiring independent verification. Therefore no step in the claimed derivation chain reduces to its own inputs.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claims are empirical and rest on the assumptions above: that VR placements transfer to real kitchen behavior, that the MetaMorph taxonomy is a valid description of the robot stimuli, that participants' stated intentions reflect their true preferences, and that the annotation scheme stayed reliable after the initial reliability check.

assumptions (4)
  • domain assumption Virtual reality behavior approximates real-world behavior
    The study relies on prior citations [9]-[11] claiming VR field studies yield largely similar behavior to real settings, without re-validating in this kitchen context.
  • domain assumption The MetaMorph model is a valid taxonomy for robot morphology
    The authors use their own MetaMorph model [6] to classify the eleven robots, treating it as a reliable standard for feature description.
  • domain assumption Participants acted on the instruction that the robot can perform any recipe step
    The study design asks participants to assume full robot capability, but Section IV-B notes they instead inferred limits from body shape, so the actual assumption is a hybrid of instruction and appearance-based inference.
  • domain assumption The annotation scheme remains reliable beyond the initial eight participants
    Intercoder reliability was measured on participants 1-8 (kappa = 0.62 to 0.85), and one annotator then coded all remaining participants; this assumes no drift.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bot App\'etit! Exploring how Robot Morphology Shapes Perceived Affordances via a Mise en Place Scenario in a VR Kitchen." pith.science (2026). https://pith.science/paper/Q5TGYGLV

@misc{pith2026250719082,
  author       = {Pith},
  title        = {Pith review of: Bot App\'etit! Exploring how Robot Morphology Shapes Perceived Affordances via a Mise en Place Scenario in a VR Kitchen},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q5TGYGLV}},
  note         = {Machine review of arXiv:2507.19082}
}
read the original abstract

This study explores which factors of the visual design of a robot may influence how humans would place it in a collaborative cooking scenario and how these features may influence task delegation. Human participants were placed in a Virtual Reality (VR) environment and asked to set up a kitchen for cooking alongside a robot companion while considering the robot's morphology. We collected multimodal data for the arrangements created by the participants, transcripts of their think-aloud as they were performing the task, and transcripts of their answers to structured post-task questionnaires. Based on analyzing this data, we formulate several hypotheses: humans prefer to collaborate with biomorphic robots; human beliefs about the sensory capabilities of robots are less influenced by the morphology of the robot than beliefs about action capabilities; and humans will implement fewer avoidance strategies when sharing space with gracile robots. We intend to verify these hypotheses in follow-up studies.

Figures

Figures reproduced from arXiv: 2507.19082 by the authors.

Figure 1
Figure 1. Example arrangements produced by human participants [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The eleven robots used in this study [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Top-down view of the kitchen environment [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Study procedure steps. Dotted lines indicate steps [PITH_FULL_IMAGE:figures/full_fig_p003_4.png]
Figure 5
Figure 5. Figure 5: Robot engagement metrics normalized by number of interacting participants, sorted by avg. task assignments (5a). [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

26 extracted references · 25 canonical work pages

  1. [1]

    Robot with humanoid hands cooks food better? effect of robotic chef anthropomorphism on food quality prediction,

    D. H. Zhu and Y . P. Chang, “Robot with humanoid hands cooks food better? effect of robotic chef anthropomorphism on food quality prediction,” Int. J. Contemp. Hosp. Manag. , vol. 32, no. 3, pp. 1367– 1383, 2020

  2. [2]

    Improving human-robot collaboration via computational design,

    J. Zhi and J.-M. Lien, “Improving human-robot collaboration via computational design,” IEEE robot. autom. lett , vol. 10, no. 2, pp. 1074–1081, 2025

  3. [3]

    Effects of anthropomorphism and accountability on trust in human robot interaction,

    M. Natarajan and M. Gombolay, “Effects of anthropomorphism and accountability on trust in human robot interaction,” in Proc. of HRI

  4. [4]

    What is human-like? decomposing robots’ human-like appearance using the anthropomorphic robot (abot) database,

    E. Phillips, X. Zhao, D. Ullman, and B. F. Malle, “What is human-like? decomposing robots’ human-like appearance using the anthropomorphic robot (abot) database,” in Proc. of HRI 2018 . ACM, 2018, pp. 105– 113

  5. [5]

    Robot career fair: An exploratory evaluation of anthropomorphic robots in various career categories,

    N. L. Tenhundfeld, E. K. Phillips, and J. R. Davis, “Robot career fair: An exploratory evaluation of anthropomorphic robots in various career categories,” in Proc. of HFES 2020 . SAGE Publications, 2020, pp. 1049–1053

  6. [6]

    Metamorph – a metamodelling approach for robot morphology,

    R. Ringe, R. Nolte, N. Zargham, R. Porzel, and R. Malaka, “Metamorph – a metamodelling approach for robot morphology,” in Proc. of HRI 2025, 2025, pp. 627–636

  7. [7]

    Understanding the impact of unexpected cobot movements on human stress levels: A time series classification task,

    L. Shilon, V . Shevtsov, N. Schillreff, A. Rother, S. Mark, and M. Spiliopoulou, “Understanding the impact of unexpected cobot movements on human stress levels: A time series classification task,” in Workshop on Embracing Human-Aware AI in Industry 5.0 , 2024

  8. [8]

    The effects of overall robot shape on the emotions invoked in users and the perceived personalities of robot,

    J. Hwang, T. Park, and W. Hwang, “The effects of overall robot shape on the emotions invoked in users and the perceived personalities of robot,” Applied Ergonomics, vol. 44, no. 3, pp. 459–471, 2013

Show all 26 references
  1. [9]

    Virtual field studies: Conducting studies on public displays in virtual reality,

    V . M¨akel¨a, R. Radiah, S. Alsherif, M. Khamis, C. Xiao, L. Borchert, A. Schmidt, and F. Alt, “Virtual field studies: Conducting studies on public displays in virtual reality,” in Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems , ser. CHI ’20. New...

  2. [10]

    Levitation simulator: Prototyping ultrasonic levitation interfaces in virtual reality,

    V . Paneva, M. Bachynskyi, and J. M ¨uller, “Levitation simulator: Prototyping ultrasonic levitation interfaces in virtual reality,” in Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems , ser. CHI ’20. New York, NY , USA: Association for Computing Ma...

  3. [11]

    Multi-agent voice assistants: An investigation of user experience,

    N. Zargham, M. Bonfert, R. Porzel, T. Doring, and R. Malaka, “Multi-agent voice assistants: An investigation of user experience,” in Proceedings of the 20th International Conference on Mobile and Ubiquitous Multimedia , ser. MUM ’21. New York, NY , USA: Association for Computi...

  4. [12]

    How people perceive different robot types: A direct comparison of an android, humanoid, and non-biomimetic robot,

    K. S. Haring, D. Silvera-Tawil, T. Takahashi, K. Watanabe, and M. Velonaki, “How people perceive different robot types: A direct comparison of an android, humanoid, and non-biomimetic robot,” Proc. of KST 2016 , pp. 265–270, 2016

  5. [13]

    How anthropomorphism affects empathy toward robots,

    L. D. Riek, T.-C. Rabinowitch, B. Chakrabarti, and P. Robinson, “How anthropomorphism affects empathy toward robots,” in Proc. of HRI

  6. [14]

    Let me tell you! investigating the effects of robot communication strategies in advice-giving situations based on robot appearance, interaction modality and distance,

    M. Strait, C. Canning, and M. Scheutz, “Let me tell you! investigating the effects of robot communication strategies in advice-giving situations based on robot appearance, interaction modality and distance,” in Proc. of HRI 2014 . ACM, 2014, p. 479–486

  7. [15]

    Judging a socially assistive robot by its cover: The effect of body structure, outline, and color on users’ perception,

    E. Liberman-Pincu, Y . Parmet, and T. Oron-Gilad, “Judging a socially assistive robot by its cover: The effect of body structure, outline, and color on users’ perception,” J. Hum.-Robot Interact. , vol. 12, no. 2, Apr. 2023

  8. [16]

    Classifying human-robot interaction: an updated taxonomy,

    H. A. Yanco and J. Drury, “Classifying human-robot interaction: an updated taxonomy,” in Proc. of IEEE SMC 2004 , vol. 3, 2004, pp. 2841–2846

  9. [17]

    An extended frame- work for characterizing social robots,

    K. Baraka, P. Alves-Oliveira, and T. Ribeiro, “An extended frame- work for characterizing social robots,” in Human-Robot Interaction: Evaluation Methods and Their Standardization . Springer, 2020, pp. 21–64

  10. [18]

    A database for kitchen objects: Investigating danger perception in the context of human-robot interaction,

    J. Leusmann, C. Oechsner, J. Prinz, R. Welsch, and S. Mayer, “A database for kitchen objects: Investigating danger perception in the context of human-robot interaction,” in Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems , ser. CHI EA ’23. N...

  11. [19]

    The perceived danger (pd) scale: Development and validation,

    J. Molan, L. Saad, E. Roesler, J. M. McCurry, N. Gyory, and J. G. Trafton, “The perceived danger (pd) scale: Development and validation,” in Proceedings of the 2025 ACM/IEEE International Conference on Human-Robot Interaction, ser. HRI ’25. IEEE Press, 2025, p. 420–428

  12. [20]

    Unity: A general platform for intelligent agents,

    A. Juliani, V . Berges, E. Vckay, Y . Gao, H. Henry, M. Mattar, and D. Lange, “Unity: A general platform for intelligent agents,” CoRR, vol. abs/1809.02627, 2018

  13. [21]

    Pybullet, a python module for physics simulation for games, robotics and machine learning,

    E. Coumans and Y . Bai, “Pybullet, a python module for physics simulation for games, robotics and machine learning,” http://pybullet. org, 2016–2021

  14. [22]

    A benchmark for recipe understanding in artificial agents,

    J. Nevens, R. De Haes, R. Ringe, M. Pomarlan, R. Porzel, K. Beuls, and P. Van Eecke, “A benchmark for recipe understanding in artificial agents,” in Proc. of LREC-COLING 2024 , 2024, pp. 22–42

  15. [23]

    General attitudes towards robots scale (GAToRS): A new instrument for social surveys,

    M. Koverola, A. Kunnari, J. R. I. Sundvall, and M. Laakasuo, “General attitudes towards robots scale (GAToRS): A new instrument for social surveys,” Int. J. Soc. Robot. , vol. 14, pp. 1559–1581, 2022

  16. [2009]

    ACM, 2009, pp. 245–246

  17. [2020]

    ACM, 2020, pp. 33–42

  18. [2023]

    Available: https://doi.org/10.1145/3544549.3585884

    [Online]. Available: https://doi.org/10.1145/3544549.3585884

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.