Pith. sign in

REVIEW 4 major objections 5 minor 76 references

Sense and Sensibility: What makes a social robot convincing to high-school students?

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper claims that a social robot's arguments moved 117 of 139 dissenting high-school answers, 34 to wrong answers, and that displayed certainty is a major driver of this alignment.

desk verdict Solid event-level evidence on robot persuasion, but the headline 75% 'beyond expected capacity' is likely inflated by an IRT guessing-parameter mismatch. read the letter →

arxiv 2506.12507 v1 pith:4UXPRB3R submitted 2025-06-14 cs.RO

classification cs.RO
keywords socialrobotseducationalroboticsconformitypersuasioninformationaltrustcertaintycuesLLMexperienceitemresponsetheory
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that high-school students are strongly persuaded by a social robot's arguments on a familiar school subject, electric circuits, even when the robot is wrong. In 117 of 139 cases where a student's preliminary answer disagreed with the robot, the student changed the final answer to match it, and 34 of those switches were to incorrect answers. Using an item-response-theory model of each student's expected performance, the authors report that 75% of students finished beyond their expected capacity, better on questions where the robot was right and worse where it was wrong, and interpret this as direct robot influence. They also report that displayed certainty matters: students aligned 94.4% of the time with a certain robot, 82.6% with a neutral one, and 71.4% with an uncertain one. Students who reported more experience with large language models were more likely to follow the robot's incorrect answers.

What carries the argument

The argument is carried by three linked devices. The first is an item-response-theory model, specifically the three-parameter logistic model, fitted to the students' diagnostic answers to estimate each student's latent ability and each question's difficulty, which defines the expected capacity that the 75% claim is measured against. The second is the alignment variable, a tri-state coding of each interaction event as agreement, resistance, or change toward the robot, with Monte Carlo simulations and Fisher's method used to decide when observed performance is beyond expectation. The third is the multimodal certainty manipulation, differences in wording, speech rate, pauses, gaze, smiles, and head movements, that defines the Certain, Neutral, and Uncertain conditions and produces the graded alignment rates.

What would settle it

Give a matched group of students the same eight true/false questions twice with no robot, or with a recorded voice instead of a present robot. If the switch rate and the rate of beyond-expected performance are as high as in the robot condition, the paper's causal reading fails; if they are substantially lower, the robot-influence claim survives.

Watch

Extended reading notes

Core claim

The central discovery is that a social robot does not need to be right to be followed: it persuaded a large majority of 40 high-school students to revise their answers on eight true/false electric-circuit questions, including when its argument was deliberately wrong on the two easiest questions. The authors frame this through the distinction between sense, the students' reliable ability to judge the arguments, and sensibility, their responsiveness to the robot's expressed certainty. They find that 75% of students performed beyond their IRT-predicted capacity, above expectation on the non-deceptive questions and below expectation on the deceptive ones, and that this shift from preliminary to final answers should be read as a direct result of the robot's influence. Displayed certainty was a decisive cue: a robot portrayed as certain drew alignment in 94.4% of disagreements, an uncertain one in 71.4%, and students rated the certain robot as most convincing. Prior experience with large language models increased alignment with incorrect answers, suggesting that familiarity with AI can raise susceptibility rather than critical resistance.

Load-bearing premise

The load-bearing premise is that the item-response-theory model fitted to the students' own diagnostic answers correctly predicts how they would have scored without the robot, and that students would not have switched answers merely from being asked again.

Editorial extensions

If this is right

  • Educational robots that argue for answers can override students' own reasoning on familiar material, so designers should treat persuasion as a safety-relevant property rather than a side effect.
  • Displayed certainty is a practical control knob: tying a robot's confidence signals to the actual reliability of its content could reduce acceptance of wrong information.
  • Students with heavy large-language-model experience appear more, not less, vulnerable to an AI's wrong answer, so AI-literacy curricula should address trust calibration rather than just tool competence.
  • Because alignment rates did not depend on measured ability, adjusting question difficulty or selecting strong students will not by itself prevent over-alignment.
  • The absence of carry-over effects suggests each robot-student exchange is its own persuasion event, so interventions to reduce overtrust may need to operate question by question.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: a matched no-robot control, with students answering the same eight questions twice, would isolate how much of the 117/139 switch rate is a retest or reflection effect rather than robot causation.
  • Beyond the paper: because the expected-capacity baseline is fitted to the same small cohort, the 75% figure could be rechecked with item parameters estimated from a larger independent sample of the same questions.
  • Beyond the paper: if the large-language-model experience effect is causal, the same susceptibility may appear with text-based chatbots, which could be tested by replacing the embodied robot with a screen agent while keeping the arguments identical.
  • Beyond the paper: a practical design rule follows, that robots should display certainty calibrated to their information's reliability, which future systems with LLM reliability metrics could implement directly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper reports an experiment with 40 Swedish high-school students who interacted one-on-one with a Furhat robot solving eight true/false electric-circuit statements. The robot argued for the correct answer on six questions and the incorrect answer on two, and it displayed three levels of certainty (Uncertain, Neutral, Certain) across groups. The main reported results are that students aligned with the robot in 117 of 139 cases of initial disagreement, including 34 cases in which the robot was wrong; that alignment was higher for Certain (94.4%) than for Neutral (82.6%) and Uncertain (71.4%); that 75% of students performed “beyond their expected capacity” as defined by a 3PL IRT model calibrated on a diagnostic test; and that self-reported LLM experience was associated with greater alignment on Deception questions. The paper interprets these results as evidence of robot-induced informational influence and discusses implications for educational robotics and AI trust.

Significance. The event-level alignment counts and the certainty manipulation are valuable and largely defensible: the certainty portrayals were validated in separate rating studies, the GLMM analyses control for ability and difficulty, and the post-test self-reports are consistent with the observed alignment patterns. If the central performance-deviation claim could be placed on a sound footing, the paper would make a strong contribution to educational robotics and to the study of trust and overtrust in AI. As it stands, however, the headline “75% beyond expected capacity” result depends on model-based expected probabilities that appear miscalibrated for the binary response format, and it relies on an in-sample calibration with no no-robot control condition. The causal interpretation of the PA-to-FA changes is therefore not yet fully supported, and the abstract-level claims go beyond what the current design can establish.

major comments (4)
  1. [Sec. III-A and Sec. IV-B] The 3PL model in Sec. III-A fixes a guessing lower asymptote c >= 0.25 because the diagnostic items had three options, but the experimental items are true/false (Sec. III-D, Fig. 2). For a two-option item, the chance-level success probability is 0.5, so p(theta) = c + (1-c)/(1+e^{-a(theta-b)}) yields expected correct probabilities that are too low for low-ability students. The Monte Carlo simulation and Fisher combined-p-value procedure in Sec. IV-B will therefore classify final-answer outcomes as “beyond expectations” too liberally; applying the same biased model to preliminary answers does not cancel the bias, since both PA and FA are evaluated with the same miscalibrated expectation. The 75% claim is not robust until the expected probabilities are re-estimated with a lower asymptote appropriate for binary items, or the model is recalibrated on the experimental response format.
  2. [Sec. IV-B and Sec. II] The interpretation of PA-to-FA changes as “a direct result of robot’s influence” is not warranted by the design. Students answer each question twice, and the second answer is always given after the robot has argued; without a no-robot control condition in which the same questions are answered twice without argumentation, retesting effects, reflection, and demand characteristics are plausible alternative explanations for the 117/139 alignment rate. This is not a minor caveat for the central causal claim: the paper needs either a control condition or a substantially more hedged interpretation of the alignment results.
  3. [Sec. III-A, Sec. III-D, Sec. IV-B] Question selection and IRT calibration use diagnostic answers from the same 40-student cohort, and the Monte Carlo p-values treat the estimated item parameters (a_i, b_i, c_i) and abilities (theta_j) as known. Because the same fitted model defines “expected capacity,” the statistical uncertainty in calibration is not propagated, and the risk of in-sample overfitting is substantial. This is especially concerning given the small cohort and the fact that item selection was itself guided by the same diagnostic data. A sensitivity analysis, such as bootstrapping item parameters or using split-half calibration, is needed before the “beyond expectations” count can be regarded as conservative.
  4. [Sec. IV-D and Table III] The LLM-usage finding, which appears in the abstract and conclusions, is selected from a large set of single-factor ANOVAs without multiple-comparison correction. The reported p = 0.038 for Deception alignment would not survive a simple Bonferroni correction across the many examined characteristics, and several other results in Table III are marginal (p values near 0.05–0.09). The abstract-level claim about LLM experience should either be supported by a confirmatory, pre-registered analysis or presented as exploratory.
minor comments (5)
  1. [Sec. III-A] There is a missing space in “c≥0.25to prevent over-fitting,” and the statement that the 3PL model is “particularly appropriate for three-option multiple-choice tests” should be reconciled with its use for binary experimental items.
  2. [Sec. III-C] The no-shows produced unbalanced groups with different raw diagnostic scores, and the paper states that IRT ability will be used as a covariate; the authors should report the actual IRT ability means per group, not only raw DA scores, to support the claim that ability was adequately controlled.
  3. [Sec. IV-A] The sentence about the non-alignment rate in Deception questions is duplicated: “The rate of non-alignment in Deception was in fact doubled compared to non-Deception” appears twice in near-identical form in Sec. V; one occurrence should be removed.
  4. [Sec. IV-B] The statement that “98 out of 117 alignments directly contributed to the deviation from expected results” is not defined precisely; the authors should specify what “directly contributed” means in terms of the Monte Carlo and Fisher procedure.
  5. [Sec. IV-C] The GLMM results for Conditioned questions report p-values but no effect sizes or confidence intervals; adding these would help readers assess the magnitude of the certainty effect.

Circularity Check

1 steps flagged · score 2.0 of 10

One localized circularity: for the four students without diagnostic answers, PA is used to estimate the very IRT expectation against which PA is later validated; the headline 75% claim and alignment results otherwise rest on independent empirical comparisons.

  1. fitted input called prediction [Sec. III-A (IRT estimation for students lacking diagnostic answers) and Sec. IV-B (expected-performance analysis)]
    "DA was lacking for four students and their ability and probable correctness per question was instead estimated through a second IRT analysis using the 40 students’ preliminary answers (PA) in the interaction with the robot."

    For the four students without DA, the same PA responses that are later benchmarked as “within expectations” in Sec. IV-B are used to estimate the IRT ability θ and per-question correctness probabilities cp. The expected PA is therefore not an independent baseline: it is fitted to the very data it is used to validate. In IRT, the ability estimate is chosen to reproduce the observed response pattern, so close agreement between PA and the model is achieved by construction rather than as evidence that the robot had not yet influenced the students. The FA comparison uses the same θ but does not fit FA, so the headline “75% beyond expected capacity” retains independent content; the circularity is limited to the PA validation step for these four students.

full rationale

The paper's central empirical claims—117/139 dissenting answers changed to align with the robot, 94.4%/82.6%/71.4% alignment across certainty conditions, and the PA-to-FA shift—are direct behavioral measurements, not quantities derived from fitted equations that contain the result. The expected-performance benchmark is calibrated on the diagnostic test taken before the interaction, so for the 36 students with DA the comparison of experimental answers against IRT expectations is an internal predictive benchmark rather than a circular fit. The one genuine circular step is the treatment of the four students with missing DA: their ability and expected correctness are estimated from their own PA, and the same PA is then validated as “within expectations” against that self-derived model. This is localized and does not by itself force the FA-based 75% result. The self-citation of the authors' own submitted exploratory study motivates the hypotheses but is not load-bearing evidence for the conclusions. The absence of a no-robot control condition is a causal-inference limitation, not a circularity. The skeptic's response-format mismatch (3PL guessing floor c≥0.25 calibrated on three-choice items applied to true/false items) is a validity threat to the expected-probability calibration and belongs to correctness risk, not circularity.

Assumptions & free parameters 5 free parameters · 6 assumptions · 0 invented entities

The central empirical claims depend on a small set of IRT-fitted parameters and on several domain assumptions about stability of ability, validity of the certainty manipulation, and causal attribution. No new physical or theoretical entities are introduced.

free parameters (5)
  • IRT difficulty parameters b_i for 8 selected questions = estimated from diagnostic answers (values not reported in text)
    Determines the expected probability p(θ) and hence the 'expected capacity' baseline for every student.
  • IRT discrimination parameters a_i = estimated from diagnostic answers
    Controls the steepness of item characteristic curves and the predicted correctness probabilities.
  • IRT guessing parameters c_i = constrained to be at least 0.25
    The 3PL guessing floor is an arbitrary constraint to prevent overfitting; it directly affects p(θ) for low-ability students and thus the beyond-expectation calculation.
  • Student ability parameters θ_j = estimated per student via IRT
    Each student's latent ability is fitted from diagnostic answers (or preliminary answers for four students) and is the key covariate.
  • Monte Carlo p-value threshold = p<0.05
    Students are classified as performing beyond expectations if the combined p-value crosses this threshold; the choice is conventional but not justified for this specific application.
assumptions (6)
  • domain assumption IRT assumptions: monotonicity, unidimensionality, local independence, invariance hold for the diagnostic test.
    Invoked in Sec. III-A to estimate abilities and item parameters; if violated, expected probabilities are invalid.
  • domain assumption Diagnostic-test ability measured one month earlier is stable and applies during the robot session.
    The expected performance baseline assumes no learning or fatigue between DA and the experiment.
  • domain assumption The certainty validation surveys generalize to the live interaction.
    The U/N/C portrayals were validated on video clips with adults, not with the actual student sample in the live setting.
  • domain assumption The four arguments per question are equivalent across certainty conditions except for the intended cues.
    Sec. III-E states the same four arguments were used, but the presentation differed when agreeing/disagreeing, which could confound condition with dialogue structure.
  • standard math Fisher's method for combining p-values is appropriate for detecting per-student deviations.
    Used in Sec. IV-B, step 5; assumes independent p-values from the Deception and non-Deception segments.
  • domain assumption Self-reported LLM usage is a valid measure of AI familiarity and is not confounded with other traits.
    Sec. IV-D uses a single self-report item to split students into above/below mean LLM experience; no validation or controls are reported for this measure.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sense and Sensibility: What makes a social robot convincing to high-school students?." pith.science (2026). https://pith.science/paper/4UXPRB3R

@misc{pith2026250612507,
  author       = {Pith},
  title        = {Pith review of: Sense and Sensibility: What makes a social robot convincing to high-school students?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4UXPRB3R}},
  note         = {Machine review of arXiv:2506.12507}
}
read the original abstract

This study with 40 high-school students demonstrates the high influence of a social educational robot on students' decision-making for a set of eight true-false questions on electric circuits, for which the theory had been covered in the students' courses. The robot argued for the correct answer on six questions and the wrong on two, and 75% of the students were persuaded by the robot to perform beyond their expected capacity, positively when the robot was correct and negatively when it was wrong. Students with more experience of using large language models were even more likely to be influenced by the robot's stance -- in particular for the two easiest questions on which the robot was wrong -- suggesting that familiarity with AI can increase susceptibility to misinformation by AI. We further examined how three different levels of portrayed robot certainty, displayed using semantics, prosody and facial signals, affected how the students aligned with the robot's answer on specific questions and how convincing they perceived the robot to be on these questions. The students aligned with the robot's answers in 94.4% of the cases when the robot was portrayed as Certain, 82.6% when it was Neutral and 71.4% when it was Uncertain. The alignment was thus high for all conditions, highlighting students' general susceptibility to accept the robot's stance, but alignment in the Uncertain condition was significantly lower than in the Certain. Post-test questionnaire answers further show that students found the robot most convincing when it was portrayed as Certain. These findings highlight the need for educational robots to adjust their display of certainty based on the reliability of the information they convey, to promote students' critical thinking and reduce undue influence.

Figures

Figures reproduced from arXiv: 2506.12507 by the authors.

Figure 1
Figure 1. Methodological flowchart outlining study phases: diagnostic test, robot [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The eight electric circuit statements with ratio of correct diagnostic [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Left: Sample question from the validation survey interface. Right: [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Experiment setup: Furhat robot (left), microphone (centre), screen with [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 6
Figure 6. Figure 6: Student alignment as a function of question difficulty: Easy (E) [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 5
Figure 5. Figure 5: (Left) Alignment matrix between students (sorted by ability) and questions. Blue and red cells (border or filled) indicate, respectively, students aligning and resisting after PA dissent. Cell numbers describe the question￾student pairs contribution to beyond expectati…
Figure 7
Figure 7. Figure 7: Comparison of Diagnostic (x-axis) and In-Experiment Performance [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

76 extracted references · 61 canonical work pages

  1. [1]

    Furhat: a back-projected human- like robot head for multiparty human-machine interac- tion

    Samer Al Moubayed, Jonas Beskow, Gabriel Skantze, and Bj ¨orn Granstr ¨om. Furhat: a back-projected human- like robot head for multiparty human-machine interac- tion. InCognitive Behavioural Systems: COST 2102 In- ternational Training School, Dresden, Germany, Febru- ary 21-26, 2011, Revised Selected Papers, pages 114–

  2. [2]

    Would you trust a robot with your mental health? the interaction of emotion and logic in persuasive backfiring

    Sidra Alam, Benjamin Johnston, Jonathan Vitale, and Mary-Anne Williams. Would you trust a robot with your mental health? the interaction of emotion and logic in persuasive backfiring. In2021 30th IEEE International Conference on Robot & Human Interactive Communication (RO-MAN), pages 384–391, 2021. doi: 10.1109/RO-MAN50785.2021.9515385

  3. [3]

    Using ChatGPT for teaching physics.The Physics Teacher, 62, 09 2024

    Karina Avila, Steffen Steinert, Stefan Ruzika, Jochen Kuhn, and Stefan K ¨uchemann. Using ChatGPT for teaching physics.The Physics Teacher, 62, 09 2024. doi: 10.1119/5.0227132

  4. [4]

    The benefits of interactions with physically present robots over video-displayed agents

    Wilma A Bainbridge, Justin W Hart, Elizabeth S Kim, and Brian Scassellati. The benefits of interactions with physically present robots over video-displayed agents. International Journal of Social Robotics, 3:41–52, 2011

  5. [5]

    The forgotten variable in conformity research: Impact of task importance on social influence.Journal of Personality and Social Psychology, 71:915–927, 11 1996

    Robert Baron, Joseph Vandello, and Bethany Brunsman. The forgotten variable in conformity research: Impact of task importance on social influence.Journal of Personality and Social Psychology, 71:915–927, 11 1996. doi: 10.1037/0022-3514.71.5.915

  6. [6]

    Participants conform to humans but not to humanoid robots in an english past tense formation task.Journal of Language and Social Psychology, 35(2):158–179, 2016

    Clay Beckner, P ´eter R ´acz, Jennifer Hay, J ¨”urgen Brand- stetter, and Christoph Bartneck. Participants conform to humans but not to humanoid robots in an english past tense formation task.Journal of Language and Social Psychology, 35(2):158–179, 2016

  7. [7]

    Social robots for education: A review.Science robotics, 3(21): eaat5954, 2018

    Tony Belpaeme, James Kennedy, Aditi Ramachandran, Brian Scassellati, and Fumihide Tanaka. Social robots for education: A review.Science robotics, 3(21): eaat5954, 2018. doi: https://www.science.org/doi/10. 1126/scirobotics.aat5954

  8. [8]

    Bender, Timnit Gebru, Angelina McMillan- Major, and Margaret Shmitchell

    Emily M. Bender, Timnit Gebru, Angelina McMillan- Major, and Margaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623. ACM, 2021

Show all 76 references
  1. [9]

    On the opportunities and risks of foundation models.arXiv preprint arXiv:2108.07258, 2021

    Rishi Bommasani et al. On the opportunities and risks of foundation models.arXiv preprint arXiv:2108.07258, 2021

  2. [10]

    A peer pressure experiment: Recreation of the asch conformity experiment with robots

    J ¨urgen Brandstetter, P ´eter R ´acz, Clay Beckner, Ed- uardo B Sandoval, Jennifer Hay, and Christoph Bartneck. A peer pressure experiment: Recreation of the asch conformity experiment with robots. In2014 IEEE/RSJ International Conference on Intelligent Robots and Sys- tems, ...

  3. [11]

    Trust matters: Examining the role of source evaluation in students’ construction of meaning within and across multiple texts.Reading Research Quarterly, 44(1):6–28, 2009

    Ivar Br ˚aten, Helge I Strømsø, and M Anne Britt. Trust matters: Examining the role of source evaluation in students’ construction of meaning within and across multiple texts.Reading Research Quarterly, 44(1):6–28, 2009

  4. [12]

    Designing social robots for older adults.Natl

    Cynthia L Breazeal, Anastasia K Ostrowski, Nikhita Singh, and Hae Won Park. Designing social robots for older adults.Natl. Acad. Eng. Bridge, 49:22–31, 2019

  5. [13]

    Designing persuasive robots: how robots might persuade people using vocal and nonverbal cues

    Vijay Chidambaram, Yueh-Hsuan Chiang, and Bilge Mutlu. Designing persuasive robots: how robots might persuade people using vocal and nonverbal cues. InPro- ceedings of the seventh annual ACM/IEEE international conference on Human-Robot Interaction, pages 293–300, 2012

  6. [14]

    Collins New York, New York, NY , 2007

    Robert B Cialdini and Robert B Cialdini.Influence: The psychology of persuasion, volume 55. Collins New York, New York, NY , 2007

  7. [15]

    Cialdini and Noah J

    Robert B. Cialdini and Noah J. Goldstein. Social in- fluence: Compliance and conformity.Annual Review of Psychology, 55:591–621, 2004. doi: 10.1146/annurev. psych.55.090902.142015

  8. [16]

    A study of normative and informational social influences upon in- dividual judgment.The journal of abnormal and social psychology, 51(3):629, 1955

    Morton Deutsch and Harold B Gerard. A study of normative and informational social influences upon in- dividual judgment.The journal of abnormal and social psychology, 51(3):629, 1955

  9. [17]

    Social robots in applied settings: A long- term study on adaptive robotic tutors in higher education.Frontiers in Robotics and AI, 9,

    Melissa Donnermann, Philipp Schaper, and Birgit Lu- grin. Social robots in applied settings: A long- term study on adaptive robotic tutors in higher education.Frontiers in Robotics and AI, 9,

  10. [18]

    A theory of social comparison pro- cesses.Human Relations, 7(2):117–140, 1954

    Leon Festinger. A theory of social comparison pro- cesses.Human Relations, 7(2):117–140, 1954. doi: 10.1177/001872675400700202

  11. [19]

    Oliver and Boyd, Edinburgh, UK, 1925

    Ronald Aylmer Fisher.Statistical Methods for Research Workers. Oliver and Boyd, Edinburgh, UK, 1925

  12. [20]

    Available at: https://docs.furhat.io/

    Furhat Robotics.Furhat SDK: Developer’s Guide, 2023. Available at: https://docs.furhat.io/

  13. [21]

    Kate Goddard, Abdul Roudsari, and Jeremy C. Wyatt. Automation bias: A systematic review of frequency, effect mediators, and mitigators.Journal of the American Medical Informatics Association, 19(1):121–127, 2012. doi: 10.1136/amiajnl-2011-000089

  14. [22]

    A meta-analysis of factors affecting trust in human-robot interaction.Human factors, 53(5):517–527, 2011

    Peter A Hancock, Deborah R Billings, Kristin E Schae- fer, Jessie YC Chen, Ewart J De Visser, and Raja Parasuraman. A meta-analysis of factors affecting trust in human-robot interaction.Human factors, 53(5):517–527, 2011

  15. [23]

    Credibility and credulity: Monitoring teachers for trustworthiness.Journal of Philosophy of Education, 41(2):207–219, 2007

    William Hare. Credibility and credulity: Monitoring teachers for trustworthiness.Journal of Philosophy of Education, 41(2):207–219, 2007

  16. [24]

    Influence of agent type and task ambiguity on conformity in social decision making

    Nicholas Hertz and Eva Wiese. Influence of agent type and task ambiguity on conformity in social decision making. InProceedings of the human factors and ergonomics society annual meeting, volume 60, pages 313–317. SAGE Publications Sage CA: Los Angeles, CA, 2016

  17. [25]

    Under pressure: Exam- ining social conformity with computer and robot groups

    Nicholas Hertz and Eva Wiese. Under pressure: Exam- ining social conformity with computer and robot groups. Human factors, 60(8):1207–1218, 2018

  18. [26]

    Oliver P. John, E. M. Donahue, and R. L. Kentle. The big five inventory – versions 4a and 54. Technical report, Berkeley, CA: University of California, Berkeley, Institute of Personality and Social Research., 1991

  19. [27]

    Conformity and trust in multi-party vs

    Alireza Kamelabad, Olov Engwall, and Gabriel Skantze. Conformity and trust in multi-party vs. individual human- robot interaction. InACM International Conference on Intelligent Virtual Agents (IVA ’24), pages 1–2, 2024. doi: 10.1145/3652988.3673954

  20. [28]

    Can you be persuaded? individual differences in susceptibility to persuasion

    Maurits Kaptein, Panos Markopoulos, Boris de Ruyter, and Emile Aarts. Can you be persuaded? individual differences in susceptibility to persuasion. InHuman- Computer Interaction–INTERACT 2009: 12th IFIP TC 13 International Conference, Uppsala, Sweden, August 24-28, 2009, Proce...

  21. [29]

    Harold H. Kelley. Attribution theory in social psychol- ogy. In David Levine, editor,Nebraska Symposium on Motivation, volume 15, pages 192–238. University of Nebraska Press, Lincoln, NE, 1967

  22. [30]

    Herbert C. Kelman. Compliance, identification, and in- ternalization: Three processes of attitude change.Journal of Conflict Resolution, 2(1):51–60, 1958

  23. [31]

    Com- paring robot embodiments in a guided discovery learning interaction with children.International Journal of Social Robotics, 7:293–308, 2015

    James Kennedy, Paul Baxter, and Tony Belpaeme. Com- paring robot embodiments in a guided discovery learning interaction with children.International Journal of Social Robotics, 7:293–308, 2015

  24. [32]

    Pardon my disfluency: The impact of disfluency effects on the perception of speaker competence and confidence

    Ambika Kirkland, Joakim Gustafson, and Eva Szekely. Pardon my disfluency: The impact of disfluency effects on the perception of speaker competence and confidence. InProceedings of INTERSPEECH, pages 5217–5221, 08

  25. [33]

    Un- reflected acceptance–investigating the negative conse- quences of chatgpt-assisted problem solving in physics education

    Lars Krupp, Steffen Steinert, Maximilian Kiefer- Emmanouilidis, Karina E Avila, Paul Lukowicz, Jochen Kuhn, Stefan K ¨uchemann, and Jakob Karolus. Un- reflected acceptance–investigating the negative conse- quences of chatgpt-assisted problem solving in physics education. InHHA...

  26. [34]

    Can ChatGPT support prospective teachers in physics task development?Phys- ical Review Physics Education Research, 19, 09 2023

    Stefan K ¨uchemann, Steffen Steinert, Natalia Re- venga Lozano, Matthias Schweinberger, Yavuz Dinc, Karina Avila, and Jochen Kuhn. Can ChatGPT support prospective teachers in physics task development?Phys- ical Review Physics Education Research, 19, 09 2023. doi: 10.1103/PhysR...

  27. [35]

    Beyond style: Synthesizing speech with pragmatic func- tions

    Harm Lameris, Jaokim Gustafson, and Eva Szekely. Beyond style: Synthesizing speech with pragmatic func- tions. InProceedings of INTERSPEECH, pages 3382– 3386, 2023. doi: 10.21437/Interspeech.2023-2072

  28. [36]

    The psychology of social impact.Amer- ican Psychologist, 36(4):343–356, 1981

    Bibb Latan ´e. The psychology of social impact.Amer- ican Psychologist, 36(4):343–356, 1981. doi: 10.1037/ 0003-066X.36.4.343

  29. [37]

    Students’ personality and susceptibility to persuasion during mathematics group- work: An exploratory study.Journal of Practical Studies in Education, 2(6):10–22, 2021

    Jieun Lee and Lillie R Albert. Students’ personality and susceptibility to persuasion during mathematics group- work: An exploratory study.Journal of Practical Studies in Education, 2(6):10–22, 2021

  30. [38]

    A systematic review of experimental work on persuasive social robots.International Journal of Social Robotics, 14(6):1339–1378, 2022

    Baisong Liu, Daniel Tetteroo, and Panos Markopoulos. A systematic review of experimental work on persuasive social robots.International Journal of Social Robotics, 14(6):1339–1378, 2022

  31. [39]

    Self- CheckGPT: Zero-resource black-box hallucination detec- tion for generative large language models

    Potsawee Manakul, Adian Liusie, and Mark Gales. Self- CheckGPT: Zero-resource black-box hallucination detec- tion for generative large language models. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceed- ings of the 2023 Conference on Empirical Methods in Natural La...

  32. [40]

    Hi robot, it’s not what you say, it’s how you say it

    J ¨ura Miniota, Siyang Wang, Jonas Beskow, Joakim Gustafson, Eva Szekely, and Andr ´e Pereiral. Hi robot, it’s not what you say, it’s how you say it. In2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), pages 307–314,

  33. [41]

    Real-time emotion generation in human-robot dialogue using large language models

    Chinmaya Mishra, Rinus Verdonschot, Peter Hagoort, and Gabriel Skantze. Real-time emotion generation in human-robot dialogue using large language models. Frontiers in Robotics and AI, 10, 2023

  34. [42]

    Montgomery.Design and Analysis of Exper- iments

    Douglas C. Montgomery.Design and Analysis of Exper- iments. John Wiley & Sons, Hoboken, NJ, 2017

  35. [43]

    doi: 10.1109/RO-MAN57019.2023.10309427

  36. [44]

    Adaptive trust calibration for human-ai collaboration.Plos one, 15(2): e0229132, 2020

    Kazuo Okamura and Seiji Yamada. Adaptive trust calibration for human-ai collaboration.Plos one, 15(2): e0229132, 2020

  37. [45]

    Trust toward robots and artificial intelli- gence: An experimental approach to human–technology interactions online.Frontiers in Psychology, 11,

    Atte Oksanen, Nina Savela, Rita Latikka, and Aki Koivula. Trust toward robots and artificial intelli- gence: An experimental approach to human–technology interactions online.Frontiers in Psychology, 11,

  38. [46]

    Subjective con- sistency increases trust.Scientific reports, 13(1):5657, 2023

    Andrzej Nowak, Mikolaj Biesaga, Karolina Ziembowicz, Tomasz Baran, and Piotr Winkielman. Subjective con- sistency increases trust.Scientific reports, 13(1):5657, 2023

  39. [47]

    The importance of the person’s assertiveness in persuasive human-robot interactions

    Raul Benites Paradeda, Maria Jos ´e Ferreira, Carlos Mar- tinho, and Ana Paiva. The importance of the person’s assertiveness in persuasive human-robot interactions. In Social Robotics: 12th International Conference, ICSR 2020, Golden, CO, USA, November 14–18, 2020, Pro- ceedin...

  40. [48]

    Persuasion strategies using a social robot in an interactive storytelling scenario

    Raul Benites Paradeda, Carlos Martinho, and Ana Paiva. Persuasion strategies using a social robot in an interactive storytelling scenario. InProceedings of the 8th Inter- national Conference on Human-Agent Interaction, HAI ’20, page 69–77, New York, NY , USA, 2020. Association...

  41. [49]

    Jin-Hwa Park and Eun Kyung Lee. Influence of pro- fessor trust, self-directed learning and self-esteem on satisfaction with major study in nursing students.The Korean Data & Information Science Society, 29(1):167– 178, 2018

  42. [50]

    The effects of educational robotics in stem education: a multilevel meta-analysis

    Fan Ouyang and Weiqi Xu. The effects of educational robotics in stem education: a multilevel meta-analysis. International Journal of STEM Education, 11, 02 2024. doi: 10.1186/s40594-024-00469-4

  43. [51]

    Springer, 1986

    Richard E Petty, John T Cacioppo, Richard E Petty, and John T Cacioppo.The elaboration likelihood model of persuasion. Springer, 1986

  44. [52]

    In-body experiences: embodiment, control, and trust in robot- mediated communication

    Irene Rae, Leila Takayama, and Bilge Mutlu. In-body experiences: embodiment, control, and trust in robot- mediated communication. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems, pages 1921–1930, 2013

  45. [53]

    Overtrust of robots in emer- gency evacuation scenarios

    Paul Robinette, Wenchen Li, Robert Allen, Ayanna Howard, and Alan Wagner. Overtrust of robots in emer- gency evacuation scenarios. InACM/IEEE international conference on human-robot interaction, pages 101–108, 03 2016. doi: 10.1109/HRI.2016.7451740

  46. [54]

    Petty and John T

    Richard E. Petty and John T. Cacioppo.Communica- tion and Persuasion: Central and Peripheral Routes to Attitude Change. Springer-Verlag, New York, NY , 1986

  47. [55]

    Humans conform to robots: Disambiguating trust, truth, and conformity

    Nicole Salomons, Michael van der Linden, Sarah Strohkorb Sebo, and Brian Scassellati. Humans conform to robots: Disambiguating trust, truth, and conformity. InProceedings of the 2018 ACM/IEEE International Conference on Human-Robot Interaction, HRI ’18, page 187–195, New York,...

  48. [56]

    A minority of one against a majority of robots: Robots cause normative and informa- tional conformity.ACM Transactions on Human-Robot Interaction (THRI), 10(2):1–22, 2021

    Nicole Salomons, Sarah Strohkorb Sebo, Meiying Qin, and Brian Scassellati. A minority of one against a majority of robots: Robots cause normative and informa- tional conformity.ACM Transactions on Human-Robot Interaction (THRI), 10(2):1–22, 2021

  49. [57]

    It would make me happy if you used my guess: Comparing robot persuasive strategies in social human–robot interaction

    Shane Saunderson and Goldie Nejat. It would make me happy if you used my guess: Comparing robot persuasive strategies in social human–robot interaction. IEEE Robotics and Automation Letters, 4(2):1707–1714,

  50. [58]

    Devel- opment and validation of a student self-efficacy scale

    Melodie Rowbotham and Gerdamarie Schmitz. Devel- opment and validation of a student self-efficacy scale. Journal of Nursing & Care, 02, 01 2013. doi: 10.4172/ 2167-1168.1000126

  51. [59]

    Saunderson and Goldie Nejat

    Stephen P. Saunderson and Goldie Nejat. Persuasive robots should avoid authority: The effects of formal and real authority on persuasion in human-robot inter- action.Science Robotics, 6(58):eabd5186, 2021. doi: 10.1126/scirobotics.abd5186. URL https://www.science. org/doi/full...

  52. [60]

    Harnessing large lan- guage models to enhance self-regulated learning via formative feedback.arXiv preprint arXiv:2311.13984, 2023

    Steffen Steinert, Karina E Avila, Stefan Ruzika, Jochen Kuhn, and Stefan K ¨uchemann. Harnessing large lan- guage models to enhance self-regulated learning via formative feedback.arXiv preprint arXiv:2311.13984, 2023

  53. [61]

    How can LLMs transform the robotic de- sign process?Nature Machine Intelligence, 5(6): 561–564, 2023

    Francesco Stella, Cosimo Della Santina, and Josie Hughes. How can LLMs transform the robotic de- sign process?Nature Machine Intelligence, 5(6): 561–564, 2023. ISSN 2522-5839. doi: 10.1038/ s42256-023-00669-7

  54. [62]

    Tormala and Richard E

    Zakary L. Tormala and Richard E. Petty. What doesn’t kill me makes me stronger: The effects of resisting persuasion on attitude certainty.Journal of Personality and Social Psychology, 83(6):1298–1313, 2002. doi: 10.1037/0022-3514.83.6.1298

  55. [63]

    Investigating strategies for robot persuasion in social human–robot interaction.IEEE Transactions on Cybernetics, 52(1): 641–653, 2020

    Shane Saunderson and Goldie Nejat. Investigating strategies for robot persuasion in social human–robot interaction.IEEE Transactions on Cybernetics, 52(1): 641–653, 2020

  56. [64]

    I am definitely certain of this! towards a multimodal repertoire of signals com- municating a high degree of certainty

    Laura Vincze and Isabella Poggi. I am definitely certain of this! towards a multimodal repertoire of signals com- municating a high degree of certainty. InEuropean and 7th Nordic Symposium on Multimodal Communication, 2016

  57. [65]

    Gender and robots: A literature review.arXiv preprint arXiv:2206.04716, 2022

    David Gray Widder. Gender and robots: A literature review.arXiv preprint arXiv:2206.04716, 2022

  58. [66]

    Ef- fective persuasion strategies for socially assistive robots

    Katie Winkle, S ´everin Lemaignan, Praminda Caleb- Solly, Ute Leonards, Ailie Turton, and Paul Bremner. Ef- fective persuasion strategies for socially assistive robots. InProceedings of 14th ACM/IEEE International Confer- ence on Human-Robot Interaction, pages 277–285, 03

  59. [67]

    Improved trust in human-robot collaboration with ChatGPT.IEEE Access, PP:1–1, 01 2023

    Yang Ye, Hengxu You, and Jing Du. Improved trust in human-robot collaboration with ChatGPT.IEEE Access, PP:1–1, 01 2023. doi: 10.1109/ACCESS.2023.3282111

  60. [68]

    Trust motivation: The self-regulatory processes underlying trust decisions

    Lisa van der Werff, Alison Legood, Finian Buckley, An- toinette Weibel, and David de Cremer. Trust motivation: The self-regulatory processes underlying trust decisions. Organizational Psychology Review, 9(2-3):99–123, 2019. doi: 10.1177/2041386619873616

  61. [69]

    A systematic review on exploring the potential of educational robotics in mathematics education.International Journal of Sci- ence and Mathematics Education, 18, 11 2018

    Baichang Zhong and Liying Xia. A systematic review on exploring the potential of educational robotics in mathematics education.International Journal of Sci- ence and Mathematics Education, 18, 11 2018. doi: 10.1007/s10763-018-09939-y

  62. [70]

    Social influence under uncertainty in interaction with peers, robots and computers.International Journal of Social Robotics, 15:249–268, 2021

    Joshua Zonca, Anna Folsø, and Alessandra Sciutti. Social influence under uncertainty in interaction with peers, robots and computers.International Journal of Social Robotics, 15:249–268, 2021. URL https://api. semanticscholar.org/CorpusID:254876990

  63. [72]

    doi: 10.1109/HRI.2019.8673313

  64. [74]

    Large language models for human- robot interaction: A review.Biomimetic Intelligence and Robotics, 3:100131, 10 2023

    Ceng Zhang, Junxin Chen, Jiatong Li, Yanhong Peng, and Zebing Mao. Large language models for human- robot interaction: A review.Biomimetic Intelligence and Robotics, 3:100131, 10 2023. doi: 10.1016/j.birob.2023. 100131

  65. [2019]

    doi: 10.1109/LRA.2019.2897143

  66. [2020]

    doi: 10.3389/fpsyg.2020

    ISSN 1664-1078. doi: 10.3389/fpsyg.2020. 568256. URL https://www.frontiersin.org/journals/ psychology/articles/10.3389/fpsyg.2020.568256

  67. [2022]

    doi: 10.3389/frobt.2022

    ISSN 2296-9144. doi: 10.3389/frobt.2022. 831633. URL https://www.frontiersin.org/journals/ robotics-and-ai/articles/10.3389/frobt.2022.831633

  68. [2023]

    doi: 10.21437/Interspeech.2023-887

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.