REVIEW 1 major objections 1 minor 40 references
Evaluating Social Engineering Risks in AI-based Interaction using Biometrics and a Gaming Setup
T0 review · 1 major / 1 minor · reviewed 2026-07-03 · grok-4.3
Pith's one-line read A gaming platform and 340 GB dataset evaluates social engineering risks in interactions with AI agents via biometrics and controlled games.
desk verdict Releases a new open multimodal dataset from LLM-mediated games with biometrics, but provides no evidence that the setup measures social engineering risks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
AIriskEval-gaming platform, which integrates configurable game templates, role-conditioned LLM agents, psychology-informed profiling, and synchronized multimodal data streams for risk evaluation.
What would settle it
A direct comparison study in which the same participants show no distinguishable biometric or behavioral patterns when exposed to actual social-engineering prompts outside the game setting versus neutral prompts would indicate the platform does not capture the targeted risks.
Extended reading notes
Core claim
The paper establishes AIriskEval-gaming as a configurable platform and accompanying dataset that supports controlled evaluation of social engineering risks in LLM-mediated multimodal interaction by combining game templates, role-conditioned agents, participant profiling, structured interaction trees, and synchronized acquisition of behavioral and biometric streams.
Load-bearing premise
Gameplay in the adapted Prisoner's Dilemma and Ultimatum Game with role-conditioned AI agents, together with the recorded biometric streams, serves as a valid proxy for social engineering risks that occur in wider real-world AI interactions.
Editorial extensions
If this is right
- Enables systematic identification of vulnerabilities in AI systems before wider deployment.
- Provides data to support protection of sensitive information during AI interactions.
- Facilitates compliance checks against evolving regulatory requirements for AI security.
- Allows comparative testing across human-human, human-AI, and AI-AI configurations.
Reading between the lines
- Patterns extracted from the six data streams may later serve as early indicators for training detection models that flag manipulation attempts in live AI conversations.
- The platform's structure could be reused with additional participant cohorts to test whether biometric signatures generalize beyond the initial 15-person sample.
- AI-AI game sessions in the dataset might surface interaction dynamics that differ from human-AI cases and warrant separate risk modeling.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces AIriskEval-gaming, an open platform and the AIriskEval-gaming-db dataset collected from 15 participants who played concatenated adapted Prisoner's Dilemma and Ultimatum Game scenarios against role-conditioned GPT-5.4 agents. It supplies 340 GB of synchronized multimodal streams (interaction logs, video, gaze, smartwatch, facial features) plus descriptive analyses of paths, outcomes, psychological profiles, and derived features, positioning the resource as enabling evaluation of social engineering risks in LLM-mediated interactions across human-human, human-AI, and AI-AI settings.
Significance. If the proxy validity of the chosen games and biometric streams for social engineering behaviors can be established, the open release of configurable templates, structured interaction trees, and this large multimodal dataset on GitHub would constitute a useful contribution to HCI and AI-safety research by supporting controlled, reproducible studies of biometric and behavioral signals. The explicit provision of raw/processed data and code is a clear strength that facilitates follow-on work.
major comments (1)
- [Abstract] Abstract and dataset description: The central claim that the platform and dataset enable 'rigorous risk evaluation' of social engineering risks rests on the untested assumption that behaviors and signals observed in the adapted Prisoner's Dilemma and Ultimatum Game with role-conditioned LLM agents serve as valid proxies for real-world tactics such as manipulation or trust exploitation. Only descriptive statistics on 15 participants are supplied; no expert annotation, external validation, or mapping to SE-relevant outcomes is presented to support this link.
minor comments (1)
- [Dataset description] The description of the six data streams and deep-learning-based feature extraction would benefit from an explicit table listing per-stream sampling rates, synchronization method, and any filtering steps applied before feature extraction.
Simulated Author's Rebuttal
We thank the referee for their constructive feedback on the positioning of our contribution. The work introduces an open platform and multimodal dataset using established game-theoretic scenarios to support research on social engineering risks in LLM-mediated interactions. We address the major comment below by clarifying scope and committing to revisions where appropriate.
read point-by-point responses
-
Referee: [Abstract] Abstract and dataset description: The central claim that the platform and dataset enable 'rigorous risk evaluation' of social engineering risks rests on the untested assumption that behaviors and signals observed in the adapted Prisoner's Dilemma and Ultimatum Game with role-conditioned LLM agents serve as valid proxies for real-world tactics such as manipulation or trust exploitation. Only descriptive statistics on 15 participants are supplied; no expert annotation, external validation, or mapping to SE-relevant outcomes is presented to support this link.
Authors: We agree that the manuscript supplies only descriptive statistics and does not include expert annotation, external validation, or explicit mappings from observed behaviors to real-world social engineering outcomes. The central positioning is that the platform and dataset provide controlled, configurable game templates (adapted Prisoner's Dilemma and Ultimatum Game), role-conditioned agents, structured interaction trees, and synchronized multimodal streams to enable future researchers to conduct such evaluations across human-human, human-AI, and AI-AI settings. These games were selected because prior literature has linked them to constructs such as trust, cooperation, and exploitation, but we do not assert that the collected signals constitute validated proxies. The abstract's phrasing that the resource enables 'rigorous risk evaluation' can be read as overstating immediate applicability. We will revise the abstract, introduction, and dataset description sections to state explicitly that the contribution supplies a foundational resource and data for subsequent validation studies rather than performing or claiming to have performed that validation. This change will be incorporated in the revised manuscript. revision: yes
Circularity Check
Dataset release paper with no derivations, predictions, or self-referential steps
full rationale
The paper introduces an open platform and dataset (AIriskEval-gaming) for collecting multimodal interaction data in adapted games with LLM agents. No equations, model fitting, predictions, or derivations are present in the abstract or described content. The work consists of platform description, data collection from 15 participants, and descriptive statistics; the central claim is the resource release itself rather than any computed result that could reduce to its inputs. No self-citations, ansatzes, or uniqueness claims are invoked as load-bearing elements. This is a standard non-circular dataset paper.
Assumptions & free parameters
assumptions (1)
- domain assumption Multimodal biometric and behavioral signals collected during role-conditioned game interactions can serve as indicators of social engineering vulnerabilities in AI-mediated exchanges
Cite this review
Pith. "Pith review of Evaluating Social Engineering Risks in AI-based Interaction using Biometrics and a Gaming Setup." pith.science (2026). https://pith.science/paper/EHNLRGPR
@misc{pith2026260617793,
author = {Pith},
title = {Pith review of: Evaluating Social Engineering Risks in AI-based Interaction using Biometrics and a Gaming Setup},
year = {2026},
howpublished = {\url{https://pith.science/paper/EHNLRGPR}},
note = {Machine review of arXiv:2606.17793}
}
read the original abstract
We introduce AIriskEval-gaming, an open platform and dataset to evaluate social engineering risks in LLM-mediated multimodal interaction through controlled games. It supports human-human, human-AI and AI-AI settings, combining configurable game templates, role-conditioned LLM agents, psychology-informed participant profiling, structured interaction trees, and synchronized behavioral and biometric acquisition, filtering, and deep-learning-based feature extraction. The dataset (AIriskEval-gaming-db) was collected from 15 participants who interacted with a role-conditioned GPT-5.4 agent in two concatenated games: an adapted Prisoner's Dilemma and an Ultimatum Game. It comprises 340 GB of raw and processed multimodal data across six streams: interaction logs, video, screen recordings, gaze logs, smartwatch signals, and game/questionnaire metadata. These data include interaction paths, written justifications, psychological profiles, subjective feedback, perceived counterpart identity, game outcomes, and derived behavioral, facial, and gaze features. Alongside the dataset, we provide descriptive analyses characterizing the resulting multimodal data. Rigorous risk evaluation is essential for the deployment of secure AI systems, as it enables the identification and mitigation of vulnerabilities, ensures the protection of sensitive data, and supports compliance with evolving regulatory and ethical standards in society. The dataset and related code are available on GitHub.
Figures
Reference graph
Works this paper leans on
-
[1]
On the Feasibility of Using Multimodal LLMs to Execute AR Social Engineering Attacks,
T. Bi, C. Yeet al., “On the Feasibility of Using Multimodal LLMs to Execute AR Social Engineering Attacks,” inProc. AAAI Conf. on Artificial Intelligence, vol. 40, no. 45, 2026, pp. 38 252–38 260
work page 2026
-
[2]
A Turing Test of Whether AI Chatbots are Behaviorally similar to Humans,
Q. Mei, Y . Xie, W. Yuan, and M. O. Jackson, “A Turing Test of Whether AI Chatbots are Behaviorally similar to Humans,” inProc. of the National Academy of Sciences, vol. 121, no. 9, 2024, p. e2313925121
work page 2024
-
[3]
Adverse Re- actions to the Use of Large Language Models in Social Interactions,
F. Dvorak, R. Stumpf, S. Fehrler, and U. Fischbacher, “Adverse Re- actions to the Use of Large Language Models in Social Interactions,” PNAS Nexus, vol. 4, no. 4, p. pgaf112, 2025
work page 2025
-
[4]
edBB-Demo: Biometrics and Behavior Analysis for Online Educational Platforms,
R. Daza, A. Moraleset al., “edBB-Demo: Biometrics and Behavior Analysis for Online Educational Platforms,” inProc. AAAI Conf. on Artificial Intelligence (Demonstration), 2023, pp. 16 422–16 424
work page 2023
-
[5]
R. Daza, A. Becerra, R. Cobos, J. Fierrezet al., “A Multimodal Dataset for Understanding the Impact of Mobile Phones on Remote Online Virtual Education,”Scientific Data, vol. 12, no. 1, p. 1332, 2025
work page 2025
-
[6]
SMARTe-VR: Student Monitoring and Adaptive Response Technology for e-Learning in Virtual Reality,
R. Daza, S. Lin, A. Morales, J. Fierrez, and K. Nagao, “SMARTe-VR: Student Monitoring and Adaptive Response Technology for e-Learning in Virtual Reality,” inACM I2M-MM, 2025
work page 2025
-
[7]
Cyber Coercion Detection Using LLM-Assisted Mul- timodal Biometric System,
A. Almehmadi, “Cyber Coercion Detection Using LLM-Assisted Mul- timodal Biometric System,”Applied Sciences, vol. 15, p. 10658, 2025
work page 2025
-
[8]
Promises and Trust in Human–Robot Interaction,
L. Cominelli, F. Feriet al., “Promises and Trust in Human–Robot Interaction,”Scientific Reports, vol. 11, no. 1, p. 9687, 2021
work page 2021
Show all 40 references
-
[9]
Playing repeated games with large language models,
E. Akata, L. Schulzet al., “Playing repeated games with large language models,”Nature Human Behaviour, vol. 9, no. 7, pp. 1380–1390, 2025
2025
-
[10]
Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing,
M. Schmitt and I. Flechais, “Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing,”Artificial Intelligence Review, vol. 57, no. 12, p. 324, 2024
2024
-
[11]
Evaluating Large Language Models’ Ability to Automate Spear Phishing,
F. Heidinget al., “Evaluating Large Language Models’ Ability to Automate Spear Phishing,”Expert Systems with Applications, 2026
2026
-
[12]
The machine psychology of cooperation: Can GPT models operationalize prompts for altruism, cooperation, competitive- ness, and selfishness in economic games?
S. Phelpset al., “The machine psychology of cooperation: Can GPT models operationalize prompts for altruism, cooperation, competitive- ness, and selfishness in economic games?”J. Phys.: Complex., 2025
2025
-
[13]
oTree—An Open-Source Platform for Laboratory, Online, and Field Experiments,
D. L. Chen, M. Schonger, and C. Wickens, “oTree—An Open-Source Platform for Laboratory, Online, and Field Experiments,”Journal of Behavioral and Experimental Finance, vol. 9, pp. 88–97, 2016
2016
-
[14]
nodeGame: Real-Time, Synchronous, Online Experiments in the Browser,
S. Balietti, “nodeGame: Real-Time, Synchronous, Online Experiments in the Browser,”Behavior Research Methods, vol. 49, no. 5, 2017
2017
-
[15]
Empirica: a Virtual Lab for High- Throughput Macro-Level Experiments,
A. Almaatouq, J. Beckeret al., “Empirica: a Virtual Lab for High- Throughput Macro-Level Experiments,”Behavior Research Methods, vol. 53, no. 5, pp. 2158–2171, 2021
2021
-
[16]
LIONESS Lab: a Free Web-Based Platform for Conducting Interactive Experiments Online,
M. Giamattei, K. S. Yahosseiniet al., “LIONESS Lab: a Free Web-Based Platform for Conducting Interactive Experiments Online,”Journal of the Economic Science Association, vol. 6, no. 1, pp. 95–111, 2020
2020
-
[17]
HARMONIC: A Multimodal Dataset of Assistive Human– Robot Collaboration,
B. A. Newman, R. M. Aronson, S. S. Srinivasa, K. Kitani, and H. Admoni, “HARMONIC: A Multimodal Dataset of Assistive Human– Robot Collaboration,”Int. J. Robot. Res., vol. 41, no. 1, pp. 3–11, 2022
2022
-
[18]
Physiological Data for Affective Computing in HRI with Anthropomorphic Service Robots: the AFFECT-HRI Data Set,
J. S. Heinischet al., “Physiological Data for Affective Computing in HRI with Anthropomorphic Service Robots: the AFFECT-HRI Data Set,” Scientific Data, vol. 11, no. 1, p. 333, 2024
2024
-
[19]
MultiPhysio-HRC: A Multimodal Phys- iological Signals Dataset for Industrial Human–Robot Collaboration,
A. Bussolan, S. Baraldoet al., “MultiPhysio-HRC: A Multimodal Phys- iological Signals Dataset for Industrial Human–Robot Collaboration,” Robotics, vol. 14, no. 12, p. 184, 2025
2025
-
[20]
The Number of Distinct Basic Values and Their Structure Assessed by PVQ–40,
J. Cieciuch and S. H. Schwartz, “The Number of Distinct Basic Values and Their Structure Assessed by PVQ–40,”Journal of Personality Assessment, vol. 94, no. 3, pp. 321–328, 2012
2012
-
[21]
The Dispositional Essence of Proactive Social Preferences: The Dark Core of Personality Vis- `a-Vis 58 Traits,
B. E. Hilbig, I. Thielmannet al., “The Dispositional Essence of Proactive Social Preferences: The Dark Core of Personality Vis- `a-Vis 58 Traits,” Psychological Science, vol. 34, no. 2, pp. 201–220, 2023
2023
-
[22]
Who is Healthier? A Meta-Analysis of the Relations between the HEXACO Personality Domains and Health Outcomes,
J. L. Pletzeret al., “Who is Healthier? A Meta-Analysis of the Relations between the HEXACO Personality Domains and Health Outcomes,”Eur. J. Pers., vol. 38, no. 2, pp. 342–364, 2024
2024
-
[23]
The Development and Testing of a New Version of the Cognitive Reflection Test applying Item Response Theory (IRT),
C. Primiet al., “The Development and Testing of a New Version of the Cognitive Reflection Test applying Item Response Theory (IRT),”J. Behav. Decis. Mak., vol. 29, no. 5, pp. 453–469, 2016
2016
-
[24]
DeepFace-Attention: Multimodal Face Biometrics for Attention Esti- mation with Application to e-learning,
R. Daza, L. F. Gomez, J. Fierrez, A. Morales, R. Tolosanaet al., “DeepFace-Attention: Multimodal Face Biometrics for Attention Esti- mation with Application to e-learning,”IEEE Access, vol. 12, 2024
2024
-
[25]
Facial expressions as a vulnerability in face recognition,
A. Pe ˜na, I. Serna, A. Morales, J. Fierrez, and A. Lapedriza, “Facial expressions as a vulnerability in face recognition,” inIEEE ICIP, 2021
2021
-
[26]
Exploring facial expressions and action unit domains for Parkinson detection,
L. F. Gomez, A. Moraleset al., “Exploring facial expressions and action unit domains for Parkinson detection,”PLoS One, vol. 18, no. 2, 2023
2023
-
[27]
RetinaFace: Single-Shot Multi-Level Face Localisation in the Wild,
J. Deng, J. Guoet al., “RetinaFace: Single-Shot Multi-Level Face Localisation in the Wild,” inProc. CVPR, 2020, pp. 5203–5212
2020
-
[28]
MediaPipe: A Framework for Perceiving and Processing Reality,
C. Lugaresiet al., “MediaPipe: A Framework for Perceiving and Processing Reality,” inProc. CVPR Workshops, 2019
2019
-
[29]
WHENet: Real-time Fine-Grained Estimation for Wide Range Head Pose,
Y . Zhou and J. Gregson, “WHENet: Real-time Fine-Grained Estimation for Wide Range Head Pose,” inProc. BMVC, 2020
2020
-
[30]
OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis,
J. Huet al., “OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis,” inProc. IEEE FG, 2025
2025
-
[31]
mEBAL2 Database and Benchmark: Image-based Multispectral Eye- blink Detection,
R. Daza, A. Morales, J. Fierrez, R. Tolosana, and R. Vera-Rodriguez, “mEBAL2 Database and Benchmark: Image-based Multispectral Eye- blink Detection,”Pattern Recognition Letters, vol. 182, pp. 83–89, 2024
2024
-
[32]
Keystroke Verification Challenge (KVC): Biomet- ric and fairness benchmark evaluation,
G. Stragapedeet al., “Keystroke Verification Challenge (KVC): Biomet- ric and fairness benchmark evaluation,”IEEE Access, vol. 12, 2024
2024
-
[33]
BeCAPTCHA-Mouse: Synthetic Mouse Trajectories and Improved Bot Detection,
A. Acien, A. Morales, J. Fierrez, and R. Vera-Rodriguez, “BeCAPTCHA-Mouse: Synthetic Mouse Trajectories and Improved Bot Detection,”Pattern Recognition, vol. 127, p. 108643, 2022
2022
-
[34]
Personalized weight loss management through wearable devices and artificial intelli- gence,
S. Romero-Tapiador, R. Tolosana, A. Moraleset al., “Personalized weight loss management through wearable devices and artificial intelli- gence,”Computers in Biology and Medicine, vol. 209, p. 111676, 2026
2026
-
[35]
Overview of Risk Assessment and Management for Intelligent Systems under the AI Act and Beyond,
J. Irigoyenet al., “Overview of Risk Assessment and Management for Intelligent Systems under the AI Act and Beyond,” inICCST, 2026
2026
-
[36]
Is my vision-language data in your AI? membership inference test (MINT) Demo 2,
D. DeAlcalaet al., “Is my vision-language data in your AI? membership inference test (MINT) Demo 2,” inIEEE COMPSAC, 2026
2026
-
[37]
Leveraging avatar fingerprinting: A photorealistic talking-head public database and benchmark,
L. Pedrouzoet al., “Leveraging avatar fingerprinting: A photorealistic talking-head public database and benchmark,”arXiv:2603.26934, 2026
2026
-
[38]
Addressing bias in LLMs: Strategies and application to fair AI-based recruitment,
A. Pe ˜naet al., “Addressing bias in LLMs: Strategies and application to fair AI-based recruitment,” inAAAI/ACM AIES, 2025
2025
-
[39]
DeepID challenge of detecting synthetic manipu- lations in ID documents,
P. Korshunovet al., “DeepID challenge of detecting synthetic manipu- lations in ID documents,” inIEEE ICCV Workshops, 2025
2025
-
[40]
AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K- 12 Educational Explanations,
J. Irigoyen, R. Daza, F. Jurado, J. Fierrez, R. Tolosana, A. Ortigosaet al., “AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K- 12 Educational Explanations,” inIEEE ICCST, 2026
2026
Reviewed July 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.