Pith. sign in

REVIEW 1 major objections 1 minor 40 references

Evaluating Social Engineering Risks in AI-based Interaction using Biometrics and a Gaming Setup

T0 review · 1 major / 1 minor · reviewed 2026-07-03 · grok-4.3

Pith's one-line read A gaming platform and 340 GB dataset evaluates social engineering risks in interactions with AI agents via biometrics and controlled games.

desk verdict Releases a new open multimodal dataset from LLM-mediated games with biometrics, but provides no evidence that the setup measures social engineering risks. read the letter →

arxiv 2606.17793 v2 pith:EHNLRGPR submitted 2026-06-16 cs.HC cs.DB

classification cs.HCcs.DB
keywords socialengineeringAIrisksmultimodaldatasetgamingplatformbiometricsLLMinteractionPrisoner'sDilemma
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces AIriskEval-gaming as an open platform that runs human-AI and other interaction settings inside two concatenated games: an adapted Prisoner's Dilemma and an Ultimatum Game. Role-conditioned LLM agents interact with 15 participants while six synchronized data streams capture logs, video, gaze, smartwatch signals, and metadata. The resulting dataset supplies interaction paths, psychological profiles, game outcomes, and derived behavioral and biometric features for analysis. Descriptive statistics characterize the collected multimodal signals. The work positions this resource as a means to identify vulnerabilities that affect secure AI deployment.

What carries the argument

AIriskEval-gaming platform, which integrates configurable game templates, role-conditioned LLM agents, psychology-informed profiling, and synchronized multimodal data streams for risk evaluation.

What would settle it

A direct comparison study in which the same participants show no distinguishable biometric or behavioral patterns when exposed to actual social-engineering prompts outside the game setting versus neutral prompts would indicate the platform does not capture the targeted risks.

Watch

Extended reading notes

Core claim

The paper establishes AIriskEval-gaming as a configurable platform and accompanying dataset that supports controlled evaluation of social engineering risks in LLM-mediated multimodal interaction by combining game templates, role-conditioned agents, participant profiling, structured interaction trees, and synchronized acquisition of behavioral and biometric streams.

Load-bearing premise

Gameplay in the adapted Prisoner's Dilemma and Ultimatum Game with role-conditioned AI agents, together with the recorded biometric streams, serves as a valid proxy for social engineering risks that occur in wider real-world AI interactions.

Editorial extensions

If this is right

  • Enables systematic identification of vulnerabilities in AI systems before wider deployment.
  • Provides data to support protection of sensitive information during AI interactions.
  • Facilitates compliance checks against evolving regulatory requirements for AI security.
  • Allows comparative testing across human-human, human-AI, and AI-AI configurations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Patterns extracted from the six data streams may later serve as early indicators for training detection models that flag manipulation attempts in live AI conversations.
  • The platform's structure could be reused with additional participant cohorts to test whether biometric signatures generalize beyond the initial 15-person sample.
  • AI-AI game sessions in the dataset might surface interaction dynamics that differ from human-AI cases and warrant separate risk modeling.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 1 minor

Summary. The manuscript introduces AIriskEval-gaming, an open platform and the AIriskEval-gaming-db dataset collected from 15 participants who played concatenated adapted Prisoner's Dilemma and Ultimatum Game scenarios against role-conditioned GPT-5.4 agents. It supplies 340 GB of synchronized multimodal streams (interaction logs, video, gaze, smartwatch, facial features) plus descriptive analyses of paths, outcomes, psychological profiles, and derived features, positioning the resource as enabling evaluation of social engineering risks in LLM-mediated interactions across human-human, human-AI, and AI-AI settings.

Significance. If the proxy validity of the chosen games and biometric streams for social engineering behaviors can be established, the open release of configurable templates, structured interaction trees, and this large multimodal dataset on GitHub would constitute a useful contribution to HCI and AI-safety research by supporting controlled, reproducible studies of biometric and behavioral signals. The explicit provision of raw/processed data and code is a clear strength that facilitates follow-on work.

major comments (1)
  1. [Abstract] Abstract and dataset description: The central claim that the platform and dataset enable 'rigorous risk evaluation' of social engineering risks rests on the untested assumption that behaviors and signals observed in the adapted Prisoner's Dilemma and Ultimatum Game with role-conditioned LLM agents serve as valid proxies for real-world tactics such as manipulation or trust exploitation. Only descriptive statistics on 15 participants are supplied; no expert annotation, external validation, or mapping to SE-relevant outcomes is presented to support this link.
minor comments (1)
  1. [Dataset description] The description of the six data streams and deep-learning-based feature extraction would benefit from an explicit table listing per-stream sampling rates, synchronization method, and any filtering steps applied before feature extraction.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for their constructive feedback on the positioning of our contribution. The work introduces an open platform and multimodal dataset using established game-theoretic scenarios to support research on social engineering risks in LLM-mediated interactions. We address the major comment below by clarifying scope and committing to revisions where appropriate.

read point-by-point responses
  1. Referee: [Abstract] Abstract and dataset description: The central claim that the platform and dataset enable 'rigorous risk evaluation' of social engineering risks rests on the untested assumption that behaviors and signals observed in the adapted Prisoner's Dilemma and Ultimatum Game with role-conditioned LLM agents serve as valid proxies for real-world tactics such as manipulation or trust exploitation. Only descriptive statistics on 15 participants are supplied; no expert annotation, external validation, or mapping to SE-relevant outcomes is presented to support this link.

    Authors: We agree that the manuscript supplies only descriptive statistics and does not include expert annotation, external validation, or explicit mappings from observed behaviors to real-world social engineering outcomes. The central positioning is that the platform and dataset provide controlled, configurable game templates (adapted Prisoner's Dilemma and Ultimatum Game), role-conditioned agents, structured interaction trees, and synchronized multimodal streams to enable future researchers to conduct such evaluations across human-human, human-AI, and AI-AI settings. These games were selected because prior literature has linked them to constructs such as trust, cooperation, and exploitation, but we do not assert that the collected signals constitute validated proxies. The abstract's phrasing that the resource enables 'rigorous risk evaluation' can be read as overstating immediate applicability. We will revise the abstract, introduction, and dataset description sections to state explicitly that the contribution supplies a foundational resource and data for subsequent validation studies rather than performing or claiming to have performed that validation. This change will be incorporated in the revised manuscript. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

Dataset release paper with no derivations, predictions, or self-referential steps

full rationale

The paper introduces an open platform and dataset (AIriskEval-gaming) for collecting multimodal interaction data in adapted games with LLM agents. No equations, model fitting, predictions, or derivations are present in the abstract or described content. The work consists of platform description, data collection from 15 participants, and descriptive statistics; the central claim is the resource release itself rather than any computed result that could reduce to its inputs. No self-citations, ansatzes, or uniqueness claims are invoked as load-bearing elements. This is a standard non-circular dataset paper.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The contribution rests on a domain assumption about the diagnostic value of the chosen games and biometrics rather than new free parameters or invented entities; no fitted constants or theoretical derivations appear in the abstract.

assumptions (1)
  • domain assumption Multimodal biometric and behavioral signals collected during role-conditioned game interactions can serve as indicators of social engineering vulnerabilities in AI-mediated exchanges
    This premise justifies the platform's design for risk evaluation as stated in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating Social Engineering Risks in AI-based Interaction using Biometrics and a Gaming Setup." pith.science (2026). https://pith.science/paper/EHNLRGPR

@misc{pith2026260617793,
  author       = {Pith},
  title        = {Pith review of: Evaluating Social Engineering Risks in AI-based Interaction using Biometrics and a Gaming Setup},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EHNLRGPR}},
  note         = {Machine review of arXiv:2606.17793}
}
read the original abstract

We introduce AIriskEval-gaming, an open platform and dataset to evaluate social engineering risks in LLM-mediated multimodal interaction through controlled games. It supports human-human, human-AI and AI-AI settings, combining configurable game templates, role-conditioned LLM agents, psychology-informed participant profiling, structured interaction trees, and synchronized behavioral and biometric acquisition, filtering, and deep-learning-based feature extraction. The dataset (AIriskEval-gaming-db) was collected from 15 participants who interacted with a role-conditioned GPT-5.4 agent in two concatenated games: an adapted Prisoner's Dilemma and an Ultimatum Game. It comprises 340 GB of raw and processed multimodal data across six streams: interaction logs, video, screen recordings, gaze logs, smartwatch signals, and game/questionnaire metadata. These data include interaction paths, written justifications, psychological profiles, subjective feedback, perceived counterpart identity, game outcomes, and derived behavioral, facial, and gaze features. Alongside the dataset, we provide descriptive analyses characterizing the resulting multimodal data. Rigorous risk evaluation is essential for the deployment of secure AI systems, as it enables the identification and mitigation of vulnerabilities, ensures the protection of sensitive data, and supports compliance with evolving regulatory and ethical standards in society. The dataset and related code are available on GitHub.

Figures

Figures reproduced from arXiv: 2606.17793 by the authors.

Figure 1
Figure 1. Overview of the ARES platform and pilot dataset. The figure illustrates the main contributions of the paper: (i) the ARES platform; (ii) psychology [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Simplified view of the game template and interaction tree used in the ARES pilot dataset. The adapted Prisoner’s Dilemma is shown at the top, and [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 2
Figure 2. Simplified view of the game template and interaction tree used in the AIriskEval-gaming dataset. The adapted Prisoner’s Dilemma is shown at the [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

40 extracted references · 40 canonical work pages

  1. [1]

    On the Feasibility of Using Multimodal LLMs to Execute AR Social Engineering Attacks,

    T. Bi, C. Yeet al., “On the Feasibility of Using Multimodal LLMs to Execute AR Social Engineering Attacks,” inProc. AAAI Conf. on Artificial Intelligence, vol. 40, no. 45, 2026, pp. 38 252–38 260

  2. [2]

    A Turing Test of Whether AI Chatbots are Behaviorally similar to Humans,

    Q. Mei, Y . Xie, W. Yuan, and M. O. Jackson, “A Turing Test of Whether AI Chatbots are Behaviorally similar to Humans,” inProc. of the National Academy of Sciences, vol. 121, no. 9, 2024, p. e2313925121

  3. [3]

    Adverse Re- actions to the Use of Large Language Models in Social Interactions,

    F. Dvorak, R. Stumpf, S. Fehrler, and U. Fischbacher, “Adverse Re- actions to the Use of Large Language Models in Social Interactions,” PNAS Nexus, vol. 4, no. 4, p. pgaf112, 2025

  4. [4]

    edBB-Demo: Biometrics and Behavior Analysis for Online Educational Platforms,

    R. Daza, A. Moraleset al., “edBB-Demo: Biometrics and Behavior Analysis for Online Educational Platforms,” inProc. AAAI Conf. on Artificial Intelligence (Demonstration), 2023, pp. 16 422–16 424

  5. [5]

    A Multimodal Dataset for Understanding the Impact of Mobile Phones on Remote Online Virtual Education,

    R. Daza, A. Becerra, R. Cobos, J. Fierrezet al., “A Multimodal Dataset for Understanding the Impact of Mobile Phones on Remote Online Virtual Education,”Scientific Data, vol. 12, no. 1, p. 1332, 2025

  6. [6]

    SMARTe-VR: Student Monitoring and Adaptive Response Technology for e-Learning in Virtual Reality,

    R. Daza, S. Lin, A. Morales, J. Fierrez, and K. Nagao, “SMARTe-VR: Student Monitoring and Adaptive Response Technology for e-Learning in Virtual Reality,” inACM I2M-MM, 2025

  7. [7]

    Cyber Coercion Detection Using LLM-Assisted Mul- timodal Biometric System,

    A. Almehmadi, “Cyber Coercion Detection Using LLM-Assisted Mul- timodal Biometric System,”Applied Sciences, vol. 15, p. 10658, 2025

  8. [8]

    Promises and Trust in Human–Robot Interaction,

    L. Cominelli, F. Feriet al., “Promises and Trust in Human–Robot Interaction,”Scientific Reports, vol. 11, no. 1, p. 9687, 2021

Show all 40 references
  1. [9]

    Playing repeated games with large language models,

    E. Akata, L. Schulzet al., “Playing repeated games with large language models,”Nature Human Behaviour, vol. 9, no. 7, pp. 1380–1390, 2025

  2. [10]

    Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing,

    M. Schmitt and I. Flechais, “Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing,”Artificial Intelligence Review, vol. 57, no. 12, p. 324, 2024

  3. [11]

    Evaluating Large Language Models’ Ability to Automate Spear Phishing,

    F. Heidinget al., “Evaluating Large Language Models’ Ability to Automate Spear Phishing,”Expert Systems with Applications, 2026

  4. [12]

    The machine psychology of cooperation: Can GPT models operationalize prompts for altruism, cooperation, competitive- ness, and selfishness in economic games?

    S. Phelpset al., “The machine psychology of cooperation: Can GPT models operationalize prompts for altruism, cooperation, competitive- ness, and selfishness in economic games?”J. Phys.: Complex., 2025

  5. [13]

    oTree—An Open-Source Platform for Laboratory, Online, and Field Experiments,

    D. L. Chen, M. Schonger, and C. Wickens, “oTree—An Open-Source Platform for Laboratory, Online, and Field Experiments,”Journal of Behavioral and Experimental Finance, vol. 9, pp. 88–97, 2016

  6. [14]

    nodeGame: Real-Time, Synchronous, Online Experiments in the Browser,

    S. Balietti, “nodeGame: Real-Time, Synchronous, Online Experiments in the Browser,”Behavior Research Methods, vol. 49, no. 5, 2017

  7. [15]

    Empirica: a Virtual Lab for High- Throughput Macro-Level Experiments,

    A. Almaatouq, J. Beckeret al., “Empirica: a Virtual Lab for High- Throughput Macro-Level Experiments,”Behavior Research Methods, vol. 53, no. 5, pp. 2158–2171, 2021

  8. [16]

    LIONESS Lab: a Free Web-Based Platform for Conducting Interactive Experiments Online,

    M. Giamattei, K. S. Yahosseiniet al., “LIONESS Lab: a Free Web-Based Platform for Conducting Interactive Experiments Online,”Journal of the Economic Science Association, vol. 6, no. 1, pp. 95–111, 2020

  9. [17]

    HARMONIC: A Multimodal Dataset of Assistive Human– Robot Collaboration,

    B. A. Newman, R. M. Aronson, S. S. Srinivasa, K. Kitani, and H. Admoni, “HARMONIC: A Multimodal Dataset of Assistive Human– Robot Collaboration,”Int. J. Robot. Res., vol. 41, no. 1, pp. 3–11, 2022

  10. [18]

    Physiological Data for Affective Computing in HRI with Anthropomorphic Service Robots: the AFFECT-HRI Data Set,

    J. S. Heinischet al., “Physiological Data for Affective Computing in HRI with Anthropomorphic Service Robots: the AFFECT-HRI Data Set,” Scientific Data, vol. 11, no. 1, p. 333, 2024

  11. [19]

    MultiPhysio-HRC: A Multimodal Phys- iological Signals Dataset for Industrial Human–Robot Collaboration,

    A. Bussolan, S. Baraldoet al., “MultiPhysio-HRC: A Multimodal Phys- iological Signals Dataset for Industrial Human–Robot Collaboration,” Robotics, vol. 14, no. 12, p. 184, 2025

  12. [20]

    The Number of Distinct Basic Values and Their Structure Assessed by PVQ–40,

    J. Cieciuch and S. H. Schwartz, “The Number of Distinct Basic Values and Their Structure Assessed by PVQ–40,”Journal of Personality Assessment, vol. 94, no. 3, pp. 321–328, 2012

  13. [21]

    The Dispositional Essence of Proactive Social Preferences: The Dark Core of Personality Vis- `a-Vis 58 Traits,

    B. E. Hilbig, I. Thielmannet al., “The Dispositional Essence of Proactive Social Preferences: The Dark Core of Personality Vis- `a-Vis 58 Traits,” Psychological Science, vol. 34, no. 2, pp. 201–220, 2023

  14. [22]

    Who is Healthier? A Meta-Analysis of the Relations between the HEXACO Personality Domains and Health Outcomes,

    J. L. Pletzeret al., “Who is Healthier? A Meta-Analysis of the Relations between the HEXACO Personality Domains and Health Outcomes,”Eur. J. Pers., vol. 38, no. 2, pp. 342–364, 2024

  15. [23]

    The Development and Testing of a New Version of the Cognitive Reflection Test applying Item Response Theory (IRT),

    C. Primiet al., “The Development and Testing of a New Version of the Cognitive Reflection Test applying Item Response Theory (IRT),”J. Behav. Decis. Mak., vol. 29, no. 5, pp. 453–469, 2016

  16. [24]

    DeepFace-Attention: Multimodal Face Biometrics for Attention Esti- mation with Application to e-learning,

    R. Daza, L. F. Gomez, J. Fierrez, A. Morales, R. Tolosanaet al., “DeepFace-Attention: Multimodal Face Biometrics for Attention Esti- mation with Application to e-learning,”IEEE Access, vol. 12, 2024

  17. [25]

    Facial expressions as a vulnerability in face recognition,

    A. Pe ˜na, I. Serna, A. Morales, J. Fierrez, and A. Lapedriza, “Facial expressions as a vulnerability in face recognition,” inIEEE ICIP, 2021

  18. [26]

    Exploring facial expressions and action unit domains for Parkinson detection,

    L. F. Gomez, A. Moraleset al., “Exploring facial expressions and action unit domains for Parkinson detection,”PLoS One, vol. 18, no. 2, 2023

  19. [27]

    RetinaFace: Single-Shot Multi-Level Face Localisation in the Wild,

    J. Deng, J. Guoet al., “RetinaFace: Single-Shot Multi-Level Face Localisation in the Wild,” inProc. CVPR, 2020, pp. 5203–5212

  20. [28]

    MediaPipe: A Framework for Perceiving and Processing Reality,

    C. Lugaresiet al., “MediaPipe: A Framework for Perceiving and Processing Reality,” inProc. CVPR Workshops, 2019

  21. [29]

    WHENet: Real-time Fine-Grained Estimation for Wide Range Head Pose,

    Y . Zhou and J. Gregson, “WHENet: Real-time Fine-Grained Estimation for Wide Range Head Pose,” inProc. BMVC, 2020

  22. [30]

    OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis,

    J. Huet al., “OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis,” inProc. IEEE FG, 2025

  23. [31]

    mEBAL2 Database and Benchmark: Image-based Multispectral Eye- blink Detection,

    R. Daza, A. Morales, J. Fierrez, R. Tolosana, and R. Vera-Rodriguez, “mEBAL2 Database and Benchmark: Image-based Multispectral Eye- blink Detection,”Pattern Recognition Letters, vol. 182, pp. 83–89, 2024

  24. [32]

    Keystroke Verification Challenge (KVC): Biomet- ric and fairness benchmark evaluation,

    G. Stragapedeet al., “Keystroke Verification Challenge (KVC): Biomet- ric and fairness benchmark evaluation,”IEEE Access, vol. 12, 2024

  25. [33]

    BeCAPTCHA-Mouse: Synthetic Mouse Trajectories and Improved Bot Detection,

    A. Acien, A. Morales, J. Fierrez, and R. Vera-Rodriguez, “BeCAPTCHA-Mouse: Synthetic Mouse Trajectories and Improved Bot Detection,”Pattern Recognition, vol. 127, p. 108643, 2022

  26. [34]

    Personalized weight loss management through wearable devices and artificial intelli- gence,

    S. Romero-Tapiador, R. Tolosana, A. Moraleset al., “Personalized weight loss management through wearable devices and artificial intelli- gence,”Computers in Biology and Medicine, vol. 209, p. 111676, 2026

  27. [35]

    Overview of Risk Assessment and Management for Intelligent Systems under the AI Act and Beyond,

    J. Irigoyenet al., “Overview of Risk Assessment and Management for Intelligent Systems under the AI Act and Beyond,” inICCST, 2026

  28. [36]

    Is my vision-language data in your AI? membership inference test (MINT) Demo 2,

    D. DeAlcalaet al., “Is my vision-language data in your AI? membership inference test (MINT) Demo 2,” inIEEE COMPSAC, 2026

  29. [37]

    Leveraging avatar fingerprinting: A photorealistic talking-head public database and benchmark,

    L. Pedrouzoet al., “Leveraging avatar fingerprinting: A photorealistic talking-head public database and benchmark,”arXiv:2603.26934, 2026

  30. [38]

    Addressing bias in LLMs: Strategies and application to fair AI-based recruitment,

    A. Pe ˜naet al., “Addressing bias in LLMs: Strategies and application to fair AI-based recruitment,” inAAAI/ACM AIES, 2025

  31. [39]

    DeepID challenge of detecting synthetic manipu- lations in ID documents,

    P. Korshunovet al., “DeepID challenge of detecting synthetic manipu- lations in ID documents,” inIEEE ICCV Workshops, 2025

  32. [40]

    AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K- 12 Educational Explanations,

    J. Irigoyen, R. Daza, F. Jurado, J. Fierrez, R. Tolosana, A. Ortigosaet al., “AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K- 12 Educational Explanations,” inIEEE ICCST, 2026

Pith tools

Reviewed July 3, 2026 · model on record in the stance chip above.