{"id":"7624bec2-842b-43b4-b471-63854898aaf6","arxiv_id":"2606.17793","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces AIriskEval-gaming platform and 340 GB dataset collected from 15 participants in two game setups with role-conditioned GPT-5.4 agent to support evaluation of social engineering risks via synchronized behavioral and biometric streams.","lead":"The authors created an open platform called AIriskEval-gaming and a 340 GB multimodal dataset from 15 participants playing adapted Prisoner's Dilemma and Ultimatum Game against a role-conditioned LLM agent, recording interaction logs, video, gaze, and smartwatch signals. A smart generalist might read it to see a concrete example of how controlled games plus biometrics can be used to probe manipulation risks in AI conversations.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Validity of adapted PD/UG games as proxies for social engineering risks remains untested","rationale":"The reader's weakest_assumption directly identifies the proxy-validity gap; the full-text placeholder does not supply counter-evidence such as validation experiments or expert mapping, so the concern stands and moves the verdict from UNVERDICTED to CONDITIONAL pending that check.","tokens_in":1748,"tokens_out":316,"duration_ms":9246,"concrete_test":"Annotate a subset of the interaction logs for presence/absence of SE tactics (e.g., false persona, urgency, reciprocity pressure) using a rubric derived from established SE taxonomies; then test whether any of the reported biometric or behavioral features statistically discriminate those segments. If no reliable discrimination appears, the proxy assumption fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the platform and dataset enable evaluation of social engineering risks in LLM-mediated interaction. This requires that the chosen games (adapted Prisoner's Dilemma + Ultimatum Game), role-conditioned LLM agents, and biometric streams (gaze, smartwatch, facial features) actually surface SE-relevant behaviors such as manipulation, trust exploitation, or information elicitation. The abstract and dataset description provide only descriptive statistics on 15 participants and derived features; no mapping, expert annotation, or external validation shows that observed interaction paths or biometric signals correspond to SE tactics that would appear outside the game setting. Without this link, the dataset supports game-behavior analysis but not the stated risk-evaluation purpose.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces AIriskEval-gaming, an open platform and the AIriskEval-gaming-db dataset collected from 15 participants who played concatenated adapted Prisoner's Dilemma and Ultimatum Game scenarios against role-conditioned GPT-5.4 agents. It supplies 340 GB of synchronized multimodal streams (interaction logs, video, gaze, smartwatch, facial features) plus descriptive analyses of paths, outcomes, psychological profiles, and derived features, positioning the resource as enabling evaluation of social engineering risks in LLM-mediated interactions across human-human, human-AI, and AI-AI settings.","tokens_in":1865,"tokens_out":409,"duration_ms":14006,"significance":"If the proxy validity of the chosen games and biometric streams for social engineering behaviors can be established, the open release of configurable templates, structured interaction trees, and this large multimodal dataset on GitHub would constitute a useful contribution to HCI and AI-safety research by supporting controlled, reproducible studies of biometric and behavioral signals. The explicit provision of raw/processed data and code is a clear strength that facilitates follow-on work.","major_comments":[{"comment":"Abstract and dataset description: The central claim that the platform and dataset enable 'rigorous risk evaluation' of social engineering risks rests on the untested assumption that behaviors and signals observed in the adapted Prisoner's Dilemma and Ultimatum Game with role-conditioned LLM agents serve as valid proxies for real-world tactics such as manipulation or trust exploitation. Only descriptive statistics on 15 participants are supplied; no expert annotation, external validation, or mapping to SE-relevant outcomes is presented to support this link.","section":"Abstract"}],"minor_comments":[{"comment":"The description of the six data streams and deep-learning-based feature extraction would benefit from an explicit table listing per-stream sampling rates, synchronization method, and any filtering steps applied before feature extraction.","section":"Dataset description"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive feedback on the positioning of our contribution. The work introduces an open platform and multimodal dataset using established game-theoretic scenarios to support research on social engineering risks in LLM-mediated interactions. We address the major comment below by clarifying scope and committing to revisions where appropriate.","responses":[{"response":"We agree that the manuscript supplies only descriptive statistics and does not include expert annotation, external validation, or explicit mappings from observed behaviors to real-world social engineering outcomes. The central positioning is that the platform and dataset provide controlled, configurable game templates (adapted Prisoner's Dilemma and Ultimatum Game), role-conditioned agents, structured interaction trees, and synchronized multimodal streams to enable future researchers to conduct such evaluations across human-human, human-AI, and AI-AI settings. These games were selected because prior literature has linked them to constructs such as trust, cooperation, and exploitation, but we do not assert that the collected signals constitute validated proxies. The abstract's phrasing that the resource enables 'rigorous risk evaluation' can be read as overstating immediate applicability. We will revise the abstract, introduction, and dataset description sections to state explicitly that the contribution supplies a foundational resource and data for subsequent validation studies rather than performing or claiming to have performed that validation. This change will be incorporated in the revised manuscript.","revision_made":"yes","referee_comment":"[Abstract] Abstract and dataset description: The central claim that the platform and dataset enable 'rigorous risk evaluation' of social engineering risks rests on the untested assumption that behaviors and signals observed in the adapted Prisoner's Dilemma and Ultimatum Game with role-conditioned LLM agents serve as valid proxies for real-world tactics such as manipulation or trust exploitation. Only descriptive statistics on 15 participants are supplied; no expert annotation, external validation, or mapping to SE-relevant outcomes is presented to support this link."}],"tokens_in":1382,"tokens_out":399,"duration_ms":21727,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this paper ships an open platform and dataset called AIriskEval-gaming. It combines configurable game templates, role-conditioned LLM agents, psychology profiling, and synchronized capture of interaction logs, video, gaze, facial features, and smartwatch signals. Data comes from 15 participants playing adapted Prisoner's Dilemma and Ultimatum Game against a GPT agent, totaling 340 GB, with code and data on GitHub plus some descriptive stats.\n\nWhat works is the concrete release. Putting the pieces together in one documented setup and making the raw and processed streams available is useful for anyone who wants to run or analyze similar controlled interactions. The structured interaction trees and multimodal synchronization look practical.\n\nThe soft spot is the gap between the stated purpose and what is actually shown. The paper frames the work as enabling evaluation of social engineering risks in LLM-mediated interaction, yet only descriptive analyses appear. No validation, expert mapping, baseline comparisons, or external checks demonstrate that game paths or biometric signals correspond to manipulation or trust exploitation tactics outside the lab. The stress-test concern holds: the proxy link remains untested. With n=15 the data is exploratory at best.\n\nThis is for HCI or AI safety researchers who need an open multimodal game dataset to build on. A reader looking for ready-to-use recordings of human-AI game play could extract value; someone expecting a validated risk measurement tool will not find it.\n\nIt deserves peer review. The dataset contribution is real and documented even if the risk-evaluation framing needs work.","headline":"Releases a new open multimodal dataset from LLM-mediated games with biometrics, but provides no evidence that the setup measures social engineering risks.","tokens_in":2381,"tokens_out":382,"would_cite":false,"duration_ms":15026,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A gaming platform and 340 GB dataset evaluates social engineering risks in interactions with AI agents via biometrics and controlled games.","keywords":["social engineering","AI risks","multimodal dataset","gaming platform","biometrics","LLM interaction","Prisoner's Dilemma"],"falsifier":"A direct comparison study in which the same participants show no distinguishable biometric or behavioral patterns when exposed to actual social-engineering prompts outside the game setting versus neutral prompts would indicate the platform does not capture the targeted risks.","tokens_in":2643,"feed_emoji":"🎮","tokens_out":631,"duration_ms":15121,"temperature":0.7,"pith_summary":"The paper introduces AIriskEval-gaming as an open platform that runs human-AI and other interaction settings inside two concatenated games: an adapted Prisoner's Dilemma and an Ultimatum Game. Role-conditioned LLM agents interact with 15 participants while six synchronized data streams capture logs, video, gaze, smartwatch signals, and metadata. The resulting dataset supplies interaction paths, psychological profiles, game outcomes, and derived behavioral and biometric features for analysis. Descriptive statistics characterize the collected multimodal signals. The work positions this resource as a means to identify vulnerabilities that affect secure AI deployment.","feed_headline":"Gaming platform and dataset test AI social engineering risks","feed_subtitle":"AIriskEval-gaming records 340 GB of biometric and behavioral data from 15 participants across two classic games to probe vulnerabilities.","key_machinery":"AIriskEval-gaming platform, which integrates configurable game templates, role-conditioned LLM agents, psychology-informed profiling, and synchronized multimodal data streams for risk evaluation.","core_discovery":"The paper establishes AIriskEval-gaming as a configurable platform and accompanying dataset that supports controlled evaluation of social engineering risks in LLM-mediated multimodal interaction by combining game templates, role-conditioned agents, participant profiling, structured interaction trees, and synchronized acquisition of behavioral and biometric streams.","pith_inferences":["Patterns extracted from the six data streams may later serve as early indicators for training detection models that flag manipulation attempts in live AI conversations.","The platform's structure could be reused with additional participant cohorts to test whether biometric signatures generalize beyond the initial 15-person sample.","AI-AI game sessions in the dataset might surface interaction dynamics that differ from human-AI cases and warrant separate risk modeling."],"forward_implications":["Enables systematic identification of vulnerabilities in AI systems before wider deployment.","Provides data to support protection of sensitive information during AI interactions.","Facilitates compliance checks against evolving regulatory requirements for AI security.","Allows comparative testing across human-human, human-AI, and AI-AI configurations."],"fun_headline_variants":["AIriskEval-gaming platform evaluates social engineering in LLM interactions","Biometric dataset from games assesses AI social engineering risks","Gaming platform collects synchronized biometric data for AI risk analysis","Structured interaction trees enable controlled evaluation of AI social risks"],"cache_read_input_tokens":64,"weakest_assumption_plain":"Gameplay in the adapted Prisoner's Dilemma and Ultimatum Game with role-conditioned AI agents, together with the recorded biometric streams, serves as a valid proxy for social engineering risks that occur in wider real-world AI interactions.","fun_headline_variants_meta":{"raw":{"variants":["AIriskEval-gaming platform evaluates social engineering in LLM interactions","Biometric dataset from games assesses AI social engineering risks","Gaming platform collects synchronized biometric data for AI risk analysis","Structured interaction trees enable controlled evaluation of AI social risks"]},"model":"grok-4.3","cost_usd":0.007506,"raw_usage":{"total_tokens":3441,"prompt_tokens":662,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":75062000,"prompt_tokens_details":{"text_tokens":662,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2716,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":662,"tokens_out":63,"duration_ms":15539,"temperature":1.0,"reasoning_tokens":2716,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-03T23:56:06.655423+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A direct comparison study in which the same participants show no distinguishable biometric or behavioral patterns when exposed to actual social-engineering prompts outside the game setting versus neutral prompts would indicate the platform does not capture the targeted risks.","supporting_citations":[],"review_version":2}