REVIEW 4 major objections 5 minor 14 references
SEAR: A Multimodal Dataset for Analyzing AR-LLM-Driven Social Engineering Behaviors
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read In a controlled study, a pipeline combining AR glasses, a multimodal LLM, and a social agent got 93.3% of participants to say they would click a photo link after a single personalized conversation, and 76.7% reported higher trust afterward.
desk verdict Useful new dataset, but the headline '93.3% phishing link clicks' is self-reported intention, not observed behavior; the dataset deserves review but the efficacy claims need reframing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The SEAR framework has three components: AR glasses that extract facial landmarks, speech transcripts, posture, and objects to build live context; a multimodal LLM that turns scraped public social-media data into a personal profile with name, contact info, interests, and life events; and a social agent that uses that profile to drive a three-stage conversation (opening, deeper conversation, trust-building topic) with iterative response refinement. The dataset records the synchronized AR video and audio, the profile, the dialogue, and the post-interaction survey. The machinery's key work is converting ambient and profile cues into personalized conversational moves that appear sincere.
What would settle it
Run a second study where, after the same SEAR conversation, participants are actually sent a photo link and a phone call and their real click and call behavior is logged; if actual compliance is far below the 93.3% and 85% self-reported rates, the central efficacy claim fails.
Extended reading notes
Core claim
The paper's core discovery is that adding real-time multimodal context to an LLM-driven conversation lets an attacker personalize a trust-building pitch enough to get high stated compliance from strangers. In three experimental conditions—bare conversation, AR plus multimodal LLM, and the full SEAR pipeline with a social agent—the full pipeline produced the largest share of positive experience ratings and the highest willingness to comply with requests such as opening photo links, adding contacts, opening SMS, and answering calls. The authors interpret this as evidence that AR-LLM systems can hijack trust quickly and uniformly across communication channels. All compliance figures are self-reported intentions captured immediately after the interaction.
Load-bearing premise
The study treats participants' post-conversation survey answers about whether they would click a link or answer a call as a measure of actual compliance; if people say yes in a survey but behave differently when the real link or call arrives, the reported vulnerability rates are overstated.
Editorial extensions
If this is right
- If the reported effectiveness holds, AR-LLM social engineering becomes a credible attack channel that combines the reach of phishing with the personalization of a live conversation.
- Defenders would need detection systems that analyze multimodal dialogue and AR context, not just email text, to flag manipulation patterns.
- The dataset gives a benchmark for comparing attack strategies and for testing defensive interventions such as real-time warnings in AR glasses.
- A single conversation appears sufficient to shift trust judgments, so defensive training and system design constraints need to assume fast trust formation.
Reading between the lines
- The headline percentages measure stated willingness, not observed clicks or calls; a behavioral follow-up could change the numbers substantially.
- The lab setting and role-play design likely understate real-world distractions and overstate the attacker's ability to hold attention, so generalization to live attacks is untested.
- A natural next experiment is to vary the social profile richness (for example, no photos or only a name) to quantify how much of the effect comes from personalization versus mere conversational fluency.
- The same pipeline could be repurposed as a defensive training tool, letting people practice resisting personalized attacks in AR.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces SEAR, a multimodal dataset of 180 AR-mediated conversations collected from 60 participants across three conditions (baseline conversation, AR + multimodal LLM, and the full SEAR pipeline with a social agent). The dataset includes synchronized AR video/audio, extracted facial/pose/audio cues, environmental context, social-media profiles, LLM-generated social profiles, interactive conversation transcripts, and post-interaction survey responses. The headline results are high self-reported compliance rates (93.3% for photo-link clicking, 85% for phone calls) and a 76.7% post-interaction trust increase, which the authors interpret as evidence that AR-LLM-driven social engineering is alarmingly effective. The authors position the dataset as a benchmark for studying and defending against such attacks.
Significance. If the reported measurements were behavioral and the trust items were correctly specified, this would be a valuable community resource: it is one of the first multimodal, synchronized datasets of AR-LLM-based social engineering, with an ablation design and IRB oversight. The paper's main risk is that its central efficacy claims rest on self-reported intentions rather than observed actions, and the printed questionnaire contains a likely typo that makes the trust-change metric uninterpretable. The dataset release and the explicit ethical safeguards are strengths, but the headline percentages as currently worded overstate the evidence.
major comments (4)
- [Abstract, §2.5, §3.2] The headline rates '93.3% phishing link clicks' and '85% call acceptance' are not supported by the measurement instrument described in Section 2.5. The Social Engineering Effectiveness Questions ask participants whether they 'will' click a shared photo link, add a friend, open an SMS, or answer a call; these are hypothetical action intentions, not logged behaviors. Section 3.2 correctly says 'indicated willingness to click email photo links' in one place, but the abstract and the opening of Section 3.2 report these as actual 'phishing link clicks' and 'call acceptance.' The manuscript should either label all such numbers as self-reported intentions throughout, or add behavioral outcome logs (e.g., sent links/SMS/calls and recorded clicks/answers) to the dataset and analysis.
- [§2.5 Trust-Before/Trust-After] The two trust questions are printed with identical wording: 'How much do you trust the person before you have the conversation?' for both Trust-Before and Trust-After. If the questionnaire actually used different before/after phrasings, the manuscript must state the correct items; as printed, the 76.7% trust surge claim in Section 3.2 cannot be interpreted.
- [§2 (study design), §3.2] The causal claim that SEAR 'boosted' compliance lacks a proper comparison. The design includes baseline and AR+LLM conditions, but Figure 8 only reports overall experience ratings, not the four action-intention items or trust-change metrics per condition. Reporting photo-link, social-app, SMS, and phone-call intention rates (and trust scores) for all three conditions is necessary to support the claim that the full SEAR pipeline increases efficacy beyond the ablations.
- [§3.2] All efficacy percentages are presented as point estimates without confidence intervals, significance tests, or participant-level variation. Given n=60 and the dramatic textual claims about 'uniformly' effective persuasion, the paper should provide at least basic inferential statistics or effect sizes, or temper the claim accordingly.
minor comments (5)
- [§2.5, §3.2] The survey item refers to 'shared photo links' and 'email photo links' interchangeably; clarify whether the link is sent by email, social app, or SMS.
- [Figure 8, §3.2] Figure 8's legend uses Great/Good/Fine/Ok/Bad, while the text describes 'Very Good', 'Fairly Good', 'Average', and 'Fairly Bad'; align the scale labels between text and figures.
- [§3.1] The text states the average age is 34 and later says ages are concentrated between 23 and 37, while Figure 6 shows a participant aged 62; clarify the summary statistics (e.g., mean vs median) and whether outliers are included.
- [Throughout] There are typographical and grammatical errors, including 'Zhao et. el.' (should be 'Zhao et al.') and 'lacking of datasets' (should be 'a lack of datasets').
- [§2.5] The eleven subjective dimensions are listed with parenthetical example questions; consider adding the actual Likert scale labels and ranges used (e.g., 1–5) for reproducibility.
Circularity Check
No derivation-chain circularity; the headline percentages are self-reported survey summaries relabeled as observed behavior, which is a measurement-validity gap rather than a circular reduction.
full rationale
The paper contains no fitted parameters, no equations, and no self-citation chain that forces its conclusions. The central 93.3%, 85%, and 76.7% figures are descriptive tabulations of the Post-Interaction Survey; Section 2.5 explicitly defines the Social Engineering Effectiveness Questions as action intentions ("Will you click...", "Will you pick up phone call..."), and Section 3.2 reports them as "indicated willingness". The abstract's shorthand "phishing link clicks" and "call acceptance" overstates the construct, and the printed Trust-Before and Trust-After items are identical, which would make the trust-surge change score uninterpretable as printed unless the released questionnaire differs. These are data-integrity and external-validity concerns that a dataset or log check can resolve, but they are not circularity: no result is equivalent to its own input by construction, and no load-bearing argument rests on a self-citation (the only overlapping reference, SD-Eval, is used to motivate a gap, not to justify SEAR's claims). Hence a low circularity score with a flagged measurement caveat.
Assumptions & free parameters
assumptions (4)
- domain assumption Self-reported behavioral intentions are a valid proxy for real social engineering compliance.
- domain assumption Simulated role-play conversations generalize to real-world AR-LLM attacks.
- domain assumption The perceptual pipeline outputs (MediaPipe landmarks, YOLO objects, Vosk transcription) are accurate enough to drive effective attacks.
- domain assumption Alternating participant roles do not contaminate trust and susceptibility measurements.
Cite this review
Pith. "Pith review of SEAR: A Multimodal Dataset for Analyzing AR-LLM-Driven Social Engineering Behaviors." pith.science (2026). https://pith.science/paper/VZIXHMLZ
@misc{pith2026250524458,
author = {Pith},
title = {Pith review of: SEAR: A Multimodal Dataset for Analyzing AR-LLM-Driven Social Engineering Behaviors},
year = {2026},
howpublished = {\url{https://pith.science/paper/VZIXHMLZ}},
note = {Machine review of arXiv:2505.24458}
}
read the original abstract
The SEAR Dataset is a novel multimodal resource designed to study the emerging threat of social engineering (SE) attacks orchestrated through augmented reality (AR) and multimodal large language models (LLMs). This dataset captures 180 annotated conversations across 60 participants in simulated adversarial scenarios, including meetings, classes and networking events. It comprises synchronized AR-captured visual/audio cues (e.g., facial expressions, vocal tones), environmental context, and curated social media profiles, alongside subjective metrics such as trust ratings and susceptibility assessments. Key findings reveal SEAR's alarming efficacy in eliciting compliance (e.g., 93.3% phishing link clicks, 85% call acceptance) and hijacking trust (76.7% post-interaction trust surge). The dataset supports research in detecting AR-driven SE attacks, designing defensive frameworks, and understanding multimodal adversarial manipulation. Rigorous ethical safeguards, including anonymization and IRB compliance, ensure responsible use. The SEAR dataset is available at https://github.com/INSLabCN/SEAR-Dataset.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Junyi Ao, Yuancheng Wang, Xiaohai Tian, Dekun Chen, Jun Zhang, Lu Lu, Yuxuan Wang, Haizhou Li, and Zhizheng Wu. 2025. SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words. arXiv:2406.13340 [cs.CL] https://arxiv.org/abs/2406.13340
arXiv 2025
-
[2]
Leyla Bilge, Thorsten Strufe, Davide Balzarotti, and Engin Kirda. 2009. All your contacts are belong to us: automated identity theft attacks on social networks. In Proceedings of the 18th international conference on World wide web. 551–560
work page 2009
-
[3]
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić. 2020. MultiWOZ – A Large- Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling. arXiv:1810.00278 [cs.CL] https://arxiv.org/abs/1810.00278
arXiv 2020
-
[4]
Marcus Butavicius, Kathryn Parsons, Malcolm Pattinson, and Agata McCormac
-
[5]
Lindsey Choo. 2025. How 2 Students Used The Meta Ray-Bans To Access Personal Information. https://www.forbes.com/sites/lindseychoo/2024/10/04/meta-ray- bans-ai-privacy-surveillance/
work page 2025
-
[6]
Polra Victor Falade. 2023. Decoding the threat landscape: Chatgpt, fraudgpt, and wormgpt in social engineering attacks.arXiv preprint arXiv:2310.05595(2023)
arXiv 2023
-
[7]
Grant Ho, Asaf Cidon, Lior Gavish, Marco Schweighauser, Vern Paxson, Stefan Savage, Geoffrey M Voelker, and David Wagner. 2019. Detecting and characteriz- ing lateral phishing at scale. In28th USENIX security symposium (USENIX security 19). 1273–1290
2019
-
[8]
Joseph O’Hagan, Pejman Saeghe, Jan Gugenheimer, Daniel Medeiros, Karola Marky, Mohamed Khamis, and Mark McGill. 2023. Privacy-enhancing technology SEAR: A Multimodal Dataset for Analyzing AR-LLM-Driven Social Engineering Behaviors ACM MM, 2025, Dublin, Ireland and everyday augmented reality: Understanding bystanders’ varying needs for awareness and consen...
work page 2023
Show all 14 references
-
[9]
Sayak Saha Roy, Poojitha Thota, Krishna Vamsi Naragam, and Shirin Nilizadeh
-
[10]
Daniel Timko, Daniel Hernandez Castillo, and Muhammad Lutfor Rahman. 2025. Understanding Influences on SMS Phishing Detection: User Behavior, Demo- graphics, and Message Attributes. (2025)
2025
-
[11]
Yicheng Zhang, Carter Slocum, Jiasi Chen, and Nael Abu-Ghazaleh. 2023. It’s all in your head (set): Side-channel attacks on{AR/VR} systems. In32nd USENIX Security Symposium (USENIX Security 23). 3979–3996
2023
-
[12]
Yiqin Zhao, Sheng Wei, and Tian Guo. 2022. Privacy-preserving reflection render- ing for augmented reality. InProceedings of the 30th ACM International Conference on Multimedia. 2909–2918
2022
-
[2016]
arXiv:1606.00887 [cs.CY] https://arxiv.org/abs/1606.00887
Breaching the Human Firewall: Social engineering in Phishing and Spear- Phishing Emails. arXiv:1606.00887 [cs.CY] https://arxiv.org/abs/1606.00887
-
[2024]
In2024 IEEE Symposium on Security and Privacy (SP)
From chatbots to phishbots?: Phishing scam generation in commercial large language models. In2024 IEEE Symposium on Security and Privacy (SP). IEEE, 36–54
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.