Pith. sign in

Paper Citation Record · LEDGER

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 11 inbound Pith citation observations for arXiv:2506.01616.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01616 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:42:34.334089Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:34:41.371201Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:28:07.459324Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0192ec2-8633-4c0c-a7e7-c28e2de56f8e · outbound

This paper cites GPT-4o System Card.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments GPT-4o System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.631650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.631650Z digest=sha256:5040a7dd79729a2fdeafa29f76be64c6a0a213a7bcca3f24a820ad1abcd4c891

Observation 51fbccb8-2d75-4f0e-bb85-4851de27047a · outbound

This paper cites Gemini 2.0 flash model card,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Gemini 2.0 flash model card,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.656884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:32.651030Z digest=sha256:acbe0ccf412b0db4f82d146854cb5c40301ba26ee3440d9c1cabb683f62d0703

Observation 1f42dbaf-c2b6-45b9-a41e-257e98bef35a · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments The claude 3 model family: Opus, sonnet, haiku,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.640303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:32.672386Z digest=sha256:d38c840003257f3875b5dbe2bced3e3d7103dd92107e6abb2a729439e8d838b0

Observation 90b58b83-1de5-45ec-880c-1581d6b3e514 · outbound

This paper cites Pixtral 12B.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Pixtral 12B

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.693528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.693528Z digest=sha256:056a25cab984c793cb97943849d809abe9b29e5ccc8bf55dad47d9fd0c10a68d

Observation d94bb0c9-0ddc-4240-ae38-b04ffcef4434 · outbound

This paper cites Chat with the Environment: Interactive Multimodal Perception Using Large Language Models.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Chat with the Environment: Interactive Multimodal Perception Using Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.714754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.714754Z digest=sha256:1a1b914f9d227ab718bdacb0413ae3df4126b41346f1d4b60d407673abdf29f0

Observation d6cd318a-e863-4908-8f01-2975ea30476c · outbound

This paper cites Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.742298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.742298Z digest=sha256:28f46c08f2c015e64bd601c5996c86af6d723e2252e2c818dc79e00974772fef

Observation 59419a14-fc70-47db-ace9-0f0f706761aa · outbound

This paper cites VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.768449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.768449Z digest=sha256:2f2e1c6b168739cc1f334ffe9b400ae1f7d49d2474a9c64e711cd7d915cd0461

Observation 8d38a11e-35f0-43d8-a017-45eca82fbdc1 · outbound

This paper cites Large Multimodal Agents: A Survey.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Large Multimodal Agents: A Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.798379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.798379Z digest=sha256:f839d646358ddca2c423f1686d989919a6b0808a2bce69f3005120a63f561191

Observation 4c05c70e-0a23-4e0f-be7a-9fe7db43e7cc · outbound

This paper cites Agent AI: Surveying the Horizons of Multimodal Interaction.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agent AI: Surveying the Horizons of Multimodal Interaction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.859252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.859252Z digest=sha256:13c0651c7b27efd42332c05a0159b6c262c8ca248b6ff5776b4d7cdd9e35cfb2

Observation 1c0f580c-e3ad-480a-b9ef-86ce9a0428c5 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.881419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.881419Z digest=sha256:da1134172f57839c8315aa8a3fac199251b84e2beac2511e38e2b04adc0fe4aa

Observation a01ae670-76c0-4411-9e86-1048cc89327c · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AppAgent: Multimodal Agents as Smartphone Users

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.902989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.902989Z digest=sha256:bad568aebdc83ca40137e90baf4b2e7d31ef480873a657e4de1fc2ea3f662a4e

Observation f6c743ab-99ed-49e8-ba22-12b1b4c53dbf · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.924240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.924240Z digest=sha256:33620b607f5707363a9576967623693c59f1950d9dc34570ee637868415d3a30

Observation 149c1d36-29e9-4477-9dd0-f1345bbf87dc · outbound

This paper cites Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.948785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.948785Z digest=sha256:7e1f0e0ea8d47c7030029306b4130b2ac901a1445a86e964b3c1900d8075b394

Observation b8b15dd0-fbfb-4c34-b722-8c5182ea0749 · outbound

This paper cites Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.973269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.973269Z digest=sha256:c6d2c634ca67047b56a88ea88e402813d7ee33415197d381d93fbabfce114932

Observation 2ea8d8c1-98c1-41c9-b0e1-8a77d39622d5 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:32.995013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:32.995013Z digest=sha256:4809949ec9adee56ef78364bb2182eccd70a3724dcb94ec1dbddfbd8fef4f9b1

Observation 58f43906-2a00-43a5-9296-71b30caad99b · outbound

This paper cites VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.016850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.016850Z digest=sha256:daf9664d63d4bdbf6552addfc4f0ae614e8c1827b924890300ff80848343de98

Observation 2b3f4c75-d2dd-4309-a2db-847a2f4f4cf2 · outbound

This paper cites WebCanvas: Benchmarking Web Agents in Online Environments.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments WebCanvas: Benchmarking Web Agents in Online Environments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.038850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.038850Z digest=sha256:f8d9f9ae6f3fd1113fd1e724984ccfe59ad038c07aa68697f2d8f8e4a0329756

Observation b6efd011-5a9b-4c64-9ad2-56c558dde720 · outbound

This paper cites Agentboard: An analytical evaluation board of multi-turn llm agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agentboard: An analytical evaluation board of multi-turn llm agents,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.623971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.077219Z digest=sha256:a51bfc60fdf5905c81f44b3e6b2158d5815a18b12ad0c357a4fc2390b51afea4

Observation 88115221-051d-428c-ad1e-1d8407fa7238 · outbound

This paper cites The Positive-Definite Completion Problem.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments The Positive-Definite Completion Problem

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:42:35.160742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.108586Z digest=sha256:9e174a5a27d4c5dfe9ba81cb094da61b8da659bb308627977d01195a39ff2766

Observation 213f8aa6-6738-490e-89d8-79fab063984b · outbound

This paper cites OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.129535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.129535Z digest=sha256:a50e5de585aef7b3686c3de2d75c31afe8fbd3bea35185afcde0ae1d274338b5

Observation 4de0b938-f923-4b0a-8e57-4b32cf438f03 · outbound

This paper cites Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.151323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.151323Z digest=sha256:3f305e5b33bb845bdd6aed64966ae944d05907b4191ac9a09926f497e264d11e

Observation 373eeb94-1219-4c1d-9fad-bde91477419e · outbound

This paper cites From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.179393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.179393Z digest=sha256:845a418278b62d435e352faf5722335fce4fc023bcb0e2f1e3d09a693ac8f4ce

Observation 444972f3-190f-4558-b62b-bf7ccd680d6c · outbound

This paper cites Agentdam: Privacy leakage evaluation for autonomous web agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agentdam: Privacy leakage evaluation for autonomous web agents,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.211463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.211463Z digest=sha256:82bd899b0513588608f3585f1b240b89f77c5d1a127e7b93dbbc394a2d3f5b06

Observation 76bd7d73-bebe-403b-81bf-2f66c536a251 · outbound

This paper cites Towards trustworthy gui agents: A survey,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Towards trustworthy gui agents: A survey,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.232181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.232181Z digest=sha256:234f23d441c25574bf19ff7005b89f3929838275cb4332fe43314f40a269331e

Observation 69a3638e-6bc9-45ab-91c3-b20690291bc4 · outbound

This paper cites Ai-powered robots can be tricked into acts of violence,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Ai-powered robots can be tricked into acts of violence,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.608567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.254892Z digest=sha256:cee037a1febd915c3deafd95ac2d528d7582b7508e324fef4dea100c1909f671

Observation ed95cab0-819a-4d59-b163-9176c90893d5 · outbound

This paper cites R-Judge: Benchmarking Safety Risk Awareness for LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.276048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.276048Z digest=sha256:cfa693ff45ffe7e81364b57791df65865b3d7de4ae59c6bf0a226c75e9221d6c

Observation 932cf7b5-82e0-4c12-9142-af15aa85b84b · outbound

This paper cites Identifying the Risks of LM Agents with an LM-Emulated Sandbox.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Identifying the Risks of LM Agents with an LM-Emulated Sandbox

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.307592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.307592Z digest=sha256:4775a53baf5c12777d84a87f135cfbc6d4495fea33925320790e1e3c7057b2d0

Observation 70068e19-bbd1-4515-b95c-696492d1afa9 · outbound

This paper cites AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.346782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.346782Z digest=sha256:e707e039b09d6f6a8f5bebdd9b0730fa1a7c3771b2f2ba92d5f8e2dadf4219f7

Observation a6a48b9f-64e0-43b3-b6fc-736c4a032d87 · outbound

This paper cites Aligned llms are not aligned browser agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Aligned llms are not aligned browser agents,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.591773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.381673Z digest=sha256:70c5426e43982f1fd5eb0c4ac14b6aa5f4a710677dbf415bf72e81b65563139a

Observation d01aae15-96e5-4d0b-9f1a-a9ea4b2c8db3 · outbound

This paper cites Figstep: Jailbreaking large vision-language models via typographic visual prompts,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Figstep: Jailbreaking large vision-language models via typographic visual prompts,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.401642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.401642Z digest=sha256:d7f7e4546ec5836f562fcbc7dac6b4ccff5a3f9c0101d079d4d06c65c73ddadd

Observation 448ccc0f-a26a-4c91-ab4a-9b4750d221e9 · outbound

This paper cites How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.433505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.433505Z digest=sha256:8a77cf14c7cb5fdb925da33aa0fc5dae1f8fac56973b9e7a3704d82d22ad3879

Observation 9e82d4df-4156-413d-8cb6-c20e668115c6 · outbound

This paper cites Red Teaming Visual Language Models.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Red Teaming Visual Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.455009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.455009Z digest=sha256:9971e8420e0c7111f69e2b1b34083af237e3a228accc4b9d88df721ada325e03

Observation a56f2271-a3df-4f32-87ea-508b0d3ac0af · outbound

This paper cites Multitrust: A comprehensive benchmark towards trustworthy multimodal large language models,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Multitrust: A comprehensive benchmark towards trustworthy multimodal large language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.565998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.476297Z digest=sha256:bdb5e37100aaa4bf9194df3136b7f41c2a3b1c7370ca69de641a7daf7d375958

Observation ee8890b0-3b36-492c-8951-bdce4cd1322a · outbound

This paper cites Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.498378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.498378Z digest=sha256:21418f8f77331e3660f18e5f5f30bb7745074af53c4b54051fd3ef7c25ea833c

Observation 9a59d0cc-58b3-45b4-98a5-4eaa03a6d04e · outbound

This paper cites Mobilesafetybench: Evaluating safety of autonomous agents in mobile device control,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Mobilesafetybench: Evaluating safety of autonomous agents in mobile device control,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.519896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.519896Z digest=sha256:8d3fbfd60c915a117e709ced82b9d367516c15e56d047eb4ed4bb46d121f5af2

Observation 223afa44-20b6-4b8e-858b-fb977338d684 · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.548948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.548948Z digest=sha256:ea73d96d1b00bd17b12ba52066d20f9c1272adefc2920b7f69ce465964863334

Observation 50298a47-6967-4c1b-befb-69fbfec31c04 · outbound

This paper cites Cognitive architec- tures for language agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Cognitive architec- tures for language agents,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.550663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.570249Z digest=sha256:3191dd881835da168a8c6ac3eee4ed044cd2f88fde95114a56cc29c58f0af9c4

Observation f64a28bd-0e6c-4b7d-9188-11c97fbbbf33 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Vipergpt: Visual inference via python execution for reasoning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.536128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.606591Z digest=sha256:2b24b0ddd2a22468452089601ab491fa5fabd76fad05a3e79531fb07ac527298

Observation db6e8ac8-d582-41c4-88cd-114c35d4d976 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Chameleon: Plug-and-play compositional reasoning with large language models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.521029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.643057Z digest=sha256:56e9e12d235a2d878fa928ce0affe3394b43e4142e6af954efb28cf23ebb68ef

Observation f8eabefb-c52c-482c-af84-86a1ee746690 · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.672003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.672003Z digest=sha256:0e7fc344c77773ba81e6fa45e321f72b0fb358f15427d1890c72336d41be2e79

Observation bd4011c7-7988-4248-8389-bf726766b610 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.692481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.692481Z digest=sha256:99ecd7e388b5a24d661f08429f06aaf8d3739964a14e6020c14c17df8ba14b79

Observation 6b589c43-ab41-4c9e-844f-8fcf586ed2dd · outbound

This paper cites An introduction to microsoft copilot,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments An introduction to microsoft copilot,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.495761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:33.714006Z digest=sha256:18b6784eab2885c5ddd408612a00919cc2b2e90f1c061a6d4bf21e4bed618dc9

Observation dd503f24-bdc1-4c00-8b14-00d50cca5717 · outbound

This paper cites GPT-4 Technical Report.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments GPT-4 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.735322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.735322Z digest=sha256:341fa2425ebd6d3f830e1245ee582059ef73de79e48bd39f9dc7cdaab2dd9ef6

Observation 78b27603-5ec6-4f11-ba85-22c42275fc17 · outbound

This paper cites Improving image generation with better captions,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Improving image generation with better captions,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.757941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.757941Z digest=sha256:761fc11ff32b3bc9ce7dcc7217a237c79836aed89213ee96e4455a7995b74441

Observation 55dd091a-4a61-48f6-9bcf-a79b6e1f5a87 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.780135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.780135Z digest=sha256:fdd26e731be5aed3d3b37296326fc386ff8815bfb8cf9864dcc40fb2bd5538bd

Observation 718ce5d8-1928-457d-a07a-b6561ba0ee61 · outbound

This paper cites Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.802260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.802260Z digest=sha256:798fa3492a751f419832d345c094e11ca102a17e4820fe479d20a18242e4317b

Observation ea697270-1886-4d70-a867-2a1346ad5472 · outbound

This paper cites Evaluating Cultural and Social Awareness of LLM Web Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Evaluating Cultural and Social Awareness of LLM Web Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.824262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.824262Z digest=sha256:202d6ea35b4fad80dd5c5a450aa8b9b2873eb1e821900ca24f38f4720140f1e5

Observation 69fe411c-866a-4592-9402-9765861f448f · outbound

This paper cites A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.844167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.844167Z digest=sha256:ff4c3ec708835e27682f498c8dee1337819d25cdae745bb70bad6459300c05c0

Observation c1a6ef74-4f30-436a-8be2-6312e39575ca · outbound

This paper cites InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.874483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.874483Z digest=sha256:6b3e82f48c89b90a2dfc5797b0e7dd7e064bdb242379ef9134f7b92fbaf7a74c

Observation 70dac990-ffe5-4634-9672-79ef68675753 · outbound

This paper cites EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.904793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.904793Z digest=sha256:441af71a987120bff7060aa2bf22b3c16a2e7332b35e29df24140c99d84f72a9

Observation 608181d0-4092-4d95-9951-316fc7a8fdc0 · outbound

This paper cites GPT-4V(ision) is a Generalist Web Agent, if Grounded.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments GPT-4V(ision) is a Generalist Web Agent, if Grounded

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.931526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.931526Z digest=sha256:08f38ca8b3a0a8c686f7413c98a82620d8cf3d289de8f925d01b13d4662c1040

Observation 9f7c962f-ba98-40db-a27c-ce9b3da6036c · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.953240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.953240Z digest=sha256:d30abe56ed245281db00656071643adfe9b1981dd0fc4e79440bbacad4724b06

Observation 453ea8eb-5968-4b02-90c7-81632c60e4aa · outbound

This paper cites Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.973959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.973959Z digest=sha256:ddf3dcda7367ce3c46f8e9561ccf660b4d1fdaadb0e3a236d488ebf3b77c1ddb

Observation a5588d20-6ca3-4cb0-8fc6-26247e9c7b34 · outbound

This paper cites ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.995822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:33.995822Z digest=sha256:9f57139924f1cb1938aefcb7b8b7c06927cc70e75ccd130cb7bb4cd572d6a827

Observation a7f792d5-1235-4c28-9446-690af4e73f9a · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.018149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.018149Z digest=sha256:a9ab857c1a6c1f8f74cfffb3393a4a32c70ae0ff8848861ca1c96c16e893641d

Observation feb6032e-6828-4491-bdee-121c427bae65 · outbound

This paper cites Dissecting adversarial robustness of multimodal lm agents,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Dissecting adversarial robustness of multimodal lm agents,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.467914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:34.043522Z digest=sha256:f01f8aa5a8ec98e0dac13d01e69fedf483baea521a5095ff1d3c3bf3e30e79f9

Observation a1b4f668-1c2c-4a02-8c9c-3535b008bb86 · outbound

This paper cites Exploring the robustness of decision-level through adversarial attacks on llm-based embodied models,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Exploring the robustness of decision-level through adversarial attacks on llm-based embodied models,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.451542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:34.064792Z digest=sha256:8a6cc054da640eee65a09e4124c3e0239c2d7fcfd712fa8f66aa7e8c1529f15d

Observation 2903d421-1a8e-4a1a-8b07-d67a46d176de · outbound

This paper cites AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.095527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.095527Z digest=sha256:4b0c14a2590fb8c8e648af99d79449694418303c033a13ad7d0dfd24c724d57e

Observation 9bf83d3d-5c2c-4003-a8f0-ffc334934005 · outbound

This paper cites AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.147867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.147867Z digest=sha256:00a574a5ba07e81fcb03b42daea567b1c0b0314b0875284166072db8ded834cd

Observation 96e618ba-0d0f-443e-a9e1-362f9ad95e1f · outbound

This paper cites A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.170693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.170693Z digest=sha256:9cb13faf19c392d96cf4bc5762ad722d6802fd684880c11c83534eb325f413e7

Observation d4f3d604-3347-46d1-9a55-af913c0b0a43 · outbound

This paper cites The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.196721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.196721Z digest=sha256:0c25b8b54eb4b241b30a46141b8acf4d39fa5ff502c582c1fddbd83653a99492

Observation 7d6a254d-c1a3-482e-b5e4-d212330de1f4 · outbound

This paper cites Large multi-modal models for strong performance and efficient deployment,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Large multi-modal models for strong performance and efficient deployment,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.436239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:34.218472Z digest=sha256:ba2a538d910ba58697561fb70100e5c484fac26dfb87ae6aadfb5f54313207ff

Observation ece3bf19-f58c-40d7-ab6a-baad12aec4c0 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.239494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.239494Z digest=sha256:97b74c94b68931cbab6f9a2b70c0370650bc43d8575f55badd5d79191f700253

Observation 35cf182c-a56d-466a-89a3-955c20040914 · outbound

This paper cites Qwen2.5-VL Technical Report.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Qwen2.5-VL Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.261444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.261444Z digest=sha256:f6c46a75b00f778be83d7a912324c911a8e34095c1f24ce353adb7cfacb38df4

Observation c8befabd-e994-4862-a7e2-f7d130544460 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments LLaVA-OneVision: Easy Visual Task Transfer

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.285617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.285617Z digest=sha256:6cef070a0ea72a956e4bb461bbd51590fc7fc763ad9a8353bbe86b61f0ee4c24

Observation e7a17782-87a1-4b0b-a5a5-5416f5b02d60 · outbound

This paper cites Phi-4 Technical Report.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Phi-4 Technical Report

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.312600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:42:34.312600Z digest=sha256:64c513444c316b460e037693151dd1f7afb1da4b6586e89528e1c0bd3120028c

Observation a1fe2f56-e41f-4c2b-bc8e-f52ccc6e6bf6 · outbound

This paper cites Elemente der exakten erblichkeitslehre. 1909,.

MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments Elemente der exakten erblichkeitslehre. 1909,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.421328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:42:34.334089Z digest=sha256:345a9f3c69adba51ca7c4519715823631c0c7ea267723330ab837f4d1513fafc

Pith citing papers

Observation 0ac778f4-c436-4eab-ab2e-8f7db66007f3 · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:42:13.989933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:99fffff09314a7a9cbbe9f8ab97389ee5ac7658513b52e4cb298d8fac3e01588

Observation 264fee2f-c180-46c9-9f0c-de7b6b451ea7 · inbound

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents cites this paper.

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:41.371201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:41.371201Z digest=sha256:73476bd9d92ac80bfcdae6f29483a2382c3b3ff1f95a79262af3fb3094e64c71

Observation 523b9fb1-1da7-4def-85b9-ddd66c16c86c · inbound

VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents cites this paper.

VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:06:42.621938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T18:06:12.349285Z digest=sha256:37bd0203c7e38cf068d0b50735c19c36b2a6aa5251c4c0760460e622a323dc64

Observation aec30d6b-d3d3-4614-91c4-54beb8bf53b5 · inbound

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation cites this paper.

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:11.014186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:11.014186Z digest=sha256:fe80b6f5a30d845123171fb814b13cebfbc00bd323b9b3b7b0e9bffc0d66546d

Observation df4f8f69-492f-4a02-8859-d6f5769654b6 · inbound

Red Teaming Large Reasoning Models cites this paper.

Red Teaming Large Reasoning Models MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:31:28.581073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T03:29:16.164423Z digest=sha256:48ff0956befee4e35e558696c51d680a5cb0a374c18a61f7205aebe13c73797a

Observation b5a7936f-cb66-4614-83a7-704b221795ca · inbound

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents cites this paper.

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:32:48.311462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T11:31:16.087579Z digest=sha256:4a6570af65501c4f291604066edc26c2ec5fac4f5bc78e78720cc1120f9374dd

Observation fd288ce0-6ab2-4df2-8f77-243cf9cb4db0 · inbound

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents cites this paper.

OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:11.643104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:51:54.310805Z digest=sha256:0ab68c3210f91315ecec0f5298cda36179f2c96ddc1fe0eabb7a04131a006d72

Observation 7b5d2f42-792d-4a94-ab49-0c71b8127d2e · inbound

Governance by Construction for Generalist Agents cites this paper.

Governance by Construction for Generalist Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:53:58.013167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T04:51:41.739358Z digest=sha256:9dee98057c9c8e0e470c7ace5236f5f2b17414af742a4842a058dcde8b8bb0a5

Observation 026806d4-6102-4363-8424-3a12856b2553 · inbound

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security cites this paper.

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:45:01.572187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T19:18:40.244556Z digest=sha256:06802fa40f331b2a5571e10f5b2b8619ccc75d97b18fd79d78fcb820a1e9dec5

Observation 32027b2f-b83e-496b-a0d2-92207bcd67a3 · inbound

CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents cites this paper.

CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:28:07.460850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:00:10.153170Z digest=sha256:00a55208f7fbf25cfc10f738b72375fae95df916783162ceac0a2c8da09528d8

Observation abc1ed42-32f8-4e7b-b357-2d381601b7ff · inbound

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion cites this paper.

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T11:50:48.700362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:50:48.700362Z digest=sha256:aaad51d1b4070f88f9fe94e4bc61eba908fd4223d380eec89a650d12cab55c58