Pith. sign in

Paper Citation Record · LEDGER

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization

As of 6 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 2 inbound Pith citation observations for arXiv:2604.09574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.09574 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T20:35:17.075890Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T02:04:51.936838Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T18:25:57.668176Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact10
  • verified fuzzy63
  • unresolved5
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 813df995-2c3e-4038-b2f2-580fd8f33bd5 · outbound

This paper cites Gpt-4 technical report.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Gpt-4 technical report

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.621691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:e0f2b35750c76d9f558a68d3afa82aca8b10aeacf390c882a5ad0a583c7cf2b4

Observation 7440554f-12af-4c91-89e7-8657b8a6c108 · outbound

This paper cites Gemini: A family of highly capable multimodal models.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Gemini: A family of highly capable multimodal models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.658886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:3bf397cf950e6df028f8ff9ecf4ecce30c2ee5a5230fcad5c3be6e2b2eb804a2

Observation 205d5136-0e2f-4924-bf67-bff4b684b442 · outbound

This paper cites Visual instruction tuning.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Visual instruction tuning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.689581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:72cb3918c25826e6a01c0e2cccecf2d50f3286f9bcc4f96a196dbf9570124127

Observation 4489f2cc-6fa8-4922-bf50-4d528463d9b9 · outbound

This paper cites Appagent: Multimodal agents as smartphone users.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Appagent: Multimodal agents as smartphone users

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.738773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:5eb8c5b0d063fe5ce9097e67f663f40411eb5d8df0c665b8bc440510d28fbc85

Observation 105422fd-a65f-48c5-94ae-e84eca2395db · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:19:28.048419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:3ca53a4202c0d475a550f8996936ed90b2c05171d2715a6134252a92d5f1beac

Observation 821cc3fd-17b5-4097-81ec-e8c9a5ba31e5 · outbound

This paper cites Cogagent: A visual language model for gui agents.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Cogagent: A visual language model for gui agents

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.724603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:6876d802e16922e0928e6b6596c51bcd273ba845e0b1cadc962234f76cd9000b

Observation 12860f18-c9b7-49f3-b57b-67af46946822 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Mind2web: Towards a generalist agent for the web

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.729609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:a5a5c79f53b0601f6f4b09c5b03eb6403305b052b123c76d40bc3e0e91bdf676

Observation 7c62b7da-0b3a-4587-8625-75f299ed1bbf · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.745945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:8011ed5c3b4441033c220e0c802a5bb38d6ca2fd653e725754d271a3504bd26e

Observation 862e1b01-ecaa-49be-8048-086df5ffbeb4 · outbound

This paper cites Superplatforms have to attack ai agents.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Superplatforms have to attack ai agents

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.661378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:5ec8713bf85793a34916c94057bdef33f234b1e16b301a90b69bb9a562d1fcd8

Observation 4d081fe3-5e99-41b1-8587-770e6a35cc2f · outbound

This paper cites What is your ai agent buying? evaluation, biases, model dependence, & emerging implications for agentic e-commerce.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization What is your ai agent buying? evaluation, biases, model dependence, & emerging implications for agentic e-commerce

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.616756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:2124350b2c3dc4973bb0005f2a9d4b40c470b9b344505cb87fcc87bdb062ffd4

Observation b8ca6cce-2c11-4257-bbcf-6b35a9b626f4 · outbound

This paper cites How can recommender systems benefit from large language models: A survey.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization How can recommender systems benefit from large language models: A survey

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.624237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:98c49621219a4b782e631930f711007f26200db08a0b49414dc17ae1afcd4dbc

Observation 05c71cfb-68eb-4473-a298-c0c03972154a · outbound

This paper cites Computing machinery and intelligence.Mind, 59(236):433–460.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Computing machinery and intelligence.Mind, 59(236):433–460

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.750505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:d789543481cc7eff94c7d7b5fc33344acb4f145d5d9c49ff024933e3a08c5870

Observation cbcd86fd-6774-4f8d-9207-5c9a9fc3cd8a · outbound

This paper cites Touch-based continuous mobile device authentication: State-of-the-art, challenges and opportunities.Journal of Network and Computer Applications, 191:103162.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Touch-based continuous mobile device authentication: State-of-the-art, challenges and opportunities.Journal of Network and Computer Applications, 191:103162

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.748307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:f9fd1d3ce3a4d657a48f113c997637e34a4a70e0eee72799f3730dc505c65207

Observation 49c14b86-b229-4462-baa2-2a0d21e136ef · outbound

This paper cites an unresolved cited work.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:36:35.666546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:4eedac177f38443c6fff1853f6bd76f07b6216d00014b6cb541626ade056e256

Observation 6cc3c084-e3fb-44b1-868a-9385c20170e9 · outbound

This paper cites AlQahtani, and Muhammad Khurram Khan.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization AlQahtani, and Muhammad Khurram Khan

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.611882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:902ea8d6906144c9e51545d7e932f2118af6f847b401cc2a4b17c915cc14f38c

Observation c322898a-7438-4b38-9913-edac983fd2d1 · outbound

This paper cites Princeton University Press, Princeton, NJ.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Princeton University Press, Princeton, NJ

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.669227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:4d09bc65fce07f3bc582df3ea7078a53ebd571a122a71ecba144be2565cef039

Observation 4e1ad2a6-1efc-4fcf-bb47-9964944a09dd · outbound

This paper cites Generative adversarial nets.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Generative adversarial nets

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.671225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:e0b34b6726001b281a6abebeca799116520bf72682e9bd13b5bd7e3bafe4bfc0

Observation 325d15d5-3843-4213-8ed2-f72f65a7fbff · outbound

This paper cites Ui-tars: Pioneering automated gui interaction with native agents.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Ui-tars: Pioneering automated gui interaction with native agents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.609032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:57a16bfd6b8f266fb9434a6fab19fcda610000c714751d8b6cfe9904d704c918

Observation a53c34e4-8b97-48ec-9824-cf4e72bf865c · outbound

This paper cites Mobile-agent-e: Self-evolving mobile assistant for complex tasks.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Mobile-agent-e: Self-evolving mobile assistant for complex tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.614193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:86eaeb021ec782622cbaafd77b831fd586b4ec22c231e76e0ab1d1a456a7bdbe

Observation 8d7d2450-6fb5-44ea-bb43-e05716791999 · outbound

This paper cites Agentcpm-gui: Building mobile-use agents with reinforcement fine-tuning.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Agentcpm-gui: Building mobile-use agents with reinforcement fine-tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.619050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:d0e51fffafa1d0d41aa2997b40ddc6d67456b114f29bf24dcc274a9e52db6dcb

Observation 56c463e0-d484-487f-b00c-a253b0e4ffc5 · outbound

This paper cites Autoglm: Autonomous foundation agents for guis.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Autoglm: Autonomous foundation agents for guis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.594162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:09335c56c6996046af2a543a0a2de204e46d8b2f6007113f876bf05549cbd58b

Observation 9d7e06c5-bc6a-44ea-9ecc-7432592b9abf · outbound

This paper cites an unresolved cited work.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:36:35.591061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:44dd4a7d6566f1f472a17e235b34aabfa17778302797e0ea03f32ca2b6877456

Observation eefea4bf-c026-4866-b3d9-b147e0e74c15 · outbound

This paper cites A mathematical theory of communication.The Bell system technical journal, 27(3):379–423.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization A mathematical theory of communication.The Bell system technical journal, 27(3):379–423

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.603431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:eb44382d7a4428e7e1055641e1b66a21118c65c79a14e1927ccde857355571dd

Observation 7632ce74-214b-4b9a-a3e2-f7bf8fd597ad · outbound

This paper cites Support-vector networks.Machine learning, 20(3):273–297.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Support-vector networks.Machine learning, 20(3):273–297

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.627084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:f9a55548dc737cd3f6f42508b88bb68754b47bbf7d58e1838f0421dc1cf573ec

Observation 7442d7f2-17bf-4c6f-add0-411697518679 · outbound

This paper cites Xgboost: A scalable tree boosting system.Cornell University.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Xgboost: A scalable tree boosting system.Cornell University

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.673426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:6704a0f7cd99bbd473fe65c5cf4969777776f7e199c82e43157f8cd4cc37cfb6

Observation f1484c49-0b6e-41de-8e42-7e5d134039df · outbound

This paper cites On calculating with b-splines.Journal of Approximation Theory, 6(1):50–62.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization On calculating with b-splines.Journal of Approximation Theory, 6(1):50–62

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.699558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:4439163444d47100ea710d69fcb3507c9b4cb43f1b2352e0af1d48004db9ac23

Observation 9840d22b-e02c-4fe8-a96d-130e3408b3fe · outbound

This paper cites Mobile-Agent-v3: Fundamental Agents for GUI Automation.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Mobile-Agent-v3: Fundamental Agents for GUI Automation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:00:48.337118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:47e3e1d407654993f8c3e36f043281ccdcff5135aae2ede9387bdfd1aed9507c

Observation 7ee9ba35-812f-4f16-bdd0-ee849b4fcc7a · outbound

This paper cites CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:36:35.223477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:e2d2fc52387027530eb8d6273be60b2bd75a78996f4249cc34992fdd144c16b4

Observation 872802e1-8ad3-478c-8b4d-160ad3c25d64 · outbound

This paper cites Mobileuse: A gui agent with hierarchical reflection for autonomous mobile operation.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Mobileuse: A gui agent with hierarchical reflection for autonomous mobile operation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:36:35.216682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:e15162930237ae9224039b031571902fd960d7f341c3bc8b3f0059f0e4ec3544

Observation e716440b-31ec-49c5-afcc-c887a300618b · outbound

This paper cites Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:36:35.237444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:37b9971be78f267a69e8ed1d68c565f2a28e9f00d6bb18e85812e36d4e5e24a2

Observation cc1575d0-a5e8-40c3-9da6-2fc47ee73da5 · outbound

This paper cites VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:36:35.226641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:774b6b2c66cef3d5c8e53f6d1ce6f56d8bb1456388545fb642412f860e4a569f

Observation b55bdfd3-e444-41b0-bf44-62b0c7678145 · outbound

This paper cites Os-kairos: Adaptive interaction for mllm-powered gui agents.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Os-kairos: Adaptive interaction for mllm-powered gui agents

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.704557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:7941dd84644f1b72c06d10cec90198a9e60b656785421086da3034a5d33495df

Observation de7c9427-68a7-4ee6-b576-a02d92b783f3 · outbound

This paper cites Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:36:35.240598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:b417eff4ed412803b86f98ea9eb243c36618d0201d1f3011e102220d09387bd3

Observation a00a8d54-94c6-4eff-8a4b-afedab8048e3 · outbound

This paper cites ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:36:35.212903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:f6e320e8170e216be5cf73c8be13d1c57c180a463c7959e5528e3e60580e9486

Observation 3b16d63d-2dae-4cc4-bb3b-33ba5da70c52 · outbound

This paper cites Mobilerl: Advancing mobile use agents with adaptive online reinforcement learning, 2025.URL https://github.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Mobilerl: Advancing mobile use agents with adaptive online reinforcement learning, 2025.URL https://github

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.727128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:d07c588a5b26d4b6a72932cc5ecf2f21a1598ef847a58d2e6ddce6ae5efa5e05

Observation 5ca1eaae-979c-4d1b-a729-75b5eac698fd · outbound

This paper cites The rise and potential of large language model based agents: A survey.Science China Information Sciences, 68(2):121101.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization The rise and potential of large language model based agents: A survey.Science China Information Sciences, 68(2):121101

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.762522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:19b46217e9b3112bb816934b4af910a52761e0a72bbdde54df109720b69199e9

Observation 260695cb-2a49-4b0d-bec9-efc1756c8175 · outbound

This paper cites Dissecting adversarial robustness of multimodal lm agents.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Dissecting adversarial robustness of multimodal lm agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.697026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:236f59ff3d00bb6c76090b200759ee821ba9b31ff457985a76a82966ed848d3b

Observation a4619541-8a85-428e-bba3-82b93596c0ba · outbound

This paper cites Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.720343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:65956bfcefb361608e4042c813985eba4ebecd3049b5bfc3156103755251d819

Observation 6dd90cd3-5347-4c87-a052-692e5581332f · outbound

This paper cites Advagent: Controllable blackbox red-teaming on web agents.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Advagent: Controllable blackbox red-teaming on web agents

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.683013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:8393399b17824e8b96c39710c39e6fab6aacd2b988171b1798d6bf6ada8850fe

Observation d4ca525f-6174-4944-8ebe-42c06a1b9dac · outbound

This paper cites Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:36:35.233975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:5e0df8d7f29315f4e9a8dd1804dc480fe551f5a6e38bd0b5f288d3ee0e633853

Observation 69b15233-0112-41ca-85df-338d162c1856 · outbound

This paper cites On the robustness of large multimodal models against image adversarial attacks.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization On the robustness of large multimodal models against image adversarial attacks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.640001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:390851f230c2804b070cda6db1c1a0f8c07fb37540005414cd1a3d3810316230

Observation c6d7064b-edcc-473f-9ad8-d105eea83a6d · outbound

This paper cites How Robust is Google's Bard to Adversarial Image Attacks?.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization How Robust is Google's Bard to Adversarial Image Attacks?

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:36:35.220145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:0446483d718a05fbe55ff8b61cf8cfde76862d29ba1e880df67fde78991001f8

Observation 20543ea7-5658-4f47-8d01-7298b7fcfe77 · outbound

This paper cites Eia: Environmental injection attack on generalist web agents for privacy leakage.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Eia: Environmental injection attack on generalist web agents for privacy leakage

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.702074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:20610f3ce48ca1734585c4aa30f643a1852f45f8088cab85a7700080c92d38e4

Observation ef0843b1-b82a-40a7-80bc-ffc03ec56760 · outbound

This paper cites Evaluating the robustness of multimodal agents against active environmental injection attacks.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Evaluating the robustness of multimodal agents against active environmental injection attacks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.694267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:aced8eb01ee11932208a12bf0c3a94e1d85cffa32c14babef9d4e698c9c95724

Observation b670936c-7241-4768-94f8-e535fb2fde3c · outbound

This paper cites The obvious invisible threat: Llm-powered gui agents’ vulnerability to fine-print injections.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization The obvious invisible threat: Llm-powered gui agents’ vulnerability to fine-print injections

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.651182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:3be34f355acdc1e42e2d291289e3df3df98604ac92cc59421f98d122aebb2140

Observation 532eace1-76b7-4c9c-bef2-7e74802796b1 · outbound

This paper cites Attacking vision-language computer agents via pop-ups.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Attacking vision-language computer agents via pop-ups

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.752624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:a837c8a49c1894dfd52389c3929aca17ba14084dc624b24badb155422cc89028

Observation e57f4e71-7340-4a97-9173-ebec127c8918 · outbound

This paper cites Clip-guided generative networks for transferable targeted adversarial attacks.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Clip-guided generative networks for transferable targeted adversarial attacks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.741178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:ffb0df8ab4a8ae5ad4018dd2631c0fa3bb5d33db92649c18f019435150abff27

Observation ddecb4b4-0dcf-4bf6-88d7-f5b3ad7e3d2c · outbound

This paper cites Qava: Query-agnostic visual attack to large vision-language models.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Qava: Query-agnostic visual attack to large vision-language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.648468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:4ebc55fac5d6025e0ea8f473545d59027c7596ca06ffddd3ae27e17eaa0a59e7

Observation dc0d18c4-5523-4a2a-8603-0c3b3df70484 · outbound

This paper cites Exploring the adversarial robustness of clip for ai-generated image detection.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Exploring the adversarial robustness of clip for ai-generated image detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.653824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:3707ac0f400c38b418ba5182fc0480ed0be618be9bd1ab6f5ccd401e7477c55b

Observation 21fcfefa-4fb9-4f5d-b122-d9b3fe3afed5 · outbound

This paper cites Badagent: Inserting and activating backdoor attacks in llm agents.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Badagent: Inserting and activating backdoor attacks in llm agents

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.656327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:f1874ed6510f100b39f95d427a2be456da32335a7b0609e7f33385ffd2dc8384

Observation 6be4dedd-a5d9-44c3-abe4-1346fae6e5c0 · outbound

This paper cites Watch out for your agents! investigating backdoor threats to llm-based agents.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Watch out for your agents! investigating backdoor threats to llm-based agents

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.664269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:78e7b4dffde825ec8c194c45666a1cff9b04d57bb23cd88a17a69892c97a9fcb

Observation c5a68547-a0a0-4099-baac-ee66562f2831 · outbound

This paper cites Foot-in-the-door: A multi-turn jailbreak for LLMs.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Foot-in-the-door: A multi-turn jailbreak for LLMs

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.633862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:8a7b4a9b6712621ba728af7a4a4e9559510f0653b04a6f354e4660ed038df1f8

Observation 643b1c95-c8af-4f04-ba17-7a66e54b6957 · outbound

This paper cites Sensor-based continuous authentication of smartphones’ users using behavioral biometrics: A survey.IEEE Access, 5:15226–15257.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Sensor-based continuous authentication of smartphones’ users using behavioral biometrics: A survey.IEEE Access, 5:15226–15257

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.722548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:fdbb53d2145ec3bf768bc80b9451112822113d87a5a197bd6b3a348d6be162ca

Observation 7a03e183-5d7e-4b4b-8e03-4da6af83f1f8 · outbound

This paper cites In27th USENIX Security Symposium (USENIX Security 18), pages 135–150.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization In27th USENIX Security Symposium (USENIX Security 18), pages 135–150

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.629496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:9c64b0444904b2ced9df81dbadf39618bfeabf2ee50478b507dd830ddda0b4c8

Observation 8532f823-f49f-4e38-b8e1-6e4fe544066b · outbound

This paper cites Browser fingerprinting: A survey.ACM Transactions on the Web (TWEB), 14(2):1–33.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Browser fingerprinting: A survey.ACM Transactions on the Web (TWEB), 14(2):1–33

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.637204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:3444252429373c95b1691b6f62e89fa37cc17b50baf14a7cdfc82c7a47d33fc4

Observation 7cc990a4-8487-4d47-aee2-268193db67b3 · outbound

This paper cites Continuous mobile authentication using touchscreen gestures.2012 IEEE Conference on Technologies for Homeland Security (HST), pages 451–456.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Continuous mobile authentication using touchscreen gestures.2012 IEEE Conference on Technologies for Homeland Security (HST), pages 451–456

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.642908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:f7fbfe4b2bff19997fd61ee2b02fe177b221a21f7840763d6d588f4b8867567b

Observation 13f10cbc-7ee4-480c-8956-f23aebbe3abf · outbound

This paper cites Kroeze and Katherine Mary Malan.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Kroeze and Katherine Mary Malan

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.645430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:e8580b7d86e587988e91cb395cbf40c401d988fc7f6049bffad462241146e384

Observation 4481aa08-2088-40b1-baf2-dd0a34f11eed · outbound

This paper cites Increauth: Incremental-learning-based behavioral biometric authentication on smartphones.IEEE Internet of Things Journal, 11:1589–1603.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Increauth: Incremental-learning-based behavioral biometric authentication on smartphones.IEEE Internet of Things Journal, 11:1589–1603

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.606197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:0fb63b5fdfd3545779a3cb99d2076554b4ceb724d7b609992476a6796bbd5f16

Observation 4cb0cd49-c73f-4f23-aebe-430eeed2c2e4 · outbound

This paper cites Mouse dynamics behavioral biometrics: A survey.ACM Computing Surveys, 56(6):1–33.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Mouse dynamics behavioral biometrics: A survey.ACM Computing Surveys, 56(6):1–33

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.678435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:a2e80bf2d788b1680ce08a4719d8efda16fd3ccb406f3d20cc9a8743eb0d4726

Observation baeb9d04-9249-4a1c-81fb-77a9ae7570f3 · outbound

This paper cites Game bot detection via avatar trajectory analysis.IEEE Transactions on Computational Intelligence and AI in Games, 2(3):162–175.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Game bot detection via avatar trajectory analysis.IEEE Transactions on Computational Intelligence and AI in Games, 2(3):162–175

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.734246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:ee7b64f8507a6f2e620aca38731096e733b4fe01d3e539b40ce5356bc31a9609

Observation ebdf6e56-7fc7-48c6-ae0e-9a6e8c2b0278 · outbound

This paper cites Forgery-resistant touch-based authentica- tion on mobile devices.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Forgery-resistant touch-based authentica- tion on mobile devices

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.764655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:c1640665e111183fcd5bee787d83b41f7379c5af52d7372b2d7eb1568495a038

Observation 1c71ac5a-f648-4d73-98ae-f94427769ba1 · outbound

This paper cites Toward robotic robbery on the touch screen.ACM Transactions on Information and System Security (TISSEC), 18(4):1–25.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Toward robotic robbery on the touch screen.ACM Transactions on Information and System Security (TISSEC), 18(4):1–25

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.757887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:ef28b1ff006924bea2689c1d055191a7f94daa71cffcabcab314084bba792a37

Observation 06aa9e0d-a404-4b2a-ad59-6066633f14c2 · outbound

This paper cites Gantouch: An attack-resilient framework for touch-based continuous authentication system.IEEE Transactions on Biometrics, Behavior, and Identity Science, 4(4):533–543.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Gantouch: An attack-resilient framework for touch-based continuous authentication system.IEEE Transactions on Biometrics, Behavior, and Identity Science, 4(4):533–543

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.731904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:63a0b8e36358b41b909af5bca0d0417b65394523c8b99094123707d91e880a96

Observation b0b54fb8-ee54-4d45-9dbd-63c50eea5ece · outbound

This paper cites A survey of ai agent protocols.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization A survey of ai agent protocols

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.736345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:f77fad9c1bee745830e5ff20db05d5175c7acda7ccdfb82916e2cbaaeab44d42

Observation de122bb7-0b20-49f3-b70c-a9721ad6b050 · outbound

This paper cites Agentic information retrieval.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Agentic information retrieval

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.743528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:14fc98738a4f52ff1262273212c99e29d0749d5f1095789d27fbf7a0a8facc51

Observation 0b0d319d-d5f3-43bb-a787-219024efde48 · outbound

This paper cites A survey of llm-based deep search agents: Paradigm, optimization, evaluation, and challenges.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization A survey of llm-based deep search agents: Paradigm, optimization, evaluation, and challenges

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.706942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:be9050eded633b80390b9b91ad406502cc93b2452da9c6b2f4259baab34d13f4

Observation 31c894d0-bf83-45b6-ad29-421df31c860e · outbound

This paper cites Evolutionary perspectives on the evaluation of llm-based ai agents: A comprehensive survey.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Evolutionary perspectives on the evaluation of llm-based ai agents: A comprehensive survey

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.755394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:f01e50ed3553bd632b6fd9ca8d1848d1f7f0bbff4e9b8a195d09406dadb5d359

Observation 5822b54b-8e57-4c2b-b71f-217379ba68df · outbound

This paper cites Long short-term memory.Supervised sequence labelling with recurrent neural networks, pages 37–45.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Long short-term memory.Supervised sequence labelling with recurrent neural networks, pages 37–45

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.687342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:d014fb8caba53d054d610a7babb4a916dbd8c09725e44589e9af7fb697f44d40

Observation 17415481-c5c9-414b-93c2-bb48ed97fbf5 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Attention is all you need.Advances in neural information processing systems, 30

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.691931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:079b4ad1d4e24d0c74eeaa8403dcdb2a0d225243f1d9132e157b6db7224e92ab

Observation 8f320a64-3416-4535-afbc-65f509ad7b43 · outbound

This paper cites see” the screen and simulate physical taps. This allows it to execute cross-app workflows without manual input, promising a “zero-touch.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization see” the screen and simulate physical taps. This allows it to execute cross-app workflows without manual input, promising a “zero-touch

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.766930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:d8980342531e600af18cb151a1a69d5c3cb40fe2e2c7cf6c3b510f08d9c222cc

Observation 5ab7c893-a788-49f8-a5da-1b7de7d7a798 · outbound

This paper cites They contend that since the user explicitly authorized the assistant, the AI acts as a legitimate digital proxy for human intent.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization They contend that since the user explicitly authorized the assistant, the AI acts as a legitimate digital proxy for human intent

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.760211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:f691854b35074ff3e68456c0e282088e49adbc2f90c8ef72198aee6c64109eaf

Observation fafa4539-7f15-4cb3-80cc-969a254a0d51 · outbound

This paper cites Turing Test on Screen.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Turing Test on Screen

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.711551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:bbea4ad4f447b55a85539a1f14276781db4bb156b1514fcccf7f1c797423d9df

Observation 0b3bc539-f7e7-4298-9f1e-bf7cb6b4d261 · outbound

This paper cites an unresolved cited work.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Unresolved cited work

Reference 73

Resolution
parse uncertain
raw_fallback, observed 2026-05-15T20:36:35.680698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:0f5c01ad36e89504911bc2af698c43923e6078dfebf6d7b02e6acb993d1f8bce

Observation 3504aebf-2922-4028-9030-f4009d4e53df · outbound

This paper cites an unresolved cited work.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:36:35.718197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:1aa13366ad39163d9594d181d950167522475fa73394a40f9889a22cd54d884c

Observation 6eb2a53c-1fe0-4c01-96d7-1c4b1684023d · outbound

This paper cites Completed contents.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Completed contents

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.716135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:ffaf347c1f69600fff48cbe8a75ba71c6d811bac46bfc2a861e986c66322da4b

Observation b1d70e28-a6c3-417d-974f-acc29f3fcb62 · outbound

This paper cites an unresolved cited work.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:36:35.713580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:1cad51d623bb543947daaefd0e414e5b97adb862544aa2bcd00f19b66210cec2

Observation f11cce20-a822-4bb0-a215-856eefc331a4 · outbound

This paper cites an unresolved cited work.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:36:35.709184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:e04b99abb993f6a875cfb4f651c15ae22f3758af9735eda97b301fd001828616

Observation bfa68072-b55c-4cd3-9c9e-2c5eeff77766 · outbound

This paper cites If the task is completed, use the "stop" action.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization If the task is completed, use the "stop" action

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.685121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:fe76480df4f096d3d61e9fbb46cf5f1109ca671334b37fdb289d1576a71145ac

Observation c56c72a0-c2af-43ff-87a9-c6f8b94f4c84 · outbound

This paper cites # Action Space - click(x, y): Tap the screen at normalized coordinates (x, y).

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization # Action Space - click(x, y): Tap the screen at normalized coordinates (x, y)

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:36:35.675581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:2cfd95ea6aca64c93125239cd77b19e548caadeb551216882ddaa665b87cd1df

Pith citing papers

Observation 1634e42e-a6d2-4040-a54d-ac2c32249646 · inbound

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability cites this paper.

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization

Reference 105

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:35:55.804574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:15:19.239355Z digest=sha256:e01df2e024d9688c8a30e4f042b9a731067d2b6e36e3a413089074f329ec580c

Observation a96b1a98-5d70-4779-a6a1-7abf5db9ce23 · inbound

It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents cites this paper.

It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-01T18:25:57.669351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T02:04:51.936838Z digest=sha256:fb384aa348accc48f5db868800a3385ebe15482328c692cdce7d859cc9a114cb