Pith. sign in

Paper Citation Record · LEDGER

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents

As of 18 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 1 inbound Pith citation observation for arXiv:2505.12632.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12632 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:34:31.209389Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T22:08:27.511727Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy38
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c3bbb260-5773-4d5b-9611-8240acda3862 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:30.937385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:30.937385Z digest=sha256:f68522d48e9d03b87a0d788204976197f32ca11475def8b7654c46956ce730f1

Observation 47143c76-d98f-4db0-9cab-253bf586ec50 · outbound

This paper cites Language Models are Few-Shot Learners.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Language Models are Few-Shot Learners

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.185438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:30.942867Z digest=sha256:67699fe498632077a809609e04be9e581ec16f3c23eefe84249428c3b7dcadfe

Observation e6d15ea8-4ba8-4b5a-8df3-524e8104bea3 · outbound

This paper cites PySceneDetect.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents PySceneDetect

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.171303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:30.947439Z digest=sha256:4e153f3c0214574757b26241df4b89777809623e864862641582dee8888b5b13

Observation b9428f4e-68e4-4473-844d-1f73626edb51 · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:30.952146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:30.952146Z digest=sha256:3f2bf0a9a2811340d5d40873a7c67acca4af6948458b86620de09ea7136831bb

Observation 2fa8da2d-5128-463f-a342-6cc69520a4e2 · outbound

This paper cites GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:30.956868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:30.956868Z digest=sha256:0f4ae22a2022f4697ae67d5c5cf1fc3d3f5ddbf635933c201a20ff8b27eacef1

Observation a9b444d0-bc96-492d-bc5c-37a29b2bc1dd · outbound

This paper cites Extracting Replayable Interactions from Videos of Mobile App Usage.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Extracting Replayable Interactions from Videos of Mobile App Usage

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:34:31.514326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:30.961626Z digest=sha256:1244aa833a764b9a3bb789ce72d56d31cf1576fcdab5860f3e4fa4ebbc9a3ac4

Observation ae9fde3e-1972-4282-9c3e-f11a3139c613 · outbound

This paper cites Towards Complete Icon Labeling in Mobile Applications.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Towards Complete Icon Labeling in Mobile Applications

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.156959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:30.966842Z digest=sha256:92ac92aa5da21a6e0f76f6368d46df83023c6e31add8bbac72f6a68762914b0f

Observation 8866b422-70f0-410d-ba2b-743d01bd7f1e · outbound

This paper cites SeeClick: Har- nessing GUI Grounding for Advanced Visual GUI Agents.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents SeeClick: Har- nessing GUI Grounding for Advanced Visual GUI Agents

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.142384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:30.971090Z digest=sha256:92fd30d09ab518bfcc2e61c7a5ad298335029e919ed5ec06f97e979c268ad04a

Observation a9bed4ef-3692-425a-8328-fe3011a294ef · outbound

This paper cites RICO: A Mobile App Dataset for Building Data-Driven Design Applications.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents RICO: A Mobile App Dataset for Building Data-Driven Design Applications

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.128396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:30.975193Z digest=sha256:f2f7c393df524210b524a82413fd1621fc81019e8e69228dbddc559da58c05df

Observation 672826e4-306e-4a09-a55e-d9fec43f8f40 · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Mind2Web: Towards a Generalist Agent for the Web

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.114624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:30.979379Z digest=sha256:76ab9d71c1422a62d38008ae1f1b266c0c105736cb0ce3b7bd69964cf5fca6ff

Observation 5910a3d5-ab4a-480c-b894-becc1e7a0f83 · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Mind2Web: Towards a Generalist Agent for the Web

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.100471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:30.983624Z digest=sha256:c8202b6b9e2065915a1053fd834468d09206ea7e73ecb86e03cddad585e77065

Observation 01cccb08-b52f-41a4-b840-47c7b2f0fbfa · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Trans- formers for Language Understanding.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents BERT: Pre-training of Deep Bidirectional Trans- formers for Language Understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.086061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:30.987641Z digest=sha256:4a663476ba2d7fe2a50ec5e8536094074b191c00bf3fe9f5a818b57ca1bc7d83

Observation 7c30bf4e-ab8e-42fb-85f0-46c881a51083 · outbound

This paper cites Video2Action: Reducing Human Interactions in Action An- notation of App Tutorial Videos.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Video2Action: Reducing Human Interactions in Action An- notation of App Tutorial Videos

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.071674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:30.991591Z digest=sha256:c7c131e056858040240b8f253d64c73165430dc90d512fb517913563ba0055c9

Observation 3ba335ad-696f-4dfb-b037-177d98e27e91 · outbound

This paper cites Fouhey, Weicheng Kuo, Alexei A.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Fouhey, Weicheng Kuo, Alexei A

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.057251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:30.995769Z digest=sha256:4d028ff170a222c86e145074d5591db3c8b7c9948d81ca2440754afc81cc2037

Observation aeb94124-ebaf-4f76-9cdf-945edf6c070e · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents LoRA: Low-Rank Adaptation of Large Language Models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.043576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:30.999994Z digest=sha256:7b82a82997eace356e90780d3df544e8d387215337636d4ca51cfc11885e613e

Observation 10aff24c-431f-4f8b-bc86-cc528b8bdd59 · outbound

This paper cites Multimodal Subtask Graph Generation from Instructional Videos.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Multimodal Subtask Graph Generation from Instructional Videos

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.028471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.004451Z digest=sha256:f0fe29c4cf93d163a2ae2e0f70c024455bfae91c069b0040bd57496aeac757ef

Observation 7aff010f-8889-4186-aa5a-3cc3746942cc · outbound

This paper cites VisualWe- bArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents VisualWe- bArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:32.011385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.008652Z digest=sha256:6d1dbf0eecf2597dc093e3f43abe8ed5b79305dc04bd8d60478a3419aeac908e

Observation 39c489b0-cf32-4e5e-a265-0a1916e2ca04 · outbound

This paper cites Tree Search for Language Model Agents.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Tree Search for Language Model Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.012778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.012778Z digest=sha256:81223e2fc015a251bdfee3ec07e5dfc610f32927f1973dffdd6a6b3606bc345c

Observation 5db60889-8941-4f09-a6fd-823f9c08d5e4 · outbound

This paper cites Benchmarking Mobile Device Control Agents across Diverse Configurations.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Benchmarking Mobile Device Control Agents across Diverse Configurations

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.996782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.017161Z digest=sha256:1466d74f4b48fae21bed2126ccb03cff099a438bca1107f0d2bca44939cc9b43

Observation e84e0ed6-0c58-47cc-ae82-0c827aeeb2f4 · outbound

This paper cites Binary Codes Capable of Correcting Deletions, Insertions, and Reversals.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Binary Codes Capable of Correcting Deletions, Insertions, and Reversals

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.983020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.021324Z digest=sha256:d0fac7e58d688d0200d4909acb687c182d9db8d8c90a0f96c012f5862141d302

Observation 06e982d0-e682-4905-8fb5-75c772d7bfd9 · outbound

This paper cites PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.025307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.025307Z digest=sha256:e9b6ae249788cee659e35f2106083fe40e8ba67c26944ebd742f5c8a38455010

Observation dd404b4d-d3c1-4269-a1a9-151de8a391da · outbound

This paper cites On the Effects of Data Scale on UI Control Agents.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents On the Effects of Data Scale on UI Control Agents

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.968880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.029797Z digest=sha256:1805411dd6ecc6d325562fe058c4cbd38683aa143505a655d9c408605bb068ca

Observation bf9628ce-406b-485c-89ab-425be8a2c339 · outbound

This paper cites Mapping Natural Language Instructions to Mo- bile UI Action Sequences.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Mapping Natural Language Instructions to Mo- bile UI Action Sequences

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.955166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.034035Z digest=sha256:e023f0449e1f24944737f23716fa0a8daa8648b112b4bb35ad51537b24c8cdcb

Observation 207cee73-1079-4796-b7ea-ebf448915b5b · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Improved Baselines with Visual Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.038069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.038069Z digest=sha256:b0d2857a310e411652f0306e1c0886139b738e07275cd76c3ab86194dac36874

Observation 18210096-9fe0-48e4-8a38-3f1115638b95 · outbound

This paper cites Visual Instruction Tuning.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Visual Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.042385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.042385Z digest=sha256:920159ee02aceb15f00a85455c3c120eb7c9d718df626ab198a6e407c803a4b8

Observation 82af6a0f-5bc9-483e-8e25-49bd0ec684a2 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.929914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.046416Z digest=sha256:2cca15e68fd280e623ae116a13aca383606a35b49ee3fc1969cd976e618f34f3

Observation 7cc8d698-64e9-4ceb-ba0d-3a97d9164222 · outbound

This paper cites Unsupervised Task Graph Generation from Instructional Video Transcripts.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Unsupervised Task Graph Generation from Instructional Video Transcripts

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.915395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.050793Z digest=sha256:13d90f7e496e8683e3b04dce8cff37fd959c9c1aec544418ea4bbc9d40a928c7

Observation 5f6cbe26-a727-4bc6-9398-37d7d39f2848 · outbound

This paper cites OmniParser for Pure Vision Based GUI Agent.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents OmniParser for Pure Vision Based GUI Agent

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.054905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.054905Z digest=sha256:fd804a216c43014ce7e7935ceb304b87c6306da964889d5350876eafec0948e0

Observation 88ca5661-2803-41fe-bb66-b97c4ebd7563 · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents MediaPipe: A Framework for Building Perception Pipelines

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.059071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.059071Z digest=sha256:e02fedc44cef0f197a339c4e64b923965646ad5b79260535bd6d6f17eddbf186

Observation 2f6b8041-f2a6-4b51-bc14-77a7dc415de6 · outbound

This paper cites PEFT: State-of-the-art Parameter-Efficient Fine-Tuning methods.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents PEFT: State-of-the-art Parameter-Efficient Fine-Tuning methods

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.899103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.063522Z digest=sha256:6a877927ebd87768b6c4925bcb025d0a3dba72dbba327184ede780b344ae8b65

Observation d3605614-6e6b-41d5-a6be-d29218027ca9 · outbound

This paper cites Llama-3.2.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Llama-3.2

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.883310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.067831Z digest=sha256:237ac128f8252b2ae6d5eaf6a5f0049b81ea0350835648564b8b7998b29dcc89

Observation 42630270-b732-4630-9012-3bdcc184b273 · outbound

This paper cites HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.076457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.076457Z digest=sha256:10ae41267191709dba52791ecb8123ad518b4475da047e76d4626d7bd860c742

Observation e26dd40b-0af0-450f-b41b-784aec4ff06e · outbound

This paper cites ScreenAgent: A Vision Language Model-driven Computer Control Agent.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents ScreenAgent: A Vision Language Model-driven Computer Control Agent

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.082873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.082873Z digest=sha256:9d61c6ecd504df68f68ffa5cf70ebb858bc870b8582b2645a845fedb13e52b98

Observation 9273a977-7ea3-457b-961f-aa1bd3ea87b0 · outbound

This paper cites GPT-4 Technical Report.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents GPT-4 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.087656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.087656Z digest=sha256:db4450ab5eb1944aed5689333badc7653b02215e73f558c412466986de32191b

Observation 07f5e651-8ee7-415c-be62-262f3f8aa7bf · outbound

This paper cites GPT-4V Limitations.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents GPT-4V Limitations

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.846680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.091744Z digest=sha256:47b39acc8d4ed4f2d081100fb795483663c9cd919b9cacf48391193967efcd04

Observation 96ea536a-4a34-4fdb-953d-5182ba0068f0 · outbound

This paper cites GPT-3.5 Instruct.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents GPT-3.5 Instruct

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.833473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.096167Z digest=sha256:146bff84344cc6ca95c4ed5b57df108a460b91e32a39a26eeb997cded063a1dc

Observation 615eefb9-52ec-4e6f-97ab-bba73ab7eebc · outbound

This paper cites an unresolved cited work.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:31.820699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.100317Z digest=sha256:2f037d1176bc2b6065cf4bf2546da7a08110ad9e088f65fa0368d807c874c47f

Observation dbf0c5f1-7a6a-4182-bdf2-0593f285d03c · outbound

This paper cites an unresolved cited work.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:31.806906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.104455Z digest=sha256:f6fe240f19f8e6cc711ad828e29beb716f373d1ecdf3833840c78f2b05f15d2a

Observation 0569f2b7-7bbb-4cf5-ba57-62f401c5024f · outbound

This paper cites Language Models are Unsuper- vised Multitask Learners.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Language Models are Unsuper- vised Multitask Learners

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.793341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.108586Z digest=sha256:2ecf9bfc4e75d2bee08558ad13f55c99356bd143b98a9689f69ce9b8490561b8

Observation a66f953e-544d-4432-94a5-78c042c5e7a6 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.779955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.112716Z digest=sha256:73dc52fde03ee3986e3425b6c6732727a84c3750ced2918f8480d4a86eb82376

Observation ac682efd-848a-4af5-81e8-4c13b30d512f · outbound

This paper cites AndroidInTheWild: A Large- Scale Dataset For Android Device Control.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents AndroidInTheWild: A Large- Scale Dataset For Android Device Control

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.766896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.116652Z digest=sha256:b7a4ccf941a864411c8391a0f9347d716f9d3efdcb02abb509403e81d2bbf938

Observation 18819b02-5f54-40e1-9913-f2d0d622fc0f · outbound

This paper cites As- sembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents As- sembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.753766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.122462Z digest=sha256:8fad576ed41b72ce389b570a0e5393680d40633d4f11f44f721e573e6f1dd4fd

Observation 71fe8954-bdc3-4ae1-9283-060fbc7294c7 · outbound

This paper cites LUSE: Using LLMs for Unsupervised Step Extraction in Instructional Videos.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents LUSE: Using LLMs for Unsupervised Step Extraction in Instructional Videos

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.740788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.126655Z digest=sha256:b8927ae3b94fb843884300263a657f6ec1daf6215a1824af0cf724055d700b55

Observation f9b597d0-ae71-451b-9c55-7ae76bac7397 · outbound

This paper cites World of Bits: An Open-Domain Plat- form for Web-Based Agents.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents World of Bits: An Open-Domain Plat- form for Web-Based Agents

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.727913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.131003Z digest=sha256:5c59e12aa52f2cf89f61a5fb8019badac3ec611452ab634a71bd954590ad4643

Observation 32122ada-a141-49be-b20a-076554ae151d · outbound

This paper cites AppBuddy: Learning to Accomplish Tasks in Mobile Apps via Rein- forcement Learning.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents AppBuddy: Learning to Accomplish Tasks in Mobile Apps via Rein- forcement Learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.714232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.135315Z digest=sha256:b6b480e3c98c4f3e57949b92f4eb686bbd92bb01083393e9a35c472dacec1a88

Observation c4764dc3-5f1f-453d-9de3-3e157451416d · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.139287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.139287Z digest=sha256:66c27e9cffb5569cb9f5afde299c9dfb779fd67f020185ab48462053d6299355

Observation 8d45d109-864a-4199-a216-768b1bacf1a4 · outbound

This paper cites Towards Better Semantic Understanding of Mobile Interfaces.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Towards Better Semantic Understanding of Mobile Interfaces

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.699960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.143625Z digest=sha256:fd3b9dc16823c5703197ea9461a054d647cda83ae49a7437c5dfa3749536882e

Observation e6ccb45e-c073-49c1-9d3b-6e4246c82502 · outbound

This paper cites COIN: A Large-scale Dataset for Comprehensive Instruc- tional Video Analysis.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents COIN: A Large-scale Dataset for Comprehensive Instruc- tional Video Analysis

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.685989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.147600Z digest=sha256:fcbadf5405831e349d5a4b78b04572a8f1e2e30692466cb20aef2d73586b4b89

Observation 2555cedd-9b09-4d7f-9f91-768b035d4ce0 · outbound

This paper cites AndroidEnv: A Reinforcement Learning Platform for Android.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents AndroidEnv: A Reinforcement Learning Platform for Android

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.151703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.151703Z digest=sha256:a3a93fadf85b8250c533ce6f4c04a398159b57526b900596a543cf379b6128d0

Observation c77d4217-ce78-429d-9daa-057ba2aac452 · outbound

This paper cites AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.156733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.156733Z digest=sha256:b98edbaf96ed8a09e538e97eeb435c22519f2050e06e698313b8153cce383e52

Observation 9b74395a-5874-44a0-842e-a066551df80d · outbound

This paper cites GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.160983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.160983Z digest=sha256:e492fec79f02066a97e370063987a34d4415078b193d76dc61b62a84b249c67b

Observation b7cce51f-59b0-4af5-825e-03d191d0c125 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.165416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.165416Z digest=sha256:ae729a97da37629d9f00e20cf87e4b19046d01bc869ad0a98060be667c693954

Observation b01155ea-fad4-4881-9082-c2cf165c4a2e · outbound

This paper cites Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.170030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.170030Z digest=sha256:8089db1bd98a887f1a1c025bf3e77e97587a68b0275c2da03710411aa4e429d0

Observation 9882fc8e-74a2-4364-bffa-18a23a6b238c · outbound

This paper cites You Only Look at Screens: Multimodal Chain-of-Action Agents.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents You Only Look at Screens: Multimodal Chain-of-Action Agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:31.174319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:31.174319Z digest=sha256:8e51b2635c4f46feda7fd344345e3b8f9d752aed0b5720e57d088d4525f8060c

Observation 5cc51e23-ea00-4d34-815c-2aa9687f0b42 · outbound

This paper cites GPT-4V(ision) is a Generalist Web Agent, if Grounded.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents GPT-4V(ision) is a Generalist Web Agent, if Grounded

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.672294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.178651Z digest=sha256:ec04f4f9fc0eb96a36babc00af72081fcad0caf93dac6d73a0facf3066ca2663

Observation db40d8fb-1f6e-46be-967a-2dda7932d43b · outbound

This paper cites WebArena: A Realistic Web Envi- ronment for Building Autonomous Agents.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents WebArena: A Realistic Web Envi- ronment for Building Autonomous Agents

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.658387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.182731Z digest=sha256:954673c4fd25e8bc84da36aad2296dfd358a5c845fe58baff96e2ffb86bd6b6f

Observation f08218b7-0280-48c7-9145-788ae52841bf · outbound

This paper cites Scalable Video-to-Dataset Generation for Cross- Platform Mobile Agents.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Scalable Video-to-Dataset Generation for Cross- Platform Mobile Agents

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.643559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.186583Z digest=sha256:9210d312717d4bb3e653647c485b71d286342552602610653c4609ce14d90e73

Observation f57f30c8-efaf-4dce-b855-70268d6ef664 · outbound

This paper cites phone screen.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents phone screen

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.628589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.191310Z digest=sha256:e22694cc90a59abfb799d364c5b7630d5c973d959e990287e260ca4420795dfc

Observation 4c36ea95-6a4f-4712-8ac3-cf648a27b0c1 · outbound

This paper cites an unresolved cited work.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:31.614162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.196435Z digest=sha256:9a4a0baa02dbab9d1daae67cbf634c55f0265eb218049671c44c6aa807297a31

Observation 24e00717-9b67-436e-83a7-d006cb61fa7e · outbound

This paper cites an unresolved cited work.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:31.599890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.200583Z digest=sha256:0d14117072fd2d9818248f800056f922e8c7045118788c344cbb8e2b8ec04e04

Observation 0e8063a3-24fc-48bf-a781-a58a04d61781 · outbound

This paper cites File: $image.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents File: $image

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.585897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.204662Z digest=sha256:4a584dc6be65931b102636cb46dce9684f8901e4b3351e732fadf02f3cf2cb4d

Observation baa988c6-0164-4f17-9f9c-c36ec82ccb14 · outbound

This paper cites How to Delete A Direct Message on Twitter.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents How to Delete A Direct Message on Twitter

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:34:31.571059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.209389Z digest=sha256:d88a9d1e47ac86422542519a2ae91d8d0c776d9f217862f501acb8e8428d1344

Observation e0e6916d-732a-45b2-ad0d-5e07f95d69c3 · outbound

This paper cites an unresolved cited work.

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:34:31.869693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:34:31.072148Z digest=sha256:bb50d81f304bb89349349f7d3103145d5636dce9e616c50c7d4eb73436d29669

Pith citing papers

Observation bf4e1a33-b013-469c-a735-1001a0f7a02e · inbound

MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding cites this paper.

MUIAnno: An Expert-Annotated Dataset and Evaluation Benchmark for Mobile UI Understanding Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:12:50.492819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T22:08:27.511727Z digest=sha256:87161225bee401928d437de5fc4fa62c17268ff6f5ff5980c909efabc3cc8b86