Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 2 inbound Pith citation observations for arXiv:2605.13527.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:05:27.821099Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T21:28:58.598336Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 27f60dd1-2323-4e7e-8558-d4e2245774ab · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Agent S: An Open Agentic Framework that Uses Computers Like a Human
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3404e158-0908-4681-bd78-52d75c500e8d · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7c5ea0d4-d852-4801-beab-83f5d2c2ce0d · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents EvoSkill: Automated Skill Discovery for Multi-Agent Systems
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bcc50edf-73f8-4ac3-9788-e5fd9fc07dc4 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Qwen3-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c9ba6989-a8b5-4648-9d8e-c9d592447fdb · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3392dde1-a5a7-4dff-9ad1-10d98d94482b · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Cua-skill: Develop skills for computer using agent
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1ac47709-125b-4b0b-a94a-6c1fe1b96c22 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents SeeClick: Harnessing GUI grounding for advanced visual GUI agents
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f9c180a3-5e0f-49ba-899a-d1423645518c · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Mind2Web: Towards a Generalist Agent for the Web
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5d0bb76-3bb8-4243-98c6-88069d3677a3 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d08272c-fa46-4bac-983d-5a6ae8850db4 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Webvoyager: Building an end-to-end web agent with large multimodal models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 94636b68-cf0c-4e6c-ae31-eed6da1a0af7 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents CogAgent: A Visual Language Model for GUI Agents
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d368f40-3266-41ea-8cdd-8b9fb6fa95df · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents lmgame-Bench: How Good are LLMs at Playing Games?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0287352e-4c26-41e5-ad64-53153b595a65 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents URL https: //doi.org/10.1162/NECO_a_00393
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 979ab9c9-eef2-4982-920c-ee23418528e0 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents XSkill: Continual Learning from Experience and Skills in Multimodal Agents
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4087e90b-091f-4724-93f4-c624695fe047 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents VisualWebArena: Evaluating multimodal agents on realistic visual web tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01b9da5e-81cb-4be0-abad-2355726a4b21 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 542aed55-042d-4ba5-a172-7ad1f4a18820 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Lost in the Middle: How Language Models Use Long Contexts
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5754cc06-b451-4cb4-a214-f528d2cfef92 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1edb1189-7738-40a4-9a4b-387635fc2c22 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents OmniParser for Pure Vision Based GUI Agent
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 20ce9619-cb9a-4aa5-a70d-2ab5746ca70d · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 542e943b-0bd4-46d0-a554-da29ca0c95b1 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents URL https://doi.org/ 10.1017/CBO9780511811678
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 17e48366-d340-4fb1-a2c8-84c073bb9bbd · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents MemGPT: Towards LLMs as Operating Systems
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 788c8ecb-23b7-4bee-bc96-179de575995e · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Generative Agents: Interactive Simulacra of Human Behavior
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8fb550e3-a08d-4eb1-aa7c-44136aaa74aa · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9405978a-5fbe-4e3a-90a0-cfe5cee7c255 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Android in the Wild: A Large-Scale Dataset for Android Device Control
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b8502d64-b27a-4b45-b0e2-b68c71a1a255 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a372cebe-cfda-4c19-a95a-4794b8b719e8 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents nips.cc/paper_files/paper/2023/hash/1b44b878bb782e6954cd888628510e90-Abstract-Conference.html
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 47b3dbea-0a21-4aac-a219-8fdd899775fb · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Kimi K2.5: Visual Agentic Intelligence
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6c43618-8764-4b86-98a2-895dde225359 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b45f1b66-cccf-415e-8da9-d36a804694f2 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c2f5d302-2ccc-409b-8ec0-593dd6d874ea · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 65319806-9555-4e7f-86a3-80378cb86217 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6eb4e3bf-a14c-45f1-82a2-4d718435a0b5 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents DeskVision: Large Scale Desktop Region Captioning for Advanced GUI Agents
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4832a7eb-97d6-44da-bc89-bc085fb57c55 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents AppAgent: Multimodal Agents as Smartphone Users
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1c3874fc-d936-40cb-b361-4736247cd03b · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents DREAM: A Dual Representation Learning Model for Multimodal Recommendation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 00a25186-d3ee-4694-b252-d75c54fde55e · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 83b02c05-c62d-40b7-aca2-02fac7f769a8 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents GPT-4V(ision) is a Generalist Web Agent, if Grounded
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 050be9a3-a594-4c99-b9c3-26a3f3550514 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d60c3f31-2c53-4485-8036-1ed7c73cb0cb · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c528b9a5-e8c8-4dcc-8d27-20012df7a473 · outbound
MMSkills: Towards Multimodal Skills for General Visual Agents Recent LLM agents have made skills a practical interface for storing and composing procedural knowledge in language-conditioned environments
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 702efe0b-ac8e-46d6-8d36-888bc44bc171 · inbound
VISUALSKILL: Multimodal Skills for Computer-Use Agents MMSkills: Towards Multimodal Skills for General Visual Agents
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d40bb2e2-fc90-4791-a14a-fdce8722497c · inbound
Progressive Agent Skill Generation via Reinforcement Learning MMSkills: Towards Multimodal Skills for General Visual Agents
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.