Pith. sign in

Paper Citation Record · LEDGER

MMSkills: Towards Multimodal Skills for General Visual Agents

As of 6 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 2 inbound Pith citation observations for arXiv:2605.13527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.13527 v3

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:05:27.821099Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T21:28:58.598336Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact13
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch24

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 27f60dd1-2323-4e7e-8558-d4e2245774ab · outbound

This paper cites Agent S: An Open Agentic Framework that Uses Computers Like a Human.

MMSkills: Towards Multimodal Skills for General Visual Agents Agent S: An Open Agentic Framework that Uses Computers Like a Human

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.539803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:366287e771213e080e2fb2c1e2d7b68d24c7acc284630c80a9c22d1654710a97

Observation 3404e158-0908-4681-bd78-52d75c500e8d · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

MMSkills: Towards Multimodal Skills for General Visual Agents Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.553070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:6c0314d42133a9a4306da2092e0b9d197ffdf59a0d8f06aebc9294b30ad116d2

Observation 7c5ea0d4-d852-4801-beab-83f5d2c2ce0d · outbound

This paper cites EvoSkill: Automated Skill Discovery for Multi-Agent Systems.

MMSkills: Towards Multimodal Skills for General Visual Agents EvoSkill: Automated Skill Discovery for Multi-Agent Systems

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.564993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:4f201572a290c5f771c7c28fd77376362713a16b96f22ec5ceadcabc10cfaa42

Observation bcc50edf-73f8-4ac3-9788-e5fd9fc07dc4 · outbound

This paper cites Qwen3-VL Technical Report.

MMSkills: Towards Multimodal Skills for General Visual Agents Qwen3-VL Technical Report

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.531773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:431582110dff8f959b5320aedae9a4834753516aad5e255f3a153f83edf3312d

Observation c9ba6989-a8b5-4648-9d8e-c9d592447fdb · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

MMSkills: Towards Multimodal Skills for General Visual Agents LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.567653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:a5fdf6714d710fa01a8383caf6db00a3b7b832f922587b379a6ffadc02d9f9b6

Observation 3392dde1-a5a7-4dff-9ad1-10d98d94482b · outbound

This paper cites Cua-skill: Develop skills for computer using agent.

MMSkills: Towards Multimodal Skills for General Visual Agents Cua-skill: Develop skills for computer using agent

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.570747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:ce9dc990f2dcdd5966fc9c475934f78b56baafa587d7f0ba2ae18d4869b0df12

Observation 1ac47709-125b-4b0b-a94a-6c1fe1b96c22 · outbound

This paper cites SeeClick: Harnessing GUI grounding for advanced visual GUI agents.

MMSkills: Towards Multimodal Skills for General Visual Agents SeeClick: Harnessing GUI grounding for advanced visual GUI agents

Reference 7

Resolution
metadata mismatch
doi, observed 2026-06-30T21:35:04.028125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:84d8db7506a7f2b67d30d811294e43109ac1efde5ae076a4d01b95ec6542b79a

Observation f9c180a3-5e0f-49ba-899a-d1423645518c · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

MMSkills: Towards Multimodal Skills for General Visual Agents Mind2Web: Towards a Generalist Agent for the Web

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.573142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:85229c0d615b57360050dc941c60439733354a92c07ca607dde55101ecf22797

Observation a5d0bb76-3bb8-4243-98c6-88069d3677a3 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.554540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:6ce19d36f9eaa5b9c2c3a7d3c7af193bbfb002a403bdd5df0a9bb4819a4d1b2e

Observation 2d08272c-fa46-4bac-983d-5a6ae8850db4 · outbound

This paper cites Webvoyager: Building an end-to-end web agent with large multimodal models.

MMSkills: Towards Multimodal Skills for General Visual Agents Webvoyager: Building an end-to-end web agent with large multimodal models

Reference 10

Resolution
metadata mismatch
doi, observed 2026-06-30T21:35:04.018795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:464f72d1dd456e30c45c6d43ad37299d03a9d8544d3845a9b1dfd6ec8d1f9d40

Observation 94636b68-cf0c-4e6c-ae31-eed6da1a0af7 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents CogAgent: A Visual Language Model for GUI Agents

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.557252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:8157404163739cee8b67e3c2e74764b4291ecf522b1af794c73f9ce63c3501f9

Observation 1d368f40-3266-41ea-8cdd-8b9fb6fa95df · outbound

This paper cites lmgame-Bench: How Good are LLMs at Playing Games?.

MMSkills: Towards Multimodal Skills for General Visual Agents lmgame-Bench: How Good are LLMs at Playing Games?

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.559898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:6f82889d5bd3a7ee04afa2e9941fa36eac996458b6c97f2d25804cd21b2a75c2

Observation 0287352e-4c26-41e5-ad64-53153b595a65 · outbound

This paper cites URL https: //doi.org/10.1162/NECO_a_00393.

MMSkills: Towards Multimodal Skills for General Visual Agents URL https: //doi.org/10.1162/NECO_a_00393

Reference 13

Resolution
verified exact
doi, observed 2026-06-30T21:35:04.020654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:bfc91439e5078a3da9d4ce8c5a2ab5afcf39db1b979c1db8b43bed1100ad1515

Observation 979ab9c9-eef2-4982-920c-ee23418528e0 · outbound

This paper cites XSkill: Continual Learning from Experience and Skills in Multimodal Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents XSkill: Continual Learning from Experience and Skills in Multimodal Agents

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:17:26.572211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:e76088bd4e49eccfa82504b26f1d928046fc8cf30075c4d64985b5da425a5b50

Observation 4087e90b-091f-4724-93f4-c624695fe047 · outbound

This paper cites VisualWebArena: Evaluating multimodal agents on realistic visual web tasks.

MMSkills: Towards Multimodal Skills for General Visual Agents VisualWebArena: Evaluating multimodal agents on realistic visual web tasks

Reference 15

Resolution
metadata mismatch
doi, observed 2026-06-30T21:35:04.034096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:9fd28a47de9167e10bd2c9695501c069cd93eaa1fac61e4dc55ee017498e1911

Observation 01b9da5e-81cb-4be0-abad-2355726a4b21 · outbound

This paper cites ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions.

MMSkills: Towards Multimodal Skills for General Visual Agents ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.025940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:249be6eb5152592873bfd66f68786501768a0bf674e295f79749131f0dc59347

Observation 542aed55-042d-4ba5-a172-7ad1f4a18820 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

MMSkills: Towards Multimodal Skills for General Visual Agents Lost in the Middle: How Language Models Use Long Contexts

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.023262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:fb3cdafbb1e7b2cb80a30a169c5bcbd11ba033319873fdbc08785d4709513b73

Observation 5754cc06-b451-4cb4-a214-f528d2cfef92 · outbound

This paper cites How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings.

MMSkills: Towards Multimodal Skills for General Visual Agents How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.543425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:d9b251278f89040e1927e588cf876016b2d38bf24388a3ce21cd3fae18185959

Observation 1edb1189-7738-40a4-9a4b-387635fc2c22 · outbound

This paper cites OmniParser for Pure Vision Based GUI Agent.

MMSkills: Towards Multimodal Skills for General Visual Agents OmniParser for Pure Vision Based GUI Agent

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.549327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:682883dc48bbae753c666a605fcc7660dac6b35057d7cf2c47bce8dbacdd706e

Observation 20ce9619-cb9a-4aa5-a70d-2ab5746ca70d · outbound

This paper cites SkillClaw: Let Skills Evolve Collectively with Agentic Evolver.

MMSkills: Towards Multimodal Skills for General Visual Agents SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.580530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:bea989f42f60b79affc63e5bc6460875ce3b6b868dfc66bcafb1a4e13c7d5330

Observation 542e943b-0bd4-46d0-a554-da29ca0c95b1 · outbound

This paper cites URL https://doi.org/ 10.1017/CBO9780511811678.

MMSkills: Towards Multimodal Skills for General Visual Agents URL https://doi.org/ 10.1017/CBO9780511811678

Reference 21

Resolution
verified exact
doi, observed 2026-06-30T21:35:04.030066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:cc35f68e785e0dcccbc0ac01dccf02c06987f00d4bd8a53a2598d77c820bb8c3

Observation 17e48366-d340-4fb1-a2c8-84c073bb9bbd · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

MMSkills: Towards Multimodal Skills for General Visual Agents MemGPT: Towards LLMs as Operating Systems

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.583018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:95b33e879228671ae070861961dff98e543e3906e3f5b6a18af745969f874c54

Observation 788c8ecb-23b7-4bee-bc96-179de575995e · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior.

MMSkills: Towards Multimodal Skills for General Visual Agents Generative Agents: Interactive Simulacra of Human Behavior

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.562519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:17ded0dbc1fa2f22cda0c36b532442fd9d07f79b1ecfe16bdfb89982a4c9ff3f

Observation 8fb550e3-a08d-4eb1-aa7c-44136aaa74aa · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.540664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:e48f52bcd7ca736588ff235ff12e335917a4d647f926cdd271ed57716434c1e2

Observation 9405978a-5fbe-4e3a-90a0-cfe5cee7c255 · outbound

This paper cites Android in the Wild: A Large-Scale Dataset for Android Device Control.

MMSkills: Towards Multimodal Skills for General Visual Agents Android in the Wild: A Large-Scale Dataset for Android Device Control

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.546582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:5f5c0f2616ed726e3405956c93551391c207a4b476a99e32db0172fa611dc116

Observation b8502d64-b27a-4b45-b0e2-b68c71a1a255 · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.578198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:bc713a2ae3c7f63f2005d86ec18b451a996562b68d4d5b2da5c25b1a7834d928

Observation a372cebe-cfda-4c19-a95a-4794b8b719e8 · outbound

This paper cites nips.cc/paper_files/paper/2023/hash/1b44b878bb782e6954cd888628510e90-Abstract-Conference.html.

MMSkills: Towards Multimodal Skills for General Visual Agents nips.cc/paper_files/paper/2023/hash/1b44b878bb782e6954cd888628510e90-Abstract-Conference.html

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T16:53:57.384410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:bace1ec5e14fb4e2bae8be53fee51d7069bc2a7d3d4134144a3f1fdd0ca30b87

Observation 47b3dbea-0a21-4aac-a219-8fdd899775fb · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

MMSkills: Towards Multimodal Skills for General Visual Agents Kimi K2.5: Visual Agentic Intelligence

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.032447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:c57acd95ae2c4d4c2966c0f50cbb6dbef4b5804f3df4579082cfbfe14705a9b2

Observation d6c43618-8764-4b86-98a2-895dde225359 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.531961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:fb2e3ae25ae90ddcc81e75d8a4346898067aff7e6357a8dba63d3cafb18547fb

Observation b45f1b66-cccf-415e-8da9-d36a804694f2 · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

MMSkills: Towards Multimodal Skills for General Visual Agents SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.575834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:3d39850618b3170ed604f71631ac61b4b403ff38e910e303a2fe566b1a815dea

Observation c2f5d302-2ccc-409b-8ec0-593dd6d874ea · outbound

This paper cites Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills.

MMSkills: Towards Multimodal Skills for General Visual Agents Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.529397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:b5fcf43400ad2dfb832065479ed0b37dc8097f3eaa0fbd620a9a98537a75a624

Observation 65319806-9555-4e7f-86a3-80378cb86217 · outbound

This paper cites Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward.

MMSkills: Towards Multimodal Skills for General Visual Agents Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.534667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:b3789175b08fc8906c57059750172ae5d8fa01a5be229292b877bef91508ea5b

Observation 6eb4e3bf-a14c-45f1-82a2-4d718435a0b5 · outbound

This paper cites DeskVision: Large Scale Desktop Region Captioning for Advanced GUI Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents DeskVision: Large Scale Desktop Region Captioning for Advanced GUI Agents

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.519678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:e1d7f135ffeb3ad97b4be8dabd92c7c964ff5909cabee64972d2c97d30536995

Observation 4832a7eb-97d6-44da-bc89-bc085fb57c55 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

MMSkills: Towards Multimodal Skills for General Visual Agents AppAgent: Multimodal Agents as Smartphone Users

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.526584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:bcb90b5902377e32a1f1ea01922efe171132dbfc27320692c7471d12465ccaa8

Observation 1c3874fc-d936-40cb-b361-4736247cd03b · outbound

This paper cites DREAM: A Dual Representation Learning Model for Multimodal Recommendation.

MMSkills: Towards Multimodal Skills for General Visual Agents DREAM: A Dual Representation Learning Model for Multimodal Recommendation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.522570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:2aabd0fa76873b1d83efbaa3de743696680eee0ad7d62cee35403afdb78dffaf

Observation 00a25186-d3ee-4694-b252-d75c54fde55e · outbound

This paper cites Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su.

MMSkills: Towards Multimodal Skills for General Visual Agents Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.537817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:86f0c9200b7fa0d141cabdda8d9cc6148a85fc3a468b7313fc2d03b35d540b27

Observation 83b02c05-c62d-40b7-aca2-02fac7f769a8 · outbound

This paper cites GPT-4V(ision) is a Generalist Web Agent, if Grounded.

MMSkills: Towards Multimodal Skills for General Visual Agents GPT-4V(ision) is a Generalist Web Agent, if Grounded

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.507748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:fe8a8de1935929aa4b5a3d2e829f1a0fab6ad058d706eb67056b0b0f9f9d0ecd

Observation 050be9a3-a594-4c99-b9c3-26a3f3550514 · outbound

This paper cites SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills.

MMSkills: Towards Multimodal Skills for General Visual Agents SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.537066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:3c8b4c63a202acac9ffc178a378c74764fe40133f8986a759e522cf33897e0ad

Observation d60c3f31-2c53-4485-8036-1ed7c73cb0cb · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 39

Resolution
malformed identifier
local_arxiv, observed 2026-06-30T21:35:04.501725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:eb48d69a86e1ddef6fa83e5074dc523d1ff82c4c9d5349a20664067d2b6a2d9d

Observation c528b9a5-e8c8-4dcc-8d27-20012df7a473 · outbound

This paper cites Recent LLM agents have made skills a practical interface for storing and composing procedural knowledge in language-conditioned environments.

MMSkills: Towards Multimodal Skills for General Visual Agents Recent LLM agents have made skills a practical interface for storing and composing procedural knowledge in language-conditioned environments

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T16:53:57.382533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:22ddeff37052b07ad5609e38448f602f9af60f13817fabbe07d69b105e1d4b98

Pith citing papers

Observation 702efe0b-ac8e-46d6-8d36-888bc44bc171 · inbound

VISUALSKILL: Multimodal Skills for Computer-Use Agents cites this paper.

VISUALSKILL: Multimodal Skills for Computer-Use Agents MMSkills: Towards Multimodal Skills for General Visual Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:28:58.599988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T00:35:02.162386Z digest=sha256:cf3e098cc8ea59cadd7642aaaa34fd2598761d6c4aaead721b4ad21ce95e2104

Observation d40bb2e2-fc90-4791-a14a-fdce8722497c · inbound

Progressive Agent Skill Generation via Reinforcement Learning cites this paper.

Progressive Agent Skill Generation via Reinforcement Learning MMSkills: Towards Multimodal Skills for General Visual Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:05:27.821099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:05:27.821099Z digest=sha256:5871bb5d59ceb7126f0083fb5dbb378cebd27a48b693240a649f3b0f2f9ac105