Pith. sign in

Paper Citation Record · LEDGER

CogAgent: A Visual Language Model for GUI Agents

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 46 inbound Pith citation observations for arXiv:2312.08914.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.08914 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 46 of 46 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T06:01:31.626071Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3214e3ab-a280-41cf-8dbd-7e197a925c31 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models CogAgent: A Visual Language Model for GUI Agents

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.446292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:4c25e972a22815004535dbfc3ae2c536ea0f79a52e0aee755e957eb53da990a9

Observation e6675e01-5d8a-410e-83a4-978dd5e5611e · inbound

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents cites this paper.

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents CogAgent: A Visual Language Model for GUI Agents

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:09:46.540356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T10:09:46.447508Z digest=sha256:e6c9a0107b8ee22416b3d6d20d51c44d316bdbf09617e33937a9e4f027cc1849

Observation 639bd103-164a-48a4-bd0d-82cbf0daa853 · inbound

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments cites this paper.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments CogAgent: A Visual Language Model for GUI Agents

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.465423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:932af175f846dfdb717929b0dd29e3276d86e4f160a3d54fcc71aed3e6c46ca0

Observation 61661f24-5128-4028-9f4a-cab87d879e8f · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites CogAgent: A Visual Language Model for GUI Agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:58.990645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:74442d3e6c88f98ad09843e7119bdedf0e6b6a187833df5a93c579141a95e875

Observation 033a0815-e370-4512-8311-f765331f23a1 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning CogAgent: A Visual Language Model for GUI Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:57.935228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:a004d9f9bc17c5ac42ebe09b5ccdff6bbb06a89f69a4061fe9575abe5b629a57

Observation daff75fb-8129-47db-9e38-ec8905cb7402 · inbound

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents cites this paper.

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents CogAgent: A Visual Language Model for GUI Agents

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:06:13.777835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T12:06:13.697487Z digest=sha256:ab6738b1be798f581db6ffccb274dcf9bd32d03c8be230198d4ebfb492da6d86

Observation d20010ca-3451-405e-8a1e-55c27a43e0da · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark CogAgent: A Visual Language Model for GUI Agents

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.172445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:e74e5cd5a3ee9dc5a54c0d2f23ac540b9bf24f87159b312e042ca469254d7af4

Observation 18d3795f-aba7-4893-9d39-29fece85d28a · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output CogAgent: A Visual Language Model for GUI Agents

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.769978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:72f14cbe128cefa50c937e1965a7e432e4f59c85ba5b065c9b5f214868890303

Observation e4f5fead-4c87-4d63-bbed-2c9ec1fb625f · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey CogAgent: A Visual Language Model for GUI Agents

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:27.968638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:998531a526173838fd1c2b25be438e2b2bcd539bef811dd0e7527c743e4a96bb

Observation 485324ff-eeda-4a5d-8817-79a990391ed7 · inbound

TRISHUL: Towards Region Identification and Screen Hierarchy Understanding for Large VLM based GUI Agents cites this paper.

TRISHUL: Towards Region Identification and Screen Hierarchy Understanding for Large VLM based GUI Agents CogAgent: A Visual Language Model for GUI Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T06:01:31.626071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T06:01:31.626071Z digest=sha256:1d7117d3dfcf42debb2e721c61ce487398cd0503a02436e1ca9076fe000d2935

Observation ae727bc1-7c64-44bf-b741-9d442477c5da · inbound

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents cites this paper.

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents CogAgent: A Visual Language Model for GUI Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T20:57:48.425527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:57:48.425527Z digest=sha256:5ba24586a7dabfbe2e7b3772eb123cc5e11231562cf8289d45718d9551cc4cd6

Observation 53784e52-3283-4c3c-ae0f-0bb1fb454dde · inbound

Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation cites this paper.

Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation CogAgent: A Visual Language Model for GUI Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:04.556039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:18:04.556039Z digest=sha256:c4e08a79dcf5cd9f340c6c25b6d4f69706f6a4a4afc57f5e0ff26d861361abe8

Observation ffb37659-bcbc-4cef-8a15-a406b0667597 · inbound

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models cites this paper.

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models CogAgent: A Visual Language Model for GUI Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:43.994161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:43.994161Z digest=sha256:c8f672055afa9cd5640e4ebf07cc6475558fcbaeda1d808a633279b9069f7555

Observation c4fb6c79-0874-4ebd-817a-ace03b950ed7 · inbound

Exploring the Potential of Metacognitive Support Agents for Human-AI Co-Creation cites this paper.

Exploring the Potential of Metacognitive Support Agents for Human-AI Co-Creation CogAgent: A Visual Language Model for GUI Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:13.264774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:13.264774Z digest=sha256:775abf26e3dcbb5bb6e4c5feb55198212fcf9c124213eed5c30cc0b4aa9a1c83

Observation cc4f8ab6-75d4-4a7d-bf8b-aa198f685dbe · inbound

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding cites this paper.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding CogAgent: A Visual Language Model for GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.212777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.212777Z digest=sha256:2a3913aa1aae98bb85ea9d922e290a82535f15a7e4d702154a01da083a447229

Observation 500ea2d7-0380-4833-8837-d97ded882cd3 · inbound

Mobile GUI Agents under Real-world Threats: Are We There Yet? cites this paper.

Mobile GUI Agents under Real-world Threats: Are We There Yet? CogAgent: A Visual Language Model for GUI Agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:02:07.944025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T06:57:25.257775Z digest=sha256:8955c3c83b5371cf7e1cd86f31633f64cdf3906e0ccb858aa1812a3f354b3408

Observation d0f6657d-6789-4ca2-a6f7-4dcdb6d7b6c8 · inbound

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding cites this paper.

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding CogAgent: A Visual Language Model for GUI Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:11.323765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:29:11.323765Z digest=sha256:4683ebe7824aec3d65af6f03a031cc9468390caa501e8e9488f671c9cb961ec3

Observation c8769967-487a-4cb0-973a-086bfb325e7f · inbound

Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation cites this paper.

Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation CogAgent: A Visual Language Model for GUI Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:08:35.405743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:08:35.405743Z digest=sha256:ab6d1839ea55ec533904014667c1ed9a14c338a42aa78ce12bf539668d280b59

Observation b9019cd6-d586-44bf-9fd8-a4568a2e5b9c · inbound

Uncertainty-Aware GUI Agent: Adaptive Perception through Component Recommendation and Human-in-the-Loop Refinement cites this paper.

Uncertainty-Aware GUI Agent: Adaptive Perception through Component Recommendation and Human-in-the-Loop Refinement CogAgent: A Visual Language Model for GUI Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T01:01:07.165776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:01:07.165776Z digest=sha256:78a5ad0dde0eb7ce02ba55d17161f958f1fbe942bf7d71fec408a3d144cd903a

Observation 174dd392-0872-481e-9294-fc014b238dca · inbound

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience cites this paper.

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience CogAgent: A Visual Language Model for GUI Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:55:48.651764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:55:48.651764Z digest=sha256:dc12627f81688ee8921b4cf5af3b5be56a161c764cc16b707b03bbcb53210832

Observation 0aeabdef-01b3-41ce-91ab-3bbc1577c9e6 · inbound

Cybernaut: Towards Reliable Web Automation cites this paper.

Cybernaut: Towards Reliable Web Automation CogAgent: A Visual Language Model for GUI Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:20.465838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:20.465838Z digest=sha256:6270a79cf53e96b5f8cb0ba05f95a0426423acadd88805aca8a8e793bd3fe3a9

Observation 0f880a2d-6315-4c62-ae82-8b20b54c5721 · inbound

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning cites this paper.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning CogAgent: A Visual Language Model for GUI Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.628996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.628996Z digest=sha256:b32d2f677c9506ca51200a770ef8b7ceb7c26701fd8fed8c7963330c9300ff0c

Observation 56f30723-c351-4607-9e16-bbe77c31f994 · inbound

MobiAgent: A Systematic Framework for Customizable Mobile Agents cites this paper.

MobiAgent: A Systematic Framework for Customizable Mobile Agents CogAgent: A Visual Language Model for GUI Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T13:34:59.368315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:34:59.368315Z digest=sha256:80a72c7cf24cdeb7f71bab8520892e923290be509bf32339b71697798b1ca96f

Observation 54e9756b-4558-43e2-8a15-858c219cc315 · inbound

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action cites this paper.

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action CogAgent: A Visual Language Model for GUI Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:03:38.142743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:03:38.142743Z digest=sha256:9e1af35440e47e27452d249fc4de8372ddebbe624a4d877f092f94dc284c2aa2

Observation 3be052ac-8eed-4fad-9a52-bf977b2ebc20 · inbound

Grounding Computer Use Agents on Human Demonstrations cites this paper.

Grounding Computer Use Agents on Human Demonstrations CogAgent: A Visual Language Model for GUI Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T23:06:04.882108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:06:04.882108Z digest=sha256:fa968a2e99d2cc166b4ebae6f2345c864ab6130dc6f26a2a2459dfbb788d9adc

Observation 0ed4a0e8-0cda-4841-95b3-4d180ee4915f · inbound

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL cites this paper.

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL CogAgent: A Visual Language Model for GUI Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T20:51:41.781289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T20:51:41.781289Z digest=sha256:c054ea145093c6392704012332fe31f8b7c68f9c8039fc8e8fd6cb63a4686ceb

Observation db48a1f3-5977-4056-9f33-511c1ac79750 · inbound

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting cites this paper.

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting CogAgent: A Visual Language Model for GUI Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T15:29:31.834566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:29:31.834566Z digest=sha256:718aba5a3e49204b813b685c6ca6bdf4c1b4567171f8ba1e7646e42c9112df92

Observation dd3a4c5c-dccb-4200-8e20-0c5f3b88b3e8 · inbound

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web cites this paper.

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web CogAgent: A Visual Language Model for GUI Agents

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:40:58.721350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:00:34.401698Z digest=sha256:45819ba7d4daeaa881170e988635dcb9a22ea5eee6c27f2be1cdaade10cc3326

Observation b3b6f78c-33cd-47bf-890b-39f1bd64ead3 · inbound

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding cites this paper.

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding CogAgent: A Visual Language Model for GUI Agents

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:10:28.874668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:06:55.472857Z digest=sha256:3fec86f438da0524b14dd97a55ec3885b4b1c0f0c80b5094b16f7b94935d6811

Observation 8a73181f-a9c7-45ae-99e5-6d60903a08d1 · inbound

OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory cites this paper.

OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory CogAgent: A Visual Language Model for GUI Agents

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:25.492638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-07T10:45:48.976501Z digest=sha256:5462feba60faff243eab71c12f450a8dd0a68fb9234eb300a4bd2bd373d4d093

Observation 85007b8e-8431-49fe-82af-c599ef7c4e98 · inbound

MMSkills: Towards Multimodal Skills for General Visual Agents cites this paper.

MMSkills: Towards Multimodal Skills for General Visual Agents CogAgent: A Visual Language Model for GUI Agents

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:07:51.413266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:05:36.511150Z digest=sha256:93d33b0645d410e07c6662d41cdd65cda3759226eba58c083bcf81ac81f3580c

Observation 27da449a-793a-4628-95e2-045219275cd9 · inbound

MMSkills: Towards Multimodal Skills for General Visual Agents cites this paper.

MMSkills: Towards Multimodal Skills for General Visual Agents CogAgent: A Visual Language Model for GUI Agents

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:59:48.175276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:59:44.669877Z digest=sha256:b9d35dc49b5a192a591f14c086d8e738527e94569a912caebe6cb3fa01702980

Observation 94636b68-cf0c-4e6c-ae31-eed6da1a0af7 · inbound

MMSkills: Towards Multimodal Skills for General Visual Agents cites this paper.

MMSkills: Towards Multimodal Skills for General Visual Agents CogAgent: A Visual Language Model for GUI Agents

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.557252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:400fe1336f3aaf5139567d7a6ea7cb242be0c10d3bc0ed5cba121e7d88d6aff6

Observation 7c800237-0192-4771-8bc4-b9d0d66db98d · inbound

ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis cites this paper.

ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis CogAgent: A Visual Language Model for GUI Agents

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:04:37.662147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T10:57:59.013811Z digest=sha256:ff63b592a940a13287aa188be7fd57b0b968153ffcde57e6addd75b8298a08cb

Observation d44059e2-596b-4f0b-b5e0-6fb567de06a8 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap CogAgent: A Visual Language Model for GUI Agents

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:02.060945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:900c524cc3908467dab0d00269f3148aa38501c88bc64acb8865f72ff4cfd7b9

Observation 546643e6-e15f-424f-94c6-5f61d0749a71 · inbound

Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark cites this paper.

Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark CogAgent: A Visual Language Model for GUI Agents

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:43:13.153036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:38:35.829605Z digest=sha256:b5f6922392a1a396c7857e38541b3e10fd5edec1c77e7d6472aaba47e9fa1343

Observation 65db81a9-6606-4ac4-a108-bae0d9e901bf · inbound

GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning cites this paper.

GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning CogAgent: A Visual Language Model for GUI Agents

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:44.696258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T22:52:04.755523Z digest=sha256:d04b3a9883db7500566724809d6ec5394c488feb47f4aba95f9dec4ab23d191b

Observation 749f98ca-d7b8-49b0-9933-53d2beb899bc · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation CogAgent: A Visual Language Model for GUI Agents

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:44.013381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:7228d156f49be206eb3fc9f7a2cd2b24dfc6d258f3bdf14ef422f1400e2c0012

Observation 84ef5e59-b519-46dd-bf4d-9381a206014c · inbound

ProCUA-SFT Technical Report cites this paper.

ProCUA-SFT Technical Report CogAgent: A Visual Language Model for GUI Agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.718222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T03:18:04.149281Z digest=sha256:b26a9811252561decbd06f2e7a73280bc068b9e51e06dec103079828272f0680

Observation bd8d87a1-6bcc-4cfd-84ef-01068cbd907e · inbound

MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents cites this paper.

MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents CogAgent: A Visual Language Model for GUI Agents

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:08:58.345016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T00:54:00.944706Z digest=sha256:c91922c4a2134787387abd4a52fe701994945103ce4a2829387d7961cbeec828

Observation fe492bcf-aa7d-4260-9aad-21ab9317c5e1 · inbound

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction cites this paper.

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction CogAgent: A Visual Language Model for GUI Agents

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:22.413201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T07:48:01.719339Z digest=sha256:8f9298dc919e8ebddd0168e2ee5e829fd791281c5701b264a7c3b18ceb6172d1

Observation 959a6d74-e4cc-420b-8379-73eb220a464b · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks CogAgent: A Visual Language Model for GUI Agents

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:14:21.193108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T07:10:38.909339Z digest=sha256:9b905c29974ff10203f2637ead687290e6c9e7f6e360f1a74605080390036622

Observation a12d611d-2687-48b6-b2bc-a2aa74e42147 · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks CogAgent: A Visual Language Model for GUI Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-15T10:24:53.345620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:24:53.345620Z digest=sha256:12b041e3f7d6e7b1de01ed15330176379a07c303efc6734b90c1d642ac04c754

Observation ab5b6383-dd99-419b-b36e-d40de6cb68c1 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report CogAgent: A Visual Language Model for GUI Agents

Reference 110

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:10147ce9d854909f7f7ebdde4080467e781e7fac477b05da13c20912870badf6

Observation 3e873257-3ded-4458-b51d-c2dc656c60c2 · inbound

Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models cites this paper.

Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models CogAgent: A Visual Language Model for GUI Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T10:42:21.217308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T10:42:21.217308Z digest=sha256:e41a2e75b87ea4171a60f81ea5a996f0b6d249a5a27cbf4440bd7b7ae08fa7ec

Observation 19c92063-4d6b-49e1-96ae-3daa64838819 · inbound

AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents cites this paper.

AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents CogAgent: A Visual Language Model for GUI Agents

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T21:40:39.900263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:40:39.900263Z digest=sha256:7b76c845b85399db4451ee4572a48a8d1a19ba7bcc7420d2748d5aaf9a9ea119