Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T21:25:57.458346Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2411.08768.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T21:25:57.458346Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 83ff7106-7c8c-4849-8f84-6d86cb4b9c59 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 00a0fe10-4be5-4260-854d-63449dcc48fe · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68732f94-8c18-40a7-b17e-6253cb1d3cd9 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Gui-world: A dataset for gui-oriented multimodal llm-based agents, 2024
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 00c757c2-c5fb-4f95-8913-8350b7f206f0 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Vg4d: Vision-language model goes 4d video recog- nition
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5e58c4ea-2cfa-4003-a831-f0dff2666657 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd, 2024
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 355b6808-dc24-4c77-982f-c2d9b5d89bac · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Videoagent: A memory-augmented multimodal agent for video understanding, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e3196b28-7e80-49e6-88a4-23ab1557d86c · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Video-of-thought: Step-by-step video reasoning from perception to cognition
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5a33fbcd-994a-41ab-b620-37c3d3653a5c · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Gemini models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6f931a34-e272-45a8-87dd-924404ff4428 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b0519d1d-4927-4f95-8f48-a11fbcd5ead7 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Smartflow: Robotic process automation using llms, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d7429d13-c4e0-4478-a73a-3e1fc0f3b75e · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Dream2real: Zero-shot 3d object rearrangement with vision-language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0fdc0360-92bd-4fab-8c77-6a1e5bd41df9 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Wolf: Captioning everything with a world summarization framework, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b4cd499a-0675-4d0a-a5c5-b0cc12bc989e · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Videochat: Chat-centric video understanding, 2023
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d95b2274-eb15-4d8d-bd30-8ffcf2d5af75 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Llava-next: Improved reasoning, ocr, and world knowledge, 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b7f4954-487a-4a59-9acd-d9488ae86eb9 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Visualwebbench: How far have multi- modal llms evolved in web page understanding and grounding?, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ff63183a-4361-4743-9227-3b5f7be2dbc9 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Online robot navigation and manipulation with distilled vision-language models, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 66795c9f-ab15-4121-aa2a-83e2b4443867 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Robollm: Robotic vision tasks grounded on multimodal large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9eece4f9-5a86-4fa3-8b58-2f8c5b8b2f6e · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Omniparser for pure vision based gui agent, 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc0d112-8c0f-4e39-a218-7e07b6d9ca02 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b7b29a5f-30c8-4001-97d6-7ae24254edaf · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e008d938-d9a3-4770-a1f0-49e6243e333f · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0661f689-ed6b-4684-a9d1-bdd96b945477 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Owl - always-on wearable ai
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ca310773-f698-4aa7-8aa2-01c334297d8d · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings bert-embedding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ad818ad9-c9f1-42ca-b44d-6e4f198ff4f4 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Learning transferable visual models from natural language supervision, 2021
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8caa44b6-4169-4161-acf9-66c196a9b58f · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Beltran-Hernandez, Masashi Hamaya, Atsushi Hashimoto, Shohei Tanaka, Kento Kawaharazuka, Kazutoshi Tanaka, Yoshitaka Ushiku, and Shinsuke Mori
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3c09affd-f4ca-4217-b55c-1514fc754e33 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Karlsson, Bo An, Shuicheng Yan, and Zongqing Lu
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2bc09ff3-9aa8-4972-8322-d1c51f42eaf2 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Sch ¨onberger, Juan Nunez-Iglesias, Franc ¸ois Boulogne, Joshua D
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 84a6865f-48a3-4b4d-aaf9-caa12ae79997 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Drive anywhere: Generalizable end-to-end autonomous driving with multi- modal foundation models, 2023
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 40d8d6d9-6bff-4f49-a4e7-88fca66f861b · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Videoagent: Long-form video understanding with large language model as agent, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 60abc502-0b87-4dc7-89d2-67dbcb3027f5 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Videotree: Adaptive tree- based video representation for llm reasoning on long videos, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fd03c251-ece6-4aee-8370-a81d11984e58 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Incorporating scene graphs into pre- trained vision-language models for multimodal open-vocabulary action recognition
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cb0ee32d-9b8f-4838-8764-11d6715bbb7c · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Vlfm: Vision-language frontier maps for zero- shot semantic navigation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fa58de06-7a6c-4e10-aeae-c06127417073 · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings Ufo: A ui-focused agent for windows os interaction, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a81000a5-0509-47cc-8ec0-22ac18b91bfe · outbound
Sharingan: Extract User Action Sequence from Desktop Recordings global_description
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.