Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:01:29.172482Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2507.04036.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:01:29.172482Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T19:41:58.268576Z
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8c888856-2c80-426c-9a26-0aae806c6ef8 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e7c8d37-d7df-4fcd-8a75-838a4cb2c463 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1983defb-0945-4889-b021-e1d98d195e20 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88a157cb-a991-44fe-a0be-9b921b80eafc · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Longformer: The Long-Document Transformer
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8208639-f3d4-418f-b57b-3510eb76ba11 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Structure-Aware Abstractive Conversation Summarization via Discourse and Action Graphs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0ed332b6-4134-4a7a-a66a-3c1564009960 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 224f5bd7-7d46-4e52-9b9a-53f049291259 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Emerging Properties in Unified Multimodal Pretraining
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a67b510-e71b-4042-ab2b-505d9b1d4613 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 94db5d95-d83a-4d52-920d-7c92e15c56fa · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation AutoPresent: Designing Structured Visuals from Scratch
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26aa9c68-8bfe-4f3d-8223-5b4d8b68b925 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e819c4f-dea6-4ea6-938f-68a682b9c7a7 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7432ab6a-fd75-46f2-abbb-4d672fb2d998 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 539fab81-7e54-484a-800f-50c400b1b6c9 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation LLaVA-OneVision: Easy Visual Task Transfer
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d594df54-74ca-4f9d-a7c0-8ed5ff2fc62f · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eee608d-98fc-44ba-9e0e-63a0d028d367 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation VideoGUI: A Benchmark for GUI Automation from Instructional Videos
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2fd21acb-63aa-437f-97d1-f4854ec9f7cb · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation ShowUI: One Vision-Language-Action Model for GUI Visual Agent
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5e48018-5dff-4b98-a175-551b82a549ee · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91da6878-55ea-4bd4-95f4-0c90c856511a · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d0b4b6f-1e3b-4012-9595-31c2f688a340 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 376c5d22-5ea4-4e77-836e-2d90a3d977d4 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 211f0bf1-80de-4007-b261-e93d1d319f58 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe2d6b6-f18a-47d0-80e3-05fa14eef9b6 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7d91f3dc-3361-40e7-9c90-5fae0056b1a9 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e9d53e-ec2e-48c7-af64-a1abd50481d2 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea3df4ac-9fa5-4d50-b165-978d2dfe9784 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Toolformer: Language Models Can Teach Themselves to Use Tools
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d92604c-22f3-4e69-ab1b-a0ab736b2cec · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3c24cbcf-69cf-4df7-b40a-5d88ccaa7dba · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86bda43d-60f2-4d9a-8427-55fae60e7051 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Hazards in Daily Life? Enabling Robots to Proactively Detect and Resolve Anomalies
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21327823-47e2-4976-a1c2-77615640adec · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce08266a-9130-4271-9094-73e58b3c01e7 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16049078-62be-42d1-a50e-9ccac96c5712 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01355dde-d6a2-459d-85e7-52f3ae7a23d3 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 67b53045-bf75-4f75-8880-745f237be6e1 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae5a6918-b9a4-4818-842d-42d62a5bbd7c · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d39aec71-f6ff-46c5-b807-f5cc87ce74b1 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e8f1e3-eb15-4360-b7c2-2deda2d51ef4 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation OpenHands: An Open Platform for AI Software Developers as Generalist Agents
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b6aed6-28ce-4f33-9182-757289dc56b1 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Foundations and Recent Trends in Multimodal Mobile Agents: A Survey
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a53afc44-f45d-41ad-9a2e-36e6e8c9b45a · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfe0fcee-d956-44f2-9803-31d2e0e4053d · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Qwen2.5-Omni Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a46645cd-3d1c-4db9-b8a3-aff88a0a87b4 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3ba96308-d4b5-40ef-8639-ee575a2b19da · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3c23a90b-6986-430d-86d7-7b7ac803d0b6 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 093e1986-e296-40c3-a1b8-2eba82a8abc1 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cb17e108-8c89-4af7-bf13-994577aa12d7 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4529811b-4832-47df-9ae5-e4180b04240e · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0e4689e-5467-46cc-959d-1fc467deadbe · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ac403f-f968-4ebe-a156-1159ee29b95b · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d525c7a2-99f6-4cc1-9d96-938a6d869311 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation KMM: Key Frame Mask Mamba for Extended Motion Generation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe9c6da-fb73-4065-b58c-08377835204f · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cdf0a49-32a4-45c5-a7c1-f182fa2ceabb · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ef6585ef-abd2-45ff-a142-5324e57784e0 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Motion Anything: Any to Motion Generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1da34a8-069d-4838-b864-111ed8088c27 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e687567a-663b-4744-babe-7ccae67079fb · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdaa49b5-5483-487c-a631-1e9abc9c2c40 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation online" 'onlinestring :=
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e030cbe4-6ff5-4983-8a81-176fb28365b1 · outbound
PresentAgent: Multimodal Agent for Presentation Video Generation write newline
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7111bdd9-9630-4eaa-b5a4-aa5db3658a83 · inbound
BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation PresentAgent: Multimodal Agent for Presentation Video Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95488604-1410-474a-8483-f99895a38273 · inbound
OmniPresent: Generating Coherent Presentation Suites from Scientific Papers PresentAgent: Multimodal Agent for Presentation Video Generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.