Pith. sign in

Paper Citation Record · LEDGER

PresentAgent: Multimodal Agent for Presentation Video Generation

As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2507.04036.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04036 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:01:29.172482Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T19:41:58.268576Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c888856-2c80-426c-9a26-0aae806c6ef8 · outbound

This paper cites GPT-4 Technical Report.

PresentAgent: Multimodal Agent for Presentation Video Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:28.433094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:28.433094Z digest=sha256:97ed637e89b4d97d130ceab71e4fb6dea7f68cddea2a6c9962c7c09e78197fe3

Observation 1e7c8d37-d7df-4fcd-8a75-838a4cb2c463 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.670856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:28.569354Z digest=sha256:e211c79776ae696f0cb8d940c0f5a0eea89b520178bf1127c6797852eb846a98

Observation 1983defb-0945-4889-b021-e1d98d195e20 · outbound

This paper cites Qwen2.5-VL Technical Report.

PresentAgent: Multimodal Agent for Presentation Video Generation Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:28.715034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:28.715034Z digest=sha256:382e336ce3edbac18f7f4cdf6780c53b8355ba0208ba5c591342b535a1e37ffd

Observation 88a157cb-a991-44fe-a0be-9b921b80eafc · outbound

This paper cites Longformer: The Long-Document Transformer.

PresentAgent: Multimodal Agent for Presentation Video Generation Longformer: The Long-Document Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:28.793825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:28.793825Z digest=sha256:308fe03b20fc59b84279c2e4b8783f8123b6b6553556c70f6b09b9b645d6eba8

Observation b8208639-f3d4-418f-b57b-3510eb76ba11 · outbound

This paper cites Structure-Aware Abstractive Conversation Summarization via Discourse and Action Graphs.

PresentAgent: Multimodal Agent for Presentation Video Generation Structure-Aware Abstractive Conversation Summarization via Discourse and Action Graphs

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:01:29.561919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:28.960148Z digest=sha256:e0d2e29eeaa3e8e2734b35be4138902a95774ffbff7ba1858bddb87462c2c031

Observation 0ed332b6-4134-4a7a-a66a-3c1564009960 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.664601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:28.993820Z digest=sha256:6527ca529a6569bbe9fc1c8d38bff7d3162a8808e4ee2dc8cb733d639ed6ca74

Observation 224f5bd7-7d46-4e52-9b9a-53f049291259 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

PresentAgent: Multimodal Agent for Presentation Video Generation Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.043152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.043152Z digest=sha256:1be43a13b47990d0520ed47823acef74e9e4d70ded7b0c61ec39ba3feef6f149

Observation 1a67b510-e71b-4042-ab2b-505d9b1d4613 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.658073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.047123Z digest=sha256:eaec5a82fe4cefda355182638ebbc1d374ba98164e73f3438dd06355998b8d39

Observation 94db5d95-d83a-4d52-920d-7c92e15c56fa · outbound

This paper cites AutoPresent: Designing Structured Visuals from Scratch.

PresentAgent: Multimodal Agent for Presentation Video Generation AutoPresent: Designing Structured Visuals from Scratch

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.060673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.060673Z digest=sha256:976d43fbd05aaa0a7a1bb2d7672c148dee0bba002b7a95f406562cde397057d3

Observation 26aa9c68-8bfe-4f3d-8223-5b4d8b68b925 · outbound

This paper cites Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation.

PresentAgent: Multimodal Agent for Presentation Video Generation Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.063984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.063984Z digest=sha256:e2e2464bb3abfcc6923f6c29ce8e80201520f023f5415dd987fdc3a0acbfa44b

Observation 9e819c4f-dea6-4ea6-938f-68a682b9c7a7 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.651742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.066441Z digest=sha256:2bbce6c1d333db638eb8f8bc72c226f11d490833c642ee8c9160dadcdd0d54cf

Observation 7432ab6a-fd75-46f2-abbb-4d672fb2d998 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

PresentAgent: Multimodal Agent for Presentation Video Generation BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.068517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.068517Z digest=sha256:18092cb3ddf1ef321c862f2971fb1070a9ebd19400a8a14280e821fa6c0cff43

Observation 539fab81-7e54-484a-800f-50c400b1b6c9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

PresentAgent: Multimodal Agent for Presentation Video Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.070705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.070705Z digest=sha256:1c59a6a963f2333d472f48ce71a53116d970f208f7aba0b347aaa00d3a2caabd

Observation d594df54-74ca-4f9d-a7c0-8ed5ff2fc62f · outbound

This paper cites VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

PresentAgent: Multimodal Agent for Presentation Video Generation VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.072774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.072774Z digest=sha256:5363b93c9292b8195e487ba9d59aabc4ffe3419af2bf55e51358cfb210b07131

Observation 2eee608d-98fc-44ba-9e0e-63a0d028d367 · outbound

This paper cites VideoGUI: A Benchmark for GUI Automation from Instructional Videos.

PresentAgent: Multimodal Agent for Presentation Video Generation VideoGUI: A Benchmark for GUI Automation from Instructional Videos

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:01:29.519212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.074771Z digest=sha256:b725c8eac3112423c8ed7cae1620db2d415cc244c3ea05cf43dabb7be2f79dd6

Observation 2fd21acb-63aa-437f-97d1-f4854ec9f7cb · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

PresentAgent: Multimodal Agent for Presentation Video Generation ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.076779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.076779Z digest=sha256:af2d1a7e86dc34ceeb5bb395f8f7dc8cd96ac4b1bfe995d0e7aa8d057a44242d

Observation b5e48018-5dff-4b98-a175-551b82a549ee · outbound

This paper cites FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion.

PresentAgent: Multimodal Agent for Presentation Video Generation FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.078845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.078845Z digest=sha256:cef4f102f2870b044ff849e471799bdf8cc94d88dc32cf010516a32ebd9d1b08

Observation 91da6878-55ea-4bd4-95f4-0c90c856511a · outbound

This paper cites OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning.

PresentAgent: Multimodal Agent for Presentation Video Generation OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.081498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.081498Z digest=sha256:2ebbd0ac292077eb6514de2e6d8b0ed5c50a73f091ca11096a2de5eb310aa6d1

Observation 7d0b4b6f-1e3b-4012-9595-31c2f688a340 · outbound

This paper cites UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction.

PresentAgent: Multimodal Agent for Presentation Video Generation UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.083800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.083800Z digest=sha256:1559731a4cdb0e2c6645e966c1d584a74a5117972ec99d7102dce3486903d0bc

Observation 376c5d22-5ea4-4e77-836e-2d90a3d977d4 · outbound

This paper cites Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition.

PresentAgent: Multimodal Agent for Presentation Video Generation Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:01:29.486456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.086266Z digest=sha256:e7fb9d9836515103466d61df45a76d310b4986a1cdaa86d02b4473ca1030b19d

Observation 211f0bf1-80de-4007-b261-e93d1d319f58 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.088530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.088530Z digest=sha256:72dc022871e09d7cfd1fb02f26d71a21fce9b52bcb7c5e31699e46d14ebc5e82

Observation 2fe2d6b6-f18a-47d0-80e3-05fa14eef9b6 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.645601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.090591Z digest=sha256:99c37b9d6ca333fdb0121c117df42536c0c6a75d99f1c1443faa97e99a05f6f0

Observation 7d91f3dc-3361-40e7-9c90-5fae0056b1a9 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

PresentAgent: Multimodal Agent for Presentation Video Generation UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.092791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.092791Z digest=sha256:53b2473995bc20cad9ec7569e4970afa832229d6b3803f71a84dcbe6fd41e24e

Observation 45e9d53e-ec2e-48c7-af64-a1abd50481d2 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.095189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.095189Z digest=sha256:5d5b6ab2568a97f35b2b1c85bb8a5bdd30e2579da2ce08f8e3426e533406a070

Observation ea3df4ac-9fa5-4d50-b165-978d2dfe9784 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

PresentAgent: Multimodal Agent for Presentation Video Generation Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.097148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.097148Z digest=sha256:29d4f686eb9fa338afe65bc2fff1d7e515f7376751022e49d698e89c2d46771d

Observation 3d92604c-22f3-4e69-ab1b-a0ab736b2cec · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.635684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.099566Z digest=sha256:17ea3c3af530801338eb7df5e000b2501f168df97b32c2873b1b06e2b7cb5963

Observation 3c24cbcf-69cf-4df7-b40a-5d88ccaa7dba · outbound

This paper cites Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models.

PresentAgent: Multimodal Agent for Presentation Video Generation Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.101657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.101657Z digest=sha256:554c448217c175cce08f1b64940737b754ba658e9cc195bb068a5ef027833145

Observation 86bda43d-60f2-4d9a-8427-55fae60e7051 · outbound

This paper cites Hazards in Daily Life? Enabling Robots to Proactively Detect and Resolve Anomalies.

PresentAgent: Multimodal Agent for Presentation Video Generation Hazards in Daily Life? Enabling Robots to Proactively Detect and Resolve Anomalies

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.103900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.103900Z digest=sha256:d00b733ea0162e6c6c0020190e965aca0fd6b9449519545ddf96f8f2460ab3f8

Observation 21327823-47e2-4976-a1c2-77615640adec · outbound

This paper cites ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models.

PresentAgent: Multimodal Agent for Presentation Video Generation ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.106548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.106548Z digest=sha256:e4a9dd116df3e1b2fa774d96a65b942b47821bbd7955c531ffaf6b819de092a9

Observation ce08266a-9130-4271-9094-73e58b3c01e7 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.108567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.108567Z digest=sha256:747e964651b9710835e11609f9143a5f210c942de84c50f571c8fc969ab58436

Observation 16049078-62be-42d1-a50e-9ccac96c5712 · outbound

This paper cites OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis.

PresentAgent: Multimodal Agent for Presentation Video Generation OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.110892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.110892Z digest=sha256:5a7a355919927c4c30095acd61e94c7dc2f5b7ca11daeb5bd7bceb1732f27ddc

Observation 01355dde-d6a2-459d-85e7-52f3ae7a23d3 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.629435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.113166Z digest=sha256:1831977dc317822761bd7a1678351ba5fe7035bb13d5ecfe3066ccded0441aae

Observation 67b53045-bf75-4f75-8880-745f237be6e1 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.115335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.115335Z digest=sha256:6b9e9bcbc3cc7e0ad2119511373d7e160881f762876351e5290480525d6db84d

Observation ae5a6918-b9a4-4818-842d-42d62a5bbd7c · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.623264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.118180Z digest=sha256:739fc7d8ecea460e441f4a59117634295d808f717612440d903a7ec722e85115

Observation d39aec71-f6ff-46c5-b807-f5cc87ce74b1 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

PresentAgent: Multimodal Agent for Presentation Video Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.120314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.120314Z digest=sha256:8fa3c3b9142f8cfb12819bb5ff7a8f6de9515d0650ae19f4b4a784b7aa562a20

Observation 49e8f1e3-eb15-4360-b7c2-2deda2d51ef4 · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

PresentAgent: Multimodal Agent for Presentation Video Generation OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.122704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.122704Z digest=sha256:fe4faaf1bf3c60f084fb16e0876ae9b3dfa1508066ed005231134762b98e6eb4

Observation 12b6aed6-28ce-4f33-9182-757289dc56b1 · outbound

This paper cites Foundations and Recent Trends in Multimodal Mobile Agents: A Survey.

PresentAgent: Multimodal Agent for Presentation Video Generation Foundations and Recent Trends in Multimodal Mobile Agents: A Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.125527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.125527Z digest=sha256:359a0157fe6bfee658bd1e877f4566a5e879923d618a15e84a0900e407ee0b5c

Observation a53afc44-f45d-41ad-9a2e-36e6e8c9b45a · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

PresentAgent: Multimodal Agent for Presentation Video Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.127674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.127674Z digest=sha256:803ede69d656c178782c24f4a5179b8596fc89333b263cf703954fc06fd62c8f

Observation cfe0fcee-d956-44f2-9803-31d2e0e4053d · outbound

This paper cites Qwen2.5-Omni Technical Report.

PresentAgent: Multimodal Agent for Presentation Video Generation Qwen2.5-Omni Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.129735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.129735Z digest=sha256:32083fc7893083d0fa2e6472b53fe7966eb15c02cc4ad289b3e2bc7d7011fcda

Observation a46645cd-3d1c-4db9-b8a3-aff88a0a87b4 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.616942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.133015Z digest=sha256:83799c4375a94878472d9ac41937da5dd0b4cb03371523de184a7c428adea5e4

Observation 3ba96308-d4b5-40ef-8639-ee575a2b19da · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.610446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.135592Z digest=sha256:005442d946145abfa2c38aa68f4c499f5e8deadad0d0ccc0412726035341017f

Observation 3c23a90b-6986-430d-86d7-7b7ac803d0b6 · outbound

This paper cites If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents.

PresentAgent: Multimodal Agent for Presentation Video Generation If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.138233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.138233Z digest=sha256:668b6adddaae3d65ca7ba4bd244176df082788806da531df2792520151efc484

Observation 093e1986-e296-40c3-a1b8-2eba82a8abc1 · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.604000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.140908Z digest=sha256:20883769bec545bbe66acc2e8c6181ce1dda4fb11c6d506bd7e9797e47c7e13d

Observation cb17e108-8c89-4af7-bf13-994577aa12d7 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

PresentAgent: Multimodal Agent for Presentation Video Generation MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.143169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.143169Z digest=sha256:15da7ea2c6a3a3995c7d8cfb7da401415a50ffd98baa933e78168e708a9facec

Observation 4529811b-4832-47df-9ae5-e4180b04240e · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

PresentAgent: Multimodal Agent for Presentation Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.145529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.145529Z digest=sha256:74132c7105186d8c31779a0acd6952413645eea9dd9e07f9c804aba7d0a8719c

Observation a0e4689e-5467-46cc-959d-1fc467deadbe · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.147965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.147965Z digest=sha256:6843914b29b7b0a02b2842cfda3620ff0aede132ce079dfafbc0866a52203e79

Observation 88ac403f-f968-4ebe-a156-1159ee29b95b · outbound

This paper cites DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search.

PresentAgent: Multimodal Agent for Presentation Video Generation DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.149835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.149835Z digest=sha256:7d49512911ad36a73c1b5a76841a1311997c92ab0c8a60733a28d469f1476d54

Observation d525c7a2-99f6-4cc1-9d96-938a6d869311 · outbound

This paper cites KMM: Key Frame Mask Mamba for Extended Motion Generation.

PresentAgent: Multimodal Agent for Presentation Video Generation KMM: Key Frame Mask Mamba for Extended Motion Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.152345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.152345Z digest=sha256:dce717efaed9ae71a9c46b6682b587a1dc272b6191faae7ae1f7e2aefa303afa

Observation cbe9c6da-fb73-4065-b58c-08377835204f · outbound

This paper cites InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation.

PresentAgent: Multimodal Agent for Presentation Video Generation InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.154828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.154828Z digest=sha256:1c54c8c017c1f2133dce1c10082e2de39b282b4cb1dd85eee418b03fd8bf21f0

Observation 7cdf0a49-32a4-45c5-a7c1-f182fa2ceabb · outbound

This paper cites an unresolved cited work.

PresentAgent: Multimodal Agent for Presentation Video Generation Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:01:29.593307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:01:29.157288Z digest=sha256:fd36a5824e8fefb1799ad28e034fc8bfd2a41ef755cd33a7dc567d7280d855b3

Observation ef6585ef-abd2-45ff-a142-5324e57784e0 · outbound

This paper cites Motion Anything: Any to Motion Generation.

PresentAgent: Multimodal Agent for Presentation Video Generation Motion Anything: Any to Motion Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.159405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.159405Z digest=sha256:61486f54deb9ec97151587b8f6a4b9ad36172c15afc70ab13149ff1c66b69c33

Observation f1da34a8-069d-4838-b864-111ed8088c27 · outbound

This paper cites Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion.

PresentAgent: Multimodal Agent for Presentation Video Generation Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.161584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.161584Z digest=sha256:a783f6eb3439bab9f18a88658073a038432854a70a68da8d647abcfb249f9b93

Observation e687567a-663b-4744-babe-7ccae67079fb · outbound

This paper cites PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides.

PresentAgent: Multimodal Agent for Presentation Video Generation PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.166158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.166158Z digest=sha256:a0320c71a2c8c7a36db92308ca590c1ded507a4f6b12f41b60256f721bf84e0b

Observation fdaa49b5-5483-487c-a631-1e9abc9c2c40 · outbound

This paper cites online" 'onlinestring :=.

PresentAgent: Multimodal Agent for Presentation Video Generation online" 'onlinestring :=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.168686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.168686Z digest=sha256:41ceae0270dd3a14a89e4bf350f0de723b8ec781f50aa7771c47722cfb9ffbbc

Observation e030cbe4-6ff5-4983-8a81-176fb28365b1 · outbound

This paper cites write newline.

PresentAgent: Multimodal Agent for Presentation Video Generation write newline

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:29.172482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:01:29.172482Z digest=sha256:119c42e1a03eca6696abd2b479f4894c17ae8534669d69bb74dd777a4303e277

Pith citing papers

Observation 7111bdd9-9630-4eaa-b5a4-aa5db3658a83 · inbound

BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation cites this paper.

BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation PresentAgent: Multimodal Agent for Presentation Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T19:41:58.268576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:41:58.268576Z digest=sha256:ebf7423e684e6bc6dc9a321d88b26a6fecf03482553f11e65369ba1f3819f8f7

Observation 95488604-1410-474a-8483-f99895a38273 · inbound

OmniPresent: Generating Coherent Presentation Suites from Scientific Papers cites this paper.

OmniPresent: Generating Coherent Presentation Suites from Scientific Papers PresentAgent: Multimodal Agent for Presentation Video Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T09:30:41.159870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:30:41.159870Z digest=sha256:77e371e4640ce06868da8cd7693d7f833d6a60b62b5c8c48bb1b32a8c58e1c9a