Pith. sign in

Paper Citation Record · LEDGER

PresentAgent-2: Towards Generalist Multimodal Presentation Agents

As of 3 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2605.11363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.11363 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T02:29:42.157339Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:44:26.876152Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact24
  • verified fuzzy15
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dadcfd4b-e521-4cf0-a5f7-c6dbd12b7147 · outbound

This paper cites Paper2poster: Towards multimodal poster automation from scientific papers,.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Paper2poster: Towards multimodal poster automation from scientific papers,

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.469529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:2bbd9242998d13cd27c7576f560a99005fb0c6a573802bbb674441868c60167d

Observation dca908e0-3930-44ee-bc54-25c4166f53c7 · outbound

This paper cites Presentagent: Multimodal agent for presentation video generation.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Presentagent: Multimodal agent for presentation video generation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.757200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:7eff861b357f064084a7ab8ca15eb44d64ba30bd7ade162956a26b98f87ab5dd

Observation bf28b32b-a132-46a0-a0c1-93e7cac92259 · outbound

This paper cites Q., and Shou, M.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Q., and Shou, M

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:32:06.475438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:e0379420bfad8602960cb2ff95f1eaf20e7d89067c0a25b01c8c86939315b099

Observation cc8e9da5-6959-4f1c-9338-4917e45097e1 · outbound

This paper cites VideoAgent: Personalized Synthesis of Scientific Videos.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents VideoAgent: Personalized Synthesis of Scientific Videos

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:32:06.466432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:4eb7da71e8d9c0223597d895953746b54772063cb7d78b99d18963a0ed3190f7

Observation ebd0c220-9ded-4d65-92c2-c9d60055bc7f · outbound

This paper cites Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:32:06.472581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:5c3ae14a2ecfa14ecb34a08be046af6562121bc9e99aa310411b9d378369a243

Observation 72818650-064f-4dcc-9249-215894a13ad0 · outbound

This paper cites Pptagent: Generating and evaluating presentations beyond text-to-slides.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Pptagent: Generating and evaluating presentations beyond text-to-slides

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.762408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:39788f84f422fd8f9eeefcd6cd1460f98ed3eb19f641154402fbe8332e0d0a97

Observation ab4f3b05-932f-4b06-b539-01fd8e637cc3 · outbound

This paper cites Auto- Slides: An interactive multi-agent system for creating and customizing research presentations.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Auto- Slides: An interactive multi-agent system for creating and customizing research presentations

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.478070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:a65ac25bad10406924941aa094807a279968585253ae15439d3c0ad331aade0c

Observation 371d8d01-54df-4256-8870-04a5cf958f9c · outbound

This paper cites Node-based editing for multimodal generation of text, audio, image, and video.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Node-based editing for multimodal generation of text, audio, image, and video

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.480754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:f0f145437b6a5f45618add4fa4e9de04d0e88d44cddb010608f15de589e1e4db

Observation 3d7b7d1c-4e97-49d1-9e0f-205daa78cebf · outbound

This paper cites PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.483612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:4003f40e65f6a28c0c5e3bb45adf49c9d5738e9e03131abe1881112b97bccecc

Observation e91fb8d3-1b09-4292-9d6b-8aa21c96918c · outbound

This paper cites Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.441607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:9cb1b5041a4b17bd794f4ebc2fe90b0d09e32637b6040c89b12ffdabf41e96a1

Observation b49932d2-ed43-4638-8890-6fc98df6d966 · outbound

This paper cites Autopresent: Designing structured visuals from scratch.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Autopresent: Designing structured visuals from scratch

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.695161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:f4496d52d91621e30e21cbac6e3dcc23faf46b97c8ac008aad4b60b76dc1613e

Observation f0312a9d-97e9-4134-bf10-12baeb471f50 · outbound

This paper cites Infinity parser: Layout aware reinforcement learning for scanned document parsing.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Infinity parser: Layout aware reinforcement learning for scanned document parsing

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.415184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:9a6b3bb404a5a8e19cf02817df676f0fb7a82f6e9295cb9acfe5a78f660fb536

Observation f796951f-2d8a-4733-b4bb-edcb5d882a95 · outbound

This paper cites Doc2ppt: Automatic presen- tation slides generation from scientific documents.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Doc2ppt: Automatic presen- tation slides generation from scientific documents

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.708349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:9b694f44e3f795f9d17045145d54de01688f135950b21288fc9c2a1efb32c823

Observation 42ef8323-5ca5-41b8-8b66-f24bdc886cb5 · outbound

This paper cites Slides agent: An intelligent agent for creating and analyzing presentations using large.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Slides agent: An intelligent agent for creating and analyzing presentations using large

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.690919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:b91b2bdc12b58355892928c1c0d31be8419ecbfffdad230e8212339463664d94

Observation 74429580-9ced-489c-a3c5-6da378344b90 · outbound

This paper cites SlideGen: Collabo- rative multimodal agents for scientific slide generation.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents SlideGen: Collabo- rative multimodal agents for scientific slide generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.407697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:44ffa2ee6569a30e8a47382b2500babbc2d0f6c1d0316502f5929b320845d600

Observation 59329108-f988-4074-b80f-c1983545e8b7 · outbound

This paper cites Presenting a paper is an art: Self-improvement aesthetic agents for academic presentations.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Presenting a paper is an art: Self-improvement aesthetic agents for academic presentations

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.457138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:149b288c4d340ddd70c03f0164e42c2e21a1b6d4fe7e6487af08857260c47352

Observation cf498312-75c4-42dd-92e9-6958df28398c · outbound

This paper cites Gpt4tools: Teaching large language model to use tools via self-instruction.Advances in Neural Information Processing Systems, 36:71995–72007.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Gpt4tools: Teaching large language model to use tools via self-instruction.Advances in Neural Information Processing Systems, 36:71995–72007

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.727964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:e5431c8751df8fa66528271e46d367ff2d2f505ae34f01c51794791325396e17

Observation c0c735ad-5fb9-4cee-b0d3-93f8f5adc520 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:17:58.977376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:d436dff04433ed45bb0a9a25421c3aa6107fbe158d1e0419bb486cd0b2074cd5

Observation ca5e12fd-76b9-45a0-9e7f-7bd3e36c8ec3 · outbound

This paper cites Os-genesis: Automating gui agent trajectory construction via reverse task synthesis.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Os-genesis: Automating gui agent trajectory construction via reverse task synthesis

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.680217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:e41e4fc1fbc01f701aeb756fa4c510983d49f04d1a738690d6a3886b2e1c238c

Observation 72a6e0f9-afd5-412c-990e-12fad2d59c2d · outbound

This paper cites VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.435877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:d6e20f169e392ed47517e36966996ba3849e00478a0b7fdf838c578feed5adad

Observation e788e6dc-dcf5-4cb9-b020-29fe7d053bc7 · outbound

This paper cites Phyt2v: Llm-guided iterative self- refinement for physics-grounded text-to-video generation.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Phyt2v: Llm-guided iterative self- refinement for physics-grounded text-to-video generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.738177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:3e11c23c2c22ab459268d4f463d04d5cfb139b337e1da4dbf9cd3d3216c36963

Observation 9a6fdb57-0869-4d03-92ee-a2f67503722a · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:32:06.444424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:c5fa60d27662f8bee2731c65318015c2af3714021f714fb759e5fd821d6ffe7c

Observation 2f886056-ac37-44db-abbc-fb5d28659806 · outbound

This paper cites Unified multimodal understanding and generation models: Advances, challenges, and opportunities.arXiv preprint arXiv:2505.02567.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Unified multimodal understanding and generation models: Advances, challenges, and opportunities.arXiv preprint arXiv:2505.02567

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.447724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:9df42fcb04c608e232f1636709712c327396c983630efa9ba3f3f106619aa995

Observation 5d738567-e7c6-46ff-91ca-6589dbf51a03 · outbound

This paper cites Qwen3.5-Omni Technical Report.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Qwen3.5-Omni Technical Report

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:32:06.418990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:b8cafc51a8bad1fb65f613f05dc7aea56b70fa392d521ddd14249bb450504fcf

Observation f6643a88-8a2c-42d7-8a54-82467eabc718 · outbound

This paper cites Motion Anything: Any to Motion Generation.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Motion Anything: Any to Motion Generation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.423233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:cfe9459eeff6c949b47eb56b6fbbeadf7315f9c39846468ddfa880572f40b4ec

Observation 497a87b8-55f1-42c9-9016-e07f2be068d3 · outbound

This paper cites InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.430047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:3ca6c295d5e5ca7a62650ca157d3e7ebcfd4a271874a2baec5516d5cd686274d

Observation 192b69c6-c65c-4666-9e76-c46796326fb2 · outbound

This paper cites KMM: Key Frame Mask Mamba for Extended Motion Generation.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents KMM: Key Frame Mask Mamba for Extended Motion Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.463488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:6ef3675394286d84fc7d343a349f7df138aef0362e186cca78894e063cb4468f

Observation 39adb5c7-7403-4651-b8b3-72cef9de1c0f · outbound

This paper cites Motion mamba: Efficient and long sequence motion generation.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Motion mamba: Efficient and long sequence motion generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.685930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:edbab59dd659551d39e9f7223f9a37f74078f7d790a2725c44963d111f6b0da0

Observation 629962e1-8b72-465e-ae00-3cdf44883a88 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Evaluating object hallucination in large vision-language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.699752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:c92b5fd7969abb7735bfcb17769a9d98438f0e2cea00d506322d76b4f7ad1836

Observation 4bcdc1dd-94f4-4e24-ac6f-01404c611947 · outbound

This paper cites Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data.Nature Communications, 16(1):7866.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data.Nature Communications, 16(1):7866

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.720574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:cf8108243fdac4a8b7411d218f712e91bfa3f263affdb9f7fc1426e5fa77a4c3

Observation 7479da98-217e-44fa-a73a-b5e9b42c56b1 · outbound

This paper cites Mavis: A multi-agent framework for long-sequence video storytelling.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Mavis: A multi-agent framework for long-sequence video storytelling

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.748374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:13068bc4b495c56518a9a7ba1ff8208d6afcf85a10a9a6062c8f075c5e4a78d1

Observation 073d4230-c1cc-4e16-9ade-febcb4b5aecb · outbound

This paper cites Multimodal content alignment with llm for visual presentation of papers.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Multimodal content alignment with llm for visual presentation of papers

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.673644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:c956c1c4080be11925f0b2b60c862da1e5899f5d39b200cf24483ec8f3b76311

Observation beb96f69-932c-425c-b445-7a1bdeb29323 · outbound

This paper cites PreGenie: An Agentic Framework for High-quality Visual Presentation Generation.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents PreGenie: An Agentic Framework for High-quality Visual Presentation Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.453852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:c45c811ddb8beeaaf63a4eb6a13831666949d097d8aec84708965e2c14343e5a

Observation da205164-c5ba-401f-a8c7-a3238b697d56 · outbound

This paper cites Present- coach: Dual-agent presentation coaching through exemplars and interactive feedback.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Present- coach: Dual-agent presentation coaching through exemplars and interactive feedback

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.460526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:7431613eafeb4f2622a60e862df58a066ad286e3ebd73664cf3513e94c07615a

Observation 9ffebeec-3f56-426a-8cb1-7dddfbadebf8 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Emerging Properties in Unified Multimodal Pretraining

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:32:06.426637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:cab10d16b7f92551ac5c90b698cba4baf184bfbb5e055f2f4697ff7b43c0d6a9

Observation 8c6ce789-a5c0-4e4a-889c-b4dc7ac5ed4b · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:32:06.438819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:ef5338580e4eab1ae646c2a077e0bc5e259a55a8dad5eb30f4953b9e097e89ec

Observation 248d91e6-138a-4823-9b10-dbb823043333 · outbound

This paper cites Showui: One vision-language-action model for gui visual agent.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Showui: One vision-language-action model for gui visual agent

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.743426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:a9d9dc3fecafcb13f84ba88fa646c68a97398e85412e8478b5d635633f2d0db6

Observation eb0ae429-ec71-4d5e-aea3-5c5c84b69a8e · outbound

This paper cites VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.450743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:eac1ef7018a249c5fec660e2951f68251f146eacb81f52c05f5055fa193be53d

Observation c1cd4cc8-e650-4584-975c-21b4dfa43538 · outbound

This paper cites Videostudio: Generating consistent-content and multi-scene videos.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Videostudio: Generating consistent-content and multi-scene videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T13:02:49.668443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:74dd32b3f565968b04fa8a368646693a9b32ac39f277bc85ab04a31ea7738076

Observation 4de520cb-cae8-439d-bf54-36ae52f0b2c4 · outbound

This paper cites LLM-grounded Video Diffusion Models.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents LLM-grounded Video Diffusion Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.411311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:a374f4999801dea4c0eaec17330413c358ca27ea4ec777a37caef3f4a7d30070

Pith citing papers

Observation 06151ab6-5875-475d-8f3d-03475fa27c04 · inbound

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog cites this paper.

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog PresentAgent-2: Towards Generalist Multimodal Presentation Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T08:44:26.876152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:44:26.876152Z digest=sha256:653cdaaba428eb823b902bab75fa7189c165ad3150309ba532be9842318d91d3