Pith. sign in

Paper Citation Record · LEDGER

Image-to-Video Diffusion: From Foundations to Open Frontiers

As of 4 August 2026, this Paper Citation Record lists 100 of 190 outbound references and 0 inbound Pith citation observations for arXiv:2605.17248.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.17248 v1

Coverage vector

measured 100 of 190 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T15:06:02.084336Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 190 outbound references displayed

  • verified exact42
  • verified fuzzy56
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e67b3fd2-7807-4c03-8bda-11054505215b · outbound

This paper cites Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM.

Image-to-Video Diffusion: From Foundations to Open Frontiers Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.017943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:78ac18f47902b9f3cefb8d71e56e9f4ec5e4fbdf51399a448b093299f4ce6ed8

Observation 36b8d225-314c-4049-9e5f-ce747b065a28 · outbound

This paper cites Levitor: 3d trajectory oriented image-to-video synthesis.

Image-to-Video Diffusion: From Foundations to Open Frontiers Levitor: 3d trajectory oriented image-to-video synthesis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.595639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:0a316e7c7b5ae8d7d1aab11459adf3abf4aa07fdd7c1852b8b2b8a53726998fd

Observation f606e89c-42d9-4b8c-8850-f94e22b8f6c0 · outbound

This paper cites Motioncanvas: Cinematic shot design with controllable image-to-video generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Motioncanvas: Cinematic shot design with controllable image-to-video generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.597679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:4e9e44b6eefaad187699523fccfc86b75badaceeaeaef121737a93197984ea9c

Observation d2b3a2d2-86c6-4dd5-9d3e-b9c204d4e716 · outbound

This paper cites Versatile transi- tion generation with image-to-video diffusion.

Image-to-Video Diffusion: From Foundations to Open Frontiers Versatile transi- tion generation with image-to-video diffusion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.572179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:5dfe3930d3e4dfd44304511bc14d9abc882dc0a02bc80eae960010818d7e74b6

Observation ba14aaba-fcb0-4aaf-9ec2-7df6539dcc39 · outbound

This paper cites I2V3D: Controllable image-to-video generation with 3D guidance.

Image-to-Video Diffusion: From Foundations to Open Frontiers I2V3D: Controllable image-to-video generation with 3D guidance

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.945708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:86c63945577ad5303eea3525f3737634c5ccf8b25e198f3ad9817177dbecec3b

Observation 2863e652-76e9-4999-821f-9f6253e99504 · outbound

This paper cites Realcam-i2v: Real-world image-to-video generation with interactive complex camera control.

Image-to-Video Diffusion: From Foundations to Open Frontiers Realcam-i2v: Real-world image-to-video generation with interactive complex camera control

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.561052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:849fcfb4bc2650e06e1b508b7853dcb625a0faaaf1a4f945e36f9cd19fc10453

Observation c7148ad5-71d6-47eb-8d88-81dcf2fdcb97 · outbound

This paper cites Extrapolating and decoupling image-to-video generation models: Motion modeling is easier than you think.

Image-to-Video Diffusion: From Foundations to Open Frontiers Extrapolating and decoupling image-to-video generation models: Motion modeling is easier than you think

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.549321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:aefbefb1377142e9155fbbf4894cd0569939521c2e62aaf282cf1f7ece290602

Observation a8fd0231-a46d-4d01-980f-c96e6e9c270d · outbound

This paper cites TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment.

Image-to-Video Diffusion: From Foundations to Open Frontiers TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-27T02:05:10.361398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:bc8f7fce5633735bb9636f2a0624553f70a3c26c0c7a9fd67d1d360f4a26cb62

Observation 0d7e91d0-8ef4-4653-bb30-84df74251c45 · outbound

This paper cites Auto-Encoding Variational Bayes.

Image-to-Video Diffusion: From Foundations to Open Frontiers Auto-Encoding Variational Bayes

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:24.904452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:3a27dfeb52fbbc3f6cd57a9e4144893b07d454ef21dc6f121b2fe5340d4ba734

Observation 193e0cdc-121c-4244-86cc-c45c3e6c1ebd · outbound

This paper cites I4VGen: Image as Free Stepping Stone for Text-to-Video Generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers I4VGen: Image as Free Stepping Stone for Text-to-Video Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.059265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:06a2a88e9ffb348cf8d346054297734e50151232f06520793420c0501f2762f8

Observation f3495b01-7d04-45d5-9b88-553a08e7dddf · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

Image-to-Video Diffusion: From Foundations to Open Frontiers Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.886841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:93e6c9463255e7de9f7380066dd10f4aaca91e96fb05592a4aa6af29945c807a

Observation b8b8d41d-065a-4a34-9a07-dda032062504 · outbound

This paper cites Rv-gan: Recurrent gan for uncondi- tional video generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Rv-gan: Recurrent gan for uncondi- tional video generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.602035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:f4c2606e13aef1961cec74dac34bcf117eac6d2b581b0e18f8d96b31a2e8a062

Observation 3a0dd6ce-d37e-44d4-8473-6290a1b2cab9 · outbound

This paper cites Styleinv: A temporal style modulated inversion network for unconditional video generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Styleinv: A temporal style modulated inversion network for unconditional video generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.592075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:58d04dc1475cd82fc79fc9688f28c39c09172a62f1d82c6abd3605de139bf8ad

Observation 6f9771aa-e97c-445f-a9d8-ac3d78b021c7 · outbound

This paper cites RealisDance: Equip controllable character animation with realistic hands.

Image-to-Video Diffusion: From Foundations to Open Frontiers RealisDance: Equip controllable character animation with realistic hands

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.993741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:61a0a51dde9ba5d479b96812eafd39d8a73353123e4aab467cd8cdcdc884b38f

Observation 31d8d594-9608-436b-b27c-258c85e8b316 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Animate anyone: Consistent and controllable image-to-video synthesis for character animation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.593878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:d870f78508e698eb61fea725d7dd78c1e1b216e0a92f62388d0daa177f3801eb

Observation 921615d0-a7ed-4614-a581-37658dc3e769 · outbound

This paper cites CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:25.015035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:e3d331bc65d2dff8e30c7f3d8ea1edde11f2475a36b75f2ea6f1dbec21a48f6f

Observation 753366d8-3d0b-485e-b266-37b2cff9b381 · outbound

This paper cites Mofa-video: Controllable image animation via generative motion field adaptions in frozen image-to-video diffusion model.

Image-to-Video Diffusion: From Foundations to Open Frontiers Mofa-video: Controllable image animation via generative motion field adaptions in frozen image-to-video diffusion model

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.585731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:f37c811195bd5ac55cc7e8e73e3367e20904e180141f6f06479f379062972373

Observation 2a146f3e-34a5-431d-89a5-f7af8da1e221 · outbound

This paper cites Champ: Controllable and consistent human image animation with 3d parametric guidance.

Image-to-Video Diffusion: From Foundations to Open Frontiers Champ: Controllable and consistent human image animation with 3d parametric guidance

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.587975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:f8602a0b3051f0047efcab5f4de349128d8e7106b446fe475b2b5251620dc94d

Observation cac2e851-c6ac-468a-a4fe-7e729446be15 · outbound

This paper cites Human video generation from a single image with 3d pose and view control.

Image-to-Video Diffusion: From Foundations to Open Frontiers Human video generation from a single image with 3d pose and view control

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.996575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:665d0d2eeafac09329292421038ac656fd7232babf99aa1e070daa3bbaf4163b

Observation 87de3532-45ad-4ba8-be0a-02b787ab86e8 · outbound

This paper cites Pia: Your per- sonalized image animator via plug-and-play modules in text-to-image models.

Image-to-Video Diffusion: From Foundations to Open Frontiers Pia: Your per- sonalized image animator via plug-and-play modules in text-to-image models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.589952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:1905787d8541a06f1e31a4d3e7f8815290b012d8f94427f34de84089896c624a

Observation 3ad9652d-2f89-4682-8a22-79022ea3db3e · outbound

This paper cites Customcrafter: Customized video generation with preserving motion and concept composition abilities.

Image-to-Video Diffusion: From Foundations to Open Frontiers Customcrafter: Customized video generation with preserving motion and concept composition abilities

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.583666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:5fbc7319db521c42f2ae43197e23681ab6e8018a4eb700b30e3e9dd2533efa12

Observation 914f5b0a-7577-441f-945c-56e29a4232ea · outbound

This paper cites U-net: Convolutional net- works for biomedical image segmentation.

Image-to-Video Diffusion: From Foundations to Open Frontiers U-net: Convolutional net- works for biomedical image segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.599679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:5555019acab792a989f2d195c00b4f5aa7fc0c78dc271c3546f49421460e95b3

Observation 8efb43ff-674e-4034-8dc9-8749e1d6bfd7 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Image-to-Video Diffusion: From Foundations to Open Frontiers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:24.836751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:e1c1f82f0337f149287e13af23019c0c35bb7d8e202935dd4d8d79f8b04004e6

Observation aee5d195-de3b-412b-ab50-00ee33455078 · outbound

This paper cites Megactor-sigma: Unlocking flexible mixed-modal control in portrait animation with diffusion transformer.

Image-to-Video Diffusion: From Foundations to Open Frontiers Megactor-sigma: Unlocking flexible mixed-modal control in portrait animation with diffusion transformer

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.569624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:15bcdec8c0a62606878c9f7a46692c900d5a6bb63fe3a851d786fb29bd1ef5d5

Observation 33d676b7-41da-4ff0-8520-737134646b04 · outbound

This paper cites Tora: Trajectory-oriented diffusion transformer for video generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Tora: Trajectory-oriented diffusion transformer for video generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.576226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:2682e8b73e8d1ba01306ff854454edca2a10ca1b3b94d7c52084c70341736622

Observation 5881a7e2-485c-4bac-a8df-f1ade38fdddb · outbound

This paper cites HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.067557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:2009dbf08e85d031da76aec04d28fd24351e2264f81bc15d2ab21f90f921f489

Observation c1e261c3-3a7b-421e-a903-81b430fcc8f9 · outbound

This paper cites A survey on video diffusion models.

Image-to-Video Diffusion: From Foundations to Open Frontiers A survey on video diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.574194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:2bfa9f6e954b6b0265722a486926fc1dbe2955720b8baa7bac55219f39337bc5

Observation a1b7d3c8-b19c-40d2-beb2-db4152aff9d0 · outbound

This paper cites Survey of video diffusion models: Foundations, implementations, and applications.

Image-to-Video Diffusion: From Foundations to Open Frontiers Survey of video diffusion models: Foundations, implementations, and applications

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.907470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:be23b059a615febcf3d8eb09a20bbfe7dc124db31f116f6b14b6b706e4fdc946

Observation 0051b4b2-b253-4fb3-816a-fe8b3c481697 · outbound

This paper cites Controllable video generation: A survey.arXiv preprint arXiv:2507.16869.

Image-to-Video Diffusion: From Foundations to Open Frontiers Controllable video generation: A survey.arXiv preprint arXiv:2507.16869

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.091177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:061c14a4b42157ef8068e9c8968c61762ae38b6a83fd344206767cd9efa0e557

Observation 252cea71-a2fe-4c2b-96ba-92c193f26b01 · outbound

This paper cites Diffusion Model-Based Video Editing: A Survey.

Image-to-Video Diffusion: From Foundations to Open Frontiers Diffusion Model-Based Video Editing: A Survey

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.877644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:d637bdea11680d5364b684e3ba8cf708e520077c97709c5343ad87812c329613

Observation 709d4a90-7986-46b4-8c7d-236a2f05d178 · outbound

This paper cites Efficient diffusion models: A comprehen- sive survey from principles to practices.

Image-to-Video Diffusion: From Foundations to Open Frontiers Efficient diffusion models: A comprehen- sive survey from principles to practices

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.699322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:ce5518a8fe5449c32f356b0372c5bcd9ad3defa83e42288e63de83f57f962af2

Observation 47dfc75c-8cb1-4943-b7be-0a6dadbc1dbd · outbound

This paper cites Bridging text and video generation: A survey.

Image-to-Video Diffusion: From Foundations to Open Frontiers Bridging text and video generation: A survey

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.895678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:a2b1a5d8a7fc2fa5283c39b3608a6eef3c5097e5aed3c07205d316532d0c4e12

Observation 53442def-81b5-482c-a2af-20333ea2ec21 · outbound

This paper cites Human motion video generation: A survey.

Image-to-Video Diffusion: From Foundations to Open Frontiers Human motion video generation: A survey

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.742894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:84e956dc3f7cf85329e02a0e1711a814578cdcd99fba743f2d7952ce6f41793a

Observation 1ac307b0-5ed1-49e9-a916-54f830d85d60 · outbound

This paper cites Diffusion models: A comprehensive survey of methods and applications.

Image-to-Video Diffusion: From Foundations to Open Frontiers Diffusion models: A comprehensive survey of methods and applications

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.607968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:750ea6aefd4fbbee947065bd03cd847879206d4d65c12775d921a52fa180c3e5

Observation 97900c18-38a0-4e88-926e-cfd18c913156 · outbound

This paper cites Humanvid: demystifying training data for camera-controllable human image animation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Humanvid: demystifying training data for camera-controllable human image animation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.628151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:582c736067b12ab89fa92a93c83d470df1a77af79ad3ce05d7998339b75c81e6

Observation 4276a0ce-03d1-493a-8389-9196f2bd4039 · outbound

This paper cites Unianimate: Taming unified video diffusion models for consistent human image animation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Unianimate: Taming unified video diffusion models for consistent human image animation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.626243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:9d001efe767b65ead44407599267b13b71c0078bc5e856c691111a5e8378d97d

Observation f1ae7e2f-9896-4f4f-b897-11cd13479180 · outbound

This paper cites Dimensionx: Create any 3d and 4d scenes from a single image with decoupled video diffusion.

Image-to-Video Diffusion: From Foundations to Open Frontiers Dimensionx: Create any 3d and 4d scenes from a single image with decoupled video diffusion

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.630020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:75aac9b09a81a27e6e3b52aecbb5db7ece97a2f71853dd4a73e045f0a6be3d67

Observation c60ef36a-043e-4396-8ea5-db360eb7b6f1 · outbound

This paper cites Phyrpr: Training-free physics- constrained video generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Phyrpr: Training-free physics- constrained video generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.847199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:e764cb7fb1bf62d7b62520132ec7d4c843dd1ce853eed76cc5a3be60f33e20b5

Observation 2d89cf5b-0feb-47bf-84e1-6f44294690f5 · outbound

This paper cites Prompt image to life: Training- free text-driven image-to-video generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Prompt image to life: Training- free text-driven image-to-video generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.567518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:8e9f997b1420bfb705923dd879defe1cb350b545934e0561d0978ac61048f7cd

Observation 9d4839fd-97af-445d-ad7f-a84ba8a67f84 · outbound

This paper cites Seedance 1.0: Exploring the Boundaries of Video Generation Models.

Image-to-Video Diffusion: From Foundations to Open Frontiers Seedance 1.0: Exploring the Boundaries of Video Generation Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:25.044183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:65f50cbfa2a9211ccf2dfe03deb45bd45977e2e60e203ce61c33440e4ac37630

Observation ff07801d-dd27-4c82-bd21-de0febbd925c · outbound

This paper cites Sana-video: Efficient video generation with block linear diffusion transformer.

Image-to-Video Diffusion: From Foundations to Open Frontiers Sana-video: Efficient video generation with block linear diffusion transformer

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.850686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:a94026139678a1d0fc09049b03a44db559583c5a28991db45b4233c6598b8b79

Observation d54cba92-b90b-4233-be10-6d9665ea6936 · outbound

This paper cites Frame context packing and drift prevention in next-frame-prediction video diffusion models.

Image-to-Video Diffusion: From Foundations to Open Frontiers Frame context packing and drift prevention in next-frame-prediction video diffusion models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.565296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:89da5c5b35e651dad495e40d97673ba89744276ea6e9b6df93b237168effb332

Observation 87999a67-c371-4800-8290-61e5181416ec · outbound

This paper cites Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance.

Image-to-Video Diffusion: From Foundations to Open Frontiers Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.999689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:8467854c8572370d80b95685587de58bbca45eda57a3fe8a52bb84159595caaa

Observation 08da3ec7-550d-4401-881e-fee6273673be · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

Image-to-Video Diffusion: From Foundations to Open Frontiers Make-a-video: Text-to-video generation without text-video data

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.579960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:3c36c3ff7b453243e47db43f8e9145e51700211418973833ac0914db43d55a83

Observation c90d52e5-8439-4256-a2d9-cdde285f89b4 · outbound

This paper cites Mvportrait: Text-guided motion and emotion control for multi-view vivid portrait animation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Mvportrait: Text-guided motion and emotion control for multi-view vivid portrait animation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.559172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:ae495aba43aa4042be864c740806d85628d850d69e4c28476414efab806e441e

Observation c769b77b-3a87-4cdf-99cd-6f8a59045b7a · outbound

This paper cites Stiv: Scalable text and image conditioned video generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Stiv: Scalable text and image conditioned video generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.563168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:a47f26e4df5073628fda4f38e19c81004e5921ac312b4340f41a28fe1b560558

Observation b59f4670-e35c-47cb-8176-e870111440fa · outbound

This paper cites Conditional image- to-video generation with latent flow diffusion models.

Image-to-Video Diffusion: From Foundations to Open Frontiers Conditional image- to-video generation with latent flow diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.557179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:9fe79c2477b80ec7507e21ebc3f6e2f9edb59645d0891388764c9b6f5f23c2f5

Observation b30dd2ac-47dc-404d-af0f-14835633f96d · outbound

This paper cites MarDini: Masked Autoregressive Diffusion for Video Generation at Scale.

Image-to-Video Diffusion: From Foundations to Open Frontiers MarDini: Masked Autoregressive Diffusion for Video Generation at Scale

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.020789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:0ed68bc5b9bebb2197361ced3f4d384db3fc6ff70f7440d98d1774faa2f5e3da

Observation aa75d62e-dd6e-42ea-84ce-5021711de27b · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

Image-to-Video Diffusion: From Foundations to Open Frontiers LTX-Video: Realtime Video Latent Diffusion

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:24.939443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:eacb03d2e5f47877f77fe8e7d267620a56a94e36305471bc2e3c84ad389d2b6f

Observation bd2ae0c6-21ff-4ae1-8707-21584d30fc06 · outbound

This paper cites LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion Models.

Image-to-Video Diffusion: From Foundations to Open Frontiers LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.898728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:197e67c20f4e87fbe6eb58dd26057bd4d50e35011777d6934fafe11c39da3946

Observation 42fc4fb1-57d1-4d35-ad35-28bc910e1473 · outbound

This paper cites Motion-i2v: Consistent and controllable image-to-video generation with explicit 14 motion modeling.

Image-to-Video Diffusion: From Foundations to Open Frontiers Motion-i2v: Consistent and controllable image-to-video generation with explicit 14 motion modeling

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.735909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:e37a86d584278207f6353b3ada467b205ab3dfa15836c887e1825e8657deeca1

Observation 77388e56-31c6-4c5e-bf48-3bd726630156 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

Image-to-Video Diffusion: From Foundations to Open Frontiers Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.551277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:9ef9ce0454cdbcb09269ecb9086bc91b1d6cbeb4f11aef8b7713a1fa57a3dba1

Observation 713a324d-588e-4ba3-867f-14646e07b39a · outbound

This paper cites Reducio! generating 1k video within 16 seconds using extremely compressed motion latents.

Image-to-Video Diffusion: From Foundations to Open Frontiers Reducio! generating 1k video within 16 seconds using extremely compressed motion latents

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.553174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:d88985873f8bd13a1e3b7ac01e3158d463ba1afc83e2a319c6217b1e1cf59871

Observation 8b9bb746-04b7-413a-b73f-c0d44a52ca97 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Image-to-Video Diffusion: From Foundations to Open Frontiers Wan: Open and Advanced Large-Scale Video Generative Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:25.050301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:96eb1039c1fd4ed188af1f6f835489c47ad05a6997f5ac427123d375aa5c7031

Observation 952a6411-6e02-4801-9322-f79a5b6cf059 · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

Image-to-Video Diffusion: From Foundations to Open Frontiers Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.547417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:05e8a650006722da177dc88f6eaaa5d6f14e60f322c4dee202f5e2e9734913a0

Observation 236a9de2-7f14-49b4-af42-24272acc0e49 · outbound

This paper cites TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models.

Image-to-Video Diffusion: From Foundations to Open Frontiers TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.053134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:c6861c07d46bbc67b3c715739a9cbe01202fd82b0b7dddcfc151af5633d4fd1c

Observation eb16766a-0513-4135-a20b-e9772afed197 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

Image-to-Video Diffusion: From Foundations to Open Frontiers I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:24.859885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:33d9daf3f25b785141985dfb2a711e4ad2e3e076ee093dce2071b522eec8ae90

Observation 4e0f72f5-7537-430e-9e14-ceaa964ba3c2 · outbound

This paper cites Trip: Temporal residual learning with image noise prior for image-to-video diffusion models.

Image-to-Video Diffusion: From Foundations to Open Frontiers Trip: Temporal residual learning with image noise prior for image-to-video diffusion models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.555060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:609b059ac1cb916723799d6facbb52e5c7481ad27f3cb3e17b3a5f9d5dfc093a

Observation 27c79576-e43c-47a0-a828-d624a12e2aaa · outbound

This paper cites AtomoVideo: High Fidelity Image-to-Video Generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers AtomoVideo: High Fidelity Image-to-Video Generation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.041355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:2bca4c3cd08bb6606cc62a087bcbf2d5094fb499e39031d15794dc9c30322d0a

Observation 25ccfc7f-3670-4f41-ba59-7101143c5c8d · outbound

This paper cites Ti2v-zero: Zero-shot image conditioning for text-to-video diffusion models.

Image-to-Video Diffusion: From Foundations to Open Frontiers Ti2v-zero: Zero-shot image conditioning for text-to-video diffusion models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.578154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:1e75f6aa9826e353fb35ab38d8af45b583e95987c6d94905f1318f834e239778

Observation b925b51b-5b2a-43ff-90fa-978e17d156c0 · outbound

This paper cites ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.892899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:eb971f56e75d5b620e48233c5723b2e402ac567da1a514eaae5c49f040804d4b

Observation 72c27364-3008-4e8f-873d-13a92854dac0 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:24.927120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:d80b8bce93e3af15b8b1dec2d990f6ed0ec41e5689ac809802afad3f47a6922d

Observation 0152ffb5-5d5d-4d88-98ed-75ffe5e63772 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

Image-to-Video Diffusion: From Foundations to Open Frontiers Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.545559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:c84f7460fcff641e02251cbb085c21dc0a4fd835692315706373f7832a3c877f

Observation d65c2ceb-7f6b-47d5-9629-fd529f0fc90d · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

Image-to-Video Diffusion: From Foundations to Open Frontiers Open-Sora: Democratizing Efficient Video Production for All

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:24.864731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:ae0dc0de564b2d4fec064759e1f5bad4017f8e65cfc9822e5b9c22799408eda2

Observation 46a88673-7abf-4a28-97a7-f456d19489c6 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Image-to-Video Diffusion: From Foundations to Open Frontiers Open-Sora Plan: Open-Source Large Video Generation Model

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:24.883653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:ddb64eff16f95d7adb4acc575d1c5be082c4f44c36170832bc8aad52ca29cdbd

Observation 7d177fc6-e2c0-42ba-8fb0-964812bed541 · outbound

This paper cites CogVideoX: Text-to-video diffusion models with an expert transformer.

Image-to-Video Diffusion: From Foundations to Open Frontiers CogVideoX: Text-to-video diffusion models with an expert transformer

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.541906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:807a9500362a9fa9025246f20489cf745e090911c85c97703a509ab43b9b89a5

Observation 86065fa7-ad9c-4285-8ca3-8819491d0f8d · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Image-to-Video Diffusion: From Foundations to Open Frontiers HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:24.880796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:82eedcaf8aff3228a993f0dd595ecac6980a91516849ab54a573b26bd2218250

Observation 0bf08bae-bc58-4366-a9e9-8e3e7f178da6 · outbound

This paper cites Mimicmotion: High-quality human motion video generation with confidence-aware pose guidance.

Image-to-Video Diffusion: From Foundations to Open Frontiers Mimicmotion: High-quality human motion video generation with confidence-aware pose guidance

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.543705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:67c1aa108be50cb0593a636c4b82ba77b5b73b7b965318017c3b91dce203e0ff

Observation 6cc6ed3a-fb32-4cb0-a98f-0008c91ca2fb · outbound

This paper cites Box- imator: Generating rich and controllable motions for video synthesis.

Image-to-Video Diffusion: From Foundations to Open Frontiers Box- imator: Generating rich and controllable motions for video synthesis

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.540073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:7abe95d20d45f1a48affaeeb0bae2d39d01af430df3cb78401221f7771770c37

Observation e4fa095e-c650-46ac-90d1-81aff07b1039 · outbound

This paper cites LMP: Leveraging Motion Prior in Zero-Shot Video Generation with Diffusion Transformer.

Image-to-Video Diffusion: From Foundations to Open Frontiers LMP: Leveraging Motion Prior in Zero-Shot Video Generation with Diffusion Transformer

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.889954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:d1243be8944ee7267b89eb43d9c890a6e465068355585e72818d460bb3736bc5

Observation 3bbe9811-308a-4866-a77e-fdd21288866f · outbound

This paper cites Movideo: Motion-aware video generation with diffusion model.

Image-to-Video Diffusion: From Foundations to Open Frontiers Movideo: Motion-aware video generation with diffusion model

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.603867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:b26b5fc65b96fdb9c5a04964b1b821aab9d56d602d769a23f149ea802499f22f

Observation 82688681-0220-45a1-b5bc-22b3d483fa7a · outbound

This paper cites Loopy: Taming audio-driven portrait avatar with long-term motion depen- dency.

Image-to-Video Diffusion: From Foundations to Open Frontiers Loopy: Taming audio-driven portrait avatar with long-term motion depen- dency

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.538247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:479624b90e461e9e098d6c15659959adf8ec675377f7ca7dfc2f0e18b27937cc

Observation 08535c7e-7e6c-4f05-835b-ab2fd1ed2a52 · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.871201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:0ddea2e37c82e5f7565d5b14d29771b92e2129d6a06f528bc44dd2778a526e54

Observation ae3e1715-0ae4-4eb5-ae40-1be6464048d9 · outbound

This paper cites Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.973274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:db087653ff35d03cf1919e744a0726719da608eccf43b923a8c54fd6e50c0a88

Observation 5e455c45-d879-4260-b307-8cf59296b479 · outbound

This paper cites Sonic: Shifting focus to global audio perception in portrait animation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Sonic: Shifting focus to global audio perception in portrait animation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.534284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:08b225a536cab7d43383e406b3130ba5d5a5044de921466180c6d19b5b033362

Observation 70c6e5a4-c7bf-442c-af7c-274447ea89e6 · outbound

This paper cites Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions.

Image-to-Video Diffusion: From Foundations to Open Frontiers Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.536232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:b2d85e7f00f0461b26f385b44b2f7235c3eddcc0af70d46e37c4e4a0b20d3305

Observation 3c1edea4-37c8-4276-bb89-6a154abf2069 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

Image-to-Video Diffusion: From Foundations to Open Frontiers AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.949010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:f738937ed5bc2de3f58736f735f733123df64a12848e599d04f842d4a1e037cd

Observation 0a3db86e-b4d9-4a68-af6a-e10d66743f1b · outbound

This paper cites StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.856783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:549fb5a1ff522ec18d24f9216928ca8c7cf1c32e6b4750c5c42f0d9d89f6c5b7

Observation bb53b63d-a6e4-40f9-b4d8-54e4e80c6fdc · outbound

This paper cites OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation.

Image-to-Video Diffusion: From Foundations to Open Frontiers OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:08:24.840385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:ec0a4002e43a312445f314c091bafb125f3687983e71002350bae0c3673a7eda

Observation ac319c22-00b7-4c7e-9cf8-a61afc71f167 · outbound

This paper cites Cyberhost: A one-stage diffusion framework for audio-driven talking body generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Cyberhost: A one-stage diffusion framework for audio-driven talking body generation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.581775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:546502ba71fb358e96250d2344dbb470ea3e905c1819ef33e5d9bdc28890adb5

Observation 2935ab38-1ffd-44af-a3e3-0a45cc479da8 · outbound

This paper cites EMO2: End-Effector Guided Audio-Driven Avatar Video Generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.976159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:a89e021cfd5722afd59bb661c5ee7bfffbea1ab8772377f521312b493d8896d4

Observation 6993e1d7-5a78-4019-a386-b96d18e52023 · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

Image-to-Video Diffusion: From Foundations to Open Frontiers HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.087991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:1726b9b8d4987afe9e5755b9e5a4c676188bca5213e9285c8e4db82102d1370d

Observation ace46b62-6745-40c0-abfa-13ec39c0e658 · outbound

This paper cites CamI2V: Camera-Controlled Image-to-Video Diffusion Model.

Image-to-Video Diffusion: From Foundations to Open Frontiers CamI2V: Camera-Controlled Image-to-Video Diffusion Model

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.844018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:73d0662073f5699c604ac8c7ef79173c55e4deeb8d91227581b7b3ea423c1243

Observation 09e8c998-c82e-473d-940e-c50a6a539b44 · outbound

This paper cites Efficient camera-controlled video generation of static scenes via sparse diffusion and 3d rendering.

Image-to-Video Diffusion: From Foundations to Open Frontiers Efficient camera-controlled video generation of static scenes via sparse diffusion and 3d rendering

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.084980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:396b5b35bc9fac2d2e3a7ed9989bb534259ac61b71acb8aeb1c93771308cb987

Observation 43ab7df5-0d97-429d-a557-753ea580146e · outbound

This paper cites Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models.

Image-to-Video Diffusion: From Foundations to Open Frontiers Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.739350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:c8052fdd91c73f7d11956e3cc0bb36b9ee50b51442627ddd08506765f9dd8779

Observation bedef5e9-ea40-47e0-b6ea-13d164060923 · outbound

This paper cites Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions.

Image-to-Video Diffusion: From Foundations to Open Frontiers Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:08:25.079111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:b82af0476448efd88794a4e6ba11095718dcf2c0b53f8eb169dfceca24c219b4

Observation 32930e22-c51b-4a93-bc8d-f2d3ded64088 · outbound

This paper cites Physgen: Rigid-body physics- grounded image-to-video generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Physgen: Rigid-body physics- grounded image-to-video generation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.744579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:6a430c002d4ba474d66b9ae94d73b93b740fed6d449ce6671b99f09aa13845dc

Observation 8ae14d54-3bf4-47db-a4cf-d6fa3bff7847 · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

Image-to-Video Diffusion: From Foundations to Open Frontiers VACE: All-in-One Video Creation and Editing

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:24.853888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:64b77931d3a295282c17e748b0f677890d7b17bcf2c207c7a803c9286e4505c6

Observation 3072df78-89af-4e22-9db2-3d534492851b · outbound

This paper cites Scalable diffusion models with transformers.

Image-to-Video Diffusion: From Foundations to Open Frontiers Scalable diffusion models with transformers

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.732005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:28e7a47b326c0df8ee9aae874db2f8df2dea21ec0ce1444f1292c8ec3f9f81e8

Observation cca5894e-0512-47b0-b770-44e5e7f80b15 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Image-to-Video Diffusion: From Foundations to Open Frontiers High-resolution image synthesis with latent diffusion models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.737628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:237ec6c680b418dba8c178e4b6daeebb244d03379175bb66feaa19303423c4ed

Observation fff0aeaf-d652-4557-871b-af0fda830e77 · outbound

This paper cites Attention is all you need.

Image-to-Video Diffusion: From Foundations to Open Frontiers Attention is all you need

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.728450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:f452f7c67d34620ed16adda384c9501d41cd8d1f3c7b8bc6516ae0597b419579

Observation 3e795b0f-8a52-481a-bf3d-b2d75be41f3d · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Image-to-Video Diffusion: From Foundations to Open Frontiers Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.726678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:020a22d22d8945e6874141c9861b991cfbd4fe9bbaa1c003721bba122a8185bf

Observation daa326e9-d757-4586-a8d7-2d2ceb0df7dd · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

Image-to-Video Diffusion: From Foundations to Open Frontiers Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.730247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:0b3b98af8da27e54150f529ec3e5dcdd82e60e40c282ce95ba388219dede6d93

Observation 7f1dbc1e-ab94-4e00-a371-652e61343673 · outbound

This paper cites Panda- 70m: Captioning 70m videos with multiple cross-modality teachers.

Image-to-Video Diffusion: From Foundations to Open Frontiers Panda- 70m: Captioning 70m videos with multiple cross-modality teachers

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.733884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:cba96bb0600642d3a2b3412841c6f0a90c362d4c10008667fcc190be415a1d50

Observation 920a0b01-eaf5-4d3e-8f98-09ee574490ef · outbound

This paper cites Magicanimate: Temporally consistent human image animation using diffusion model.

Image-to-Video Diffusion: From Foundations to Open Frontiers Magicanimate: Temporally consistent human image animation using diffusion model

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.746347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:aa5c5bb44962b1327c9acf28d6ff7c92b93d7c54f474c41dd904e09bf1360671

Observation 6d53814f-8387-422a-a285-c58f3ffdc237 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Image-to-Video Diffusion: From Foundations to Open Frontiers Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.722679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:1bfebfae1838e4661cf2aaa6704b380f62f9c0831477a0f0ba2ad75ebd1b6f8a

Observation e786b949-1868-41bf-850e-f709df767b6c · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b-en.

Image-to-Video Diffusion: From Foundations to Open Frontiers Laion coco: 600m synthetic captions from laion2b-en

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.724521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:13a3d09b4ae808d1ca34768bc666198324b7200d79e8f44d484e3747d893c390

Observation 911eacd2-6850-477d-b84d-fb21be659649 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.

Image-to-Video Diffusion: From Foundations to Open Frontiers Journeydb: A benchmark for generative image understanding

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.720877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:1b9f1bc114638de79f4da6c7bffe592fd7730fb74f6fd916ffa2c2b27d1b59e2

Observation 27f8bfe7-0efe-4bc4-8efa-fee60dcb5816 · outbound

This paper cites Make pixels dance: High-dynamic video generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers Make pixels dance: High-dynamic video generation

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T15:08:25.718948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:ce2b5a8a04ceee2c005ea8a13f5adc4a7068b25128bb9acbea24264bc7d61e1f

Observation a9e7a628-5e75-4739-a9e1-224028a44cee · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

Image-to-Video Diffusion: From Foundations to Open Frontiers OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:08:24.867893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:6401456a5910d0c93ddc14d5de2369ca961c68aa48e496a68be5856079513242

Pith citing papers

No inbound Pith citation observations are available.