Pith. sign in

Paper Citation Record · LEDGER

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

As of 21 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 2 inbound Pith citation observations for arXiv:2501.03059.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03059 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:01:57.121114Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:19:17.425943Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:05:51.755490Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a79eec42-9284-4ab7-b538-b06d8183c268 · outbound

This paper cites Stochastic Interpolants: A Unifying Framework for Flows and Diffusions.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Stochastic Interpolants: A Unifying Framework for Flows and Diffusions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.903581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.903581Z digest=sha256:4c408d6b16990875e369d7396afadc178d9713e58251a72b023fb0be7428a616

Observation 53b577ea-56c4-4119-b4a2-dea4fddab751 · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.908137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.908137Z digest=sha256:c86c1802e6f367b4d4697b907d6179ee3e642120d634a09d5bf24d4dad520a9f

Observation 026bdaca-e3e6-419f-bedd-d90ba08bcb12 · outbound

This paper cites Spatext: Spatio-textual representation for con- trollable image generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Spatext: Spatio-textual representation for con- trollable image generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.762151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:56.912145Z digest=sha256:0210c407166528719888dbad15b4e1a163eda9a7f1918df2e5d14a78713e9359

Observation 6545e6c9-05f6-45cc-b5ff-d5e1d6aec3b7 · outbound

This paper cites Improving image generation with better captions.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Improving image generation with better captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.916116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.916116Z digest=sha256:386c600d5726c07591d0664ec8f5329c49321559415cab9524186eed9601dd3f

Observation 1e61b7d5-daca-40de-a087-2c92a7c4e001 · outbound

This paper cites Understanding object dynamics for in- teractive image-to-video synthesis.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Understanding object dynamics for in- teractive image-to-video synthesis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.739754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:56.919447Z digest=sha256:c124afe622cd9ce63f11f80532f753b4ef1f8593ce87287460c96e7d7a123adc

Observation 3068459e-e59b-4a28-883a-7bf4ccd5a9f3 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.922872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.922872Z digest=sha256:2e5b7512abe566ba99ab501c7e63c50cfcb945cea9cd004cf7ec52290ed7d744

Observation bb540953-09a0-4a6d-9f82-963835373585 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.926465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.926465Z digest=sha256:f1b9d851f17c8ed0d3e12657a99aa622171fc311adbb4df72f3ecfd3132de805

Observation b4d8372c-59c9-43c0-89a9-54498e30b8bd · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Instructpix2pix: Learning to follow image editing instructions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.710141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:56.930307Z digest=sha256:c5f056946f765631f37e18caf5a3fe6aba660a359f09f8a07fc34b18e634f69b

Observation d176fb41-f0cb-4dae-bb5b-44aed1de8e5c · outbound

This paper cites Videocrafter1: Open diffusion models for high-quality video generation, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Videocrafter1: Open diffusion models for high-quality video generation, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.696417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:56.933896Z digest=sha256:3f12794254869dda90b7328c4dc3b91e6a64b6f755b824e6aaaeb63cfc2f5c07

Observation d4752850-775e-48dd-b57c-61b559bcec21 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.683744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:56.937556Z digest=sha256:fc042438b07881419ab77b7867d1cc70fc87b1f67d8fb0a4b259d88ed282159f

Observation 59486bbc-3e52-4d5f-8ad1-798f0469ce3d · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.941099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.941099Z digest=sha256:69034da4d87827029f81f9a861a46cca73f14cfe6c2d88e68dc2ecf529ad58fb

Observation 241d01f1-fc55-4112-b15e-f50fd3eb7c06 · outbound

This paper cites Animateanything: Fine- grained open domain image animation with motion guid- ance, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Animateanything: Fine- grained open domain image animation with motion guid- ance, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.667603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:56.944984Z digest=sha256:07426056274e0fbcb9544e5730f117cb84b37a8b8f5b0b40ae5063abad90214c

Observation 339b29cf-d819-4082-a1ad-ba821192fe25 · outbound

This paper cites The llama 3 herd of models, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation The llama 3 herd of models, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.651638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:56.948143Z digest=sha256:5820e7ab5f83c9e1e7e0f0d9431fe5d1821e9f2b2ea23218348be6b4dc18d237

Observation cf8d510e-3057-4111-bc9d-7b91f7cf6eda · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.641005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:56.951413Z digest=sha256:e0e30197c7c356368253ca5f2c26474815cd4f89ffe09f57158324e3e76c1553

Observation 4e319397-44a7-4626-88ff-29f3d1e94af5 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Preserve your own correlation: A noise prior for video diffusion models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.954856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.954856Z digest=sha256:c5a9af7cb163e0bc204099317bc0a684f9a22389189ad727965193068daa1686

Observation f1465803-95ba-4758-900e-5490b1021b54 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.958288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.958288Z digest=sha256:89ff942372982fbba77685bbaf889932e76a25219bde8ce4d252ff6929651fcd

Observation 6f791340-10ec-4842-ab9f-a80fa5685c64 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.623359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:56.961966Z digest=sha256:dfb9b7d7adc11cf016e10bd9f78063d4d88062befa4e4876c5896bf07e6d2d9e

Observation 794d29f6-db4c-4e8d-b549-c34e33193d3d · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.965670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.965670Z digest=sha256:656012df94b1bef7546321351cba5340c58cc0bba9eec4b45cf7952e28c272cb

Observation 59eca75f-cf40-44da-8858-40a590a90645 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Classifier-Free Diffusion Guidance

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.969711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.969711Z digest=sha256:ad5a9ba530472ac3979d8306151e18a24ada4f684823d7a6ece15ac2af61e141

Observation 43a9cf77-7e3b-4d9c-ae18-32865dfe6879 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Denoising dif- fusion probabilistic models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.973238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.973238Z digest=sha256:deee971952610b70134feae48f554a42040d21f6e7439fa02e22e2d24b912e89

Observation 35d94a97-3fac-43a3-b083-25da50e759fe · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.976611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.976611Z digest=sha256:7c46b2fa41e846b751e6e608ea9cb2d2450c386125bf50c9677e65dc4508b862

Observation 6bb44c0d-11ac-497a-b775-f095f605d541 · outbound

This paper cites Video dif- fusion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Video dif- fusion models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.980424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.980424Z digest=sha256:6c83ab0b66378d8352c9e98ca34e288fa5b9ddd57e356ca27081b168fe7c4b73

Observation 8bf0d250-e71a-4f02-91e9-a6dde78bcb7c · outbound

This paper cites Auto-Encoding Variational Bayes.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Auto-Encoding Variational Bayes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.983661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.983661Z digest=sha256:7da692f379a7bfa31c84db09cf056a70eb4074177772f500262e417c3068a54b

Observation 1e2dcf34-dc34-409c-9fd3-2beb71745ce1 · outbound

This paper cites VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.986615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.986615Z digest=sha256:c730cab4a1dd408962a5478c060b4a20341db43d106694a9e7072f8f3e42dfbd

Observation fd5142c5-1752-4337-8487-7e606761e9d2 · outbound

This paper cites Gligen: Open-set grounded text-to-image generation, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Gligen: Open-set grounded text-to-image generation, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.599310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:56.989461Z digest=sha256:43f6bb917a42a5daa4f24dfa0570815c65cf60b8f570ca00eec20759f5fc43b5

Observation ba225d87-1e3a-4e6f-bf43-f36c23d0f811 · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Common diffusion noise schedules and sample steps are flawed, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.587743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:56.992375Z digest=sha256:2da07f23db001609624549798fa0b99df1cc627079d6b6c156c488e22ec9b7ee

Observation b4a4f9be-6999-4bef-8900-5f709a9707fd · outbound

This paper cites an unresolved cited work.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.995450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.995450Z digest=sha256:f4dda4d193b5bc4922fb60b9a8755e42f1ec3781a79075fbd7f0dd700a4c0a57

Observation 465be81b-491b-4b76-a33d-52988a93c56e · outbound

This paper cites Rectified Flow: A Marginal Preserving Approach to Optimal Transport.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Rectified Flow: A Marginal Preserving Approach to Optimal Transport

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.998298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.998298Z digest=sha256:b68dc3b563ad2debb2f67ef1dc6ff3d75c29e4e797d6c89d6a9b1acb61370fe1

Observation 26e1a02e-49a8-4d0f-a258-a86e57f40f0b · outbound

This paper cites Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.001678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.001678Z digest=sha256:01c4c20d8ab81052e2f64585b79b72e49e77ec6c5d0293b9d407150a895bb236

Observation 1c944c83-5a2c-4b2b-ba12-af961d62b9c7 · outbound

This paper cites Cinemo: Consis- tent and controllable image animation with motion diffusion models, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Cinemo: Consis- tent and controllable image animation with motion diffusion models, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.563305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.004917Z digest=sha256:e876bd194d7d07478a4477ec9855a32c16efc5e3da102a9638c1dc432dff47de

Observation 2b2397d9-42a8-460d-a8d7-1570d3ddb895 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Latte: Latent Diffusion Transformer for Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.009271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.009271Z digest=sha256:6bcdf7987af07d9da7646f48c7a7ac1f8fe609d827ce35730f7c0a2471891c8d

Observation 2e2167d8-cb15-4b8d-954a-04158ab401ab · outbound

This paper cites Snap video: Scaled spatiotemporal transformers for text-to-video synthesis.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Snap video: Scaled spatiotemporal transformers for text-to-video synthesis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.552663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.013032Z digest=sha256:bd24f3954ce4e31632259c3782b3ec19b8d341909377d79ffcf172567d53699d

Observation 25e0d8ad-9146-4f52-8e0b-26e12dd0aa7d · outbound

This paper cites Compositional text-to-image gen- eration with dense blob representations, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Compositional text-to-image gen- eration with dense blob representations, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.542735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.016249Z digest=sha256:a41f632b008ec2a3915f70649275455cf24c137e72e0a2ae5f710203b17dadee

Observation bbdd92b7-fdd3-4aad-b416-20b7ce96c201 · outbound

This paper cites Video generation models as world simula- tors.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Video generation models as world simula- tors

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.532320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.019827Z digest=sha256:8babf1aee0863328736d0af73a9f34ef9d09b8e49af547df5101f149c8921a29

Observation 89497550-e4ca-4777-808e-45c596c84c8b · outbound

This paper cites Video generation from sin- gle semantic label map.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Video generation from sin- gle semantic label map

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.521339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.023517Z digest=sha256:d6493aa23b164e5a4f398d8970fecb76300a8c077eda23f338a698ca43a71d35

Observation 87207992-9770-43ec-9487-d23ab702bce9 · outbound

This paper cites Scalable diffusion models with transformers.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Scalable diffusion models with transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.026576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.026576Z digest=sha256:6368fcc6c5f89e44b9722af5758a24726b4f83f6413f8e0997e292c11ec726e4

Observation 4a42d2a5-590c-46f4-beeb-241a526e16e5 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.030059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.030059Z digest=sha256:fb01529079fa1367fe6a4c90d55cce2f2d4c367e0ba530753d31f492513268e2

Observation 7dd62e78-f0a7-4dc4-9b9e-005dbf124fe3 · outbound

This paper cites Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.503944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.033703Z digest=sha256:860f1c48d75f380f3c10c63bdc987d85dffd655afbe62d32c18d7f48d56f8c15

Observation f453eeed-d4f3-4658-be80-4da6ce43f96f · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Learning transferable visual models from natural language supervision, 2021

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.036908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.036908Z digest=sha256:51ed18250259b057536235fd29d2c63e246445c2762c6b49b669c49d4edfb5f5

Observation 20d15cec-fdb4-4863-997d-6ba41dcf599b · outbound

This paper cites Sam 2: Segment anything in images and videos,.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Sam 2: Segment anything in images and videos,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.040095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.040095Z digest=sha256:6bd30929f61cf050d6151e09db9de9c184a9b30e52c2fda955b60e5dd1001efa

Observation 6a0e9035-f195-4696-ba21-d4e43d5c5ee0 · outbound

This paper cites Consisti2v: Enhancing visual consistency for image-to-video generation, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Consisti2v: Enhancing visual consistency for image-to-video generation, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.480654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.043751Z digest=sha256:49a74d2bf95c8857c70bdc28e7311ae2f29dd248983fcf28cb513c840a73b7b8

Observation e05b1915-93e3-4c75-9963-b40e97d5e580 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation High-resolution image synthesis with latent diffusion models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.470475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.047344Z digest=sha256:2b48df3676fbd96c5e202c26f6b3c7204c34814bc4ec7265d70895ec93f8d9b7

Observation 660f2410-fdf5-4973-b83a-6a28e00e162a · outbound

This paper cites Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.460028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.051467Z digest=sha256:7b603ad55d73bf8be8695dc89e32633007fd35150788684774fd2afce4caa3ad

Observation 7c554da8-0fcd-4a0f-b428-c38a9cd680d0 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data, 2022.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Make-a-video: Text-to-video generation without text-video data, 2022

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.450239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.054893Z digest=sha256:82f35cf96451b2ec74128f316c52674ab45e97c175ae4ac30bc58545f5e5c8e1

Observation 556750ab-c2a3-4f8c-91aa-0055e0b78c60 · outbound

This paper cites Denoising Diffusion Implicit Models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Denoising Diffusion Implicit Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.058078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.058078Z digest=sha256:4b11a21b21111bba8b71b51f0a0df0202877820e45a15d526ea5f5959bb8f4c0

Observation f0bf9f6c-004e-4650-9f91-47a4fd3ef298 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Score-Based Generative Modeling through Stochastic Differential Equations

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.061614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.061614Z digest=sha256:5b691d062b5be35d6fa5721f3700a10e25978265d9cd4838177271627c4717b0

Observation fb07f7b0-7508-4810-994a-3049dd0e15da · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow, 2020.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Raft: Recurrent all-pairs field transforms for optical flow, 2020

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.439948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.065341Z digest=sha256:731770dd35f0491bb11a1ad334fb7e44c97cec3fb1af7ea64669af8a01e130b5

Observation 979d90b6-12b3-4aa0-b422-001609b75440 · outbound

This paper cites To- wards accurate generative models of video: A new metric & challenges, 2019.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation To- wards accurate generative models of video: A new metric & challenges, 2019

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.068199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.068199Z digest=sha256:14d8727cb1707ecbe9472c97ed7f176eea0eb4e5e9ccc22740a87962af06a348

Observation 152c2249-a9ce-4d79-8163-89c3773a278d · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation ModelScope Text-to-Video Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.071166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.071166Z digest=sha256:b18c87c94fe47d709a74980374ddb4b1974743db5788c51a4a790f7313bb98f0

Observation 0060c327-9d8f-41d8-adbe-7add5337125d · outbound

This paper cites Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.074328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.074328Z digest=sha256:bf6c8cf80ffd69271083aaafd0aad02458d4f7192fbaf185d8f0084b2b19be53

Observation e504a121-af29-4c76-9b00-cb055b0f80da · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.078986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.078986Z digest=sha256:cba129089443fb6142de02b1f68ae84ff8573d4342ff9a78e974bdea620b7ca5

Observation c74e4b38-fd33-4197-b3c8-0f0724d36a24 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.423665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.082626Z digest=sha256:eb65a8906d696f25bf62a2fc2d837a1447fb50b0e1ccf4bb1f15568440ee8a9e

Observation a6138c38-078e-4e60-9570-b216a2aeeaea · outbound

This paper cites Cvpr 2023 text guided video editing competition, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Cvpr 2023 text guided video editing competition, 2023

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.086032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.086032Z digest=sha256:2f4e194ac2684b8be6401afa1d17578efd3847da1053f0f4def8599db443f861

Observation 3e80e4cc-1dec-4c74-a97f-0de74b0da988 · outbound

This paper cites Dynamicrafter: Animating open-domain im- ages with video diffusion priors, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Dynamicrafter: Animating open-domain im- ages with video diffusion priors, 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.407001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.089488Z digest=sha256:29489512289278fb224397b81b691fcd09d243f336498940b5c7f8177729e3e9

Observation ee53ac5c-e98f-45e2-81dc-a4a741fb77d6 · outbound

This paper cites I2vgen-xl: High-quality image-to-video synthe- sis via cascaded diffusion models, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation I2vgen-xl: High-quality image-to-video synthe- sis via cascaded diffusion models, 2023

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.396106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.092834Z digest=sha256:6d11e19144dd98b603c4a035c85c17b953ee2f67f5944af34250ffff05790df9

Observation f2cdff7d-03cc-484b-98ba-aca83a4d4f89 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.096246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.096246Z digest=sha256:e98e8181a82fc0d6934ec28289aa5f6a1b556928f4cca62c2cb67442e2463dbf

Observation 45863543-3017-4676-844a-782aeaa0ff15 · outbound

This paper cites Qualitative Comparison of Masked Attention Mechanism Fig.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Qualitative Comparison of Masked Attention Mechanism Fig

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.386219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.100257Z digest=sha256:6d1a0cf5e7d8db7a1640608efd8969f46489e21d4073c8531360b02cdc88b0ae

Observation 2c53be84-e420-4376-9e28-fbfd94990343 · outbound

This paper cites 3.1, our pre-processing pipeline ex- tracts a motion-specific prompt, cmotion, from the input text c, using a pre-trained LLM.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation 3.1, our pre-processing pipeline ex- tracts a motion-specific prompt, cmotion, from the input text c, using a pre-trained LLM

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.374870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.103605Z digest=sha256:ccfe34ace182d76a0ec1e1d72720d21ddf04ebb450eaafd6e9d5613e50021835

Observation a1db1218-8197-43f9-ba37-0f12a174fafb · outbound

This paper cites 3.1, the pre-processing process be- gins with extracting motion-capable object prompts from the global prompt c.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation 3.1, the pre-processing process be- gins with extracting motion-capable object prompts from the global prompt c

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.362362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.107151Z digest=sha256:f9dd4dd32684d901a0c9b9bddeaf9474422949edd13dd266d43cf195ea7eeeab

Observation 4ca3f9d3-c478-45b9-890a-0d95840fbab9 · outbound

This paper cites First, the initial segmenta- tion s(0) is extracted from x(0) using SAM2 [40].

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation First, the initial segmenta- tion s(0) is extracted from x(0) using SAM2 [40]

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.351616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.110565Z digest=sha256:392ff1369874208d2614ebda6c7a400990f05365acf1acabd6792af9e8a1f891

Observation 9ed908a2-eb82-4c35-8cdd-cba69f8a898a · outbound

This paper cites The first is the U-Net architecture.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation The first is the U-Net architecture

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.341696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.113763Z digest=sha256:96a16dbfe6eaa52733673d0822307374a374912596281a42aa61a775226945f0

Observation b01e8fac-5167-4c7d-9b03-3d4a8a2b7c53 · outbound

This paper cites The filtering of 128 videos, out of the full SA-V dataset, involved several steps.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation The filtering of 128 videos, out of the full SA-V dataset, involved several steps

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.331486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.117552Z digest=sha256:5c0484628d20e9b67c9c0999743f864b51dbd91758a69e86986a0ebc1b6905a2

Observation 16d604f0-777d-47f1-8ccf-fb2d0fb1ceb4 · outbound

This paper cites description of overall motion.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation description of overall motion

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.318405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:01:57.121114Z digest=sha256:198dd29f06315905ee5f1104c84951536f5d361c3793008dd16aae15ed0d6bff

Pith citing papers

Observation 988a81d2-fcf8-43f8-b760-b76e2c8df26f · inbound

Seeing Voices: Generating A-Roll Video from Audio with Mirage cites this paper.

Seeing Voices: Generating A-Roll Video from Audio with Mirage Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:17.425943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:17.425943Z digest=sha256:7c19d87f129aae539ecf9cacdd55b35dcbb1da034ab8907cfe281f52b7db8029

Observation 22e46e91-26c5-42ff-a16e-a8bcfdf52b99 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

Reference 227

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.761473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:95f9c9665b65937034f7b322eb170725f8ec2ab119fccbc0e9b8b5049007ce19