Pith. sign in

Paper Citation Record · LEDGER

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

As of 16 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 2 inbound Pith citation observations for arXiv:2501.03059.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03059 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:01:57.121114Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:19:17.425943Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:05:51.755490Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a79eec42-9284-4ab7-b538-b06d8183c268 · outbound

This paper cites Stochastic Interpolants: A Unifying Framework for Flows and Diffusions.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Stochastic Interpolants: A Unifying Framework for Flows and Diffusions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.903581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.903581Z digest=sha256:32598b6eedd1df98e52cc7eeed0584b70c4986a47eec6b2300ad67930fac3c85

Observation 53b577ea-56c4-4119-b4a2-dea4fddab751 · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.908137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.908137Z digest=sha256:f4a18cf62b80cf6b15c23b763fa005754ce5013eb118a2b82737a9de5ad587f7

Observation 026bdaca-e3e6-419f-bedd-d90ba08bcb12 · outbound

This paper cites Spatext: Spatio-textual representation for con- trollable image generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Spatext: Spatio-textual representation for con- trollable image generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.762151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.912145Z digest=sha256:d18eb2797973801e687fe6c6fb329b754017d3492e0856cd71c0f31179286e48

Observation 6545e6c9-05f6-45cc-b5ff-d5e1d6aec3b7 · outbound

This paper cites Improving image generation with better captions.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Improving image generation with better captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.916116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.916116Z digest=sha256:01aca5da302d670681d8ec271325b072307b37cb165acdd61398a205497f6e2f

Observation 1e61b7d5-daca-40de-a087-2c92a7c4e001 · outbound

This paper cites Understanding object dynamics for in- teractive image-to-video synthesis.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Understanding object dynamics for in- teractive image-to-video synthesis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.739754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.919447Z digest=sha256:a6d3e0a04b0c54841a76d19f132675b9f1dba6e12dbcd995da7ecffc468d1f15

Observation 3068459e-e59b-4a28-883a-7bf4ccd5a9f3 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.922872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.922872Z digest=sha256:485169f14e6dbb6ad580b29d07742f0f990047773c026ebccef34707668df578

Observation bb540953-09a0-4a6d-9f82-963835373585 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.926465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.926465Z digest=sha256:b82208983b52cb1bc284ea868a4c37a786be0b91c9e27228749b3bb0f9c550cd

Observation b4d8372c-59c9-43c0-89a9-54498e30b8bd · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Instructpix2pix: Learning to follow image editing instructions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.710141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.930307Z digest=sha256:77b95c899a5fb0f2b58e94aef1ed17078e1f1d66a3fcb705368c0b18b4159ebd

Observation d176fb41-f0cb-4dae-bb5b-44aed1de8e5c · outbound

This paper cites Videocrafter1: Open diffusion models for high-quality video generation, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Videocrafter1: Open diffusion models for high-quality video generation, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.696417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.933896Z digest=sha256:16a106147aea10fe6eec18ebb9ceae57e60567eacad914a284ce1f59983d8cff

Observation d4752850-775e-48dd-b57c-61b559bcec21 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.683744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.937556Z digest=sha256:4e83e1337a472f0c918bbe10b9e4ff23a0e6901079b7f5fb25abe6f348eed357

Observation 59486bbc-3e52-4d5f-8ad1-798f0469ce3d · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.941099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.941099Z digest=sha256:fff2fa591e0b77e95522acfea8c859e22af6348eaa69bb193e5cdb26cedf16a5

Observation 241d01f1-fc55-4112-b15e-f50fd3eb7c06 · outbound

This paper cites Animateanything: Fine- grained open domain image animation with motion guid- ance, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Animateanything: Fine- grained open domain image animation with motion guid- ance, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.667603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.944984Z digest=sha256:62dc5e20b1e0d201466eefcdfff9b7751ae3c2c9753dad7566ce4b4fc3468481

Observation 339b29cf-d819-4082-a1ad-ba821192fe25 · outbound

This paper cites The llama 3 herd of models, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation The llama 3 herd of models, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.651638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.948143Z digest=sha256:c637930f379a5417b97acabf496b8529490c640ec5eeaef8c71dc29579967aa6

Observation cf8d510e-3057-4111-bc9d-7b91f7cf6eda · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.641005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.951413Z digest=sha256:8577005b043600fa3afbc2867e4aee43c0cbfd1fed9c716c32e2a90006f9f36e

Observation 4e319397-44a7-4626-88ff-29f3d1e94af5 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Preserve your own correlation: A noise prior for video diffusion models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.954856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.954856Z digest=sha256:33edf080a22d0734caba1d8e26f6146b21864b9c2d23d558b5462432215ff8e7

Observation f1465803-95ba-4758-900e-5490b1021b54 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.958288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.958288Z digest=sha256:e5e0be58309c97bff1245a162a5ad15349df60df6905f40f794bf36295dc5f13

Observation 6f791340-10ec-4842-ab9f-a80fa5685c64 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.623359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.961966Z digest=sha256:fca5506571bc108f0bf8ccefcc958128fa088fbc17934ac0fe88d3a0e08842c1

Observation 794d29f6-db4c-4e8d-b549-c34e33193d3d · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.965670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.965670Z digest=sha256:6ae557b96f98c344cd2fb70cefcd69b1eb9036643116e9a2c3957e79d8bfae41

Observation 59eca75f-cf40-44da-8858-40a590a90645 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Classifier-Free Diffusion Guidance

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.969711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.969711Z digest=sha256:579cecdad10417561e06649bd29b91dc29fe69f77cb33c11c6aeaf9db3b0020f

Observation 43a9cf77-7e3b-4d9c-ae18-32865dfe6879 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Denoising dif- fusion probabilistic models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.973238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.973238Z digest=sha256:5bc27b835ca773464ca17f7c621308e2d23f6efcd293cbb3756fdf468113484f

Observation 35d94a97-3fac-43a3-b083-25da50e759fe · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.976611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.976611Z digest=sha256:566e8de1fbe69e95ed45ac5dae9a83c12c09417206e65dd8ee82486fc8969278

Observation 6bb44c0d-11ac-497a-b775-f095f605d541 · outbound

This paper cites Video dif- fusion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Video dif- fusion models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.980424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.980424Z digest=sha256:10c315e4b97ab71cba7845b6dbecde1086f451b851c6c28f64ea90f6696cbda6

Observation 8bf0d250-e71a-4f02-91e9-a6dde78bcb7c · outbound

This paper cites Auto-Encoding Variational Bayes.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Auto-Encoding Variational Bayes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.983661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.983661Z digest=sha256:60065798b94ef845b9ae629e4780258ce02bf6e1da728bbe55286698ec4c1c20

Observation 1e2dcf34-dc34-409c-9fd3-2beb71745ce1 · outbound

This paper cites VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.986615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.986615Z digest=sha256:cc87c93f728068fba5c9472530c29c78605e0973053d38182ff1ef453f89819c

Observation fd5142c5-1752-4337-8487-7e606761e9d2 · outbound

This paper cites Gligen: Open-set grounded text-to-image generation, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Gligen: Open-set grounded text-to-image generation, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.599310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.989461Z digest=sha256:cdf9424d8a5e6bbbaa58e9b2e055d061fe4e0ebb71a7d7fdd5157646c98bcd3e

Observation ba225d87-1e3a-4e6f-bf43-f36c23d0f811 · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Common diffusion noise schedules and sample steps are flawed, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.587743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.992375Z digest=sha256:a580702cb3304253c08139cdbe73a89189e271d8a2668f9468273b9bb261ef27

Observation b4a4f9be-6999-4bef-8900-5f709a9707fd · outbound

This paper cites an unresolved cited work.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.995450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.995450Z digest=sha256:0dcf2161ffb243e7d78f6a385afed6927da4579d2088547c068208024365b68e

Observation 465be81b-491b-4b76-a33d-52988a93c56e · outbound

This paper cites Rectified Flow: A Marginal Preserving Approach to Optimal Transport.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Rectified Flow: A Marginal Preserving Approach to Optimal Transport

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.998298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.998298Z digest=sha256:ffaec88bb3aa0619d46be5178062bee0c385b3b45e2fde41b7f324313bfa83ec

Observation 26e1a02e-49a8-4d0f-a258-a86e57f40f0b · outbound

This paper cites Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.001678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.001678Z digest=sha256:19c0936086a36c21fd8e22c3833d64279ffb2d483cd57441fb5f1981fdcac016

Observation 1c944c83-5a2c-4b2b-ba12-af961d62b9c7 · outbound

This paper cites Cinemo: Consis- tent and controllable image animation with motion diffusion models, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Cinemo: Consis- tent and controllable image animation with motion diffusion models, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.563305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.004917Z digest=sha256:79627529cc3baac37cc1ab7f0412caeeedae4fa842c9dacfc723516f6308c192

Observation 2b2397d9-42a8-460d-a8d7-1570d3ddb895 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Latte: Latent Diffusion Transformer for Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.009271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.009271Z digest=sha256:4c9d3cefd07eefbc2856410d1650e4c5e53280b779954ed9b2b1785a9d9bfc13

Observation 2e2167d8-cb15-4b8d-954a-04158ab401ab · outbound

This paper cites Snap video: Scaled spatiotemporal transformers for text-to-video synthesis.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Snap video: Scaled spatiotemporal transformers for text-to-video synthesis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.552663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.013032Z digest=sha256:3dcf89c0720add2ca3b1aa5bb1e1f9f1d7c5e8283304aee33c26cd36d120f1bb

Observation 25e0d8ad-9146-4f52-8e0b-26e12dd0aa7d · outbound

This paper cites Compositional text-to-image gen- eration with dense blob representations, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Compositional text-to-image gen- eration with dense blob representations, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.542735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.016249Z digest=sha256:5bec2671ba9980666d47b6105f8b4813e84a1462cfe5bb5ecff5234b9379390b

Observation bbdd92b7-fdd3-4aad-b416-20b7ce96c201 · outbound

This paper cites Video generation models as world simula- tors.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Video generation models as world simula- tors

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.532320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.019827Z digest=sha256:e48c2c701c04cd8fddca4441970d5cadd160f332d342aeef7106456ff10eecb0

Observation 89497550-e4ca-4777-808e-45c596c84c8b · outbound

This paper cites Video generation from sin- gle semantic label map.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Video generation from sin- gle semantic label map

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.521339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.023517Z digest=sha256:72404febd9b23107886a48dd2310b8e89690dedd51319cb1a44a74728f83dfed

Observation 87207992-9770-43ec-9487-d23ab702bce9 · outbound

This paper cites Scalable diffusion models with transformers.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Scalable diffusion models with transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.026576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.026576Z digest=sha256:4770257d051a8ec8edcfe1299ecf7da90aea67e7d5cf5f14ef30a8b23d088cfd

Observation 4a42d2a5-590c-46f4-beeb-241a526e16e5 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.030059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.030059Z digest=sha256:b77e39d25e42e258ddcba02606e23c3efc8e2fdd728295aa1fd2c3f60832dfd9

Observation 7dd62e78-f0a7-4dc4-9b9e-005dbf124fe3 · outbound

This paper cites Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.503944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.033703Z digest=sha256:f7d57b1686898a62658b4c81d92ace1e27db07a51662cb4bad8e1320e437499d

Observation f453eeed-d4f3-4658-be80-4da6ce43f96f · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Learning transferable visual models from natural language supervision, 2021

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.036908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.036908Z digest=sha256:93371e3f58e47a764e045277944e0183546c670e540087c34ecfac9729396c76

Observation 20d15cec-fdb4-4863-997d-6ba41dcf599b · outbound

This paper cites Sam 2: Segment anything in images and videos,.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Sam 2: Segment anything in images and videos,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.040095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.040095Z digest=sha256:90ed87ca73a5739c692193c782fb1e30072a4099ee9f366c5c81cfa6d0fda3f0

Observation 6a0e9035-f195-4696-ba21-d4e43d5c5ee0 · outbound

This paper cites Consisti2v: Enhancing visual consistency for image-to-video generation, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Consisti2v: Enhancing visual consistency for image-to-video generation, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.480654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.043751Z digest=sha256:1f8cf62698ee4620a6393c9d0661d8a4f9e262b0764e3195f0c5dc79d637fe72

Observation e05b1915-93e3-4c75-9963-b40e97d5e580 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation High-resolution image synthesis with latent diffusion models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.470475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.047344Z digest=sha256:02bae9b052333f3f4f206786302069ac6098c9d9680860c1188a238e503a5f84

Observation 660f2410-fdf5-4973-b83a-6a28e00e162a · outbound

This paper cites Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.460028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.051467Z digest=sha256:56c74d82414bfabbc1fc9f966afe967f58ad88c99c9678425f5308e61409a87e

Observation 7c554da8-0fcd-4a0f-b428-c38a9cd680d0 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data, 2022.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Make-a-video: Text-to-video generation without text-video data, 2022

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.450239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.054893Z digest=sha256:18c4cb1021593ad7a1f2d2654683caca48619c26d519c53a2598632efcdd9963

Observation 556750ab-c2a3-4f8c-91aa-0055e0b78c60 · outbound

This paper cites Denoising Diffusion Implicit Models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Denoising Diffusion Implicit Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.058078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.058078Z digest=sha256:5c6bb630805e397cbdec78aad088a48c227a7fac8d5ae25c7a46a4753e0a88a1

Observation f0bf9f6c-004e-4650-9f91-47a4fd3ef298 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Score-Based Generative Modeling through Stochastic Differential Equations

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.061614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.061614Z digest=sha256:bb08a009a27d1e323d8472cfe9ce507de6ba759e978a775c021d8f8b9e4ca4cd

Observation fb07f7b0-7508-4810-994a-3049dd0e15da · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow, 2020.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Raft: Recurrent all-pairs field transforms for optical flow, 2020

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.439948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.065341Z digest=sha256:fc6f6c811c9652350f7c5e20e77fbde109fabdfffb1c9d9e9673211a4ec6510a

Observation 979d90b6-12b3-4aa0-b422-001609b75440 · outbound

This paper cites To- wards accurate generative models of video: A new metric & challenges, 2019.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation To- wards accurate generative models of video: A new metric & challenges, 2019

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.068199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.068199Z digest=sha256:633075f8270677216f889572d79de7d1ad736a7dfa08b02424d1dc60ed18d62d

Observation 152c2249-a9ce-4d79-8163-89c3773a278d · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation ModelScope Text-to-Video Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.071166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.071166Z digest=sha256:68b61d9def3841aad9c76fcc8f47a23ff8029d22db658cc2da6fadfdcad2984f

Observation 0060c327-9d8f-41d8-adbe-7add5337125d · outbound

This paper cites Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.074328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.074328Z digest=sha256:52fda7b117a866af171bd41b122fa687c325ada443a2bd0b0262e07e9b1251ac

Observation e504a121-af29-4c76-9b00-cb055b0f80da · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.078986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.078986Z digest=sha256:d30903016eb57d5a673e9ff1443ee376f4c4093cabe9b1830e4fd9767625fc25

Observation c74e4b38-fd33-4197-b3c8-0f0724d36a24 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.423665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.082626Z digest=sha256:0e1e779a0aff998855ad92897006c70c68a3929dd7ba54f048006bf82f582202

Observation a6138c38-078e-4e60-9570-b216a2aeeaea · outbound

This paper cites Cvpr 2023 text guided video editing competition, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Cvpr 2023 text guided video editing competition, 2023

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.086032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.086032Z digest=sha256:9ac7f06d776fce7b6c285c7df14c0f198806c2b66e6e392c244d6b3fa4fa1ad7

Observation 3e80e4cc-1dec-4c74-a97f-0de74b0da988 · outbound

This paper cites Dynamicrafter: Animating open-domain im- ages with video diffusion priors, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Dynamicrafter: Animating open-domain im- ages with video diffusion priors, 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.407001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.089488Z digest=sha256:54844895149ff511e5f84fcaa63ec02c3e85cb49e1f578b2392a56ff6848510f

Observation ee53ac5c-e98f-45e2-81dc-a4a741fb77d6 · outbound

This paper cites I2vgen-xl: High-quality image-to-video synthe- sis via cascaded diffusion models, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation I2vgen-xl: High-quality image-to-video synthe- sis via cascaded diffusion models, 2023

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.396106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.092834Z digest=sha256:c656784f9cfcf086ee6301e506f74d29688f1292e14549a62d6a2f4cd070b80c

Observation f2cdff7d-03cc-484b-98ba-aca83a4d4f89 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.096246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.096246Z digest=sha256:d9e9d3b6fd56208ed1ece6766d82214283ff713e9a9081a36e9cec943abbbf59

Observation 45863543-3017-4676-844a-782aeaa0ff15 · outbound

This paper cites Qualitative Comparison of Masked Attention Mechanism Fig.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Qualitative Comparison of Masked Attention Mechanism Fig

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.386219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.100257Z digest=sha256:199f8a2a0b1170128a633aadec10a9179f03dc24aa688faa7f7a9604366ed3a6

Observation 2c53be84-e420-4376-9e28-fbfd94990343 · outbound

This paper cites 3.1, our pre-processing pipeline ex- tracts a motion-specific prompt, cmotion, from the input text c, using a pre-trained LLM.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation 3.1, our pre-processing pipeline ex- tracts a motion-specific prompt, cmotion, from the input text c, using a pre-trained LLM

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.374870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.103605Z digest=sha256:d2dbfff98a104226ef79688fabff371a70ab8b9875134e113e7827a009d52a61

Observation a1db1218-8197-43f9-ba37-0f12a174fafb · outbound

This paper cites 3.1, the pre-processing process be- gins with extracting motion-capable object prompts from the global prompt c.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation 3.1, the pre-processing process be- gins with extracting motion-capable object prompts from the global prompt c

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.362362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.107151Z digest=sha256:1f3bc2ddf403d7f65c613479c852aa1d98457a9246a8fe479ddc2ffdf6006740

Observation 4ca3f9d3-c478-45b9-890a-0d95840fbab9 · outbound

This paper cites First, the initial segmenta- tion s(0) is extracted from x(0) using SAM2 [40].

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation First, the initial segmenta- tion s(0) is extracted from x(0) using SAM2 [40]

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.351616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.110565Z digest=sha256:a187c94279b8aaa35d1307ee971d335550d20ef87c59e61a93f9d7c38e641ce9

Observation 9ed908a2-eb82-4c35-8cdd-cba69f8a898a · outbound

This paper cites The first is the U-Net architecture.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation The first is the U-Net architecture

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.341696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.113763Z digest=sha256:bea0a9fe57c08c6eacf9348a486fdc31e8e258a7237051de50b46e826ddc2f53

Observation b01e8fac-5167-4c7d-9b03-3d4a8a2b7c53 · outbound

This paper cites The filtering of 128 videos, out of the full SA-V dataset, involved several steps.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation The filtering of 128 videos, out of the full SA-V dataset, involved several steps

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.331486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.117552Z digest=sha256:6807a3ee7012363ab06cbead30cee5f791ace7367d243b6be9aaf3f2cdec238e

Observation 16d604f0-777d-47f1-8ccf-fb2d0fb1ceb4 · outbound

This paper cites description of overall motion.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation description of overall motion

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.318405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.121114Z digest=sha256:cf029129b9036378c75c2d54b61586ea9ad9b6339bd1247ed0e40aebae5b10b8

Pith citing papers

Observation 988a81d2-fcf8-43f8-b760-b76e2c8df26f · inbound

Seeing Voices: Generating A-Roll Video from Audio with Mirage cites this paper.

Seeing Voices: Generating A-Roll Video from Audio with Mirage Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:17.425943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:17.425943Z digest=sha256:85e11d3a6cc3ef38040dc8a3d0bad281290230f0653d5e9393fef043dc7f4d55

Observation 22e46e91-26c5-42ff-a16e-a8bcfdf52b99 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

Reference 227

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.761473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:8f5da59231202950e81e1e9b4861bd4fccd5e4580d98066341f656400bdc0b7e