Pith. sign in

Paper Citation Record · LEDGER

MotiF: Making Text Count in Image Animation with Motion Focal Loss

As of 13 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2412.16153.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16153 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:49:06.719292Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved36
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f8feaa9-cf3b-43e9-a30c-c2facc3da2bd · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.451901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.451901Z digest=sha256:595c58146bf66243e699536ef6cb12315fda19a84f5483d3c0e761d8d854723a

Observation da23cae8-7d2e-4b07-afec-6a12ae46f14f · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.458641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.458641Z digest=sha256:a94f962f12a30f85a42a6f22c5863fd588bd76a52364cc699fd2d811468f2d85

Observation 74fbbf72-d4d4-4804-86e7-bebc7489e8c6 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.463622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.463622Z digest=sha256:feed5df2e3148713651aa369695ebb0635d62454f2a97cba5b324b2154d7de30

Observation b9e03d42-cbe0-4baf-a636-f1f8e651acf7 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.469399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.469399Z digest=sha256:409d2322e92078548a201b39ee0a4e4b03d7f42025360309ee8b97c17b51ba5c

Observation 4bf75c4d-7f4b-4156-a00c-cea98946bcc6 · outbound

This paper cites Video generation models as world simulators.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Video generation models as world simulators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.474602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.474602Z digest=sha256:1f6acb861577086bf25e3436554c3e26216c78da218ecafae75321912e7980a1

Observation bf199bd8-4545-4d3c-8709-440476e7bc59 · outbound

This paper cites Animat- ing general image with large visual motion model.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Animat- ing general image with large visual motion model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.539067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.479805Z digest=sha256:d23c49b379a6e6b9593850283346178373f0d1963d987d06909c45442f05b2f5

Observation 32733eb9-4450-4ecb-b3f0-6df34be1be7f · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.484693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.484693Z digest=sha256:635bac5f7278e3415b698e1c6ddff6d95ebd099b6e6c41a7b67d315d6a8e1a8a

Observation d54ab77c-b092-421e-a38a-7ca45abfbd3e · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.519339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.492435Z digest=sha256:ffee873472e766142286f2c57a1829b9b7468b1c15ee9c5db455fe4b0e9bfa8b

Observation 7bce2c6f-c0a6-4870-b07e-767fb72f33a2 · outbound

This paper cites Seine: Short-to-long video diffu- sion model for generative transition and prediction.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Seine: Short-to-long video diffu- sion model for generative transition and prediction

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.493575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.497289Z digest=sha256:ab1ee8f7719ffd8fb40cc13ece1ad743088fd4dfcc5bd533a3145bfb1e9d975b

Observation 17e117f3-301d-4a44-baff-cb78e7f84eef · outbound

This paper cites Livephoto: Real image animation with text-guided motion control.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Livephoto: Real image animation with text-guided motion control

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.473555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.503127Z digest=sha256:6d88789bd1859af16e83451646d59bbda316eacc268ecb2b7da2896b3076eb2f

Observation f4ed73d5-72b2-471f-8811-223aa048cb00 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.508047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.508047Z digest=sha256:3ea260b7deb07d28eee5a1be56c03ebc7a8135470304423bdf41dfa43cfe258a

Observation 49397f0e-404f-4899-81a9-5da62a0eb0a4 · outbound

This paper cites Animateanything: Fine- grained open domain image animation with motion guid- ance.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Animateanything: Fine- grained open domain image animation with motion guid- ance

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.455046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.513281Z digest=sha256:6ad3d3140ccb9a1b59e1e4661f3f432f07bfe3c6a2a766c4d056580ea4b34227

Observation 0627cba8-ed4e-43c1-bd08-dc47d10d7ccb · outbound

This paper cites AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI.

MotiF: Making Text Count in Image Animation with Motion Focal Loss AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.518012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.518012Z digest=sha256:4ff9e269e7e023716ea0da23761b613d1bb69f2100b61abd6ae65bf6f5ef21ad

Observation 88a32a85-c904-41fc-94e1-e40979e04783 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Preserve your own correlation: A noise prior for video diffusion models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.523213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.523213Z digest=sha256:0734c800ba5c3702c8e0b6984790f87e029ca1725c8fcfa852a1f2629d0ebdac

Observation 586eaac0-db1c-4445-b16a-69e6df17719f · outbound

This paper cites Emu video: Factoriz- ing text-to-video generation by explicit image conditioning.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Emu video: Factoriz- ing text-to-video generation by explicit image conditioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.428994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.528000Z digest=sha256:35c5bd6d79bd2881ddb7ad3fd7a2191778d56fc8b55b21474c2a58eb297fb1b9

Observation 7955499d-8107-4545-9870-6a393211c43b · outbound

This paper cites I2v-adapter: A general image-to-video adapter for diffusion models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss I2v-adapter: A general image-to-video adapter for diffusion models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.532316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.532316Z digest=sha256:1ced31ff0525d517dbea7ccae20bbe754bc7694ee01b96616ac06be21d1f8390

Observation aa24b8f4-5bac-4bf0-8c80-dbabc9fbd52c · outbound

This paper cites Denoising dif- fusion probabilistic models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Denoising dif- fusion probabilistic models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.536820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.536820Z digest=sha256:5200cbb39e7e5d01232625a90000a9a2519b3900b4d9688f889efcffc7b6e375

Observation 560fb056-dcc2-4556-998d-8a1405e52b7e · outbound

This paper cites Video dif- fusion models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Video dif- fusion models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.541111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.541111Z digest=sha256:14d43cc1dc84456e64316c7f8dfaebb3654263110c819b949048f23100ee5303

Observation d04e818a-69d6-49fa-9824-370b937f0d58 · outbound

This paper cites Make it move: Controllable image-to-video generation with text descrip- tions.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Make it move: Controllable image-to-video generation with text descrip- tions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.375401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.545065Z digest=sha256:80c038e9d974a8287c98774728d605835b2dd772a1b1ab428e300a0d87408b68

Observation 12675c13-be6a-4cb3-bc47-a716c29613bf · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss VBench: Com- prehensive benchmark suite for video generative models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.358045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.549104Z digest=sha256:f11cc3abedc02bfe6c8a1368482a745817b75b70a4b6a9c0540657274808bcd3

Observation 198c63aa-93f9-469f-a78b-ad352445e59b · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Vbench: Comprehensive bench- mark suite for video generative models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.341698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.553667Z digest=sha256:ed35c974137c561ef632d041c6cedc4b81e0539303d65611a5e2c49ef45b9f31

Observation c1e73ecd-7829-48da-bb5b-4573dde9b8d0 · outbound

This paper cites VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.558534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.558534Z digest=sha256:89616f2f1db269de5f0e8ecbab058fa08bac2f38353fae505b8472f49db010c1

Observation 343e7d14-1382-4eb7-a5b8-5cf077a2a8b9 · outbound

This paper cites Physgen: Rigid-body physics-grounded image- to-video generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Physgen: Rigid-body physics-grounded image- to-video generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.324414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.563595Z digest=sha256:817b2d4eb01594ce312f7ec2da472c9ee7f0bf1ec55bce42eea059bee20da36e

Observation ffc46977-a43e-49ae-a806-a95416d3611a · outbound

This paper cites Evalcrafter: Benchmarking and eval- uating large video generation models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Evalcrafter: Benchmarking and eval- uating large video generation models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.567962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.567962Z digest=sha256:c6ffb8d50d0e7ce75486d2e482a384c8b7e50bba673be9974f431b8134d9205c

Observation a7fca664-894b-4a25-8566-d6cff4a41ead · outbound

This paper cites Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.573171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.573171Z digest=sha256:41115e5497e655557293c45416693d3e77c412972eed7708d2992652bead1ed4

Observation 75213b1f-7a3a-41b8-b899-4ea7df769d5f · outbound

This paper cites Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.578398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.578398Z digest=sha256:7027b6138d588dbcb6c54502b3a9145fc9b5699f6ad2fb5523d891886a7ab896

Observation 5c4a0bb4-37e9-4cd9-b919-a87fdd65bcee · outbound

This paper cites Sync-draw: Automatic video generation using deep recurrent attentive architectures.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Sync-draw: Automatic video generation using deep recurrent attentive architectures

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.295710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.583317Z digest=sha256:cbe84295829003e63ada1a2be0af54cc87bd70cac8ffdda23921a500ce4f4e08

Observation ce76bf80-9736-453e-bc6a-cf06adfcfbb1 · outbound

This paper cites Ti2v-zero: Zero-shot image condition- ing for text-to-video diffusion models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Ti2v-zero: Zero-shot image condition- ing for text-to-video diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.277813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.588321Z digest=sha256:10548a5abe80bd833534025c01dfc34f451548af338d01394e7d1cec6006fd33

Observation 1ee9b71f-6541-4f7d-b6e0-ea914530cfaf · outbound

This paper cites Scalable diffusion models with transformers.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Scalable diffusion models with transformers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.593227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.593227Z digest=sha256:2f2fe5a60544071115b129eae4d246aa403d13df6f91784f3edb7dff3c3d887c

Observation 3b34f342-e379-4337-8539-e45aa07638f1 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Movie Gen: A Cast of Media Foundation Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.598188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.598188Z digest=sha256:a9c1b53c0c546aa70536acaaca0dec463a7c7d670709d9abfb8381bdcd42fcfc

Observation 0c3b48a3-4cec-4060-935e-db410c87f3c0 · outbound

This paper cites Hier- archical spatio-temporal decoupling for text-to-video gener- ation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Hier- archical spatio-temporal decoupling for text-to-video gener- ation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.603898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.603898Z digest=sha256:32382593b7d4c8f65ca8d2ac74e19719d3c5885fb71f46ce79edd42aa560f667

Observation 456fba28-d5e7-40d7-a7c2-35b61bdecc12 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

MotiF: Making Text Count in Image Animation with Motion Focal Loss SAM 2: Segment Anything in Images and Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.609625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.609625Z digest=sha256:b6e759aefc3d5473ed9ae10cd3cb9ed9b1e5da23ad776d2660c081579f90645b

Observation 1851ffcd-1ad3-4806-8e0e-b75bdf7dc961 · outbound

This paper cites ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.615135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.615135Z digest=sha256:8af660422d2b8d723c42118e5ab9a749407f3b380926a0213d445867a559cf63

Observation c855f6b7-85a3-4d4b-944e-4428cd9f4637 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss High-resolution image synthesis with latent diffusion models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.620722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.620722Z digest=sha256:f678beb76631b1c8c76c287b65faafcda3c09275e4633bd36a1894eced407e1a

Observation 95ab1552-a16b-4c80-b701-d3549cbb9a92 · outbound

This paper cites Focal loss for dense ob- ject detection.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Focal loss for dense ob- ject detection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.625570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.625570Z digest=sha256:f8aa74bba0bcea04e07f2603538932b5c1e62e6f08e9b0f1dd58fcaee429e63d

Observation fb6607c7-825f-4d07-9c73-fb26afa6bdc2 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Progressive Distillation for Fast Sampling of Diffusion Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.630409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.630409Z digest=sha256:df53e66bf82f700d26bc02251c23634eecafb06e98b44cb706c338b65e4bc495

Observation d27b60d5-83a2-431d-9a44-6b2ee793c7d2 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.635388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.635388Z digest=sha256:a1014a7649182dfb978846f2277b8afe83cbd6d9694d1a9fac99d86f46d88f9e

Observation da056e74-be47-4220-b25d-961c3f577954 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Deep unsupervised learning using nonequilibrium thermodynamics

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.640425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.640425Z digest=sha256:fccc4d57408ff5af09d5f64476a4236bfbf245d3f6594bbf8e183be780325f84

Observation c621ec9d-17c4-485e-9852-666146dbef3d · outbound

This paper cites Denois- ing diffusion implicit models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Denois- ing diffusion implicit models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.645239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.645239Z digest=sha256:77df645c931f5e4b9d88c971125828384de49d8747ff0a833a8d08927351bbcd

Observation 5466c700-6529-41f7-b510-725dfd94d289 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Score-Based Generative Modeling through Stochastic Differential Equations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.649998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.649998Z digest=sha256:32396701d87dfbeac1ed16096b68ab8fd81090d19d1fb70c36250b38a9121ec6

Observation 4618c896-8899-46b6-a223-673af3e2432e · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MotiF: Making Text Count in Image Animation with Motion Focal Loss UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.654412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.654412Z digest=sha256:bb2ce43ce57caa1063082b2ebf53be362f5bccdf4cb0465ae0d4bb2c294c6a23

Observation 9b8b7f74-b3ab-44d4-bd90-2016e500b1f2 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Raft: Recurrent all-pairs field transforms for optical flow

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.659181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.659181Z digest=sha256:8fb834c9d96a059f3a682e534a6503a5f54359f7af9407557909c8aacda4af2d

Observation 78a763d8-6539-4f4b-a808-3efc8e5682ad · outbound

This paper cites Microcinema: A divide-and- conquer approach for text-to-video generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Microcinema: A divide-and- conquer approach for text-to-video generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.187114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.663915Z digest=sha256:2495d9fa21dacd266224598f004de726cbcc3f00e2a02a2eacee5a438b117389

Observation 44c9d339-6e15-4eaf-b087-3df805681924 · outbound

This paper cites Dreamvideo: Composing your dream videos with customized subject and motion.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Dreamvideo: Composing your dream videos with customized subject and motion

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.668128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.668128Z digest=sha256:ecc0afe8e2e83edcab16e0852e20212643517b68a39173fc22361fb66c8e7845

Observation 3f8302a7-dfa1-4c62-8f72-61e94fd9df06 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.672568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.672568Z digest=sha256:a1c5b5bc2b660afb7ccda8c74e5168e60a24bf1977ee8bbaeb803a54a6798a19

Observation 0ceb8851-790f-4283-8b4a-c5c16195bb4b · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.149060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.677016Z digest=sha256:17ccab1d6d5d2f8d2f88e4daa35f4d60a328e6e90af2211d3385ccf7d0090d46

Observation 07a1458d-3778-4676-8bc7-48a26c7abdd2 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Msr-vtt: A large video description dataset for bridging video and language

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.131692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.681700Z digest=sha256:a90d35d2b38a19a60d174209d718f7ffd625693d22aedfc6020c7cb59ecb7c52

Observation 5cb96f43-0e99-4955-a58a-98b8842049d8 · outbound

This paper cites Motion-Conditioned Image Animation for Video Editing.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Motion-Conditioned Image Animation for Video Editing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.686047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.686047Z digest=sha256:2f1f5fbe3b5e53fc37dee1da35bf850060dd25fc95552ae737d07bd930da6ad6

Observation f3641f0b-fa18-4d10-9283-0702a1bc8f2c · outbound

This paper cites Zero-shot controllable image-to-video animation via motion decomposition.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Zero-shot controllable image-to-video animation via motion decomposition

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.113704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.691065Z digest=sha256:d2ed4e6ee71eca02ef14efd66e2f8bd2eb3821eac2c6197210fc08aac55313a5

Observation f28483e1-c41b-4d32-a134-fc6d89f71144 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.696189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.696189Z digest=sha256:db9b399ae5d3d6bfe02f2be646aa9a9cd62674b4e918da2734b944801949f0dd

Observation 168f8b80-745a-4b15-995b-140bdf2a3a62 · outbound

This paper cites Pia: Your personalized image animator via plug-and-play modules in text-to-image models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Pia: Your personalized image animator via plug-and-play modules in text-to-image models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.095893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T10:49:06.701883Z digest=sha256:ef793e9910abbf465b12c82f754633525ba2eb1230faf4955d16427304ad9547

Observation 0f481d93-ab49-4145-9808-ea9fe352c893 · outbound

This paper cites Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified Model.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.709009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.709009Z digest=sha256:fb0f3dfe412d472e1ce5c3454165883fcfb70dd7296c22fbf9421e9ca6723009

Observation e25f6935-7bbd-4cc5-abc3-6af1b4c76252 · outbound

This paper cites Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.714131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.714131Z digest=sha256:8562a64dd1ae9026b2fcef5b5f7e5669ab64958279a5d709673b43d3cacb377d

Observation 2672eec0-28a5-49b0-943f-c08c0e4192e4 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 54

Resolution
malformed identifier
no resolver link, observed 2026-08-11T10:49:06.719292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.719292Z digest=sha256:e1dddb2e70385b03f4adbe4c74849b05e0a1a28da48bf125d04cb244551a8e90

Pith citing papers

No inbound Pith citation observations are available.