Pith. sign in

Paper Citation Record · LEDGER

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

As of 13 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 30 inbound Pith citation observations for arXiv:2502.02492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02492 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T12:04:47.390095Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:54:12.520467Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.389947Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb2fc609-8a88-4ff2-8ce5-9f40dc1f3ad7 · outbound

This paper cites write newline.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.955949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.955949Z digest=sha256:d1e45c0516cb276a636619557886faea5c9946f1b771a3324398aa65d7dd784a

Observation 3b54b379-bdad-499d-bb19-005e50082bb9 · outbound

This paper cites write newline.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.965726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.965726Z digest=sha256:07d70ac689fd16f58eb39f20556d1657b40430d19b20d674f19995d06945a748

Observation b3fd3627-3c77-41da-982a-ef8173818147 · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.973669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.973669Z digest=sha256:aeaecc14949a7cb3a1307d6e4feaca4d7a579e2aeb015680864afc968a0330d7

Observation 20ae7674-3293-4492-952c-a667b5745596 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.982926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.982926Z digest=sha256:058815e39e6740815b7da5d30c44d6f3117b175178dea6bedf691fb25a653c4b

Observation 6fc56778-e1f4-4097-9f39-c269b6c784d0 · outbound

This paper cites FLUX , 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models FLUX , 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.731711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:46.990934Z digest=sha256:4240f070678dbdc0eac2ef936b3af93a12512e64ff886560368331f702c7eadb

Observation 0e6ff5d8-b059-49b1-841d-617ded0a920d · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.998289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.998289Z digest=sha256:1092bf5ea5b79d7301db6908fe6cd7ab21ef7e7a7332e79336d826956bfc2c25

Observation 1e97bc72-4b71-4d3e-8980-5c0f012ab992 · outbound

This paper cites W., Fidler, S., and Kreis, K.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models W., Fidler, S., and Kreis, K

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.705637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.005852Z digest=sha256:5e87b529914771c738b3a574a54ea7bc76ef0c84a44028f3720a586ca9dd0156

Observation e9bce902-5ef0-48d9-a70a-da36699859a4 · outbound

This paper cites an unresolved cited work.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:04:48.685580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.012846Z digest=sha256:b6268ad417f613358901bebfb447a858c54bfde1c7fdc155bdb3e20b80c2898a

Observation 00c48ed4-15f6-4282-a5d7-53efe0339952 · outbound

This paper cites Video generation models as world simulators.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Video generation models as world simulators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.020437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.020437Z digest=sha256:d4e9a34ee341279815c559f8f10ba0d9d1f68b9d1239d6da18d76300446f5c81

Observation b11d7e93-109e-4b49-823f-d1063822c9a0 · outbound

This paper cites The hidden language of diffusion models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models The hidden language of diffusion models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.652827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.031598Z digest=sha256:e1e21c082ab4778194413d4d2d9395cb6c4c6db94524903a62f3a29570ce7779

Observation c8df70af-e99a-4a45-b005-41b3576f1b01 · outbound

This paper cites Still-Moving: Customized Video Generation without Customized Video Data.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Still-Moving: Customized Video Generation without Customized Video Data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.038836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.038836Z digest=sha256:af2c86a67f1d315cf78c514ae902b37f95a1de2d5a0a088d29f6e49c5a3f775c

Observation c7b0eaa4-99aa-4dd1-b524-df3d3eb0f7d8 · outbound

This paper cites FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.045946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.045946Z digest=sha256:b1e754393945927749fec567b0332032cb574481b24ed39dcf7ae9372981074b

Observation 9b813077-7f04-4bf5-b216-b3b651c9f26a · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.053445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.053445Z digest=sha256:6f4fed4778c99beb2ff17397de3cdddeeb12c30c8ee98855ffa7724b1fe96633

Observation c92ee6ee-8c57-42d4-b1f2-2cdbaca38738 · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Diffusion Models Beat GANs on Image Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.062406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.062406Z digest=sha256:4515a75a5b3f2dee448b99e285018cd2fa1b28611a060ac59bc21dbb9d2427a4

Observation a9690037-55bb-4b8e-8a0a-34d7ee51fc23 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.069229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.069229Z digest=sha256:6af15c776d10e37ff9448703c7e6d8b0b1a1c026db2006ec4441a3b3cdb662c6

Observation 5802ce2b-5c00-4135-aaa4-ef23014dfc7a · outbound

This paper cites Motion prompting: Controlling video generation with motion trajectories, 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Motion prompting: Controlling video generation with motion trajectories, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.634312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.076608Z digest=sha256:06f2c1b53d923a9d622ad8b3ebbe4383e9a5fa0f04f666d03a84a4da90731027

Observation 462eb414-94be-4d97-a5f6-0265445f844f · outbound

This paper cites an unresolved cited work.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:04:48.614770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.082935Z digest=sha256:55a7bd72c736b1c2d58b09011fcaecddf15c726cfd0301fa21a84c1e1e15776b

Observation 0410430b-8358-4279-bfd1-a170355415da · outbound

This paper cites S., Shah, A., Yin, X., Parikh, D., and Misra, I.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models S., Shah, A., Yin, X., Parikh, D., and Misra, I

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.593746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.089154Z digest=sha256:24cdf51aae7b78e20d832b3f7bfda71813dc80e8cfb17bf23ffe1791adf92fbf

Observation c19d4bc7-9270-48c2-b27d-4dd49618a665 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.102846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.102846Z digest=sha256:48a0865db81c43fdab847c5108a9ab8beae3f83ee9ebec34e779dab370350dd4

Observation be521d57-e6eb-4338-be8a-ce4316659fd5 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models LTX-Video: Realtime Video Latent Diffusion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.109845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.109845Z digest=sha256:7c7854d66f160943420da1acec3126c571ea935e2f5889c6ca7eb184363954e6

Observation 886536c6-0bdb-44ae-9568-19d34c0669f9 · outbound

This paper cites Classifier-Free Diffusion Guidance.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Classifier-Free Diffusion Guidance

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.117338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.117338Z digest=sha256:4ed5a00a8d8b3b1f28619b6d85403ec9804e704ab2125d1445f8732ce9ba4864

Observation 0f8b8cf5-c4e6-4738-ab4c-f8674d54989d · outbound

This paper cites Denoising diffusion probabilistic models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Denoising diffusion probabilistic models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.124758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.124758Z digest=sha256:cd542d0c9579d1dc71d75eb80f87446d356dfcaef22a945cc0a64746573e7615

Observation 0fc08a37-7298-4ac8-a968-fc873aaa90a3 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Imagen Video: High Definition Video Generation with Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.137721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.137721Z digest=sha256:44b040eb344411387ff8f342705344177caf12474801b10492ba7f97aa0e7212

Observation b0b13fc9-5d76-49f8-b1d9-fdcd9f4291c7 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.144778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.144778Z digest=sha256:0544c76cc32af83c0c7f6f9ec74bc570ffb7429756828e42611424f452337b75

Observation c31c91d1-ba7b-480c-947c-a1554dde9248 · outbound

This paper cites VBench : Comprehensive benchmark suite for video generative models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models VBench : Comprehensive benchmark suite for video generative models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.537112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.151679Z digest=sha256:612784d0aa321bf356e0cdc1aadfd671e23778a4deae79d5735e41138ea9893f

Observation 09de8e98-60fe-4c79-b13a-4608e640291c · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling, 2024 a.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Pyramidal flow matching for efficient video generative modeling, 2024 a

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.517911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.157464Z digest=sha256:d87959fa3e6269761ecd56db2c5b64b59455f4586ac8947417036c4eba815f3d

Observation 7c0fc3d1-122b-444c-a288-31572fa68350 · outbound

This paper cites Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.493106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.163568Z digest=sha256:bd98940f4a7964c38bad9c4c26f52076818169ddec89a0d7c9e9e029d27f272b

Observation e4639c7f-dfac-475d-9c2e-ab49c6876d3a · outbound

This paper cites How far is video generation from world model: A physical law perspective, 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models How far is video generation from world model: A physical law perspective, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.474814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.170092Z digest=sha256:aa69b565f21389958d88a43ab9a8d319c13bf56a972e1427bae440386dd338ea

Observation 053c16d2-de68-4394-8ecb-67bed69fee26 · outbound

This paper cites Kling AI , 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Kling AI , 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.452135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.182859Z digest=sha256:861fdba1ce0e7bbf4f688d8808e063dc15801605d65fcb4056f6e5b491753e07

Observation 8905310e-95d7-464a-896a-19c5a7fbfadc · outbound

This paper cites an unresolved cited work.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:04:48.433408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.189116Z digest=sha256:b31f308b50587ca963784cd8a7e3cd0184c92b2998586e82c1e5af64a9e67448

Observation 2987bede-615a-42d4-961b-8fbf7cedf034 · outbound

This paper cites Compositional Visual Generation with Composable Diffusion Models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Compositional Visual Generation with Composable Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.195759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.195759Z digest=sha256:8d1535d3aa7ee886fde64462d965b5e6127299fabfd2278d2f34ef422d377357

Observation 38d61803-4c6c-487b-81f4-5287cdb3ea3c · outbound

This paper cites Physgen: Rigid-body physics-grounded image-to-video generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Physgen: Rigid-body physics-grounded image-to-video generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.410728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.202883Z digest=sha256:a18a925990fa3c113a01147d6243f662bb7d8acb9840996679a97ad6307c944b

Observation 9ffc3dfe-8fb2-4db8-86d8-2563d60760aa · outbound

This paper cites HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.210037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.210037Z digest=sha256:0e37f73a451ec7f6a0cff498c37ee11fe03ba77d4801a55997ab1e7a458b92f2

Observation 8a81af78-bd11-479b-8c02-c890c89495bb · outbound

This paper cites Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024 b.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024 b

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.387761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.217395Z digest=sha256:ecaca1852dd91321103a2a34a119a560abc4c05c794968924223c00c0aa2393b

Observation 6cf3b0bc-69a1-4c8f-b36e-cdb42c97fcdb · outbound

This paper cites K., Lewis, J.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models K., Lewis, J

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.362582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.224980Z digest=sha256:db516f76897ec611aa934266c1e823338a9e47bec3b3f0bc730165f85df50428

Observation 3284c59f-af94-4184-ab4a-99fdd23085b9 · outbound

This paper cites SDEdit : Guided image synthesis and editing with stochastic differential equations.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models SDEdit : Guided image synthesis and editing with stochastic differential equations

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.338774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.231751Z digest=sha256:887ed81d87c832f61b1835b5182f28cb88ab9611dfeb87ba09e7aaed3d1bc13e

Observation 15175173-fb54-4991-95fd-78223bc2ae7e · outbound

This paper cites Motioncraft: Physics-based zero-shot video generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Motioncraft: Physics-based zero-shot video generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.304086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.239604Z digest=sha256:bc754631a52970556e125937860e73a7b7ade9651aea81a9a3c3eaea1c9f9b22

Observation b42784c2-2786-44d7-8846-830f7ee23092 · outbound

This paper cites Dall-E 3 , 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Dall-E 3 , 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.282485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.248404Z digest=sha256:8f546ada88bca9536e9ee3bf394d140d2f5c7e365273a3f37c3f08991db3040d

Observation 6837dce2-b85b-4ccc-99a3-f563cb951282 · outbound

This paper cites and Xie, S.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models and Xie, S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.261801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.258880Z digest=sha256:a2e7c3214e8f81aaa07d1b3307a2ddfabf2ce87b1d566daf6d31c46226aac889

Observation 6a24c95c-e5d0-48b4-9362-644121881fda · outbound

This paper cites K., Zhang, P., Vajda, P., Duval, Q., Girdhar, R., Sumbaly, R., Rambhatla, S.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models K., Zhang, P., Vajda, P., Duval, Q., Girdhar, R., Sumbaly, R., Rambhatla, S

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.238350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.267406Z digest=sha256:53ab157df24ade26aec7b940852e1a93f3b79e1ecdaa17088a5b6b07d0309408

Observation 95fc7424-e68e-4e7b-ae4d-9a6ddb5f337d · outbound

This paper cites Hierarchical spatio-temporal decoupling for text-to-video generation, 2023.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Hierarchical spatio-temporal decoupling for text-to-video generation, 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.218282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.277593Z digest=sha256:822fd5f8ea900481033d74cf5aadaf4bf867670f29414a2667e00003a4fd993e

Observation d528d5f8-aa23-4da5-bab4-9eaedcc3b43d · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models High-resolution image synthesis with latent diffusion models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.285025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.285025Z digest=sha256:30c05ef9f65c4c25b0917a356dbdc1b2fd223ef983dba8f359307e42c38a8733

Observation 185e09bf-0d55-4417-8560-1237961dad62 · outbound

This paper cites Enhancing motion in text-to-video generation with decomposed encoding and conditioning, 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Enhancing motion in text-to-video generation with decomposed encoding and conditioning, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.184817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.292878Z digest=sha256:09dbaacd4164819958cdcfb2d761e328da4bf9f191066414e4be34b816a1c52a

Observation faf01b81-852b-4b01-9ed1-9ee4d869b5a8 · outbound

This paper cites DreamBooth : Fine tuning text-to-image diffusion models for subject-driven generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models DreamBooth : Fine tuning text-to-image diffusion models for subject-driven generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.165218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.299580Z digest=sha256:6f72df1d96c5e0418a47d524e1de9654f42816209422efec2def1702d268543c

Observation 6d07b195-a7f5-4c05-a50b-6954504c5d32 · outbound

This paper cites Gen-3 Alpha , 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Gen-3 Alpha , 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.145960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.307142Z digest=sha256:9c6e9cb9c31d84e60aa1b7272aef2b9b22d8b7e9b4abf0b90abd52291dfd3eac

Observation c07ad25f-1522-46f4-b83b-7d1a27265bff · outbound

This paper cites Decouple content and motion for conditional image-to-video generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Decouple content and motion for conditional image-to-video generation

Reference 47

Resolution
verified exact
doi, observed 2026-08-09T12:04:47.449872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.313539Z digest=sha256:2e2282a3c3ad3e1a6547cd7548ca3d753d753475e684c0cf87ab64c7740d40f1

Observation 8f39b790-6f7a-422b-8b89-4ebaf84d6829 · outbound

This paper cites C., See, S., Qin, H., Dai, J., and Li, H.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models C., See, S., Qin, H., Dai, J., and Li, H

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.126001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.320663Z digest=sha256:95c14eaa4015a61e69e26f36bf4d0e0a65b9fb1f002d9acc2daaf546b54db345

Observation 2f689545-a1a7-49e8-9b73-7538fd782d5e · outbound

This paper cites Make-A-Video : Text-to-video generation without text-video data.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Make-A-Video : Text-to-video generation without text-video data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.107102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.327444Z digest=sha256:9db5108e43a31ccb05ff35123ad2fe600ec8e9061ee7b68fca75b121bf9f3c32

Observation 1727dd0f-d530-4044-8fe9-fed7c92c0a7f · outbound

This paper cites UL2: Unifying Language Learning Paradigms.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models UL2: Unifying Language Learning Paradigms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.336723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.336723Z digest=sha256:a98435c7114d4e8a2e8fee2a5be37a9eb249852ccd54ea2ea4577aa5e1b8deac

Observation c1a8b496-3a6c-42aa-8748-c751e71c2ecf · outbound

This paper cites and Deng, J.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models and Deng, J

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.087835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.344122Z digest=sha256:4453b2abb70d1faf5cfe652035d2b6895ac44affcfa13a850dc2343be1c6a723

Observation 5f4f0caa-2d7f-4bc1-bdb7-33721262f9c5 · outbound

This paper cites MoCoGAN : Decomposing motion and content for video generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models MoCoGAN : Decomposing motion and content for video generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.070826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.350819Z digest=sha256:bb2e5e4c2b1336b5d135747baca2db1f169c670ad6c36c22ae094026948d15ae

Observation d3275e23-66b0-4ff6-8e8b-84d242941f6c · outbound

This paper cites ModelScope Text-to-Video Technical Report.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models ModelScope Text-to-Video Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.359552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.359552Z digest=sha256:102be3f2b504eae75015566a2fe9b6ece8b748ce14c22a103523615e6538ee4b

Observation 013abcbc-723c-424e-b43b-44ae7e72fdce · outbound

This paper cites Motif: Making text count in image animation with motion focal loss, 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Motif: Making text count in image animation with motion focal loss, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.049634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.368235Z digest=sha256:208b1cea5178a2e50832abf6edc789ce0a8febe1de8b646d37906a337b98016c

Observation 7394f5d3-49a2-497e-aff7-3604e1924e09 · outbound

This paper cites Z., Ge, Y., Wang, X., Lei, S.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Z., Ge, Y., Wang, X., Lei, S

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.027826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.375593Z digest=sha256:2a0e798f2f49ccccbb405c18634cc3d17e48c4aa1c08232cd354ecdef34031e8

Observation 3104c1c0-3a9a-4f8d-905e-c33383543a60 · outbound

This paper cites Demystifying CLIP Data.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Demystifying CLIP Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.381741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.381741Z digest=sha256:0f291834da6954a5fa8b497a874960ba2e6e783f6108b25ad8f5c100a2bc66f8

Observation da6c45d6-f6a4-4101-912f-577d24941d56 · outbound

This paper cites ByT5 : Towards a token-free future with pre-trained byte-to-byte models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models ByT5 : Towards a token-free future with pre-trained byte-to-byte models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.007015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.390095Z digest=sha256:11b22508ebece3b2c8d7d867bec9a2b2b14ad995e171087343fca39165895c37

Pith citing papers

Observation 3564ce24-f399-4a15-af4e-d0d840e1adc3 · inbound

MusicInfuser: Making Video Diffusion Listen and Dance cites this paper.

MusicInfuser: Making Video Diffusion Listen and Dance VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:32:15.635560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T23:28:51.767845Z digest=sha256:5854c9232e93c14fd85499b5f6488e8d37fd30d9cda269b4dce6ec9aabe082a8

Observation 8ba1652b-0765-49c2-8246-f68df6ddb1f5 · inbound

FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation cites this paper.

FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:12.520467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:12.520467Z digest=sha256:97330661c8a91402879243079f5db76414b46d8bc84e9c81edc1a6eeb41b0b46

Observation ab1affd8-b904-4a21-afde-00c6dd068898 · inbound

LumosFlow: Motion-Guided Long Video Generation cites this paper.

LumosFlow: Motion-Guided Long Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:12.386394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:27:12.386394Z digest=sha256:b2f85eba254a37a45e6072ca97247ae87f29aa644b7ff1dc7bd31a2dfcaf9fb6

Observation f96657f5-9e9c-4fdd-a597-dc00e1d115ed · inbound

UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting cites this paper.

UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:33.701446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:33.701446Z digest=sha256:a58325c0269e47314cbe3d3b1f35616d9d98e4a0d523db1b6223afcb62d25665

Observation 0280e3ab-442d-4bc2-b5e5-35c36b402062 · inbound

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly cites this paper.

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:04.394521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:04.394521Z digest=sha256:6ba0928a033ed4e5f2beacc89cb8cd6739cc7f1080ed296579d43bfbcf306223

Observation 77f362a1-7732-49d8-b983-322f2ad649a1 · inbound

LuxDiT: Lighting Estimation with Video Diffusion Transformer cites this paper.

LuxDiT: Lighting Estimation with Video Diffusion Transformer VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:02.115858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:02.115858Z digest=sha256:6d997769b79a0365cf25d54871e29c1de32d5b014ed0e017a73d3d8a434532fb

Observation 38f8dc10-0612-4b39-8cff-48529173ae1b · inbound

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders cites this paper.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.883059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.883059Z digest=sha256:3752af7b5681340fd9e5bd9651d4219b7196e380de12368c9349b31bb2994961

Observation e2de9bea-2e30-4ed5-b17b-54075e236bf4 · inbound

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation cites this paper.

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:39:54.294494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T22:39:53.995700Z digest=sha256:6e1ebeceec4529ab03c57e78f1d7aadff05a21a2eff964edb5ef8054bee9603f

Observation 1af90540-05d7-4c30-ad29-e85fa9ad3195 · inbound

Motus: A Unified Latent Action World Model cites this paper.

Motus: A Unified Latent Action World Model VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:44:36.734675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T18:44:36.636455Z digest=sha256:f387c3a1d6da88cde2586c21cea253a507f77667317433c4d064375a5daa256a

Observation 0b855597-dbdd-4373-9433-b6114998bd80 · inbound

REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion cites this paper.

REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T15:34:54.195615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:34:54.195615Z digest=sha256:9e86bd3bf728bb0043d3128e8eba24eb377f9c4337574d98b2a32293f4b92e08

Observation 8d6fb0af-9848-4bb3-8cf2-efad3a500ad8 · inbound

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos cites this paper.

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:47:57.542476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T13:43:26.460480Z digest=sha256:46c0c7ceacbe93783ffb9183907693ba3795735ac2f7267d3f094d18042956c8

Observation b241a86c-5a20-4bb5-aeee-dda2a521b46d · inbound

Reward-Forcing: Autoregressive Video Generation with Reward Feedback cites this paper.

Reward-Forcing: Autoregressive Video Generation with Reward Feedback VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:52:49.288157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T11:52:43.943226Z digest=sha256:3231f778f39537d9fabdf0b1be03a85668e69a1919f119b5fa49f100bc8548d4

Observation a5237902-803f-4c63-88f0-1a1ce5e7505b · inbound

Olaf-World: Orienting Latent Actions for Video World Modeling cites this paper.

Olaf-World: Orienting Latent Actions for Video World Modeling VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T01:20:04.762038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:20:04.762038Z digest=sha256:40d846bac445d22c2537ef7647d77ba3261beab575f3148540d435db748c502c

Observation feae8ad8-bf67-48e7-881c-dfd64b3de699 · inbound

HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation cites this paper.

HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.155549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T20:22:02.186361Z digest=sha256:947a2f520dc75a3338118ed3020429a891af86125131bb54b82a9fb30167cfd0

Observation 7e30ce3f-34d1-4348-b70f-8a9c59eec9d9 · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:41:04.207907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:381ae889b7676a275e1515a2031d383603e48cc5a5e9e429a26e11ff96c1c180

Observation 6109990c-7d5a-48fa-9cb1-1a0c1e99444d · inbound

From Priors to Perception: Grounding Video-LLMs in Physical Reality cites this paper.

From Priors to Perception: Grounding Video-LLMs in Physical Reality VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:08.267707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T17:41:23.233366Z digest=sha256:13280b74d80ddd7c4511e58a84478176f07fc521b4056e1a0f0f62930bf99c9b

Observation 17d67360-8438-420c-bfa0-88bde63bf99f · inbound

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking cites this paper.

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:29:28.858619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:28:12.151547Z digest=sha256:5a66168f477e1d8173e4c18d61e92de929176df11bfca8e6e3dd403d26c0ba3d

Observation df95399c-099d-4deb-a662-963f0d237a07 · inbound

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation cites this paper.

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:49:41.385833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T02:49:21.291716Z digest=sha256:5603608332a13a0e0d0b6ce29ab8e950d3c352f4533b66f7ce118c6fd301a135

Observation 85714818-bb9f-4531-ab90-c05f19eef78c · inbound

Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes cites this paper.

Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:10:22.241236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-25T05:07:14.227821Z digest=sha256:0ff854907f52e5b9663111fbd2fe339aa2422843d568de831b93d3965d1e2cd0

Observation e1b303de-5490-4d8a-8431-4d70e41e240a · inbound

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation cites this paper.

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:45:20.469846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-25T04:42:32.717968Z digest=sha256:155674e72d8cce28a2e8818279597d9f1e14b6cc05f166b133740cb287d51596

Observation 6fbf81cd-6226-4dbd-bf60-a242f417d914 · inbound

Tempered Self-Similarity Alignment for Physically Plausible Video Generation cites this paper.

Tempered Self-Similarity Alignment for Physically Plausible Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:44:38.411893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T11:39:06.597513Z digest=sha256:b87dba54e52ce387a34ad95df8299725b2ee63263cfc03dabe2700bd8c9c01bf

Observation 5e1d9312-98de-4133-b35b-3fab77002416 · inbound

OptiWorld: Optimal Control for Video World Generation under Physical Constraints cites this paper.

OptiWorld: Optimal Control for Video World Generation under Physical Constraints VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:32:35.541970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T19:02:51.848742Z digest=sha256:8d8ae1eadfcfa9eacf40c29b6e42f7ae7a93bb84f132891dc5c75d718155833d

Observation 96b52b38-a148-4b51-af97-2297a40c1add · inbound

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation cites this paper.

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.508999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T15:25:22.778550Z digest=sha256:f2f96363b7582cac83a34e9686e27cc7f098c00bfa02c9d9aa37a9ac418832d4

Observation 2a65feb2-1926-42da-ad60-7c563ff866e8 · inbound

SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation cites this paper.

SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.807908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T10:00:35.608696Z digest=sha256:393e3620cec6e73737dbc7297891e64fb741f9de918d6663c0f77478fa053cea

Observation 5d54411b-d156-4d00-9759-015cd3cdd3d0 · inbound

NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics cites this paper.

NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:18:43.391827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T04:25:27.067362Z digest=sha256:8e7a2f814e37fcac963f8e3833c2d90db56a2d16a56aa7fe755888418385f9bf

Observation 480316a2-93a5-4890-b86b-3749654d193a · inbound

HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers cites this paper.

HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:24:32.115411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T09:24:04.427234Z digest=sha256:67f61f7807ae8459959f7de3ca3ea8f5d0b2d04994b5ffa96533dc786f134b5a

Observation 2dc13058-177f-4fdd-9c28-8145f41e7066 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T01:59:43.167178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:59:43.167178Z digest=sha256:418c232a83098ac562ec21bdb62ad184cf7b5f908e4b39f5584ce6272e5878d4

Observation eb4a1a26-ee33-49dc-a82f-f77a0ff436c8 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T15:10:13.757131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:10:13.757131Z digest=sha256:f21344fd548bd9a297a700e849ed77af2a4995a345b5f524bf0862b349d49498

Observation 819d11f0-4f0b-48fd-a282-f70213438dfb · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T07:35:22.332436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:35:22.332436Z digest=sha256:02cdb3ca526093d8612048d6615c1e8014a0c011bd71c7f2999b1edbe42dc856

Observation be199bd7-a337-4f11-a865-fa77c59a3271 · inbound

DreamWAM: Beyond RGB Future Prediction for World Action Models cites this paper.

DreamWAM: Beyond RGB Future Prediction for World Action Models VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:28.004817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:59:28.004817Z digest=sha256:df477407b972933131cf265e918d40a3c5aa630ff5afc13791ed106661919680