Pith. sign in

Paper Citation Record · LEDGER

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation

As of 21 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2605.23610.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.23610 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T04:57:13.479165Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:41:46.594758Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact15
  • verified fuzzy3
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92d1b717-84f2-413f-b758-9505b40d3340 · outbound

This paper cites Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Di- dac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Di- dac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.371292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:0a3f1d31ea4a9ceda0ef31e150e6aa739b1e16056087e6e2faa60daf916211d3

Observation f2a8f08a-9e23-4058-a6eb-57b57866918e · outbound

This paper cites SAM 3: Segment Anything with Concepts.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation SAM 3: Segment Anything with Concepts

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.354179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:6388f199c5b5e860ac4a80b380ae924265352f4424be55df843cd12abf7ac721

Observation 9d28e263-4cfb-4f3c-b862-6d752477acfc · outbound

This paper cites Seedance 1.0: Exploring the Boundaries of Video Generation Models.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Seedance 1.0: Exploring the Boundaries of Video Generation Models

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.359546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:39d5eab03d6fbe3aa4ea6d6843458a111295fd0dc586f2da5a8eb1bfe35904bc

Observation f29817ed-b14d-4853-a33e-34511bee3852 · outbound

This paper cites Long Context Tuning for Video Generation.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Long Context Tuning for Video Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.344529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:de3d5dd74d7a518be92aee5c75d182b7f0f67055cc2a364c6ee72ef2eb985b7e

Observation f7f5049d-314c-4cf7-b4ab-5d1cd429cca8 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation LTX-Video: Realtime Video Latent Diffusion

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.349392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:1d19c2405f1939e204315c6cabc299b1a16c0ca12d6ad9a1de9cc896fe35ab7e

Observation 26aaf508-9ee6-4e2f-bdf8-8cfda720bf03 · outbound

This paper cites StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.365290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:90fba87e7088c900f759f37b7db77795461f226a58b68e885e75f4549a6b20ce

Observation 4466a625-7cd4-488a-83ab-9f697e297425 · outbound

This paper cites Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.433579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:2cc4e15c2fbcb1fa5dfd96882441510529e8df1405c28b1bf0039ea745472230

Observation a42a244f-5b48-43ad-8485-712bbc25adb9 · outbound

This paper cites Jiaxiu Jiang, Wenbo Li, Jingjing Ren, Yuping Qiu, Yong Guo, Xiaogang Xu, Han Wu, and Wangmeng Zuo.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Jiaxiu Jiang, Wenbo Li, Jingjing Ren, Yuping Qiu, Yong Guo, Xiaogang Xu, Han Wu, and Wangmeng Zuo

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.438537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:dba697718737c44794c66a49abf5119c2177bcc0a72b7b9c636e8b9855ceea81

Observation 9371b6bb-a56e-4b8f-883b-2f2005ae3c56 · outbound

This paper cites LoViC: Efficient Long Video Generation with Context Compression.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation LoViC: Efficient Long Video Generation with Context Compression

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.453810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:c0ec666c9156aa68b657f7aec13928008750c2c4c2e778d511505991c089be24

Observation 275dac04-d344-4289-a29e-11c8bba72c3e · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:00:22.410894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:134988d28b6068ee2eb8ac1674fe4eda5846e7ae8d93b7d8591733570d65d243

Observation ad387fd1-38a3-49bb-9d74-977093202338 · outbound

This paper cites Lijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen, Jiawei Liu, Gen Li, Siyu Zhou, Qian He, and Xinglong Wu.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Lijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen, Jiawei Liu, Gen Li, Siyu Zhou, Qian He, and Xinglong Wu

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T12:57:02.474847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:bbc7644cdef510ed68e868b353d339d19186c2931d6d85462d74093877f1ae3d

Observation 048d38ed-a3bc-4ad6-a5c0-aa4a4bded87a · outbound

This paper cites Phantom: Subject-consistent video generation via cross-modal alignment.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.425250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:690594e848bc67ddc6a42be44b08bbc5e6a039ab8d7214f76f71c34aee1a5813

Observation e19dd601-3df6-417f-b524-9a595c603bc4 · outbound

This paper cites Yihao Meng, Hao Ouyang, Yue Yu, Qiuyu Wang, Wen Wang, Ka Leong Cheng, Hanlin Wang, Yixuan Li, Cheng Chen, Yanhong Zeng, et al.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Yihao Meng, Hao Ouyang, Yue Yu, Qiuyu Wang, Wen Wang, Ka Leong Cheng, Hanlin Wang, Yixuan Li, Cheng Chen, Yanhong Zeng, et al

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.416607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:3c3f2aff5815d9c34afab8585d8582f99a857d6877a5d14497aa92605db32194

Observation df9f3999-920c-4a97-b040-30a351d15a71 · outbound

This paper cites Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, et al.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, et al

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.448662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:731b7624f30a002736ac2a29b597239e72c41dc66c27cfc142f7be1453acd870

Observation 0e74b0de-6789-4086-ae9f-8ad9f2aa1c4a · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.406491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:4d2f5a9b2b2256308c4f4540dab8429adce2816ea3fb3c9fe326da282ec00561

Observation 1192408d-f51e-4641-9240-eb5249c495a0 · outbound

This paper cites Xierui Wang, Siming Fu, Qihan Huang, Wanggui He, and Hao Jiang.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Xierui Wang, Siming Fu, Qihan Huang, Wanggui He, and Hao Jiang

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.386844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:7a4e813dcc68cb1b6f0bee8ec709d6a8042108ff25725024c8d233e4cd970b67

Observation 524961f5-bf00-4db4-8037-e997ead0bfa9 · outbound

This paper cites MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.391492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:07c1f3ec1e71b3aebf72cee40740266d2b7fa8ce26076f0b9966fe95e51eba34

Observation 983cb006-f158-4bb8-8bb6-e9c1df261456 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.396412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:06984ebf3924b5ae47c2bfcfcc07169403ed9bc2222d72b7eae3c3068a194d28

Observation cef2477f-1136-4ecc-9c98-119b7c78b6fd · outbound

This paper cites Captain Cinema: Towards Short Movie Generation.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Captain Cinema: Towards Short Movie Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.443351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:14721f1e68f5544d1066cedb0a54123d333fc25566e8232bfb36a7e126bf4caa

Observation 435360c2-7592-435d-bc53-b612f130e17d · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.401300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:3cdde4b2eb9b1d9be5c8553f1fc152cdd247549dd8feb864124316054cca526c

Observation 65022ada-aeee-4d40-b802-efc9e5a212fb · outbound

This paper cites Lvmin Zhang, Shengqu Cai, Muyang Li, Chong Zeng, Beijia Lu, Anyi Rao, Song Han, Gordon Wetzstein, and Maneesh Agrawala.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Lvmin Zhang, Shengqu Cai, Muyang Li, Chong Zeng, Beijia Lu, Anyi Rao, Song Han, Gordon Wetzstein, and Maneesh Agrawala

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.376667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:11fe9b5385800c4ea3a4371feb34aec67e896f86250f00fa3bd168d28fd1dc10

Observation 0c18829f-edee-45b2-a013-1f4da83bd6b7 · outbound

This paper cites Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, and Qibin Hou.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, and Qibin Hou

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.381836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:cba40cc932a0b3a4698c0c3b752010b9d23f02042238645c6551343a35c42c0d

Observation 1180b002-3b58-43c7-8a30-7c6fb3df0111 · outbound

This paper cites Advances in Neural Information Processing Systems37 (2024).

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Advances in Neural Information Processing Systems37 (2024)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T12:57:02.478402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:f7b509839ffbf4a35e8cea1985dd00f4918dfe997355de744b4c417a977d0050

Observation 217d35d4-cd64-4b8f-82fd-319417d5dd43 · outbound

This paper cites Statistic Value Scene Statistics Indoor / outdoor shots (%) 25.7 / 74.3 Avg.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Statistic Value Scene Statistics Indoor / outdoor shots (%) 25.7 / 74.3 Avg

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T12:57:02.482406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:8dec7b28c36858a105f9e4af845e5a3263e8341df403594066a7dc2d26270018

Pith citing papers

Observation 2723e283-e2d4-4ff5-85f8-e3ca0acfd853 · inbound

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling cites this paper.

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T13:41:46.594758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:41:46.594758Z digest=sha256:acfc68a9fd9b7fe6289ad60c6fb710b51ec92356f3f74bed73ea079e70234376