Pith. sign in

Paper Citation Record · LEDGER

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation

As of 21 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2605.23610.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.23610 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T04:57:13.479165Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:41:46.594758Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact15
  • verified fuzzy3
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92d1b717-84f2-413f-b758-9505b40d3340 · outbound

This paper cites Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Di- dac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Di- dac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.371292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:3ccab06568ba6ba4484e5b7ae9926a7efadcde8f1731aca3b224bb807d52e14b

Observation f2a8f08a-9e23-4058-a6eb-57b57866918e · outbound

This paper cites SAM 3: Segment Anything with Concepts.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation SAM 3: Segment Anything with Concepts

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.354179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:cb69baa700ba173372798d180c5e73735492b2d9b9fb7fdc27896b675bc47394

Observation 9d28e263-4cfb-4f3c-b862-6d752477acfc · outbound

This paper cites Seedance 1.0: Exploring the Boundaries of Video Generation Models.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Seedance 1.0: Exploring the Boundaries of Video Generation Models

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.359546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:94ba39ca2e366469fb7ff66e43518b89812b96d4dc8c794606941cd8cf393f7b

Observation f29817ed-b14d-4853-a33e-34511bee3852 · outbound

This paper cites Long Context Tuning for Video Generation.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Long Context Tuning for Video Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.344529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:69ebf22ecf1546a18a3d77babb6af7562d78463b3ba1d718c3d073f15be3702a

Observation f7f5049d-314c-4cf7-b4ab-5d1cd429cca8 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation LTX-Video: Realtime Video Latent Diffusion

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.349392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:706f17428325fc40a806d91d2471e47f0fb3db4abd998e92b87fad0630229861

Observation 26aaf508-9ee6-4e2f-bdf8-8cfda720bf03 · outbound

This paper cites StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.365290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:ce0aff5d6efb752dcb97942dd89842cd05a9a915081c7fb97efea20b80b06314

Observation 4466a625-7cd4-488a-83ab-9f697e297425 · outbound

This paper cites Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.433579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:035d1bf344ada046e8c6a53bddfdcd1289ceb8d9f03f39106626fdd8c34c9dad

Observation a42a244f-5b48-43ad-8485-712bbc25adb9 · outbound

This paper cites Jiaxiu Jiang, Wenbo Li, Jingjing Ren, Yuping Qiu, Yong Guo, Xiaogang Xu, Han Wu, and Wangmeng Zuo.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Jiaxiu Jiang, Wenbo Li, Jingjing Ren, Yuping Qiu, Yong Guo, Xiaogang Xu, Han Wu, and Wangmeng Zuo

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.438537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:1d78bcccdc4587dbe205c9fcc12e7aa5d24ea7239b27a5a9e75140cfdb3e4c85

Observation 9371b6bb-a56e-4b8f-883b-2f2005ae3c56 · outbound

This paper cites LoViC: Efficient Long Video Generation with Context Compression.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation LoViC: Efficient Long Video Generation with Context Compression

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.453810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:cac6e4be94f9cbb884fec2d4c91ad9ad2742abe25b3b8e1f6b81e1cf1af9da51

Observation 275dac04-d344-4289-a29e-11c8bba72c3e · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:00:22.410894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:fe39268d5bb141910dcb2cafdc4e43bc181ba12a4c5fad7fb0c0168b171ff71c

Observation ad387fd1-38a3-49bb-9d74-977093202338 · outbound

This paper cites Lijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen, Jiawei Liu, Gen Li, Siyu Zhou, Qian He, and Xinglong Wu.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Lijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen, Jiawei Liu, Gen Li, Siyu Zhou, Qian He, and Xinglong Wu

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T12:57:02.474847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:d32dc5f03148caccb6aac088194b843e6344c064115e0370547d50ed7b403b84

Observation 048d38ed-a3bc-4ad6-a5c0-aa4a4bded87a · outbound

This paper cites Phantom: Subject-consistent video generation via cross-modal alignment.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.425250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:f25e03b262943643c91c2374194b017eff3dca4ea2579340d97f742d2356340d

Observation e19dd601-3df6-417f-b524-9a595c603bc4 · outbound

This paper cites Yihao Meng, Hao Ouyang, Yue Yu, Qiuyu Wang, Wen Wang, Ka Leong Cheng, Hanlin Wang, Yixuan Li, Cheng Chen, Yanhong Zeng, et al.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Yihao Meng, Hao Ouyang, Yue Yu, Qiuyu Wang, Wen Wang, Ka Leong Cheng, Hanlin Wang, Yixuan Li, Cheng Chen, Yanhong Zeng, et al

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.416607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:87db029905761230feebe1f81aa864c5dac006b74fb8431ff5e6a071dd0fda3a

Observation df9f3999-920c-4a97-b040-30a351d15a71 · outbound

This paper cites Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, et al.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, et al

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.448662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:0958413d43ce0ca74c5542771e6fe9a7ba6e4f7e1603fe8c188a402ecb3cd42e

Observation 0e74b0de-6789-4086-ae9f-8ad9f2aa1c4a · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.406491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:1dffc97927fa85ad9f98c9ccc0bcb1625013f224f5fc5744b68ff328d26c1c3e

Observation 1192408d-f51e-4641-9240-eb5249c495a0 · outbound

This paper cites Xierui Wang, Siming Fu, Qihan Huang, Wanggui He, and Hao Jiang.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Xierui Wang, Siming Fu, Qihan Huang, Wanggui He, and Hao Jiang

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.386844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:e3568c13af9daaf39d5cf4f298985eee4dc7701562edef235d6d6845053891ad

Observation 524961f5-bf00-4db4-8037-e997ead0bfa9 · outbound

This paper cites MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.391492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:17c77e78dd121b42e2c4d582f588eb836b140270857e83a83a09c5daebce8e9b

Observation 983cb006-f158-4bb8-8bb6-e9c1df261456 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.396412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:f5903115550d8d2699be9934ef43098f886a57e74e6ac3636842f52f56122a2e

Observation cef2477f-1136-4ecc-9c98-119b7c78b6fd · outbound

This paper cites Captain Cinema: Towards Short Movie Generation.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Captain Cinema: Towards Short Movie Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.443351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:4a705100ad7ff2d05b459f31a23877c8d0b114cb6fc37d134ba2630ec0ce5141

Observation 435360c2-7592-435d-bc53-b612f130e17d · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:22.401300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:2d3883b90f3a8f49838866914d8d32d5763079161ecb1ec807e459aa60bb4f6b

Observation 65022ada-aeee-4d40-b802-efc9e5a212fb · outbound

This paper cites Lvmin Zhang, Shengqu Cai, Muyang Li, Chong Zeng, Beijia Lu, Anyi Rao, Song Han, Gordon Wetzstein, and Maneesh Agrawala.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Lvmin Zhang, Shengqu Cai, Muyang Li, Chong Zeng, Beijia Lu, Anyi Rao, Song Han, Gordon Wetzstein, and Maneesh Agrawala

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.376667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:839be70413297cd4d1559bde67bc37267f21681c025a2e1f3ae2a9d51ad323cb

Observation 0c18829f-edee-45b2-a013-1f4da83bd6b7 · outbound

This paper cites Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, and Qibin Hou.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, and Qibin Hou

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.381836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:ac467b65f33ba8e0ef9a9475a49496d211894f2f9ca98907e2f785e5f684d105

Observation 1180b002-3b58-43c7-8a30-7c6fb3df0111 · outbound

This paper cites Advances in Neural Information Processing Systems37 (2024).

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Advances in Neural Information Processing Systems37 (2024)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T12:57:02.478402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:0f956d3fb436361bcefb8c201410343a10ca3413552ea98c496ea0b7b88b1713

Observation 217d35d4-cd64-4b8f-82fd-319417d5dd43 · outbound

This paper cites Statistic Value Scene Statistics Indoor / outdoor shots (%) 25.7 / 74.3 Avg.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Statistic Value Scene Statistics Indoor / outdoor shots (%) 25.7 / 74.3 Avg

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T12:57:02.482406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:456ba5b1b993c55ffd70f3a8a24bad18bfab9e9874d7c785822376cd833c2c67

Pith citing papers

Observation 2723e283-e2d4-4ff5-85f8-e3ca0acfd853 · inbound

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling cites this paper.

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T13:41:46.594758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:41:46.594758Z digest=sha256:acfc68a9fd9b7fe6289ad60c6fb710b51ec92356f3f74bed73ea079e70234376