Pith. sign in

Paper Citation Record · LEDGER

Phantom: Subject-consistent video generation via cross-modal alignment

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2502.11079.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.11079 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:32.071589Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:30.378145Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e8ca3bea-1703-4a77-b703-ae6b0d629a7b · inbound

VACE: All-in-One Video Creation and Editing cites this paper.

VACE: All-in-One Video Creation and Editing Phantom: Subject-consistent video generation via cross-modal alignment

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T00:53:54.003853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T00:53:53.855965Z digest=sha256:2d1adf0f93827670bde604262871bd9021ab02429c97c0b991e190a20ecaaaac

Observation dab58e75-016b-4c38-b61a-e1ac198ec9ce · inbound

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute cites this paper.

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute Phantom: Subject-consistent video generation via cross-modal alignment

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T17:51:54.313500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T17:51:41.947939Z digest=sha256:c6483883f5dda8b964015d09379e546cac9c877274ce08155a1e0c7a8ee6fddd

Observation 8b4584ab-6c87-4bf6-a062-9600001f2d39 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.071589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.071589Z digest=sha256:c91e0af18095b26a97ab86cf32d833627270515562fa9effe32074a7a82b3a8a

Observation 820661d1-558e-4c66-b64a-1d735b45ff39 · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:10.052472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:10.052472Z digest=sha256:aea4bbf7de48d6993dc7f896d448632c434f2cdcf1963aea41f3989b8dfb9a6d

Observation e62295d4-8fa2-4396-8deb-fac1b7d52bf6 · inbound

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement cites this paper.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.974027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.974027Z digest=sha256:39f7b403b3a587b76215168ce4c5de2941b7dbb200f2eb28455578919fe9f212

Observation 762ce83a-a2b1-4b2b-b0d3-7452fde4b615 · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers Phantom: Subject-consistent video generation via cross-modal alignment

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:39.046915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:39.046915Z digest=sha256:b37339f0865f73c5430b902da9295d319cdd7b4ca54432f0eb1460e90c28a1dc

Observation cde2f2fb-60aa-483a-8704-ee42bd252af6 · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:09.885970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:09.885970Z digest=sha256:a582698c1cecee9d4509931f54937c57d4372a7afd506de8159c190872270b41

Observation 78f0d332-7e44-4678-bd3e-d7f2777dd5a5 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Phantom: Subject-consistent video generation via cross-modal alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.649430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.649430Z digest=sha256:474fe491a178a891b0c90159cd85db9be06c99e96069ee614ff7993943ee8f70

Observation 5684ad4b-79c4-42a7-8685-928f47c503ff · inbound

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers cites this paper.

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers Phantom: Subject-consistent video generation via cross-modal alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.070911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:39:36.070911Z digest=sha256:3ab85043af9073072cc10c8b67f156bb7df585d3ebbf3a657f8cef9583569c5a

Observation d07100f1-754d-40de-b195-94151d481fd2 · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:47.800330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:47.800330Z digest=sha256:a9c01c79c16e39b3a5da155bde7dd1e6805da30b9f83d1186dd58dedc42ccb56

Observation eeeba291-e2d0-47b9-8c95-6822177149b3 · inbound

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement cites this paper.

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement Phantom: Subject-consistent video generation via cross-modal alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:42:40.992340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:42:40.992340Z digest=sha256:6ac6ef34dc4a3196c54938c32cfb63b4f027ee27c245883d89a3d2383f97686a

Observation 91ed0832-9ef6-4f2a-a878-5ce4e3355301 · inbound

UniVideo: Unified Understanding, Generation, and Editing for Videos cites this paper.

UniVideo: Unified Understanding, Generation, and Editing for Videos Phantom: Subject-consistent video generation via cross-modal alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.589583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.589583Z digest=sha256:56b6170164e5735aa80c234c07b1855d71c20f3a8a4a55fb921c00a6106a85fc

Observation e5d041c1-8351-482f-9b6d-34fad7652cad · inbound

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer cites this paper.

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T18:25:28.396016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T18:25:22.486891Z digest=sha256:14b904f5e5d3db17ba55539ce73d1ff648c934f39bbf6d71907dcc17799aed6a

Observation b1e13b4c-d732-427f-b594-34f56251ba95 · inbound

CustomX: Unified Character, Action, and Scene Customization in Video World Models cites this paper.

CustomX: Unified Character, Action, and Scene Customization in Video World Models Phantom: Subject-consistent video generation via cross-modal alignment

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:55.722892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:55.722892Z digest=sha256:07092ebb02e580d01358b1431b2cc69d569bc58703f7954b37c48cc50d217a7d

Observation 04a016b9-f2f9-4cfa-a631-94fba03f29ea · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:20:22.455970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:9ce043e3b31605de74574b2043e003a077c3dfc4be71a1d872f9c7b4a8b8cd17

Observation b16c105f-f6af-4215-ad0b-6fb297f91cdc · inbound

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model cites this paper.

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model Phantom: Subject-consistent video generation via cross-modal alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:18.232473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:51:18.232473Z digest=sha256:bff53c0a83c413f37b2671db94fbc4d69844423d5b9bd60125b753b5969f1a29

Observation 60704d15-2282-4390-b433-5a5fe85931b1 · inbound

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision cites this paper.

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision Phantom: Subject-consistent video generation via cross-modal alignment

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:55:49.210641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T20:33:25.422350Z digest=sha256:424571d3f7d9f07f73bf7b62abc279a8eafd7f596a4b0bbbead8be0a83dd73c4

Observation a5ba3bef-7a30-4b49-b8ba-dfc6e2dd5f94 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Phantom: Subject-consistent video generation via cross-modal alignment

Reference 205

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:05:51.328734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:97006eda55ac41aebbed663d46407631eec13a29a375262143cc64e54ac470f1

Observation d4177208-82fc-49ec-9bf1-bed27d5d2af9 · inbound

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation cites this paper.

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:01:00.416886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:42:35.693519Z digest=sha256:54553c1c5f02d989d243e1007e884f2fe59836a82bdf3426ae083d0fab6a2249

Observation c5df77d7-6e3f-49ed-8acf-5e3c5ddccbe4 · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding Phantom: Subject-consistent video generation via cross-modal alignment

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:03.692220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:9e4a4e37b19a172856e70191116a1d5ef95429243e3215e4e42e9e03d35bc576

Observation 8990056c-c55a-4965-b135-a1b03ef9f68b · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Phantom: Subject-consistent video generation via cross-modal alignment

Reference 141

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:56.950219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:b2a4d150d3900c345637b13f21e23afbf39f6500e0d09abdb7f632a7817b8f17

Observation a504db6f-6ef3-4462-bd37-eae9a89ad9b7 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Phantom: Subject-consistent video generation via cross-modal alignment

Reference 120

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:af663da6b541dee353074a5e5d80be38c10f758d920e5bf823904bc69aa47122

Observation 78cd8348-c81b-4c6c-8519-b0c3160d0004 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:06:05.893388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:62189b80faa44f12f965dc50ee724984ef3d54bf0dae197efdac9fa70a001ea7

Observation 8020a8c1-a8f7-4b1e-bd94-51dde555e8a1 · inbound

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis cites this paper.

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis Phantom: Subject-consistent video generation via cross-modal alignment

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:01:26.592322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:24:57.882447Z digest=sha256:551b2bd0165ba986bb839eae37e903e86eb96fa3b538769ad8a26bcc5bd4ca15

Observation 9c0c0c03-b5a1-479c-bb2b-3fd35485b6c0 · inbound

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation cites this paper.

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:21:09.635204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:40:00.224358Z digest=sha256:fb6e7b03f1655db2f40d51234c0a8aa5e00810ce918a5efe87ab4d8c598eb1a2

Observation 871d318c-ba5c-438c-ae15-f110cd8a7bac · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization Phantom: Subject-consistent video generation via cross-modal alignment

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:13:43.610050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T20:11:33.277349Z digest=sha256:d50d09c8ddef75e02ba1f002fdb757b42584733bf4331e30e4ff45f743c12719

Observation 66e35ba8-be81-4f43-982a-63b0a576b155 · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization Phantom: Subject-consistent video generation via cross-modal alignment

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T19:25:00.562996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T19:23:07.044659Z digest=sha256:d02006f3eda2ff4d141b13697222372aa74bfc4d85c4fe4f8296d65a5a23078e

Observation b6a1476b-f6fc-4200-b0fc-087b41a9a896 · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion Phantom: Subject-consistent video generation via cross-modal alignment

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T06:41:10.542258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:c5be61c9abccfccb9487cb40521a56b3289accc6d609e6b7041e3dcafbc709e7

Observation 048d38ed-a3bc-4ad6-a5c0-aa4a4bded87a · inbound

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation cites this paper.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.425250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:8ac68e75ea45edcec5f16ab7c58d2454b29275014cd4ffd648ab6d80397f71e1

Observation 54d0a592-7bba-47d0-b7f3-a32634311248 · inbound

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data cites this paper.

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data Phantom: Subject-consistent video generation via cross-modal alignment

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.244659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:52:30.406683Z digest=sha256:00ea7e32e0dd2636ce1e740a9537c9d293144fab6665b58649b512939521e1ae

Observation 610c77d1-d687-469a-abaf-59b53dac25f1 · inbound

Streaming Video Generation with Streaming Force Control cites this paper.

Streaming Video Generation with Streaming Force Control Phantom: Subject-consistent video generation via cross-modal alignment

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.387385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:14:32.465663Z digest=sha256:28d64a5fc3d852c1490d78fc3d6690b6de6e24639a56c5990d5114cc2099bbaf

Observation f48bf4c6-21ea-43b5-b88f-86cac9cca637 · inbound

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation cites this paper.

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:37:36.648289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:49:50.272650Z digest=sha256:5b0037b5f2dd7ab36f1bd590e44639e1911f57ffde31d6da1028e6632867d6d1

Observation 24afdf48-3e48-4b2e-a930-603545850b29 · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:17:45.947052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:31882223245b6f2720e902de751af9096701e7eead6012538b00ce756c4675e6

Observation 68f86b7d-16e3-4865-a3fa-70fb1297d568 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling Phantom: Subject-consistent video generation via cross-modal alignment

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:30.381652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T18:15:51.862192Z digest=sha256:cfd5b05cf843e523fb897995f784ecf7e4ec8731db580b3f6af0608204d7431f

Observation 73bda7db-f3f1-424c-89bf-43f5dcb8f4a5 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling Phantom: Subject-consistent video generation via cross-modal alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:05.537673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:05.537673Z digest=sha256:4122ba5ae4854cfb6bdec205e025eab096d670683da181c47d57a39e227ced54

Observation ecf247e4-33a2-4343-bb22-f847989a56a7 · inbound

A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks cites this paper.

A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks Phantom: Subject-consistent video generation via cross-modal alignment

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:45:48.892408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T00:59:50.478584Z digest=sha256:28afb4c3d8e4709c20fb7f368873d13148fa7fd1d0f36e20a70f828299724359

Observation 95c4ed9a-fded-45a9-8f3e-332d66048293 · inbound

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation cites this paper.

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T16:30:02.172800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:30:02.172800Z digest=sha256:b3e90f5256ef528ac7d465055f9d55c86e68d6d39734e2551584d2a2cf926d77

Observation d111140e-0430-4dbc-a018-a4527dfbfc31 · inbound

ID-V2V: Identity-Preserving Video Restylization cites this paper.

ID-V2V: Identity-Preserving Video Restylization Phantom: Subject-consistent video generation via cross-modal alignment

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-01T04:27:31.566712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:27:31.566712Z digest=sha256:d15024483141070dca46eef13a6b1ec044bd7de70d1c1cb3533b1fdae24f766e