Pith. sign in

Paper Citation Record · LEDGER

Phantom: Subject-consistent video generation via cross-modal alignment

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2502.11079.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.11079 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:32.071589Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:30.378145Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e8ca3bea-1703-4a77-b703-ae6b0d629a7b · inbound

VACE: All-in-One Video Creation and Editing cites this paper.

VACE: All-in-One Video Creation and Editing Phantom: Subject-consistent video generation via cross-modal alignment

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T00:53:54.003853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T00:53:53.855965Z digest=sha256:e65d892268f0ab7d7f0fbb458b1b5d7a03cb53718a8662747c8c7a72817b000b

Observation dab58e75-016b-4c38-b61a-e1ac198ec9ce · inbound

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute cites this paper.

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute Phantom: Subject-consistent video generation via cross-modal alignment

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T17:51:54.313500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T17:51:41.947939Z digest=sha256:fb53789d18d6916f699ddae46a1d826134f53f79e9450e7635a0238cfa8bc2cb

Observation 8b4584ab-6c87-4bf6-a062-9600001f2d39 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.071589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.071589Z digest=sha256:f65742edda918bbd81ecb68a73b64a40f84772578bcf2e5229d4ec2139b3b7d2

Observation 820661d1-558e-4c66-b64a-1d735b45ff39 · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:10.052472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:10.052472Z digest=sha256:aea4bbf7de48d6993dc7f896d448632c434f2cdcf1963aea41f3989b8dfb9a6d

Observation e62295d4-8fa2-4396-8deb-fac1b7d52bf6 · inbound

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement cites this paper.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.974027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.974027Z digest=sha256:66e0fc511b9a09d4404f13e26dddb93681bdf2b882594f5fe9fef0ac696a687a

Observation 762ce83a-a2b1-4b2b-b0d3-7452fde4b615 · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers Phantom: Subject-consistent video generation via cross-modal alignment

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:39.046915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:39.046915Z digest=sha256:82c0210176c68306a5f991b37bd50891b3672048bdcba0f84bd7057457019873

Observation cde2f2fb-60aa-483a-8704-ee42bd252af6 · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:09.885970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:09.885970Z digest=sha256:a582698c1cecee9d4509931f54937c57d4372a7afd506de8159c190872270b41

Observation 78f0d332-7e44-4678-bd3e-d7f2777dd5a5 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Phantom: Subject-consistent video generation via cross-modal alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.649430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.649430Z digest=sha256:d18d143c36f33b39a4b92587ad4a13d7d9cb87ead15c21428459c14f705b96c1

Observation 5684ad4b-79c4-42a7-8685-928f47c503ff · inbound

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers cites this paper.

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers Phantom: Subject-consistent video generation via cross-modal alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.070911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:39:36.070911Z digest=sha256:e8ecbfc088f783adeaa2bbb6469d9723171ef8cab38a2614943faa7fb4d22f8d

Observation d07100f1-754d-40de-b195-94151d481fd2 · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:47.800330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:47.800330Z digest=sha256:a9c01c79c16e39b3a5da155bde7dd1e6805da30b9f83d1186dd58dedc42ccb56

Observation eeeba291-e2d0-47b9-8c95-6822177149b3 · inbound

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement cites this paper.

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement Phantom: Subject-consistent video generation via cross-modal alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:42:40.992340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:42:40.992340Z digest=sha256:8a7307c78e6378435970cbce92ea1106083763e9d69e87dd853c5b140c122fef

Observation 91ed0832-9ef6-4f2a-a878-5ce4e3355301 · inbound

UniVideo: Unified Understanding, Generation, and Editing for Videos cites this paper.

UniVideo: Unified Understanding, Generation, and Editing for Videos Phantom: Subject-consistent video generation via cross-modal alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.589583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.589583Z digest=sha256:56b6170164e5735aa80c234c07b1855d71c20f3a8a4a55fb921c00a6106a85fc

Observation e5d041c1-8351-482f-9b6d-34fad7652cad · inbound

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer cites this paper.

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T18:25:28.396016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T18:25:22.486891Z digest=sha256:78577b7cf63d219baf40fa6f72ec3f7c51a6dc2e2e587c7d75d542ce23f4a894

Observation b1e13b4c-d732-427f-b594-34f56251ba95 · inbound

CustomX: Unified Character, Action, and Scene Customization in Video World Models cites this paper.

CustomX: Unified Character, Action, and Scene Customization in Video World Models Phantom: Subject-consistent video generation via cross-modal alignment

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:55.722892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:55.722892Z digest=sha256:ecff93c3a0a6f4677064daa6314f400b09d64e10eccac7c0381caaa4e07f9100

Observation 04a016b9-f2f9-4cfa-a631-94fba03f29ea · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:20:22.455970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:22d248b978cf8f7831a4ff3fbee7b8a1def6896e31bf06dfcdacb71a7d2689be

Observation b16c105f-f6af-4215-ad0b-6fb297f91cdc · inbound

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model cites this paper.

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model Phantom: Subject-consistent video generation via cross-modal alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:18.232473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:51:18.232473Z digest=sha256:ec1c8ce547e358be10607893b074b3a71e7a45555de4f581ded2049faca95ecc

Observation 60704d15-2282-4390-b433-5a5fe85931b1 · inbound

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision cites this paper.

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision Phantom: Subject-consistent video generation via cross-modal alignment

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:55:49.210641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T20:33:25.422350Z digest=sha256:2cebd389e2afb0ceff3c4f30820798bf506064b09feaab755e964d730d8c1085

Observation a5ba3bef-7a30-4b49-b8ba-dfc6e2dd5f94 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Phantom: Subject-consistent video generation via cross-modal alignment

Reference 205

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:05:51.328734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:6c0f3e5019ad7cd5bb9964f4eb295fc47f4df45a86c4995c4a90afa8addfaede

Observation d4177208-82fc-49ec-9bf1-bed27d5d2af9 · inbound

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation cites this paper.

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:01:00.416886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T15:42:35.693519Z digest=sha256:1d9ad0e4389bdf5eddea4eb58c23823f41a4f661b1294294033c379bf9b605a0

Observation c5df77d7-6e3f-49ed-8acf-5e3c5ddccbe4 · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding Phantom: Subject-consistent video generation via cross-modal alignment

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:03.692220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:840d79e35c890a1bd2615e1a1646063a4e9816294e586d42ae2817f89ba86f1d

Observation 8990056c-c55a-4965-b135-a1b03ef9f68b · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Phantom: Subject-consistent video generation via cross-modal alignment

Reference 141

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:56.950219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:02a42ca2638332084d5de4341d25ef2ac60c88d01b1dffbc1f322dcc6bf08f2f

Observation a504db6f-6ef3-4462-bd37-eae9a89ad9b7 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Phantom: Subject-consistent video generation via cross-modal alignment

Reference 120

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:af663da6b541dee353074a5e5d80be38c10f758d920e5bf823904bc69aa47122

Observation 78cd8348-c81b-4c6c-8519-b0c3160d0004 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:06:05.893388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:39f5e2de0c8ff3e713149a1a4e36617b45490e0ffec895937d7e8b318accc42b

Observation 8020a8c1-a8f7-4b1e-bd94-51dde555e8a1 · inbound

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis cites this paper.

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis Phantom: Subject-consistent video generation via cross-modal alignment

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:01:26.592322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T02:24:57.882447Z digest=sha256:733087fbaeaef1d6688f61e5f53686b33dcad0f10f5ba0aabbb1e05dfff8a84b

Observation 9c0c0c03-b5a1-479c-bb2b-3fd35485b6c0 · inbound

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation cites this paper.

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:21:09.635204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T17:40:00.224358Z digest=sha256:59ec18053a98d5446c5283bf78346aa3d7f3e31208f2f2a8e4418882daa08afe

Observation 871d318c-ba5c-438c-ae15-f110cd8a7bac · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization Phantom: Subject-consistent video generation via cross-modal alignment

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:13:43.610050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T20:11:33.277349Z digest=sha256:76f13ada867b23d122432c0497eccbc91a0d403a6057e17766761daed962f7e2

Observation 66e35ba8-be81-4f43-982a-63b0a576b155 · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization Phantom: Subject-consistent video generation via cross-modal alignment

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T19:25:00.562996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T19:23:07.044659Z digest=sha256:b68dfd28ddf6d87841737f875b2adca1e27631f7db0e4ca27c99ad37ceaf91dc

Observation b6a1476b-f6fc-4200-b0fc-087b41a9a896 · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion Phantom: Subject-consistent video generation via cross-modal alignment

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T06:41:10.542258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:13de7143874bb967b0560994cc1a00d6ad985532db76a593de29f5ce1a09cca8

Observation 048d38ed-a3bc-4ad6-a5c0-aa4a4bded87a · inbound

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation cites this paper.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.425250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:f6b96947773de82258c4dd373a7b47b79c2053484496bf53e644789674c55d1b

Observation 54d0a592-7bba-47d0-b7f3-a32634311248 · inbound

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data cites this paper.

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data Phantom: Subject-consistent video generation via cross-modal alignment

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.244659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T14:52:30.406683Z digest=sha256:22b284b32073dfacbe4e5067a932fef61c56b4bad5b4b8a6cebafa81e72c08ad

Observation 610c77d1-d687-469a-abaf-59b53dac25f1 · inbound

Streaming Video Generation with Streaming Force Control cites this paper.

Streaming Video Generation with Streaming Force Control Phantom: Subject-consistent video generation via cross-modal alignment

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.387385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T22:14:32.465663Z digest=sha256:37d712c6a088d5881b0819a418b53a792bb419829a6ebdac23b91d470a00f380

Observation f48bf4c6-21ea-43b5-b88f-86cac9cca637 · inbound

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation cites this paper.

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:37:36.648289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T13:49:50.272650Z digest=sha256:c5850b453e188fb17de7e765dcf19455bbb5a41e01ad93ee540cf4767b753264

Observation 24afdf48-3e48-4b2e-a930-603545850b29 · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:17:45.947052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:5604cd5db6c9ffc10391d8b2523adb4a7ddfbd141c02cde3d9ba2a0e8bc3be57

Observation 68f86b7d-16e3-4865-a3fa-70fb1297d568 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling Phantom: Subject-consistent video generation via cross-modal alignment

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:30.381652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T18:15:51.862192Z digest=sha256:05cd3680e56528b7ee47cbf039497207730d655114c3a21ae44efadd45427a2e

Observation 73bda7db-f3f1-424c-89bf-43f5dcb8f4a5 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling Phantom: Subject-consistent video generation via cross-modal alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:05.537673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:05.537673Z digest=sha256:4122ba5ae4854cfb6bdec205e025eab096d670683da181c47d57a39e227ced54

Observation ecf247e4-33a2-4343-bb22-f847989a56a7 · inbound

A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks cites this paper.

A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks Phantom: Subject-consistent video generation via cross-modal alignment

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:45:48.892408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-30T00:59:50.478584Z digest=sha256:813d523b3b0d421ac48a6fc8aa6ea0df27cc2b9ac71673caea015cd7a2035d99

Observation 95c4ed9a-fded-45a9-8f3e-332d66048293 · inbound

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation cites this paper.

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T16:30:02.172800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:30:02.172800Z digest=sha256:b3e90f5256ef528ac7d465055f9d55c86e68d6d39734e2551584d2a2cf926d77

Observation d111140e-0430-4dbc-a018-a4527dfbfc31 · inbound

ID-V2V: Identity-Preserving Video Restylization cites this paper.

ID-V2V: Identity-Preserving Video Restylization Phantom: Subject-consistent video generation via cross-modal alignment

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-01T04:27:31.566712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:27:31.566712Z digest=sha256:85fb09bf0a1ed3ca65791ada6808b02bd97ff4f486b21f24d924ca6d09a15cdb