Pith. sign in

Paper Citation Record · LEDGER

CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2503.10391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.10391 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:29.343261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.682483Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7344fd96-2a37-4a56-9bff-dc864b12df8a · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.343261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.343261Z digest=sha256:acf66465529e0ca9112331bec15593755ba2c64611ab47990f2bfb0d4264ad07

Observation 10bbdf14-7a25-4c48-bb27-3789b2a20813 · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:08.227600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:08.227600Z digest=sha256:e2972c63738e4c6ac548bda9570001c287b99dee43afd874e05e297908be4195

Observation aa4e6714-16c5-4dc3-9c33-4edf724954cd · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:36.010543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:36.010543Z digest=sha256:eb542c86180c081427da24ae8e69493515bff9558a63fe2cb2bf87c5d8d6c56a

Observation 3e6fad62-7218-485c-ae16-187812f6a501 · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:08.256451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:08.256451Z digest=sha256:cefb818fbfdac994467155840ec0de0c3665d930c7623edd543f0d71a0aa7159

Observation 6a475883-7016-4dfc-a2ec-e9afb0e3bc01 · inbound

HOComp: Interaction-Aware Human-Object Composition cites this paper.

HOComp: Interaction-Aware Human-Object Composition CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:09:51.852399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:09:51.852399Z digest=sha256:ee14d4e50191af17d89c01fa1c12f9e213d232ae00c971a98746fb3e5e28e2ea

Observation 6b52e863-a64e-45b6-8a0b-47dda76c78a3 · inbound

RefAlign: Representation Alignment for Reference-to-Video Generation cites this paper.

RefAlign: Representation Alignment for Reference-to-Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T18:01:19.034578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:01:19.034578Z digest=sha256:0ee97f9f21d5689e198247fd05ac9f0bbe35e267e2ba9580a0628e80913ee617

Observation a6452147-1bd9-46b1-b879-ee5b15d83f5b · inbound

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation cites this paper.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.912925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:242a34bb688bbff8fbaaae8846568521cee879e498a029dc6a8fc1eef347d62a

Observation 54963091-81c4-4cea-8fc1-ebc352845893 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.977015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T05:53:21.851578Z digest=sha256:c04e2d24fad8cf4dfd77f6ce547c52e9156c8bc58436ce1eea02dae3e6cc5d16

Observation b80fc3bc-d5c4-4034-9335-58cd8047c016 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:03:03.343949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T22:00:01.349754Z digest=sha256:6ec4fddfe77909301090bd9e6e6591c8dbbd6e1a14e57fa6a1ec5d6906e13333

Observation 38fef8f3-95d7-4a33-bf2c-af5c51c038ba · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.976977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:6d7d4c04ba5bff91bf2152ecc0324e3c672e0fe74e9db61b8c6314188daf5021

Observation 3d29057a-593a-41ea-925b-47a9414aae75 · inbound

Customizing Video Portraits via Identity-ActionDecoupling cites this paper.

Customizing Video Portraits via Identity-ActionDecoupling CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:42.380535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T10:56:11.553110Z digest=sha256:817a6e3c7679f546ceb96023f45f81ed2a22e173d45b20bfacb794bf57c6cb33

Observation dcce61f5-bf18-4c2d-9ad4-326d741f59bf · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.684393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:f9c3491b99690d2aded84e339bc3517d053ac0f6e00efc4cbf8779e43b29fe31

Observation 75e11631-dd5b-46f4-bfa7-df8687893a99 · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:35df8ac7480ec3bc9b93e4c826cfbd55d41ce3896422705dbad8956292d747f2

Observation 4a50d647-fe85-45be-b497-2b84e4849794 · inbound

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement cites this paper.

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:19.025878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:19.025878Z digest=sha256:c2fb58a0e8867a8e2357688a9c4ed29f2d1c4f1babc307027b355d637a773ef5

Observation c2ea64d7-54ec-4967-92ac-74862def1255 · inbound

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry cites this paper.

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:04.839943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:04.839943Z digest=sha256:d0dd7e0092cc92312a609815baf81058cd1c4c2c726e4c11c38f5b4b90f06ce4