Pith. sign in

Paper Citation Record · LEDGER

CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2503.10391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.10391 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:29.343261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.682483Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7344fd96-2a37-4a56-9bff-dc864b12df8a · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.343261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.343261Z digest=sha256:faf790b2ea6948d1701d60901d5ebae52a5593095ddd8cf63938536a769189ef

Observation 10bbdf14-7a25-4c48-bb27-3789b2a20813 · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:08.227600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:08.227600Z digest=sha256:1ea4fc0c2c0d7f4c0a84572bd8c28399b6621d8194f54f04b5227483f8a6a26b

Observation aa4e6714-16c5-4dc3-9c33-4edf724954cd · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:36.010543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:36.010543Z digest=sha256:64f603f98a25835cdc3882b1750d1de72e23258e05d228eed31b2f7c14925f36

Observation 3e6fad62-7218-485c-ae16-187812f6a501 · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:08.256451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:08.256451Z digest=sha256:422666a64cf8b086d045279eeadfb58c6e3915babf285a1d5ea0736a67237a0d

Observation 6a475883-7016-4dfc-a2ec-e9afb0e3bc01 · inbound

HOComp: Interaction-Aware Human-Object Composition cites this paper.

HOComp: Interaction-Aware Human-Object Composition CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:09:51.852399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:09:51.852399Z digest=sha256:e6612dac0748bbdc9d805d9e3d888e128facc466c33b49ce80aa89e7292b1d99

Observation 6b52e863-a64e-45b6-8a0b-47dda76c78a3 · inbound

RefAlign: Representation Alignment for Reference-to-Video Generation cites this paper.

RefAlign: Representation Alignment for Reference-to-Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T18:01:19.034578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:01:19.034578Z digest=sha256:d0aa1f9b096f6e662d377314171610ed1e8e74c19a75e8b448a0b485171347ca

Observation a6452147-1bd9-46b1-b879-ee5b15d83f5b · inbound

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation cites this paper.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.912925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:e54167e0a64a007a2c2b869b8788a7791d6c1f4013afd820c5980d5360d1ce75

Observation 54963091-81c4-4cea-8fc1-ebc352845893 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.977015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:53:21.851578Z digest=sha256:46e560aec45678f27b7ae6b802718616910c333cf8d289283bddc0b35e22a809

Observation b80fc3bc-d5c4-4034-9335-58cd8047c016 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:03:03.343949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T22:00:01.349754Z digest=sha256:48f8619ea0efca4cbc2fab22d947d981d99f0b41532f71a811856ea89a879efb

Observation 38fef8f3-95d7-4a33-bf2c-af5c51c038ba · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.976977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:d020adf17a9186287355fa8ca3f97dbbc7a13863f9cd0d818d06973facf0c3c7

Observation 3d29057a-593a-41ea-925b-47a9414aae75 · inbound

Customizing Video Portraits via Identity-ActionDecoupling cites this paper.

Customizing Video Portraits via Identity-ActionDecoupling CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:42.380535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T10:56:11.553110Z digest=sha256:f815c323313e496d57b44c9f2d8284b61c92122aeff84a0c5a3cfbe93ae986e9

Observation dcce61f5-bf18-4c2d-9ad4-326d741f59bf · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.684393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:d9b6118c8be28fb9f4aa53caa4b3be1aea790768f77f9248e5c267e703e0ed11

Observation 75e11631-dd5b-46f4-bfa7-df8687893a99 · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:42c5d470af6bd2842383e3c61fe0d7c05c23c03d66f318f724b8e74f8a8c14f3

Observation 4a50d647-fe85-45be-b497-2b84e4849794 · inbound

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement cites this paper.

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:19.025878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:19.025878Z digest=sha256:7ef1357c5345528c33fc65eede97d60d78f02c5129a499682ff1a63d5cac18fd

Observation c2ea64d7-54ec-4967-92ac-74862def1255 · inbound

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry cites this paper.

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:04.839943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:04.839943Z digest=sha256:a877d1c4b8cdcc11f911226b49fd612cc120c727efdb7309b7f75aff26c278ae