Pith. sign in

Paper Citation Record · LEDGER

HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2505.04512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04512 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:16:44.665393Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.669744Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8dd27455-0272-48ec-a382-50a76bb3bf6f · inbound

Hunyuan-Game: Industrial-grade Intelligent Game Creation Model cites this paper.

Hunyuan-Game: Industrial-grade Intelligent Game Creation Model HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:30.549695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:30.549695Z digest=sha256:8636fc31d3131fb58dd5584226686c0e03b288fdaf51399ceafeea6e5d66fed0

Observation 63e9d5a2-120e-4bf2-8687-2cb8e04ecf70 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.485928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.485928Z digest=sha256:b7375dcd4d4e8054bec5c87fbe95a05921baf7ba45e3c65ca72980e0765811f6

Observation abfbb6ad-57e0-46c4-8e59-66bedea81e69 · inbound

OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation cites this paper.

OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:51.666724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:51.666724Z digest=sha256:8508f3863ad8a79f79bba8e5e265ea54542efb3c4492b2823f8ed6c3dbaa865b

Observation 788063b8-c336-4922-8764-25d8cbfc7f70 · inbound

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement cites this paper.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.934408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.934408Z digest=sha256:496eb85992a6280f2f0b1afc0cb57586b87f35b90ae28568a9f4445aef16677a

Observation 074be566-ef03-4cca-86e2-8188ae16870a · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:37.363932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:37.363932Z digest=sha256:c162a46e84527e184e19c6eb73f30aa7f650ab7d9c83beead1e23f3c1ab71dca

Observation 4fe3aad9-1022-4d98-a760-24ab29b51e31 · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:09.055319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:09.055319Z digest=sha256:aae5c27e8887aab644d7cadda8ab11c6b5a96a6a17b47b47101723c4696fd668

Observation b641c452-348b-4b38-88ec-e82f018c0546 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:26.356204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:26.356204Z digest=sha256:e56835cf0abf3e1b4ddf715cdb191c7e1e504f671db21a75a54fe507259ed9f9

Observation 0a61444a-cf7f-48d6-94ed-a824c5c1e72d · inbound

DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing cites this paper.

DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T18:34:25.012542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:34:25.012542Z digest=sha256:4762741d7cc9c323689f56e72c8edf1989ca974568d6671c78ada3daca6fbbca

Observation 812b033c-b4c9-46d7-a44b-6d9dd7a06258 · inbound

InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion cites this paper.

InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T15:20:18.764867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:20:18.764867Z digest=sha256:3f2a5658a31b4f90b90b8cd848cf3e8aac8097ce13b8cc6ff8008a89f63be92c

Observation e5bf8773-d8bb-4ee5-9162-fd11c56f9371 · inbound

CustomX: Unified Character, Action, and Scene Customization in Video World Models cites this paper.

CustomX: Unified Character, Action, and Scene Customization in Video World Models HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:52.239638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:52.239638Z digest=sha256:6c6390a4cd4ac144f5dce2a6eab82e4921ca8517b942d41fd28986ecdc818fd4

Observation 6ee46033-c95f-4400-813a-8b67dfa1477b · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:12.422942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:12.422942Z digest=sha256:3459a384e2b1fdfc950d70cfe52fb6a1fc7523721ad388da75a8010cf57831d1

Observation c97ce011-c67f-4cbf-a1d6-0403d5d9fecc · inbound

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model cites this paper.

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:18.169044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:51:18.169044Z digest=sha256:a2910f9878817f7ab021fdf568d86e3e702f086728e77b19d5d381378bacf752

Observation d24af3bf-84ac-43b0-ae30-3b1dd448ad90 · inbound

RefAlign: Representation Alignment for Reference-to-Video Generation cites this paper.

RefAlign: Representation Alignment for Reference-to-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T18:01:19.034578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:01:19.034578Z digest=sha256:e533cc29bc018459e773989793330075b266fa7cce3635e4eaadf0eb8efa66ed

Observation e4ce79ab-ccc8-420b-b23d-c466bd122994 · inbound

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation cites this paper.

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T12:23:05.876881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:23:05.876881Z digest=sha256:0c6d856039674acadc815427e73ff9aa60e30db8deb9ad3eb06a1856c63cc6b8

Observation 0163d288-73d9-4b31-97ba-08b701df773b · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 204

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.753975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:588ecf569c5e9c57e592220327f7898c628faa9cb3a1ffc52c7d05a8b6b0b0f0

Observation 04fe287e-dfd8-4469-bcbb-8a9e6bc57b04 · inbound

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation cites this paper.

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:01:00.395392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:42:35.693519Z digest=sha256:8384fa424802a15d0adf47235dcd4f7dc7e6d1b8b000f52ef12e8d99e575298f

Observation f08526f3-b1aa-49ac-a63c-5947baa19432 · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:03.662032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:8c38a904b19eb92e024b5d155c40ddd66c1fa1ed71bdc52f87a1db062c3b7eff

Observation c1d8d2a0-a7d1-48f9-9bb7-bff428615352 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.806138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:9616c226d31645c73de535b37eea73d4e76869d9207562b5a29663dd837281b2

Observation 952199a1-ebad-4157-a3ce-83405553fdcc · inbound

Controllable Video Object Insertion via Multiview Priors cites this paper.

Controllable Video Object Insertion via Multiview Priors HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:45:20.706056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:44:17.033051Z digest=sha256:ff882df1d428f0b91e48977a423d566a462a30161f1511cdaf2292b971fcf9ee

Observation 65333d83-84de-4bec-87d3-da71d12a5d45 · inbound

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation cites this paper.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:53:29.903427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:afe5247239268db60f77fd916fc85416ec1d9f4dfa9c0d1a32b57ae3008b16ac

Observation eacb0a72-0353-47e2-931e-c83cd4fbc94b · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:28.018405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:263ace49b2dbc682ddb20b9e50151243a171954714165e68e7c140293e8a618f

Observation f3ac5160-9ad1-4dfb-b028-b606168c8028 · inbound

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation cites this paper.

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:09.751609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:40:00.224358Z digest=sha256:73f53ea72456bbd53b91b52c42f82818694dbc5c8b364a7d19697d8c0d28c43d

Observation d89586ac-4a70-4d6b-8094-c6a6994a90a8 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.937729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:53:21.851578Z digest=sha256:2baa9a08ae006498d1319a1f8c223414a0214d11fc80bf2c1675468405bf8689

Observation 586ba538-ab1b-4b26-8fe9-2a6b78fff3c1 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:03:03.463171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T22:00:01.349754Z digest=sha256:abc64cc993c6fbe17fd984c5095d73baf77c8d701f5665390f3f2cf72f383539

Observation 0f54d741-72e8-4725-929e-dbc3f06fb308 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.373265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:04b86be08fe0bd8381c83fe6c0f7433686761bf7b10bfbb111e3b466d7a7bbe6

Observation 2af97305-875a-456f-9aab-b10078e1f81f · inbound

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation cites this paper.

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.501082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T15:25:22.778550Z digest=sha256:23ec26989405290745ef2e3f5696e420df9b9f0ec183a103d2add4734e577d0b

Observation a795f952-7525-4266-a9d2-855d4aa99514 · inbound

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data cites this paper.

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.256912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:52:30.406683Z digest=sha256:9f6f347394651926b75861d94ad8703df320133178eab725c16a928fc278b5db

Observation e690e724-3c34-452d-9f64-4b51f8002a07 · inbound

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation cites this paper.

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:07:28.122494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:30:25.371658Z digest=sha256:d30b658f4aae6155b61f4c9e6154de3b71ef10f5ce6f1120f2f93d80ca492e1e

Observation 8d18100d-2885-40e1-892b-3bdca5f5cedb · inbound

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation cites this paper.

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:36.656864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:49:50.272650Z digest=sha256:dfebd07e812e66a5bd1b36fcad226433976c8a3b4442d8a4cd0cd3932191467f

Observation 0c8b23d7-dc6d-456c-9fb3-b8ba38c205fc · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.961436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:985e6b9ec68bffc207d2bba0b680bfd784362df1b03575c4f205a9814b57e430

Observation 7f39caec-ed05-4f26-aee3-ac525f169cd6 · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.671686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:5e94297266f9fcc205f6f440687615ef719063a1c60ca0bc4126e21deba2ddda

Observation bf6267c4-094e-4ab3-9066-b21b8598b93f · inbound

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models cites this paper.

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.123698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T13:16:16.676244Z digest=sha256:3e0b8691c669952bb10e7a8b67e1ce9bbd3b91533dcbe921f5dd12183d99823b

Observation fd6a0b46-6194-4af3-baa6-6848e97a5470 · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:5d86447036f4ab812df6dddf3c5a93c8c205bb2e0d433af5c6b49b957826487e

Observation 63374138-3231-45c3-9427-fbcf3959b7d1 · inbound

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation cites this paper.

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T16:30:01.157173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:30:01.157173Z digest=sha256:f26c336a82db3b0190da3650fe3041866798c8907d5576a2da73b6d39d68f410

Observation 18cd2add-2d16-4bc1-8d33-4e7a7b66130b · inbound

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement cites this paper.

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:20.695859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:20.695859Z digest=sha256:cb8e10326cc08e97b164669bcec096cb73e28e81b8fcde8f39883f3337c9eb1e

Observation f34af647-a454-47b8-99ea-6fd7f628fb95 · inbound

UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation cites this paper.

UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T17:55:19.705417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:55:19.705417Z digest=sha256:4863a58e5489b01d6bf00e8c38b4fe641b8087f5322404e6b8b37cceec33db0f

Observation f7d9a41d-ccf7-4740-93f2-eabb73556743 · inbound

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing cites this paper.

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T12:16:44.665393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:16:44.665393Z digest=sha256:d5dc4c7fa315379e062126de0991542ccec60a149a6f7ed6b6f5bd4eb526e72a

Observation 7b68b13c-0de6-45ee-807e-dc1472fd2363 · inbound

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation cites this paper.

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:58.654243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:58.654243Z digest=sha256:815ffcd9daf9943992c6f8f62d62d39df14b3de257c2161c0a94ae9baac4a8a9