Pith. sign in

Paper Citation Record · LEDGER

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

As of 11 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 25 inbound Pith citation observations for arXiv:2509.08519.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08519 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:35:30.373373Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:04:14.380575Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.676485Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f6e52b1-42c4-4c68-8d10-025ce616f829 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Qwen2.5-vl technical report, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.233495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.233495Z digest=sha256:fdd7d5780428561febe9d1a28662594a8bd9c9533957b1d12bbb62a5eaff5119

Observation 6cff7969-361a-4ced-bf5a-e069987fa478 · outbound

This paper cites Goku: Flow Based Video Generative Foundation Models.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Goku: Flow Based Video Generative Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.238979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.238979Z digest=sha256:3210f1eb696c922d6cf318a3d3e900dae9af021fac282ca75ff30cf4ade8e143

Observation c7a744eb-4b69-4c52-a723-e40767a1831a · outbound

This paper cites Phantom-data : Towards a general subject-consistent video generation dataset, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Phantom-data : Towards a general subject-consistent video generation dataset, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.244265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.244265Z digest=sha256:ce5225afbbac0f859e267720189b959685e1487fc580b4ed298d49d497d5a85c

Observation 01de29e7-9c07-4e5c-9b3b-f2d0a4f64690 · outbound

This paper cites Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.248586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.248586Z digest=sha256:cce35b3c9df2cd81c866459a194c21fe54f7edb27f70f36f8b202a54b073a792

Observation 73526c4d-fef5-4337-a088-f151684a4563 · outbound

This paper cites Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.252509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.252509Z digest=sha256:0efcd0e7af3ae51640abe201170ef6a274127e73bba1c6ffbb3999d724c88825

Observation 1982d1ed-fb0e-4aa2-a0c1-7a6d565eae51 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, page 1–1, 2021.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Arcface: Additive angular margin loss for deep face recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, page 1–1, 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.256166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.256166Z digest=sha256:d43daddce7ffcbe8c5d40637dc00383ccde77f75f897abeb1d9304b31d41c156

Observation e04f3538-3bc5-484b-ac68-160620b20b12 · outbound

This paper cites Magref: Masked guidance for any-reference video generation, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Magref: Masked guidance for any-reference video generation, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.259941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.259941Z digest=sha256:a4111de439839f52d6c85540f1a47d4eb53d347636d31afe10f98257bbdf307e

Observation 44100d2c-905c-4bed-8b66-757f6c16f9e7 · outbound

This paper cites Seedream 3.0 technical report, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Seedream 3.0 technical report, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.264055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.264055Z digest=sha256:0d9053f67c378c3771e548a813617fd2725e71ba9e4dc14922a965f9b2f7ae8e

Observation 74c364d4-1825-4b38-87bc-97a4f1befbf4 · outbound

This paper cites Seedance 1.0: Exploring the Boundaries of Video Generation Models.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Seedance 1.0: Exploring the Boundaries of Video Generation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.268006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.268006Z digest=sha256:c2ab4fed3aac1718a589590f4e5ceeddb2a4c576282f7d039e11dd1bd5c6da7c

Observation 5c884a46-d3b8-45aa-acca-7cf61f7e5515 · outbound

This paper cites ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.271855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.271855Z digest=sha256:08c2c789688a021f2030bb1e43bab90bbbd7f4fb7ed3c4ff2038e458382feb90

Observation 3b99b558-d953-4af1-bf04-df27b6da5e97 · outbound

This paper cites Hunyuancustom: A multimodal-driven architecture for customized video generation, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Hunyuancustom: A multimodal-driven architecture for customized video generation, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.275819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.275819Z digest=sha256:7ad178bfdcc29a9d359b8e814fc1548ea4cdb9c81e2afe7b6f63ef3800d4579b

Observation 863638f1-8ad9-4be2-a9a8-6b58c0f34f5c · outbound

This paper cites Curricularface: Adaptive curriculum learning loss for deep face recognition.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Curricularface: Adaptive curriculum learning loss for deep face recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.279526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.279526Z digest=sha256:45cff2c351bfedebb6e55eb449cdbd4fdc71d6c65053f768cefcecce31dc3637

Observation c5c714b8-89c6-4e78-8a3d-a5fdac551aec · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Vbench: Comprehensive benchmark suite for video generative models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.283029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.283029Z digest=sha256:14b820993ed2ee5e59b7321e5fccd7422c7a8b2f97bcf340a50ea6e24de9d6e4

Observation fdbdf8e8-125d-48c5-959f-175d01f7c997 · outbound

This paper cites Multi-reference images to video generation feature.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Multi-reference images to video generation feature

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.286812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.286812Z digest=sha256:1e9be27184928c6e6589d4163b390028a1b071f247d25b7e7845a247986239de

Observation e02591c5-48db-4050-9bf1-7b32005ecb64 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.290001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.290001Z digest=sha256:78d12f065946cfea8ee57d3d4ff35b0907dc52f55760561e13f92c204bd85e5d

Observation e85d97a2-7601-46d2-a88a-5a5e715ee43d · outbound

This paper cites Let them talk: Audio-driven multi-person conversational video generation, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Let them talk: Audio-driven multi-person conversational video generation, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.293710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.293710Z digest=sha256:8e0d0ceb3b6fa5402d1cb5b2702b01b8f8410d06bb93caadc7dfd5c6d0c5df9b

Observation e8056d4d-a20d-4e3a-82c0-26d76dd245ec · outbound

This paper cites Latentsync: Taming audio-conditioned latent diffusion models for lip sync with syncnet supervision,.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Latentsync: Taming audio-conditioned latent diffusion models for lip sync with syncnet supervision,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.297026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.297026Z digest=sha256:4b86abbebc02cb9c10e48867b89c44c92e92234043508bcbc1eb2937389ee590

Observation 0fe98423-05f1-4ca0-ad34-4ad69779c0d0 · outbound

This paper cites Openhumanvid: A large-scale high-quality dataset for enhancing human-centric video generation.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Openhumanvid: A large-scale high-quality dataset for enhancing human-centric video generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.304093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.304093Z digest=sha256:fc176a938a5d5c59ab5c779b196524d4139f7d06f8ddad9531700844450b969b

Observation 6d9df52a-35b1-4f13-a357-79421bbbcf0b · outbound

This paper cites Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.307584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.307584Z digest=sha256:c703b1289a30c168166e5974fdbc081f08214f2a74361b58bd474181b8945da1

Observation 54b180c7-8a38-4f89-8ad7-61606f1148ea · outbound

This paper cites an unresolved cited work.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.311182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.311182Z digest=sha256:ace10b448a09e21f0caae08be206fc965f33cc38505005397a610ceaadbaa698

Observation 19d52d6e-60b5-4514-95eb-4cfa52ede0b2 · outbound

This paper cites Improving video generation with human feedback, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Improving video generation with human feedback, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.314780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.314780Z digest=sha256:ced2dac48973117f10f2a33ef77557d2e64a089bf86a244cc50c6d23de0c53cc

Observation 89dac79d-2b20-4695-8017-7d86e6f3cecd · outbound

This paper cites Phantom: Subject-consistent video generation via cross-modal alignment.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Phantom: Subject-consistent video generation via cross-modal alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.318096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.318096Z digest=sha256:0079d0ed3512cf9906f08dd576b6582821ac2ccb8f1f5aa71f14f2e7c26882fc

Observation 8edc8a2e-5f5e-405f-83ba-f097b569c2e0 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.321358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.321358Z digest=sha256:85b3ea3412a48fbea080a1c8b27bab2a95b4e21cd2682247cf1632e161bd65ac

Observation dcca47f5-00a3-4da8-9eea-bfe8a32261de · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Movie Gen: A Cast of Media Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.324506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.324506Z digest=sha256:8487c58570337faa0d8a05e0febabbb43040d956777f7cd2708b87712b4b9282

Observation dd519e46-b700-4561-93c2-7c6a130b742e · outbound

This paper cites Learning transferable visual models from natural language supervision.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Learning transferable visual models from natural language supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.328222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.328222Z digest=sha256:2d65d13bfbf270ee34d0e1c37f4a65d86cb85f8efbf794ff19b702a6ff60966b

Observation e35adaba-2d5d-45de-a961-6add66bae3d7 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Robust speech recognition via large-scale weak supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.331572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.331572Z digest=sha256:d37b4b39c1bc0553fdd596114af155356c5bf7d6e0b0e56e0754f7be0eed56cb

Observation 1342596d-7656-46e5-9770-fa20a0badf5b · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.335032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.335032Z digest=sha256:ecc205b5de024c409f970ddb6b9fb49609f5e523685bbb0e76ab1b672200996d

Observation 404e1143-193f-435e-a946-c7821ae3966b · outbound

This paper cites Seitz, and Ira Kemelmacher-Shlizerman.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Seitz, and Ira Kemelmacher-Shlizerman

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.338475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.338475Z digest=sha256:4567a5a29618bbf2a267e33c42d33d601b98b4cfe4f19f9ee6947ba737417005

Observation 66a12392-8187-48c7-99b0-49621ece1022 · outbound

This paper cites Gemini 2.5 flash image.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Gemini 2.5 flash image

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.341527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.341527Z digest=sha256:75591124e858ff91682c7c67ee4ad8e966890cb74d4e2e92d9bd62e9df61c246

Observation 5015b6f4-0fdf-4cc8-99f7-331a36e5d40e · outbound

This paper cites Gemini 2.5: Our most intelligent ai model.https://blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Gemini 2.5: Our most intelligent ai model.https://blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.344720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.344720Z digest=sha256:0e0fe82514a53710fd356b055b9d8b1c3fe4cc7463da352c0137c67f05b77122

Observation 6761de8e-cee4-4ef1-8e5c-626d1aece653 · outbound

This paper cites Wan: Open and advanced large-scale video generative models, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Wan: Open and advanced large-scale video generative models, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.348072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.348072Z digest=sha256:da8f4896623af164ffb3ac96348e041d4bbec9e0feddab284a2a6224963fbc9b

Observation 535abdab-5554-4da2-b606-24ca809d2cbc · outbound

This paper cites Fantasytalking: Realistic talking portrait generation via coherent motion synthesis.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Fantasytalking: Realistic talking portrait generation via coherent motion synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.352280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.352280Z digest=sha256:8b1674df03422fa34aba247c02f2bcf7d552d46204fe0c07936353ea9cbb6725

Observation e298fb75-861e-4e5f-b1d5-32b3bd102811 · outbound

This paper cites Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.355742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.355742Z digest=sha256:ca03c6bb869dd86b5df4369325f955b90367dece3de407dc4224dc5d74ed2c86

Observation b517a61d-d8bc-4ca9-ae08-daba4705f3d7 · outbound

This paper cites Interacthuman: Multi-concept human animation with layout-aligned audio conditions, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Interacthuman: Multi-concept human animation with layout-aligned audio conditions, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.359072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.359072Z digest=sha256:c6ab15bb98599d539d7f323b0da8986c959b4e8127907458fbc77b28403a9bfd

Observation ca611332-8519-4922-aa86-96592f562257 · outbound

This paper cites Mocha: Towards movie-grade talking character synthesis, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Mocha: Towards movie-grade talking character synthesis, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.362761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.362761Z digest=sha256:d75801f0ebcddec50d94caae9527dd783168bd862cce1f561d59ee60b078a0ce

Observation 5d0c4ebe-897b-41e1-9b3f-afff7b7f7a52 · outbound

This paper cites Magicinfinite: Generating infinite talking videos with your words and voice, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Magicinfinite: Generating infinite talking videos with your words and voice, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.366342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.366342Z digest=sha256:d3ae65e7a6689ec170b79b2c28d80c40039e011f00b96ca288a2ee6c78acf033

Observation 200ba8ae-e51a-4cd2-b076-cb23f14777c6 · outbound

This paper cites Opens2v-nexus: A detailed benchmark and million-scale dataset for subject-to-video generation, 2025.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Opens2v-nexus: A detailed benchmark and million-scale dataset for subject-to-video generation, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.369765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.369765Z digest=sha256:5b1a1522b0090d884ad7b11a811f017a264b85a570baca38b1d0b624fb96394e

Observation 437606c5-bc8c-40c1-ac7f-c34dc4d3a8af · outbound

This paper cites Identity- preserving text-to-video generation by frequency decomposition.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Identity- preserving text-to-video generation by frequency decomposition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.373373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.373373Z digest=sha256:eaa7611be8a999a6cb2a7f907033b817deadaf1e58039706de07b6bdac894383

Observation d05a74d5-491e-44ba-a394-a08056f9ec64 · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.300454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.300454Z digest=sha256:791ebefec12a59cd7008e884dd0669ebb7a4d17e48c1aa95e655eef7bb8f0a3f

Pith citing papers

Observation cc157498-875d-48b4-a842-beae10033b8d · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:03:15.211260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:7735e8aeef68f5fed0eefe0967d5fcee9c6372d30cf177d4e7e46724eac080cd

Observation b2177093-0791-4c79-9f92-221d3a59a587 · inbound

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection cites this paper.

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:48:58.476215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T03:48:26.807495Z digest=sha256:85b3fe01ad8949749f728ae0154919edfa65d4c506f31eba1d57ac2bfee28789

Observation d19122a2-98f1-43bd-a2c9-84765f842cc7 · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.907924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T19:43:37.604351Z digest=sha256:2d5decfa878d08be941e353f8c997ffd2d6028fbb2833f34295a2648a3d3cb4e

Observation f09ff69a-813c-46be-8db5-611a115730d0 · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-21T16:14:15.230425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T16:10:31.015783Z digest=sha256:574f1a76b411b3baa8f6dd477aef95d144f985f3696fb0b8716559068a1dca56

Observation c43748d0-94d2-4373-8142-ea99c47050ad · inbound

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos cites this paper.

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:47:57.511227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:43:26.460480Z digest=sha256:864e3563c9a23e9cec95fe4da9090b775aabe833c6c690f470c9351cfa5e8212

Observation 7617bc1d-d26a-4f65-90f6-a0409ef9d3ad · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:10.037400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:10.037400Z digest=sha256:bdcc6e06b429caa46cc70515bcc1cacf5653832465b220ca8aa5c153fbc8955e

Observation 64187c61-3b4e-499a-ae99-90f38bc8c8d4 · inbound

What if? Emulative Simulation with World Models for Situated Reasoning cites this paper.

What if? Emulative Simulation with World Models for Situated Reasoning HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-15T13:51:30.008232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:51:30.008232Z digest=sha256:cf680e78e07720cbca81f3c94648974144b0cd7770fb26c4bdfa57fbec6d92e9

Observation e55e8956-b8f0-4f28-8a80-3d82ee835944 · inbound

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model cites this paper.

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:18.110781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:51:18.110781Z digest=sha256:1aa8176a067e9f479a0000839889d9db170a72e07de3a4d3629565bfb7a1dcfd

Observation c0aa2cf3-d9d9-4912-b478-6bb81d717bd9 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.846071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:a2cf17a0fe3e975c6aee242ec9e074a24c4d612bd277d25ed5c8f7f0525aaa62

Observation 3fbbb5f7-a58b-429f-a599-051e32a9f910 · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:41:04.305402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:660474d5290d8a97754de1ab2407b04e5a67e7245b9a6b30be36dda53e9b9e26

Observation 6a0795f9-e1a4-4d76-8c84-64a7468ddddd · inbound

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis cites this paper.

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:02.046272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T02:24:57.882447Z digest=sha256:102eda3ec271ead254c95f84adf7ebb8abca3b95d556035e260bd259ba683c20

Observation 4c170827-95c6-4e97-848b-14e456023853 · inbound

Generate Your Talking Avatar from Video Reference cites this paper.

Generate Your Talking Avatar from Video Reference HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:29.856032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T05:32:04.519820Z digest=sha256:fc0632956153db9892e0b435c6ae93e93b1d6ed2d02c74467199c9eb49bb4ecd

Observation 43191b27-afa2-4dd0-ab05-1b3066bfbcb7 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.416194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:b4024def9ac1c8cc58ed16d1cdcd9741d80c2d5b0fd6e242c65df017fe936e6c

Observation c6630e7a-2ec7-401d-a29e-2a57d085da97 · inbound

Aurora: Unified Video Editing with a Tool-Using Agent cites this paper.

Aurora: Unified Video Editing with a Tool-Using Agent HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:48:12.712658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T10:47:05.038308Z digest=sha256:23348aebdd719cb60c8fb5cb0617cacc33a37389f8e06ae6052b1cc5fdc7e2b2

Observation ecb2003d-3b8c-4c9f-bf19-2a13a20370de · inbound

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models cites this paper.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:44:41.246048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:93ffb0a45eac69cdfaab3f8400d6f6e2a8ef4f380f04e39d7b2df02650441a0e

Observation b3cb5746-2fbf-4d8a-9bfd-8afa6b38967f · inbound

LongCat-Video-Avatar 1.5 Technical Report cites this paper.

LongCat-Video-Avatar 1.5 Technical Report HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.636837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T18:28:16.968339Z digest=sha256:18247d806b9ed3ea2cf8c12dee40b7ace3716e4b6a30a152239c3b30d62c4248

Observation b2f10b8f-5a9c-4371-9d50-d5d54890b41c · inbound

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation cites this paper.

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.473972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T15:25:22.778550Z digest=sha256:6fe72f572b8aa7fbb0dcfc3c8da81c8fce3d16078a4cb7253501fd740fd0f2ae

Observation 23737e46-c8ca-483b-a636-1a1101528f32 · inbound

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation cites this paper.

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:36.651576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T13:49:50.272650Z digest=sha256:0c50a7d26f7cf85b900f27fc95c278331d2f75aab3d72019580f36bae7732b59

Observation f6df4e3b-1a50-4e2e-a38f-4b7da48f6258 · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.678040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:666012f9231af7d251275d45d6c14424f9f0727738ad689063044f88022adc8e

Observation 8f0d2e04-551f-41c2-9952-bb51c793bb71 · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:b8747feda8c472d726679acdd7ed1bbd9aa81f84b61c9010bfd739a36457e5f0

Observation a7effad9-82da-4ba1-ae87-441f648ab486 · inbound

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation cites this paper.

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T10:37:57.061257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:37:57.061257Z digest=sha256:7b843a598b02c0796907e9af8679979bab82465488d18f992825307a6b947672

Observation 65597ed5-91f9-48d8-8830-5a7d2d18c6a9 · inbound

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment cites this paper.

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T05:27:15.266478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:27:15.266478Z digest=sha256:3a02419061b086f6780c52ff77d5436882d4cdf4f70afd92c16c6eef0939ccb6

Observation 8917a336-66f4-4a14-b688-44e063684b20 · inbound

ID-V2V: Identity-Preserving Video Restylization cites this paper.

ID-V2V: Identity-Preserving Video Restylization HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-01T04:27:31.607132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:27:31.607132Z digest=sha256:cffb7a407b0ef437bc796afb9ba233d5e45060c38e069cbf4bd8be2028b4f9f5

Observation 99433b28-261f-4369-903d-a30387fef95f · inbound

Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction cites this paper.

Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-30T20:31:38.681716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T20:31:38.681716Z digest=sha256:d9506d5b4c664b1cb085d1ec43ef077d4654b926323fff5b6320ff50f86817dc

Observation ff8d26ba-0435-4682-9be0-1eaf459188d7 · inbound

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing cites this paper.

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T13:04:14.380575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:04:14.380575Z digest=sha256:1ef7e1b8186e73078a228e15fa91cd23de3e1da345fcba2ba9bccbfa9da55d21