Pith. sign in

Paper Citation Record · LEDGER

HunyuanVideo: A Systematic Framework For Large Video Generative Models

As of 5 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 100 inbound Pith citation observations for arXiv:2412.03603.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03603 v6

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T07:41:58.617477Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 593 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:32:30.841102Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact31
  • verified fuzzy67
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

6
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation f87e7832-d83c-4b53-811d-12b611e78139 · outbound

This paper cites GPT-4 Technical Report.

HunyuanVideo: A Systematic Framework For Large Video Generative Models GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.505184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:6d5e41a2c7a946d6db45c506447f7869164f7fc703d89ec7577576fbe34dbff8

Observation b3bff811-f25c-452b-b7c9-f6c95c4232ed · outbound

This paper cites PaLM 2 Technical Report.

HunyuanVideo: A Systematic Framework For Large Video Generative Models PaLM 2 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.529917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:346ea257ffbd284679b410a8e9a5e4eccb1b5b62d23d4bc8c2d7f2747f731540

Observation dc180955-82c4-43b4-99f8-f223cb660db5 · outbound

This paper cites All are worth words: A vit backbone for diffusion models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models All are worth words: A vit backbone for diffusion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.766053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:276ee52e4528c02b26f713f99d97d5aac3eb6367d0a363a427cecea9a154ffac

Observation 4c7ac835-9c54-4331-9e14-13c4d1d1969a · outbound

This paper cites Improving image generation with better captions.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Improving image generation with better captions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:47:43.151619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:08eaddd69022f0d59f9ab893d0a2bbb866757ec0ed7bb0d7bdb263e4adadc7c0

Observation 28483edd-24a5-435f-bdad-3934449300e5 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Stable video diffusion: Scaling latent video diffusion models to large datasets

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.752242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:72beb3fb3574ab7477ee46556813ff5bb67a2275bf96fc53e0e96567569fe824

Observation 2d1a36f8-6d60-4fbd-987a-25fb561eae33 · outbound

This paper cites Large Scale GAN Training for High Fidelity Natural Image Synthesis.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Large Scale GAN Training for High Fidelity Natural Image Synthesis

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.493380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:37e6b09685eac14aac5e5c23da55fe95e9790287de4d61378f5493d9874f183e

Observation d3cc1320-25b4-4966-8e61-2d9a63f93f80 · outbound

This paper cites Video generation models as world simulators.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Video generation models as world simulators

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.760996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:95c041a94d8790c8649ef5049fcc2e83380bbc2c599716fe18d8e77b4841c5f7

Observation 95d9111a-e91a-4aa0-ab45-5ff1235b8891 · outbound

This paper cites Language Models are Few-Shot Learners.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Language Models are Few-Shot Learners

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.600927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:b76e90feb265b15803dba9c4e8c0dd95a8032fbaade2fb459e36507918c6b2d5

Observation 0a7f01ee-53b4-4875-8d9e-c8fb5a13e918 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

HunyuanVideo: A Systematic Framework For Large Video Generative Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.543335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:9bc025f504ca381a861340173cfb24d7741659cc705a2c9834df9beb5fd6483c

Observation 111ec42d-5624-46bc-9f4d-494adccba1d7 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

HunyuanVideo: A Systematic Framework For Large Video Generative Models ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.549699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:6b64b49b5425c48148366937124652667d69c7332e009f20f126a22aca9e9b64

Observation 83549f5d-b81c-47a5-96ee-f2ad7e40d315 · outbound

This paper cites OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model.

HunyuanVideo: A Systematic Framework For Large Video Generative Models OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.415858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:a6c9aa8620a541e00a047b5a545c4009094f4679baee63fe39dc6052e525e1c1

Observation 4e3563ef-8bf6-4984-872d-62fc3c5b0134 · outbound

This paper cites Follow-Your-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Follow-Your-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.429546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:12d55c831745cbff72b514f3ff17568b2f1074ec0d2e9df91e5d4fd3d6debdb1

Observation d254dd77-5fb5-49a0-a84f-0edac3d8b9a0 · outbound

This paper cites Neural ordinary differential equations.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Neural ordinary differential equations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.735161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:174d17212fa88d0bbd4787ba2fa5891ab5356c6d671f4d872fc2fadaa42e5c31

Observation 31feb07e-70da-4df7-9e17-e7b5140ce30e · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.727799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:92ef91d74550a511fa15044ded7ad43964d655b694bd0ba9a7debbc7fe090702

Observation adc7da84-8c73-4ada-92ca-56f5a9e694e8 · outbound

This paper cites Seine: Short-to-long video diffusion model for generative transition and prediction.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Seine: Short-to-long video diffusion model for generative transition and prediction

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.731479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:752608e0cd80d275cc6a787c6cc810cafc539b364972c76c8fff0cd3ea4f0414

Observation 1c22e319-d0ba-4505-95bf-b858d41972fa · outbound

This paper cites EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions.

HunyuanVideo: A Systematic Framework For Large Video Generative Models EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.511917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:e1ff5c70c72e2765a4e207613a9dea4a0be5456e9782a76341dbb1a480f8a3f8

Observation 90944d98-7c00-4f48-b91c-70eb56d38e1e · outbound

This paper cites Xtuner: A toolkit for efficiently fine-tuning llm.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Xtuner: A toolkit for efficiently fine-tuning llm

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.738946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:5a7ad3c5290879d1d3f195d3a6b8fcf27c79dac121f340ed75e367f402e8f515

Observation 7de844a2-4a85-4229-ad96-51c2392558b1 · outbound

This paper cites an unresolved cited work.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-23T07:45:30.755853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:5305e2a167c56e2286414894e2091d73f11d910e6cf8a0ebebe76b09ac6935ad

Observation c0da432b-82e0-426a-be60-90e1840a1f7b · outbound

This paper cites Pyscenedetect.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Pyscenedetect

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.779625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:6f76633310500d9dc7041b96435f10910c8b61ab8abec1609f43dca6392e8b83

Observation ef3c6ed6-235b-405c-8a20-632aec7bb33a · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

HunyuanVideo: A Systematic Framework For Large Video Generative Models LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.525030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:8774a84cb080c00178a654d696813ecb2571e89a9d7ae76a953c25eaf7580a47

Observation ed61aa5e-c8cd-4a80-aa89-57007cb92c20 · outbound

This paper cites Scaling rectified flow transform- ers for high-resolution image synthesis.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Scaling rectified flow transform- ers for high-resolution image synthesis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.772333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:aaecd478d3336a8a5d3285d2f1b6fd9643f30054e156849b6a04b73414359142

Observation 83289cb7-8a21-47b6-b599-5400eef718a4 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Taming transformers for high-resolution image synthesis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:47:43.148772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:309bd7549a8411846941b681456035a8ee305e6fe7d252164648764f3ff006f8

Observation 1a8be5e9-fde3-4620-85d7-e12ba2b9effc · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.487031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:4a57418b6e6a62de210a2a12ecbbaa45cbd0c1602b689fe2ad0e07bf2a675a30

Observation 97818f2d-4076-4acc-b36b-43d9bfd8a2c6 · outbound

This paper cites YOLOX: Exceeding YOLO Series in 2021.

HunyuanVideo: A Systematic Framework For Large Video Generative Models YOLOX: Exceeding YOLO Series in 2021

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.473250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:8fb76b3fc676f721fba1a0d58c916fe074971a846f65d96319af7cf3e44b43fc

Observation 681ba316-c6ff-45ab-9fc8-5b1b6a570203 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.581649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:754aa9524be5d7913bb117f5717d1d8f86ece0eab431c495bdd9afc9a3b9f412

Observation 9e025068-4a03-4ac7-8c21-1b76f093d448 · outbound

This paper cites Chatglm: A family of large language models from glm-130b to glm-4 all tools.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Chatglm: A family of large language models from glm-130b to glm-4 all tools

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.694293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:4113d70294c4bd420f6b40937966c122542d3a5b0cb7173dae71ced71b6a3082

Observation 2efde1ec-cbb7-403f-97e4-ee617f57fa7a · outbound

This paper cites Sparsectrl: Adding sparse controls to text-to-video diffusion models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Sparsectrl: Adding sparse controls to text-to-video diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.680885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:d311d1be9ea21803d146ac1672f53eaa9a303fb65754f5a677d430fb55bad098

Observation 1ee511d2-1cc7-48ba-ae11-6016c51a4f55 · outbound

This paper cites Animatediff: Animate your personalized text-to-image diffusion models without specific tuning.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Animatediff: Animate your personalized text-to-image diffusion models without specific tuning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.686783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:62da050f1f04fa7389504c6964690d7c82df8b2072a5a240bbdde1557eabe692

Observation 62c79e91-2061-402c-a6bf-6f4d89e674ee · outbound

This paper cites Taming Data and Transformers for Audio Generation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Taming Data and Transformers for Audio Generation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.395027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:ddad9bedc3b5b56d818067345749a1c8f12d0047f0804161f030a82ffb8b8fa1

Observation c5c2acad-0f60-4a7a-95d0-0e30f961b19f · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.479249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:bdeed7184f94e5e8ed54f8d270f1dd473fdfdadda3cc4931965a1b30748da68c

Observation ea725561-f526-4fd6-9de3-e9e16f5ad336 · outbound

This paper cites Animate-a-story: Storytelling with retrieval-augmented video generation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Animate-a-story: Storytelling with retrieval-augmented video generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.667337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:e5dee83fca0fba37068e8bb61e3fea9b131bc92f97c65b83f18760f59bdd24e3

Observation 3429933b-ea54-410d-b472-b6e66343360d · outbound

This paper cites Video diffusion models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Video diffusion models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.676589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:eb6eaa62dcf5a5072648a3b5cf254a254e6679696364fec8fdadd4768c48d3ec

Observation 6c6d7cc5-ea49-41a1-a29a-a980ef84283e · outbound

This paper cites Imagen video: High definition video generation with diffusion models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Imagen video: High definition video generation with diffusion models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.690521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:86e8d6489a922c6dfa734869e444af802cfe4f09a6b091be10304cd3141715b4

Observation 00443d7e-a206-45e0-b0c3-26f88b5e67e3 · outbound

This paper cites Denoising diffusion probabilistic models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Denoising diffusion probabilistic models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.701677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:91eae3a7627db5c8faf468403b6838e6e7e1e8fe5bcd77f853d97706527e11f5

Observation 1da26971-7e2c-4dc9-94d2-94617802165d · outbound

This paper cites Classifier-free diffusion guidance.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Classifier-free diffusion guidance

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.717015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:f1559c7ddbbb60168f86db6d441b6ace587fad6d3b27a7118fb39275d80058d6

Observation 3748e25a-ea43-4cb6-b32a-6aaf71edd70f · outbound

This paper cites Training Compute-Optimal Large Language Models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Training Compute-Optimal Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.385871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:4516da12303238dc14dbf632f7ee3cd8c7e68038d8d7527df360ddcbf59a1193

Observation 163631e8-7e64-4d52-b1b8-ab8153047d06 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Animate anyone: Consistent and controllable image-to-video synthesis for character animation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.654328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:6b17933668bdd2ae20f24daa859ae416aa4b095b86e40c0766f640b5bceeab4d

Observation 85bf1a45-8951-4726-89a7-18bba89e0f87 · outbound

This paper cites A large tv dataset for speech and music activity detection.

HunyuanVideo: A Systematic Framework For Large Video Generative Models A large tv dataset for speech and music activity detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.628834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:dac5e455019b629816199b1f0c44228923a4fab318b35d2aa324ac6fba0a63df

Observation 7442c48a-3ff7-4b31-b28e-ae63811ee1b3 · outbound

This paper cites General data protection regulation (gdpr), n.d.

HunyuanVideo: A Systematic Framework For Large Video Generative Models General data protection regulation (gdpr), n.d

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.632107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:c041db2ea82aefe509a3ec06df932dfc6c6b797be8d6c8bec13820da36995ad5

Observation 503240b0-3920-426b-843f-cae5c39b30ab · outbound

This paper cites Text2performer: Text-driven human video generation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Text2performer: Text-driven human video generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.632669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:44795462f5102d987eb6842b63f9a5fd690f543fce2c23fdd0f85c23cb237a36

Observation 1ee66aff-535f-4e44-b3d6-89b09351a20b · outbound

This paper cites Scaling Laws for Neural Language Models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Scaling Laws for Neural Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.422916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:dfe10ba8a44d9b9e5bfd2337e461662805631e92deff18cdeec0501cdb43b499

Observation 9ef91b16-b4fb-4f37-b5f9-d44985a5c925 · outbound

This paper cites Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.589084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:3a1c5cca5f4cfca4e751dc52af98f9c2236dc6378227c9820bd41f0143dc6eff

Observation 29838383-0f70-4178-9bec-d5b8f37e14d2 · outbound

This paper cites Re-ex: Revising after explanation reduces the factual errors in llm responses.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Re-ex: Revising after explanation reduces the factual errors in llm responses

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.648012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:f16afaeef44f0f2a50296b4c720a549d8b62ddb17e91db9d13868b220e570c89

Observation 34651ef2-a3cc-4ef0-af76-d55f969319f7 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.654154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:55a3f0018dc01c9df45a604e60f470ca781e85bc465efa1eb68a30830eb5cbb9

Observation 77f0c4c5-23e5-4e4f-baac-8cdc25dcf84c · outbound

This paper cites Reducing activation recomputation in large transformer models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Reducing activation recomputation in large transformer models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.680561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:3cd64f89687fd40ad29154f57d2331b7b3b4f47249a3528a2d34f85017f8bd87

Observation 733d18cf-51b5-4020-bf26-9ad6657567ce · outbound

This paper cites Open-sora-plan, April 2024.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Open-sora-plan, April 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.688665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:7a5d17819375fcd2b50398ecf983ded05ac10db595ca753216b4c76a5f2608cf

Observation 75ee155b-368f-4054-b6af-a92fe22e9457 · outbound

This paper cites Flux.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Flux

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.692197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:133b923901ae1df77527e47c9f61b3bd5e76ad2121d29a8bfa9d6c86563c6d73

Observation 7b917617-b89a-4519-a312-d894cc1ef6b9 · outbound

This paper cites Tccl: Co-optimizing collective communication and traffic routing for gpu-centric clusters.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Tccl: Co-optimizing collective communication and traffic routing for gpu-centric clusters

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.664094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:5daf1b7227baad5f90a2878f5a0941691c58eedc145e4f33d4674f3b4aa5d46a

Observation 9b26706e-bd1e-4d71-99cb-b0068e78e994 · outbound

This paper cites On the scalability of diffusion- based text-to-image generation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models On the scalability of diffusion- based text-to-image generation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.592965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:632e73a0e450c44fd8ae4bde5fc4b9ead9895a2a8103617c4dd9582abc11426b

Observation 73ef7e2a-3e32-4558-a044-a6faec100e48 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.667703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:86effb02f54627d69b3650e97b3fc9358d05867a1dae51f987c057deb8b02522

Observation 65626552-a074-4b41-8f63-7dcb03c465da · outbound

This paper cites Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understanding.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.671594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:ae6efa3446878c6d9a939403bbd582598f49a881ea4fab90dfd7abef30b6e895

Observation a0c5f80b-bb98-445e-b31b-da4750850e1f · outbound

This paper cites Flow Matching for Generative Modeling.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Flow Matching for Generative Modeling

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.453642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:b647e9750647521fcf434e0f9dc2adc9ca2cb6ae1dee3cced5fdbd6bfcc465dc

Observation e1730fd4-76c0-4b73-b5e2-b63f3cdd53c6 · outbound

This paper cites Visual instruction tuning.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Visual instruction tuning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.592214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:7ff99025c1a109e177298555b553ba1b7de967a60e857c094fbdc5978c409dc9

Observation 5c4688c3-2a54-4e6a-9e9a-8e9542e764f9 · outbound

This paper cites Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Diff-foley: Synchronized video- to-audio synthesis with latent diffusion models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.600539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:6da4d559f086d85649176952b2f7ec0694980a077492ddd9e6fbc74ec097f81f

Observation 0bb5ee22-f104-4072-9403-c58cfb3fea3b · outbound

This paper cites Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.556148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:959a6000d6e7bbe07d2fa56716204c135871bb8faea32fe7134eb4fafb8f02c8

Observation 0ad29f0b-ea33-4bae-9c16-742e9bf01671 · outbound

This paper cites Follow your pose: Pose-guided text-to-video generation using pose-free videos.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Follow your pose: Pose-guided text-to-video generation using pose-free videos

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.621616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:44a9d8c346814d737c6dd50a5bbb94ef6c77af7bc9d8c3de7673b4f6d73519cf

Observation 405e66aa-e59d-42b0-9120-41c0515e4381 · outbound

This paper cites Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.562270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:6fd4a37bc0dfeb1ce47631801b97ef08987559153cc4aac197a1c2b4269680c2

Observation 8fbf0ae3-2d93-4596-a5bc-febc0b7728b7 · outbound

This paper cites Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.460841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:f74c45240a895fe12a6432f2bf1ace6a59c5bf19d2b6b1416b9925af8209d754

Observation 94cadbfa-a056-4350-9480-5ec5d94df178 · outbound

This paper cites Some methods for classification and analysis of multivariate observations.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Some methods for classification and analysis of multivariate observations

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.624970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:d7ceab56e4ea9f06d232a41f11dfb722cf84c555d5f4038f82fd85f616c434fb

Observation 3de46200-e716-4788-839a-c2dcf41feed9 · outbound

This paper cites On distillation of guided diffusion models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models On distillation of guided diffusion models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.567906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:1a52237d03ea259c464a2ebade09b452b7d86950cd0cad2fffe6abc2859e9299

Observation a822031d-26ef-48cc-a32d-ceffe019581e · outbound

This paper cites Conditional image-to-video generation with latent flow diffusion models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Conditional image-to-video generation with latent flow diffusion models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.705216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:6f2b931f2405629be5d3244647c8dda13c32cca555702a82335c18739b915463

Observation 47627293-853b-410e-ab9d-4efd1514e583 · outbound

This paper cites Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.536547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:298abe8169473b666dcc8474dc6dbc3d3b3ec63a5f1074b55519abfce7fdd03b

Observation 73117fa4-bde3-4d63-9e6c-caf57fea83c1 · outbound

This paper cites Context parallelism overview.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Context parallelism overview

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.708639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:0de037f9f812b8bbb025c00078828492279be6df28a45df4c95373e8b5f14e67

Observation fd890e38-a8ff-421b-a99b-c7dd10d6058b · outbound

This paper cites Cosmos-tokenizer.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Cosmos-tokenizer

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.607741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:271425f6d2113462b7a57e9fde66ef2fbaba95fc2ed4584fe3cf80e377084b18

Observation 025bb97d-38da-44bf-bea7-df8dc5e7a94c · outbound

This paper cites Scalable diffusion models with transformers.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Scalable diffusion models with transformers

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.571430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:89ca1e4f95dc2d85078d5fe162f84892b9f398ec332595c64732fc8582463d85

Observation 63ad5918-fb7a-4fdd-a401-bcadfc4e3f2d · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.578971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:ae0a016da96d0a49719d0b100326ebfc7620ff58941b69b63ddc992f9025e9e9

Observation 92fe95a1-d7dd-4183-b56e-83c7f1a21dff · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Movie Gen: A Cast of Media Foundation Models

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.402396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:ffc70a03c22f63cd572661f7ad72037cb1817632843b830a67c76207739093f5

Observation 219e9a0a-7b4c-49f7-b326-75422074460b · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

HunyuanVideo: A Systematic Framework For Large Video Generative Models A lip sync expert is all you need for speech to lip generation in the wild

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.551586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:5ff213e5ded05c2f4f597dd91fba033563eb66430b21c456765750ad37636cf0

Observation 4cdcf38d-4934-4d01-bf85-214d923adbcf · outbound

This paper cites Learning transferable visual models from natural language supervision.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Learning transferable visual models from natural language supervision

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.574835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:0fdd35ccb552cfbb9ef2156e085af2a2fed33373a425e7dd18058aefe1ef9671

Observation 613d70e3-9d0d-4d9e-a84e-af797ed0951a · outbound

This paper cites Learning transferable visual models from natural language supervision.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Learning transferable visual models from natural language supervision

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.604277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:a769fd0ab5e63a246f1d24c2be53aeeefe259761851e4fc4af374831c164ad69

Observation f54297b9-4d21-4e4f-9964-1aaf99a0a296 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.582436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:1178d021166e2b372b7da85ea16b1c2a38ee34eacbc6b56573858e266dcd2c9b

Observation 46286e98-2719-426c-b2a7-77f31d88d3f3 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models High-resolution image synthesis with latent diffusion models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.742616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:4344bd4dd08c5c41f46367ed3fb6a2e29d67977d3cb722876bf171ef3a8ed1fe

Observation 3c16f9b3-94f9-4fb8-b4d9-ca5fe0be41aa · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Progressive Distillation for Fast Sampling of Diffusion Models

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.409132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:bdb210d941988eb784d3dfa289dd5d494f0075ca7eb360b272606b73c26071ec

Observation a5652ae3-f68e-4c14-84de-88f417e5fa37 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.467254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:9f34aa4dcc8cfa544f043d11826a2a1d39927ffc5988887766097ee38a9464e9

Observation 6f0e9c00-2fee-41cc-9458-d9f4f180d937 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Make-a-video: Text-to-video generation without text-video data

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.718025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:32acbea9386695224fb93881bc13d49c7f8763bb17107f38f1a0b4d0a9ebef98

Observation fe2c8e8a-b8a9-42d5-bc3e-50beb37ef41d · outbound

This paper cites TransNet V2: An effective deep network architecture for fast shot transition detection.

HunyuanVideo: A Systematic Framework For Large Video Generative Models TransNet V2: An effective deep network architecture for fast shot transition detection

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T07:42:43.518678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:c8ffb52943c3e019df17986e3cc313deee3b7968686c23b295799b9981fcf3eb

Observation fc07dbaf-3ac6-49f0-b891-83791ae9a1a1 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Roformer: Enhanced transformer with rotary position embedding

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.542794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:2d4dcaa2026f901d7ae4ee7d07bcbda420fa24f0286f2460d87d05e8c61c7ba8

Observation eab2c802-dbcb-423b-8a37-f050989a5c89 · outbound

This paper cites Hunyuan-large: An open-source moe model with 52 billion activated parameters by tencent.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Hunyuan-large: An open-source moe model with 52 billion activated parameters by tencent

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.547864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:eb4c4b2cd54b465c2b7167274dc5e611f5351ef0fab991a3cdca8a3920edb947

Observation 109665dc-14c2-489e-9312-06ae507da4e5 · outbound

This paper cites Mochi 1: A new sota in open-source video generation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Mochi 1: A new sota in open-source video generation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.585921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:741d25e251afe73994628b7822add5739caaadd7411d6c2dae39f2bc0b85a1c6

Observation d60f8fc1-faf6-4600-ac83-fd8495fc472d · outbound

This paper cites Emo: Emote portrait alive – generating expressive portrait videos with audio2video diffusion model under weak conditions.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Emo: Emote portrait alive – generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.545034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:f93f38a13d4bd4544cb764119976a5390f9acb886b2a5d9e479c7d98d9b95cc0

Observation 522f56cc-09e7-4be3-b119-cf51e2050af0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models LLaMA: Open and Efficient Foundation Language Models

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.574426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:cba127570a7ffd8240a4f0a1c2e0ba1f6c21dea1f33349979cd3e0c1295acdfd

Observation ca1ab31b-09a4-416f-b965-d6e3527f791f · outbound

This paper cites Modelscope text-to-video technical report.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Modelscope text-to-video technical report

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.531232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:4dd8f210c3d992b03965524c45a7f03d9ce445cf635ec4edf049a951ae97c84e

Observation e5625999-4b75-4bc6-b6ee-d9a9a544ad89 · outbound

This paper cites Disco: Disentangled control for referring human dance generation in real world.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Disco: Disentangled control for referring human dance generation in real world

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.534857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:087f87127e4662bc4e9e73580d305df41734c9aead87bf68d0cd78df679bb26c

Observation e012f27f-3640-42ac-8c20-32fe93a5a6f9 · outbound

This paper cites Lavie: High-quality video generation with cascaded latent diffusion models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Lavie: High-quality video generation with cascaded latent diffusion models

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.575048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:4939101df54b5ab75a407a6177e10b27d1cdb0ce6cf4f7ff47892302af2b5c19

Observation ab0240ff-9e55-46cd-bd93-0e93a9050387 · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Exploring video quality assessment on user generated contents from aesthetic and technical perspectives

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.713336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:c62bb86fc3c2d5cff3b96e194542fa48af71d5885325d1098b0d4076a902d431

Observation 22dccad6-4b74-40b2-b2de-6ec4348a3492 · outbound

This paper cites Lamp: Learn a motion pattern for few-shot-based video generation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Lamp: Learn a motion pattern for few-shot-based video generation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.720923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:09514737bb82f95ca9b7b1f1ad755900b76a921be93e7f245256ba9ca9fb86fd

Observation cd43debe-5541-4263-8851-668201f28254 · outbound

This paper cites Hallo: Hierarchical audio-driven visual synthesis for portrait image animation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Hallo: Hierarchical audio-driven visual synthesis for portrait image animation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:45:30.628547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:d386b8fd95391c5fba1381efbd7e5a3c0a7368e03aca3bd519f06b926eb7f636

Observation 3f856197-4b28-4730-8dcd-60b72f236c56 · outbound

This paper cites VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time.

HunyuanVideo: A Systematic Framework For Large Video Generative Models VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.499945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:59c7eefafb63122a2a6da65adb0c6b01c5e4e31a27ef22750b3b2cfd5b0b2f4a

Observation 85c39580-26f2-4795-b242-d89b3f0d8509 · outbound

This paper cites Magicanimate: Temporally consistent human image animation using diffusion model.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Magicanimate: Temporally consistent human image animation using diffusion model

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:47:43.166913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:de9d99833b18b9f6622c35659dade215f251455d4eb829270b542db7e2968fe2

Observation e6a9df93-bf42-42d2-b046-6a56d05374bd · outbound

This paper cites Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.434632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:a5061a62355938390e1ec680d6a8fbaf052a3640918eeaee9ef60c23df3d95a7

Observation d037db2e-f21d-4d2a-bd1a-93f3b10156c5 · outbound

This paper cites Probabilistic adaptation of text-to-video models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Probabilistic adaptation of text-to-video models

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:47:43.158459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:6897c7432978aed588d594552905737e3330def871d678babfa49a6b562d0eca

Observation 8cbe08e1-538a-4907-b13c-26fc5483d42b · outbound

This paper cites Effective whole-body pose estimation with two-stages distillation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Effective whole-body pose estimation with two-stages distillation

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:47:43.139696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:c7aede6975882ca367296a0b1d6947e2bfaa04bb4b284f0b834d3f4823a8271f

Observation 99807af0-6961-4840-912f-b909bf817533 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

HunyuanVideo: A Systematic Framework For Large Video Generative Models CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.568261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:503654456879ff294107bde75cd679983666dfbdf023d219cd8138e7c28b6fa3

Observation 077a7a8f-a87b-418a-ac0e-fe4d87044a24 · outbound

This paper cites Mimictalk: Mimicking a personalized and expressive 3d talking face in minutes.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Mimictalk: Mimicking a personalized and expressive 3d talking face in minutes

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:47:43.176668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:284c0e5b88dcd2fb3476a0e532362f3448f138cd4e1ad89bb1895a25fe289366

Observation 59659b5e-f1e0-40d6-9c41-f52e293a3c1b · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.594956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:5ea3933092d1e7a15ff8ef69839de606d906ec269bfa18f060c5a7f22e6303de

Observation 103c2871-7c49-4886-b672-224701d4cd14 · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Show-1: Marrying pixel and latent diffusion models for text-to-video generation

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:47:43.143193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:64e8f61c04c6e09bb811f4ed71b537c5b744b0fe2c6d3bffbf6a56649ee39956

Observation 125693be-6e1f-4129-b910-edca402e066d · outbound

This paper cites Dinet: Deformation inpainting network for realistic face visually dubbing on high resolution video.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Dinet: Deformation inpainting network for realistic face visually dubbing on high resolution video

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:47:43.155401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:c59e15bdcd334724eefd4a2302b42af7aed7de048686e659b1e70c4cc577b72a

Observation d9151573-43e4-4377-9de5-c696a57b9044 · outbound

This paper cites Unsupervised representation learning from pre- trained diffusion probabilistic models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Unsupervised representation learning from pre- trained diffusion probabilistic models

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:47:43.161443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:be3c4c95b7cff02065a0609e8c9b178a9421f34a8f6cbf59b88c72086076818d

Observation d5c316bd-3dbf-4ef5-8ef1-87f738ed02e8 · outbound

This paper cites Shiftddpms: exploring conditional diffusion models by shifting diffusion trajectories.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Shiftddpms: exploring conditional diffusion models by shifting diffusion trajectories

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:47:43.164041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:d0d29053f620713216a5ec744462c8bbff391716522eb80158ebb96aa4d7f3e6

Observation c59ceede-0061-4ce5-b0be-995d1557dc75 · outbound

This paper cites Motiondirector: Motion customization of text-to-video diffusion models.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Motiondirector: Motion customization of text-to-video diffusion models

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:47:43.173586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:efe109db32f064461deb651769346fc7f04a49cadb812e8037b71102dbf018b2

Pith citing papers

Observation 84700953-d876-43b1-94a6-2d095bb6f923 · inbound

Latte: Latent Diffusion Transformer for Video Generation cites this paper.

Latte: Latent Diffusion Transformer for Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:45:35.793876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T21:45:35.754742Z digest=sha256:1814901354abf121f223cb3206bee26ac71221089b5f3471d3490ff8213b411d

Observation 9e7fac36-4113-4dfd-9c5c-7b332578cd53 · inbound

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference cites this paper.

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:18:42.366516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T01:17:11.261301Z digest=sha256:b8c852422687fb980082180397d606e4aabfff01838c64970c1c61752d9fff83

Observation 79c0bd49-a129-4041-a8c3-c95a5db26552 · inbound

LTX-Video: Realtime Video Latent Diffusion cites this paper.

LTX-Video: Realtime Video Latent Diffusion HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:36:12.260358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T10:36:12.058970Z digest=sha256:28644d290c9f558426016ac9f62580597984d21882b1ac75d2bd9f7a67a2b8e7

Observation 75c603e5-8230-4bf1-8e31-a2bfc1722210 · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:38:45.615723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:222a577bb5d0827dd5b1fc6ec6a2033ea3f584f55d09851c4038763d5139fb87

Observation 89f803ed-4b2d-4538-bb4c-7480acaa1ffd · inbound

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation cites this paper.

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:55:24.895404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T04:53:58.465445Z digest=sha256:8badbec59b9e67162d3e2563d86bfe4aa4341031387f9d7f2f7931613ff1d7ad

Observation d66e9e1c-1e90-4a2a-95eb-9a9db11f5dd8 · inbound

Improving Video Generation with Human Feedback cites this paper.

Improving Video Generation with Human Feedback HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T15:30:02.704365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T15:30:02.578430Z digest=sha256:0446951a20eafa0ee227677e16ce35b61d2c9582d93cc7d1f4bf83b19ccb75de

Observation ad71c6a6-74c2-4721-a295-821ec4004ad1 · inbound

History-Guided Video Diffusion cites this paper.

History-Guided Video Diffusion HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:00:14.744304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T12:00:14.672729Z digest=sha256:067408cac397b441760accc38cca68c85e29be000c57e0a8fb0f0491974b2154

Observation ce52b633-263e-4d43-b961-5582384821aa · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:23.931507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:73bf98753fdea98868682c5a81132a6be272f5463096d07cd0bde6726b3f8c5d

Observation e29feea0-d941-44e2-b3b8-ec7206d553f7 · inbound

VACE: All-in-One Video Creation and Editing cites this paper.

VACE: All-in-One Video Creation and Editing HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-16T00:53:53.940174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T00:53:53.855965Z digest=sha256:fbf790ed770f4aecdd93ac8a3936e372576039d95597782dc422b5503a707e4a

Observation 134fbd94-6400-4bfa-99eb-84289c2fb8b7 · inbound

Long-Context Autoregressive Video Modeling with Next-Frame Prediction cites this paper.

Long-Context Autoregressive Video Modeling with Next-Frame Prediction HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:05:17.255739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:05:17.201790Z digest=sha256:08bd6e7b116a3ab04e6d18cade079d0f2e3781c6f3f7d9cd50073177bf599321

Observation c245c804-1333-4874-890a-9798ce7c57da · inbound

Wan: Open and Advanced Large-Scale Video Generative Models cites this paper.

Wan: Open and Advanced Large-Scale Video Generative Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:07:14.392735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T23:05:32.595632Z digest=sha256:6caf35aef7c897cf4b866f4e689184f9e6121c4da25734f4002b0a7bb4837cd7

Observation 6ddd7a80-58c8-4ad5-bfff-70d90d25a405 · inbound

We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback cites this paper.

We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:56:58.336898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:56:39.735334Z digest=sha256:2470840a7ae50ac5b350dbbd003450c8eaa1f6fbf1ed20e678ea33ae5518b5bd

Observation 2c569132-841d-466c-8002-09f7ccb4f6b0 · inbound

Step1X-Edit: A Practical Framework for General Image Editing cites this paper.

Step1X-Edit: A Practical Framework for General Image Editing HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:36:41.658854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T14:36:41.467429Z digest=sha256:6cdceccea0df181803bbad4f33529c023644bb43e6271dd562c5003ab6b707e7

Observation f31e647e-fd2d-4bac-8514-6a3729d1e7e9 · inbound

Flow-GRPO: Training Flow Matching Models via Online RL cites this paper.

Flow-GRPO: Training Flow Matching Models via Online RL HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:45:16.868292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T18:45:16.641012Z digest=sha256:7b3beeec9f2773b1f42b2e3f27601da1ace114cd5bd48cb888260c092b686686

Observation 3ab6331c-aa81-4c44-b3ef-d421b22191e2 · inbound

DanceGRPO: Unleashing GRPO on Visual Generation cites this paper.

DanceGRPO: Unleashing GRPO on Visual Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:28:26.762461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T22:28:24.929046Z digest=sha256:ced033c1972459b1deef7e6b25aebe533cbaaeff96b90bec928c82649eec2b57

Observation b174ec06-e8f6-40a6-b3b1-1ed888a51823 · inbound

DreamGen: Unlocking Generalization in Robot Learning through Video World Models cites this paper.

DreamGen: Unlocking Generalization in Robot Learning through Video World Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:50:45.447431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T23:50:45.332466Z digest=sha256:53c72c54969ee625c6bd895297277c200d9067b0b6918c61bb5441d4f5b454bb

Observation f65506e8-04b8-4284-85d3-668d7cecf64a · inbound

MAGI-1: Autoregressive Video Generation at Scale cites this paper.

MAGI-1: Autoregressive Video Generation at Scale HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:31:15.805001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T20:31:15.700943Z digest=sha256:2c390dac8442cb5c14f425d3e4dd384691227cf702549630e5553c7ec842a7ae

Observation 43fc7807-9237-4cae-b9f5-66485888a76e · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.227621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:87377a33acfb0c483158a6df67c1d6c715efbc58136306a29795474520bd7a30

Observation 2537f70f-da02-4681-bfd4-6bbaeb52fb21 · inbound

Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation cites this paper.

Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:32:17.794853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:31:11.886858Z digest=sha256:bb2924fbac6e16a06ea484e604679059800b6fd8043f5483acddc9ccbbb81ecf

Observation 611dfa89-06ed-4ec1-ab3c-c1103e286331 · inbound

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion cites this paper.

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:36:53.103420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:36:53.029590Z digest=sha256:73fb62790fd82e56f9135f2775674d5038f8d4377536474dbc7742d6251b9331

Observation f286e790-868a-4b31-baf7-0838d7eb9f09 · inbound

Seedance 1.0: Exploring the Boundaries of Video Generation Models cites this paper.

Seedance 1.0: Exploring the Boundaries of Video Generation Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:09:57.514485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T12:09:56.836351Z digest=sha256:a1712a24795aad2afe0c27d74bd423aeb6188edba3029b30ba0ad1dc77935515

Observation 753aaacc-db8e-416a-b1e3-8f9fd9bb2c10 · inbound

Listener-Rewarded Thinking in VLMs for Image Preferences cites this paper.

Listener-Rewarded Thinking in VLMs for Image Preferences HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:42:09.222937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:38:52.273903Z digest=sha256:a434749682e32d22ba9dfe468059269b967c9aad6e49e4189dba938823d18e5e

Observation 761f4169-abc1-442e-b9b1-f9172730d5af · inbound

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation cites this paper.

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:02:10.298464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:59:49.398271Z digest=sha256:c6db24aa3116bb76c4356fd259c399cce43d1273c08f8d7943f86c099cc102ce

Observation 8cef9893-fe22-4c92-961b-232a7f7931b9 · inbound

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling cites this paper.

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:17:06.748987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T05:13:28.767788Z digest=sha256:0587e61e53f34edf48bb12129c1051da830d41c8a88c331cccea634b0918676a

Observation 309ec80c-d386-427c-9532-5799d2598f39 · inbound

Vidar: Embodied Video Diffusion Model for Generalist Manipulation cites this paper.

Vidar: Embodied Video Diffusion Model for Generalist Manipulation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:54:28.298405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T09:54:28.271928Z digest=sha256:504b4e2e1b649602c9ec5fc01aed3c31ed53dd1bca04df8b20b610cfce5ba6e4

Observation f1905e0f-6e67-447a-9a6d-271ff85b10a5 · inbound

Qwen-Image Technical Report cites this paper.

Qwen-Image Technical Report HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:29:06.998061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:29:06.883874Z digest=sha256:90166be2b512eba1ac7bb7905b8ba12f41f77eea7838261208dee38223eeeb92

Observation a9ff98aa-1958-41dd-b0bf-798a44317501 · inbound

Matrix-game 2.0: An open-source real-time and streaming interactive world model cites this paper.

Matrix-game 2.0: An open-source real-time and streaming interactive world model HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:36:53.228732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T22:36:31.044743Z digest=sha256:c22de030ea2075d8ccdfdcc46bc9dc6383e7749d5ebaeedd131135ffa09dc07f

Observation aa3dac1f-fb0c-4b63-b8fb-ecd279bfcce3 · inbound

HERO: Hierarchical Extrapolation and Refresh for Efficient World Models cites this paper.

HERO: Hierarchical Extrapolation and Refresh for Efficient World Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:11:52.834147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T22:10:16.645510Z digest=sha256:1ff25b3883c7e43b7a855d281a5cda19b3ce3f63565afcc08ad057197e134bd7

Observation b0133e61-3f1b-41b0-869b-5c97632f62cc · inbound

ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency cites this paper.

ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:30.841102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:30.841102Z digest=sha256:220096b83e65d4bc2446e2bff1ba5bd2ea23a0fc94f38aac071d618ef9059ae9

Observation 7fb76f94-98a9-4e52-b57a-a60ac840fe62 · inbound

Wan-S2V: Audio-Driven Cinematic Video Generation cites this paper.

Wan-S2V: Audio-Driven Cinematic Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.824740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.824740Z digest=sha256:98d4f9c5ac94cb3d3532f1b5e9575150ca24795402e91e656cc7567167601452

Observation 2e1a7866-4c6d-4c1a-90a4-0b68584b0c53 · inbound

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance cites this paper.

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:22.264641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:22.264641Z digest=sha256:8d1997917a147fa8f15cb948ed152db94d4147f7077be0fc7757764d0be6ad55

Observation e79f3ce6-b458-4db6-a974-3be0c80eac10 · inbound

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement cites this paper.

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:42:40.545632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:42:40.545632Z digest=sha256:4c66ea0c52b3361395a279e242edb83385273a63ef40e8fbb7b9218369f3959d

Observation 1750f22e-10b1-440f-a70b-d49a49f0a5d5 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:46.492137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:46.492137Z digest=sha256:d338f9d05e52d2b62be0fc6a01cc97ddf82c035943f865ab3db02abb980ba72b

Observation e02591c5-48db-4050-9bf1-7b32005ecb64 · inbound

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning cites this paper.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.290001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.290001Z digest=sha256:7da64a0bade9c7855520ae52bdebe0528a9ae071c889911bf8f2a902b45cd811

Observation 836d2260-1b27-41f8-82c9-82b7380589d3 · inbound

RewardDance: Reward Scaling in Visual Generation cites this paper.

RewardDance: Reward Scaling in Visual Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T20:08:57.666163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:08:57.666163Z digest=sha256:5d1513ca27ae19c2a71bdcf2cf61dac7a9f1ae5a8820b80eb0966ba1e7d7e405

Observation e0f7e400-8ebc-4ee8-9383-21a235fa7205 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 257

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:24.647347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:2f9cdbee6b440c88adff83b3ad32c33158c4e6d91861e7dfe6a1f4e4c09b3932

Observation a74e50d2-734c-49bc-bfd3-f7faa0312de4 · inbound

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders cites this paper.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.939891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.939891Z digest=sha256:fc4bb84a28b8d2415917d3a10901775984c0ae5410838aebe1bd54bb62c9cad4

Observation 1299a658-34d4-4a53-b674-62e4a1b9e299 · inbound

FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models cites this paper.

FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:02.039951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:02.039951Z digest=sha256:33e6c639ed7c4300e08f86c8c015c3396746d9b06fefe85137513d4e339c0681

Observation 629a8416-51ca-4dbd-9cea-206891ce50cc · inbound

LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation cites this paper.

LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T15:57:09.017436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:57:09.017436Z digest=sha256:74bd74417215c67eafe1f1f016391cfdeea0c9acf6617fc13a13c99398f51052

Observation 7b6498c4-3adb-49cd-b5c3-b290dfe548ab · inbound

EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning cites this paper.

EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:46:25.831875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T13:46:05.547669Z digest=sha256:193d3e6aa065ca735dca19ad84de6ba77a39a426338fc565f32d5482adfeac0f

Observation faad3cd3-396f-4c35-9273-fb1519a9133d · inbound

LongLive: Real-time Interactive Long Video Generation cites this paper.

LongLive: Real-time Interactive Long Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:52:59.529743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T03:52:59.287555Z digest=sha256:2e9fc6ca48243d204dcb7b5128540c7a963d8196dc72669d9646e6651b24f483

Observation ceb4ec91-c080-4fa6-aeea-b787520464a3 · inbound

Sample-Efficient Optimisation over the Outputs of Generative Models cites this paper.

Sample-Efficient Optimisation over the Outputs of Generative Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:01:23.721555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T12:58:33.173828Z digest=sha256:ab9aeeb399a9867ea93fff00dc06227726e888e376773e07b0908f2e6941cd3d

Observation ed7604c9-d1f4-43a2-81a7-e1817c14df89 · inbound

HunyuanImage 3.0 Technical Report cites this paper.

HunyuanImage 3.0 Technical Report HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:02:32.847972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:02:32.806844Z digest=sha256:86dab36d9d05b0edbaf8553e26800620270d4fcf93c16c074e513e9263adf403

Observation 86f357cf-c9ff-4caf-82dc-612523d9f1cb · inbound

HunyuanImage 3.0 Technical Report cites this paper.

HunyuanImage 3.0 Technical Report HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:14.426058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:42:14.426058Z digest=sha256:cd4513dc4187df7f2a3c447d21427646c2d1030143b68f32b02c2bbdc06358a5

Observation f647d21b-fc35-4af8-a8c0-5cc8cf6368a5 · inbound

Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution cites this paper.

Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:51:20.383078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T11:47:22.725217Z digest=sha256:1ef1b7b81064ebc19665b520c2a2eecf4f9dd3072b181e6adc6d4d588877279c

Observation d79cef52-2b7c-426b-895c-43e038b7b525 · inbound

Rolling Forcing: Autoregressive Long Video Diffusion in Real Time cites this paper.

Rolling Forcing: Autoregressive Long Video Diffusion in Real Time HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:15:29.145086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T11:15:29.102090Z digest=sha256:245ea0dd071fe9d11e320d1340ac18d951188508b42b20795fc0e09cb5848563

Observation 9a20bef6-f011-4a87-b2a5-9ca7114d999f · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:03:15.244132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:50c01fd7f4f5e3fb7da312895988b08bb4ec2a1888ff57e408b5ce835e2d895a

Observation 6d82c558-2ac6-4f03-894e-078c41c0b990 · inbound

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation cites this paper.

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:39:54.070210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T22:39:53.995700Z digest=sha256:897826022d87151596fbbaabacdd8fcc42a281617350bc24be6b22b622a20166

Observation d0ec8403-afbd-4bd5-8376-38d561ef2c3f · inbound

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction cites this paper.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.460371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.460371Z digest=sha256:392dc7063193e28d411224af564bb86f57cfd98c780d66e92d9cf580842285f1

Observation cbe1dc13-452b-4167-b290-dd39e60f4179 · inbound

UniVideo: Unified Understanding, Generation, and Editing for Videos cites this paper.

UniVideo: Unified Understanding, Generation, and Editing for Videos HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.115320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.115320Z digest=sha256:3f2f2839d07931ad965c4b2ef57b9b20fe24add67fd5ed2283523c5f11bc4f08

Observation 0f3d864e-139b-49dd-afe5-816e2fd1a42a · inbound

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning cites this paper.

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:45:58.886091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:45:58.886091Z digest=sha256:63220f2f6866e3dc57613c1605ffb8dbb866005567a52632fb310ad191c8f712

Observation 597b68e7-fdc0-455f-8f3a-978c4e3c6503 · inbound

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning cites this paper.

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:09.131224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:09.131224Z digest=sha256:275c988912613ed32c39ec5c3b8b9a8cdbca03d359847f451b9209b9754b8ab2

Observation ee0587aa-c9b9-464c-afcd-dfaf06f68c69 · inbound

Demystifying Transition Matching: When and Why It Can Beat Flow Matching cites this paper.

Demystifying Transition Matching: When and Why It Can Beat Flow Matching HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:31:36.530182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T13:27:09.343975Z digest=sha256:92d9f7d82be6088f886f9d97d3435bc454537dc3f91dc15318f4e925e8b639a8

Observation 1cbaa0e5-0270-4eb1-95a2-5f66c6f04b5a · inbound

OmniNWM: Omniscient Driving Navigation World Models cites this paper.

OmniNWM: Omniscient Driving Navigation World Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T08:57:10.411369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:57:10.411369Z digest=sha256:9c50ed5215f6af34dccde080b4b82b461d1cf4999710abac3355f992940a7566

Observation f08137bb-858b-4c51-b4aa-feaf40f2d22e · inbound

RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling cites this paper.

RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:15:54.504794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T05:13:42.934115Z digest=sha256:6a032d3f0811fd5749043adc142701bae20bf5d9057545e572ab6db596b1b477

Observation 62294a1f-e490-429a-80c0-87f0bf0a9a53 · inbound

Epipolar Geometry Improves Video Generation Models cites this paper.

Epipolar Geometry Improves Video Generation Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T08:20:08.446380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:20:08.446380Z digest=sha256:cc57c33f0bd6ae80db83c6bdce33aef226815459bec56dc3f32a0e82061d464c

Observation 9254576b-b5d8-4ea8-a881-8fe662e59d44 · inbound

Generative View Stitching cites this paper.

Generative View Stitching HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T02:55:46.631739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T02:53:03.898500Z digest=sha256:a1092cf7b63aec0ba0e9af3d937fff20db64db0d7f700b8346e5854602eb396f

Observation 665fad89-9f5f-469c-a9e5-a035537641f5 · inbound

World Simulation with Video Foundation Models for Physical AI cites this paper.

World Simulation with Video Foundation Models for Physical AI HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.624574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:b59e24ea91e15023c26786a05118a01e166a396ddd167debe09e36607aa66f74

Observation 33abd66d-501a-4660-9dca-bb67f03953cf · inbound

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm cites this paper.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.113178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:31f773ba3e1cddde0d2bee278ba3a50e5778270eb152cec64373e1f4e1efe3f3

Observation 6b19ba5a-aac7-4f23-9844-69623a89d02d · inbound

CGCE: Classifier-Guided Concept Erasure in Generative Models cites this paper.

CGCE: Classifier-Guided Concept Erasure in Generative Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T23:27:48.846912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:27:48.846912Z digest=sha256:9492ae16964303f259ed0e621d87802e1da387a750c3ddf58088ff041634e511

Observation 50d456e2-f19e-402d-934e-b5c4cc5a3a59 · inbound

Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space cites this paper.

Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T22:12:43.829418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:12:43.829418Z digest=sha256:c4e96e00b14f56ebe50734a23b6f53440b3400c6fd37cfab82fc6393ef51e390

Observation 3ad9ba29-fd5d-4b66-9abd-9d67ea5f28ae · inbound

PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention cites this paper.

PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T21:01:08.144786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:01:08.144786Z digest=sha256:5b0aaeadcbb40c7997e5eae77df9bc0470f601ea5e8ba8d9159911403c50bb19

Observation ace2c002-f65f-4b86-93ff-17534c32308f · inbound

Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation cites this paper.

Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T19:55:09.960106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T19:54:35.190865Z digest=sha256:e76787138d2da8d822139a878b840def8107c4d1d0f59bfe02976a903717506a

Observation 2822162d-e4c3-4759-a83f-4f402d5bb595 · inbound

HunyuanVideo 1.5 Technical Report cites this paper.

HunyuanVideo 1.5 Technical Report HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:31:35.082213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:de4b026818cad5e9e8fbb4c3c9de0b80337be65eebdb21795a48e4a3116f6c94

Observation d3d85427-7f60-48d7-82cd-40bc2d32a586 · inbound

SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation cites this paper.

SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:24:18.309415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T18:21:08.172735Z digest=sha256:b0d94693837c756c35d6c3aa4b1f1997229ca2c6812160ed447bd7bd060b8da9

Observation dfe3484d-e872-494c-8a88-fb3fa698eb21 · inbound

Rethinking Reward Signals in Video GRPO: When Scores Become Targets cites this paper.

Rethinking Reward Signals in Video GRPO: When Scores Become Targets HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:36:06.225043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:36:06.225043Z digest=sha256:940d6486c2071d3d41ed05aa76cb6aaac00389e33aa6b9d400c30772fdf94a3a

Observation 79ca9d28-9f9c-48b2-b3d3-ea0458dc9a7c · inbound

Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations cites this paper.

Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:55.200351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:55.200351Z digest=sha256:30c98f8f6e1604fd4c6be7eaf166afe8f46bb11cdc27d13c53e863475b834019

Observation bef57812-93d2-4e07-a8cf-e4998f49a69b · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:05.784126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:05.784126Z digest=sha256:97924c5c0665374ca245530c8e7563801eab0fde4f6e4db77abef890ec33df17

Observation f96a9cb3-76d4-475d-b607-ed08ad924797 · inbound

WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation cites this paper.

WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T19:54:48.562703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:54:48.562703Z digest=sha256:f246a1429017a6a12e50db5b461be88ad8c28585586942b85f8dce30d77278b9

Observation e2fbe6af-36b9-4237-9bef-baa4124df15c · inbound

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer cites this paper.

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:25:28.343463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T18:25:22.486891Z digest=sha256:4ee46d8392ef927fff158410498ef510d2f65a6decf67f580e1d479b4d6e36c7

Observation 3419c7af-9678-46fd-a89a-f333b3941dd7 · inbound

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos cites this paper.

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T19:11:49.466700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:11:49.466700Z digest=sha256:6e911e80aa93bde094106e4f30e2076a3671b35255e82e51a7931e521092bb0d

Observation 4fa19502-a264-4a5f-874b-17340f71bcf9 · inbound

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models cites this paper.

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:00:27.297144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T17:57:57.263574Z digest=sha256:6ddc4895b1678cac7801cc527b1ff97850f7d820001bc6a8d11fc4945c4bad34

Observation f04f5eee-251a-4dd1-a099-a025906e6801 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:58:51.454329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T01:56:44.123092Z digest=sha256:dddf8da403cda96fd3a0ef5d1b98610c80d1c4e43082352620ce3a086486fc8c

Observation abc910a8-b0b3-480e-828e-57e52457be0a · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T18:39:41.373827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:39:41.373827Z digest=sha256:ee0ae1358cc5c21e43d894f357d04c7a34582f20bac9bced0eb3001452888ccf

Observation ee29985e-7d6e-4b31-93ed-5d24e60c9a78 · inbound

Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation cites this paper.

Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:17:55.181751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T18:17:54.943863Z digest=sha256:741d6f7156eda5de966356e0a28fd83e9cde9cdc8ba395f9e7c72e534c1710c8

Observation 00edf025-6713-4902-a718-aa54e6d3ab5a · inbound

ProPhy: Progressive Physical Alignment for Dynamic World Simulation cites this paper.

ProPhy: Progressive Physical Alignment for Dynamic World Simulation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:08:47.993615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T01:05:08.087136Z digest=sha256:df2616c6873888c50ef7cfe54529ca0f73d0559b395709a973cd5e0405a6ff67

Observation 9fd97775-b1cd-4578-bad0-c80f1a3e7aff · inbound

InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem cites this paper.

InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T18:26:34.119227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:26:34.119227Z digest=sha256:b9acbc772d9795267bd34c2c58c7c279510f136efe5cbbc567f6040991806a48

Observation ea62bf59-d49e-45b8-b50f-5c7fbf00fdb6 · inbound

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment cites this paper.

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T18:09:49.547789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:09:49.547789Z digest=sha256:5333537fa577865126634a5a07d18062b3f7972e4231612ec928b2be479a648d

Observation 93f590a6-4fe2-488f-b5eb-3577149b11c0 · inbound

VideoCoF: Unified Video Editing with Temporal Reasoner cites this paper.

VideoCoF: Unified Video Editing with Temporal Reasoner HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:08:43.104090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:08:09.479706Z digest=sha256:2e22b1c443deb2d944b133ef7589156e642d1013c3f601d4acfaa8416f88d67b

Observation bdf091f5-c608-45e1-9cae-fd84df69a343 · inbound

VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification cites this paper.

VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:31:21.971293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:30:14.969895Z digest=sha256:5c0fb9157b16a85ba0d5969bed2ac407597889e04e0a38c204d4a5ef64bd8518

Observation 55361951-6ae8-4684-b1f0-7dbecfdc3109 · inbound

AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner cites this paper.

AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:28:40.788123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:26:21.855485Z digest=sha256:4d0ea497fba7d2dce522852500d4bbaf0ea28090e08f7deeb05222de3c795a94

Observation d86a5b82-9294-4adf-aaa8-c2705b28c545 · inbound

Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns cites this paper.

Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T19:59:06.214617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:59:06.214617Z digest=sha256:9092f13da6d061f57c13cd0aee434dbfdd2b2788525cad0a4d4ef255afc2d02a

Observation 31881e81-69eb-4538-ae80-d1fe6afc964d · inbound

Setting the Stage: Text-Driven Scene-Consistent Image Generation cites this paper.

Setting the Stage: Text-Driven Scene-Consistent Image Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:44:17.307598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T17:40:25.779794Z digest=sha256:27434ec41e280ce6a9c885784f5cd40177b74ab440005703b8e2c4da125ba178

Observation 808651a2-bde0-4c82-9029-654d3ead2f1b · inbound

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits cites this paper.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:35.074851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:35.074851Z digest=sha256:8ec45401dca5ddefd5085990895cfa5418fa702490368da39cabdec90367e031

Observation bfe6faaf-de51-4bd5-8b36-5fd5c5a1b3f5 · inbound

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? cites this paper.

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:48:34.451321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T21:46:43.353305Z digest=sha256:ce94d85b944d189f965c0c9a22b3a9b81a2b692a6cf9aecd582a23ee0023fe78

Observation 576ba46b-a0b7-4efe-a2b4-84cfe3b974c7 · inbound

Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model cites this paper.

Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T01:35:37.877410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T01:35:37.817082Z digest=sha256:2020110501c6db868aa342c4da887c1784eede4de49c410c0e5cd8f73ffd128f

Observation 11c20f79-fd10-4c1e-b34b-bba0579a3fc9 · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.736283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.736283Z digest=sha256:50a4a2bdcba6dc928520145586d8a7f78c6ae5fa2808ae4b494b9e7c84095703

Observation 9024fe7d-5d1e-4016-acdd-90ad88f4b1ef · inbound

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling cites this paper.

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:29:56.389990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T14:29:56.348733Z digest=sha256:b137d712161b7b021ed99fee2e03dbce31219d9e16bb95fb2590090e55eb3a16

Observation f637feb6-232e-4e94-a8d4-c2a9a8ce9ad6 · inbound

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling cites this paper.

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T16:11:12.825647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:11:12.825647Z digest=sha256:97b8b41d60ca5f3e3c0568eb0df1325a36ea369aa9dec0d354c89188f2d09b04

Observation e2675f78-181a-42c4-8c29-b475e1bfaad9 · inbound

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning cites this paper.

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-21T16:44:16.012591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T16:43:11.995960Z digest=sha256:0e8c2f99f28e6b29cd81e0fca128e90ede90bc83c69575c8fc774a45ffb293e1

Observation 0662f10a-f8a5-4274-83de-6542cceca7fa · inbound

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling cites this paper.

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T15:46:08.841274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:46:08.841274Z digest=sha256:ac6f1c74ff7bf926e67a3e69a032af1bc8e9158048db918210b551d3ab43bbc9

Observation d95b703e-7b6c-4353-a75d-f6217517ad01 · inbound

Large Video Planner Enables Generalizable Robot Control cites this paper.

Large Video Planner Enables Generalizable Robot Control HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:28:33.943330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T21:26:32.048309Z digest=sha256:5698f22e00353de79088c1536b52fd93db72ef9ac80a9370aee264e9e0792934

Observation ce27b337-1791-4b72-9574-29ba28556629 · inbound

CustomX: Unified Character, Action, and Scene Customization in Video World Models cites this paper.

CustomX: Unified Character, Action, and Scene Customization in Video World Models HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:53.569578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:53.569578Z digest=sha256:25c89a806938e552f289849f443a7f715cc0d96c7c24407caba4094ce0427674

Observation 5072f660-e94e-4547-84f7-5b4202deaca8 · inbound

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection cites this paper.

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:08:22.948024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T20:06:30.109866Z digest=sha256:b2ff6d937b88a7c46b57c8bfa115260f7e0ce939e8d55afb167bb0ef8bea5a2d

Observation 701615fe-de36-4a32-b6f8-a275d82f03d0 · inbound

Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking cites this paper.

Probing and Leveraging Video Diffusion Transformer Features for Robust Point Tracking HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:49.944520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:24:49.944520Z digest=sha256:9ec1fd957ea2cc1656d51ac8d50aa3ef3cde14e5526178ea4ff666d1f1c93bbd

Observation fb6c9ec9-1add-43d7-9563-6a433f84bd5c · inbound

GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure cites this paper.

GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:11:13.702288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T20:08:23.319013Z digest=sha256:c307f06276662f500fc2950b418639dc82b2bdee98f531595c3d7b783572530a

Observation e50ab9c7-99be-4212-886c-6d84e069dc1a · inbound

GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure cites this paper.

GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T14:09:40.361628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:09:40.361628Z digest=sha256:03d88a689093008229ca3577a6b064c1fb5ca28e7073bc696432ff288b8c2a77

Observation 94a02d24-31b0-48ac-9755-b53912b3faa5 · inbound

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World cites this paper.

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:38:21.039279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:34:39.518649Z digest=sha256:a69280a5ef6071fa869a64ab78cbc06621a0dad86a45c29a3fa1ebdb83f21261

Observation e7a8a474-bba7-4b3d-a4fe-14427df89878 · inbound

TinyHistory: Lightweight Video History Embeddings via Two-Stage Context Learning cites this paper.

TinyHistory: Lightweight Video History Embeddings via Two-Stage Context Learning HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T13:35:49.942787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:35:49.942787Z digest=sha256:6eada576284c2cc6b856a41aa5f7c4fa84873a6f3c65feb8cdb8e95b8a4b1f6c

Observation aafbc018-a56c-440c-b588-7b13d9fff2a4 · inbound

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation cites this paper.

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T13:24:45.974419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:24:45.974419Z digest=sha256:8c77604ea644d6e9ac82485f373fc3fbfd8af2182e207f3a422b9f0c321fde23