Pith. sign in

Paper Citation Record · LEDGER

HunyuanVideo 1.5 Technical Report

As of 25 July 2026, this Paper Citation Record lists 22 of 22 outbound references and 92 inbound Pith citation observations for arXiv:2511.18870.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.18870 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T02:31:34.927101Z

measured 114 of 114 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-24T06:31:00.690269+00:00

measured 92 of 92 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T22:56:23.337858Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:46:40.932396Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact9
  • verified fuzzy12
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abbb5b02-510e-438a-a4ef-898e49fa6929 · outbound

This paper cites Kling 2.5 turbo.

HunyuanVideo 1.5 Technical Report Kling 2.5 turbo

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T02:31:35.099816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:31f1dd0a3b4a36e12b09ee0c2b4553c72bd74a261903870aea73fafd00d2b152

Observation ca220547-9405-4cdc-8e26-0a2721c40d37 · outbound

This paper cites Veo 3.1.https://deepmind.google/technologies/veo/.

HunyuanVideo 1.5 Technical Report Veo 3.1.https://deepmind.google/technologies/veo/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T02:31:35.104706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:11927e448fb942e7eeffabe12beeb30a676050b0d7be5b0696795d78da928b26

Observation b813c304-dc66-4bef-b07f-db5b610e8e5f · outbound

This paper cites Sora 2.https://openai.com/sora.

HunyuanVideo 1.5 Technical Report Sora 2.https://openai.com/sora

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T02:31:35.107218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:a4d5d453b6f35e2994a23aebfe89d2755d2b8c5cd4ad68ec5ad2daddf51579b6

Observation 2822162d-e4c3-4759-a83f-4f402d5bb595 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

HunyuanVideo 1.5 Technical Report HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:31:35.082213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:26f9f4031c8185130cc25b2beba1f9aa3153406ee3f31f116958d6c04084a7c3

Observation e52145d8-551f-44ff-963f-decce4ea2148 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

HunyuanVideo 1.5 Technical Report Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:0ee0b80e06407e146c212243eb25345b8df8bad3813fae46dd162df306b274c9

Observation 7ceba260-f889-4af7-9865-9bdaa8037902 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

HunyuanVideo 1.5 Technical Report Wan: Open and Advanced Large-Scale Video Generative Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:31:35.094328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:d577485c4e015ed01ae0e1b818da52810adef912549fca3b6d2c5793882901c9

Observation 0f1a2b6a-3f42-4e8a-9a11-91018c7f07c5 · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

HunyuanVideo 1.5 Technical Report FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:45:36.571697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:6ef9074191222aeae505bc68da0ade3b83f1ae72abd7f814a1173536d93b2047

Observation 4a51cd60-77dc-4f30-bc7d-bfa46a4f0e87 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

HunyuanVideo 1.5 Technical Report Kimi K2: Open Agentic Intelligence

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:31:35.074047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:2917f0f3effc374832b5754524b744dfa35602e18e0e26e03085dd69f17ca8e3

Observation b5cc9bd6-3b26-44b8-9c17-2fe92737ee70 · outbound

This paper cites HunyuanImage 3.0 Technical Report.

HunyuanVideo 1.5 Technical Report HunyuanImage 3.0 Technical Report

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:02:32.945705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:dd5f4144b07ae1704a01fcb9303b2147843c1467dc050720f4967db2df1040ad

Observation c6dc4143-9cdb-4a32-ba40-18121d295eb3 · outbound

This paper cites Pyscenedetect: A python library for video scene detection.

HunyuanVideo 1.5 Technical Report Pyscenedetect: A python library for video scene detection

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T02:31:35.109670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:d961b210255aae7d0f274fe4624652b98343a7eaec4d793c2c8836c08fc92265

Observation da1cfeba-56a3-4515-8532-c9121ddf667d · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives.

HunyuanVideo 1.5 Technical Report Exploring video quality assessment on user generated contents from aesthetic and technical perspectives

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T02:31:35.112196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:798a066a31e2dc75ceaef73e4b4e7768946e26ecc97500552b9b15507f00f262

Observation e679c718-595f-438a-9d05-a23ab068ce7a · outbound

This paper cites Mitigating hallucina- tions in large vision-language models via dpo: On-policy data hold the key.

HunyuanVideo 1.5 Technical Report Mitigating hallucina- tions in large vision-language models via dpo: On-policy data hold the key

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T02:31:35.114927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:bbd26cc04041026951f333c6d83d15c411de0641ab2149c35d78f6740ffdc475

Observation d0078e77-e31e-4181-b3f8-9357aed782e5 · outbound

This paper cites Qwen2.5-vl technical report.

HunyuanVideo 1.5 Technical Report Qwen2.5-vl technical report

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T02:31:35.117579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:a789895d646bc6334a93d713218febc9c201c9e19976757c5c79c66944087cc5

Observation ca0c21a5-5328-43db-919d-1a6f8851cd57 · outbound

This paper cites Qwen2.5-VL Technical Report.

HunyuanVideo 1.5 Technical Report Qwen2.5-VL Technical Report

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:31:35.061741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:283b3dcc3941bc6dc2162c8f2542170de11c2051585c7de45f37c5a5d3224936

Observation 718f4ae6-c5c1-403f-8f50-088c7a54fdc7 · outbound

This paper cites Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering.

HunyuanVideo 1.5 Technical Report Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.065831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:83c8ffc5d2752bd64545daf346833770f0fa1ad29b52977b3dd78fbdfb2f3d32

Observation 4eac0e83-1ee5-4226-ad6d-468d9fb9bdeb · outbound

This paper cites Fast Video Generation with Sliding Tile Attention.

HunyuanVideo 1.5 Technical Report Fast Video Generation with Sliding Tile Attention

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:31:35.070439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:8b8147f068e785fa526c394ac9ef75a18b07ce45997966d27eb050da290e35d6

Observation b096ec8e-31c9-4b26-af91-c8997878217d · outbound

This paper cites flex-block-attn: an efficient block sparse attention communication library.

HunyuanVideo 1.5 Technical Report flex-block-attn: an efficient block sparse attention communication library

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T02:31:35.119796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:cbe0986d04bbe9a57949dafc1d6bbbd5a0cded0b3f46ef45d686ef7a3789b8f2

Observation 74f98148-9944-40e3-a888-474eb78c1940 · outbound

This paper cites Spector, Simran Arora, Aaryan Singhal, Daniel Y.

HunyuanVideo 1.5 Technical Report Spector, Simran Arora, Aaryan Singhal, Daniel Y

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T02:31:35.122078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:670b5e63904ccd57c3108df9124e6d122e036e3381a17dcf6014de9b00c31d04

Observation 1f51a624-17a1-4ad0-987f-482923111294 · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

HunyuanVideo 1.5 Technical Report Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T02:31:35.124922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:f6386738592df038737dd3bd22ef517fbb64ac8e1b77cf0b788a230d664d6750

Observation a53b66a3-d14b-4c2e-9068-81b22b044828 · outbound

This paper cites MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE.

HunyuanVideo 1.5 Technical Report MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:31:35.086217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:a35798b687ff6dd119358efe0fd4f0ce6899ce0e42a16e7694356d9cc2fdaa76

Observation a38b6fba-204b-43d4-b002-30c1d75d7fcd · outbound

This paper cites Diffusion model alignment using direct preference optimization.

HunyuanVideo 1.5 Technical Report Diffusion model alignment using direct preference optimization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T02:31:35.097504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:db6d951971e8022a513384ea881b785056da270723d21f1821756310bb4abfd0

Observation dde183c8-8d77-4d11-98eb-8c5fb64de53f · outbound

This paper cites Sageatten- tion2: Efficient attention with thorough outlier smoothing and per-thread int4 quantization.

HunyuanVideo 1.5 Technical Report Sageatten- tion2: Efficient attention with thorough outlier smoothing and per-thread int4 quantization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T02:31:35.102097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T02:31:34.927101Z digest=sha256:e0cab67fc8fca6f6cb78e71ce92b49128abdc13797eaee5d9b94697b424ef3d6

Pith citing papers

Observation c3c14651-8f78-453b-a76e-1ae5ee9090de · inbound

Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model cites this paper.

Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model HunyuanVideo 1.5 Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T01:35:37.934166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T01:35:37.817082Z digest=sha256:6e5ab9afcbbec3165ff2bae678416f160d40f5828943c440ed7850be61d717fc

Observation ea475063-bd54-4d56-8573-ace881f7231f · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation HunyuanVideo 1.5 Technical Report

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:48:21.886862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T19:43:37.604351Z digest=sha256:fd063b0e6c9bd68000b2be005c448c19baeffcfcfa64bfb33e1f3208c72e4ad0

Observation e24e5d36-135a-496a-b8ca-199d364dc033 · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation HunyuanVideo 1.5 Technical Report

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-21T16:14:15.305066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-21T16:10:31.015783Z digest=sha256:b18245c8fa2a86aef085e9ccc079398c1fcfb913df630d64671d550399f81b20

Observation dd9ccffc-1869-4ec5-93e5-694d55861d8a · inbound

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization cites this paper.

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization HunyuanVideo 1.5 Technical Report

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:00:44.696647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-16T07:58:24.456859Z digest=sha256:9904569cc7efda046f54e75fce0f9ed145e595a48d5e513a40dc35640cb5c6e5

Observation 92020938-9e8a-4ebd-9fa8-15238fff0e27 · inbound

Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints cites this paper.

Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints HunyuanVideo 1.5 Technical Report

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T12:20:00.572071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T12:18:45.538658Z digest=sha256:5a681cc5e78ba47447a5936f95679d580b3e9bad006c7c30ebcf498f74e9c319

Observation fe7062fc-2f92-4826-8fe2-0881415960f2 · inbound

Event-Driven Video Generation cites this paper.

Event-Driven Video Generation HunyuanVideo 1.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T22:56:23.337858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:56:23.337858Z digest=sha256:56cad383b5cde772938f041db07497deb0e3d095fbf82f93829b4d65c6df547d

Observation 320f5560-124d-4af2-9cc2-49ad315cf1a4 · inbound

AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization cites this paper.

AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization HunyuanVideo 1.5 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T23:09:50.674516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:09:50.674516Z digest=sha256:509920082f36d91b0319b6b3cc15a157f18534aef38ef754e695a98de9cc7a6f

Observation b31b73c2-245a-44cf-a9cb-8580ac71cb2d · inbound

AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors cites this paper.

AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors HunyuanVideo 1.5 Technical Report

Reference 132

Resolution
unresolved
no resolver link, observed 2026-07-13T22:49:03.259461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:49:03.259461Z digest=sha256:a852bae34f667bbc0d1a8a256314a2cfd7f7f73f739f1fcca65c2b1b75f93ee1

Observation 043c7744-cfda-4d64-b20c-5d9d47f0f56b · inbound

RefAlign: Representation Alignment for Reference-to-Video Generation cites this paper.

RefAlign: Representation Alignment for Reference-to-Video Generation HunyuanVideo 1.5 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T18:01:19.034578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:01:19.034578Z digest=sha256:87b9df4b2a2221c2b4197c5409c8cca41b7fa21fb72410664d8cc599dd3112bf

Observation 894e4985-5e14-452c-9bec-8a3323f95371 · inbound

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms cites this paper.

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms HunyuanVideo 1.5 Technical Report

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-14T01:35:14.878069Z digest=sha256:b4d2ee9759a614596090e5bb2d0fba2b9bf17823bdd3acb5be5ab7a01762aee6

Observation 65b25897-9858-4db5-b300-f4241f263d40 · inbound

MMPhysVideo: Scaling Physical Plausibility in Video Generation via Joint Multimodal Modeling cites this paper.

MMPhysVideo: Scaling Physical Plausibility in Video Generation via Joint Multimodal Modeling HunyuanVideo 1.5 Technical Report

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-13T20:20:16.538920Z digest=sha256:d3d389369141845c69ae82f5e61378b223294ab38f37d16430e76349f96f5910

Observation 23c06a9f-cea7-417d-8a06-53492e61d630 · inbound

InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation cites this paper.

InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation HunyuanVideo 1.5 Technical Report

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T17:39:59.791758Z digest=sha256:6cf241703ede2745cbcd8f83e4d24a7762f83e9a2ed1585b2414154ca150089f

Observation b9539934-7b00-4ba7-b8a8-243a28e219ed · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HunyuanVideo 1.5 Technical Report

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:4ddb996409de14f49261ec5757864793e83db7c61f5bbd448609dde6c3dd1881

Observation 2549a6e4-658e-4178-9629-ab67c34c697e · inbound

AnimationBench: Are Video Models Good at Character-Centric Animation? cites this paper.

AnimationBench: Are Video Models Good at Character-Centric Animation? HunyuanVideo 1.5 Technical Report

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T11:47:27.131584Z digest=sha256:a327b5b369c5f521c96e27d097e4f8820afcdd2f75ca3d3115c2a43cfc08c307

Observation dbfd3aa6-ccf0-4b84-a10c-9040a5ecd5e8 · inbound

Efficient Video Diffusion Models: Advancements and Challenges cites this paper.

Efficient Video Diffusion Models: Advancements and Challenges HunyuanVideo 1.5 Technical Report

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T08:28:29.706249Z digest=sha256:dac0b1d30cff8c618705f854d913ef29fcc113ce0e1a34e1f3563db0a7f9e26e

Observation cebacc3b-6b36-4b0b-b7e7-2feb1f3b4bb8 · inbound

Motif-Video 2B: Technical Report cites this paper.

Motif-Video 2B: Technical Report HunyuanVideo 1.5 Technical Report

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T15:50:04.001693Z digest=sha256:f9a0aeb2dab96f10962418a4bc6abf65d143ce5d7602e48697f6888b17d67cee

Observation d8d1b169-a855-49c2-a604-3fdfd7e9485e · inbound

Motif-Video 2B: Technical Report cites this paper.

Motif-Video 2B: Technical Report HunyuanVideo 1.5 Technical Report

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-21T00:13:53.056813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-21T00:12:23.096146Z digest=sha256:37b6cffff165bba086168212686d28b3b650dd2b0d80a0f66d5d850ef1c014e6

Observation 2a7e5d68-3cbe-4502-a603-69b025ad4836 · inbound

Grokking of Diffusion Models: Case Study on Modular Addition cites this paper.

Grokking of Diffusion Models: Case Study on Modular Addition HunyuanVideo 1.5 Technical Report

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T05:28:08.221886Z digest=sha256:29095c77db586876ddcc2edf92aa27fbd78797ebdb603dd411d16b3a7ef7041c

Observation eba4af88-2a4d-4965-98be-67e626308349 · inbound

GS-STVSR: Ultra-Efficient Continuous Spatio-Temporal Video Super-Resolution via 2D Gaussian Splatting cites this paper.

GS-STVSR: Ultra-Efficient Continuous Spatio-Temporal Video Super-Resolution via 2D Gaussian Splatting HunyuanVideo 1.5 Technical Report

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T04:46:39.643664Z digest=sha256:5fc259d63334dc2bfae7105c824c82bf4c33fbf6ff24e8767a3dfce74fbcaf4b

Observation 8101dc12-79ce-4ad2-9ed3-8358d041fe1d · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation HunyuanVideo 1.5 Technical Report

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T03:02:26.084185Z digest=sha256:a50061d857716261362fe92a80708dd5cc6ea6c135bf328fdbf96c3a843bef49

Observation aea06abc-2a99-47b2-b350-2f0f81b4f5bb · inbound

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation cites this paper.

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation HunyuanVideo 1.5 Technical Report

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:19:50.189549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T06:15:32.881140Z digest=sha256:c1df1d22ec68be091c3f38d89909270cfb994785cd32c209de0f4530f36d65f2

Observation 06ff4caa-f75b-473e-bc92-ddbc2fb7deae · inbound

How Far Are Video Models from True Multimodal Reasoning? cites this paper.

How Far Are Video Models from True Multimodal Reasoning? HunyuanVideo 1.5 Technical Report

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:a4640c7c338b8015256085583537fbacb119a4ee72f1d62891e32115bbc7d685

Observation 69810b9d-620a-4e03-94dd-99a7dd562332 · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation HunyuanVideo 1.5 Technical Report

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:1246510f257198d64548c553a2ce561f919020b36c477e91a0aed40c047554cc

Observation c7755a19-b87e-430a-8478-cf9a389f5ba3 · inbound

HumanScore: Benchmarking Human Motions in Generated Videos cites this paper.

HumanScore: Benchmarking Human Motions in Generated Videos HunyuanVideo 1.5 Technical Report

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-10T01:17:16.513512Z digest=sha256:58a97f53646fd09200d01839fe665e6f72a77e365bc4dc3d3dbf6056b7743b62

Observation 14281051-45a9-478f-94e8-38d542af2b94 · inbound

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE cites this paper.

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE HunyuanVideo 1.5 Technical Report

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-08T18:26:58.696936Z digest=sha256:79aef9be54e1a07556a6a0205377391b2ba6519d4bda7ce6ceb08ca31005fe86

Observation a9579c60-9668-48c0-b4a7-66cbc4e9aa29 · inbound

LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention cites this paper.

LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention HunyuanVideo 1.5 Technical Report

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:25:09.391588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=arxiv_source observed=2026-07-01T00:24:30.573122Z digest=sha256:b7ff9d02bb5f24515bbd52d4f5a83abf994e9111858140e8c0b9aae960a19ef7

Observation 38408528-28da-4693-9509-d3f009e5d03a · inbound

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation cites this paper.

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation HunyuanVideo 1.5 Technical Report

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-08T13:28:05.540719Z digest=sha256:26754ebb6dfc1765b3742b1715d40f2d5f4a5a03b43b1b61b2e5ea6c7037f359

Observation bb43f715-51cf-4357-aea3-39246c09534e · inbound

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation cites this paper.

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation HunyuanVideo 1.5 Technical Report

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-12T01:50:30.588220Z digest=sha256:2e276efe5f1c8dfbc8e895a1b39479b4e57b530887cc9e2d9eb0b98ed337447d

Observation db06403d-836f-438b-8591-fb889395c645 · inbound

DCR: Counterfactual Attractor Guidance for Rare Compositional Generation cites this paper.

DCR: Counterfactual Attractor Guidance for Rare Compositional Generation HunyuanVideo 1.5 Technical Report

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-08T13:06:56.964670Z digest=sha256:d9cead26da7f6405802ff11cd74cc8a10cdfaf0adfbdd689aa510ca8e35a21d9

Observation 03ebb172-73e1-4737-a43b-41d0ead91a9f · inbound

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models cites this paper.

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models HunyuanVideo 1.5 Technical Report

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-11T02:21:52.861714Z digest=sha256:3cdd77b7185feaa59cef79d9abeca9777f1f3ea617c9a5449976f1f1606874f7

Observation 0324c496-34fd-4377-94d0-8098f2cc8e37 · inbound

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models cites this paper.

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models HunyuanVideo 1.5 Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:15:07.981524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-30T23:13:31.195562Z digest=sha256:2b39f8dbe6fde4d3be27163072d308a07a21a44de80963b0f08c8d9864cfa898

Observation 05854b68-a6f5-4c83-90d4-ebf0765c6b90 · inbound

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm cites this paper.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm HunyuanVideo 1.5 Technical Report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-13T06:02:39.114660Z digest=sha256:e30eb4275f1c59ce2e6d029cf85d465b7cd3583b8573ca9a97a464cc4622539c

Observation 7378e38a-87a4-4805-80b1-215eec5cfcbb · inbound

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm cites this paper.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm HunyuanVideo 1.5 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.675275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:587b34beb29c99d7a0982b9de4283b6eb237539e16f99f17fb3738f5bc71979f

Observation 7a749445-6ff2-4355-a954-171b4f1d4dc4 · inbound

Qwen-Image-VAE-2.0 Technical Report cites this paper.

Qwen-Image-VAE-2.0 Technical Report HunyuanVideo 1.5 Technical Report

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-14T19:53:25.849049Z digest=sha256:e9d08c1dd25bb9c516d6dbf1db98310942fc840738656706e30e42a2ca6309d6

Observation 116f539c-be3c-430f-8b23-5da29894b766 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation HunyuanVideo 1.5 Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-15T01:52:14.874049Z digest=sha256:f8e61067ee62aa0384ece76378eadc3e3f47b9c8cc30755deda4ad3a8390120c

Observation 4e8a8b6e-85d7-4865-878c-2cf9b72ee414 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation HunyuanVideo 1.5 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:59:06.473511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-20T21:54:33.902256Z digest=sha256:621c1381980c7e7e4da509bcc88b0e800f434b6a5a39a8ddc75b35d0d28e2fcf

Observation 4180eef5-2db7-4a8a-9a1a-52276d6857a5 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation HunyuanVideo 1.5 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:14:05.603127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-21T09:12:35.777810Z digest=sha256:562655541d3fe726ba10cefbd093ac53ea59daf501dd7e989dca9e12fa35f95b

Observation 35981e8a-d7a5-4290-b0e2-875d5b4c2c69 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation HunyuanVideo 1.5 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:45:05.487201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-30T21:44:21.427570Z digest=sha256:b4aae79f33888d8eeda8ce2d4e57b563e6e7467f15224f72d36e2d948d73c6b8

Observation 731143ad-0172-4744-a0c9-9f6aa2021052 · inbound

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO cites this paper.

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO HunyuanVideo 1.5 Technical Report

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:35:47.239870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-30T20:38:28.328436Z digest=sha256:eebc7accb51728c61d71ccfb75b0ddb1d5aaca55ceadd5bad8e783ec57ce5542

Observation 533edb80-44f4-4873-b739-37ecd2602d54 · inbound

Video Models Can Reason with Verifiable Rewards cites this paper.

Video Models Can Reason with Verifiable Rewards HunyuanVideo 1.5 Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:07:37.149746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-19T15:03:14.894952Z digest=sha256:dd902117e4fd5e1e5ffb39a315ca9c47d31a2bf4e111fdbe6d022eeaa97e6e9d

Observation 20a04c84-8abd-4f68-aa9c-21c6a07dd67b · inbound

WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation cites this paper.

WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation HunyuanVideo 1.5 Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-20T17:48:48.678477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-20T17:47:32.953903Z digest=sha256:ed3b00a54a82b7d6c41bcfa3159869c52f3ab3e41a81271a061f6530f9be75c8

Observation 9c94c22f-7410-44c9-9f72-acc05bc989a6 · inbound

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization cites this paper.

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization HunyuanVideo 1.5 Technical Report

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:28:52.886558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-20T18:28:06.253200Z digest=sha256:4a7dcb265cbf4492566af3d595b65a70342308397b944506b283dde5ce5c8521

Observation 4335b79f-d33f-41f7-82b7-546c41e35a2b · inbound

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization cites this paper.

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization HunyuanVideo 1.5 Technical Report

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:45:00.993154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-30T19:37:18.563704Z digest=sha256:9e79cfcc4aad6618cb4613941955d167f6725eca88f605dd9cd3c697cdd044b1

Observation e5c4828c-7aaf-4598-b6b4-797770656d9a · inbound

AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling cites this paper.

AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling HunyuanVideo 1.5 Technical Report

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T18:28:53.048976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-20T18:26:57.735246Z digest=sha256:0b5717d6f3948108244c256dad9ec683a9648d2f5b7042eb253d1d55a86d662f

Observation 3790b864-45ec-45d6-9c9a-6e2571fc5b43 · inbound

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration cites this paper.

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration HunyuanVideo 1.5 Technical Report

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:18:18.296922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-20T13:15:54.413960Z digest=sha256:63e815c37ffe430ef4dc9b47002e207cbfd91e5f031d60a122d2ec5155024502

Observation 42345b3c-f0ed-406d-9511-7115948825bd · inbound

Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion cites this paper.

Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion HunyuanVideo 1.5 Technical Report

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:23:14.172519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-20T11:19:21.721494Z digest=sha256:14c2186be071270ec89d7e18e0c16c5133d4dd8c4c73408f2aa748b196fe3662

Observation f0377d96-986e-4d56-8eaa-8b19bde9ce0e · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy HunyuanVideo 1.5 Technical Report

Reference 122

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:48:14.793795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-20T11:46:52.658984Z digest=sha256:3c8b3685e057fbe609828077ed010738b2c5394053be3302ebba3c6c8a992f4b

Observation 0d136636-b823-402b-817a-141c3379105c · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy HunyuanVideo 1.5 Technical Report

Reference 123

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.493452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:01bb1b5dfdb7a5bd7050fad60918c9ce21993a4411d8df1042c5b9431fedf119

Observation d55ed59d-2b9f-486a-b9dd-11f512311cee · inbound

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models cites this paper.

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models HunyuanVideo 1.5 Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:58:04.987325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-20T05:57:44.096941Z digest=sha256:38690908a28c151dafafc5ace93e730493b4a72cbb20536490a9eb33c3199451

Observation c2d581a7-e343-40c7-b4c4-00abce1b1f0c · inbound

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models cites this paper.

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models HunyuanVideo 1.5 Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.816126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-21T07:55:40.616049Z digest=sha256:ed597ecde4eb1adab05167262582359a85f5e651f1064c8e6c920fa1f8222d43

Observation 7d702089-65c2-43e6-a76a-6ac12c768e73 · inbound

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models cites this paper.

Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models HunyuanVideo 1.5 Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:35:00.134242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-30T18:33:44.480873Z digest=sha256:0de190dfde9a331170c06746038f1147f00a72b852b56ccc31040ec30a563dd1

Observation 11860ab0-da1c-4fc4-b38f-776e703622a9 · inbound

Dynamic Video Generation: Shaping Video Generation Across Time and Space cites this paper.

Dynamic Video Generation: Shaping Video Generation Across Time and Space HunyuanVideo 1.5 Technical Report

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:43:58.833433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-21T05:42:41.925474Z digest=sha256:b066eb418e6851e9f7ed605141b58a9652d68e69ffcc9dfdb82b153053c49d1d

Observation a7a83167-a7b5-4fea-8bc1-4c19df01d55f · inbound

Q-ARVD: Quantizing Autoregressive Video Diffusion Models cites this paper.

Q-ARVD: Quantizing Autoregressive Video Diffusion Models HunyuanVideo 1.5 Technical Report

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T05:29:39.498841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-21T05:29:04.061943Z digest=sha256:4a46f6a42ea2caf3e1bbd8e79d16c25951bb4364bdfd70bfa421ddf3ec2911d6

Observation cadac9a3-cb6d-4ad6-bf7b-8cd27e5ed40f · inbound

Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework cites this paper.

Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework HunyuanVideo 1.5 Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:35:21.950304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-25T04:30:20.593882Z digest=sha256:b84b1d53c3bbed8cfa40660be1600731bd1f56819fb25491a3a36d571251f619

Observation 92a5747c-166e-44e6-8b77-cfdf55f6c109 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap HunyuanVideo 1.5 Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.962319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:2aa702587ed1a3563d8433a50b77c381cef6a2472d2dcad634da978637fc93da

Observation fd5da662-09d6-472c-bc06-3e42efb75423 · inbound

PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution cites this paper.

PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution HunyuanVideo 1.5 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:34:02.336354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T22:24:50.679950Z digest=sha256:1c21ac732596d2b3f476f00f178bca445a554f546079d120299b541ed0965fb2

Observation 9053dbd1-5cce-43c2-936a-35aa59f5743c · inbound

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV cites this paper.

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV HunyuanVideo 1.5 Technical Report

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:54:00.770105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T22:52:38.330851Z digest=sha256:7fe885e30aa213b815e2c6485e9f24caf5b2ca671396270bd43f3ed3aa4c17fa

Observation 02e07769-2bae-43ab-9b93-72549519475a · inbound

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios cites this paper.

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios HunyuanVideo 1.5 Technical Report

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.370134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T18:23:22.987086Z digest=sha256:8fd433e93ceec2a41b8fdf76a5f41705d1cc4c5675a44ac7d0932175ff598dbc

Observation 439620f1-8ac6-4676-949b-d140e014f53d · inbound

Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation cites this paper.

Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation HunyuanVideo 1.5 Technical Report

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.614749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T12:55:24.689338Z digest=sha256:2e8a78170b69fc83e97656319d14d02fd1f7cfab11b073e5aa1a3c8c8873e66f

Observation cefbdf58-6d8b-4293-aa95-26680fe87449 · inbound

OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning cites this paper.

OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning HunyuanVideo 1.5 Technical Report

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:13:27.511276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T13:06:02.201607Z digest=sha256:eaf9a16efcc432d868db333316824edfde488398cda46fe5847ba77b03fd2296

Observation 893374a4-eee2-4e68-b34c-4b7a94016c6a · inbound

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation cites this paper.

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation HunyuanVideo 1.5 Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:33:15.209905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=arxiv_source observed=2026-06-29T08:30:00.438334Z digest=sha256:0802682fb1c05cab7e15a3f8922d076170ad24e17cc3445a9b7f5ff4c27d56cf

Observation 9eb56e12-8f3e-4051-bb77-598ad2a43bca · inbound

SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation cites this paper.

SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation HunyuanVideo 1.5 Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:23:15.593683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T08:16:26.055531Z digest=sha256:ce719f86553fefb39b7015f6122005f6a4d96fb44593f314f2f95b08170a6d53

Observation a3309dd5-dde4-48af-8b5b-c94f2b2e4eca · inbound

Veda: Scalable Video Diffusion via Distilled Sparse Attention cites this paper.

Veda: Scalable Video Diffusion via Distilled Sparse Attention HunyuanVideo 1.5 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.684920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T07:54:13.390911Z digest=sha256:eee8aa6403f91e9c7a3d30610287b4c834aec9a85782172529b140c5b2eabfea

Observation bd9a3eb0-f8fc-451d-be40-fe037314d86e · inbound

OptiWorld: Optimal Control for Video World Generation under Physical Constraints cites this paper.

OptiWorld: Optimal Control for Video World Generation under Physical Constraints HunyuanVideo 1.5 Technical Report

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:32:35.501898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-28T19:02:51.848742Z digest=sha256:f6cf9e5b1f20328e370db56ab250024d46b873468b8ac614503bebc74352bb1a

Observation 3687c8b7-7fb8-4512-b102-88a8532d02cb · inbound

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning cites this paper.

ImagineUAV: Aerial Vision-Language Navigation via World-Action Modeling and Kinodynamic Planning HunyuanVideo 1.5 Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:14.921309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-28T17:13:10.713616Z digest=sha256:0bc57e545df4d4eb2e69f6c7ba219aedbe2814e26f6597e7cf9644c45ea9d169

Observation 3994fd89-0ad8-410d-a5e0-b3a34d81fb3c · inbound

Knowledge-Intensive Video Generation cites this paper.

Knowledge-Intensive Video Generation HunyuanVideo 1.5 Technical Report

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:06:13.328812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-28T17:33:07.941924Z digest=sha256:a08df0cb9a6ab5edfd19e0c4b71f17d670053235288ac44eacdc5f5fdf722187

Observation 85443a78-47eb-44f9-976d-5ebfaaff036d · inbound

Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition cites this paper.

Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition HunyuanVideo 1.5 Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:16:16.541461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-28T15:33:30.182150Z digest=sha256:dfbdf69f0add0fb3bcd18f074b3d918fe5ad7a229350d7aaddc65b320cd4bbe2

Observation 44c5b63b-14a6-490a-b656-14428afb31ea · inbound

Jailbreaking Multimodal Large Language Models using Multi-Clip Video cites this paper.

Jailbreaking Multimodal Large Language Models using Multi-Clip Video HunyuanVideo 1.5 Technical Report

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:36:17.110021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-28T15:16:48.957645Z digest=sha256:ea05d6c085adadb7936f48c21b58007a8618a834f5c9a172309cc050ded68e52

Observation aebb9130-59e5-41ee-b521-7724d831771f · inbound

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation cites this paper.

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation HunyuanVideo 1.5 Technical Report

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.455204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-28T15:25:22.778550Z digest=sha256:db75f76347bedb24c5ab8082c654db9c453862d8e050801611146fb1ba18ad6c

Observation cbe2bc01-c21a-40de-8179-12492ca3b449 · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization HunyuanVideo 1.5 Technical Report

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.284174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-28T15:26:21.284810Z digest=sha256:7849b6858f6b4288af83033ba305f10e43bfd88f65262ab72837f7fe33b25067

Observation f089b9a7-c3f6-4acc-9969-812f1bfda836 · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization HunyuanVideo 1.5 Technical Report

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:44:36.863984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-30T10:38:22.619277Z digest=sha256:3214d09af9024284c6ef0ac084604c530f93a396d66cdf202bb9b855e9f16891

Observation f41710e2-ee9a-4359-b726-735e5a8ec7e7 · inbound

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO cites this paper.

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO HunyuanVideo 1.5 Technical Report

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T16:17:08.919704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-27T22:54:21.740892Z digest=sha256:09a34034f19c4e2b1e9420dc99ca6cd7c9a081e46cc040e73d3dd8dc58f9aac4

Observation 18a90366-213c-4a09-9293-5ecc600e4f47 · inbound

Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions cites this paper.

Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions HunyuanVideo 1.5 Technical Report

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:37:29.763218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=arxiv_source observed=2026-06-27T17:07:57.743780Z digest=sha256:681e9a45fa7de3c58a24be5fdc6c71de5239e0146c227eacf6ea7e632a8e80f1

Observation 1a5c88ee-4b45-416b-8bc4-aa873ba680dd · inbound

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation cites this paper.

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation HunyuanVideo 1.5 Technical Report

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:07:28.081103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-27T17:30:25.371658Z digest=sha256:6bb4505b5785d4fcc9d3eb978533666be879803414e89af3dd111600983cd740

Observation cfb34154-509c-417f-9c85-875e4035cc33 · inbound

SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation cites this paper.

SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation HunyuanVideo 1.5 Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.760937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-27T10:00:35.608696Z digest=sha256:5f191892a9984730eec1ba2aabef860ee18ad4db09494bd5f4c29064434b4e35

Observation 5f945e09-2b15-422a-8d1b-65c8d98dcf10 · inbound

World Model Self-Distillation: Training World Models to Solve General Tasks cites this paper.

World Model Self-Distillation: Training World Models to Solve General Tasks HunyuanVideo 1.5 Technical Report

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:07:55.835992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-27T10:16:35.511426Z digest=sha256:70e31d0f7ffdf0c021d0ed33e2fe8ae7e33b84280a43b572945bfda6acc16bfe

Observation 04453cd5-5eb9-4505-adda-7c43a1c49891 · inbound

Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering cites this paper.

Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering HunyuanVideo 1.5 Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:48:45.908421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-27T03:44:36.071488Z digest=sha256:5e5fe1d73516955790b16cea44f39045222c35010beaa749b40b7e6304ec88e9

Observation a1900dbd-2862-4953-90ca-7f747d6346cb · inbound

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model cites this paper.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model HunyuanVideo 1.5 Technical Report

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.745718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:aa3f253e9cf6c5c8406555a42e01177391884f0d6d180c2126445d4ac0416f6e

Observation 04000111-bb93-476d-8c54-cb4a4d5a7549 · inbound

SurgVista: Long-Horizon Surgical World Modeling with Plausible Instrument-Tissue Dynamics cites this paper.

SurgVista: Long-Horizon Surgical World Modeling with Plausible Instrument-Tissue Dynamics HunyuanVideo 1.5 Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:19:31.471412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-26T18:09:48.872541Z digest=sha256:f4bea0cb3ba5b219ff23edbd67840af6f0b7242dddc14c96471676e4b97189c4

Observation 4325faab-2047-45b1-88ca-6603811d1f2d · inbound

Robot Critics that Sweat the Small Stuff cites this paper.

Robot Critics that Sweat the Small Stuff HunyuanVideo 1.5 Technical Report

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:39:37.252551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-26T14:20:37.905355Z digest=sha256:c044dbe162118a5fe7b2139b48319f44574452ab8727efde0de97ad92eb7e494

Observation e9abb76b-3fa6-484b-8efb-32d7f8269478 · inbound

Compression and Retrieval: Implicit Memory Retrieval for Video World Models cites this paper.

Compression and Retrieval: Implicit Memory Retrieval for Video World Models HunyuanVideo 1.5 Technical Report

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:29:45.459303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-26T08:46:28.619102Z digest=sha256:9bf39d89b1f94864c207629a0deddc2a2153b3b8290a099757596a8ba17aae01

Observation d5299c41-1b7d-466b-90a6-c55cc07675be · inbound

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation cites this paper.

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation HunyuanVideo 1.5 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:49:41.553281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-26T11:05:44.972128Z digest=sha256:449091bd7c429b9f03ea249d62636ed4f076969615462e23bd90784602f5439d

Observation 982f6187-0bcc-4bdf-830e-0630c0ffa1a8 · inbound

GeoT2V-Bench: Benchmarking 3D Consistency in Text-to-Video Models via 3D Reconstruction cites this paper.

GeoT2V-Bench: Benchmarking 3D Consistency in Text-to-Video Models via 3D Reconstruction HunyuanVideo 1.5 Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:39:57.134619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-26T00:28:33.734545Z digest=sha256:763b5c065a0aebf39717c1a2fcaaa8374a19f3aae8f11955b68c57ff562ec19c

Observation 555ea8b6-a318-4410-ac26-1b030a4e1c37 · inbound

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation cites this paper.

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation HunyuanVideo 1.5 Technical Report

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.901700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T04:34:38.286863Z digest=sha256:e0ec82852540bcd6910b74f497edcdc948271a452c51585672f7c0e195e8f392

Observation 7385868b-901f-4737-9c47-64d2a2bce09d · inbound

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control cites this paper.

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control HunyuanVideo 1.5 Technical Report

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:24:21.605409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-30T07:18:20.501369Z digest=sha256:b1f3233312805c32d1be694a7e8baf3f36551b6e1f01c4eaf7a29ff8dd5d05dd

Observation 77cf3235-a7e9-4882-86ed-331e5a8d681b · inbound

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control cites this paper.

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control HunyuanVideo 1.5 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-14T17:03:01.432145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:03:01.432145Z digest=sha256:ba727ec4359368ed80c22159ce55e035ef0f8c50c928c0fa7cee4e9815b45fc1

Observation ab8d08a0-d781-4073-a655-4e8c70149b23 · inbound

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model cites this paper.

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model HunyuanVideo 1.5 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T01:59:48.820112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:59:48.820112Z digest=sha256:be2a736db073b8b3c6fe95918c7fd60400598e67eadfa3ae9cde395cb17f95db

Observation 713685f0-0768-4bc2-abdd-126c5f9b0bb2 · inbound

Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control cites this paper.

Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control HunyuanVideo 1.5 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T22:42:31.529313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:42:31.529313Z digest=sha256:72e16c916b284c19f9f1cc6c41d24eb5d30ea2d38c1f409085fff1d5e983a8a0

Observation 92dec690-f942-4ca3-bad2-b96f9f3fe8bc · inbound

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence cites this paper.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence HunyuanVideo 1.5 Technical Report

Reference 113

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.325480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:82bd6b5b0dcaf929b8e006b42f03ede116c117796247597f9e5cf2f1c1ea4802

Observation c0047a30-0832-4b0f-8df6-e7a5a46116af · inbound

OpenCoF: Learning to Reason Through Video Generation cites this paper.

OpenCoF: Learning to Reason Through Video Generation HunyuanVideo 1.5 Technical Report

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.933515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:78e0c1107f966fa7c24af6b4fe200f11b789d346290d9fc265000862fa3ed1d9

Observation ca15b249-4a57-4b0a-98c4-1f7588e0ada2 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation HunyuanVideo 1.5 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T01:59:43.167178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:59:43.167178Z digest=sha256:cf27cc09487444e3f8478f129762da59dd4e933186af4fcd8fa9f2feb4af7783

Observation 94f7c0cc-586d-470f-b013-4de1c34a86af · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation HunyuanVideo 1.5 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T15:10:13.757131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:10:13.757131Z digest=sha256:570fe4d4f296a01add7427e36308b8d3d66b39da7816c7a732e9b7dab337138a