Pith. sign in

Paper Citation Record · LEDGER

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

As of 10 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 19 inbound Pith citation observations for arXiv:2507.20939.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20939 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:12:41.078055Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:03:28.047281Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:49:30.275005Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09faad0a-b366-49d8-9cb6-2caf2d9a1b54 · outbound

This paper cites Qwen2.5-VL Technical Report.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.360764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.360764Z digest=sha256:d09462e96287a68209a1a2ef15562955eb6b0fcf772dadd12a08de79fd706356

Observation 7548790d-90d4-4c51-8750-4ca490328737 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.953417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.953417Z digest=sha256:4033c6f7f9f80c6a0a15bc4408dc9061485622bf10c6ca667a02f4d1c7cd7ffd

Observation 9b6c2843-4cbc-4135-a943-43880b395e63 · outbound

This paper cites EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.150961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.150961Z digest=sha256:468df03aea790437955c75b72251d2e92fec2baba6b30b22fbf47c1f9a8f8f52

Observation b4a0a645-4785-4ba4-a8bd-573947db35c1 · outbound

This paper cites Audio-Visual LLM for Video Understanding.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Audio-Visual LLM for Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.249511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.249511Z digest=sha256:640f09af397969ff2f88bcb37986c3e8e73ce749a3200ed7c2c74e5bc9cebd2e

Observation 121a0ef3-6a52-41a4-83b9-bcc9bd4d0e63 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.387502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.387502Z digest=sha256:22173b7be77ac9f0de019a547f2eb844caddc3e4ee01a2feac536aadc3a8dec6

Observation f9929eab-f34d-4ddf-99e4-742369172e99 · outbound

This paper cites video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.512824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.512824Z digest=sha256:82a7fc2d49912196a40bf1b69175eb71378c1e03f2560215a6b5bfadc3e4fece

Observation cd8c78b8-5b4e-411d-a750-cbaf2e8ec76f · outbound

This paper cites video- salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220,.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts video- salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.594915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.594915Z digest=sha256:c260fe1510803897c5cde2d19871127dddaec5169e5daf70b5192fda2089783d

Observation 36e2cfcd-788a-4d8c-85b5-7d2114fb4bea · outbound

This paper cites Kwai Keye-VL Technical Report.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Kwai Keye-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.675805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.675805Z digest=sha256:c1e2e9d670d1a75570b22c9be14ba9a8359c60db695a0b83dec38218c10ebe1b

Observation d90faea1-3e4d-4f84-9b8d-1851eb24d599 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.870798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.870798Z digest=sha256:54f220eff5340a224fa3bc7e21c609a03f0f19ca644fecd596eec6a302eff811

Observation 7e81c92a-5039-4fec-b214-1c1dc739921b · outbound

This paper cites Self-supervised product title rewrite for product listing ads.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Self-supervised product title rewrite for product listing ads

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:12:41.674677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:12:40.971685Z digest=sha256:5baf9d99107cc857f856399f6b81d397b733fcf1244243f61d4ec1ffcc5b168c

Observation 0330b3b7-885f-4f0f-b3ed-c34643bedbe1 · outbound

This paper cites Pre-trained language model based ranking in baidu search.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Pre-trained language model based ranking in baidu search

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:12:41.484033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T13:12:41.078055Z digest=sha256:e69da556cb6308aff269001d0937776d9b32f64272a600773f9ec0c01ab5dd54

Observation 19e9010d-dacf-4543-9715-976def34db69 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.428823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.428823Z digest=sha256:4b040110f528bb69e5043e58c944ac98774e14fe849276b186136ba7b4e232c4

Observation 83b99e30-bf69-4500-bbe8-019717c45630 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.878853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.878853Z digest=sha256:d7a9a3534f5595ec3188998ae2e1115a9dc1a5352ebdb803469caed87d4f4a57

Observation 8adc1884-5841-4d92-9623-872b8dd0785b · outbound

This paper cites CLIP2Video: Mastering Video-Text Retrieval via Image CLIP.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.766732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.766732Z digest=sha256:550d35e4c5bb140fa90108b55bbc4b6ca10fb7f2cfcb9f1ea99902cf89fad90c

Observation 4bb1c417-c44c-49eb-b3bd-7193a9df94e0 · outbound

This paper cites VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.063831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.063831Z digest=sha256:a948e6741ef8ddd574801c598d45b1094649f053771a2330463127e5319ad6f6

Observation 929118e3-4139-417a-8134-569457a97108 · outbound

This paper cites Qwen2.5-Omni Technical Report.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Qwen2.5-Omni Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.757999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.757999Z digest=sha256:c8580382fae1c6ad7430a17aac126c53de23fb50b3973fab6bc29db47a53a72c

Observation ef3698a5-928b-43b8-b734-8e05bca71011 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.499439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.499439Z digest=sha256:2d1328bda4b6a6ae4a9dda0ca209e101735544d3bf2d985ecadc713d1f40b868

Observation 17226eaf-2214-427b-b571-910ed26a187a · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.601562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.601562Z digest=sha256:b9aec29a9a85d3eb143694ac09549c5a296a5fc90edf909ee6e61dfea0d47f48

Observation 7da6f7e7-2ed8-4976-b4ca-a5e4d7ee1be7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.679158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.679158Z digest=sha256:327878e9793209b93ab1170f8d791de8d4d4e07ed432c82a1cd2ab8d74bbeef8

Pith citing papers

Observation d82c0732-c3fe-44a9-bd50-e5bb579ddea6 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:28.047281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:28.047281Z digest=sha256:c602b4d7efff24f21424cd0b29e5c7d0e2590cf755334aa2c5af3cbe67aff69c

Observation 016ec883-792f-4898-8ee7-42fd8c75a35d · inbound

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models cites this paper.

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:15.022646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:45:37.418493Z digest=sha256:47ec0a226ed56ef4634ce8e1b66af925b87bcb4a90d1dd9db6dc887f25dea65e

Observation 431cd667-1fe8-4634-a1e5-be3605422d64 · inbound

AdaTooler-V: Adaptive Tool-Use for Images and Videos cites this paper.

AdaTooler-V: Adaptive Tool-Use for Images and Videos ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:34.444944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T21:23:33.598026Z digest=sha256:72f9a0ddffafc4226155cfe2607f787d46f7c8987ef8f8d2fbed6cacd0c4e178

Observation 895f2136-3282-4deb-91f7-27e2bff74873 · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.809969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:2fc9deb62b3efe816ccfa825e41735f626651cc86ef8da7cea27dfe0ba1de814

Observation 30240866-491b-453d-8a29-00fb300b939d · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:40:46.361256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:38:49.075457Z digest=sha256:f63b0fecf0c7f7161160199b75349497de3bceefae2bd4b3e5b372617d690abc

Observation 0d5a81a4-f4a7-4ce7-9d4f-f670f02cb94f · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:23.065987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:14:23.065987Z digest=sha256:502cf08ef1fbc947da746dfa146f61cc3713f4bd43c4069eed803d3908663bdd

Observation 2f3689b5-1689-4621-8177-8e9a4589739a · inbound

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video cites this paper.

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:13.095482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:35:23.842691Z digest=sha256:e3589c77ca55a875423672efd2c243aa85e3aa098a339f51477b441d725be071

Observation 110d5eda-42cd-4a0c-a706-ced4bcebc4c6 · inbound

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning cites this paper.

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:13.636648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T08:04:15.238840Z digest=sha256:56af12985c8cb3504fb1c09652c18ded339eae5665737685b6e4344e4b199946

Observation 87b80d37-3a54-4c85-9892-6f7cf8472eb7 · inbound

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models cites this paper.

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:52:16.273746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T04:52:03.076788Z digest=sha256:88ea46100287ca6033722ce0db314c5687e022dc89b56e48b4c129af4180e913

Observation a522d834-08cb-4dae-914a-d4ee04358f2d · inbound

Stage-adaptive Token Selection for Efficient Omni-modal LLMs cites this paper.

Stage-adaptive Token Selection for Efficient Omni-modal LLMs ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:08:04.999290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T06:05:44.739490Z digest=sha256:2c7fb4d171200e866e49e08ed5f7b6494509a1ce6f12ad1656ac47635a9478ba

Observation 0f3610da-8dd6-4be3-aadc-32053dd315be · inbound

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding cites this paper.

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:43:50.865844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T18:36:29.376048Z digest=sha256:2e4239cd1e9cdb9c51d3457fe89d482c567c2d81b14f703a5c49e3e504d6ea0c

Observation f4ef4a4c-eac4-4efb-a123-54a804582313 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.276794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:c00b8eb196c8bc41fe92eff825f378ddd54853acc6a54ee32508fd6ec48478d0

Observation f5d9aaa1-58f1-450d-aa96-2d51e362a60a · inbound

Learning to Deny: Action Denial in Multimodal Large Language Models cites this paper.

Learning to Deny: Action Denial in Multimodal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:39.578572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:21:09.996386Z digest=sha256:99ce09d311fb89bf996f4a5ddfbfa209f5ab326b350dead1a1d393fe7ad09d0d

Observation b1162d31-7bdc-41ae-b96f-3fdc4ef55fc5 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.611522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:d94555a446d349724a00bb9a838c62cd8f64f085f585b13f461e3135113f5cd4

Observation 182b0ee0-427f-43da-aee5-3e1c50ac9f1e · inbound

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning cites this paper.

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T05:48:27.255331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:48:27.255331Z digest=sha256:8e6aca0bdb5e9c9e18d6138d15ed6d610f432d25450657c9951747c527f325a5

Observation e773c642-522a-4375-8026-1b9b52077a05 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.165633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.165633Z digest=sha256:596d04638de761f7f09422362099a4709536c495caa31049b44de8422d9dd0fe

Observation 6d25759c-5f23-49ae-9276-95bc2c367031 · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.828792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.828792Z digest=sha256:1f89f9351e3bdd9fa1d39e14679ff99ca3621dc55dd10a56b1643a42dee4f3f2

Observation 70c59f4e-51a2-44fa-bca1-f081081e2100 · inbound

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs cites this paper.

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:37.839045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:37.839045Z digest=sha256:78cb21226b620617d83dbc7400fbe93ecb39f11a3c817d2d20290525674229f0

Observation df338285-baa2-473b-a284-b677b5a98a8e · inbound

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models cites this paper.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.785575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.785575Z digest=sha256:9c87221fdce7d1713b5a5a402320225ff3db57931692e9b927814d1146d4d851