Pith. sign in

Paper Citation Record · LEDGER

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

As of 21 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 20 inbound Pith citation observations for arXiv:2507.20939.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20939 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:12:41.078055Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:13:44.392418Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:49:30.275005Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09faad0a-b366-49d8-9cb6-2caf2d9a1b54 · outbound

This paper cites Qwen2.5-VL Technical Report.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.360764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.360764Z digest=sha256:658511468d15c4715cd2603ea29ef2eb8937611e8fba07f26e182776e3b016ed

Observation 7548790d-90d4-4c51-8750-4ca490328737 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.953417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.953417Z digest=sha256:599956004815c645514e841e3c9c18667b5850566a12b2b4c6e307581eaf45de

Observation 9b6c2843-4cbc-4135-a943-43880b395e63 · outbound

This paper cites EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.150961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.150961Z digest=sha256:87f7fcc59a413a14a18c4666854160249637f8478d1d80b0f2843a355efdcc06

Observation b4a0a645-4785-4ba4-a8bd-573947db35c1 · outbound

This paper cites Audio-Visual LLM for Video Understanding.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Audio-Visual LLM for Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.249511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.249511Z digest=sha256:1f8dd526db181e83909eb08b3e87bd498fec4f841dc73bdc0720837c16adc3bc

Observation 121a0ef3-6a52-41a4-83b9-bcc9bd4d0e63 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.387502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.387502Z digest=sha256:10e02ea5b59cc330454a121f145eec98fa1de748753fff0bebfc3608881ca90a

Observation f9929eab-f34d-4ddf-99e4-742369172e99 · outbound

This paper cites video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.512824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.512824Z digest=sha256:f29bfc3daec195e400273911f87e5c012a3bb146b45ddc7bce05eec93ec756d6

Observation cd8c78b8-5b4e-411d-a750-cbaf2e8ec76f · outbound

This paper cites video- salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220,.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts video- salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.594915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.594915Z digest=sha256:a89bea407b92a67fb37cc70a58c766cd0ff2e7fe69b6f9e13e7f39cf4a083b2a

Observation 36e2cfcd-788a-4d8c-85b5-7d2114fb4bea · outbound

This paper cites Kwai Keye-VL Technical Report.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Kwai Keye-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.675805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.675805Z digest=sha256:d9c18247fdaafa496d8d7f1242c5df777d647a6f94f07de54e2338945ddb3b85

Observation d90faea1-3e4d-4f84-9b8d-1851eb24d599 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.870798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.870798Z digest=sha256:25934bbee0b16c5c213f8bf0f9862ac1095b22479b0714d4483c8d510750322b

Observation 7e81c92a-5039-4fec-b214-1c1dc739921b · outbound

This paper cites Self-supervised product title rewrite for product listing ads.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Self-supervised product title rewrite for product listing ads

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:12:41.674677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:12:40.971685Z digest=sha256:95bb2562fe0d333939fbb02d886228db75c26df868777f3a03219dfb55c39635

Observation 0330b3b7-885f-4f0f-b3ed-c34643bedbe1 · outbound

This paper cites Pre-trained language model based ranking in baidu search.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Pre-trained language model based ranking in baidu search

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:12:41.484033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T13:12:41.078055Z digest=sha256:115a922441a5a1eafdad5c236d5176a04efbfa70289f58852cc7f272421f3d50

Observation 19e9010d-dacf-4543-9715-976def34db69 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.428823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.428823Z digest=sha256:49f26c0619eefa3dae413d9b6f503542612b08df2c13fce8aba03e71a48f28e8

Observation 83b99e30-bf69-4500-bbe8-019717c45630 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.878853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.878853Z digest=sha256:62fb7ff334ae90e57d6aab9136660ebed035a4750e9f0cc54156309feb057b88

Observation 8adc1884-5841-4d92-9623-872b8dd0785b · outbound

This paper cites CLIP2Video: Mastering Video-Text Retrieval via Image CLIP.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.766732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.766732Z digest=sha256:0410819c39014048d7c1b5ad5aa1514d8b6920fa001440b5d48fd713a30b9e7f

Observation 4bb1c417-c44c-49eb-b3bd-7193a9df94e0 · outbound

This paper cites VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.063831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.063831Z digest=sha256:8b36a5ff2ff25c0720225ae43b35a6272fb98eefb057da7257b412f17e4e4c83

Observation 929118e3-4139-417a-8134-569457a97108 · outbound

This paper cites Qwen2.5-Omni Technical Report.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Qwen2.5-Omni Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.757999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.757999Z digest=sha256:aca818b8721b9967a8419b6551370091bf09b14e95121eb559195f1dc4887044

Observation ef3698a5-928b-43b8-b734-8e05bca71011 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.499439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.499439Z digest=sha256:f0c1f847b649187d9475f613a64c6735e909a000cbc691b9bb5396141f917dbb

Observation 17226eaf-2214-427b-b571-910ed26a187a · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.601562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.601562Z digest=sha256:66b413f82c1c1e7196f3148e3b15fddf7e9c43d7d9cfc943a1824336aba9f834

Observation 7da6f7e7-2ed8-4976-b4ca-a5e4d7ee1be7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.679158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.679158Z digest=sha256:f837f8b4ed1565fb84c177b83c1e21203a107585c7ede7fff289586a9fe705cc

Pith citing papers

Observation d82c0732-c3fe-44a9-bd50-e5bb579ddea6 · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:28.047281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:28.047281Z digest=sha256:5a24306f06aefd5023942536367f1a9571b899c28ea5ff2c4e99eb6305eb6b24

Observation 016ec883-792f-4898-8ee7-42fd8c75a35d · inbound

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models cites this paper.

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:15.022646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T20:45:37.418493Z digest=sha256:38305da43b2e77212b9bfbcc3939ce8d02bd19eed85301f2d795dceed4249875

Observation 431cd667-1fe8-4634-a1e5-be3605422d64 · inbound

AdaTooler-V: Adaptive Tool-Use for Images and Videos cites this paper.

AdaTooler-V: Adaptive Tool-Use for Images and Videos ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:34.444944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T21:23:33.598026Z digest=sha256:8d36056143cb5315aeb523f47a7912e53bfcfdf1752541e7c45e0529475aecb8

Observation 895f2136-3282-4deb-91f7-27e2bff74873 · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.809969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:2a1d95c9ba4815294ea542021b928f42cb319fc55de38f78c42292c253aa9cb4

Observation 30240866-491b-453d-8a29-00fb300b939d · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:40:46.361256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T08:38:49.075457Z digest=sha256:6802d98e01563d789432126cedf9649e51b3b28ab85c8082851b57e2faa0b380

Observation 0d5a81a4-f4a7-4ce7-9d4f-f670f02cb94f · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:23.065987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:14:23.065987Z digest=sha256:85a5d360486219a111e60ca0ce052b495446d1b5345545472baccdad2129e45c

Observation 2f3689b5-1689-4621-8177-8e9a4589739a · inbound

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video cites this paper.

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:13.095482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:35:23.842691Z digest=sha256:ea87832ad8953650099ab4248bd5110a46ffbd521741bafc5ab873a407619504

Observation 110d5eda-42cd-4a0c-a706-ced4bcebc4c6 · inbound

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning cites this paper.

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:13.636648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T08:04:15.238840Z digest=sha256:ba3ba4e27d618e27b7d7aa273369ec9283f6008f8365ca915c45b837b7803ce5

Observation 87b80d37-3a54-4c85-9892-6f7cf8472eb7 · inbound

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models cites this paper.

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:52:16.273746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T04:52:03.076788Z digest=sha256:423936b4208a448331e080c9132e19c56cf23afca96db049dd6d3438a8ca943c

Observation a522d834-08cb-4dae-914a-d4ee04358f2d · inbound

Stage-adaptive Token Selection for Efficient Omni-modal LLMs cites this paper.

Stage-adaptive Token Selection for Efficient Omni-modal LLMs ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:08:04.999290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T06:05:44.739490Z digest=sha256:c5abb420f7c5f2169c4f3626192b0335bb0835d5cd5588ca7122341ee84838cf

Observation 0f3610da-8dd6-4be3-aadc-32053dd315be · inbound

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding cites this paper.

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:43:50.865844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-29T18:36:29.376048Z digest=sha256:8985ec1fffa14b4954419d190952627f64e0c4a2213c2ceda4f3adadea5a95f7

Observation f4ef4a4c-eac4-4efb-a123-54a804582313 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.276794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:b00f7ffba69fcf313458c7ba8b5bfb5a6b9ede99880ba6834e4b8a32758fa5ec

Observation f5d9aaa1-58f1-450d-aa96-2d51e362a60a · inbound

Learning to Deny: Action Denial in Multimodal Large Language Models cites this paper.

Learning to Deny: Action Denial in Multimodal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:39.578572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T06:21:09.996386Z digest=sha256:8775ec10484b8f915cb2176c2eccaa803f76d920d777f0e327e16cb02046ff6d

Observation b1162d31-7bdc-41ae-b96f-3fdc4ef55fc5 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.611522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:1ae20ec855fbc22bef138b7af089034eba7b75fdb1b5286b3d7876e45855919a

Observation 182b0ee0-427f-43da-aee5-3e1c50ac9f1e · inbound

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning cites this paper.

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T05:48:27.255331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:48:27.255331Z digest=sha256:d17962c876a281ff4b6f3adededd7a036cbdad965e2c0ba7d0ad69f484735b8e

Observation e773c642-522a-4375-8026-1b9b52077a05 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.165633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.165633Z digest=sha256:ce4abbec7daba1448d1defa0210f76cdd894030e56eb909a7ce6fe60db30c6f7

Observation 6d25759c-5f23-49ae-9276-95bc2c367031 · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:19.828792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:19.828792Z digest=sha256:01c7b86122ea59671fc44cc2d16a640fb2b7c88736f402dc8d1cb9bb9d109494

Observation 70c59f4e-51a2-44fa-bca1-f081081e2100 · inbound

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs cites this paper.

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:37.839045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:37.839045Z digest=sha256:a275c65d470c51330ab7e52d49a8e1b5ac975d394fb2bfbce0c4616858b1e7d8

Observation df338285-baa2-473b-a284-b677b5a98a8e · inbound

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models cites this paper.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.785575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.785575Z digest=sha256:8b2dfc3c9ffaf353c13166a5fe5011e037eec87c521be4ca1b3e1f8c635a7711

Observation 8aeb4fe0-9bc3-4f5c-9714-ed4492a388df · inbound

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward cites this paper.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.392418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.392418Z digest=sha256:19af5d4d874eccf4ab671b865eb3258a5a26935864ca6847d82f03aad32b8387