Pith. sign in

Paper Citation Record · LEDGER

Audio-Visual LLM for Video Understanding

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2312.06720.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.06720 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:36:29.216841Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:36.866746Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9c3355a9-bca8-4cc7-8895-0fa71688eca1 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs Audio-Visual LLM for Video Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.634793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:ff4847df20fd5dae46240cd6020cd1ea732c9d4e268870a8d174b97d040fe93a

Observation f56fd9f8-008a-434e-98bf-c4f74ea804ec · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Audio-Visual LLM for Video Understanding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.785427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:27d8ca3aa0eac80f1a1d10b8da666cd3d08d6c48ae12fff23b061f806f8c02df

Observation d6b5824c-d880-4145-867f-cfec6ee61f70 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Audio-Visual LLM for Video Understanding

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.157256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:0a69eb9044fd56e1a28509faaf829c2702a7f0d9af1611bb863a14f4ebbc7f45

Observation bf0cfcfd-0c56-4fce-a101-93c74b334982 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Audio-Visual LLM for Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.428279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:0f08141410eed54fdc321bd903f652af5ad6c18df44a773dad70c603720d91b1

Observation bc1ed0eb-b562-42ae-a323-31a5a96752f1 · inbound

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning cites this paper.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Audio-Visual LLM for Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.216841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.216841Z digest=sha256:49996e42f3aa51a9fe9a13efb298ae3576984b87eff283aba41feac3fa09ff70

Observation b4a0a645-4785-4ba4-a8bd-573947db35c1 · inbound

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts cites this paper.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Audio-Visual LLM for Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.249511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.249511Z digest=sha256:fa29815f9e3753fc27f64b957446aa62815fe0619a031bf88308f258742265d6

Observation 0a5b208f-05a0-45e2-9930-fc1c196af9fd · inbound

AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering cites this paper.

AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering Audio-Visual LLM for Video Understanding

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.403906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T05:26:07.722390Z digest=sha256:3c2ce76783c931a24d0e5b6bb5f3a79c6dc85cb007c20328e55cc85302069b5a

Observation 52e6c9c0-0ce5-47fc-96f6-0e6a2b0a3fb1 · inbound

AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering cites this paper.

AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering Audio-Visual LLM for Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:26.153084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:55:26.153084Z digest=sha256:a48e405f50a591ee46f25fc58537ecb0430655ca2f3452e72916092d61c2f459

Observation e99330e2-0e58-43c7-a6f6-f913315cf8e1 · inbound

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning cites this paper.

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning Audio-Visual LLM for Video Understanding

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.868245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T13:53:53.520545Z digest=sha256:eedb21ca07ed39a6c627d31acdd68984eb97ad74d334de1356eddb9748868aff