Pith. sign in

Paper Citation Record · LEDGER

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2406.15704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.15704 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:06:49.812070Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:41.262186Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9086d914-ea15-4b8d-9bdf-78ee8ccc431c · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.780200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:88e3ba8f287fac058d827806e9b2d2a32c4783d34c8482690b20495d80d77edd

Observation 5bb29447-6df8-40f9-ac9b-062ec74a97f6 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.769347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:61b08bf3b5560f513a409d307a13c8124c10035c69092767b5ab3457b767cf5d

Observation ad3b723d-8515-4fca-9182-b78bbe47eea3 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.429149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:5afb4c0a1b68d8a59ce9fedf015c24fbcaf299d99f35a5c7e64f3a3854981044

Observation e34c96dd-e627-42eb-a36b-329e70e30d8b · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.884144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:5a9ffdc238319b0d7f1d2c496e668889cf73c2d3a6711c33375d76b001de3475

Observation 622d2af9-d6df-4961-af0d-9b68a07042d5 · inbound

Qwen2.5-Omni Technical Report cites this paper.

Qwen2.5-Omni Technical Report video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:54:03.428921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:54:03.225439Z digest=sha256:d8d4f2022e6cd37afb6abd8f595d8f6c1b5fd7264c130061367e4ff5143e3de7

Observation abae598a-bd1a-4bf1-b652-9477b1db0b36 · inbound

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion cites this paper.

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:49.812070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:49.812070Z digest=sha256:8ca26465ccddbd87afd44e04f76916cfaea7058301a0b09e97c9d8f0aeda6ef5

Observation 3a2e00d1-6dff-4ecc-9867-4cf0fc9abd08 · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.563770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.563770Z digest=sha256:d2339cd9d415465299fdce6dc1e8ee2fc850b8af42b712518f3e4484de59748e

Observation 75b41b82-8f33-443b-a107-50688a16e005 · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.321734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.321734Z digest=sha256:ff2187359e00eed5900686c45cf9e4679cfbbc24345d7dd98074fd2ef9de1450

Observation 121a0ef3-6a52-41a4-83b9-bcc9bd4d0e63 · inbound

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts cites this paper.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.387502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.387502Z digest=sha256:653442875c64827831c89b997de52cae7add2fb03d11d4fbcf0ba9c5a6fcaf3e

Observation ca268976-1cb6-4ddb-9811-eec61b099fe3 · inbound

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models cites this paper.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:45:56.083406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T05:45:07.700571Z digest=sha256:6fcc731ad03ef24a8d02b780b27294cc66f19509b506eb2ec117cc3e636bbe18

Observation 9e80efd4-13a0-4024-a623-69ec58345c4a · inbound

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models cites this paper.

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:15.076824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:45:37.418493Z digest=sha256:991e869ab1a40673e885209735210fa9228b328eed8a9aa9a9f2e71408f55403

Observation ec66fa16-1fb0-4246-a5c9-9226ad4ccf26 · inbound

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models cites this paper.

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:13:52.158743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T02:12:55.170296Z digest=sha256:49e979a0e53758d5dfbd8049258054ac2434d0a5574c3af552ba63a722a02f2d

Observation 3e6f9156-1428-40c7-a443-58f5f2f11e38 · inbound

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation cites this paper.

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T14:31:59.486559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:31:59.486559Z digest=sha256:72dd64c39eb3c6fc074426b6a481d134d5b2291ccbd17a5c72da762ffd9e3a6d

Observation 0eab07e7-a0fc-4847-b52c-70ee7a568482 · inbound

Do Audio-Visual Large Language Models Really See and Hear? cites this paper.

Do Audio-Visual Large Language Models Really See and Hear? video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.814841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:56:19.815569Z digest=sha256:a0b93f5705ab4588796efc665a2eb40d2b1deb357797ac7bda0035f9e2173a98

Observation 117f93a6-94a4-43bb-840d-11d718dadabc · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:03.707399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:2924561e757da05a663a3827c349598dca16790981f5a5eaff48dbfb1cc2a501

Observation fbf1a986-5b57-4442-a0e3-722dde0c147d · inbound

EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness cites this paper.

EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:42:00.943353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T19:24:23.119301Z digest=sha256:b8078a3fb99ae8ef84a1cfd31b3a44743fe70c4363b4b77f2fbea63e4c73b380

Observation ba12e385-0a6d-4efe-8d44-e926dfc201b7 · inbound

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models cites this paper.

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:52:16.308057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T04:52:03.076788Z digest=sha256:1a5442f6e58d2739d22bac6bfd94b81593fc70ba8ae29958320d75680f3c5d43

Observation 32598e10-9203-4fcb-9211-f2c0e35caee2 · inbound

V-LynX: Token Interface Alignment for Video+X LLMs cites this paper.

V-LynX: Token Interface Alignment for Video+X LLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.180082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T19:00:18.702121Z digest=sha256:35d2ee052d623476e39ecba7a335320f8799c23ab0edade4633cbd5761a32371

Observation cbfa1188-5433-46fc-ad43-3954fedc130e · inbound

Sandboxed Coding Agents are Competitive Omni-modal Task Solvers cites this paper.

Sandboxed Coding Agents are Competitive Omni-modal Task Solvers video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.203166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T18:59:51.362554Z digest=sha256:d58c2fbe0af28270e624a571a8d4a888e6de18d00e377ea025a99e5a482b6443

Observation cc3dad78-4e04-49cc-b6bf-f569a1205fa5 · inbound

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales cites this paper.

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:41.263959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:09:01.026544Z digest=sha256:3d231459b846973df578441d92b8a20e0e46e617d6f0559883651e1e77bdce81

Observation 4eb56518-0312-4c00-aa23-69c13833f974 · inbound

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships cites this paper.

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:23.278012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:23.278012Z digest=sha256:4c637e1c620da7f4ecb46472f8cb07923e78321a2ff52f95a930af3a2d70cd1a

Observation 81ebca45-58da-47b9-9c4a-b1bb9cbe8de0 · inbound

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos cites this paper.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.603969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.603969Z digest=sha256:77d9bd1c13ace74c3dd9bc18efe72c15a2e4e11cd55531ee82af5c07b620ca14

Observation 9ef2309b-b41d-4c52-8a47-59b364c4fd80 · inbound

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models cites this paper.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T03:25:31.574785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:25:31.574785Z digest=sha256:f57b20d438506b9ae89df2827c87160e9f7409e16ad27619aad0d22530567a90

Observation 9942ec5b-7557-407f-b7d9-2993050d133e · inbound

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models cites this paper.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.482439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.482439Z digest=sha256:804bf36a648bbda1b466703754b8cc76651f6011878d5a14f2786c13cab96022