Pith. sign in

Paper Citation Record · LEDGER

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2406.15704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.15704 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:59:33.626864Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:41.262186Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d89cc173-cdc5-465c-8831-c5c3fb1aef2c · inbound

Detecting Children with Autism Spectrum Disorder based on Script-Centric Behavior Understanding with Emotional Enhancement cites this paper.

Detecting Children with Autism Spectrum Disorder based on Script-Centric Behavior Understanding with Emotional Enhancement video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T20:43:48.276913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:43:48.276913Z digest=sha256:f8edddd133d7e2e0667c049ed8b59060ba68e3357eca498fa50a59c33e796799

Observation 91e725b0-70a0-457e-ada8-afc99f40a5ba · inbound

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? cites this paper.

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T23:19:11.643010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:19:11.643010Z digest=sha256:bc2e38a60eca7e39c14dab2fdd880802a962536c356c490e3bda49e8cee03109

Observation 9086d914-ea15-4b8d-9bdf-78ee8ccc431c · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.780200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:9d20c6ee1d26ebef6803d3578d0fe8bc106a27bd56b409c115cd7f71923575ad

Observation d47bb43c-1a55-4c6e-918b-a3548116c8cc · inbound

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs cites this paper.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.511386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.511386Z digest=sha256:9c3e9b4bb247ad24b4bdf41dc73645d8a8193b80fae296486bc4bdd1ab22d5d4

Observation 5bb29447-6df8-40f9-ac9b-062ec74a97f6 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.769347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:251a6683af68bfe3c47fbad4c42005d4938b0c8cb9c3d21d479e0c5b92e93177

Observation ad3b723d-8515-4fca-9182-b78bbe47eea3 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.429149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:ae1b14423eb2b935fb147e73aeb265df1ab05c3b87c894308ed53243e269144e

Observation 0584da7b-ea17-4e07-a9b1-c028c05681c0 · inbound

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM cites this paper.

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T21:12:22.824098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:12:22.824098Z digest=sha256:151b6a8cd520327844d4632cbefab061891955054eab7b8344bf5ff79e6243d8

Observation e34c96dd-e627-42eb-a36b-329e70e30d8b · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.884144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:7e168551ffc0bb0a2bd0a00d6346617f58165f68997adb1b9a37446bdcf67a3e

Observation 622d2af9-d6df-4961-af0d-9b68a07042d5 · inbound

Qwen2.5-Omni Technical Report cites this paper.

Qwen2.5-Omni Technical Report video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:54:03.428921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:54:03.225439Z digest=sha256:7c0831411184a19054aa3a359904e96edf95e5bc23c4e575d2dd50690f3ceb8a

Observation ec3dd096-487a-4ce7-8741-b02ab8275c7c · inbound

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models cites this paper.

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:33.626864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:33.626864Z digest=sha256:cf0d068127e93d66115a0e69787d0f0898ceda9f7d1f649c7c466363c66a462e

Observation abae598a-bd1a-4bf1-b652-9477b1db0b36 · inbound

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion cites this paper.

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:49.812070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:49.812070Z digest=sha256:3a9deafddc64b285a7180b5041fecf3067dd92291e3968546b37dfe45e807362

Observation 3a2e00d1-6dff-4ecc-9867-4cf0fc9abd08 · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.563770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.563770Z digest=sha256:be834462025371961800dc82dbf8bd535973695a7289bfc5f12ed0760c9e8ec4

Observation 75b41b82-8f33-443b-a107-50688a16e005 · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.321734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.321734Z digest=sha256:88a4bbdfb05d7e0f975e74d99633de6cb9b4fa3fe7128a0b9c354d6587a2a9ed

Observation 121a0ef3-6a52-41a4-83b9-bcc9bd4d0e63 · inbound

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts cites this paper.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.387502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.387502Z digest=sha256:10e02ea5b59cc330454a121f145eec98fa1de748753fff0bebfc3608881ca90a

Observation ca268976-1cb6-4ddb-9811-eec61b099fe3 · inbound

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models cites this paper.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:45:56.083406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T05:45:07.700571Z digest=sha256:76adbff17621bb7db2e93dfda25ee300b14898b017bdc5bdb5b524a52031d487

Observation 9e80efd4-13a0-4024-a623-69ec58345c4a · inbound

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models cites this paper.

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:15.076824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T20:45:37.418493Z digest=sha256:8eaf904d29e68293ebd18e56ba69c486f10ffce0f908b6c63b39bb76dd2e4756

Observation ec66fa16-1fb0-4246-a5c9-9226ad4ccf26 · inbound

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models cites this paper.

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:13:52.158743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T02:12:55.170296Z digest=sha256:c10b50d4c30602212c643a4aa8feee34f38ca5d4314cd05a34db5eefe2321004

Observation 3e6f9156-1428-40c7-a443-58f5f2f11e38 · inbound

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation cites this paper.

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T14:31:59.486559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:31:59.486559Z digest=sha256:1e3bd7306fada2181feee8451f5dea302a6db695f528b3add4343ca69101afd5

Observation 0eab07e7-a0fc-4847-b52c-70ee7a568482 · inbound

Do Audio-Visual Large Language Models Really See and Hear? cites this paper.

Do Audio-Visual Large Language Models Really See and Hear? video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.814841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T20:56:19.815569Z digest=sha256:b966010ead16b5946e6a65afebc1ff0985075c19c77fd58ef36ea0c1d8aa9909

Observation 117f93a6-94a4-43bb-840d-11d718dadabc · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:03.707399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:fff33e55a642199af85fa31c0930d3ad886602bcd33866b5a221ba260ae73701

Observation fbf1a986-5b57-4442-a0e3-722dde0c147d · inbound

EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness cites this paper.

EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:42:00.943353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T19:24:23.119301Z digest=sha256:e58f3898d169c4e76e4190ec482b3628e2232659e97033ec87d06b9d85637186

Observation ba12e385-0a6d-4efe-8d44-e926dfc201b7 · inbound

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models cites this paper.

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:52:16.308057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T04:52:03.076788Z digest=sha256:ee646f37068422836c6be67c909436a270bead0ff921c57daf717995352f3fd1

Observation 32598e10-9203-4fcb-9211-f2c0e35caee2 · inbound

V-LynX: Token Interface Alignment for Video+X LLMs cites this paper.

V-LynX: Token Interface Alignment for Video+X LLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.180082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T19:00:18.702121Z digest=sha256:d576a973beeb18f7cea02b11cbf564834f43e4ec7c0bd9e25d45f8847c852910

Observation cbfa1188-5433-46fc-ad43-3954fedc130e · inbound

Sandboxed Coding Agents are Competitive Omni-modal Task Solvers cites this paper.

Sandboxed Coding Agents are Competitive Omni-modal Task Solvers video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.203166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T18:59:51.362554Z digest=sha256:729c00c5185ed11fb5d4f7a618f654b2e32e8b3fff0c84c1c220a19d0be7f956

Observation cc3dad78-4e04-49cc-b6bf-f569a1205fa5 · inbound

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales cites this paper.

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:41.263959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T12:09:01.026544Z digest=sha256:e0bbc1533e0e3fb39a955d7d230a32fd91d5c2a28f73b6509baf1d3a913f0631

Observation 4eb56518-0312-4c00-aa23-69c13833f974 · inbound

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships cites this paper.

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:23.278012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:23.278012Z digest=sha256:0f1e1fdb89a01387a0112ad158da062b16b61b53d6ed52bb688988d880eed2a8

Observation 81ebca45-58da-47b9-9c4a-b1bb9cbe8de0 · inbound

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos cites this paper.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:41.603969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:41.603969Z digest=sha256:f4ed1bd2c45cb70871e1c0712b434a7c7098149dd50d6bf21c13fe2acf0442ab

Observation 9ef2309b-b41d-4c52-8a47-59b364c4fd80 · inbound

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models cites this paper.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T03:25:31.574785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:25:31.574785Z digest=sha256:175257d348f113077f994adead58f5f7e5a2c02f8f61f814cb537fae276c75df

Observation 9942ec5b-7557-407f-b7d9-2993050d133e · inbound

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models cites this paper.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.482439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.482439Z digest=sha256:f59ff7bc7a9cd2cdf09c00910c6c2770ab5499b305b908efa9a98912a4a9ba53