Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2406.15704.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:06:49.812070Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:09:41.262186Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 9086d914-ea15-4b8d-9bdf-78ee8ccc431c · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5bb29447-6df8-40f9-ac9b-062ec74a97f6 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad3b723d-8515-4fca-9182-b78bbe47eea3 · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e34c96dd-e627-42eb-a36b-329e70e30d8b · inbound
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 622d2af9-d6df-4961-af0d-9b68a07042d5 · inbound
Qwen2.5-Omni Technical Report video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation abae598a-bd1a-4bf1-b652-9477b1db0b36 · inbound
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a2e00d1-6dff-4ecc-9867-4cf0fc9abd08 · inbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75b41b82-8f33-443b-a107-50688a16e005 · inbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 121a0ef3-6a52-41a4-83b9-bcc9bd4d0e63 · inbound
ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca268976-1cb6-4ddb-9811-eec61b099fe3 · inbound
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e80efd4-13a0-4024-a623-69ec58345c4a · inbound
OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ec66fa16-1fb0-4246-a5c9-9226ad4ccf26 · inbound
See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e6f9156-1428-40c7-a443-58f5f2f11e38 · inbound
Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eab07e7-a0fc-4847-b52c-70ee7a568482 · inbound
Do Audio-Visual Large Language Models Really See and Hear? video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 117f93a6-94a4-43bb-840d-11d718dadabc · inbound
Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fbf1a986-5b57-4442-a0e3-722dde0c147d · inbound
EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ba12e385-0a6d-4efe-8d44-e926dfc201b7 · inbound
OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32598e10-9203-4fcb-9211-f2c0e35caee2 · inbound
V-LynX: Token Interface Alignment for Video+X LLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cbfa1188-5433-46fc-ad43-3954fedc130e · inbound
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cc3dad78-4e04-49cc-b6bf-f569a1205fa5 · inbound
CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4eb56518-0312-4c00-aa23-69c13833f974 · inbound
ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81ebca45-58da-47b9-9c4a-b1bb9cbe8de0 · inbound
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ef2309b-b41d-4c52-8a47-59b364c4fd80 · inbound
OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9942ec5b-7557-407f-b7d9-2993050d133e · inbound
OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.