Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2412.02611.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:56.875342Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T03:49:30.296211Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c0de47b9-1277-447a-b849-99fe422276a8 · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 141c6743-16af-4f7a-bccc-bcbb9b7fe1d3 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 260
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation added75e-72c0-4bd8-8f88-2349c968ab3d · inbound
Qwen2.5-Omni Technical Report AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 75bd58d3-d36d-490a-9111-a79eda5b8246 · inbound
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f714a50a-37d6-47ff-820f-dc548c484757 · inbound
Learning Sparsity for Effective and Efficient Music Performance Question Answering AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 484f6ce8-fdd9-4447-b52b-7e1f6f78b46d · inbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20fc6335-e5a8-4b6b-95a5-7aac665adb73 · inbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 122
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfc23237-ecb4-4048-85aa-77b5b166f59c · inbound
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e9fd2fa-bc3b-44c4-aef5-f66678c16da2 · inbound
Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9bff972c-7da0-4f12-9ad8-98248f38af4e · inbound
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a5b971d3-a744-429a-acc7-eb4a60f1a7a9 · inbound
Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6c7a6a4b-b71b-4f7c-95a7-53a6023051c7 · inbound
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6add56f3-eb07-4c52-8e7e-da8f291eee42 · inbound
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3362eafb-b804-4ffc-9e0c-5044270b27b0 · inbound
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b30d95c-6c4d-472c-819b-5bca454f3202 · inbound
SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 596a8cd0-8b46-4d82-9f49-2728b5ab5c91 · inbound
From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 894cfb0a-48d4-4c8e-82ad-491bcf5ef139 · inbound
CogniRoute: Learning to Route Social Evidence in Omni-Modal Models AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 194
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 966b6e6d-3c38-41f6-add6-84fb95b73292 · inbound
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.