Pith. sign in

Paper Citation Record · LEDGER

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2412.02611.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02611 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:56.875342Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:49:30.296211Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c0de47b9-1277-447a-b849-99fe422276a8 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.417668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:2ce18429fdaa2bf00230cfb6bfadf4daeba5755c74f8a47b504e4bfc00d0b7aa

Observation 141c6743-16af-4f7a-bccc-bcbb9b7fe1d3 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 260

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.590020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:14f352438ab081d59827c685afc396af2f5e0474308902d128138cd3c138cc05

Observation added75e-72c0-4bd8-8f88-2349c968ab3d · inbound

Qwen2.5-Omni Technical Report cites this paper.

Qwen2.5-Omni Technical Report AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:54:03.340882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:54:03.225439Z digest=sha256:f539fa3b26a574766ff0631f55110432c889f07c6518ae2d71a4da5a6e878d80

Observation 75bd58d3-d36d-490a-9111-a79eda5b8246 · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:56.875342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:56.875342Z digest=sha256:6e6b0a044a41f120614145f852ccb3c59bd498667bd379838ff20546153cd1ff

Observation f714a50a-37d6-47ff-820f-dc548c484757 · inbound

Learning Sparsity for Effective and Efficient Music Performance Question Answering cites this paper.

Learning Sparsity for Effective and Efficient Music Performance Question Answering AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:41.293828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:50:41.293828Z digest=sha256:4c3a3dc2e74d2c42825b8291acd9debe7d0821fba0de5e96c95b6bcbdde42ca3

Observation 484f6ce8-fdd9-4447-b52b-7e1f6f78b46d · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.674331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.674331Z digest=sha256:29dad97defc0e31357b13510f0fcbd10d070ae8866c04c73c87bf86a5dc62293

Observation 20fc6335-e5a8-4b6b-95a5-7aac665adb73 · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.597241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.597241Z digest=sha256:57dcf4c91bc20e95a987b9df92420a3fa98352c9ee0894b60b37450cdceff719

Observation bfc23237-ecb4-4048-85aa-77b5b166f59c · inbound

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models cites this paper.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:45:56.100282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T05:45:07.700571Z digest=sha256:e78529b64c05afaafa56f4d90a2b4066ead1b608e99fc5671f5a9e7061e313ea

Observation 7e9fd2fa-bc3b-44c4-aef5-f66678c16da2 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.046832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:82a9490af541c459cb2bddbb61b25b990d2f1fc2199f39484742787efa853390

Observation 9bff972c-7da0-4f12-9ad8-98248f38af4e · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.354000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:9a97e290374fb286027c165b4620d0e3f1cd9caef16d5fae638630a3eeb7c80a

Observation a5b971d3-a744-429a-acc7-eb4a60f1a7a9 · inbound

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search cites this paper.

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:27.463941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:21:52.972107Z digest=sha256:09cf0fc0c9a71d8b85204a8e045ca766bf53012a3a149dc3551da03cb3974dc1

Observation 6c7a6a4b-b71b-4f7c-95a7-53a6023051c7 · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.681606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T03:49:58.240883Z digest=sha256:e7bd17fd8c19c5a46d3a19a3006bd056b2f6e1f78591c34c2cd455ad4463bc8e

Observation 6add56f3-eb07-4c52-8e7e-da8f291eee42 · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:09:50.271349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T06:06:20.658030Z digest=sha256:23ad09cfa1d1dcb166bc1fb3a9cb7229eeb05e7ee667252b4346d9f77a30d0c9

Observation 3362eafb-b804-4ffc-9e0c-5044270b27b0 · inbound

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding cites this paper.

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:24.991923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T05:51:49.390597Z digest=sha256:cf27ab0db297e686c4c4f853f30593d0e53e7d02d73032a84683a88c88499cc8

Observation 9b30d95c-6c4d-472c-819b-5bca454f3202 · inbound

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models cites this paper.

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:46:15.354338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T16:24:43.104448Z digest=sha256:84ce7d21349234f544d5ef210bb6f6307d3847b8075117710fd64c067ef03642

Observation 596a8cd0-8b46-4d82-9f49-2728b5ab5c91 · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.319448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:663e6b530f55a4d25ac2389bb4d055035975350cd991153cb8c67a7bb8ebbe8b

Observation 894cfb0a-48d4-4c8e-82ad-491bcf5ef139 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:30.298437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:3111050ed64c34d4843e30f4531ab338ed83b501c029ed377e728940c293b653

Observation 966b6e6d-3c38-41f6-add6-84fb95b73292 · inbound

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs cites this paper.

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:34:19.509749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:25:38.593423Z digest=sha256:cee4dc9c586b807348d176edf5e457a111a31c8a03a78e64aaeb7e13474b7839