Pith. sign in

Paper Citation Record · LEDGER

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

As of 9 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 7 inbound Pith citation observations for arXiv:2506.05328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05328 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:27:05.708277Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T22:09:56.270516Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T05:45:56.102421Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact3
  • verified fuzzy15
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a684ffce-ddae-4dfa-b02a-13da027d9496 · outbound

This paper cites Qwen2.5-VL Technical Report.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.464312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.464312Z digest=sha256:d7c163889a2998b1d8003721e2b7e0f1361feb10d6be9fd2e26b49c09740dd30

Observation 6be6bd9a-0430-44a7-9f86-3859b73daa87 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.468595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.468595Z digest=sha256:0be7c8944178364c715c952db782b7026130bec54c9cc5f116377e6670125856

Observation 22c657ba-6fc6-4910-9e5a-f65fe28c6634 · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.471940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.471940Z digest=sha256:2208840cc1457ac5c132e9189045eba0706133a833a74788477333826c3b01a9

Observation e2307a22-5132-418a-bd8d-d6bf2f3373d7 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.475164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.475164Z digest=sha256:c567a885e8832b3113120d6acae1be242be82a7dcc54bd2fa7a2f7ec7a8e9525

Observation 77bb1d4b-203d-486f-8adf-71fbb007e644 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.565887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.478496Z digest=sha256:847777911e119401664a6048b1d90b903534fbc6d4b22b426c726fb96ee6a463

Observation aec4a5df-6f37-466a-955a-a43cfada4361 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.482122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.482122Z digest=sha256:61ab98194ae104beef95b07a91609dd0f434ae893eb7730e521be26926dc1cc5

Observation e816cc37-4f83-46ff-b32a-26b4eaafe715 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.486127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.486127Z digest=sha256:1119220c0da29efc0c285da5e29f4e54a8fa8163cdf764e21dde8ae459028712

Observation dbda99d1-540f-4f58-b15b-70c985135e06 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.489310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.489310Z digest=sha256:dcbbfa3388539c8bdaf5d7f0a67e768cef52b3cebc0f847764063c0ee9bfc39e

Observation fbd5d1e7-aff8-4073-980f-3d9d450d7a83 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.492480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.492480Z digest=sha256:4294ffb49c632c543b3f74490de7f4ecbaa368079cbf7105a12dcb068f3fe7da

Observation 0ba4e3ee-4df0-48c8-b7f6-ca213a40175a · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.495697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.495697Z digest=sha256:da78d4a5be18228ab929466aa9a6d38755041e55164966e95edee3bb27763943

Observation 5c096c6d-ac30-465c-a225-b79619168202 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.499105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.499105Z digest=sha256:dae2f5853b0f5f31eb84fb78f4a18d7476b184ecd9f30533131986e8772b8231

Observation 09444582-a7b0-46d5-9633-3341ba5b0202 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Cogvlm: Visual expert for pretrained language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.502419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.502419Z digest=sha256:a50d38740f01e94a3778df643bb0c90134d99eb77c0295447f99fda19fc71cde

Observation 46c9b6b1-1bdf-47dc-9e13-ea4ac9dc1f34 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CogVLM2: Visual Language Models for Image and Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.505880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.505880Z digest=sha256:6d48931026cc27aff0e706795820e811d4b7e4604c5b752f634f25bb3af31c91

Observation 5e2d9d77-a5ee-42a6-8377-2f5e94b38832 · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoLLM: Modeling Video Sequence with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.509262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.509262Z digest=sha256:df0a8fdc3578e03c2d2bb91dd2c146e9db30cb0734ed13900e08a7cd5cfb9096

Observation 5e625726-1323-43d1-b2b6-16c47c50e13f · outbound

This paper cites OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.512719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.512719Z digest=sha256:afd6daa72f8d5f78682a720f7cf5851217ac1cd9f8a190b957ea64d7a43b5e7c

Observation 300bb223-9ce7-44fe-9ffb-674a621388d9 · outbound

This paper cites Dvd-counting, 2025.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Dvd-counting, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.548934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.516090Z digest=sha256:57f197fe1951a25dd7c0bdaf7bf8990f8b18ec6d95eb4b9d172d56945da194a2

Observation 3a54d292-f766-44d4-a434-7cda7ddb0235 · outbound

This paper cites Counting out time: Class agnostic video repetition counting in the wild.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Counting out time: Class agnostic video repetition counting in the wild

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.519281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.519281Z digest=sha256:725db439387ef339ce8f3d9b76a0cd9f4a5e631657cea39f9cc19dd229e700b6

Observation 3fdb8348-6f9e-4a57-b55f-82e31989d5c2 · outbound

This paper cites Repetitive activity counting by sight and sound.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Repetitive activity counting by sight and sound

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.532936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.522916Z digest=sha256:03ef4262cc4ab5314da62462a0bc090ca18463344eaa0e724c010bf53ed33493

Observation 771c62c1-f2e7-4a53-aedc-b111cfc00a7f · outbound

This paper cites TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.526108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.526108Z digest=sha256:17142752519fc39abf004c683d8cf7e892869c28c29c02bee345f196737dc5d8

Observation 02223deb-a7f6-4379-8ff1-a37f4262b12d · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.529740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.529740Z digest=sha256:9265cb50fd5f4808069cba587b242e0d2904e620c5ea4c113017bd2b8b60b794

Observation be396407-bbd2-4d64-bb62-6ae72aab8163 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.533023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.533023Z digest=sha256:a347aaf9b2b90f473635dbc7679abf1b200f4acab3ac68db4aaf2701b0db1a8a

Observation 806f6ff7-5b7b-4842-9630-86556c9f42b7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.536387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.536387Z digest=sha256:71d83a2da138c000ed16d2877fa34ba4cf562fed927d9840be5e616ddef8dcbb

Observation d411759c-8069-4827-b3b0-eabe637af90a · outbound

This paper cites Curriculum learning for reinforcement learning domains: A framework and survey.Journal of Machine Learning Research, 21(181):1–50, 2020.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Curriculum learning for reinforcement learning domains: A framework and survey.Journal of Machine Learning Research, 21(181):1–50, 2020

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.539317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.539317Z digest=sha256:0f85fcd63f8a526ed335e9fc3005c8da0a241fbe8a2f08d4754ce3e273b9ecf3

Observation c7b9bcc2-ec63-4c1a-b970-4323e5492402 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.542467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.542467Z digest=sha256:5554885da66b7a3fcb41d6164c8a4e37fec6868edadf63948ab8cea7f20b4755

Observation 9e0e43b2-8ea1-4d4b-9536-cd3afc7f67f2 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.545846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.545846Z digest=sha256:c9b78b0016eafc0b7f7d6f3c8ee2b7d68c3bcaeab723922dcc250114d7e00e34

Observation 686c5244-6f31-443a-a4b1-34f22aab81aa · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.549332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.549332Z digest=sha256:02268870fbe0562ce2c57a3fb766d94dacf20c73fd4ce853b12654ec0b66c2cc

Observation addbeab7-af44-472e-9131-91bcf2215c69 · outbound

This paper cites OMCAT: Omni Context Aware Transformer.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs OMCAT: Omni Context Aware Transformer

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:06.046811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.553172Z digest=sha256:a1e820942b8b3e44917ef8cf537db7dbbaf7c371983bf097904eb8db923f6ccc

Observation 9ca97b63-d6e3-4293-87f8-684f8ed388ba · outbound

This paper cites Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.557034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.557034Z digest=sha256:16068bfe08f2ead03db14b63db4bb27775b90af22db2cc92397c63e7b31cdfcf

Observation 18af2022-4246-4367-be8f-081bb7661fe2 · outbound

This paper cites Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.560273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.560273Z digest=sha256:bf078cecc25e2765e8cadcf575747a639fd6ba033eada8d8e09342780636f229

Observation 3a2e00d1-6dff-4ecc-9867-4cf0fc9abd08 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.563770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.563770Z digest=sha256:f1f2729802903236e2a601f300435ee6d866215da0cbf35f7d660c8a85fbd7ab

Observation 5ecc1de9-e393-4d1f-bf8e-82aeb980cc0d · outbound

This paper cites Meerkat: Audio-visual large language model for grounding in space and time.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Meerkat: Audio-visual large language model for grounding in space and time

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.567346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.567346Z digest=sha256:155a4189d71c2eb1ef2076fbe2a950525f19e3d8c5a477fb7defd3287a270828

Observation e178f6d4-3cc7-4795-9423-6e16a9a8f312 · outbound

This paper cites PAVE: Patching and Adapting Video Large Language Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs PAVE: Patching and Adapting Video Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:05.978456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.570467Z digest=sha256:0e56ec7416b18d0beeab4876738e2523e6b3c0ec511fdbe96129ef0a2d374a60

Observation bf31b9ad-6fc1-401a-be87-937002b6cf52 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.574186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.574186Z digest=sha256:f6b87135f36d19576c6310226603913603d40b62ea21a56d04ecfd8ea0e97460

Observation c6254dd5-267a-4a5b-a5c4-fc16f2043bbb · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.503625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.578385Z digest=sha256:f582ce3353b63915e9bd77191381c02f978fb9d90cd9b27f4b6afbaf36b11861

Observation 84bd407e-a595-4e5a-a0a5-5fa754d236f5 · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.581493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.581493Z digest=sha256:5ae771de394fce677b9771356ec1f74f021f81ba5401fac7ff58488ece9a222e

Observation 0b61ec07-d356-4e35-894e-09aab9193737 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.584562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.584562Z digest=sha256:72fe9eee6871cfade6735a88bea1713f2aac0970ed36ae7ec01436c96f3e63b9

Observation 34bc1aaf-6807-4143-893e-f983d1033b79 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.587891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.587891Z digest=sha256:4ea76cece8376430e804f44f855c1cf2b0125bb13bf2285818a0cde9c07fcc04

Observation 7d06622a-75a5-4ce8-a4ce-e48d40e83409 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.591545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.591545Z digest=sha256:53dc5ae9e0b57aa52e3ef017c1ce051986f5a399d5257c501818356390229339

Observation 79e302ca-b92a-4cb5-9391-af5c5397bfa4 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.594716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.594716Z digest=sha256:5a2a4130673f47f97097f25ff598d3cb51bc471dc2a22ce34463de28a268513d

Observation 8debd7c1-08a0-4b71-8179-253c635383c6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.598090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.598090Z digest=sha256:4773de9fc7345f676af69c427c418b1128d8586ac97c7a66ea85ea22f317cbca

Observation 6435e326-212d-41e0-b84e-7df8a253e2bb · outbound

This paper cites Introducing gpt-4.1 in the api, 2025.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Introducing gpt-4.1 in the api, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.487558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.601374Z digest=sha256:f720e39c630718503b74323485c04ded32a5653213be10ed6e8088d0625e6c14

Observation 6d514cce-0015-4127-85dc-032e853a8a3c · outbound

This paper cites GPT-4o System Card.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs GPT-4o System Card

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.604689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.604689Z digest=sha256:c121a3517350565181d776c8b55a5112222e7c55135c4c68ea633185b88b2699

Observation 2d5d455d-553e-4098-aa0f-ffc0607bc6b7 · outbound

This paper cites Seed1.5-VL Technical Report.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Seed1.5-VL Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.608221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.608221Z digest=sha256:3f25e4bff47ec0e913b094627bb55d735d497a36029f52b4988dc26b75c895bc

Observation 8ee657c0-142e-46e5-abb7-0f73ee36fc1f · outbound

This paper cites Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.611836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.611836Z digest=sha256:dd5d06517fc910b87afd80ddcf72489b5e5a020aeb6b2c9da4fcebd4f19419d6

Observation 8cf9e3ab-02f8-44a7-a346-1c46941f6bbd · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.618482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.618482Z digest=sha256:ab3dee21f25732a7d0a630ceed1ce6d73a02661e076fc31c62c9fac120968fad

Observation 1746ded8-6795-49d0-871f-472260bf428f · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.621901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.621901Z digest=sha256:ce7cc3194a7584516ccfa2f16ce3092bcd2e9a67964c1e42edd69c308e908e32

Observation dc908336-520b-4e60-8078-7f6460e1aaf8 · outbound

This paper cites Qwen2.5-Omni Technical Report.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Qwen2.5-Omni Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.625899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.625899Z digest=sha256:ba0e2655e28d8e5d1eb982ff0d73dd135affc0043b80b120f42ac846781332be

Observation 27005343-43db-4f9d-864c-8a8fd57809ca · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Avqa: A dataset for audio-visual question answering on videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.629557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.629557Z digest=sha256:ace3a1a172d9d1142d0a5ac68e273e987b9be014eb6e4b8cea69e3a7b6303d5b

Observation 5a2c7e4d-0bfa-46c4-be84-db3f5867773a · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Learning to answer questions in dynamic audio-visual scenarios

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.633066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.633066Z digest=sha256:f10a52a02769f367aeeefea88444475852934a1912d237a22a5ce382ec3512c2

Observation 872115b7-75bf-42ad-94ca-e787115101ca · outbound

This paper cites Audio-visual event localization in unconstrained videos.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Audio-visual event localization in unconstrained videos

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.463851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.636748Z digest=sha256:aea12f7d25a0b9102e24aa333516c76cc13111df48eaf728c76fea2d47b97876

Observation 2fcdf031-2400-4683-a35c-832a7b127ab2 · outbound

This paper cites Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.453872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.639954Z digest=sha256:b2f3305cf355a0f1a25f7161016b7de3f457bf581b48b9238ca97572dd2ea81d

Observation 3ad08ab3-41c1-486e-8ab4-6402bbae9db1 · outbound

This paper cites Cross-Modal learning for Audio-Visual Video Parsing.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Cross-Modal learning for Audio-Visual Video Parsing

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:27:05.828372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.643589Z digest=sha256:69b7d631e8c3f45cc25d56ec10c41a452e34d65d08ee75475c33f97de5e60c91

Observation fe57144d-5fa1-42e1-8c75-b318be294a27 · outbound

This paper cites Audio-visual segmentation with semantics.International Journal of Computer Vision, pages 1–21, 2024.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Audio-visual segmentation with semantics.International Journal of Computer Vision, pages 1–21, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.443673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.647139Z digest=sha256:f43aa147ca898c4d81744e1c9c1ae9ed1754c3e21a0c7939acb2db7a6c19d37a

Observation 100a54ff-3995-45cd-b92f-440fe0f4a08f · outbound

This paper cites Enhancing geometric factors in model learning and inference for object detection and instance segmentation.IEEE transactions on cybernetics, 52(8):8574–8586, 2021.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Enhancing geometric factors in model learning and inference for object detection and instance segmentation.IEEE transactions on cybernetics, 52(8):8574–8586, 2021

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.433369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.650416Z digest=sha256:9e684971444cdbd271ff1111389ca1a4c66c70d5de9d594395afd959d4551d6b

Observation cc351f6b-87d6-47af-a2fa-68c710048516 · outbound

This paper cites URL https: //www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_ Card_Claude_3.pdf.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs URL https: //www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_ Card_Claude_3.pdf

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.423501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.653489Z digest=sha256:053dfc851ed60d4a7cf6298ef72af4d690088717b02eccbf7d97fd18b920d966

Observation ab692d4d-79fa-40b9-b859-571f1388fd5e · outbound

This paper cites Onellm: One framework to align all modalities with language.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Onellm: One framework to align all modalities with language

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.657817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.657817Z digest=sha256:ee30ddab154f1a2050b747163fdcdd9bb09c47c912581272c183ad394c4347bc

Observation 71987732-166b-49b7-b8f0-00bddd142d99 · outbound

This paper cites GroundingGPT:Language Enhanced Multi-modal Grounding Model.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs GroundingGPT:Language Enhanced Multi-modal Grounding Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.661256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.661256Z digest=sha256:b5ee32809eef6e7058b001d141aff5983dde2b550f7f2efc40f7ed0251eebd11

Observation 6005d5e9-493a-467d-9ed5-a8d3c08c7d5d · outbound

This paper cites Valor: Vision-audio-language omni-perception pretraining model and dataset.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Valor: Vision-audio-language omni-perception pretraining model and dataset.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.407310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.664610Z digest=sha256:1b515d8a6a145c04ad7a6c0011fcda988bd0f32b58c2b5287d80e3b799770e84

Observation a29cf447-a518-4291-b3eb-d092abc80390 · outbound

This paper cites X-instructblip: A framework for aligning image, 3d, audio, video to llms and its emergent cross-modal reasoning.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs X-instructblip: A framework for aligning image, 3d, audio, video to llms and its emergent cross-modal reasoning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.397918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.667783Z digest=sha256:d4e92947e5fd1bcf5cf1d55ca8128417ce0498fb0e79d3544a31240512ba5bed

Observation c550d55b-4f9b-4f2d-89a0-50e4a988721a · outbound

This paper cites Avicuna: Audio-visual llm with interleaver and context-boundary alignment for temporal referential dialogue.arXiv e-prints, pages arXiv–2403, 2024.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Avicuna: Audio-visual llm with interleaver and context-boundary alignment for temporal referential dialogue.arXiv e-prints, pages arXiv–2403, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.387217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.670950Z digest=sha256:bde374cbe5102710d3849495304343f74c4ef7ea32ad781723cd238f8a211513

Observation 484f6ce8-fdd9-4447-b52b-7e1f6f78b46d · outbound

This paper cites AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.674331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.674331Z digest=sha256:29dad97defc0e31357b13510f0fcbd10d070ae8866c04c73c87bf86a5dc62293

Observation 0b5835ec-d2b5-4b49-8558-da2721948a9c · outbound

This paper cites Omnibench: Towards the future of universal omni-language models.arXiv preprint arXiv:2409.15272, 2024.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Omnibench: Towards the future of universal omni-language models.arXiv preprint arXiv:2409.15272, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.677898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.677898Z digest=sha256:04a4fd04061bb24ff3e54b6477453fb9c7cb830d68b966fd696d8e88a32b19c7

Observation 5dc4b245-3c20-459a-9398-0aedc32eeaa4 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.680865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.680865Z digest=sha256:91d2f0d8279790121f101c2b606c63ac19f1fda70dd8eef418e21e8524e14381

Observation 0b62d6ce-670d-4aac-9cc7-e14526b3520e · outbound

This paper cites How many people spoke in the scene showing the conference table?.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs How many people spoke in the scene showing the conference table?

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.377743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.684003Z digest=sha256:b8ad55bf206fff69113b8d84d1d1a5666e251618c87596cafa35b1805a61c9a9

Observation b2e7becb-f94b-4077-a055-9a04c361c860 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.367572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.687707Z digest=sha256:09128f6f5d192212e6ed48e1892f584f231906b0bcee5d864258bbaf5929ad7c

Observation 466bba4f-c527-4377-93ee-848320097d79 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.357940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.691318Z digest=sha256:3271d6de7abdab7ea392b6cb6e3422b9f678f968e0510ec04e3212f9ee4d6b77

Observation 6181835e-8342-4651-8f75-1f781fa56004 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.347782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.694655Z digest=sha256:ceda8366468cdf7291f90ddeac5da6dfbe19f53cb823623325604c5cb300b4f0

Observation 24fcc273-cade-443a-a195-8b8c300f5146 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.338960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.697979Z digest=sha256:6c6ec707b7e7c3c8186447f8e529350d22b5fbb9676f32bf9ac55e16407250cd

Observation 4cc85d94-89fc-4989-bc42-0e2dc4b9a040 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.328979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.701456Z digest=sha256:28092cd27d0d0b2e6137bdcb7ca44a38829f79648fdbf1185b51bd293274f7e3

Observation 30856085-49c4-4a07-912b-1ea40b179069 · outbound

This paper cites an unresolved cited work.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:27:06.319383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.705052Z digest=sha256:cb3bbfdf7d1b40f6e708e256fecb726211a87f1665f3cb419764a01bcbec1c18

Observation 476f044f-235e-4f69-b4dd-82d77e9b3a79 · outbound

This paper cites question.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs question

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:27:06.309524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:27:05.708277Z digest=sha256:b670fdbf301ac5eb13620811880ec58da2ca6b549fb20878071f58845e50fabd

Pith citing papers

Observation 9c2be27d-d82f-4e75-b00e-629425ae9f37 · inbound

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models cites this paper.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:45:56.105299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:45:07.700571Z digest=sha256:f4dc91ffcbcbbe506d610d07da2dec46ca383f56f63c506e4f94f2d2de0d304e

Observation e3189c56-f1bb-4f92-8354-b7a926ffdf6e · inbound

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance cites this paper.

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T22:09:56.270516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:09:56.270516Z digest=sha256:3bb6aba1c9f872bd2fe3f26ae6606f58a0fde5a00947c24e6c7d222dff9c1dad

Observation e4baaf95-ed9e-476e-b0ff-3904fcbc7312 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.142010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:c4548c3a3d7d833407137d6a484f8875d27aca3d2d0fb3b25349edcbc64dd272

Observation da73e082-7283-4c33-aaae-e153c39db0af · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:18:32.125415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:48fe2821cf0072068c9ba3b2868f342edb52bcbbab13e74b5208dce3d7de701f

Observation 11430591-9fb8-4725-9863-066575a13626 · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.698673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T03:49:58.240883Z digest=sha256:e4f6e8df4fa7ef035c660c342b97d98d1b0544bf054c3ed4439ba9868c1c3306

Observation 2449202f-3603-4a76-9f9c-6f41ce858e3d · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:09:50.259437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T06:06:20.658030Z digest=sha256:bfcfdccb868dee36e21ec4aa809ed4b49e6b83e422b8568825df18e6a3e8569a

Observation 15ec56f8-173e-4780-b416-249258d469bf · inbound

Empowering Long-form Omni-modal Understanding with Robust Audio Perception cites this paper.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:f25bccfdf72c90b26a0c62fe2dc2de19ac03619232cc590208094fcb6bfb5718