Pith. sign in

Paper Citation Record · LEDGER

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning

As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2508.07470.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.07470 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:09:16.309676Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:04:02.263300Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:05:03.827879Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf48b8a6-817a-4536-915f-b4ca919a677b · outbound

This paper cites EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T22:09:16.793393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:09:15.243340Z digest=sha256:28b7f6d0ca85cf29ae7fe70268605d6c835de303017038c41e58d65e874e9452

Observation 3de3c3d9-f400-40bc-8f61-ad4adf1bb35d · outbound

This paper cites Girdhar, R.; El-Nouby, A.; Liu, Z.; Singh, M.; Alwala, K.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Girdhar, R.; El-Nouby, A.; Liu, Z.; Singh, M.; Alwala, K

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.343828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.343828Z digest=sha256:81b69f3a281d5eda39ef930d06e6632ce15dcaa48c165d9be68841430e9e24b5

Observation 5dc66f6d-baec-4bf6-9559-c827d76c459a · outbound

This paper cites Multimodal Pretraining for Dense Video Captioning.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Multimodal Pretraining for Dense Video Captioning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.603717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.603717Z digest=sha256:0a0664a824f5340e522334ae684b73dde86a36130c5b06dc208034e6aa39c1a3

Observation 2b12457c-5b01-447d-ac71-1375b9ea22f4 · outbound

This paper cites Jina CLIP: Your CLIP Model Is Also Your Text Retriever.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.681915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.681915Z digest=sha256:cf6f8bfd4285179fefec63f7e1d9c26459ef8d6673554b323af202be81ce8c95

Observation 32fe63cb-5832-4136-91ac-6194506ec5c5 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.767519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.767519Z digest=sha256:5ff628157e459a798164ea2dbf2ba229795305028ca655f722da1cbcc45b29ea

Observation 6ce1a547-b414-41aa-948d-4eef6ec84102 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.844231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.844231Z digest=sha256:d463d06d3c99d0e47171b5d11af89db75907af4593ef89301725008a0f5b5352

Observation 04d8c846-5ee5-46db-86a6-f8a3f6ead5d3 · outbound

This paper cites Qwen2.5 Technical Report.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Qwen2.5 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:16.061859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:16.061859Z digest=sha256:7fee52133d991764853213407bfb642e652b45bf00c2242344abb4fe6d1538eb

Observation b8227ad1-5869-42b2-abee-d7d250b9f5d2 · outbound

This paper cites UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T22:09:16.487647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:09:16.128976Z digest=sha256:a9db7a5ea56f4acfaca80a2368273b063ca5f8e3e4adda6dd81a3e379dd42b3d

Observation 0b428c00-5abe-47ad-9039-1b4c412a6a56 · outbound

This paper cites Video Question Answering: Datasets, Algorithms and Challenges.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Video Question Answering: Datasets, Algorithms and Challenges

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:16.203189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:16.203189Z digest=sha256:dcc322340a56390f216bb184df943f38daaa1c3a368c52c391825c2b5dcca48f

Observation 3deea570-b5f0-447b-afb2-615e92829ac0 · outbound

This paper cites In Inter- national Conference on Learning Representations (ICLR).

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning In Inter- national Conference on Learning Representations (ICLR)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:09:17.047351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:09:16.309676Z digest=sha256:99b45f9ed01a4e83de049addde4b2496d8250b1ede273b1311d389a92f848747

Observation 9838723f-089b-415e-9c58-314d24b06794 · outbound

This paper cites Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.943570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.943570Z digest=sha256:44eb6c462da8fe6869d28682a04a7e2016860a7cf0ded47f7301118ec39b4eeb

Observation 92a9f263-beb7-438d-bfcf-52840dfc59a3 · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning ImageBind-LLM: Multi-modality Instruction Tuning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.507027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.507027Z digest=sha256:71e2a1f2f7632a0512fe937ea389918c55c84889a0a4d280c32086717c750651

Observation ad6c6200-5aad-49e1-8cbb-e7cb7e6f0e10 · outbound

This paper cites In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 11287–11297.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 11287–11297

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:09:17.257252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:09:15.424425Z digest=sha256:08b77e141991de28f603d6cff1eafb187fe227a1b0e8bdea846a259f0ce83b2c

Observation 15aa2a3b-5e70-45df-922d-dd1d97003e78 · outbound

This paper cites AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:14.941204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:14.941204Z digest=sha256:9f482ca3c6ca1478bb5f8c9aaa8e2d02a37b49599277408a6196a6b0451e79b4

Observation ea2b809f-564d-4508-92df-bd75ce360936 · outbound

This paper cites See https://vicuna.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning See https://vicuna

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:09:17.434401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:09:15.153338Z digest=sha256:938e69cd76f5036d26505f3751fed70ef8f33e93d05db53e0b27abe619419011

Observation ce3dab9b-d7de-4e10-903b-f2b647d01ca7 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.067940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.067940Z digest=sha256:b59e7d1b7967c31b14ffdd3177f5002c4fdc855a52acb84988e7fc389e7d779f

Observation de74c809-b302-4a26-9fd4-f05ed2a14e2f · outbound

This paper cites Non-invertible Symmetries in 2D from Type IIB String Theory.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Non-invertible Symmetries in 2D from Type IIB String Theory

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:14.844607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:14.844607Z digest=sha256:9a7c1921c57ac74fcef9e87f587ee97bbeca8e75be23f2379e5d860a7b5fdc6d

Pith citing papers

Observation 7d8ab9f7-1791-41e7-8467-231f65af8c49 · inbound

Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces cites this paper.

Uncovering the Representation Geometry of Minimal Cores in Overcomplete Reasoning Traces AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning

Reference 79

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T21:05:03.829322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:04:02.263300Z digest=sha256:334fdb221a2e87dd292b6317ab0e70e036bb42e82809e1f23f75ba13bdc4a78e