Pith. sign in

Paper Citation Record · LEDGER

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

As of 7 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 3 inbound Pith citation observations for arXiv:2506.05414.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05414 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:50:53.648115Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T23:39:22.070629Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T23:45:08.222067Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 51c70d91-d0d3-4485-9597-c4635738810b · outbound

This paper cites Shelton and Timothy P.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Shelton and Timothy P

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.218328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:52.562058Z digest=sha256:9eda3d9856745b578d1e4f6daccf09bfa07c1af50baf38605923cb9b6bfe505e

Observation 54604aca-b920-4463-a66d-cc2af4ca5b7f · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.568788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.568788Z digest=sha256:8d8119b36f916b6d59ee4ef9eb27e9aa81e8d9793bc12f0921f55a78ece9ce16

Observation 8d839f7c-d13a-4593-8741-4824627e8516 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.575432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.575432Z digest=sha256:dbd7a308b54393ad1b25719bf2ff929a54b3c0ff4c023028c569c585ea4bb6d7

Observation d9994da5-2214-47aa-967d-1499f0e1d1ef · outbound

This paper cites Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.581847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.581847Z digest=sha256:93a96f824e23941a7c619f1beec2241294530c6d43654ec862cf641c4f46906e

Observation 7ab239f0-022e-42d6-bc5d-4a37fbd8672b · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.588064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.588064Z digest=sha256:0b61a3260ae0a6d9c2a0ba646518278b758cd018868ad473da66c46caaa8eafa

Observation 30493e97-6e53-4e79-89a1-589441b5d889 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.596168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.596168Z digest=sha256:16703765c7e18299066ba3a6942f1e0f54b8bf99031718bc45fe953183918401

Observation c43455d5-fd59-416f-a947-1c9b60b35c83 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Openeqa: Embodied question answering in the era of foundation models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.603103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.603103Z digest=sha256:ed5c0eb17f42e3fb321ba0c207fdfe0210dda7bd28b57f2bb8fe1625e6d03f6a

Observation d85400df-0ead-4fef-bad4-b78914fb1e5b · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Learning to answer questions in dynamic audio-visual scenarios

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.610231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.610231Z digest=sha256:4b855a7e1e0013d510a79414d69bf77323838e989f5d4a2e0d16ce8ea130cd3a

Observation a4d67649-81ae-4855-8570-07441948562b · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ego4d: Around the world in 3,000 hours of egocentric video

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.617297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.617297Z digest=sha256:6d20b6aace7c48706123e32c6c1fe731a27f0dbdc2eadbb2f297a295bf8f1383

Observation e126f361-8d8a-4cd4-9614-0a94ea7d0776 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.624322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.624322Z digest=sha256:704e7fd9236f412ba2a1b305913f1311af1556829ad308784404f64538da1a7c

Observation 5815ef23-37f8-41bd-a5f7-b5a5bbffcd90 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.630717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.630717Z digest=sha256:b37da8b0218cdde28927786ed5ca61e7641e1e04cb5b1199b2140d7b77609ac8

Observation 1b379fc0-4ef3-45e3-93b6-51201ac64726 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.638441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.638441Z digest=sha256:15c7170ec52304ccaa1c3c16d71ca82d819778fd4b42562e84dd9ead53fa42c9

Observation d7856dab-999d-4b83-a02d-48a4ba13866f · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing VideoChat: Chat-Centric Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.645520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.645520Z digest=sha256:d387472a4164440485ac14563225d96aa66dd9082fc4f9f48ab61a1cce9c0904

Observation 961cccdb-6df9-40ec-8327-66d0ad9ea2b4 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.652513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.652513Z digest=sha256:3be168dc500413dbf55a769e7ba3668194be0283264323a76ebec7389803d067

Observation 64518349-8e1d-4ed3-8a5c-6eb46e3d765a · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.658655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.658655Z digest=sha256:c9c4ad26d235bd7f1d62327ad817182d4af5f5e5d05fb35be43c2d2a09146c06

Observation 3d34628c-381a-41de-ae2a-00fadf19fe1e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.664757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.664757Z digest=sha256:80c3e9d8f4b160845da333e39ef62d5d6522c23e4c33764932324b3103d9125e

Observation b19e6aaa-7840-46aa-9110-6e641672ca3c · outbound

This paper cites Listen, Think, and Understand.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Listen, Think, and Understand

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.670777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.670777Z digest=sha256:48f8716fd09786216b554244f2bcffcd38af0ff29df6421a674cc525f6e79f5b

Observation fbef4dad-59d5-4be8-ba15-309b10b28eb1 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.677692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.677692Z digest=sha256:4a6bc8515bef6c55fd961593274b0b45c4140e82adc182e156ec03b9450fe221

Observation 2f83e2ec-4967-4a44-8c46-197267c26a52 · outbound

This paper cites Kimi-Audio Technical Report.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Kimi-Audio Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.684460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.684460Z digest=sha256:13e0544f1836741b09b0e2f3b7b6eb2b33971eff44c6b673969cbff1e98c5f5f

Observation 95790d61-6845-4e0e-9e17-22c7694b5a54 · outbound

This paper cites Audio-reasoner: Improving reasoning capability in large audio language models, 2025.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Audio-reasoner: Improving reasoning capability in large audio language models, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.135905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:52.691260Z digest=sha256:816f318f324b78f749d4d3ca1be4082b7b85847fcdd041713e9474f8f545f104

Observation f773e9ef-4b51-40a0-b55a-e8f37e6c6c5d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.698238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.698238Z digest=sha256:7a10236d166ac5e34d2bba9b32a91bd36d810f2d62c1bdb36517743232a98393

Observation f5dcc519-42f4-4923-9c3a-02a13a5e89ac · outbound

This paper cites LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.705026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.705026Z digest=sha256:4b31a7908d593e156746d7583e4948c8ff38ab2a3e63c53d7df5545b74f80c62

Observation 1286094a-ef5e-48ac-9c06-67421896a221 · outbound

This paper cites Egolife: Towards egocentric life assistant, 2025.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egolife: Towards egocentric life assistant, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.712377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.712377Z digest=sha256:d67b09c5cf72151613439e4ef9c39af409b3948070b5c1faf177920eaa6f6801

Observation 757897fa-3c40-4072-bca9-bb4530bc2848 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.719441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.719441Z digest=sha256:2f83cf67ac9657b508b571c7d6574da7f5c0757d2017370c346b61cc299154f1

Observation a410c64e-8898-4152-9fd8-a7bd08e38da9 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.728028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.728028Z digest=sha256:bcb44a7da5a3136d045320ab4e800d02035290454ec6008563600965aeec0472

Observation e8db80ae-0bf3-485d-b398-3fc0753f61e4 · outbound

This paper cites video-SALMONN: Speech-enhanced audio-visual large language models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing video-SALMONN: Speech-enhanced audio-visual large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.110689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:52.734466Z digest=sha256:85908d7dc1faff770a57b3a47256aa09ab74983d28a210a1151a8b521cb6fcae

Observation d2d7f427-c064-4870-b456-31241f56a0db · outbound

This paper cites Onellm: One framework to align all modalities with language.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Onellm: One framework to align all modalities with language

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.094678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:52.741622Z digest=sha256:0d7bd1bf5bf1eaad17e98808131bcec9903139f9d967488c913b6155da3dff38

Observation b7dc4d74-c4f0-4614-b6be-86224e0a7b4e · outbound

This paper cites X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent cross-modal reasoning, 2023.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent cross-modal reasoning, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.076047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:52.749258Z digest=sha256:0c435c187a4ce55fd898cd8882c09943655b13742458726ced785bcb7cf7d190

Observation 20a70e97-adf7-4fe3-8876-f3fb30f6b31e · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.756287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.756287Z digest=sha256:8ef1525ff9f5c895fc7ace5a4d76bf214a1769c55bd575022529d47de9a8237d

Observation c633c6df-ccb1-4ba4-9085-88e1fd92b35f · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.761992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.761992Z digest=sha256:bdc15bc062a2ac05b16dce1bc3609f3a52538e74824b6f65eee6496c1679a19f

Observation 061cff8b-33b5-4e83-a9cd-f71616a5ae0e · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.768322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.768322Z digest=sha256:3f8e8c90e80d2087314d140382223919576eed67f6717b2e312597f82caee3c3

Observation c14d7703-8bd6-4d2b-88ae-881a8999a0dd · outbound

This paper cites Robohop: Segment-based topological map representation for open-world visual navigation.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Robohop: Segment-based topological map representation for open-world visual navigation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.054845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:52.775435Z digest=sha256:95eb1dd3d15a6de0dfa4eb6d39d72d067bec43cbf5279a0543665cd6e8cace8f

Observation 26c96cbc-7529-41bd-b782-88ab2c7aae1d · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.782291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.782291Z digest=sha256:4579f528a13f39a7187f4c9744a741bc39c7d13fcaa0b6bcafa85572c87d489b

Observation 5615034d-7b85-4e48-83af-89adf22e580d · outbound

This paper cites 3d-mem: 3d scene memory for embodied exploration and reasoning, 2024.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing 3d-mem: 3d scene memory for embodied exploration and reasoning, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.021240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:52.789671Z digest=sha256:9324964e47e5a99af09d25eff6609fc869dd1e7184d2d1563ad44637856c02dc

Observation cfa85fcf-2cdf-4e1d-8934-7ba0409f6ef4 · outbound

This paper cites 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.796197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.796197Z digest=sha256:c6a7c59d8982580e074638ca4350e06cfe12cd9e6afecdf4590d5f57fe1e1b1e

Observation 3eabc801-a4d0-4b7b-ae40-b2bc915861bf · outbound

This paper cites Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.802988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.802988Z digest=sha256:a2690301622c14a57be3180577a0d4e9d4066b802b7e4a6190a32fc85a5456cf

Observation 279161f6-8053-4e84-a3bf-af5e8044a284 · outbound

This paper cites Gridmm: Grid memory map for vision-and-language navigation.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Gridmm: Grid memory map for vision-and-language navigation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.809764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.809764Z digest=sha256:416aeb95693eec91be4fe0090aeb1192692b19f270a596eed8f60d48fbb91e21

Observation 1f6e8217-2a87-4753-8dee-63be6fbd7fbc · outbound

This paper cites Visual language maps for robot navigation.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Visual language maps for robot navigation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.985179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:52.816719Z digest=sha256:b61386804d9ca3e873206a8059769e2c7ea253a2cbfdca57bdc7c15e7a0e87c2

Observation 88f90bd5-fb71-45df-912b-dd854a51244c · outbound

This paper cites ChatSplat: 3D Conversational Gaussian Splatting.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing ChatSplat: 3D Conversational Gaussian Splatting

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.823519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.823519Z digest=sha256:a3eaa8ca54e51aa406933a1eaa0d3226cb60f2e7f529e2ba6c0faa5806aa296e

Observation b1d1f0eb-5cd9-4f36-aa76-5afbb6c206db · outbound

This paper cites Language embedded 3d gaus- sians for open-vocabulary scene understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Language embedded 3d gaus- sians for open-vocabulary scene understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.965124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:52.831954Z digest=sha256:0d5e977d2c859fae0e34d6a5bd655f575d4f47eca6ef771abc9e7e4b4c208429

Observation c43b3d02-3cd2-42ba-af71-3d520818c96b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.840774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.840774Z digest=sha256:924dddb5c12716651c5d4545e4f206f2f7869bfa6992c0735a085397d86548a9

Observation fff3a9a3-9b5c-44a7-98dd-53a5720c8082 · outbound

This paper cites Sound event localization and detection of overlapping sources using convolutional recurrent neural networks.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Sound event localization and detection of overlapping sources using convolutional recurrent neural networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.945823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:52.848124Z digest=sha256:283920190cec85c38dd058a03ba21216f3fefb71f1ab609f4af3284123a20130

Observation 604093e5-88f7-44b7-8c2a-be60f54dca3f · outbound

This paper cites Robust sound source tracking using srp-phat and 3d convolutional neural networks.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:300–311, 2020.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Robust sound source tracking using srp-phat and 3d convolutional neural networks.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:300–311, 2020

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.927560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:52.854282Z digest=sha256:70c2163168fb6643b0270c65ea5762c64bb7e163c5bbd8d1647a909b8a28c787

Observation efc636f5-7775-4511-97f9-5a2ddf7bfe92 · outbound

This paper cites BAT: Learning to Reason about Spatial Sounds with Large Language Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing BAT: Learning to Reason about Spatial Sounds with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.861092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.861092Z digest=sha256:b3178be3c4438e3148cfb5e3f9c2323a40a65a2a45898b08d986f959b8257cd5

Observation bb8ad452-b6a5-4306-b951-3c05b1870867 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.869518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.869518Z digest=sha256:ed12dc90c6214b4cd88e5813b485c1708b0199ae7e078e466022d7097d2090d8

Observation bc12e5ea-1e8b-4171-9691-9bf612a8fdfb · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.909983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:52.877319Z digest=sha256:22b12e6171a9c88626bdfa55867539acf2050c10edc9d21b9e17514210293190

Observation 635378ff-4dab-49c3-9e8d-0bf0d600d337 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.882782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.882782Z digest=sha256:ca0a5c25781f0c83d6dc94e1c8de6630e10d5f669ce8c7c3ad1acb6f26cd31bb

Observation 7015e766-957a-4caa-ae1a-f5ee80c334fc · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.887571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.887571Z digest=sha256:747033e9ea8d24599dfe21696dfcd239ff9151daf9db5531ab7e10ed36eba699

Observation b6c4f11c-2b5b-40bf-8e1d-df83df2d7f6e · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.895136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.895136Z digest=sha256:5a9c9ff4962c1c331a777bd4b4c62173f73cab91d378d491b7a610913af2dfd5

Observation 1b58dd93-e748-4d80-b495-5a2a619c530e · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.925613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.925613Z digest=sha256:f97c23bf897787134e8d1a9aaa6f9764419deac7432be83080725dd8d895b79e

Observation 310fc69f-3eaf-45d3-aa84-59b71bcd55fb · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egotaskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.725195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:53.049207Z digest=sha256:a1590c1040dbdf19fabec7a916c03b4b6d65772fefea43e43cf16df9f33e3057

Observation 19258998-a866-4f9e-8f3f-d97f04b0a596 · outbound

This paper cites Vitatecs: A diagnostic dataset for temporal concept understanding of video-language models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Vitatecs: A diagnostic dataset for temporal concept understanding of video-language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.186552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.186552Z digest=sha256:f3b960c2f5f669110f8000c567f16a1ccadacabb28efe6da2109899b3030ae07

Observation 9c23d341-727b-45fe-a305-5e5ebaa9831e · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing TempCompass: Do Video LLMs Really Understand Videos?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.344906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.344906Z digest=sha256:b42e6b537fbe6c4dc610f6d85a53e4a509630651713cd9e2b63cdc7f087f1e66

Observation de2fd68e-6b4e-4389-82a8-f67a378ad5e7 · outbound

This paper cites Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.504475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.504475Z digest=sha256:72f94551e0bf00a04c24843accc3649966025fe5b75bbce35160c8fae331ca76

Observation 5bb61c32-e1aa-46ba-831d-2129db380b73 · outbound

This paper cites Tackling data bias in music-avqa: Crafting a balanced dataset for unbiased question-answering.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Tackling data bias in music-avqa: Crafting a balanced dataset for unbiased question-answering

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.669457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:53.539032Z digest=sha256:d7d08e56cf11732159046d6fbcd093aa1303ac8383c841c9fca90985696c3fb3

Observation def4e580-b561-4b44-a3c7-3a91abb81f92 · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Avqa: A dataset for audio-visual question answering on videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.543428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.543428Z digest=sha256:90367198689eaf290b61dd9028e51d6aeb7fed9fe7e9d4bcf5725afe76521617

Observation 49b099fd-7b1b-4ebf-95bc-c124ef5a7b27 · outbound

This paper cites Pano-avqa: Grounded audio-visual question answering on 360deg videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Pano-avqa: Grounded audio-visual question answering on 360deg videos

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.416844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:53.548005Z digest=sha256:bd94ce61108470f24212bfd4ae1cb7a74c93de84de9f679fb6fd1da2c0f97550

Observation 729992a3-cf9f-4247-93aa-6672fe9925b3 · outbound

This paper cites Egocentric audio-visual object localization.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egocentric audio-visual object localization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.377983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:53.553335Z digest=sha256:32c9a14ba45dab97b9b24e3be02d84bf5dbfd0e5d553d6f118384d48bc983b36

Observation aa926751-849e-4ffe-be82-aada78b03281 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Scanqa: 3d question answering for spatial scene understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.558636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.558636Z digest=sha256:eecddcc892587893dfac9cb7cb7f942f232121b60704e4a55be37785bb50915f

Observation cdab30bd-6328-42ec-a59f-65cd428c537e · outbound

This paper cites Aria Everyday Activities Dataset.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Aria Everyday Activities Dataset

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.578305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.578305Z digest=sha256:1640dbb922b98b21d1d26723e4cca20e1e13220a376d2340be313eff68b82d8f

Observation 535037db-2991-4d05-9af7-42db1e362124 · outbound

This paper cites EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric Foundation Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric Foundation Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.582860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.582860Z digest=sha256:3c9f06dbebae4ba3cb36f5bcf2613688b302849166d6cc8ace1b47c7d91394f1

Observation 807f7348-d4b2-45c7-ba36-a2e8112298f8 · outbound

This paper cites Photo-slam: Real-time simul- taneous localization and photorealistic mapping for monocular stereo and rgb-d cameras.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Photo-slam: Real-time simul- taneous localization and photorealistic mapping for monocular stereo and rgb-d cameras

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.344003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:53.589037Z digest=sha256:70517b2bf6ce26936e079a1c941d9ed75d2bffc4f37fb5f72686f9bd21b7d158

Observation 1fc0ec9a-3838-4591-b4fa-0b8096f1dfec · outbound

This paper cites Image segmentation using text and image prompts.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Image segmentation using text and image prompts

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.322048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:53.594292Z digest=sha256:e911284246bad66ce07931d671385e4a94c3cfebc7c0571e60143c694d874ac2

Observation 3bca59fb-6576-4c28-8f48-23d558462aba · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing SAM 2: Segment Anything in Images and Videos

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.598703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.598703Z digest=sha256:65a0374b9baf355345a81b93bfeff13ca5c4543297464278f4e9638c63d182bd

Observation d4850f78-4ad4-4792-a478-502cc15fb36c · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.604010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.604010Z digest=sha256:56f96c35dc503e90994fa30a9bc31754c70d067477d4f2aadda1806a26f272ff

Observation 9f425254-a64c-4200-8c14-fc01a2b8f709 · outbound

This paper cites Brown University, 2000.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Brown University, 2000

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.300866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:53.610157Z digest=sha256:492cdd91511dc95f866983b42011d98a6311fd10869bc38d2cf6c4eef47c44c9

Observation 5eb5c881-4c66-4058-9e95-398f0fe18c18 · outbound

This paper cites Coherent-to-diffuse power ratio estimation for dere- verberation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 23(6):1006– 1018, 2015.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Coherent-to-diffuse power ratio estimation for dere- verberation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 23(6):1006– 1018, 2015

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.282138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:53.616226Z digest=sha256:913e2dc58e978b3a6e6d01806c452a673e0c7c4823023feb166c000b9ad29e8b

Observation 91ed6964-f625-4b2b-8f82-fc619039816a · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.621308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.621308Z digest=sha256:143566b1aacb388d4b2067d0f00d88179a7e52f773c92f3c49ef0cec1d79a92c

Observation 02188b04-effb-453d-b0b7-83a0edf3fc75 · outbound

This paper cites A new approach to linear filtering and prediction problems.Transactions of the ASME–Journal of Basic Engineering, 82(Series D):35–45, 1960.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing A new approach to linear filtering and prediction problems.Transactions of the ASME–Journal of Basic Engineering, 82(Series D):35–45, 1960

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.254676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:53.628005Z digest=sha256:2fef61d833e60e89a9d6b63d96ca66fc0bcde04e0413e1323ea8b70fbdb34feb

Observation c8411998-4f37-4049-accb-4eef08f76a8e · outbound

This paper cites The pascal visual object classes (voc) challenge.International Journal of Computer Vision, 88:303–338, 06 2010.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing The pascal visual object classes (voc) challenge.International Journal of Computer Vision, 88:303–338, 06 2010

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.635225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.635225Z digest=sha256:ead7a60ea5a675fdbb47cc2499c89a48196f30da79643d5dcaebace55f702c98

Observation aa89350e-5ea6-4487-b41b-6fccf2ed5fd6 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.641537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.641537Z digest=sha256:b59659623bbc4936d911cf76b5231d6289d30e21b5b11bc82af3f1a456d54a12

Observation 5d13a14a-8e9a-4e43-adc4-bf1ea64118bd · outbound

This paper cites an unresolved cited work.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:50:54.224692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:50:53.648115Z digest=sha256:4e8ffade265ab3ce684772c493ba63ff308cc657412586a800e4106799d86ec0

Pith citing papers

Observation ccdd4417-f720-4d73-a485-7bcd61c5b45d · inbound

EgoSound: Benchmarking Sound Understanding in Egocentric Videos cites this paper.

EgoSound: Benchmarking Sound Understanding in Egocentric Videos SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:46:42.583892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T21:44:47.636912Z digest=sha256:8df336c82447382242a8bf02f0e0770c751a122182aa6a08b27ebef3c94a84ad

Observation 162511d4-657a-4d71-8e94-51887078ef76 · inbound

Do Joint Audio-Video Generation Models Understand Physics? cites this paper.

Do Joint Audio-Video Generation Models Understand Physics? SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.223574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:cb675c0c5697c034e7af228ce8c393b9307f0a57a1863c1f437415f7ca353ebc

Observation 1105d00a-a7c5-41f2-b56c-abe588f4d4b2 · inbound

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks cites this paper.

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:13:11.900409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T10:10:31.059095Z digest=sha256:090c07ec3199efc6a64a41258c9afbb73cee761600fa5a93e260fc9dc7273cba