Pith. sign in

Paper Citation Record · LEDGER

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

As of 22 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 3 inbound Pith citation observations for arXiv:2506.05414.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05414 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:50:53.648115Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T23:39:22.070629Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T23:45:08.222067Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 51c70d91-d0d3-4485-9597-c4635738810b · outbound

This paper cites Shelton and Timothy P.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Shelton and Timothy P

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.218328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:52.562058Z digest=sha256:e1eef7b498dc6ecdb2652638f6ca3ab7a0b4da680091ca9d8ebf3f4dff8b2179

Observation 54604aca-b920-4463-a66d-cc2af4ca5b7f · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.568788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.568788Z digest=sha256:7ede0457220721b450422fd23cbb853e52183467cf3ea766355287fd43778806

Observation 8d839f7c-d13a-4593-8741-4824627e8516 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.575432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.575432Z digest=sha256:d56afe66b88d0fd13b82657669bed1d001dc10b2d92cab926871323425284381

Observation d9994da5-2214-47aa-967d-1499f0e1d1ef · outbound

This paper cites Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.581847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.581847Z digest=sha256:2da58e9aff01324835ae06c90bbea9136951a399b5c7aed520890d235e5dea3e

Observation 7ab239f0-022e-42d6-bc5d-4a37fbd8672b · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.588064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.588064Z digest=sha256:9309b1d7420c0b31e1261486ea4f73714013d30d6c0ec4ccd533178c1e45e5c9

Observation 30493e97-6e53-4e79-89a1-589441b5d889 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.596168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.596168Z digest=sha256:b2719720f157160ecace5bee50c4cd187c5f2a27040ee941ea603e88bfa0af8c

Observation c43455d5-fd59-416f-a947-1c9b60b35c83 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Openeqa: Embodied question answering in the era of foundation models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.603103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.603103Z digest=sha256:18f8a21b74249511627b7952f783c7b103e58590766a05b6aebb6f81d607cd20

Observation d85400df-0ead-4fef-bad4-b78914fb1e5b · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Learning to answer questions in dynamic audio-visual scenarios

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.610231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.610231Z digest=sha256:8687724b8a91c24e826630ff2c0b98cf4e73857a7e619d8347a77ca2af21136a

Observation a4d67649-81ae-4855-8570-07441948562b · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ego4d: Around the world in 3,000 hours of egocentric video

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.617297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.617297Z digest=sha256:3f53f3d835cce60eada6bca1c2bc7260ad1f6336820fe877f18a4e754d5e4bc1

Observation e126f361-8d8a-4cd4-9614-0a94ea7d0776 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.624322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.624322Z digest=sha256:4016736340307c78d33bd6145c40d2d91a8cd6dbf3412e38de278d7b88a8123e

Observation 5815ef23-37f8-41bd-a5f7-b5a5bbffcd90 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.630717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.630717Z digest=sha256:4b0642e036b515cbcca3b654fef8fd6b4af8836dda614abca36fc75efa0e1408

Observation 1b379fc0-4ef3-45e3-93b6-51201ac64726 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.638441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.638441Z digest=sha256:b18b919fb3f441a14942969bb06d708d635301a303796e86e2a9a6afbe43bcef

Observation d7856dab-999d-4b83-a02d-48a4ba13866f · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing VideoChat: Chat-Centric Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.645520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.645520Z digest=sha256:07433f6f6eba275ce4dddb6226179ec92bc6e33acf63fc176c372160dc0e732c

Observation 961cccdb-6df9-40ec-8327-66d0ad9ea2b4 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.652513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.652513Z digest=sha256:8cee4aa7d6114c30163c31b81a780bbca3b0b32ed72d2eaa5904b579ad7b5623

Observation 64518349-8e1d-4ed3-8a5c-6eb46e3d765a · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.658655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.658655Z digest=sha256:23a85e507ad9f10eafa3770258bc44b808c84a48a7eeeac7f7fdcd1752c0c44f

Observation 3d34628c-381a-41de-ae2a-00fadf19fe1e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.664757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.664757Z digest=sha256:3e57d942235ea90f16c31b2886a4f39c5d49b6ae73307074824e78de46a49261

Observation b19e6aaa-7840-46aa-9110-6e641672ca3c · outbound

This paper cites Listen, Think, and Understand.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Listen, Think, and Understand

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.670777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.670777Z digest=sha256:3107276467b186c690c22e6557c39c0874c6ccd3c856a3bc0e498ada239f32c7

Observation fbef4dad-59d5-4be8-ba15-309b10b28eb1 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.677692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.677692Z digest=sha256:d913ed4da7d4c6ce6a1e1fa4094328f3fc5ea8a785e22333d65d6abfdee5b29a

Observation 2f83e2ec-4967-4a44-8c46-197267c26a52 · outbound

This paper cites Kimi-Audio Technical Report.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Kimi-Audio Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.684460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.684460Z digest=sha256:51faf8f41b59a072362bcff821dc102675278995d198a4aa2d5302291ad36003

Observation 95790d61-6845-4e0e-9e17-22c7694b5a54 · outbound

This paper cites Audio-reasoner: Improving reasoning capability in large audio language models, 2025.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Audio-reasoner: Improving reasoning capability in large audio language models, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.135905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:52.691260Z digest=sha256:68663b30df2547b06bbeff10c9f77107a63f7c4a1071f5775b72e25d2f1ab836

Observation f773e9ef-4b51-40a0-b55a-e8f37e6c6c5d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.698238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.698238Z digest=sha256:7c300deb32625dec478d8408aa2220b8697dedcce8526d99e18495a601e426f2

Observation f5dcc519-42f4-4923-9c3a-02a13a5e89ac · outbound

This paper cites LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.705026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.705026Z digest=sha256:dac46a7e6718dd53c8f619a90f0477614ef6430b6a89115fbca81f30119fb34e

Observation 1286094a-ef5e-48ac-9c06-67421896a221 · outbound

This paper cites Egolife: Towards egocentric life assistant, 2025.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egolife: Towards egocentric life assistant, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.712377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.712377Z digest=sha256:c18bf91df92dfb3303b34bd5498d985f149ce5dc5f05f02dcd902a46d868638c

Observation 757897fa-3c40-4072-bca9-bb4530bc2848 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.719441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.719441Z digest=sha256:042547e5a0cf73456e37d3e2ce1a73ede5e07abcc5b5d0366bef47b489bc2e6e

Observation a410c64e-8898-4152-9fd8-a7bd08e38da9 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.728028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.728028Z digest=sha256:f9cb08dc3357667989135d86b4b96892e21a1f657120d7d0793bf829b8bfa009

Observation e8db80ae-0bf3-485d-b398-3fc0753f61e4 · outbound

This paper cites video-SALMONN: Speech-enhanced audio-visual large language models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing video-SALMONN: Speech-enhanced audio-visual large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.110689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:52.734466Z digest=sha256:0207929f6e8944c1c5788621b1b25dd014d68b78c9032c63d8989230acf85a10

Observation d2d7f427-c064-4870-b456-31241f56a0db · outbound

This paper cites Onellm: One framework to align all modalities with language.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Onellm: One framework to align all modalities with language

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.094678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:52.741622Z digest=sha256:f46a381f8d94b1b0c45e16438bce442e41a6b6d64916bbe441b80bed169cc2ee

Observation b7dc4d74-c4f0-4614-b6be-86224e0a7b4e · outbound

This paper cites X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent cross-modal reasoning, 2023.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent cross-modal reasoning, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.076047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:52.749258Z digest=sha256:3e65391d5278e73bb4afd0f6c9922f7d39ffb8aa588222f15077e9d324b65d50

Observation 20a70e97-adf7-4fe3-8876-f3fb30f6b31e · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.756287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.756287Z digest=sha256:3c047df2237d1ec4eb503b93d74085495702bf84edc64b531187d62a3ffc7e70

Observation c633c6df-ccb1-4ba4-9085-88e1fd92b35f · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.761992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.761992Z digest=sha256:f767f6f085748e1660989412f7375fb561140e252558f8889b866ed3deeff879

Observation 061cff8b-33b5-4e83-a9cd-f71616a5ae0e · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.768322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.768322Z digest=sha256:ec7076069d316d7dd2e41b2fdedeab4e940c211e5848c51c7b8424057fa02e68

Observation c14d7703-8bd6-4d2b-88ae-881a8999a0dd · outbound

This paper cites Robohop: Segment-based topological map representation for open-world visual navigation.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Robohop: Segment-based topological map representation for open-world visual navigation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.054845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:52.775435Z digest=sha256:5e6d25bb5069bec4e4f0b9fc0272efc27c8c7a1dbd7db43f0690e8b6873b8fe6

Observation 26c96cbc-7529-41bd-b782-88ab2c7aae1d · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.782291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.782291Z digest=sha256:e18fdc164a744072dc547752abd9fb1ab617fdd9ce68f8afeed7cef15f59d4f4

Observation 5615034d-7b85-4e48-83af-89adf22e580d · outbound

This paper cites 3d-mem: 3d scene memory for embodied exploration and reasoning, 2024.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing 3d-mem: 3d scene memory for embodied exploration and reasoning, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:55.021240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:52.789671Z digest=sha256:bc1c14773811d21da0d42bfc626af8e3c2dc98726e046066c7d8fc77dc1f2e18

Observation cfa85fcf-2cdf-4e1d-8934-7ba0409f6ef4 · outbound

This paper cites 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.796197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.796197Z digest=sha256:eb4bb44ccaae4d98cffc841a0c2c1dff5283b1d9b1a15b17f64232698cfed7aa

Observation 3eabc801-a4d0-4b7b-ae40-b2bc915861bf · outbound

This paper cites Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.802988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.802988Z digest=sha256:e20ae7a5fc593be0eff4da3782e13a547ef4e01fcd180ed9b8a33883d38ce1fa

Observation 279161f6-8053-4e84-a3bf-af5e8044a284 · outbound

This paper cites Gridmm: Grid memory map for vision-and-language navigation.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Gridmm: Grid memory map for vision-and-language navigation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.809764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.809764Z digest=sha256:cc2172c3bdbd797412c240e46de56695240126a4a5fa75fed813f066872d6633

Observation 1f6e8217-2a87-4753-8dee-63be6fbd7fbc · outbound

This paper cites Visual language maps for robot navigation.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Visual language maps for robot navigation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.985179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:52.816719Z digest=sha256:6d4e8444d4b131fbccddd48e6ecaf179f75982099909205b7c362f6892a6a777

Observation 88f90bd5-fb71-45df-912b-dd854a51244c · outbound

This paper cites ChatSplat: 3D Conversational Gaussian Splatting.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing ChatSplat: 3D Conversational Gaussian Splatting

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.823519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.823519Z digest=sha256:04d18e077630b97ef17a9bf72ffdf789aee5af89d4ec09adf253401468966837

Observation b1d1f0eb-5cd9-4f36-aa76-5afbb6c206db · outbound

This paper cites Language embedded 3d gaus- sians for open-vocabulary scene understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Language embedded 3d gaus- sians for open-vocabulary scene understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.965124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:52.831954Z digest=sha256:1f120fb8af7e7f124c8e3bc31876f731fdbf0ebb057cddba5184fe743f74dbe0

Observation c43b3d02-3cd2-42ba-af71-3d520818c96b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.840774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.840774Z digest=sha256:27afec708e1b7dc53ffbbb4dcf36443d74bfb8da2bc57a7c740eb5ffcdea9c97

Observation fff3a9a3-9b5c-44a7-98dd-53a5720c8082 · outbound

This paper cites Sound event localization and detection of overlapping sources using convolutional recurrent neural networks.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Sound event localization and detection of overlapping sources using convolutional recurrent neural networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.945823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:52.848124Z digest=sha256:57ecf36e8433cd237db57635fbd714a9137bb98f8267f1b8d7bd29187ae1f463

Observation 604093e5-88f7-44b7-8c2a-be60f54dca3f · outbound

This paper cites Robust sound source tracking using srp-phat and 3d convolutional neural networks.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:300–311, 2020.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Robust sound source tracking using srp-phat and 3d convolutional neural networks.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:300–311, 2020

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.927560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:52.854282Z digest=sha256:1303d66ea7f04ae9f78989e6241ea01d8b47ce3a350194c406975d66d92ed202

Observation efc636f5-7775-4511-97f9-5a2ddf7bfe92 · outbound

This paper cites BAT: Learning to Reason about Spatial Sounds with Large Language Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing BAT: Learning to Reason about Spatial Sounds with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.861092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.861092Z digest=sha256:43aef7ad605ed504413365db61bbd3438fb64f3e81d19e71ed8755fd99304192

Observation bb8ad452-b6a5-4306-b951-3c05b1870867 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.869518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.869518Z digest=sha256:816fa896f4d603badb3d79b6b8155ca7a84e49f33be277e07ace13bfd5067cfd

Observation bc12e5ea-1e8b-4171-9691-9bf612a8fdfb · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.909983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:52.877319Z digest=sha256:f10df388d9041a405f3ea040bdf42e6c5340d21e3db0e727a4996d1652f162a5

Observation 635378ff-4dab-49c3-9e8d-0bf0d600d337 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.882782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.882782Z digest=sha256:a1b9c1867b2f68c62b18631448d4dad56241860988ea2d233e302d5948a9fce7

Observation 7015e766-957a-4caa-ae1a-f5ee80c334fc · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.887571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.887571Z digest=sha256:02b14d329e83ad8e03a975322385add2aec5d456f1114f249ef0d04c0a9ffe0f

Observation b6c4f11c-2b5b-40bf-8e1d-df83df2d7f6e · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.895136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.895136Z digest=sha256:345c80c3db92c82064df46fc16c586770e920260165d6927eb53f60e8f69c01b

Observation 1b58dd93-e748-4d80-b495-5a2a619c530e · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.925613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.925613Z digest=sha256:e122f345fe1a6125fb3816c2b22cd4fb9cd9375e3605a8332da6d7510ee141b8

Observation 310fc69f-3eaf-45d3-aa84-59b71bcd55fb · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egotaskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.725195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:53.049207Z digest=sha256:7005f26c735893cd074ce52bb7b75d78d6061dfd8b1f107fdd8404ff8f78afb0

Observation 19258998-a866-4f9e-8f3f-d97f04b0a596 · outbound

This paper cites Vitatecs: A diagnostic dataset for temporal concept understanding of video-language models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Vitatecs: A diagnostic dataset for temporal concept understanding of video-language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.186552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.186552Z digest=sha256:342a3294b763e4f2b542f26d10135828531bb97b97b3b23469fe24b13a6ea229

Observation 9c23d341-727b-45fe-a305-5e5ebaa9831e · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing TempCompass: Do Video LLMs Really Understand Videos?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.344906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.344906Z digest=sha256:79c5884ab5336235f95fc5649f8f6cee974c824220f6fcc21adf33cc2aee90c3

Observation de2fd68e-6b4e-4389-82a8-f67a378ad5e7 · outbound

This paper cites Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.504475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.504475Z digest=sha256:f8f44191880f299069ed8435fe4d4c2b146bc02027d55a1a4ddf4ca93c33be20

Observation 5bb61c32-e1aa-46ba-831d-2129db380b73 · outbound

This paper cites Tackling data bias in music-avqa: Crafting a balanced dataset for unbiased question-answering.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Tackling data bias in music-avqa: Crafting a balanced dataset for unbiased question-answering

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.669457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:53.539032Z digest=sha256:7e91b686391a2443ea9608abd5f1eb7cb23394c13d12ebd5a64c7fc817198dd6

Observation def4e580-b561-4b44-a3c7-3a91abb81f92 · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Avqa: A dataset for audio-visual question answering on videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.543428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.543428Z digest=sha256:1da0c4c883988df6d1110bcafc563dbc55c78f380e27961fa4e8ce63340488c6

Observation 49b099fd-7b1b-4ebf-95bc-c124ef5a7b27 · outbound

This paper cites Pano-avqa: Grounded audio-visual question answering on 360deg videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Pano-avqa: Grounded audio-visual question answering on 360deg videos

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.416844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:53.548005Z digest=sha256:3971c9acf02e0a5399eedb5e6c1fd3e7f09b2eccce73506f3d92d68559ee3d5a

Observation 729992a3-cf9f-4247-93aa-6672fe9925b3 · outbound

This paper cites Egocentric audio-visual object localization.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egocentric audio-visual object localization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.377983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:53.553335Z digest=sha256:92470a93f0eb3a35f66f1982dbfb3f5abc298c3a224d774a64dd28404292036e

Observation aa926751-849e-4ffe-be82-aada78b03281 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Scanqa: 3d question answering for spatial scene understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.558636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.558636Z digest=sha256:939b0895fa0bf9a95fe67a02481b8e34a50720a41710923ce1baed5bac7fc77d

Observation cdab30bd-6328-42ec-a59f-65cd428c537e · outbound

This paper cites Aria Everyday Activities Dataset.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Aria Everyday Activities Dataset

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.578305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.578305Z digest=sha256:4645b33bec2c0daa6da36e94c3930b46577a8582e4470bcb12e7d4eae9f9aea2

Observation 535037db-2991-4d05-9af7-42db1e362124 · outbound

This paper cites EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric Foundation Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric Foundation Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.582860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.582860Z digest=sha256:22a20a93046bdd5564c65f1cca5b7792d80059c01a06bba016f959d87f668a28

Observation 807f7348-d4b2-45c7-ba36-a2e8112298f8 · outbound

This paper cites Photo-slam: Real-time simul- taneous localization and photorealistic mapping for monocular stereo and rgb-d cameras.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Photo-slam: Real-time simul- taneous localization and photorealistic mapping for monocular stereo and rgb-d cameras

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.344003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:53.589037Z digest=sha256:b0fe5df8b733b689700b15bc8ee122c34084cdb1e67856f8b3971cbd1d53f0ec

Observation 1fc0ec9a-3838-4591-b4fa-0b8096f1dfec · outbound

This paper cites Image segmentation using text and image prompts.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Image segmentation using text and image prompts

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.322048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:53.594292Z digest=sha256:abaefe8fb10d7df6d297e5f63623ff872f5936113eda1ebf1dd641b6ae00418d

Observation 3bca59fb-6576-4c28-8f48-23d558462aba · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing SAM 2: Segment Anything in Images and Videos

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.598703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.598703Z digest=sha256:5cf6f8990f9fde84e346450662ba32748bb5a98e8bf9d46b8dece6d405fcd04d

Observation d4850f78-4ad4-4792-a478-502cc15fb36c · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.604010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.604010Z digest=sha256:5a8a437b303c059976e21c87424692f590be93bf87ca3946e794ea02f4ee03bf

Observation 9f425254-a64c-4200-8c14-fc01a2b8f709 · outbound

This paper cites Brown University, 2000.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Brown University, 2000

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.300866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:53.610157Z digest=sha256:82a53cb3ee2f8a619fcadb4c5320c5ce6e25d38f08e563902e93b16d86a22367

Observation 5eb5c881-4c66-4058-9e95-398f0fe18c18 · outbound

This paper cites Coherent-to-diffuse power ratio estimation for dere- verberation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 23(6):1006– 1018, 2015.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Coherent-to-diffuse power ratio estimation for dere- verberation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 23(6):1006– 1018, 2015

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.282138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:53.616226Z digest=sha256:2bcdc3d09d7dc98ce2f12699a71af178ecfbcabb223783a516d4cd0ed7507787

Observation 91ed6964-f625-4b2b-8f82-fc619039816a · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.621308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.621308Z digest=sha256:d81b419f7cf868236e496436f5bd30dbc585e2f754a98f8a53738c751249ebfe

Observation 02188b04-effb-453d-b0b7-83a0edf3fc75 · outbound

This paper cites A new approach to linear filtering and prediction problems.Transactions of the ASME–Journal of Basic Engineering, 82(Series D):35–45, 1960.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing A new approach to linear filtering and prediction problems.Transactions of the ASME–Journal of Basic Engineering, 82(Series D):35–45, 1960

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:50:54.254676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:53.628005Z digest=sha256:1e9bdf687bd904a7eb7631fccf8c0a8346e5f9bb2b93c7db4b1f2700a4e681d2

Observation c8411998-4f37-4049-accb-4eef08f76a8e · outbound

This paper cites The pascal visual object classes (voc) challenge.International Journal of Computer Vision, 88:303–338, 06 2010.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing The pascal visual object classes (voc) challenge.International Journal of Computer Vision, 88:303–338, 06 2010

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.635225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.635225Z digest=sha256:09ee403b14efeb22827f4095cb1a353a6a0a78bcc7bbe7bf1d757242602f19ea

Observation aa89350e-5ea6-4487-b41b-6fccf2ed5fd6 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:53.641537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:53.641537Z digest=sha256:d497624c6fdad0bfab44bbdf3d6705a48c61b9c5f3d56b660c494e5047462c09

Observation 5d13a14a-8e9a-4e43-adc4-bf1ea64118bd · outbound

This paper cites an unresolved cited work.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:50:54.224692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:50:53.648115Z digest=sha256:7229a660cb3dcc7301e21b3389c8f0d619b451826d0a0a5dbed0ba1a078e6337

Pith citing papers

Observation ccdd4417-f720-4d73-a485-7bcd61c5b45d · inbound

EgoSound: Benchmarking Sound Understanding in Egocentric Videos cites this paper.

EgoSound: Benchmarking Sound Understanding in Egocentric Videos SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:46:42.583892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T21:44:47.636912Z digest=sha256:2dad551b6a374afd02bcabebb9a38490b80c40cf148d62989077c74be60812f3

Observation 162511d4-657a-4d71-8e94-51887078ef76 · inbound

Do Joint Audio-Video Generation Models Understand Physics? cites this paper.

Do Joint Audio-Video Generation Models Understand Physics? SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.223574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:efa7eb46e885a368a2aacb0206d0e6796cd4e6b8415c37044823bd995764be97

Observation 1105d00a-a7c5-41f2-b56c-abe588f4d4b2 · inbound

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks cites this paper.

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:13:11.900409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:10:31.059095Z digest=sha256:2e5cae9ed130b8880b105f2f7741be139eedccfd3eadee802cf14359240f0958