Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:50:53.648115Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 3 inbound Pith citation observations for arXiv:2506.05414.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:50:53.648115Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T23:39:22.070629Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T23:45:08.222067Z
72 of 72 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 51c70d91-d0d3-4485-9597-c4635738810b · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Shelton and Timothy P
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54604aca-b920-4463-a66d-cc2af4ca5b7f · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d839f7c-d13a-4593-8741-4824627e8516 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9994da5-2214-47aa-967d-1499f0e1d1ef · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab239f0-022e-42d6-bc5d-4a37fbd8672b · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30493e97-6e53-4e79-89a1-589441b5d889 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c43455d5-fd59-416f-a947-1c9b60b35c83 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Openeqa: Embodied question answering in the era of foundation models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d85400df-0ead-4fef-bad4-b78914fb1e5b · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Learning to answer questions in dynamic audio-visual scenarios
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d67649-81ae-4855-8570-07441948562b · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ego4d: Around the world in 3,000 hours of egocentric video
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e126f361-8d8a-4cd4-9614-0a94ea7d0776 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5815ef23-37f8-41bd-a5f7-b5a5bbffcd90 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b379fc0-4ef3-45e3-93b6-51201ac64726 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7856dab-999d-4b83-a02d-48a4ba13866f · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing VideoChat: Chat-Centric Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 961cccdb-6df9-40ec-8327-66d0ad9ea2b4 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64518349-8e1d-4ed3-8a5c-6eb46e3d765a · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d34628c-381a-41de-ae2a-00fadf19fe1e · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LLaVA-OneVision: Easy Visual Task Transfer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19e6aaa-7840-46aa-9110-6e641672ca3c · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Listen, Think, and Understand
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbef4dad-59d5-4be8-ba15-309b10b28eb1 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f83e2ec-4967-4a44-8c46-197267c26a52 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Kimi-Audio Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95790d61-6845-4e0e-9e17-22c7694b5a54 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Audio-reasoner: Improving reasoning capability in large audio language models, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f773e9ef-4b51-40a0-b55a-e8f37e6c6c5d · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5dcc519-42f4-4923-9c3a-02a13a5e89ac · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1286094a-ef5e-48ac-9c06-67421896a221 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egolife: Towards egocentric life assistant, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 757897fa-3c40-4072-bca9-bb4530bc2848 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ola: Pushing the Frontiers of Omni-Modal Language Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a410c64e-8898-4152-9fd8-a7bd08e38da9 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8db80ae-0bf3-485d-b398-3fc0753f61e4 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing video-SALMONN: Speech-enhanced audio-visual large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2d7f427-c064-4870-b456-31241f56a0db · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Onellm: One framework to align all modalities with language
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7dc4d74-c4f0-4614-b6be-86224e0a7b4e · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent cross-modal reasoning, 2023
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20a70e97-adf7-4fe3-8876-f3fb30f6b31e · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c633c6df-ccb1-4ba4-9085-88e1fd92b35f · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 061cff8b-33b5-4e83-a9cd-f71616a5ae0e · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c14d7703-8bd6-4d2b-88ae-881a8999a0dd · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Robohop: Segment-based topological map representation for open-world visual navigation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26c96cbc-7529-41bd-b782-88ab2c7aae1d · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5615034d-7b85-4e48-83af-89adf22e580d · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing 3d-mem: 3d scene memory for embodied exploration and reasoning, 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfa85fcf-2cdf-4e1d-8934-7ba0409f6ef4 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eabc801-a4d0-4b7b-ae40-b2bc915861bf · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 279161f6-8053-4e84-a3bf-af5e8044a284 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Gridmm: Grid memory map for vision-and-language navigation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f6e8217-2a87-4753-8dee-63be6fbd7fbc · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Visual language maps for robot navigation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88f90bd5-fb71-45df-912b-dd854a51244c · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing ChatSplat: 3D Conversational Gaussian Splatting
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d1f0eb-5cd9-4f36-aa76-5afbb6c206db · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Language embedded 3d gaus- sians for open-vocabulary scene understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c43b3d02-3cd2-42ba-af71-3d520818c96b · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff3a9a3-9b5c-44a7-98dd-53a5720c8082 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Sound event localization and detection of overlapping sources using convolutional recurrent neural networks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 604093e5-88f7-44b7-8c2a-be60f54dca3f · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Robust sound source tracking using srp-phat and 3d convolutional neural networks.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:300–311, 2020
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation efc636f5-7775-4511-97f9-5a2ddf7bfe92 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing BAT: Learning to Reason about Spatial Sounds with Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb8ad452-b6a5-4306-b951-3c05b1870867 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc12e5ea-1e8b-4171-9691-9bf612a8fdfb · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 635378ff-4dab-49c3-9e8d-0bf0d600d337 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7015e766-957a-4caa-ae1a-f5ee80c334fc · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6c4f11c-2b5b-40bf-8e1d-df83df2d7f6e · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b58dd93-e748-4d80-b495-5a2a619c530e · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 310fc69f-3eaf-45d3-aa84-59b71bcd55fb · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egotaskqa: Understanding human tasks in egocentric videos.Advances in Neural Information Processing Systems, 35:3343–3360, 2022
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19258998-a866-4f9e-8f3f-d97f04b0a596 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Vitatecs: A diagnostic dataset for temporal concept understanding of video-language models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c23d341-727b-45fe-a305-5e5ebaa9831e · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing TempCompass: Do Video LLMs Really Understand Videos?
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de2fd68e-6b4e-4389-82a8-f67a378ad5e7 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb61c32-e1aa-46ba-831d-2129db380b73 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Tackling data bias in music-avqa: Crafting a balanced dataset for unbiased question-answering
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation def4e580-b561-4b44-a3c7-3a91abb81f92 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Avqa: A dataset for audio-visual question answering on videos
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49b099fd-7b1b-4ebf-95bc-c124ef5a7b27 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Pano-avqa: Grounded audio-visual question answering on 360deg videos
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 729992a3-cf9f-4247-93aa-6672fe9925b3 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Egocentric audio-visual object localization
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa926751-849e-4ffe-be82-aada78b03281 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Scanqa: 3d question answering for spatial scene understanding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdab30bd-6328-42ec-a59f-65cd428c537e · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Aria Everyday Activities Dataset
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 535037db-2991-4d05-9af7-42db1e362124 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric Foundation Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 807f7348-d4b2-45c7-ba36-a2e8112298f8 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Photo-slam: Real-time simul- taneous localization and photorealistic mapping for monocular stereo and rgb-d cameras
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1fc0ec9a-3838-4591-b4fa-0b8096f1dfec · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Image segmentation using text and image prompts
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3bca59fb-6576-4c28-8f48-23d558462aba · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing SAM 2: Segment Anything in Images and Videos
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4850f78-4ad4-4792-a478-502cc15fb36c · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f425254-a64c-4200-8c14-fc01a2b8f709 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Brown University, 2000
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5eb5c881-4c66-4058-9e95-398f0fe18c18 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Coherent-to-diffuse power ratio estimation for dere- verberation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 23(6):1006– 1018, 2015
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91ed6964-f625-4b2b-8f82-fc619039816a · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing A density-based algorithm for discovering clusters in large spatial databases with noise
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02188b04-effb-453d-b0b7-83a0edf3fc75 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing A new approach to linear filtering and prediction problems.Transactions of the ASME–Journal of Basic Engineering, 82(Series D):35–45, 1960
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8411998-4f37-4049-accb-4eef08f76a8e · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing The pascal visual object classes (voc) challenge.International Journal of Computer Vision, 88:303–338, 06 2010
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa89350e-5ea6-4487-b41b-6fccf2ed5fd6 · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d13a14a-8e9a-4e43-adc4-bf1ea64118bd · outbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ccdd4417-f720-4d73-a485-7bcd61c5b45d · inbound
EgoSound: Benchmarking Sound Understanding in Egocentric Videos SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 162511d4-657a-4d71-8e94-51887078ef76 · inbound
Do Joint Audio-Video Generation Models Understand Physics? SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1105d00a-a7c5-41f2-b56c-abe588f4d4b2 · inbound
Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.