Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T23:45:16.890460Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2607.15265.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T23:45:16.890460Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 087e9a15-00a5-4ddc-839e-9e842e1ba234 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa155ee4-435a-4479-b5ee-1b76757b423b · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language SA VVY: Spatial awareness via audio-visual LLMs through seeing and hearing
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fa58045-f767-4434-8878-e1ed7b319cee · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Learning transferable visual models from natural language supervision
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 359e75a2-fb47-48a7-8558-ba205e3e8922 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Sigmoid loss for language image pre-training.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11941–11952, 2023
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba4ffdb9-582b-4d7b-8197-3c05cea8f469 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84882c8d-fd77-4feb-b39f-f6a4487ce965 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9be0394-ee91-4bde-aa42-f2d4c75dd318 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language M2d-clap: Masked modeling duo meets clap for learning general-purpose audio-language representation, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6afbe198-dbec-4a86-bf96-0a76ba2dd749 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Imagebind: One embedding space to bind them all
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a251de8-209c-4b92-a886-28c4e134ea44 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Soundspaces 2.0: A simulation platform for visual-acoustic learning.Advances in Neural Information Processing Systems, 35:8896– 8911, 2022
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ca474f-e31f-4530-9065-3826b6a9080e · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Spatial-clap: Learning spatially-aware audio–text embeddings for multi-source conditions.arXiv preprint arXiv:2509.14785, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cefcb73f-99ff-412c-8cc5-c91d2f27ccf3 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Audioclip: Extending clip to image, text and audio, 2021
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a135e44-0c66-4de4-abfc-b2c8d919dca7 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5a72fea-267f-4099-a288-860f71dff475 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Onellm: One framework to align all modalities with language
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90422da9-de6d-4bea-bbd1-dc5e5ee793ce · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent cross-modal reasoning, 2023
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f574047-85e9-44d2-84f0-254ef50bc28f · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Robohop: Segment-based topological map representation for open-world visual navigation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96adb037-51b2-492c-9bda-a63cfc9a6986 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e45cf3d6-2fb8-4b09-a2c2-0aa05a72d06f · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 867591e1-9f8b-42ee-8337-0496957f2db6 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language 3d-mem: 3d scene memory for embodied exploration and reasoning, 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d4d01b-c508-4578-999c-518ff00825f7 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Gridmm: Grid memory map for vision-and-language navigation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18c5015a-b2ec-49a1-97b8-65d6b0aa6f33 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Visual language maps for robot navigation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 787cd91f-2c69-49fc-86e6-e138bae2a2b4 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language ChatSplat: 3D Conversational Gaussian Splatting
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a330364-c9bb-485f-927e-3b7752c4b4a4 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Language embedded 3d gaus- sians for open-vocabulary scene understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 869d5ad7-b0cd-42a4-87b9-95c7485f50b6 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Grounded sam: Assembling open-world models for diverse visual tasks, 2024
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ada12380-da4a-4a51-b8a8-ac326d4bbc3c · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cec6af44-1c2d-496f-8afd-5e9fa93481bb · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9d4588e-2aea-4595-bcd1-081a310578c6 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language End-to-end object detection with transformers
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7735c57b-7f8d-4085-a0fe-1f3999ea89e5 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Deformable DETR: deformable transformers for end-to-end object detection
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c277ad5-efd2-441b-837d-408a9db48d12 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Visual scene graphs for audio source separation.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 1184–1193, 2021
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6931c8d2-7ac8-4d2a-869b-c1eddec8864a · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Learning audio-visual dynamics using scene graphs for audio source separation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaefe87c-0218-4f4e-8df5-f278b90afa59 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Audio-visual grouping network for sound localization from mixtures
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42abd712-08d6-4308-a2d0-37dcafdc1f65 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Lavss: Location-guided audio-visual spatial audio separation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d06e21e3-55e7-4a05-a746-eda9ef93222e · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Audio-visual scene analysis with self-supervised multisen- sory features
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb297ea-d6c8-4832-bbea-d2450a83228b · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Discriminative sounding objects localization via self-supervised audiovisual matching
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a079e94-ac57-4bac-b5aa-ee37f30e3cca · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Cyclic co-learning of sounding object visual grounding and sound separation.2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2744–2753, 2021
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6375b6c3-6475-4cd7-b519-374fc9abb232 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Sound event localization and detection of overlapping sources using convolutional recurrent neural networks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6ae3561-fd26-4391-96ca-f791e2cea1c7 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Robust sound source tracking using srp-phat and 3d convolutional neural networks.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:300–311, 2020
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9efd7d7b-ca93-45ea-899b-a67e4b12152c · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Bat: Learning to reason about spatial sounds with large language models.International conference on machine learning, 2024
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03cd5d5f-eb1f-4398-8f33-4fa02f26aef6 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Learn- ing spatially-aware language and audio embeddings
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33a8d984-0ce8-4804-8617-7277ee0e7e9d · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Phasecoder: Microphone geometry-agnostic spatial audio understanding for multimodal llms, 2026
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ede0104-3ccb-4c0a-aa2e-0a61176afc18 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Hear you are: Teaching llms spatial reasoning with vision and spatial sound
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cce96e4-6ae2-4f9a-bc45-40baf052a953 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Representation Learning with Contrastive Predictive Coding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77dc329a-2f41-42a6-ad31-1e9efd935d37 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c8a13d-4907-4ed5-a2bb-fc325b4189c6 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Audiocaps: Generat- ing captions for audios in the wild
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ea4721-43c5-4328-acb0-40570d21e1bf · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Microsoft coco: Common objects in context
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 603ca818-19a3-4c3d-a5e5-95370ad2288a · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce4cf90-12d6-4bba-b2e4-185c3480e0b3 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language OmniAudio: Generating Spatial Audio from 360-Degree Video
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 353e9768-a41e-4bb2-9a0f-a4678a78b367 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Egocentric audio-visual object localization
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 383336a3-3d67-4805-9280-e08d372edd49 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Learning to localize sound source in visual scenes
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d48dc4b-3740-4396-a076-f6b30842fbb9 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Space-Time Memory Network for Sounding Object Localization in Videos
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95c9fc40-ea70-4699-b065-03355a706977 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Localizing visual sounds the hard way
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9fa7dce-9b23-457d-b5ef-961ed5ad3755 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d66b6d27-f23d-43b8-864a-167f465af2da · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Mix and localize: Localizing sound sources in mixtures
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 232d415c-9c5b-4bd9-8fb3-b16ac595c070 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Seeing speech and sound: Distinguishing and locating audio sources in visual scenes
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b12a34dd-9798-4d1c-9803-70969eea25d0 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Cnn architectures for large-scale audio classification
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13992702-ac53-4d69-bdf4-6bda3fd4bb2e · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Hrtf measurements of a kemar dummy-head microphone
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 512491a8-737e-40c7-8d43-bf1cc8de3d4b · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language - visual_only: if visible, but it is silent or not synchronized with any sound
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfe70c16-5a9d-47bf-9cde-368b0eb8ba21 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language - Duration Constraint: Events must be short atomic instances
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7de46d3a-8f6d-4e9f-b90d-23ff514e4446 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18c585ce-c824-4bd8-bd52-92b5e851d144 · outbound
SceneBind: Binding What and Where Across Vision, Audio and Language - semantic_anno (6 to 8 words): Describe WHAT the object is doing/being
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.