Pith. sign in

Paper Citation Record · LEDGER

SceneBind: Binding What and Where Across Vision, Audio and Language

As of 8 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2607.15265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15265 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T23:45:16.890460Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 087e9a15-00a5-4ddc-839e-9e842e1ba234 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

SceneBind: Binding What and Where Across Vision, Audio and Language Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:09.832430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:09.832430Z digest=sha256:541a1915c3b33e746f3bb9c7ded5aa50d51b2ee3c3ede388262cfbdf60a68892

Observation fa155ee4-435a-4479-b5ee-1b76757b423b · outbound

This paper cites SA VVY: Spatial awareness via audio-visual LLMs through seeing and hearing.

SceneBind: Binding What and Where Across Vision, Audio and Language SA VVY: Spatial awareness via audio-visual LLMs through seeing and hearing

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:09.937044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:09.937044Z digest=sha256:2692fc02498216b982f8218c5e21aca56fe049694ba4abcadbec6cd28c8e8baf

Observation 9fa58045-f767-4434-8878-e1ed7b319cee · outbound

This paper cites Learning transferable visual models from natural language supervision.

SceneBind: Binding What and Where Across Vision, Audio and Language Learning transferable visual models from natural language supervision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:10.049426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:10.049426Z digest=sha256:bb1e43c189da8859590f0de040d6186caf6bda0a2a9c1b4667d8fbe6611a2dbd

Observation 359e75a2-fb47-48a7-8558-ba205e3e8922 · outbound

This paper cites Sigmoid loss for language image pre-training.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11941–11952, 2023.

SceneBind: Binding What and Where Across Vision, Audio and Language Sigmoid loss for language image pre-training.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11941–11952, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:10.179811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:10.179811Z digest=sha256:7d81ea9c380d850c656a597121e46795587144baf130fa14460546e6401001e5

Observation ba4ffdb9-582b-4d7b-8197-3c05cea8f469 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

SceneBind: Binding What and Where Across Vision, Audio and Language SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:10.278101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:10.278101Z digest=sha256:63bdb32fbe4cbc59e3f0c56d4cb990be133e92e6f200e99f2155124b521f8c09

Observation 84882c8d-fd77-4feb-b39f-f6a4487ce965 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation.

SceneBind: Binding What and Where Across Vision, Audio and Language Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:10.394732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:10.394732Z digest=sha256:52abcd866e4c8dec1badb5138ab5a5011eb6b6784dbf19487c0fc2e00ba054eb

Observation c9be0394-ee91-4bde-aa42-f2d4c75dd318 · outbound

This paper cites M2d-clap: Masked modeling duo meets clap for learning general-purpose audio-language representation, 2024.

SceneBind: Binding What and Where Across Vision, Audio and Language M2d-clap: Masked modeling duo meets clap for learning general-purpose audio-language representation, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:10.532974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:10.532974Z digest=sha256:964cfd2798314ba50e592fd5945c0d2887d8fe56731413564101a2986c70beaa

Observation 6afbe198-dbec-4a86-bf96-0a76ba2dd749 · outbound

This paper cites Imagebind: One embedding space to bind them all.

SceneBind: Binding What and Where Across Vision, Audio and Language Imagebind: One embedding space to bind them all

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:10.614491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:10.614491Z digest=sha256:26dcb9e3243c4727438c48ab8b595d1528528ea71777c183e7063393f723e2f4

Observation 0a251de8-209c-4b92-a886-28c4e134ea44 · outbound

This paper cites Soundspaces 2.0: A simulation platform for visual-acoustic learning.Advances in Neural Information Processing Systems, 35:8896– 8911, 2022.

SceneBind: Binding What and Where Across Vision, Audio and Language Soundspaces 2.0: A simulation platform for visual-acoustic learning.Advances in Neural Information Processing Systems, 35:8896– 8911, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:10.782196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:10.782196Z digest=sha256:c75cb0cf33c97a0d36385781c9c7cb4ff4882ed46408a4c79205db30f868551d

Observation 37ca474f-e31f-4530-9065-3826b6a9080e · outbound

This paper cites Spatial-clap: Learning spatially-aware audio–text embeddings for multi-source conditions.arXiv preprint arXiv:2509.14785, 2025.

SceneBind: Binding What and Where Across Vision, Audio and Language Spatial-clap: Learning spatially-aware audio–text embeddings for multi-source conditions.arXiv preprint arXiv:2509.14785, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:10.929539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:10.929539Z digest=sha256:e1e00570ba4714e4df8077c24bf8f4292307d9f7143e7f2ab4b2effcec8f402d

Observation cefcb73f-99ff-412c-8cc5-c91d2f27ccf3 · outbound

This paper cites Audioclip: Extending clip to image, text and audio, 2021.

SceneBind: Binding What and Where Across Vision, Audio and Language Audioclip: Extending clip to image, text and audio, 2021

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:11.050820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:11.050820Z digest=sha256:eb019442c0da608d602038bb048f385e62e8cc73cb3b6d7ef78c18b633d412ce

Observation 7a135e44-0c66-4de4-abfc-b2c8d919dca7 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

SceneBind: Binding What and Where Across Vision, Audio and Language Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:11.136712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:11.136712Z digest=sha256:505272342839b0d77b8927a78a631dd71f3d45d89e3e968547e60755cda38b0f

Observation d5a72fea-267f-4099-a288-860f71dff475 · outbound

This paper cites Onellm: One framework to align all modalities with language.

SceneBind: Binding What and Where Across Vision, Audio and Language Onellm: One framework to align all modalities with language

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:11.276028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:11.276028Z digest=sha256:6f0a0b41a76b81e72309a52fca9ef3bb41a297b1a080f1642ab4c00370824aa8

Observation 90422da9-de6d-4bea-bbd1-dc5e5ee793ce · outbound

This paper cites X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent cross-modal reasoning, 2023.

SceneBind: Binding What and Where Across Vision, Audio and Language X-instructblip: A framework for aligning x-modal instruction-aware representations to llms and emergent cross-modal reasoning, 2023

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:11.437574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:11.437574Z digest=sha256:1826f34ed663106139d5682d18f3850cfd575bab58a72114e34b84d4a003c8d4

Observation 5f574047-85e9-44d2-84f0-254ef50bc28f · outbound

This paper cites Robohop: Segment-based topological map representation for open-world visual navigation.

SceneBind: Binding What and Where Across Vision, Audio and Language Robohop: Segment-based topological map representation for open-world visual navigation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:11.563238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:11.563238Z digest=sha256:8cea6d606e949214219e3bd0265c7394e8fcea597d49c262b194bb78d3ccc1d6

Observation 96adb037-51b2-492c-9bda-a63cfc9a6986 · outbound

This paper cites Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba.

SceneBind: Binding What and Where Across Vision, Audio and Language Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:11.725079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:11.725079Z digest=sha256:0f8bcabde4dcbcd9443ba6f1dda4f6ff8dbd8e3deb8b827327d435ba44bbd112

Observation e45cf3d6-2fb8-4b09-a2c2-0aa05a72d06f · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning.

SceneBind: Binding What and Where Across Vision, Audio and Language Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:11.860092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:11.860092Z digest=sha256:7401d2b3c0e8c10d91121e8bb5a7d9f41a1d5b30f06bf0df7a663823e9a6263c

Observation 867591e1-9f8b-42ee-8337-0496957f2db6 · outbound

This paper cites 3d-mem: 3d scene memory for embodied exploration and reasoning, 2024.

SceneBind: Binding What and Where Across Vision, Audio and Language 3d-mem: 3d scene memory for embodied exploration and reasoning, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:12.002085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:12.002085Z digest=sha256:2b6207f7c65eab9756797899dced1b46d80e4fe56e7629228348ea78bca1c690

Observation d5d4d01b-c508-4578-999c-518ff00825f7 · outbound

This paper cites Gridmm: Grid memory map for vision-and-language navigation.

SceneBind: Binding What and Where Across Vision, Audio and Language Gridmm: Grid memory map for vision-and-language navigation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:12.133995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:12.133995Z digest=sha256:4e2e884d95eb8e25436abf265a218f4e1904eb8cc5bbb8727c318650a085242e

Observation 18c5015a-b2ec-49a1-97b8-65d6b0aa6f33 · outbound

This paper cites Visual language maps for robot navigation.

SceneBind: Binding What and Where Across Vision, Audio and Language Visual language maps for robot navigation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:12.276313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:12.276313Z digest=sha256:0b0b27581d7b45dff8d99e49044c4b5f2922cc30e39f1572cadaa9c4e8d6604a

Observation 787cd91f-2c69-49fc-86e6-e138bae2a2b4 · outbound

This paper cites ChatSplat: 3D Conversational Gaussian Splatting.

SceneBind: Binding What and Where Across Vision, Audio and Language ChatSplat: 3D Conversational Gaussian Splatting

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:12.402273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:12.402273Z digest=sha256:d431156f911629b5a37d7db8fb4a9b47ffebca4c30028f0a49913101f14613b2

Observation 7a330364-c9bb-485f-927e-3b7752c4b4a4 · outbound

This paper cites Language embedded 3d gaus- sians for open-vocabulary scene understanding.

SceneBind: Binding What and Where Across Vision, Audio and Language Language embedded 3d gaus- sians for open-vocabulary scene understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:12.593488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:12.593488Z digest=sha256:3942fd688977230beb362408feb538f193ea97fdf61fc4eb321002c6c5329daf

Observation 869d5ad7-b0cd-42a4-87b9-95c7485f50b6 · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks, 2024.

SceneBind: Binding What and Where Across Vision, Audio and Language Grounded sam: Assembling open-world models for diverse visual tasks, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:12.735472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:12.735472Z digest=sha256:f08ffe1797ee99c6ef73aab56b0ede9ef8f37b462549e356e0b9ef5c5fa38a2e

Observation ada12380-da4a-4a51-b8a8-ac326d4bbc3c · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

SceneBind: Binding What and Where Across Vision, Audio and Language Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:12.869585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:12.869585Z digest=sha256:d810c61900340d2282c6871f3f0cf77f17c386547bcb817f36b5b834167273bc

Observation cec6af44-1c2d-496f-8afd-5e9fa93481bb · outbound

This paper cites Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection.

SceneBind: Binding What and Where Across Vision, Audio and Language Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:12.978003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:12.978003Z digest=sha256:35b5e31bc8f3ca01238bbbae54874a25947eeb4d34b69b0193b0f9f4f2c8a8e7

Observation a9d4588e-2aea-4595-bcd1-081a310578c6 · outbound

This paper cites End-to-end object detection with transformers.

SceneBind: Binding What and Where Across Vision, Audio and Language End-to-end object detection with transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:13.129125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:13.129125Z digest=sha256:b93738d5b150fb033a4b79a49dfe4002c8ac4120e9850bdcd12f3e9e936532ba

Observation 7735c57b-7f8d-4085-a0fe-1f3999ea89e5 · outbound

This paper cites Deformable DETR: deformable transformers for end-to-end object detection.

SceneBind: Binding What and Where Across Vision, Audio and Language Deformable DETR: deformable transformers for end-to-end object detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:13.230943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:13.230943Z digest=sha256:ceacf516db8509acacb57e450323b3ce5295c27c700e7775b51d22a57d474ee7

Observation 5c277ad5-efd2-441b-837d-408a9db48d12 · outbound

This paper cites Visual scene graphs for audio source separation.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 1184–1193, 2021.

SceneBind: Binding What and Where Across Vision, Audio and Language Visual scene graphs for audio source separation.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 1184–1193, 2021

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:13.355094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:13.355094Z digest=sha256:024597c0fe383ca199fe405731b0eda70860491d7fa6bb0feb79c900cd99b8c4

Observation 6931c8d2-7ac8-4d2a-869b-c1eddec8864a · outbound

This paper cites Learning audio-visual dynamics using scene graphs for audio source separation.

SceneBind: Binding What and Where Across Vision, Audio and Language Learning audio-visual dynamics using scene graphs for audio source separation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:13.471187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:13.471187Z digest=sha256:dc91f8f3651a3f0c61c78dbed46a1a88686f2fac54a7ccd40e0862ab25bd9dd7

Observation eaefe87c-0218-4f4e-8df5-f278b90afa59 · outbound

This paper cites Audio-visual grouping network for sound localization from mixtures.

SceneBind: Binding What and Where Across Vision, Audio and Language Audio-visual grouping network for sound localization from mixtures

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:13.619588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:13.619588Z digest=sha256:b709ffd753fb944c6a0f8e2b45c303fe62b46941f527b005b891bedbab6c246e

Observation 42abd712-08d6-4308-a2d0-37dcafdc1f65 · outbound

This paper cites Lavss: Location-guided audio-visual spatial audio separation.

SceneBind: Binding What and Where Across Vision, Audio and Language Lavss: Location-guided audio-visual spatial audio separation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:13.774025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:13.774025Z digest=sha256:52e1e880417565064618c547521ccfaa2918c1fc4bc25f786d4b2379631a4cdc

Observation d06e21e3-55e7-4a05-a746-eda9ef93222e · outbound

This paper cites Audio-visual scene analysis with self-supervised multisen- sory features.

SceneBind: Binding What and Where Across Vision, Audio and Language Audio-visual scene analysis with self-supervised multisen- sory features

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:13.876437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:13.876437Z digest=sha256:ce0c318b276b41acc52d710ad279aa65f0a3ad71c6548f8a9f9db895783278eb

Observation bbb297ea-d6c8-4832-bbea-d2450a83228b · outbound

This paper cites Discriminative sounding objects localization via self-supervised audiovisual matching.

SceneBind: Binding What and Where Across Vision, Audio and Language Discriminative sounding objects localization via self-supervised audiovisual matching

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:13.994141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:13.994141Z digest=sha256:1e830fdaa9cd628468c2251a922206636d4adfe793e64579acf24f9c72b7bcb4

Observation 4a079e94-ac57-4bac-b5aa-ee37f30e3cca · outbound

This paper cites Cyclic co-learning of sounding object visual grounding and sound separation.2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2744–2753, 2021.

SceneBind: Binding What and Where Across Vision, Audio and Language Cyclic co-learning of sounding object visual grounding and sound separation.2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2744–2753, 2021

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:14.119034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:14.119034Z digest=sha256:1e1f18cf50289eaf4a75bf753d389f208a8e64743d964c8d82596e757c7658b7

Observation 6375b6c3-6475-4cd7-b519-374fc9abb232 · outbound

This paper cites Sound event localization and detection of overlapping sources using convolutional recurrent neural networks.

SceneBind: Binding What and Where Across Vision, Audio and Language Sound event localization and detection of overlapping sources using convolutional recurrent neural networks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:14.234913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:14.234913Z digest=sha256:330fd41ef11dd31a563ccac9e11cb8ec29b94dde7582a994e0c6ddf9a95b4fb2

Observation a6ae3561-fd26-4391-96ca-f791e2cea1c7 · outbound

This paper cites Robust sound source tracking using srp-phat and 3d convolutional neural networks.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:300–311, 2020.

SceneBind: Binding What and Where Across Vision, Audio and Language Robust sound source tracking using srp-phat and 3d convolutional neural networks.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:300–311, 2020

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:14.359911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:14.359911Z digest=sha256:5c85252004d798fde876909d669e28415b8a2b04b1eb322b7774ff2e43de96c8

Observation 9efd7d7b-ca93-45ea-899b-a67e4b12152c · outbound

This paper cites Bat: Learning to reason about spatial sounds with large language models.International conference on machine learning, 2024.

SceneBind: Binding What and Where Across Vision, Audio and Language Bat: Learning to reason about spatial sounds with large language models.International conference on machine learning, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:14.478774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:14.478774Z digest=sha256:badd379739df8b9e0d71f611a349a91626a581325cdd5f98da55eb333168d523

Observation 03cd5d5f-eb1f-4398-8f33-4fa02f26aef6 · outbound

This paper cites Learn- ing spatially-aware language and audio embeddings.

SceneBind: Binding What and Where Across Vision, Audio and Language Learn- ing spatially-aware language and audio embeddings

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:14.603604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:14.603604Z digest=sha256:a0b3414fa2d2d10130b8cdde86899843c2fb5f6c795a466fdb266d5c8904a428

Observation 33a8d984-0ce8-4804-8617-7277ee0e7e9d · outbound

This paper cites Phasecoder: Microphone geometry-agnostic spatial audio understanding for multimodal llms, 2026.

SceneBind: Binding What and Where Across Vision, Audio and Language Phasecoder: Microphone geometry-agnostic spatial audio understanding for multimodal llms, 2026

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:14.756020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:14.756020Z digest=sha256:4ea1f34e98845229efa71d01f60c225a56093314aba8f138c467616b67b46749

Observation 6ede0104-3ccb-4c0a-aa2e-0a61176afc18 · outbound

This paper cites Hear you are: Teaching llms spatial reasoning with vision and spatial sound.

SceneBind: Binding What and Where Across Vision, Audio and Language Hear you are: Teaching llms spatial reasoning with vision and spatial sound

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:14.865953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:14.865953Z digest=sha256:5a3f6034b827ae39533bb4d24707289247a67e5c454c30dc37790228efec8d91

Observation 6cce96e4-6ae2-4f9a-bc45-40baf052a953 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

SceneBind: Binding What and Where Across Vision, Audio and Language Representation Learning with Contrastive Predictive Coding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:14.924921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:14.924921Z digest=sha256:9bcb769f29388fdd81fe9c1806528c9288380859911f4d086421bef882a0e599

Observation 77dc329a-2f41-42a6-ad31-1e9efd935d37 · outbound

This paper cites an unresolved cited work.

SceneBind: Binding What and Where Across Vision, Audio and Language Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:14.994531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:14.994531Z digest=sha256:3d9dae44faf8a3367053eedcf8586a89e4571dd64a91e499f39cf1fa104895c8

Observation 53c8a13d-4907-4ed5-a2bb-fc325b4189c6 · outbound

This paper cites Audiocaps: Generat- ing captions for audios in the wild.

SceneBind: Binding What and Where Across Vision, Audio and Language Audiocaps: Generat- ing captions for audios in the wild

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:15.088113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:15.088113Z digest=sha256:447c4b22eded41711373b1eb772a37d4e605aab31a9a45284a3c76dc6af54b3a

Observation a1ea4721-43c5-4328-acb0-40570d21e1bf · outbound

This paper cites Microsoft coco: Common objects in context.

SceneBind: Binding What and Where Across Vision, Audio and Language Microsoft coco: Common objects in context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:15.208143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:15.208143Z digest=sha256:5c266011833cf887efeaa78d627db2e2d4c08d4f3158a8d74dbf54d5152b14dc

Observation 603ca818-19a3-4c3d-a5e5-95370ad2288a · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

SceneBind: Binding What and Where Across Vision, Audio and Language Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:15.291684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:15.291684Z digest=sha256:053b0530aff29fa46c9259e75f14245c627dec09371504bc242d608d2b98fafc

Observation bce4cf90-12d6-4bba-b2e4-185c3480e0b3 · outbound

This paper cites OmniAudio: Generating Spatial Audio from 360-Degree Video.

SceneBind: Binding What and Where Across Vision, Audio and Language OmniAudio: Generating Spatial Audio from 360-Degree Video

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:15.384571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:15.384571Z digest=sha256:636f47439ec8b54235e95f8d49c7b328f392f1bcb2c54f258508297c43eed7d8

Observation 353e9768-a41e-4bb2-9a0f-a4678a78b367 · outbound

This paper cites Egocentric audio-visual object localization.

SceneBind: Binding What and Where Across Vision, Audio and Language Egocentric audio-visual object localization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:15.502818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:15.502818Z digest=sha256:ee7be47e97ad0a250ba166cc10f3c9bdfca95013e226ddf3b153da46be840584

Observation 383336a3-3d67-4805-9280-e08d372edd49 · outbound

This paper cites Learning to localize sound source in visual scenes.

SceneBind: Binding What and Where Across Vision, Audio and Language Learning to localize sound source in visual scenes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:15.653444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:15.653444Z digest=sha256:8469fc84a32e1e03ad7375f4890b0d2a4638102fa7743f44c0573f654e561d13

Observation 3d48dc4b-3740-4396-a076-f6b30842fbb9 · outbound

This paper cites Space-Time Memory Network for Sounding Object Localization in Videos.

SceneBind: Binding What and Where Across Vision, Audio and Language Space-Time Memory Network for Sounding Object Localization in Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:15.752844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:15.752844Z digest=sha256:0af21298685e705f648639bebb3c9b50a46cfc84fbb106d6fc1d22faebd64238

Observation 95c9fc40-ea70-4699-b065-03355a706977 · outbound

This paper cites Localizing visual sounds the hard way.

SceneBind: Binding What and Where Across Vision, Audio and Language Localizing visual sounds the hard way

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:15.887730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:15.887730Z digest=sha256:3b10f04d90268aedd7f688aba3286a6ac25c150ae1f0d0217d3f695d15389157

Observation b9fa7dce-9b23-457d-b5ef-961ed5ad3755 · outbound

This paper cites Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes.

SceneBind: Binding What and Where Across Vision, Audio and Language Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:16.031689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:16.031689Z digest=sha256:a32281a362f1454591b4647c3ff4a7a42133081873cefbfd2e0a6e7085596b31

Observation d66b6d27-f23d-43b8-864a-167f465af2da · outbound

This paper cites Mix and localize: Localizing sound sources in mixtures.

SceneBind: Binding What and Where Across Vision, Audio and Language Mix and localize: Localizing sound sources in mixtures

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:16.142632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:16.142632Z digest=sha256:8468fca70e6e2284845f6ebd69d46b4ac90bd09c9ef4da0324ca3fbfd69f9647

Observation 232d415c-9c5b-4bd9-8fb3-b16ac595c070 · outbound

This paper cites Seeing speech and sound: Distinguishing and locating audio sources in visual scenes.

SceneBind: Binding What and Where Across Vision, Audio and Language Seeing speech and sound: Distinguishing and locating audio sources in visual scenes

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:16.265301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:16.265301Z digest=sha256:f4b17bf5ac6689a9c28887f03e220bc3bbe56b9546b74c2ec8c285f4bdf001c6

Observation b12a34dd-9798-4d1c-9803-70969eea25d0 · outbound

This paper cites Cnn architectures for large-scale audio classification.

SceneBind: Binding What and Where Across Vision, Audio and Language Cnn architectures for large-scale audio classification

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:16.368315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:16.368315Z digest=sha256:c046ad9293ac9657aa220791a6423e92d05f82a1f1fe6c2134d926f8df0da214

Observation 13992702-ac53-4d69-bdf4-6bda3fd4bb2e · outbound

This paper cites Hrtf measurements of a kemar dummy-head microphone.

SceneBind: Binding What and Where Across Vision, Audio and Language Hrtf measurements of a kemar dummy-head microphone

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:16.485987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:16.485987Z digest=sha256:96cdc7cce3d8baf4e1d132765e5e28c7d417de29ac3f571843e55f3daa110c27

Observation 512491a8-737e-40c7-8d43-bf1cc8de3d4b · outbound

This paper cites - visual_only: if visible, but it is silent or not synchronized with any sound.

SceneBind: Binding What and Where Across Vision, Audio and Language - visual_only: if visible, but it is silent or not synchronized with any sound

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:16.596325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:16.596325Z digest=sha256:a1554ecddab25951d06278a02af2df9a4ad830c3f22d9185468755a320ebbaad

Observation bfe70c16-5a9d-47bf-9cde-368b0eb8ba21 · outbound

This paper cites - Duration Constraint: Events must be short atomic instances.

SceneBind: Binding What and Where Across Vision, Audio and Language - Duration Constraint: Events must be short atomic instances

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:16.701974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:16.701974Z digest=sha256:e9625bc43554895b30d94466147f25bb6036d514a8508083bdd13a1aba350110

Observation 7de46d3a-8f6d-4e9f-b90d-23ff514e4446 · outbound

This paper cites an unresolved cited work.

SceneBind: Binding What and Where Across Vision, Audio and Language Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:16.815500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:16.815500Z digest=sha256:7f8f18ac311964193b071478a2dc580308624c67226b63529760bcb298f3166f

Observation 18c585ce-c824-4bd8-bd52-92b5e851d144 · outbound

This paper cites - semantic_anno (6 to 8 words): Describe WHAT the object is doing/being.

SceneBind: Binding What and Where Across Vision, Audio and Language - semantic_anno (6 to 8 words): Describe WHAT the object is doing/being

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T23:45:16.890460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:45:16.890460Z digest=sha256:9e9f7690db3444dfa33a90ac10097831181107f5bbf2e543d63b167da78962da

Pith citing papers

No inbound Pith citation observations are available.