Pith. sign in

Paper Citation Record · LEDGER

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models

As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2607.24786.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.24786 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:33:04.329045Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 835c4629-c3d2-4f32-84ab-bee119b7687b · outbound

This paper cites Localizing visual sounds the easy way.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Localizing visual sounds the easy way

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.210067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.210067Z digest=sha256:fcc9fe665d37c46d439d6586a82d48d6c6ea291b3122a514247a8b21a70e6ab3

Observation 2f1d469d-193c-4185-8c50-4c281b6261dc · outbound

This paper cites A closer look at weakly-supervised audio-visual source localization.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models A closer look at weakly-supervised audio-visual source localization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.247659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.247659Z digest=sha256:7b4fe891205c5fa8cfc6134cfec33c3b75fac3fe8b8bd4af60292d1f345e7e00

Observation 5058dc0b-c6cb-4b7e-b9df-e479783d343e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Learning transferable visual models from natural language supervision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.303952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.303952Z digest=sha256:6409dcb21c55e62fbe1970e27589e221b4ddfbac5de1190307098b2fc069a314

Observation 962945e0-c51d-4de9-a9f3-08b09fba1d04 · outbound

This paper cites Sigmoid loss for language image pre-training.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Sigmoid loss for language image pre-training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.359248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.359248Z digest=sha256:44e4f1961dd75483bbe4be48cbb86bcc73a35df7a743a72d26f69f88718966ea

Observation 674ee0e8-fa75-49ea-8d8f-564ae12addc2 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Imagebind: One embedding space to bind them all

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.396527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.396527Z digest=sha256:da416a0166588272c38b4fdc123aa86dbe254b3f21abd8956be9cf5603f9ace6

Observation 12b7909e-9f69-40c9-83e3-e2182c7e6571 · outbound

This paper cites Pushing the frontier of audiovisual perception with large-scale multimodal correspondence learning.arXiv preprint arXiv:2512.19687, 2025.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Pushing the frontier of audiovisual perception with large-scale multimodal correspondence learning.arXiv preprint arXiv:2512.19687, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.451568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.451568Z digest=sha256:5e3d52f94df51533dbb356b932d7a678114e15e617161a597eaa7c7c46687eaa

Observation 8b43dc88-62e8-43f7-a76c-02b8f432bbe1 · outbound

This paper cites Learning audio-visual source localization via false negative aware con- trastive learning.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Learning audio-visual source localization via false negative aware con- trastive learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.504059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.504059Z digest=sha256:433bdc6444f7e00706dcdee2f9a16c9af0c2a432e610372d89a4fc0afb0cbe86

Observation 1914a910-90cf-4540-9dec-4e7ed1c94bda · outbound

This paper cites What’s making that sound right now? video-centric audio-visual localization.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models What’s making that sound right now? video-centric audio-visual localization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.571128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.571128Z digest=sha256:0ff4b14b923911a8c1bf311ff4832a91fa0a6a9eb3d396d237eb22469943d28c

Observation 3c334cf1-f7ff-4899-850b-595d1d4f0471 · outbound

This paper cites Flair: Vlm with fine-grained language-informed image representations.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Flair: Vlm with fine-grained language-informed image representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.624637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.624637Z digest=sha256:fa8d4364fa34270874cbc43cecb4b449e0c335df46c15618e8007bde7cb85941

Observation 8d9c5fc8-d090-4aa0-a83f-433f07159929 · outbound

This paper cites Learning to localize sound source in visual scenes.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Learning to localize sound source in visual scenes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.662240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.662240Z digest=sha256:186f087e0d8402e6b3880b7e15289dd7cf41eb7409ca091aa882e80e59a5006e

Observation cc124740-8156-4d87-b69c-7ea9404c3c51 · outbound

This paper cites Exploiting transformation invariance and equivariance for self-supervised sound localisation.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Exploiting transformation invariance and equivariance for self-supervised sound localisation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.729151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.729151Z digest=sha256:79a8342cffd97c6bc5dfaf4ba65ef49861b9e035e5d4139b75d75203883b2f77

Observation 3dd7d787-5a65-4373-a33a-271be860b14e · outbound

This paper cites Marginnce: Robust sound localization with a negative margin.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Marginnce: Robust sound localization with a negative margin

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.822503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.822503Z digest=sha256:ce1c0b355dca86c1785bd053e7b9baa234a445d09835617db2a6a1caf07a00a8

Observation a8606333-4089-4702-b633-3409d51f09fe · outbound

This paper cites Sound source localization is all about cross-modal alignment.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Sound source localization is all about cross-modal alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:02.923850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:02.923850Z digest=sha256:be80e345787550ffef11c5c2cc4db7ac0b11a9e98da7722c25d8d52cd9b64426

Observation 56bc5e8a-24d9-4227-bc06-13b43f5116dd · outbound

This paper cites Audio–visual segmentation.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Audio–visual segmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:03.021479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:03.021479Z digest=sha256:dc4c37de161402bc57001c6bd2ef3085baa231da800fa0c885d6455789d1c41c

Observation 8a061c02-1bc4-455e-9767-df5fd6890f55 · outbound

This paper cites AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:03.127024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:03.127024Z digest=sha256:1a4eefbee2c6fd4c9a32bb482a66b1d1634976bddf8b36dd99d65f4f5491bb89

Observation 70302193-7606-46af-b827-b1f51e739d43 · outbound

This paper cites Audio visual segmentation through text embeddings.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Audio visual segmentation through text embeddings

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:03.235286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:03.235286Z digest=sha256:8525279e86e173735b40575cb0ca66ae2ba98160fedd6e0ff50edfda4b0b13c0

Observation ed72577c-daa9-44b3-9e75-a2373db1cfd4 · outbound

This paper cites Open-vocabulary audio-visual semantic segmentation.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Open-vocabulary audio-visual semantic segmentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:03.328242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:03.328242Z digest=sha256:b31007152f160f315e4d04dcbfbea488b2a6a4a4038daf365a265011810e1f46

Observation 7e8140a3-35d2-4ed4-ae5b-3128f75467c3 · outbound

This paper cites Taco: Training-free sound prompted segmentation via semantically constrained audio-visual co-factorization.Transactions on Machine Learn- ing Research, 2026.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Taco: Training-free sound prompted segmentation via semantically constrained audio-visual co-factorization.Transactions on Machine Learn- ing Research, 2026

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:03.409702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:03.409702Z digest=sha256:a9b3140da309ecff896cef2accf7887ed626f7c271481f42616005e8efdb19d8

Observation 44173815-9d10-4c8f-ae3a-b9b7237766da · outbound

This paper cites Perception encoder: The best visual embeddings are not at the output of the network.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Perception encoder: The best visual embeddings are not at the output of the network

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:03.539332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:03.539332Z digest=sha256:2f29ef1801f528289f5aa4edfef63c78742770b8a002edf72b63695c59cbc11a

Observation dc16366d-6174-4812-a251-ea90bd3ecda2 · outbound

This paper cites Vision transformers need registers.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Vision transformers need registers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:03.647463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:03.647463Z digest=sha256:0ec6555d4fee051ea9c242af02a78c3b6eccb0847decbb956017c1f10f5b3138

Observation 96ae56ee-03eb-4b9d-8084-2d5b29dde6aa · outbound

This paper cites Gradient-based learning applied to document recognition.Proceedings of the IEEE, 86(11):2278–2324, 2002.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Gradient-based learning applied to document recognition.Proceedings of the IEEE, 86(11):2278–2324, 2002

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:03.775098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:03.775098Z digest=sha256:d8388a18d6694f073776f6fae8bcac215e9978377989312826d70e00a5b8f8da

Observation fcc6ce4b-0ca1-4088-a735-14e7aaab3e9e · outbound

This paper cites Parameter-efficient transfer learning for nlp.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Parameter-efficient transfer learning for nlp

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:03.884963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:03.884963Z digest=sha256:4393b0d5e66871985ba63a151cab35577de3953df6b3f319f7691e6de09db9e1

Observation 0902e3d4-21db-44f2-9fc0-7c4032e3bb85 · outbound

This paper cites chirp" from the.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models chirp" from the

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:03.992678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:03.992678Z digest=sha256:55d60aac2ea240464f40eaef1dad0b93daab2418764d871da2e55e6c3278d242

Observation 2faff45e-d3eb-48be-981a-02427295ff7f · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Vggsound: A large-scale audio-visual dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:04.059504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:04.059504Z digest=sha256:7e3eb8041074b89ac07968efa8c65038bc559110701f151eafcafbbc3a6057d7

Observation 7ff947f8-7e51-45f2-818c-0293bc9479c2 · outbound

This paper cites Transfer learning from audio-visual grounding to speech recognition.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Transfer learning from audio-visual grounding to speech recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:04.103719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:04.103719Z digest=sha256:087d0f2650e6d8195bc7e59a1b3a81f5c4b17847bfcb16a62da00d140d7f5ffc

Observation e79ce49c-91aa-4a0b-9ee2-afc02c60a9d6 · outbound

This paper cites Contrastive audio-visual masked autoencoder.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Contrastive audio-visual masked autoencoder

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:04.150073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:04.150073Z digest=sha256:996015de76d14cb2736da7824c24989e5a1d2b22e9a414784331a1461fedb3e4

Observation 6dcc3903-aa82-4f68-8eba-c72ec0734ce1 · outbound

This paper cites Cav-mae sync: Improving contrastive audio-visual mask autoencoders via fine-grained alignment.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Cav-mae sync: Improving contrastive audio-visual mask autoencoders via fine-grained alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:04.186922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:04.186922Z digest=sha256:1d4c33ec8381a793a5b893dbe8a9da2ac8a801c19e9a003031822624b229dd2f

Observation 740fd3d5-3be0-41bc-a799-88fc3db0ef27 · outbound

This paper cites Convolutions die hard: Open- vocabulary segmentation with single frozen convolutional clip.Advances in Neural Information Processing Systems, 36, 2023.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Convolutions die hard: Open- vocabulary segmentation with single frozen convolutional clip.Advances in Neural Information Processing Systems, 36, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:04.220152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:04.220152Z digest=sha256:9370ae710f07f85794196b658fe5f10356ee0d85f4e57757f8494054dd08c6f4

Observation 5844169b-cf35-4fac-99e3-629622c96a78 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models Reproducible scaling laws for contrastive language-image learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:04.274503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:04.274503Z digest=sha256:ea6608869500c1af0b35fd2f831ff8452006a222fdcb38264e3849e67d8a2d93

Observation 5d0a6d57-062f-41af-a1f4-69d217dd6f40 · outbound

This paper cites LAIP two poolers.

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models LAIP two poolers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T10:33:04.329045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:33:04.329045Z digest=sha256:30a15b6fa0dc23d2cafa4b0688ffed3dac533d0a872d49aa06e42cb2e4b3b9be

Pith citing papers

No inbound Pith citation observations are available.