Pith. sign in

Paper Citation Record · LEDGER

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

As of 7 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2506.06537.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06537 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:58:02.357429Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:58:02.208194Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:58:02.433864Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c324de4-f448-45d4-9973-c134069f57ec · outbound

This paper cites This capability is essential for applica- tions in video understanding, surveillance, human-computer in- teraction, and autonomous systems.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models This capability is essential for applica- tions in video understanding, surveillance, human-computer in- teraction, and autonomous systems

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.137883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.197732Z digest=sha256:c8ad40d9d7d5e8717d1c55b5b3630ad5c3975270bbfc926ad04f2e97c392588d

Observation ba2b96cc-0daf-4343-abb5-a88b9eeb2afc · outbound

This paper cites Many approaches leverage cross-modal attention [1, 2, 3] combined with contrastive learning [4, 5] to align audio and visual features.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Many approaches leverage cross-modal attention [1, 2, 3] combined with contrastive learning [4, 5] to align audio and visual features

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.124094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.203606Z digest=sha256:ad2cd760b46832ae48856b434437041444e947fa51782ccee25c9436c2eafd90

Observation 5b467b46-f159-4823-b8ab-669a8d04bc28 · outbound

This paper cites Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:58:02.440658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.208194Z digest=sha256:a3308b1cb67b3b1100bd38428c409c5b5474d66bd36b65b6013aaade071a326f

Observation 05360814-6668-4846-b41d-7cdc1d1a6efe · outbound

This paper cites Models Below, we elaborate on the final models built for each approach.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Models Below, we elaborate on the final models built for each approach

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.110256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.213087Z digest=sha256:6f6509ada5d0e84a9bd4829efd504502be7a0667515feef2796365bc6256609e

Observation 7266a948-ce5a-4a2f-97ae-f12a5361c794 · outbound

This paper cites By systematically integrating audio, vision, and text representations, we bridge modality gaps and achieve fine-grained segmentation without requiring A VS-specific annotations.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models By systematically integrating audio, vision, and text representations, we bridge modality gaps and achieve fine-grained segmentation without requiring A VS-specific annotations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.096865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.218012Z digest=sha256:18eda2d209a8d1f438c810593e5dafce1d0d0384445b2549be0c55be1aea9efa

Observation 45b681cf-81c4-47a9-8c43-0945a0339f6f · outbound

This paper cites an unresolved cited work.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:58:03.083300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.222340Z digest=sha256:180a418b7dfe3fb1ad9ff23e70669e67a2b718840d5b602747bf08e132e66db5

Observation 51b6b9a9-a56e-40c4-a262-d9929e24d7fe · outbound

This paper cites Learning to localize sound sources in visual scenes: Analysis and applications,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning to localize sound sources in visual scenes: Analysis and applications,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.069950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.227428Z digest=sha256:9775fcc07f01c8d85d568c47c02b0cb8b84425cdc73207035f2055ec109ba68d

Observation dbc8d8e5-6e64-4c80-844f-d0602e249a49 · outbound

This paper cites Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.231835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.231835Z digest=sha256:640b096b3de41d203ddfc6f3527b61f0608cafd1fc2b994cbdd7be938840bf4f

Observation ea2187a2-3ab3-4113-9ef0-27d3b535398a · outbound

This paper cites Exploiting transformation in- variance and equivariance for self-supervised sound localisation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Exploiting transformation in- variance and equivariance for self-supervised sound localisation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.057441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.236552Z digest=sha256:2d36e2abd286e03750ba974373d5e4b778642ac9220da3ba2d06ed1fc586134a

Observation 61908570-0221-42fd-89a2-e50426373aeb · outbound

This paper cites Localizing visual sounds the easy way,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Localizing visual sounds the easy way,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.044572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.241086Z digest=sha256:9c0e0940e35dfa87ab130df5ac460838dee57099ceea60aadf3fd223bfabf5c0

Observation 187b50d7-1b94-47c1-a2b1-a173a5565f20 · outbound

This paper cites Localizing visual sounds the hard way,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Localizing visual sounds the hard way,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.032265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.246202Z digest=sha256:d1a010ca1523a863f058f70944272c2e0f11edbfd7aec8a813ef8f60111d8001

Observation 8330852f-ff89-4c02-b9ba-acd7124ae7d0 · outbound

This paper cites Learning audio-visual source local- ization via false negative aware contrastive learning,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning audio-visual source local- ization via false negative aware contrastive learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.019857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.250492Z digest=sha256:70bec79fbdc5a37a2c0f5fe4c50bd571daba69d9fe2d67425d120f4226b720c9

Observation fe06a59f-5fdb-47ab-a449-d45ab740654f · outbound

This paper cites Marginnce: Robust sound localization with a negative margin,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Marginnce: Robust sound localization with a negative margin,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.006909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.255202Z digest=sha256:4dc5a980cc3679e2705c264f10185a75b22703e7f6444b902833899661563930

Observation c2d24cae-a9bd-49af-86f2-29dc969e99db · outbound

This paper cites Audio–visual segmen- tation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Audio–visual segmen- tation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.994261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.259409Z digest=sha256:d2f61e28fb0d499887e411d777e34886a4b326855c7b26a7e3d3b4e53a613809

Observation a97a0451-fba9-402e-b8a0-3e358bed884f · outbound

This paper cites Improving audio-visual segmentation with bidirectional generation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Improving audio-visual segmentation with bidirectional generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.981778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.263449Z digest=sha256:9f6b829536fd8b944132d5ee91650c15d61ce5667975af79eb3cfff630da0c05

Observation ea52bc2d-7569-4054-80fd-e8d55ad2b4fb · outbound

This paper cites Selm: Selective mechanism based audio-visual segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Selm: Selective mechanism based audio-visual segmentation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.910728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.267534Z digest=sha256:3d63b4f992fd673ecd09c4c14e3ab26c84c6a60bd41e4a4df749c7c2d3f6e248

Observation 110a07f7-5a14-41a5-a5a0-4d25e7cd41e2 · outbound

This paper cites Bavs: Bootstrapping audio-visual segmentation by integrating foundation knowledge,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bavs: Bootstrapping audio-visual segmentation by integrating foundation knowledge,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.801448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.271953Z digest=sha256:8745cb32b25dd1be3bffda093af340f45977cef2d1be9699fe4c6e91fe574836

Observation b0f3a086-d99a-4b58-add1-057deb845d7d · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning transferable visual models from natural language supervision,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.755960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.276213Z digest=sha256:1dbad1de1490c4bc32e395f302937e4d88817f044f565dba387a32bf262299e1

Observation b97ce92b-dd08-458a-a942-8ba4a2e442af · outbound

This paper cites Natural Language Supervision for General-Purpose Audio Representations.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Natural Language Supervision for General-Purpose Audio Representations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.280500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.280500Z digest=sha256:ec55eaa8621c54a15d63ba56c4a969d242ffcc3091a936ab6e30632a6224105a

Observation b957c4a4-d464-45c3-8a17-463ab759847e · outbound

This paper cites WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language mul- timodal research,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language mul- timodal research,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.714559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.285239Z digest=sha256:e34c03b40fa18302da76e6ce855a968835f6768467971d0faed9308d48895afe

Observation b8a97d62-a16f-4e8d-9b2a-5cfb5cca0335 · outbound

This paper cites AudioCaps: Generating Captions for Audios in The Wild,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models AudioCaps: Generating Captions for Audios in The Wild,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.682158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.289939Z digest=sha256:b2e7e60fa5e6128b347101d815bb42b38675ca182154edf9b8ecb531807c3cd4

Observation 5d9c8752-8a7a-40f4-a793-eda42895a494 · outbound

This paper cites Microsoft coco: Common objects in context,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Microsoft coco: Common objects in context,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.646312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.294346Z digest=sha256:699afe87f51595ca5c167b61ac21847056ff079c2227e0c6c3f45335e5c438d1

Observation c4704ecb-6a7c-4a03-bfd6-782595e87095 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.626754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.298465Z digest=sha256:7dae7d92decfd00dbddf41b3eb737a51a61a853603fe2b5c6130c3f0414f790d

Observation 3cead3a2-851f-41a0-9133-c6a0c9ce0374 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.302544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.302544Z digest=sha256:92a6a2c698795d1512a1754e5f3b45fc41078332b825aed07449ac343c858873

Observation cdcf4a11-9322-438d-b2be-c3d27ffbf35f · outbound

This paper cites Learning to visually localize sound sources from mixtures without prior source knowl- edge,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning to visually localize sound sources from mixtures without prior source knowl- edge,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.613063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.307043Z digest=sha256:fba0154ebf89a900a317e83aadda20d5b7b3f844acc460974f3d5ce157995393

Observation ef3a48c4-44d5-4ca7-923b-66227211d06f · outbound

This paper cites Adaptive selection based referring im- age segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Adaptive selection based referring im- age segmentation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.599272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.310995Z digest=sha256:8522fe84ec36612b197f58c4630cfa5e76c498691ed17e7ae40578a2fecd0661

Observation a1d42340-6c74-4b35-a8e2-918242b51871 · outbound

This paper cites Beats: Audio pre-training with acoustic tokenizers,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Beats: Audio pre-training with acoustic tokenizers,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.585955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.315084Z digest=sha256:515abef165b7754b06944fa63f22c42e8d2553b161dcc58a5ebf6df20ebf8f0b

Observation 3955a45d-4ea2-4ede-af26-ed42f347f5f7 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Audio set: An ontology and human-labeled dataset for audio events,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.319071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.319071Z digest=sha256:a6ce38b9f9a51daa186ae6abfee5865876438c8041a1928aa2ea49dadbc9be6d

Observation 7e948701-6ee2-4222-99c9-14d9ced8b76b · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Clotho: An audio cap- tioning dataset,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.563661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.323482Z digest=sha256:0e853623e5c6093cdd86d14c65f9d43e38f383682f400b0bf6ffb4bee6fe4a06

Observation 809d8bfc-ead5-43fe-8513-94e268eb35f4 · outbound

This paper cites spaCy 2: Natural language under- standing with Bloom embeddings, convolutional neural networks and incremental parsing,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models spaCy 2: Natural language under- standing with Bloom embeddings, convolutional neural networks and incremental parsing,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.550108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.327698Z digest=sha256:5869ae76f769ac31fe56d9bf87277ad95ea359c6ac9a739221eb14af734e64bc

Observation 99888b15-3398-4842-943c-7ca0b879f59c · outbound

This paper cites High-resolution image synthesis with latent diffusion models,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models High-resolution image synthesis with latent diffusion models,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.331753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.331753Z digest=sha256:e6419a26792787bc2a47e0ad261db9055354a8a0616fdbacd4a9839e7090d8bf

Observation 406847a3-bac8-415e-b23d-bce3dce0f5e7 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Vggsound: A large-scale audio-visual dataset,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.526857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.335799Z digest=sha256:e046391a9d5a19d32178f6bbde9dc986ce078070e35bbe5189f4679c6fff9784

Observation 96a2c20c-d4bb-4263-b611-9c364e4e3414 · outbound

This paper cites Segment anything,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Segment anything,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.512415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.340053Z digest=sha256:88fd619422fa304bd1f5dc38faf8b39918dea56692e2752134dd7dff9e637af2

Observation 3dec38ed-1032-4ad8-8887-aca36f518f87 · outbound

This paper cites Unraveling instance associations: A closer look for audio-visual segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Unraveling instance associations: A closer look for audio-visual segmentation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.497818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.344260Z digest=sha256:cfe81a29d795b04825c5192f6a0bf91af0ca5ac7b271adf7f118709a94847629

Observation 005345c1-4e4b-48e5-b82e-acfb4c9df8b5 · outbound

This paper cites A closer look at weakly-supervised audio-visual source localization,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models A closer look at weakly-supervised audio-visual source localization,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.483950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.348596Z digest=sha256:30ed7cdfefaf657486c239a27caa18043a027824fae6c800b0e9d7427f433e93

Observation 9f67a6e9-214e-4beb-8c99-694d45256cc4 · outbound

This paper cites Bridg- ing vision and language encoders: Parameter-efficient tuning for referring image segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bridg- ing vision and language encoders: Parameter-efficient tuning for referring image segmentation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.469077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.352781Z digest=sha256:2a57515a5d18d66e2bc7cba7136ffe3862129237f3727e91a397ac2935732078

Observation f16500da-de98-4eac-9740-46fc5134d510 · outbound

This paper cites Cris: Clip-driven referring image segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Cris: Clip-driven referring image segmentation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.455112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.357429Z digest=sha256:4f7be70bb94cf5073186d47e4c9f4b562c21369633dc927c25ff6964d9000839

Pith citing papers

Observation 5b467b46-f159-4823-b8ab-669a8d04bc28 · inbound

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models cites this paper.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:58:02.440658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:58:02.208194Z digest=sha256:a3308b1cb67b3b1100bd38428c409c5b5474d66bd36b65b6013aaade071a326f