Pith. sign in

Paper Citation Record · LEDGER

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes

As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.01558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01558 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:44:12.200891Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f657d51-1ebb-49b7-a1c1-a31055a06513 · outbound

This paper cites Self-calibrated clip for training-free open-vocabulary segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Self-calibrated clip for training-free open-vocabulary segmentation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.322434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.322434Z digest=sha256:f3db71e8b61f531cb477c31086b946dac26b5f63580a713167e11a837ead4582

Observation 4106bdea-8ebf-42b7-b522-168cf3635064 · outbound

This paper cites Unraveling in- stance associations: A closer look for audio-visual segmenta- tion.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Unraveling in- stance associations: A closer look for audio-visual segmenta- tion

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.348892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.348892Z digest=sha256:e85e00a68c64c999fa4e833d2999e7279757dd0630f61110e0badc5fd82c71fb

Observation 28ddbb4c-1743-43ca-a927-991312d9307f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.369535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.369535Z digest=sha256:31d2977b831722068fb6db518611b461e2e659bd99814592aae6f3eb2f00bf7c

Observation 9dbd80ee-966b-44cf-9bd8-158540a1080f · outbound

This paper cites Avsegformer: Audio-visual segmentation with trans- former.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Avsegformer: Audio-visual segmentation with trans- former

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.736426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.393623Z digest=sha256:1253f1431c4691a18d5e81c8ef7eb56f7b8af58694175040ee9d4d0b4d59570d

Observation 20ec1905-41b2-4500-bd1c-255311445b5a · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio set: An ontology and human- labeled dataset for audio events

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.699455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.422743Z digest=sha256:ebc9c3415e4d3c87c477389c64fd618bdc7ceb1d469ef4be1348940252a9fdf1

Observation 36ef3bfc-767c-4f91-9234-e25c73b26324 · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Cnn archi- tectures for large-scale audio classification

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.667367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.443369Z digest=sha256:0663250bfc753190da5bbab4df3661fd26c59f6f301f9aeaf5f106f01e0bbbc7

Observation 0bd5aab6-3c94-4e93-9d15-3be7573625d0 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.631580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.467929Z digest=sha256:351fc10108c26de299471109f2c24685f8079e9045791e64fe487c1a7e81725e

Observation 112943e8-e328-4570-85bf-0f711cd5c1a1 · outbound

This paper cites Segment any- thing.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Segment any- thing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.487126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.487126Z digest=sha256:62b6888c0e97c78e25a917e3f2a938b44ebf340eda84a6a5ea16d6b11026fd3f

Observation 44aa1544-d441-432f-a99f-8a2339efc465 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Lisa: Reasoning segmentation via large language model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.507306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.507306Z digest=sha256:178ef0947236afc9cfe2dd901a5f4faf5f8dc3b7e657231129694e6d209d18a4

Observation a8e998dc-9bfe-4169-a895-9b7f0e2d6e10 · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Learning to answer questions in dynamic audio-visual scenarios

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.580940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.527309Z digest=sha256:bc1d3bc0b0679ef6f02baade078536402abc331aebe360326ca74d541f92644d

Observation dd7a614f-769c-40c2-a826-1465021b808b · outbound

This paper cites Robust referring video object segmentation with cyclic structural consensus.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Robust referring video object segmentation with cyclic structural consensus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.551397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.551397Z digest=sha256:e46e0e9273e17fe13db3be9646388f20f65e742aaeb29f4f5963b740fa5e23a1

Observation 68ac84f6-97a6-4d9e-9619-a4d26330f4de · outbound

This paper cites Libero: Benchmarking knowl- edge transfer for lifelong robot learning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Libero: Benchmarking knowl- edge transfer for lifelong robot learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.532561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.568454Z digest=sha256:706788d7aa8dae2c57ffb9d9bc584daa0fffd4c5c1bf3af1a489ac5c81511676

Observation b7696e5c-1259-4a90-a136-bc97a26f253b · outbound

This paper cites GRES: Gen- eralized referring expression segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes GRES: Gen- eralized referring expression segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.498870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.584902Z digest=sha256:578a11c195028a1e203e72801870b66127d902217044c5650bb0353c04270545

Observation abedcc9d-ea78-471c-b193-3feea0248855 · outbound

This paper cites Improved baselines with visual instruction tuning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Improved baselines with visual instruction tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.601043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.601043Z digest=sha256:2d77ae418bb7a22511f1db7b80bd5b9b77260e028910a23878f19015286ee8aa

Observation e108d5a6-4276-4322-8943-95722e94cc8f · outbound

This paper cites Visual instruction tuning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.621005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.621005Z digest=sha256:aa76a2d4124b279e3831bf06d7bf0ddab77dd52165144caa2191d1b2c30deb30

Observation a2c1f8d2-72ed-4b08-804b-9a28108aa5a3 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.639745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.639745Z digest=sha256:84ef0e2da8a682e93ca11dedaadda7702cea2c85591fe91c8ab9b343db850786

Observation e1580b59-e6a2-4b27-adb5-7382a9c7bf06 · outbound

This paper cites Open-vocabulary segmentation with semantic-assisted calibration.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Open-vocabulary segmentation with semantic-assisted calibration

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.446323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.660057Z digest=sha256:5f4e265d4fd05988d18b6d33abef14f97c1a1845f689d5407de45de5e5bf4e0a

Observation d3be685d-3edd-4cb2-a54e-cce2bfe266c4 · outbound

This paper cites Universal segmentation at arbi- trary granularity with language instruction.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Universal segmentation at arbi- trary granularity with language instruction

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.399232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.677555Z digest=sha256:740edad43299836670de7432d8f01b7142fac4322fcbd95dfefd289c08b9a6f4

Observation f94e348d-ca8a-47fc-ace8-4688e635702d · outbound

This paper cites Decoupled weight de- cay regularization.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Decoupled weight de- cay regularization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.704295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.704295Z digest=sha256:f3cfef50a3f359a6dc0f354ce3ac2d3e41173b7619382cb72639acee0031a66c

Observation 09e0cbf6-df4a-49e9-94f7-25215f04a837 · outbound

This paper cites Soc: Semantic-assisted object cluster for referring video object segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Soc: Semantic-assisted object cluster for referring video object segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.354324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.721300Z digest=sha256:a5f178c44ddb29cff05c58d0669d072a2d7147aa77d63c40f310c49da5936a4c

Observation 16f510cc-11dc-4c5a-a067-d6c1aef562ba · outbound

This paper cites Mod- eling context between objects for referring expression un- derstanding.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Mod- eling context between objects for referring expression un- derstanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.313045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.738165Z digest=sha256:c9d62c361d5b31ac027032f85b326306db97ae69350490cfd0391de7c0783f20

Observation b645e32b-6c58-4403-94c1-1563de57e8aa · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes SAM 2: Segment Anything in Images and Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.756990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.756990Z digest=sha256:f24f8b2794762522a4696e08922619c43faa06ad184dc2306cb591a42fc2eb9c

Observation 247f7548-6d8e-430a-9382-b699f0dee7a6 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Pixellm: Pixel reasoning with large multimodal model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.269637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.778749Z digest=sha256:e18cf0fab407894053b0c4efe19e8f600b08223567e73e2a820871bfff5586a8

Observation e51c9301-f233-417f-95a4-7afa4d1929a7 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.802462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.802462Z digest=sha256:5b7e8856f23758a7b893285f3eade64a22b584098d260ac9b12396960114744a

Observation 3028eb4f-1c46-4812-92eb-c73560c00a93 · outbound

This paper cites Efficient attention: Attention with lin- ear complexities.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficient attention: Attention with lin- ear complexities

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.225927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.832827Z digest=sha256:915f8cd4854f459e252468a35e939dc67b0ce12a2aececec9e2b95a7f88894fe

Observation 5878c5f0-72c0-41ed-a211-55eac3e2713c · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Roformer: Enhanced transformer with rotary position embedding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.848405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.848405Z digest=sha256:31157033e49a48c5331e054aca6e1409ddeda9c1da8b65a54321b29e3a9f8fc5

Observation bc52a43d-4679-4847-ad56-7721d435448c · outbound

This paper cites Auto- acd: A large-scale dataset for audio-language representation learning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Auto- acd: A large-scale dataset for audio-language representation learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.182438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.871678Z digest=sha256:a94d8297a07142c4c4d01b57ed5920fca99fabcdb3c6cd4337fbc7a72cf1eede

Observation f31b91df-7213-47e0-aca7-bd32ddac1e79 · outbound

This paper cites Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.892307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.892307Z digest=sha256:b5cdb0ba3bcf88e58f6d89497047ba02dbffe0cd88b1bd2d19942f6afcc1a681

Observation 64f558e1-147b-468f-878c-483ff6fa7bbe · outbound

This paper cites Efficient remote sensing transformer for coastline detection with sentinel-2 satellite imagery.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficient remote sensing transformer for coastline detection with sentinel-2 satellite imagery

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.151753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.907122Z digest=sha256:2b8820652c962a730d091c309f1f3ef54953d6273220188c59b4f66b3f08c01f

Observation 58ab5cbc-bd89-45e4-b97f-f47dc87a431e · outbound

This paper cites Prompting segmentation with sound is gen- eralizable audio-visual source localizer.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Prompting segmentation with sound is gen- eralizable audio-visual source localizer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.117916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.924209Z digest=sha256:7e121ca5659bde33d84b48f2f3746d373d0879076340c39119d79d2f9092711f

Observation 2e111730-5219-4b15-8ebf-ad71dd6e9371 · outbound

This paper cites Ref-avs: Refer and segment objects in audio-visual scenes.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Ref-avs: Refer and segment objects in audio-visual scenes

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.060729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.945322Z digest=sha256:a5a98f3d750da7acde4b933bce75b9a5444b7fc9e5c4749c64ea3ae9bfa65383

Observation 493dc4ad-5486-4df8-804b-0238b2ed74bc · outbound

This paper cites Convolution meets trans- former: Efficient hybrid transformer for semantic segmenta- tion with very high resolution imagery.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Convolution meets trans- former: Efficient hybrid transformer for semantic segmenta- tion with very high resolution imagery

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.026308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:11.959188Z digest=sha256:bb61751a8beee1da8b9dac652a282429eb5a6667d10944929039b3c7813e6037

Observation 5cb92975-e83f-4199-a288-d262b10ca69c · outbound

This paper cites IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.980254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.980254Z digest=sha256:a95805b0fb52a9584de03add3817985ec9d9447fc7100af28e4096b59f594447

Observation 4138baf4-d4f3-423a-bb44-b3b7c22d4c7c · outbound

This paper cites Onlinerefer: A simple online baseline for referring video object segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Onlinerefer: A simple online baseline for referring video object segmentation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.997831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.997831Z digest=sha256:8d206c491c715ba83049a9f58eea7d49720d659e9445c32af53f3d689c4df41b

Observation b4aace26-db3c-4949-8b24-73b9f780121b · outbound

This paper cites Language as queries for referring video object seg- mentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language as queries for referring video object seg- mentation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.014816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.014816Z digest=sha256:07813feed949ea88b41d6a9c73fefd59aa6d58f2ebfd5b3e5be965e30d7a9bff

Observation 0882748c-073e-47fa-aa91-48331e4454a4 · outbound

This paper cites Language as queries for referring video object seg- mentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language as queries for referring video object seg- mentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:13.972399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:12.032542Z digest=sha256:18bb22c16989c09b89b09539e8724a30ddbe5081c14e415bd8c4c45ecb2d64e3

Observation 4d8b0770-911d-4830-955e-49fbd5ba5748 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Gsva: Generalized segmentation via multimodal large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:13.935682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:12.045116Z digest=sha256:7b2968452a6977f8177c13e7b05ba11e1aafff607b1e716e4a1da9436034a397

Observation d7575162-9125-4f72-a778-1645e4c6b5a8 · outbound

This paper cites Efficientsam: Leveraged masked image pretraining for efficient segment anything.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficientsam: Leveraged masked image pretraining for efficient segment anything

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.061070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.061070Z digest=sha256:785beb8320446717c83f0b07cef5f7e9e17f9041aaee3c2884ccc3092b5aaf75

Observation 338baf62-16e1-439a-bd25-27965cef8bea · outbound

This paper cites Avqa: A dataset for audio- visual question answering on videos.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Avqa: A dataset for audio- visual question answering on videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:13.878090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:12.078141Z digest=sha256:fb47ce65d0e5295214cd3d726e567dde9e6dd03807b5a52aaa200718df113327

Observation 9556304f-7048-4819-9f84-f36361c06f50 · outbound

This paper cites Lavt: Language-aware vision transformer for referring image segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Lavt: Language-aware vision transformer for referring image segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:12.578490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:12.092128Z digest=sha256:f36ae1eb61f8f8fc51cfcef61126c31ce4fb084ed1938773203250509341eaae

Observation 16a50437-db4c-45f7-9e15-127705cd219f · outbound

This paper cites Language- aware vision transformer for referring segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language- aware vision transformer for referring segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:12.541709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:12.110003Z digest=sha256:449444d96401a1e1ade50a9ef880411914ed4610d39dd458d54d5516ff6d5bc4

Observation fd632782-eadc-4606-8d53-4ee658adf5c9 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.122600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.122600Z digest=sha256:b9718c73cbffa147ad9eeafffc2a3746ffa78a42dfb3c8d4088f01139f432abb

Observation d7f2f6e5-a0fd-4d05-91d9-d1a3c3424bb8 · outbound

This paper cites EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.137620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.137620Z digest=sha256:6136ea0c4688eb0326b606e05e24b9e8ccc183725960521e190d561d277322e5

Observation 6d85d712-5c57-410d-8660-822dd4315a3c · outbound

This paper cites Fast Segment Anything.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Fast Segment Anything

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.151305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.151305Z digest=sha256:b29c87328bcc97875e0a4351e2bc2e0c7c8706107f342b65822ed122aef6d542

Observation 98aa6ac5-3a2e-4830-9d09-ed2199079c42 · outbound

This paper cites Audio-visual segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio-visual segmentation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:12.500281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:12.170613Z digest=sha256:f34a32b256fe4b78e714b5265013b97744d63081022e7f6190ee43f3963b704a

Observation 3321ace5-26c3-49e0-8a45-7ab64cb4f444 · outbound

This paper cites Audio-Visual Segmentation with Semantics.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio-Visual Segmentation with Semantics

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.186738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.186738Z digest=sha256:75bcbc98d806af8ba29ec2a138e0c6bbb7526022b806fa9f4447d0b036f8b097

Observation 518e8f9d-885c-4021-83d6-3e19bb0a794f · outbound

This paper cites Generalized decoding for pixel, image, and lan- guage.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Generalized decoding for pixel, image, and lan- guage

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.200891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.200891Z digest=sha256:d62b4d9b001f2f8d2358cf701f7714fb2f9d684f5ed23690dc276e6b2e9ae531

Pith citing papers

No inbound Pith citation observations are available.