Pith. sign in

Paper Citation Record · LEDGER

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes

As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.01558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01558 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:44:12.200891Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f657d51-1ebb-49b7-a1c1-a31055a06513 · outbound

This paper cites Self-calibrated clip for training-free open-vocabulary segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Self-calibrated clip for training-free open-vocabulary segmentation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.322434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.322434Z digest=sha256:f3db71e8b61f531cb477c31086b946dac26b5f63580a713167e11a837ead4582

Observation 4106bdea-8ebf-42b7-b522-168cf3635064 · outbound

This paper cites Unraveling in- stance associations: A closer look for audio-visual segmenta- tion.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Unraveling in- stance associations: A closer look for audio-visual segmenta- tion

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.348892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.348892Z digest=sha256:e85e00a68c64c999fa4e833d2999e7279757dd0630f61110e0badc5fd82c71fb

Observation 28ddbb4c-1743-43ca-a927-991312d9307f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.369535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.369535Z digest=sha256:31d2977b831722068fb6db518611b461e2e659bd99814592aae6f3eb2f00bf7c

Observation 9dbd80ee-966b-44cf-9bd8-158540a1080f · outbound

This paper cites Avsegformer: Audio-visual segmentation with trans- former.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Avsegformer: Audio-visual segmentation with trans- former

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.736426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.393623Z digest=sha256:674eb4a56411bef2d4bcde9a6baf8af83834abbed568ba3d2c087dce5b4a101b

Observation 20ec1905-41b2-4500-bd1c-255311445b5a · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio set: An ontology and human- labeled dataset for audio events

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.699455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.422743Z digest=sha256:be7002da5dc80601b6d4253eb5f75d4ed944eb1616e45c5399c055f3af002595

Observation 36ef3bfc-767c-4f91-9234-e25c73b26324 · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Cnn archi- tectures for large-scale audio classification

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.667367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.443369Z digest=sha256:f4f613d34ebe10c6d247800007baf3bd223ed152b4410b1eb79268d577481ada

Observation 0bd5aab6-3c94-4e93-9d15-3be7573625d0 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.631580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.467929Z digest=sha256:de4eae7682e70764ccda3018b75831cab46a60d323dda59703d67c538ba0d15d

Observation 112943e8-e328-4570-85bf-0f711cd5c1a1 · outbound

This paper cites Segment any- thing.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Segment any- thing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.487126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.487126Z digest=sha256:62b6888c0e97c78e25a917e3f2a938b44ebf340eda84a6a5ea16d6b11026fd3f

Observation 44aa1544-d441-432f-a99f-8a2339efc465 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Lisa: Reasoning segmentation via large language model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.507306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.507306Z digest=sha256:178ef0947236afc9cfe2dd901a5f4faf5f8dc3b7e657231129694e6d209d18a4

Observation a8e998dc-9bfe-4169-a895-9b7f0e2d6e10 · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Learning to answer questions in dynamic audio-visual scenarios

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.580940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.527309Z digest=sha256:2264b97039fa73e6788c2af3e50b59d8f1fb9c5ef6ed9737db968e8532555ff6

Observation dd7a614f-769c-40c2-a826-1465021b808b · outbound

This paper cites Robust referring video object segmentation with cyclic structural consensus.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Robust referring video object segmentation with cyclic structural consensus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.551397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.551397Z digest=sha256:e46e0e9273e17fe13db3be9646388f20f65e742aaeb29f4f5963b740fa5e23a1

Observation 68ac84f6-97a6-4d9e-9619-a4d26330f4de · outbound

This paper cites Libero: Benchmarking knowl- edge transfer for lifelong robot learning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Libero: Benchmarking knowl- edge transfer for lifelong robot learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.532561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.568454Z digest=sha256:dbc8dbb80fe991a84a45a10631c6f5ea8c46f8cd98f9cc6ecbd6f660f5012642

Observation b7696e5c-1259-4a90-a136-bc97a26f253b · outbound

This paper cites GRES: Gen- eralized referring expression segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes GRES: Gen- eralized referring expression segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.498870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.584902Z digest=sha256:993ae2787c8a2a580461d03dcc1e993214d1bb9d087b1114d927200a6fa24122

Observation abedcc9d-ea78-471c-b193-3feea0248855 · outbound

This paper cites Improved baselines with visual instruction tuning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Improved baselines with visual instruction tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.601043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.601043Z digest=sha256:2d77ae418bb7a22511f1db7b80bd5b9b77260e028910a23878f19015286ee8aa

Observation e108d5a6-4276-4322-8943-95722e94cc8f · outbound

This paper cites Visual instruction tuning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.621005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.621005Z digest=sha256:aa76a2d4124b279e3831bf06d7bf0ddab77dd52165144caa2191d1b2c30deb30

Observation a2c1f8d2-72ed-4b08-804b-9a28108aa5a3 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.639745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.639745Z digest=sha256:84ef0e2da8a682e93ca11dedaadda7702cea2c85591fe91c8ab9b343db850786

Observation e1580b59-e6a2-4b27-adb5-7382a9c7bf06 · outbound

This paper cites Open-vocabulary segmentation with semantic-assisted calibration.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Open-vocabulary segmentation with semantic-assisted calibration

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.446323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.660057Z digest=sha256:2cb777e899bd30619bfd47358c06416970253eb326dc6fa7d003d227b995cbbb

Observation d3be685d-3edd-4cb2-a54e-cce2bfe266c4 · outbound

This paper cites Universal segmentation at arbi- trary granularity with language instruction.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Universal segmentation at arbi- trary granularity with language instruction

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.399232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.677555Z digest=sha256:555074b03ae6ab15b7d0990c5320bcd74e2f038ba3b1ee788a85fd1935f1820d

Observation f94e348d-ca8a-47fc-ace8-4688e635702d · outbound

This paper cites Decoupled weight de- cay regularization.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Decoupled weight de- cay regularization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.704295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.704295Z digest=sha256:f3cfef50a3f359a6dc0f354ce3ac2d3e41173b7619382cb72639acee0031a66c

Observation 09e0cbf6-df4a-49e9-94f7-25215f04a837 · outbound

This paper cites Soc: Semantic-assisted object cluster for referring video object segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Soc: Semantic-assisted object cluster for referring video object segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.354324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.721300Z digest=sha256:0f8fd48b854dbc558910a95a6d15a52855408b39a237b1904a9be1e97fe13535

Observation 16f510cc-11dc-4c5a-a067-d6c1aef562ba · outbound

This paper cites Mod- eling context between objects for referring expression un- derstanding.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Mod- eling context between objects for referring expression un- derstanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.313045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.738165Z digest=sha256:ace9481264849e75af640314a65a621d0d4fde717f9cc6d6932abd3a1be1678c

Observation b645e32b-6c58-4403-94c1-1563de57e8aa · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes SAM 2: Segment Anything in Images and Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.756990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.756990Z digest=sha256:f24f8b2794762522a4696e08922619c43faa06ad184dc2306cb591a42fc2eb9c

Observation 247f7548-6d8e-430a-9382-b699f0dee7a6 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Pixellm: Pixel reasoning with large multimodal model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.269637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.778749Z digest=sha256:a2832a6bd03c35e088afc7b4cfca4eab901992781c11b127a98a3efad36ad76e

Observation e51c9301-f233-417f-95a4-7afa4d1929a7 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.802462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.802462Z digest=sha256:5b7e8856f23758a7b893285f3eade64a22b584098d260ac9b12396960114744a

Observation 3028eb4f-1c46-4812-92eb-c73560c00a93 · outbound

This paper cites Efficient attention: Attention with lin- ear complexities.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficient attention: Attention with lin- ear complexities

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.225927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.832827Z digest=sha256:cd90472d7b1f3dcc951790485963441892fdf85c26507a23114f257815ca9069

Observation 5878c5f0-72c0-41ed-a211-55eac3e2713c · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Roformer: Enhanced transformer with rotary position embedding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.848405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.848405Z digest=sha256:31157033e49a48c5331e054aca6e1409ddeda9c1da8b65a54321b29e3a9f8fc5

Observation bc52a43d-4679-4847-ad56-7721d435448c · outbound

This paper cites Auto- acd: A large-scale dataset for audio-language representation learning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Auto- acd: A large-scale dataset for audio-language representation learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.182438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.871678Z digest=sha256:0d62c73f6dec06a87c21709ed651418cb9ee89c849fed744c3007e507f5d2d1e

Observation f31b91df-7213-47e0-aca7-bd32ddac1e79 · outbound

This paper cites Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.892307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.892307Z digest=sha256:b5cdb0ba3bcf88e58f6d89497047ba02dbffe0cd88b1bd2d19942f6afcc1a681

Observation 64f558e1-147b-468f-878c-483ff6fa7bbe · outbound

This paper cites Efficient remote sensing transformer for coastline detection with sentinel-2 satellite imagery.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficient remote sensing transformer for coastline detection with sentinel-2 satellite imagery

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.151753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.907122Z digest=sha256:565ff10c3205f9b64a1fe35092741f0b6d8bbfcd5f49f80844500ce2a7035530

Observation 58ab5cbc-bd89-45e4-b97f-f47dc87a431e · outbound

This paper cites Prompting segmentation with sound is gen- eralizable audio-visual source localizer.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Prompting segmentation with sound is gen- eralizable audio-visual source localizer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.117916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.924209Z digest=sha256:b2a5d11913fd1278d7ad0cc3686e715ea6a088f806bc15d9750637d889b95a83

Observation 2e111730-5219-4b15-8ebf-ad71dd6e9371 · outbound

This paper cites Ref-avs: Refer and segment objects in audio-visual scenes.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Ref-avs: Refer and segment objects in audio-visual scenes

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.060729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.945322Z digest=sha256:c218fac08a04949f64b2dd33f7682f687af395890f9feae21579460274cab566

Observation 493dc4ad-5486-4df8-804b-0238b2ed74bc · outbound

This paper cites Convolution meets trans- former: Efficient hybrid transformer for semantic segmenta- tion with very high resolution imagery.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Convolution meets trans- former: Efficient hybrid transformer for semantic segmenta- tion with very high resolution imagery

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.026308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:11.959188Z digest=sha256:968ba2cf443d2c0ffa6294a3949c64930e7fd45029c6316bfbebeb4b461fbf38

Observation 5cb92975-e83f-4199-a288-d262b10ca69c · outbound

This paper cites IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.980254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.980254Z digest=sha256:a95805b0fb52a9584de03add3817985ec9d9447fc7100af28e4096b59f594447

Observation 4138baf4-d4f3-423a-bb44-b3b7c22d4c7c · outbound

This paper cites Onlinerefer: A simple online baseline for referring video object segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Onlinerefer: A simple online baseline for referring video object segmentation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.997831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.997831Z digest=sha256:8d206c491c715ba83049a9f58eea7d49720d659e9445c32af53f3d689c4df41b

Observation b4aace26-db3c-4949-8b24-73b9f780121b · outbound

This paper cites Language as queries for referring video object seg- mentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language as queries for referring video object seg- mentation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.014816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.014816Z digest=sha256:07813feed949ea88b41d6a9c73fefd59aa6d58f2ebfd5b3e5be965e30d7a9bff

Observation 0882748c-073e-47fa-aa91-48331e4454a4 · outbound

This paper cites Language as queries for referring video object seg- mentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language as queries for referring video object seg- mentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:13.972399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:12.032542Z digest=sha256:4c436dfd379be62ade796ce2e5e2e52af3466cc3ed8b487e33f6c55a8b30097e

Observation 4d8b0770-911d-4830-955e-49fbd5ba5748 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Gsva: Generalized segmentation via multimodal large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:13.935682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:12.045116Z digest=sha256:1043e52295e92ce3a1e935292fd161fac5af7e727778ba5b7ac28ae62f2895c7

Observation d7575162-9125-4f72-a778-1645e4c6b5a8 · outbound

This paper cites Efficientsam: Leveraged masked image pretraining for efficient segment anything.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficientsam: Leveraged masked image pretraining for efficient segment anything

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.061070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.061070Z digest=sha256:785beb8320446717c83f0b07cef5f7e9e17f9041aaee3c2884ccc3092b5aaf75

Observation 338baf62-16e1-439a-bd25-27965cef8bea · outbound

This paper cites Avqa: A dataset for audio- visual question answering on videos.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Avqa: A dataset for audio- visual question answering on videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:13.878090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:12.078141Z digest=sha256:ecfbffe3301af94456358facc23211b86d9c70a363ff14734833ac3f1ff7b7b1

Observation 9556304f-7048-4819-9f84-f36361c06f50 · outbound

This paper cites Lavt: Language-aware vision transformer for referring image segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Lavt: Language-aware vision transformer for referring image segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:12.578490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:12.092128Z digest=sha256:a49b5f648c28a93d37aec7ddd5e44c7f1950d26d0f69ff432f017a24973427f6

Observation 16a50437-db4c-45f7-9e15-127705cd219f · outbound

This paper cites Language- aware vision transformer for referring segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language- aware vision transformer for referring segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:12.541709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:12.110003Z digest=sha256:3e24b109a0e33d12b09979b308463d99792e0665c00d4c4aad593a50ccf015bb

Observation fd632782-eadc-4606-8d53-4ee658adf5c9 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.122600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.122600Z digest=sha256:b9718c73cbffa147ad9eeafffc2a3746ffa78a42dfb3c8d4088f01139f432abb

Observation d7f2f6e5-a0fd-4d05-91d9-d1a3c3424bb8 · outbound

This paper cites EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.137620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.137620Z digest=sha256:6136ea0c4688eb0326b606e05e24b9e8ccc183725960521e190d561d277322e5

Observation 6d85d712-5c57-410d-8660-822dd4315a3c · outbound

This paper cites Fast Segment Anything.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Fast Segment Anything

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.151305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.151305Z digest=sha256:b29c87328bcc97875e0a4351e2bc2e0c7c8706107f342b65822ed122aef6d542

Observation 98aa6ac5-3a2e-4830-9d09-ed2199079c42 · outbound

This paper cites Audio-visual segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio-visual segmentation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:12.500281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:12.170613Z digest=sha256:1003fce871551595bb3c604a6146c977700446ca11fc8a9513374f695d71b432

Observation 3321ace5-26c3-49e0-8a45-7ab64cb4f444 · outbound

This paper cites Audio-Visual Segmentation with Semantics.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio-Visual Segmentation with Semantics

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.186738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.186738Z digest=sha256:75bcbc98d806af8ba29ec2a138e0c6bbb7526022b806fa9f4447d0b036f8b097

Observation 518e8f9d-885c-4021-83d6-3e19bb0a794f · outbound

This paper cites Generalized decoding for pixel, image, and lan- guage.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Generalized decoding for pixel, image, and lan- guage

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.200891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.200891Z digest=sha256:d62b4d9b001f2f8d2358cf701f7714fb2f9d684f5ed23690dc276e6b2e9ae531

Pith citing papers

No inbound Pith citation observations are available.