Pith. sign in

Paper Citation Record · LEDGER

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes

As of 22 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2506.01558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01558 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:44:12.200891Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f657d51-1ebb-49b7-a1c1-a31055a06513 · outbound

This paper cites Self-calibrated clip for training-free open-vocabulary segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Self-calibrated clip for training-free open-vocabulary segmentation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.322434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.322434Z digest=sha256:b942bc2d2f867b24222d8ded2b93a2e3f050a69b3770d7e73a41cc47773d0a0b

Observation 4106bdea-8ebf-42b7-b522-168cf3635064 · outbound

This paper cites Unraveling in- stance associations: A closer look for audio-visual segmenta- tion.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Unraveling in- stance associations: A closer look for audio-visual segmenta- tion

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.348892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.348892Z digest=sha256:81cde7582a1d15bea1b81ea7af4d6757db636de37ed33ff99c0c8c4bcf29d177

Observation 28ddbb4c-1743-43ca-a927-991312d9307f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.369535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.369535Z digest=sha256:27fb00c1bd61f3a8aae1a585c0b406c1635e2acf28fa78d5c02577ae5e430442

Observation 9dbd80ee-966b-44cf-9bd8-158540a1080f · outbound

This paper cites Avsegformer: Audio-visual segmentation with trans- former.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Avsegformer: Audio-visual segmentation with trans- former

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.736426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.393623Z digest=sha256:10cad99d2b847b38f5b6109a04614f6b940b4a82831e48fb0df3677dcabe1b96

Observation 20ec1905-41b2-4500-bd1c-255311445b5a · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio set: An ontology and human- labeled dataset for audio events

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.699455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.422743Z digest=sha256:4312fd59a49bf081e147f73ba9f4cc516901762e3764c06e9ae19664d9dc2f18

Observation 36ef3bfc-767c-4f91-9234-e25c73b26324 · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Cnn archi- tectures for large-scale audio classification

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.667367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.443369Z digest=sha256:9f129b6a96846e34fe65a2019efe6896531bbe089b57c4f2cf4649f45036109a

Observation 0bd5aab6-3c94-4e93-9d15-3be7573625d0 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.631580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.467929Z digest=sha256:db38062c6a3a897c5f7b6d9a4e3540b631955a9abfc5c3c0170ca051ac96db6a

Observation 112943e8-e328-4570-85bf-0f711cd5c1a1 · outbound

This paper cites Segment any- thing.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Segment any- thing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.487126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.487126Z digest=sha256:decf70983a7b8a13387450c3ca45bc49c3d3653e2f90b21ac2d02c69ba355cfc

Observation 44aa1544-d441-432f-a99f-8a2339efc465 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Lisa: Reasoning segmentation via large language model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.507306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.507306Z digest=sha256:ca96dabf06ba01ce3845df672a38fea8a24ed0a1f8366621f5adb876d8b9fd27

Observation a8e998dc-9bfe-4169-a895-9b7f0e2d6e10 · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Learning to answer questions in dynamic audio-visual scenarios

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.580940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.527309Z digest=sha256:bd6a263778884b62d57513d106a13eeca8321488a098bea9c2b6b0c4d7d76f02

Observation dd7a614f-769c-40c2-a826-1465021b808b · outbound

This paper cites Robust referring video object segmentation with cyclic structural consensus.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Robust referring video object segmentation with cyclic structural consensus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.551397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.551397Z digest=sha256:3664a09c9ac96e644d6254c5479e795b65fa9f852b5860acca5d9c9b32d113d4

Observation 68ac84f6-97a6-4d9e-9619-a4d26330f4de · outbound

This paper cites Libero: Benchmarking knowl- edge transfer for lifelong robot learning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Libero: Benchmarking knowl- edge transfer for lifelong robot learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.532561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.568454Z digest=sha256:b50eee03380fc161ca2fc47d734e3f2e4764d12e9164a1214c9bece049e0ddf6

Observation b7696e5c-1259-4a90-a136-bc97a26f253b · outbound

This paper cites GRES: Gen- eralized referring expression segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes GRES: Gen- eralized referring expression segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.498870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.584902Z digest=sha256:a6c61befa92925ba39507afd002796cc1cb1596ea78f3bca0e6be7eae3b9ef53

Observation abedcc9d-ea78-471c-b193-3feea0248855 · outbound

This paper cites Improved baselines with visual instruction tuning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Improved baselines with visual instruction tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.601043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.601043Z digest=sha256:131d687de01b4d6c0299b6dc325cf6132cba027e37a3161a219fe0baefedddd5

Observation e108d5a6-4276-4322-8943-95722e94cc8f · outbound

This paper cites Visual instruction tuning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.621005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.621005Z digest=sha256:b29d25d2548db572dce8538e1a73ea7b137fa1acf86aec92748e39677122dea5

Observation a2c1f8d2-72ed-4b08-804b-9a28108aa5a3 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.639745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.639745Z digest=sha256:12ca887fd79520a5d36780df340fef5f17ea2bac8c55a098bf2a57cb5ce647f6

Observation e1580b59-e6a2-4b27-adb5-7382a9c7bf06 · outbound

This paper cites Open-vocabulary segmentation with semantic-assisted calibration.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Open-vocabulary segmentation with semantic-assisted calibration

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.446323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.660057Z digest=sha256:1789c4fc020463785c6b2fbc0345fe487ca29c7a2584430a569f1b63c32bb7c8

Observation d3be685d-3edd-4cb2-a54e-cce2bfe266c4 · outbound

This paper cites Universal segmentation at arbi- trary granularity with language instruction.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Universal segmentation at arbi- trary granularity with language instruction

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.399232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.677555Z digest=sha256:6a3a8349511cc287cb81619c848ea68a105f1fd29017343dc9b5bab789e34918

Observation f94e348d-ca8a-47fc-ace8-4688e635702d · outbound

This paper cites Decoupled weight de- cay regularization.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Decoupled weight de- cay regularization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.704295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.704295Z digest=sha256:315c481876a44d7bedd4fb15c0384e150efe4acfb5410ad27586a11cce0a9e69

Observation 09e0cbf6-df4a-49e9-94f7-25215f04a837 · outbound

This paper cites Soc: Semantic-assisted object cluster for referring video object segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Soc: Semantic-assisted object cluster for referring video object segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.354324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.721300Z digest=sha256:57d220b13d09c783ad2606b687c9ef0648e871eeede242a65f2bfe51bc88d6b6

Observation 16f510cc-11dc-4c5a-a067-d6c1aef562ba · outbound

This paper cites Mod- eling context between objects for referring expression un- derstanding.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Mod- eling context between objects for referring expression un- derstanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.313045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.738165Z digest=sha256:788bd5407ffc867b540aa45f7c795349405b70e93a8deaf6b4ed3cfd0e52320f

Observation b645e32b-6c58-4403-94c1-1563de57e8aa · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes SAM 2: Segment Anything in Images and Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.756990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.756990Z digest=sha256:435e69ee3905afc7967dd61c996af4f743cad00245ae503e78858fc5a002e00a

Observation 247f7548-6d8e-430a-9382-b699f0dee7a6 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Pixellm: Pixel reasoning with large multimodal model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.269637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.778749Z digest=sha256:66099844431150929be005cb073b5f85c27e227c5bdac9cc3022f80bc1dd6900

Observation e51c9301-f233-417f-95a4-7afa4d1929a7 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.802462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.802462Z digest=sha256:484cfd8019aa9e21abc3de6f3ddf017c7e82264f9776e65838d51d55d82281db

Observation 3028eb4f-1c46-4812-92eb-c73560c00a93 · outbound

This paper cites Efficient attention: Attention with lin- ear complexities.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficient attention: Attention with lin- ear complexities

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.225927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.832827Z digest=sha256:73f85ee141f7391a58d3b98574e1517debf616fdd6b1fba72bc120ee30c58f8a

Observation 5878c5f0-72c0-41ed-a211-55eac3e2713c · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Roformer: Enhanced transformer with rotary position embedding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.848405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.848405Z digest=sha256:67c29cce2bbc823474c850656be471b16c5469e59cf1194358543aa1e69a7b9d

Observation bc52a43d-4679-4847-ad56-7721d435448c · outbound

This paper cites Auto- acd: A large-scale dataset for audio-language representation learning.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Auto- acd: A large-scale dataset for audio-language representation learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.182438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.871678Z digest=sha256:cf0028c28c4e99b50dc104090ca1b422654648c4422eba4d4929faf3b93afa41

Observation f31b91df-7213-47e0-aca7-bd32ddac1e79 · outbound

This paper cites Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.892307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.892307Z digest=sha256:44c41195bd17fcd2721dd85f606ac39a1b7051a602b8e4036e4b9bd460e4890c

Observation 64f558e1-147b-468f-878c-483ff6fa7bbe · outbound

This paper cites Efficient remote sensing transformer for coastline detection with sentinel-2 satellite imagery.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficient remote sensing transformer for coastline detection with sentinel-2 satellite imagery

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.151753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.907122Z digest=sha256:9bb5add806787f79e6ca53ea91b8c618d75986474094ae4e4ef20e8a5b668d12

Observation 58ab5cbc-bd89-45e4-b97f-f47dc87a431e · outbound

This paper cites Prompting segmentation with sound is gen- eralizable audio-visual source localizer.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Prompting segmentation with sound is gen- eralizable audio-visual source localizer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.117916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.924209Z digest=sha256:6865858ba0221cdd0c497810e63d98dce5e597d455033d9e3e447366a3d5a391

Observation 2e111730-5219-4b15-8ebf-ad71dd6e9371 · outbound

This paper cites Ref-avs: Refer and segment objects in audio-visual scenes.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Ref-avs: Refer and segment objects in audio-visual scenes

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.060729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.945322Z digest=sha256:eed9cd6cc781d97b195980e90cca58b86b7ec761c924c12dc226fc9da8b5e0b5

Observation 493dc4ad-5486-4df8-804b-0238b2ed74bc · outbound

This paper cites Convolution meets trans- former: Efficient hybrid transformer for semantic segmenta- tion with very high resolution imagery.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Convolution meets trans- former: Efficient hybrid transformer for semantic segmenta- tion with very high resolution imagery

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:14.026308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:11.959188Z digest=sha256:6810c35b2950e384ff5469110b389b6985924b8d28e1e483c19b9ea41904aff2

Observation 5cb92975-e83f-4199-a288-d262b10ca69c · outbound

This paper cites IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.980254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.980254Z digest=sha256:a96932fb107e3ae7fa815014ce4ac55302ac9681a8ca5b807a03fd3994fbee03

Observation 4138baf4-d4f3-423a-bb44-b3b7c22d4c7c · outbound

This paper cites Onlinerefer: A simple online baseline for referring video object segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Onlinerefer: A simple online baseline for referring video object segmentation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:11.997831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:11.997831Z digest=sha256:a5d2fbb7344126ef688bbd495ca6ac6933bc3ce16afe602af57c517a4c81d0bd

Observation b4aace26-db3c-4949-8b24-73b9f780121b · outbound

This paper cites Language as queries for referring video object seg- mentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language as queries for referring video object seg- mentation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.014816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.014816Z digest=sha256:4d5b6f1035da4b6b27dcb0189e3a6001dd8d0ec86fcff76958c5647d7417a43e

Observation 0882748c-073e-47fa-aa91-48331e4454a4 · outbound

This paper cites Language as queries for referring video object seg- mentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language as queries for referring video object seg- mentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:13.972399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:12.032542Z digest=sha256:57b06fcb946dac33111e602a76fa3129ffb75a398e0cd9932ddd8a081630d472

Observation 4d8b0770-911d-4830-955e-49fbd5ba5748 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Gsva: Generalized segmentation via multimodal large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:13.935682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:12.045116Z digest=sha256:e3d9990d5e05276d9a0d7be8ddc1f20026c74d3685e0be6c2c49c9252ef37e6c

Observation d7575162-9125-4f72-a778-1645e4c6b5a8 · outbound

This paper cites Efficientsam: Leveraged masked image pretraining for efficient segment anything.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Efficientsam: Leveraged masked image pretraining for efficient segment anything

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.061070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.061070Z digest=sha256:0b7b348a517299c2a8749c94e154fc36a2d585772fbd2b0ea51d958a2d5f5407

Observation 338baf62-16e1-439a-bd25-27965cef8bea · outbound

This paper cites Avqa: A dataset for audio- visual question answering on videos.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Avqa: A dataset for audio- visual question answering on videos

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:13.878090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:12.078141Z digest=sha256:314c7a71ca9be471c8a314c38294aae7fb621356740691125ca91383d0dfeba7

Observation 9556304f-7048-4819-9f84-f36361c06f50 · outbound

This paper cites Lavt: Language-aware vision transformer for referring image segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Lavt: Language-aware vision transformer for referring image segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:12.578490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:12.092128Z digest=sha256:ca395d5fd53c350de02a9734820b8da7ed8d456ddc44e7c779979296ffdd1a22

Observation 16a50437-db4c-45f7-9e15-127705cd219f · outbound

This paper cites Language- aware vision transformer for referring segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Language- aware vision transformer for referring segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:12.541709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:12.110003Z digest=sha256:ff2410487afac030af2712511d789cc3e9bd5729221b923ed14d9708fdceeb8e

Observation fd632782-eadc-4606-8d53-4ee658adf5c9 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.122600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.122600Z digest=sha256:9a7ea03b310c5a55a41a521807127416577d42c8653d32e2b76f5aa1c3c00dc8

Observation d7f2f6e5-a0fd-4d05-91d9-d1a3c3424bb8 · outbound

This paper cites EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.137620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.137620Z digest=sha256:6f3e73e1ccf4cda0f7883fb3a41e65cfdc6cdf75f367c6706b952cca0d540eb2

Observation 6d85d712-5c57-410d-8660-822dd4315a3c · outbound

This paper cites Fast Segment Anything.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Fast Segment Anything

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.151305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.151305Z digest=sha256:483b961e4ce94b037a7eb305d5cb8c3312304c5663f0e9ed985e72eec9a8e0dd

Observation 98aa6ac5-3a2e-4830-9d09-ed2199079c42 · outbound

This paper cites Audio-visual segmentation.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio-visual segmentation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:12.500281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:44:12.170613Z digest=sha256:329a7d3fc7a947ebc591bf7f09282b5f5ebd95f8243ac4cfe3e08917cef7e5f1

Observation 3321ace5-26c3-49e0-8a45-7ab64cb4f444 · outbound

This paper cites Audio-Visual Segmentation with Semantics.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Audio-Visual Segmentation with Semantics

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.186738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.186738Z digest=sha256:91b20e798e0f9ba24538c7033d90bab9ecf8a49de6ff65c7c704d4069eaf0dda

Observation 518e8f9d-885c-4021-83d6-3e19bb0a794f · outbound

This paper cites Generalized decoding for pixel, image, and lan- guage.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes Generalized decoding for pixel, image, and lan- guage

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.200891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.200891Z digest=sha256:4d8037b64c7ac9384f36c4a20f46bc3ad86cd906b5d0d4c1c4d1c103606a4746

Pith citing papers

No inbound Pith citation observations are available.