Pith. sign in

Paper Citation Record · LEDGER

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation

As of 22 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2411.17646.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17646 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:57:39.826709Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:55:33.862576Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:00:27.770659Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 516c5e4d-9897-4587-aee1-6c22e5d883ce · outbound

This paper cites One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.556739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.556739Z digest=sha256:a6fdbc5aea8ff1004b09b85ed8a43f505b99d8010fb3e4ab23f9e90741f59867

Observation b2370a6f-83d9-4f52-8d11-cf30b25deefc · outbound

This paper cites A closer look at referring expressions for video object segmentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation A closer look at referring expressions for video object segmentation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.657885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.562887Z digest=sha256:675cb80e42e888d8948187bb1630e35fde524ff728206008d28b4f1587cf25a1

Observation 541e76cf-e891-411f-be86-56e93fdb1ebc · outbound

This paper cites End-to-end referring video object segmentation with multi- modal transformers.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation End-to-end referring video object segmentation with multi- modal transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.642608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.567658Z digest=sha256:e9f41b6785521a741de61429c4faa68caf37115caaa45f8be1284bb6232a9cbe

Observation 78d454f5-8243-42b1-b08d-509cdea5bb5b · outbound

This paper cites End-to- end object detection with transformers.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation End-to- end object detection with transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.572684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.572684Z digest=sha256:939bd0ae8ab056567a7bf9c9397bf5ada82a7e9d6ba67929dcbd0430de9a98b7

Observation 3d98ff22-6168-4b76-b39d-b8fc5aba6c7a · outbound

This paper cites Adaptformer: Adapting vision transformers for scalable visual recogni- tion.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Adaptformer: Adapting vision transformers for scalable visual recogni- tion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.618218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.577594Z digest=sha256:5e31e5fea852ec93aefa13dc2ca2c82eb323d1055898cd57cfa99fad09d40e3d

Observation 4d075d69-8d21-4120-9430-ffad1b609b9a · outbound

This paper cites Mask grounding for referring image seg- mentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Mask grounding for referring image seg- mentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.603663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.582948Z digest=sha256:70aa26d3ebd124d803085f84ea06726e6763bed0e1b09cd89d03cb39895d0c4a

Observation 9ea197c8-2289-405f-828d-d08f0eb8357d · outbound

This paper cites Vlt: Vision-language transformer and query generation for referring segmentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Vlt: Vision-language transformer and query generation for referring segmentation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.589035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.588107Z digest=sha256:e4ac92ebe7cbe71da81af581cadb173c5f573131ae02b770512aa7e83358fc9b

Observation 6f9c8153-47c1-4c8d-b5ae-b3d65a2ec9d5 · outbound

This paper cites MeViS: A large-scale benchmark for video segmentation with motion expressions.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation MeViS: A large-scale benchmark for video segmentation with motion expressions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.574541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.593215Z digest=sha256:6ff14e4236f980861cd842282cc4a3ec5c1a180e789c7b8a6c1315f60d2501fd

Observation 71fd608c-563d-440c-b48d-e2eed12e28d1 · outbound

This paper cites Progressive multimodal interaction network for referring video object segmentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Progressive multimodal interaction network for referring video object segmentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.559534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.599048Z digest=sha256:73b08d857e1598115d9964d3f3512f323ed02cae26135aa85f32dd8b5573e426

Observation f7945b1b-0468-4edc-9cc5-a09043bf2a3c · outbound

This paper cites Actor and action video segmentation from a sentence.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Actor and action video segmentation from a sentence

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.544815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.603698Z digest=sha256:872be84074fdb5993d298aa4090c0645c969e1820361367834cf83d08cedc955

Observation 6ea7e5fe-6e1c-4f56-8abc-7b316ee1d832 · outbound

This paper cites Html: Hybrid temporal-scale mul- timodal learning framework for referring video object seg- mentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Html: Hybrid temporal-scale mul- timodal learning framework for referring video object seg- mentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.530019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.608181Z digest=sha256:7b099e6f929ace59493249bb5da7c94a4ab68564c5abcdaa5f1b391b40e63d81

Observation cac50cba-3509-4b3f-8b16-99db2b89610e · outbound

This paper cites Decoupling static and hier- archical motion perception for referring video segmentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Decoupling static and hier- archical motion perception for referring video segmentation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.514815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.613532Z digest=sha256:d75cb35d4846e4573bcf981b836894a2a3b92044a1958f9a85f19eab93d4c9bd

Observation 140dee2d-5561-4fb6-bc74-668417b2a6b1 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Parameter-efficient transfer learning for nlp

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.499890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.618027Z digest=sha256:b1dde7692ce5d4aef9ef6a9ddd0b34fe452bd0007ce2002f683837746cf689bb

Observation a87941e8-20ec-4e7b-af0c-19fed35ac66a · outbound

This paper cites Lora: Low-rank adaptation of large language models.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Lora: Low-rank adaptation of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.485669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.622466Z digest=sha256:3e760fb9b53d23e5bd4cf10d963e56876bd19c9d5be4426d335456d149ec662f

Observation 466e9787-6ee4-416e-8d9a-e3eb37500fad · outbound

This paper cites Temporal context enhanced referring video object seg- mentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Temporal context enhanced referring video object seg- mentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.471563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.626909Z digest=sha256:d62e469afcff5869fe7439e1eca5fc99026f3312b39fb38f3753c3d3202e0835

Observation 4d7f97a3-867c-4736-be6a-ee39a597c6df · outbound

This paper cites Cross-Modal Adapter for Vision-Language Retrieval.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Cross-Modal Adapter for Vision-Language Retrieval

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.631384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.631384Z digest=sha256:06eae21c3ba1d7fece803ab36bf8cacfacde5cf99ea6faef55981781805c3566

Observation 84876c47-bd51-41aa-b828-8ade5a3478c0 · outbound

This paper cites Mv-adapter: Multimodal video transfer learning for video text retrieval.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Mv-adapter: Multimodal video transfer learning for video text retrieval

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.456795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.636161Z digest=sha256:a7c2b8162f51adae5f789f75c48d597f2f8f9eec72b1ac46dc48d81ea3a8fcd5

Observation c78c4f96-38b4-452e-9d6a-a4aa8468d4ec · outbound

This paper cites Video object segmentation with language referring expressions.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Video object segmentation with language referring expressions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.640632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.640632Z digest=sha256:d441ab258df257f92f936776ac32beeb8c6f6c21c35a74eac0746af4f4d375ce

Observation 810f7553-7f1d-4e2f-b1ea-f63058485c0d · outbound

This paper cites Segment any- thing.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Segment any- thing

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.644779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.644779Z digest=sha256:648778920faf0c97a7f7c0774e0516c6d7e3b6b93c25175be22c1ae1b175e564

Observation 691b5412-517b-47e9-a6ad-431d33919417 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Lisa: Reasoning segmentation via large language model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.424311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.649695Z digest=sha256:7936da181fe1de699e603e28b3d6dfce2bbe90eb74355fb3dccef129741d87ae

Observation 48a8285a-b5f5-41a6-b965-d2664864eddd · outbound

This paper cites Refsam: Efficiently adapting segmenting anything model for referring video object segmentation, 2024.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Refsam: Efficiently adapting segmenting anything model for referring video object segmentation, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.410565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.654201Z digest=sha256:e1a2eadfe4d1f384abd3cc7a5f0f2f9d98e5ba7faa852c80155ecd539c5150bb

Observation e35b0d9e-14d8-451a-a520-459b362616eb · outbound

This paper cites Visual instruction tuning.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Visual instruction tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.658450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.658450Z digest=sha256:2f5810f12fc3a973791174f17cf0bbf3167855e8f7d3710a394b6073deeed5d1

Observation 8fc8d0cc-02d8-4c8c-a73b-8d4cf2e4eec5 · outbound

This paper cites Revisiting temporal modeling for clip-based image-to-video knowledge transferring.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Revisiting temporal modeling for clip-based image-to-video knowledge transferring

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.385423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.662656Z digest=sha256:7ddda4c5f7b83813fd28afde678ba4009254be307df39fd17ea6c1862e807a40

Observation a80647c7-d3fb-44d6-a8b5-6da044ddc740 · outbound

This paper cites Li, Ying Shan, and Ge Li.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Li, Ying Shan, and Ge Li

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.371623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.667089Z digest=sha256:21bcf365767c9445036b80d9c9d2ccab9bcecaad4271c9c49a2537876b69780d

Observation 7d9faebc-ca8d-48b9-a71f-c7623e49190c · outbound

This paper cites Cross-modal progressive comprehension for referring segmentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Cross-modal progressive comprehension for referring segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.357641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.671628Z digest=sha256:3b4f692c4254934943d508b01de2f1b264171c997867ff00c64de91d30ccd8d5

Observation ff93141b-9bf1-4316-b4c7-fb686ec2b4fd · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.675900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.675900Z digest=sha256:f4399eef26a68842afddb963f30837289a33b72f14851b084af7a22ce87122aa

Observation 63124c46-3ff1-4b7b-b075-57a9e6769e26 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.680657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.680657Z digest=sha256:ad7515b41f24506367ca56e30a005c1d47b31fae142bda3f8badec2b8e9a04c2

Observation d605d979-d476-4a63-ad82-382b00ac0d62 · outbound

This paper cites UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.685781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.685781Z digest=sha256:9908639365624fd2ed55b4de56270cce19c206c1e8f5de1f177c57b444651270

Observation 2acdb8c9-dc74-4c71-a272-c204e037213a · outbound

This paper cites Soc: Semantic-assisted object cluster for referring video object segmentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Soc: Semantic-assisted object cluster for referring video object segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.343549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.690545Z digest=sha256:aa6000ddf1a3f1a0be1a5982f1221c9f9e9f6beeb6da983872f2ff9454698851

Observation 63715c28-5d61-4434-b223-f66160ec8ead · outbound

This paper cites Spectrum-guided multi-granularity referring video object segmentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Spectrum-guided multi-granularity referring video object segmentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.329300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.694978Z digest=sha256:98b99e2dafd7105e5caac3941c5cecd005fc3e33deeec303d55a13cc26f35a1d

Observation d7b82916-2165-4d07-a0b9-6308c9769951 · outbound

This paper cites Mod- eling context between objects for referring expression under- standing.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Mod- eling context between objects for referring expression under- standing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.699480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.699480Z digest=sha256:fd2b1180986ad3db7bd29601250216dbb0eab16490103481a1e6f5ba4836484c

Observation 5e6de772-05cd-4184-a3f3-6b6ef50cdec5 · outbound

This paper cites Video object segmentation using space-time memory networks.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Video object segmentation using space-time memory networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.304598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.705174Z digest=sha256:d841aaab90bd5d35a27f7faadc72319deca8b5962643d8a83bbf1c6b4e3c36f3

Observation dca92c30-1485-4102-95b3-9485cc440d68 · outbound

This paper cites Keeping your eye on the ball: Tra- jectory attention in video transformers.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Keeping your eye on the ball: Tra- jectory attention in video transformers

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.289633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.709517Z digest=sha256:93bf15d7c4e73c1a03a6659e11c39ad09094e42b3684a758b67770b791e75022

Observation b5eedeb7-9430-40d4-aa34-cf8620a0103f · outbound

This paper cites To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.714035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.714035Z digest=sha256:298330e773286b8d14cfc0a96f9e76f104c67b02c96987d911b73920e695bc76

Observation 010d4f8a-ef35-4387-ad7f-254a69d76852 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Learning transferable visual models from natural language supervi- sion

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.718791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.718791Z digest=sha256:a843b6f608def930acb65e025bd20a7dcc31dddc48c0e21d63b6d12024c9013f

Observation 04624361-602d-4ab3-8a09-ba9ce2260989 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation SAM 2: Segment Anything in Images and Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.723468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.723468Z digest=sha256:098acefd4919d76b03f956250c4ec73bc1ce47e41db20b465d63ed771ad8af7d

Observation f30ed969-f7af-4646-a08d-068f85be13d3 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.728081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.728081Z digest=sha256:b33380b60489c5800d90270f011d5e38f889247e35d68abd30a7c33d028b20de

Observation 3dbc1fcf-a728-4b8b-b3ab-b0a38ade82d2 · outbound

This paper cites Hi- era: A hierarchical vision transformer without the bells-and- whistles.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Hi- era: A hierarchical vision transformer without the bells-and- whistles

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.265542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.734083Z digest=sha256:9ec918cd59a944ada570cfaa3d5c2f23dfcb3c3e319467d2f66d79f40fbb510f

Observation 88394d67-f7fd-46ee-af16-acb67bec6cbd · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.250370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.739537Z digest=sha256:6ebd3070b880f8c114352ee03f084641e34adc90bad3fc652fb61a653122c925

Observation fb185e65-5305-4519-bb2f-12c299858945 · outbound

This paper cites Temporal collection and distribution for referring video object segmentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Temporal collection and distribution for referring video object segmentation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.744073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.744073Z digest=sha256:936e59651caa5c912887dfdef88efaf2e60c42603456e81824ea52adcb69d5b1

Observation 79e3d4e4-603d-4c1f-8acb-236521fcd79d · outbound

This paper cites A multimodal, multi-task adapting frame- work for video action recognition.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation A multimodal, multi-task adapting frame- work for video action recognition

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.224392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.748463Z digest=sha256:916f31da2c6b978758f1630a20c699c97bef2491679c65ba16b6b3ef0ba36fae

Observation 35e7943b-daad-4ccb-9813-16c9273bd9e4 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.752915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.752915Z digest=sha256:b78ef6fbd9d16c6ace4ba7aa979790f3ac9cc87fb4cce74bca6f8314d79ad959

Observation b1a83b00-7f14-4ae0-bc63-f10e334fc25a · outbound

This paper cites OnlineRefer: A simple online baseline for referring video object segmentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation OnlineRefer: A simple online baseline for referring video object segmentation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.207621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.757584Z digest=sha256:46695060f6cca0049264cf31974a8e0d36e21dd02a65c915ffa74e0a9a93f2f4

Observation 84a8429a-e192-4658-9c77-b8d506d2fe95 · outbound

This paper cites Language as queries for referring video object seg- mentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Language as queries for referring video object seg- mentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.192518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.761981Z digest=sha256:1989672893816ce96401fb922b266f7a7484294fc5b119922c20b0257bd5dc6b

Observation 2993c173-7e0b-4ebb-8249-6dd11ab558c3 · outbound

This paper cites Bridging vision and language en- coders: Parameter-efficient tuning for referring image seg- mentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Bridging vision and language en- coders: Parameter-efficient tuning for referring image seg- mentation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.766344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.766344Z digest=sha256:d414618c74dd22ae7276210974ee50ea8ecf70995b561dacc26382ba4861677e

Observation cd71f7c3-71bd-412e-a9dd-a0f84fe12063 · outbound

This paper cites VISA: Reasoning Video Object Segmentation via Large Language Models.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.770468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.770468Z digest=sha256:90e878add28e7f5d87322f0d31e711b7e4c98c6fc13d29cffe5a4e40c866e29c

Observation 918be086-2a92-4ee7-b03d-36d87db88b6e · outbound

This paper cites Referred by multi-modality: A unified tem- poral transformer for video object segmentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Referred by multi-modality: A unified tem- poral transformer for video object segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.168569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.774967Z digest=sha256:2f70d515f1d59f5a3d525c149e0af9db462abe777ecc880880fcba3a21c13059

Observation 7e133695-56fd-42e4-ae69-cc7d29db6e6d · outbound

This paper cites Mma: Multi-modal adapter for vision-language models.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Mma: Multi-modal adapter for vision-language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.779235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.779235Z digest=sha256:8ea9dcf37f4eb0b7ba60991c70eaeaa5dcc76e9c7ec3def851d46546fd730f01

Observation 12e28ed4-30b9-4a1c-8d3c-324b36d56156 · outbound

This paper cites Cross-modal self-attention network for referring image seg- mentation.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Cross-modal self-attention network for referring image seg- mentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.144384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.784148Z digest=sha256:0465528f33975f264de2a781e40b4004d8cb749f84b2faebdae1d8dcd09577c9

Observation 2539c044-6010-45e9-b9be-b0cb02eae43f · outbound

This paper cites Modeling context in referring expres- sions.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Modeling context in referring expres- sions

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.788349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.788349Z digest=sha256:758ea4d46859cb3fa697415f9baad7680f9cde613b24cfd468292c38a371e8ee

Observation 668c52a9-896a-432b-8f8f-de6e0860adc6 · outbound

This paper cites EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.792553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.792553Z digest=sha256:0f92407bfd21c7c8c9f19c295c86d9ef344c440528ebe38933e0661debd91427

Observation 6c6b5c92-fd5e-444c-8136-dda76604c062 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.798025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.798025Z digest=sha256:8e13cb70eadd92cbbcb317c721e9e672be073c1670c5beead0a6f99673d692ac

Observation 455a2c65-5df8-4f43-a02f-7be9e89abed4 · outbound

This paper cites We train our Conditional Mem- ory Encoder (CME) via self-supervision.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation We train our Conditional Mem- ory Encoder (CME) via self-supervision

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.118564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.802627Z digest=sha256:5d0c6bbdd424951970a40ecfca6103a4b1edc269bca8d5290164d4f12fe1aecf

Observation b14436c1-85e2-40bb-9806-18330b73a69a · outbound

This paper cites an unresolved cited work.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:57:40.100592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.807721Z digest=sha256:d353716897056813a7c9afe2c52e18be9707e7d7af3de743b95bf52ff5bf8e14

Observation 0d1142a9-4702-4ad2-b0b9-0f0778483173 · outbound

This paper cites 9, where we plot the memory features.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation 9, where we plot the memory features

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.085543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.813264Z digest=sha256:1d19286e8a04839e480194031f025e81ac5d5f0df28b0aea5cf194573b345fc2

Observation 534038b6-1136-4aff-98e2-7b9f0ca3c58d · outbound

This paper cites an unresolved cited work.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:57:40.071205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.817588Z digest=sha256:5581b3662d35104f8109cedd6a68be159d85e5559e682a9580587667e680681c

Observation 38af42ba-b164-422e-9779-3920a460769c · outbound

This paper cites an unresolved cited work.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:57:40.057450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.822518Z digest=sha256:29690b04a9c469b980f24272fc9478c5ee9a04e7ff4e375f04c0b072229cccdd

Observation fc991d70-0bfd-4eb9-8997-ba63c15263f5 · outbound

This paper cites 10, we present qualitative examples from the MeViS dataset that highlight the effectiveness of SAMWISE.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation 10, we present qualitative examples from the MeViS dataset that highlight the effectiveness of SAMWISE

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:57:40.042874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T11:57:39.826709Z digest=sha256:331d7c928117403ef1fd6c7066ff23bd979a1e4496db01681ad40de1d14d2499

Pith citing papers

Observation 56a71c8e-6ca6-49e5-baea-48b64b72f8a2 · inbound

ReSurgSAM2: Referring Segment Anything in Surgical Video via Credible Long-term Tracking cites this paper.

ReSurgSAM2: Referring Segment Anything in Surgical Video via Credible Long-term Tracking SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:55:33.862576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:55:33.862576Z digest=sha256:4a760de284be53fa317e27839b4b4d0bc25c2f95f33d301bc28f9b8286c1974f

Observation 946344be-ebcb-4169-8c1b-b29664be802d · inbound

BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos cites this paper.

BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:00:27.774972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:00:27.151946Z digest=sha256:cfb74bb369dcd466e82660bb9cbf90e7ec46a2494bc68c46e8c2c495c0481e74