Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:57:39.826709Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2411.17646.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:57:39.826709Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:55:33.862576Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T23:00:27.770659Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 516c5e4d-9897-4587-aee1-6c22e5d883ce · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2370a6f-83d9-4f52-8d11-cf30b25deefc · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation A closer look at referring expressions for video object segmentation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 541e76cf-e891-411f-be86-56e93fdb1ebc · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation End-to-end referring video object segmentation with multi- modal transformers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 78d454f5-8243-42b1-b08d-509cdea5bb5b · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation End-to- end object detection with transformers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d98ff22-6168-4b76-b39d-b8fc5aba6c7a · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Adaptformer: Adapting vision transformers for scalable visual recogni- tion
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4d075d69-8d21-4120-9430-ffad1b609b9a · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Mask grounding for referring image seg- mentation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9ea197c8-2289-405f-828d-d08f0eb8357d · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Vlt: Vision-language transformer and query generation for referring segmentation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6f9c8153-47c1-4c8d-b5ae-b3d65a2ec9d5 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation MeViS: A large-scale benchmark for video segmentation with motion expressions
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 71fd608c-563d-440c-b48d-e2eed12e28d1 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Progressive multimodal interaction network for referring video object segmentation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f7945b1b-0468-4edc-9cc5-a09043bf2a3c · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Actor and action video segmentation from a sentence
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6ea7e5fe-6e1c-4f56-8abc-7b316ee1d832 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Html: Hybrid temporal-scale mul- timodal learning framework for referring video object seg- mentation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cac50cba-3509-4b3f-8b16-99db2b89610e · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Decoupling static and hier- archical motion perception for referring video segmentation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 140dee2d-5561-4fb6-bc74-668417b2a6b1 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Parameter-efficient transfer learning for nlp
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a87941e8-20ec-4e7b-af0c-19fed35ac66a · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Lora: Low-rank adaptation of large language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 466e9787-6ee4-416e-8d9a-e3eb37500fad · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Temporal context enhanced referring video object seg- mentation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4d7f97a3-867c-4736-be6a-ee39a597c6df · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Cross-Modal Adapter for Vision-Language Retrieval
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84876c47-bd51-41aa-b828-8ade5a3478c0 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Mv-adapter: Multimodal video transfer learning for video text retrieval
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c78c4f96-38b4-452e-9d6a-a4aa8468d4ec · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Video object segmentation with language referring expressions
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 810f7553-7f1d-4e2f-b1ea-f63058485c0d · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Segment any- thing
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 691b5412-517b-47e9-a6ad-431d33919417 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Lisa: Reasoning segmentation via large language model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 48a8285a-b5f5-41a6-b965-d2664864eddd · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Refsam: Efficiently adapting segmenting anything model for referring video object segmentation, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e35b0d9e-14d8-451a-a520-459b362616eb · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Visual instruction tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fc8d0cc-02d8-4c8c-a73b-8d4cf2e4eec5 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Revisiting temporal modeling for clip-based image-to-video knowledge transferring
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a80647c7-d3fb-44d6-a8b5-6da044ddc740 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Li, Ying Shan, and Ge Li
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7d9faebc-ca8d-48b9-a71f-c7623e49190c · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Cross-modal progressive comprehension for referring segmentation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ff93141b-9bf1-4316-b4c7-fb686ec2b4fd · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63124c46-3ff1-4b7b-b075-57a9e6769e26 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d605d979-d476-4a63-ad82-382b00ac0d62 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2acdb8c9-dc74-4c71-a272-c204e037213a · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Soc: Semantic-assisted object cluster for referring video object segmentation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 63715c28-5d61-4434-b223-f66160ec8ead · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Spectrum-guided multi-granularity referring video object segmentation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d7b82916-2165-4d07-a0b9-6308c9769951 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Mod- eling context between objects for referring expression under- standing
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e6de772-05cd-4184-a3f3-6b6ef50cdec5 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Video object segmentation using space-time memory networks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dca92c30-1485-4102-95b3-9485cc440d68 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Keeping your eye on the ball: Tra- jectory attention in video transformers
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b5eedeb7-9430-40d4-aa34-cf8620a0103f · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 010d4f8a-ef35-4387-ad7f-254a69d76852 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Learning transferable visual models from natural language supervi- sion
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04624361-602d-4ab3-8a09-ba9ce2260989 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation SAM 2: Segment Anything in Images and Videos
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f30ed969-f7af-4646-a08d-068f85be13d3 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dbc1fcf-a728-4b8b-b3ab-b0a38ade82d2 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Hi- era: A hierarchical vision transformer without the bells-and- whistles
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 88394d67-f7fd-46ee-af16-acb67bec6cbd · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Urvos: Unified referring video object segmentation network with a large-scale benchmark
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fb185e65-5305-4519-bb2f-12c299858945 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Temporal collection and distribution for referring video object segmentation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e3d4e4-603d-4c1f-8acb-236521fcd79d · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation A multimodal, multi-task adapting frame- work for video action recognition
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 35e7943b-daad-4ccb-9813-16c9273bd9e4 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a83b00-7f14-4ae0-bc63-f10e334fc25a · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation OnlineRefer: A simple online baseline for referring video object segmentation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 84a8429a-e192-4658-9c77-b8d506d2fe95 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Language as queries for referring video object seg- mentation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2993c173-7e0b-4ebb-8249-6dd11ab558c3 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Bridging vision and language en- coders: Parameter-efficient tuning for referring image seg- mentation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd71f7c3-71bd-412e-a9dd-a0f84fe12063 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation VISA: Reasoning Video Object Segmentation via Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 918be086-2a92-4ee7-b03d-36d87db88b6e · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Referred by multi-modality: A unified tem- poral transformer for video object segmentation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7e133695-56fd-42e4-ae69-cc7d29db6e6d · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Mma: Multi-modal adapter for vision-language models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12e28ed4-30b9-4a1c-8d3c-324b36d56156 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Cross-modal self-attention network for referring image seg- mentation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2539c044-6010-45e9-b9be-b0cb02eae43f · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Modeling context in referring expres- sions
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 668c52a9-896a-432b-8f8f-de6e0860adc6 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c6b5c92-fd5e-444c-8136-dda76604c062 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 455a2c65-5df8-4f43-a02f-7be9e89abed4 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation We train our Conditional Mem- ory Encoder (CME) via self-supervision
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b14436c1-85e2-40bb-9806-18330b73a69a · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0d1142a9-4702-4ad2-b0b9-0f0778483173 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation 9, where we plot the memory features
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 534038b6-1136-4aff-98e2-7b9f0ca3c58d · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 38af42ba-b164-422e-9779-3920a460769c · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fc991d70-0bfd-4eb9-8997-ba63c15263f5 · outbound
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation 10, we present qualitative examples from the MeViS dataset that highlight the effectiveness of SAMWISE
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 56a71c8e-6ca6-49e5-baea-48b64b72f8a2 · inbound
ReSurgSAM2: Referring Segment Anything in Surgical Video via Credible Long-term Tracking SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 946344be-ebcb-4169-8c1b-b29664be802d · inbound
BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.