Pith. sign in

Paper Citation Record · LEDGER

EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2406.20076.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.20076 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:59:29.719733Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T08:06:57.772386Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5c63839f-a28c-486d-b7c6-d57edf0e1e06 · inbound

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding cites this paper.

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:25:28.430275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T07:24:01.527093Z digest=sha256:dda0acb1c489b01530edf5fef3effb1bfc457fce93341416bd5e37c851b2e95e

Observation ba6c847b-5d56-4dda-ab62-a23bebb4c965 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.556638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:ca0319a4f6dd8e29e69531dc6d9b9066450a582c35f08c5f097e963226644bc3

Observation 6c734f9e-ea84-4ac8-a2e5-80d9356fc318 · inbound

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping cites this paper.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:29.719733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:29.719733Z digest=sha256:4d5c4827c697acf14d24bb7db83c9814e41f92c92905cb6808d627600dd83eee

Observation cb421d72-2444-4831-b123-426b0b3e7b25 · inbound

Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning cites this paper.

Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T22:40:36.054758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:40:36.054758Z digest=sha256:894954f0876e22f743ed876a7f7e42c678982f90979cedd13d726f17918cf5ea

Observation 62c7771d-d80f-4243-ae70-dab1685d75c1 · inbound

CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation cites this paper.

CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-06T21:54:43.687107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:54:43.687107Z digest=sha256:5117a3834b5a1e0e9733a2bd1171dfc8f2c7efd364ce018f848b2e40c3df1d5d

Observation f18181be-c505-4aa3-bef9-046aac89f670 · inbound

SAM 3: Segment Anything with Concepts cites this paper.

SAM 3: Segment Anything with Concepts EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.420980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:487e29fb28cc0cbab82413b57ec1bab29bb60b1d3bc0d9e4ee3830d8f702deed

Observation 276b06cd-f6f2-4a4f-a6f1-143db866ae1c · inbound

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation cites this paper.

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:21:01.008627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:11:13.376684Z digest=sha256:dc67a8a2d5afa56fb66bb3ae371f9c8ab9b4020d67e3d5676375294846c1ec74

Observation 0aaeb7f0-5d2e-466a-96bd-bdc82a654fc6 · inbound

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation cites this paper.

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:43:49.426512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:10:44.608959Z digest=sha256:a4c38020b7c3fbadd2f3198d60a974db9a40f338a9e1e0b3ad5127ff988b7534

Observation ad041107-2c10-4d15-b05d-cf91ba74fd24 · inbound

SIGMA-ASL: Sensor-Integrated Multimodal Dataset for Sign Language Recognition cites this paper.

SIGMA-ASL: Sensor-Integrated Multimodal Dataset for Sign Language Recognition EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:11.352200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T07:04:47.169148Z digest=sha256:41c2dfa30c954a230fb42b3a59ca5984e458add25d1d343d3b3402c2ab662e3f

Observation fe24e2de-87e2-49c2-b47e-0b30bf539301 · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.383299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:bcdb1805c4e06765bd0c2394aa655b72d8d1efb9e002ef0401994bf6046f0ed8

Observation c36a9fb7-f75a-4e05-bf97-86dff3564869 · inbound

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction cites this paper.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:05.240709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:1ec263c8442657f072f52e49a1f81d7224088356cb9615a68473485369d741ac

Observation 66b2baea-c095-4da5-8163-0da46a97be86 · inbound

IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring cites this paper.

IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:34:40.563286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T13:28:01.889147Z digest=sha256:53263ebe31fe2f3fd8cacc28ba3284a74ae56bb065c8d5d694b06176985150a6

Observation 0414d691-0d81-40c8-b5b1-36ef31c4b4b9 · inbound

InstructSAM: Segment Any Instance with Any Instructions cites this paper.

InstructSAM: Segment Any Instance with Any Instructions EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:02.029740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:34:00.442420Z digest=sha256:72b28ac4c396cffbf880f932535a921909a3fa86cdc254936084de06cab6033c

Observation 94431ffb-eb37-4d23-8da3-bd26278c199b · inbound

MAOAM: Unified Object and Material Selection with Vision-Language Models cites this paper.

MAOAM: Unified Object and Material Selection with Vision-Language Models EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 121

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.324793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T11:08:59.900161Z digest=sha256:d421294caadcfe4a578af66ff810f73d8d380948af082527c59150c10bcb5cf7

Observation 5523097c-b5eb-4508-b58f-4697d7e5016d · inbound

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation cites this paper.

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-10T08:06:57.773790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T08:06:03.952787Z digest=sha256:e937c43571d9ade0903feb2744cf93ca756ddaf38c7520eff2d80bd02fa0a3db

Observation 843ff5ba-65e2-4102-a554-a43190bc32e3 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:cedcfba82946aca00d5f1cb25b2ee7d28639cbae369dc8ec7dbf2b8f0ece2e15

Observation 77b526f0-3e14-49f8-a679-e6cd4bd18095 · inbound

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation cites this paper.

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T00:48:46.730648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:48:46.730648Z digest=sha256:16506e62bdec1d4131ce9cb3764f08841dce70424a5593ee735c5666e61b9ae4