Pith. sign in

Paper Citation Record · LEDGER

Segment and Track Anything

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2305.06558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.06558 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:34:20.234594Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:13:49.304240Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 99305b33-8589-4c07-9df4-c17ffe475cc8 · inbound

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory cites this paper.

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory Segment and Track Anything

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:58.147228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T13:03:57.828598Z digest=sha256:8e0e52a70502bda702412b850f46364a1756172f5bf0845d48b916ab61601efc

Observation 0ff62b57-678d-45e7-b2a5-9c5b361a04b4 · inbound

SAM 2: Segment Anything in Images and Videos cites this paper.

SAM 2: Segment Anything in Images and Videos Segment and Track Anything

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:56:25.371685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:56:25.331304Z digest=sha256:7d5905a7ec6b886845fecd62a12f0d812e8e4892f957fedeb57b848113cfd50a

Observation aad9540f-c6dd-42f1-a3a4-4ce68d8109d5 · inbound

On Efficient Variants of Segment Anything Model: A Survey cites this paper.

On Efficient Variants of Segment Anything Model: A Survey Segment and Track Anything

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.490908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:42:24.122342Z digest=sha256:3aa1a42d7e391e0ed7346f87928369d88be5c4fa66a1d8d964afa90962e9cc35

Observation dbfacbdb-dc75-4c8f-93df-277c442ea72b · inbound

Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement cites this paper.

Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement Segment and Track Anything

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T08:25:29.380265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T08:25:01.468957Z digest=sha256:4419571d8b44847906b914037b28c2542961378ff8a609225bafb139d98c5999

Observation 736f90f5-e8ac-4755-8c8e-9cca77784074 · inbound

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects cites this paper.

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects Segment and Track Anything

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:20.234594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:20.234594Z digest=sha256:8cbae7109669471a6a437afd918081f853adb02d6b7454c8bd616b2c9b74d803

Observation 9fb9494b-f604-4bec-bb6a-b100b8452b68 · inbound

STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing cites this paper.

STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing Segment and Track Anything

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:20.179134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:20.179134Z digest=sha256:9603249078689688ea995f016615745e4a52e69e4427d0af309c509c788c750f

Observation ee9369df-9b15-406f-94ae-1f7bc5d66af0 · inbound

CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation cites this paper.

CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation Segment and Track Anything

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:32.123934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:32.123934Z digest=sha256:fec2050512d5a63ac20da7ac11e793923d2dc9266c3ed3fab8d9a7a621ffa1c3

Observation fe661157-e11b-41fb-83bf-d5599062b8c6 · inbound

ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation cites this paper.

ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation Segment and Track Anything

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:40:06.711380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:40:06.711380Z digest=sha256:da16c862a2001d1677e5069555a79f129dfc07b8618b9912ced933da5ee65952

Observation 0dc380bd-5ec3-4a86-9b19-408a2abb5427 · inbound

CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios cites this paper.

CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios Segment and Track Anything

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:50.844918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:31:50.844918Z digest=sha256:f4a3f7ade5fc6aac11dac8f3396cbd7fae147196d14ad21aeb0957f65cb2e51f

Observation 9cdf3218-26e9-44f6-b31c-a370dfcdc5a4 · inbound

High-fidelity 3D Gaussian Inpainting: preserving multi-view consistency and photorealistic details cites this paper.

High-fidelity 3D Gaussian Inpainting: preserving multi-view consistency and photorealistic details Segment and Track Anything

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:58.782053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:44:58.782053Z digest=sha256:210e3f1ae9c89e5607a075b0f3b2b96c5d117564fe26218221b36c8c8a59b148

Observation e23316d3-68b9-4291-b465-fdcda0340cb3 · inbound

Grouped Speculative Decoding for Autoregressive Image Generation cites this paper.

Grouped Speculative Decoding for Autoregressive Image Generation Segment and Track Anything

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:09.363139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:57:09.363139Z digest=sha256:bca835d0facff783b420458613d748afe2cf1deb1b7986a0583e4912226730ee

Observation 538a1d52-85c3-4e25-be6c-d8fa0183b243 · inbound

Representative Volume Element: Existence and Extent in Cracked Heterogeneous Medium cites this paper.

Representative Volume Element: Existence and Extent in Cracked Heterogeneous Medium Segment and Track Anything

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:29:52.721427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:29:52.721427Z digest=sha256:d57fbbf1ce9c186f7596aa88fef80e79cb2355b22a59fe0953f347c7907b6f8d

Observation f9e0e54b-bce0-41ea-a5b5-a435ff99900e · inbound

ViPE: Video Pose Engine for 3D Geometric Perception cites this paper.

ViPE: Video Pose Engine for 3D Geometric Perception Segment and Track Anything

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:41:08.664360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T16:41:08.620285Z digest=sha256:c77eb4c61dda540fb88a6353356d3fbcf14be9e6ed391ee8c0f715c8420596cb

Observation bb55ce6e-db0c-41a4-9ca2-2c958af88378 · inbound

DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing cites this paper.

DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing Segment and Track Anything

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T18:34:23.512742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:34:23.512742Z digest=sha256:8e5ebec131a070faf2b9e59efbf98f65695174613af0d75f855fd3a03329f21c

Observation d3d34e96-2a6a-4207-a531-14fd3f85fee9 · inbound

VoCap: Video Object Captioning and Segmentation from Any Prompt cites this paper.

VoCap: Video Object Captioning and Segmentation from Any Prompt Segment and Track Anything

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:16.517409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:16.517409Z digest=sha256:22ab02d65caebdb11bed559751ee3c767e539713a32b48851a3d2a4f9d0b0c84

Observation 22ef0701-514c-41be-a049-6a883b793cc2 · inbound

Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control cites this paper.

Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control Segment and Track Anything

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:11.934372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:11.934372Z digest=sha256:451ff85b9f116f6ffa862a463d9851964f527f54c26414e8660442f78e3e5b7a

Observation 31a06d06-832b-4879-b307-bd7cd3992c96 · inbound

Grasp Like Humans: Learning Generalizable Multi-Fingered Grasping from Human Proprioceptive Sensorimotor Integration cites this paper.

Grasp Like Humans: Learning Generalizable Multi-Fingered Grasping from Human Proprioceptive Sensorimotor Integration Segment and Track Anything

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T20:44:40.124432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:44:40.124432Z digest=sha256:bf999e315cbdfa7dc4d04930cde3dc013ad968505dad012ce411708f38e076d2

Observation 51f54a4b-81d8-4b67-b2c5-cbcaa0599749 · inbound

ViSTR-GP: Online Cyberattack Detection via Vision-to-State Tensor Regression and Gaussian Processes in Automated Robotic Operations cites this paper.

ViSTR-GP: Online Cyberattack Detection via Vision-to-State Tensor Regression and Gaussian Processes in Automated Robotic Operations Segment and Track Anything

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T17:25:30.813136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:25:30.813136Z digest=sha256:087a345004b49a648f463d0a1be511af39ca49d34a67222c949710b57e9795de

Observation 8f735f20-1c5b-47cb-8360-7803ec16e0e9 · inbound

Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation cites this paper.

Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation Segment and Track Anything

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:56:01.667572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T06:53:21.438159Z digest=sha256:16d76aec78a3e37b1367594e0b4d09299c23475f224500932269271b48d2631b

Observation 638f907f-54be-4905-87a0-ba068f0665b0 · inbound

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents cites this paper.

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents Segment and Track Anything

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:58:31.814788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:56:58.770875Z digest=sha256:3a75bb2e693aaf744c7b6667eec309d990e04783d229a6d81b4833969bcd6234

Observation c67fe298-7b32-47f0-aff3-30da96baf5c1 · inbound

Efficient Segment Anything with Depth-Aware Fusion and Limited Training Data cites this paper.

Efficient Segment Anything with Depth-Aware Fusion and Limited Training Data Segment and Track Anything

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:11.238949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:11.238949Z digest=sha256:a45a97654faa98593db13864dc07783bd05ecc768681816694035ada8913cc7a

Observation 8ec5f0ab-ef26-47f4-8ee7-153c2e331ded · inbound

ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis cites this paper.

ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis Segment and Track Anything

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T18:24:58.428701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:24:58.428701Z digest=sha256:bd771d5cd6b7946848e8b1fef7665e6f4ca96bff75dc31822f09c556025b992f

Observation 521aac84-4e08-43cf-ad2c-a24e5a4ea3b5 · inbound

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models cites this paper.

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models Segment and Track Anything

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:53.007612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:39:26.420463Z digest=sha256:c4598a099f10c7a7177564a6864412cf82b22c1dbb1144bc8bf69c867527d202

Observation b755e582-44c9-4a1b-85f9-40414dcdad08 · inbound

AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking cites this paper.

AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking Segment and Track Anything

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:34:47.308672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T00:34:26.106672Z digest=sha256:d4bb084713edbbc8121361c8b265c69569a180f4dbff77ef17bb932a0b59e8de

Observation 119424ae-d507-4a02-a73c-9eb699106c1d · inbound

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition cites this paper.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Segment and Track Anything

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:11.807561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:5a8536c4be59686c225b87f06f8277074a77e8bd27bb166b8c327db182e3fa51

Observation 456f231d-bb91-41ed-b7cc-04ac797e0193 · inbound

Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection cites this paper.

Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection Segment and Track Anything

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:13:49.306052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-29T05:13:51.469018Z digest=sha256:f2e23128c5b5f01c966452cb906e771c15c3e70891791c241cfabff80c202e24

Observation bfbb65b8-fc5d-4102-8d79-1dc80c1f94fd · inbound

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation cites this paper.

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation Segment and Track Anything

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:48:38.659879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:48:38.659879Z digest=sha256:3cc8c7b89176cf03d4db2c56e85843f2acc63a0c35119957624da36a3bd20162