Pith. sign in

Paper Citation Record · LEDGER

Track Anything: Segment Anything Meets Videos

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2304.11968.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.11968 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:29:00.725533Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

95
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 59e406f9-dff0-4aa7-87ed-8e702d7c6fa3 · inbound

Faster Segment Anything: Towards Lightweight SAM for Mobile Applications cites this paper.

Faster Segment Anything: Towards Lightweight SAM for Mobile Applications Track Anything: Segment Anything Meets Videos

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:41:43.492900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:41:43.411128Z digest=sha256:0c23ffddf4f53557a24d60761602936fd9261b5cd5341f1ea4b9f4b79b3f497c

Observation 222e03e3-3644-4541-849b-0d4cb32b831b · inbound

On Efficient Variants of Segment Anything Model: A Survey cites this paper.

On Efficient Variants of Segment Anything Model: A Survey Track Anything: Segment Anything Meets Videos

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.551644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T19:42:24.122342Z digest=sha256:294e22415840c7232c120f77daa3eb0edf05c4d37d90f39a802695fad6d91082

Observation d2c6632d-85fc-494d-a12a-fb32a701b46e · inbound

RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos cites this paper.

RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos Track Anything: Segment Anything Meets Videos

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T08:02:43.822894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T07:58:54.015749Z digest=sha256:e722f78b671b3354cfa3a717549546eaf755c8e225f6def857f1d5e8af108ac0

Observation 914e81ab-71f6-44d6-a9ed-d965447b7a9e · inbound

SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement cites this paper.

SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement Track Anything: Segment Anything Meets Videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T14:29:00.725533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:29:00.725533Z digest=sha256:fe88945ef2cef73fc7a0cedd432463e66d87bc81ff12f6f708bcc3b0a93c059a

Observation 5efc0986-52ca-47c7-b1d9-10c3274f858e · inbound

CU-Multi: A Dataset for Multi-Robot Data Association cites this paper.

CU-Multi: A Dataset for Multi-Robot Data Association Track Anything: Segment Anything Meets Videos

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:19.959554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:19.959554Z digest=sha256:1ad5252cfe498d5443e25b87defae267dbd6d07b0caafe5de7ec69de13aa07e2

Observation 59f9617f-a5a4-4939-a621-9dcedf8a1970 · inbound

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory cites this paper.

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory Track Anything: Segment Anything Meets Videos

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T12:57:17.913143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T12:54:31.765909Z digest=sha256:230567e86a43f89a39a0edc24b96e5700b7d2cad3816f419ef13c965b8257bee

Observation 9ac3aeba-a833-4a41-8e9e-70ccdfed6daa · inbound

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost cites this paper.

SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost Track Anything: Segment Anything Meets Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:49.056362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:49.056362Z digest=sha256:c8f76e4f7af68ec9b39d305f90520b71b5f02fd7bf854180afc9fbf7f9df3f91

Observation 340f00e3-ff2c-42a8-af1f-05f4ad41d936 · inbound

UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting cites this paper.

UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting Track Anything: Segment Anything Meets Videos

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:16.361815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:16.361815Z digest=sha256:8b445ac0534e92c976b7004e5981858eb24166cb246074d614507ad322f15d0e

Observation 3da692c5-cc8d-489c-a1d9-cee14765e42b · inbound

UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References cites this paper.

UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References Track Anything: Segment Anything Meets Videos

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:06.851144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:06.851144Z digest=sha256:84327cdbb97bdda78039b9f4070d5414e2ed41352aa42b762d448365c123cd5f

Observation a3034808-69f0-46d6-b451-7ef6a9d78fd2 · inbound

R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision cites this paper.

R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision Track Anything: Segment Anything Meets Videos

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:57.411759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:47:57.411759Z digest=sha256:f8a7b7fabfae21660467e5ee2207931df111a6cf42b21a4c1fa4e225ccb5775c

Observation 5962e6ba-592a-4d4b-a0da-765b77fa12db · inbound

ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs cites this paper.

ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs Track Anything: Segment Anything Meets Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:19:48.831105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:19:48.831105Z digest=sha256:2cc3c342a85d770798e8ce5366c91c4931f307e25f24a7c4af7cd19e3ae7a39b

Observation 68a3a034-4946-4916-8007-372c479b0fe2 · inbound

CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning cites this paper.

CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning Track Anything: Segment Anything Meets Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:37.920077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:37.920077Z digest=sha256:706aa62832d1c2c8b9b30670bb6ca4b19381fcb7fe21042260fe7401ce1d468b

Observation c66b9629-eb4d-4b1e-b7c0-17d197ee9f7f · inbound

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation cites this paper.

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation Track Anything: Segment Anything Meets Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:57.025806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:57.025806Z digest=sha256:1856870be15ad034fcd87cd8b5a136b5be57c5c4f6bb1d606671a0163f0959f7

Observation 9543d116-36d2-4807-8b26-bcfeb0505aa5 · inbound

Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation cites this paper.

Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation Track Anything: Segment Anything Meets Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:46.955280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:46.955280Z digest=sha256:672bf69ebd850a6fc9495e85c635bbc7b4bdf89189763a656ed06e17d756b4ba

Observation 33b1a0ff-8793-40ba-b7e6-137eb96c29a9 · inbound

Continuous Marine Tracking via Autonomous UAV Handoff cites this paper.

Continuous Marine Tracking via Autonomous UAV Handoff Track Anything: Segment Anything Meets Videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:36.662024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:36.662024Z digest=sha256:a80c13f21c67907c8ab732419b4685ab3f2cdced8b36f3702dfd04589769fa75

Observation 4d88af1e-0a19-425c-8225-e696cc2273ec · inbound

SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition cites this paper.

SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition Track Anything: Segment Anything Meets Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:41.282511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:41.282511Z digest=sha256:7988005b64eda62eed1ffb1cd23a16f6e75c170b58827809707143f294a30895

Observation be47a831-ff2e-4435-9969-65cef60567e7 · inbound

3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction cites this paper.

3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction Track Anything: Segment Anything Meets Videos

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T22:22:05.096000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:22:05.096000Z digest=sha256:b7a3b3a504361d287c341e43a750f90f1b45086a11f1cff0857da1f8c01c2f81

Observation e579c1b4-5e9d-4862-84a0-bcbba4f1976e · inbound

Grouped Speculative Decoding for Autoregressive Image Generation cites this paper.

Grouped Speculative Decoding for Autoregressive Image Generation Track Anything: Segment Anything Meets Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T21:57:10.519532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:57:10.519532Z digest=sha256:9cd319f0f3bc82e3cafd36fd8c5cf387541cc689474feaccda9373cb13ed239c

Observation dc07d712-d243-49bb-9167-63d0ab8cde11 · inbound

VoCap: Video Object Captioning and Segmentation from Any Prompt cites this paper.

VoCap: Video Object Captioning and Segmentation from Any Prompt Track Anything: Segment Anything Meets Videos

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:18.011590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:18.011590Z digest=sha256:292218a08a930d21cbafa21f18d579060b501cb303f5cccd9c53e4865ef204f0

Observation 6ebd9c47-9370-431a-b61d-da4d31740695 · inbound

Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation cites this paper.

Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation Track Anything: Segment Anything Meets Videos

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:56:01.690261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T06:53:21.438159Z digest=sha256:9b9fc1e570b60e27346f2028acc406927c03128a4a3f6fa4b17b67020c8ac6fe

Observation 2cd0e4bc-c030-468d-8127-afb1d449134e · inbound

SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors cites this paper.

SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors Track Anything: Segment Anything Meets Videos

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:11:34.847331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T06:10:24.636998Z digest=sha256:6bf02af2773217331ce2a2471d70b441ba74cd6aebfe3ab3a2ce3876b904652f

Observation 0a4b6bb3-d735-4366-98cd-29d83974c513 · inbound

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents cites this paper.

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents Track Anything: Segment Anything Meets Videos

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:58:31.849678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T20:56:58.770875Z digest=sha256:4157c8cecc07641be91642624febe20154c510c73dfa22573c68f89d637157f9

Observation ed35b158-6198-4fd3-b124-f41879769591 · inbound

Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping cites this paper.

Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping Track Anything: Segment Anything Meets Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.945749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T19:51:58.756605Z digest=sha256:55c19653391e001504fa4ea40751c16ccd4e276d39d1f2217075bba875f8b1a5

Observation 0067fb33-9d7f-4214-83d9-ac1afe7a2e3c · inbound

ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis cites this paper.

ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis Track Anything: Segment Anything Meets Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T18:24:58.428701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:24:58.428701Z digest=sha256:f8308985a9789d76f90c6ec73518dbdf372dd3a2edfb18524faf58bfad6f4224

Observation 97ac1bc3-7562-46b6-a5b3-5d9aec6c6867 · inbound

Learning to Synergize Semantic and Geometric Priors for Limited-Data Wheat Disease Segmentation cites this paper.

Learning to Synergize Semantic and Geometric Priors for Limited-Data Wheat Disease Segmentation Track Anything: Segment Anything Meets Videos

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:05:51.506643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:20:04.690220Z digest=sha256:5e6c7c58db92b4b89dd68dfd41132410f962a4984aa54c89cc915ab14c23c6b2

Observation 8d1f8330-4715-4b6e-8022-fd92d7f194f9 · inbound

DP-DeGauss: Dynamic Probabilistic Gaussian Decomposition for Egocentric 4D Scene Reconstruction cites this paper.

DP-DeGauss: Dynamic Probabilistic Gaussian Decomposition for Egocentric 4D Scene Reconstruction Track Anything: Segment Anything Meets Videos

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:00.941092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:20:54.028846Z digest=sha256:42b3a08b35336b1d09c3768c54a7fb35b3ff76b1afcb30627b9a439c55fd40d6

Observation 09bf1170-c3c3-41d3-875f-f9003e323fbb · inbound

Do Instance Priors Help Weakly Supervised Semantic Segmentation? cites this paper.

Do Instance Priors Help Weakly Supervised Semantic Segmentation? Track Anything: Segment Anything Meets Videos

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:03.251934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:23:15.172519Z digest=sha256:b2bd416731daa3a4e84d5dba6a4e88ffd6d52fe4a824feb29f506d9eb0c73de4

Observation 9f126d71-9aa3-43ad-b798-6358d6c3dcea · inbound

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition cites this paper.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Track Anything: Segment Anything Meets Videos

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:11.879222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:8760054c9b832f0da6444532c9f1e8f884ffb7466aa7d48f241cab7ae5c427bb

Observation 27a0bc2a-a6db-4545-9e7c-e500ce02400a · inbound

VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement cites this paper.

VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement Track Anything: Segment Anything Meets Videos

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:11:20.977325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:09:34.131634Z digest=sha256:6a5c9040885dac6c0117cef2e3fe0330c658f5fac9f8a5fc59151608447febf0

Observation fe5c1906-bfee-4183-a40a-62fd8c924c85 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models Track Anything: Segment Anything Meets Videos

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:26.687252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:394d7a422b438a6e98089d161327efc0bcfaba48b73a2d8a7c659115e291d114

Observation f6be40f8-8ff4-4b30-8f4d-dee55b3d97e6 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models Track Anything: Segment Anything Meets Videos

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:57:28.258081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:9eff67b96a91b4d28b8d4303a9e57c2c7dada5e875972bb9bc0280f5073c9020

Observation f19ecb62-6f13-4fc1-b92c-ab6e4b28bd36 · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models Track Anything: Segment Anything Meets Videos

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:36.587306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:6385b0279862ffc5a28c2fb4273daf8914e3d23cc639e8f768738f2065ddcae8