Pith. sign in

Paper Citation Record · LEDGER

Referring Video Object Segmentation via Language-aligned Track Selection

As of 22 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2412.01136.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01136 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:42:30.341132Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:08:16.065856Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T05:09:31.997736Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72b3f13b-d469-49cc-af3f-03b302024ed0 · outbound

This paper cites GPT-4 Technical Report.

Referring Video Object Segmentation via Language-aligned Track Selection GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.187102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.187102Z digest=sha256:db81f69aa4458dd766643bfb3180f70e9b3173dc0855be3c80ce4df48062586e

Observation 7aedcaaa-2530-4ba2-8bfd-3eca2b38a316 · outbound

This paper cites End-to-end referring video object segmentation with mul- timodal transformers.

Referring Video Object Segmentation via Language-aligned Track Selection End-to-end referring video object segmentation with mul- timodal transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:31.055324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.192902Z digest=sha256:73937ca2b679776ca04d67637d1605061eb8dbf1d533fabb2d8d4648bb3bda25

Observation 2e5a6799-73e8-4236-97bb-42b1e1f4627f · outbound

This paper cites Vision-language transformer and query generation for refer- ring segmentation.

Referring Video Object Segmentation via Language-aligned Track Selection Vision-language transformer and query generation for refer- ring segmentation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:31.039986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.198030Z digest=sha256:56555230c757bb110f9ddf2b049b6357d56e535f614539b43ab7d352625adc97

Observation 021f977b-b978-4745-96c1-a0bf0c45470d · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

Referring Video Object Segmentation via Language-aligned Track Selection Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:31.023971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.203375Z digest=sha256:9b001d830f7bb1885f0043542d9e2de51699e7756f34048c4c50720e49b7098f

Observation e8af9f4e-02b6-4310-813e-d1c9fe1ff8ca · outbound

This paper cites Language-bridged spatial-temporal interaction for referring video object segmentation.

Referring Video Object Segmentation via Language-aligned Track Selection Language-bridged spatial-temporal interaction for referring video object segmentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:31.003957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.208440Z digest=sha256:59ea8516cf365ec74698a6f33d863f2de6a7848efbc3615ea8c94df51154ce84

Observation ee750bb6-3a96-4573-bdf1-84b2f0fed86c · outbound

This paper cites The pascal visual object classes (voc) challenge.

Referring Video Object Segmentation via Language-aligned Track Selection The pascal visual object classes (voc) challenge

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.985638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.213445Z digest=sha256:6da14eeb0effe6a2d4ebbc04f1b06485e3d1adf652af68e9968f04198d129ae5

Observation 2e25aaf5-ba70-495e-9386-7bcf1ceb2269 · outbound

This paper cites Actor and action video segmentation from a sen- tence.

Referring Video Object Segmentation via Language-aligned Track Selection Actor and action video segmentation from a sen- tence

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.969164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.220559Z digest=sha256:29e305dbe2b45783ccc7b826e9a482a5e0747dd75ab9c833faf0d2d0e73275bd

Observation ce2302a8-491a-4ef8-8d50-72b4a0c643a7 · outbound

This paper cites Html: Hybrid temporal-scale mul- timodal learning framework for referring video object seg- mentation.

Referring Video Object Segmentation via Language-aligned Track Selection Html: Hybrid temporal-scale mul- timodal learning framework for referring video object seg- mentation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.951065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.228708Z digest=sha256:57cba048a4be53ecdd5776c8b8414a70f1a7d2cd39015b32cc02aa328d2905e8

Observation 5ba59491-4c49-46a1-8788-8e9f46a22f65 · outbound

This paper cites Decoupling static and hier- archical motion perception for referring video segmentation.

Referring Video Object Segmentation via Language-aligned Track Selection Decoupling static and hier- archical motion perception for referring video segmentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.934513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.233906Z digest=sha256:737b0fe3d71dbe471aa6150bd7162984abb558376186c026d75d366728dc17e8

Observation 72a94c27-7bf6-4408-8487-2d83cf546cdf · outbound

This paper cites Benchmarking Neural Network Robustness to Common Corruptions and Perturbations.

Referring Video Object Segmentation via Language-aligned Track Selection Benchmarking Neural Network Robustness to Common Corruptions and Perturbations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.238720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.238720Z digest=sha256:f68e770c7d7c0d04dfb75da96f65fa806ebbae611a69a5048dca5f77e152a785

Observation 41c97bf8-dcb4-4d16-ac06-e2152538a441 · outbound

This paper cites Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation.

Referring Video Object Segmentation via Language-aligned Track Selection Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.243704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.243704Z digest=sha256:88d85ad65b242510b3c7ab805f8d8842788bdd901317512e8a657f6411fcc123

Observation 4b34fb5a-31c7-4a10-a5f1-1511cd520211 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Referring Video Object Segmentation via Language-aligned Track Selection Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.915946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.249752Z digest=sha256:eb4420bc7249206c4b7bba3feea99669ed6f79b4e3d742f3c4620dc4ba04c0e9

Observation 19fee968-32f3-468c-a432-a56361293f7d · outbound

This paper cites Video object segmentation with language referring expressions.

Referring Video Object Segmentation via Language-aligned Track Selection Video object segmentation with language referring expressions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.899243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.254626Z digest=sha256:e63e309818f412953cd02c9c22c0371a4c1fc7526eab989435754fc5d0a2759c

Observation 5c37430b-569c-4528-b8ca-035f1ecc81c4 · outbound

This paper cites Video object segmentation with language referring expressions.

Referring Video Object Segmentation via Language-aligned Track Selection Video object segmentation with language referring expressions

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.881871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.259230Z digest=sha256:932b168a372d582e6263d18ac29d7807520b6e9bb573b6a9f025a1a1377b43f5

Observation a9ab9048-f549-47f7-aca3-3a4c9a07bc0a · outbound

This paper cites Segment any- thing.

Referring Video Object Segmentation via Language-aligned Track Selection Segment any- thing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.264270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.264270Z digest=sha256:94bd18c3875c1c9beb9d5a06b3b1a5066eaf6140ed032564894de301943176dc

Observation 2102bee9-3c31-493a-a017-c3540bfab65e · outbound

This paper cites Referring image seg- mentation via recurrent refinement networks.

Referring Video Object Segmentation via Language-aligned Track Selection Referring image seg- mentation via recurrent refinement networks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.855992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.268630Z digest=sha256:d489170f32d306e732fac92e61969583f0c038229eede732b6c04df0df14ca7f

Observation c5ebdb6e-1be1-4721-acbe-1f04f1af376b · outbound

This paper cites Robust referring video object segmentation with cyclic structural consensus.

Referring Video Object Segmentation via Language-aligned Track Selection Robust referring video object segmentation with cyclic structural consensus

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.838830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.273139Z digest=sha256:a46350944c43f604911dbf4dd1cbbbc72e4d78586844836d2143c0d8765e02b1

Observation c5366fd6-a2d9-4667-8ec4-f8cfbac51261 · outbound

This paper cites RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation.

Referring Video Object Segmentation via Language-aligned Track Selection RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.278910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.278910Z digest=sha256:a265d2fff45595b720ad6faf29d583098fe477517d907140ab7ccc2d28927f35

Observation c146b9a3-2a06-4a09-b7ce-677443afd363 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Referring Video Object Segmentation via Language-aligned Track Selection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.285147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.285147Z digest=sha256:1fcda86d774bcb94e412a3e9b3834687e72ee8701163c6156690ca435c1bddde

Observation 70c4a044-c2b8-473a-85f6-7cd0364320c3 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Referring Video Object Segmentation via Language-aligned Track Selection RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.292818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.292818Z digest=sha256:86f1956a9a87d5520ee9145e1564c4af19a872b651ce24c4bbaa2e6638beb9f6

Observation af574789-eae5-4493-ab62-1d0c2fe1bb50 · outbound

This paper cites Temporally consistent referring video object segmentation with hybrid memory.

Referring Video Object Segmentation via Language-aligned Track Selection Temporally consistent referring video object segmentation with hybrid memory

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.821262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.298412Z digest=sha256:9c23205d6e7b2000f8d5350c112b0e0189cadb6d16f60bd0fc4bebf1bc0a40c6

Observation 74a8702f-3397-492b-9705-167cb268bcbe · outbound

This paper cites Efficient non- maximum suppression.

Referring Video Object Segmentation via Language-aligned Track Selection Efficient non- maximum suppression

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.801716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.304055Z digest=sha256:aedde0e8cfd8bb0e8c401afcede1d8e1edc03fb32a46873810089d815dc38857

Observation 6f332e0b-e8e8-422c-8616-623c1c5c75c7 · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

Referring Video Object Segmentation via Language-aligned Track Selection A benchmark dataset and evaluation methodology for video object segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.783350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.309190Z digest=sha256:a16044c0800f67b2150a217d3bd30e776eac982c5180d890b91cae8787a9f619

Observation deaf41a7-9c1d-4c4d-90c4-d639b97bbee4 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Referring Video Object Segmentation via Language-aligned Track Selection SAM 2: Segment Anything in Images and Videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.313812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.313812Z digest=sha256:f7602e2580bf27ff7d3002e720fbd2f3654a9845de37df6eab565284faefb7ab

Observation 4a8dcfc2-c299-46a1-9f6d-86faa7caf827 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

Referring Video Object Segmentation via Language-aligned Track Selection Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.763510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.320392Z digest=sha256:6a3de33557d10b81d1b891fac9af27e83be0ff4366aeb3ef4524ce847954e00f

Observation 1617197a-0fdd-4668-afc2-5ef4222d579f · outbound

This paper cites Attention is all you need.

Referring Video Object Segmentation via Language-aligned Track Selection Attention is all you need

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.325968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.325968Z digest=sha256:b69273e51ae93d8948e58cc487d70cc6a631c6525f4fdc82f311ae50045730f1

Observation 77670a91-7112-4f11-89a8-e3b7ef0d0497 · outbound

This paper cites Data-efficient mul- timodal fusion on a single gpu.

Referring Video Object Segmentation via Language-aligned Track Selection Data-efficient mul- timodal fusion on a single gpu

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.735719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.330343Z digest=sha256:5eae4052ba524803eae007fdff4e605ef7ef234f6ef5b73636151d184568374a

Observation 4a000abf-1e12-427a-b2a3-5943585bfade · outbound

This paper cites Language as queries for referring video object segmen- tation.

Referring Video Object Segmentation via Language-aligned Track Selection Language as queries for referring video object segmen- tation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.719106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.336155Z digest=sha256:892cffa4285076c26efa87ec74706a0b718a00a76370679d49a0a7bb11d20d12

Observation c50724b1-ee13-4332-949a-76b6a6232ef8 · outbound

This paper cites Going right.

Referring Video Object Segmentation via Language-aligned Track Selection Going right

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.695296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:42:30.341132Z digest=sha256:3c056a5403e7da931a4386aa0f68ca33f1af10d72f5c2c9d24b6336e8ea540bd

Pith citing papers

Observation ec8a01c8-6981-4e8b-8e64-4b8fa22f65bb · inbound

MOVE: Motion-Guided Few-Shot Video Object Segmentation cites this paper.

MOVE: Motion-Guided Few-Shot Video Object Segmentation Referring Video Object Segmentation via Language-aligned Track Selection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:16.065856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:16.065856Z digest=sha256:c3c371f5c0e50d96bd95d64fb4f5d77b023b27b61dbdec4d171b4f9640c74f8d

Observation cff9ce97-b9ba-401e-82fe-fa7684435c61 · inbound

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation cites this paper.

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation Referring Video Object Segmentation via Language-aligned Track Selection

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:09:32.003729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T05:09:31.671155Z digest=sha256:f4514559fef41698dcb65ef5e3f593cfdfa4b8b9cde22694f7df51e49f67d47d