Pith. sign in

Paper Citation Record · LEDGER

VISA: Reasoning Video Object Segmentation via Large Language Models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2407.11325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.11325 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:51:23.416348Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T04:47:33.548752Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9aa97154-862d-4086-b7a4-c2f730a0bf0c · inbound

Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level cites this paper.

Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:46.611013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:46.611013Z digest=sha256:7b84e61854e2a238b1aac424957c10382a4550fe97afdb37192841a326d3ac3e

Observation 9688d33e-2224-49f6-8754-3d2e40f34147 · inbound

HyperSeg: Towards Universal Visual Segmentation with Large Language Model cites this paper.

HyperSeg: Towards Universal Visual Segmentation with Large Language Model VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T12:03:54.218521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:03:54.218521Z digest=sha256:54e8be7144c21d9f1c0730bff7f682dfa96438ab0d390c4500efae5c0d519fc5

Observation cd71f7c3-71bd-412e-a9dd-a0f84fe12063 · inbound

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation cites this paper.

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T11:57:39.770468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:57:39.770468Z digest=sha256:90e878add28e7f5d87322f0d31e711b7e4c98c6fc13d29cffe5a4e40c866e29c

Observation aafbff1c-ae5b-4abe-8544-3ac70e8c391e · inbound

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models cites this paper.

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:37:59.656152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:37:59.656152Z digest=sha256:9cf0af0f4b7b4e381e113e7e517a387ad2f18d631f0956fb4d90f1533f1d932b

Observation 9938afde-aaa0-4c5f-9cfd-741d2486861c · inbound

Efficient Reasoning with Hidden Thinking cites this paper.

Efficient Reasoning with Hidden Thinking VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:47:33.552448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T04:45:38.608009Z digest=sha256:47f386a458858a9d7d274ad77ec3e751c4f346a815a07e7597c4ee968fc2348a

Observation 03755c2a-6af6-4934-a85d-088efd0306c3 · inbound

RVTBench: A Benchmark for Visual Reasoning Tasks cites this paper.

RVTBench: A Benchmark for Visual Reasoning Tasks VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:51:23.416348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:51:23.416348Z digest=sha256:4a8d9ce2dc7cd91be921be1e4eaa41e8706f79e13ea1126b52b0aa8399e482e2

Observation 7dc7196d-1d7d-4a5b-9b90-e1fa8f024dc9 · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.905262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.905262Z digest=sha256:fe4080b504e6f8af4a6bd530799baa07b0c6e8e8c525e29878864cfa7b293b1e

Observation 555d31c1-f140-4e16-a401-0627bf54e0cf · inbound

Contact-Rich and Deformable Foot Modeling for Locomotion Control of the Human Musculoskeletal System cites this paper.

Contact-Rich and Deformable Foot Modeling for Locomotion Control of the Human Musculoskeletal System VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:34.029272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:29:34.029272Z digest=sha256:c5e4ef64065a3305a45a7bbf4edcdcdd3ef903e383bdecb133445f54950378fb

Observation 6ca5da01-d6a6-4697-9abc-24aa8c456706 · inbound

JVLGS: Joint Vision-Language Gas Leak Segmentation cites this paper.

JVLGS: Joint Vision-Language Gas Leak Segmentation VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T16:55:56.992932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:55:56.992932Z digest=sha256:7d5c2d39f46e52f85448139ca931f9dc06c56a341f39ced0888e53fe4f0168a7

Observation 9ea3d715-7a0c-43e1-96a9-31bf0a037463 · inbound

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation cites this paper.

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:31.758998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:09:31.758998Z digest=sha256:2be1b068852f2d68cf9423caaa20c84289e51a94f8719fa76107034191c9b8a1

Observation ccd7c3eb-cf59-4a92-b883-3e796c194a95 · inbound

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track cites this paper.

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.548072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T03:29:12.783993Z digest=sha256:8c1b60037b69b62c7c8293122ca0ddf6e55b4eef3d77be1135772a6504b4e610

Observation eff002f5-1923-4ad0-853c-d032ceecd1ff · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 159

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.277951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.277951Z digest=sha256:7b6b90621f63b952c80647a633237d44e9734863e8c2cc7b3407574df5903258

Observation 326989ef-ee08-4941-a8c3-e4ee0f13e581 · inbound

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding cites this paper.

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T04:55:07.042876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T04:55:07.042876Z digest=sha256:a8e8f1041de3b73d1e5b83dcced7cae91016fe55bd5bc4e98bbfc3cb4fbcae31