Pith. sign in

Paper Citation Record · LEDGER

ViLLa: Video Reasoning Segmentation with Large Language Model

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2407.14500.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.14500 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:34:25.104905Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.216469Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ab1af148-e5f5-45e1-86d7-1042694e9af6 · inbound

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects cites this paper.

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:25.104905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:25.104905Z digest=sha256:8de669b89422fb9f10d5d8af51c658a0251cf8f742bb13d370bec2a1bf10f89f

Observation d69fd75a-9aa7-4bd6-928c-b428aaaf0e05 · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.326601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.326601Z digest=sha256:08420de81ba7f8086bf1dde19df11e86c7e4b1d0ffee6389473dc825d5b21a2b

Observation 51c74104-c7d0-4941-bc6f-b74ebc17b510 · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:56.438861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:56.438861Z digest=sha256:740171625371b51b1d2dedd91b1c1f2428aaa635dc8f5b973579608a46710626

Observation f35ef5b4-7b2b-4406-9312-0d74b021a834 · inbound

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation cites this paper.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.460552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.460552Z digest=sha256:31cfb040ba876d4e91e05a12c1f83cc4ebfc33931b6f374af0d5bcc888d5d00a

Observation dd204a70-d312-42b5-b8a4-8b8f3464d726 · inbound

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation cites this paper.

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:31.781108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:09:31.781108Z digest=sha256:8721d11e961a3005ac69dc25d570bcccd11f3c5b22ef7bb08681756c29c61ff6

Observation b829bba1-a116-4cf6-a6c0-05e336f5e66c · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:08.070769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:446a165afa2720263c4507f36e7b6f1694fd4035a32123a9f1caa48cc380351d

Observation 9752265a-1874-4753-a54e-5512d78a2de7 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.143330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T09:20:54.635375Z digest=sha256:b6cd10e35e7ab452aee321eb6faa359bb84787ef9880bcc39ec3f09ff15eb652

Observation 41eec131-2845-444a-84fa-8165ba0eec00 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:02.936737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:02.936737Z digest=sha256:5167a39ef55311cd947fa6422b0166655cb8a98ee14ecd2e086f0bc0fdc8de1c

Observation 24c37d64-4d13-4d55-bf73-140cbe1b3500 · inbound

Weakly-Supervised Referring Video Object Segmentation through Text Supervision cites this paper.

Weakly-Supervised Referring Video Object Segmentation through Text Supervision ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:56:11.089619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:54:33.227975Z digest=sha256:f9be4dadb1b42c1a6b7b9c001df4569ff93785520d8350037f40b197e8a07fcc

Observation 88f36ba8-30bd-4fa6-9265-6d15d5212724 · inbound

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track cites this paper.

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.565364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T03:29:12.783993Z digest=sha256:e142a33f484d4b83e87132a3f958ea1efd650095ca672f33c9a7910f8e0a36f9

Observation 3b99a5eb-097e-40f2-838c-751e1f2a28ae · inbound

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method cites this paper.

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:33:41.857762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:14:25.423302Z digest=sha256:b841cde55d1597cb918d44af5652994cbbcfc391e0132e58c8684379340241bf

Observation 6e2118be-67cd-429d-b53a-d9e35271a891 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:59.882045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:d5b3ed54d2f036872e57e7671b563422c8a25ee988aa489b6d7effc44e1ce19a

Observation 7c45d39c-2791-40cc-8f92-915504edf8e5 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.217966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:f8a205501440610ad46f56244fe1d2b57eb36405c39b8c4b486f4813a5ac5d35