Pith. sign in

Paper Citation Record · LEDGER

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

As of 14 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 7 inbound Pith citation observations for arXiv:2506.04277.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04277 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:08:47.365596Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:10:17.325958Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:47:30.631926Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e5f23e84-ba59-44de-bd41-5be8ded3fc81 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Hallucination of Multimodal Large Language Models: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:40.402070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:40.402070Z digest=sha256:e6d4613cd7fbac51db77aeb429bb726dba9d95a5099927873a25455825620016

Observation ba55ffd2-d882-44ca-85b4-7107ee658769 · outbound

This paper cites Reframe Anything: LLM Agent for Open World Video Reframing.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Reframe Anything: LLM Agent for Open World Video Reframing

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:08:48.258624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:40.474413Z digest=sha256:27d5063ce1c94370d2b0421606fc2e5ecbc3e2b5e322aad8d70baa0e776f5ccb

Observation af604303-f121-4e89-bf79-d0a64c592fa4 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:40.556277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:40.556277Z digest=sha256:5abe0b54cfe7da140c366e5c18bfa4b007465e0deb5bcdae502dafaffd2a6e6b

Observation c4bbb67e-1a05-4f59-a731-b4446a15aba3 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:54.729787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:40.681671Z digest=sha256:5aa3e4681e72edbb83fffbce857fbe72e40df8807e503a005b0879d1edb7e584

Observation e46b9ce0-3507-42ed-9fa0-5445fb7a13ac · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:54.514876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:40.787454Z digest=sha256:a511d84c64c1d0d3586c567b0cbd88e92cb6714f0292abbc407a5129a10f817e

Observation 2ec218c6-b22e-4362-881e-4b7d6aff0071 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:54.264144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:40.984131Z digest=sha256:761ea26f51eaa1456cf78d6c31f46f1ab996d5f894638b99413e6a5ad958ce29

Observation 69a6d90d-2fa3-43ed-9d7c-6d693f9152a7 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought LoRA: Low-Rank Adaptation of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.110594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.110594Z digest=sha256:e605748607f72a862630401d15fd07543cdbf748b14f78dbb4b9aab4e28e9906

Observation f75fec81-34f7-44c8-b523-8cb03ab78cd9 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:54.046756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:41.266042Z digest=sha256:6f87dfbc4baec9c80db28404f3203fd57cd52d2a076b828c7bc3e8cbad9fbc41

Observation 332913ec-3213-4238-b852-00678b9d6229 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Doll \'a r, and Ross B.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Berg, Wan-Yen Lo, Piotr Doll \'a r, and Ross B

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:08:53.818075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:41.417931Z digest=sha256:c59b386b10578d39c8431beebc0e7325d15b3d2830d2137a23d31b8d917f015f

Observation 6ca3f72d-c73e-40b3-a3bc-0b126bcfcb13 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Gonzalez, Hao Zhang, and Ion Stoica

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.567535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.567535Z digest=sha256:06e40e8895d3bf73ede5bb73583e2260014b33d81acb33612fdcd9dd55a541a2

Observation b7ba99d1-bc73-4c31-8fa0-5837d3cb84a8 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.796732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.796732Z digest=sha256:f3ee8e380c49737156d8023588bb331b2b748d2fd861b2d3209f2e20ecd13898

Observation 86bf2cb7-94da-403f-8ed4-05db4472013b · outbound

This paper cites VEU-Bench: Towards Comprehensive Understanding of Video Editing.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought VEU-Bench: Towards Comprehensive Understanding of Video Editing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.965687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.965687Z digest=sha256:d46db812ba8da554f496e98c09192d62e1d28c108dfc3786757acaabc2f0abc2

Observation 8aa6eeb3-2245-4d99-8f1c-4cef689c819a · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.177990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.177990Z digest=sha256:0134ed5bcdf57842c707fa27a4042314e1a637cc9cd2b720bd43cc4a21e40143

Observation e32e10fc-095e-4f50-a471-72f556a82b9a · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:53.515400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:42.377897Z digest=sha256:0b6f5bb5c6789741da135546fa133d08d57e5810c2f2cdbb8a8990b7b3f816e4

Observation a1aa0c70-a29b-4b96-b5e4-a09febec9bff · outbound

This paper cites Visual Instruction Tuning.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Visual Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.538130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.538130Z digest=sha256:713fdba37366cb4694dc21aefc98cb35ed8736098f8e5cdbf1f16d0620bc8931

Observation 0b696a43-dd1a-4874-b45b-28cd24fca905 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.665200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.665200Z digest=sha256:a24c70132c85d32e8473a85757876af00753aad77193874e204f3c7eb379d3fd

Observation 02ce7b08-8488-445d-88ce-899b1600cbcb · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.819240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.819240Z digest=sha256:f896084919829ac3407678d8aea741873a95f867e2e4b62066441bcb74717605

Observation d78528e8-4679-4c96-9299-0e685a999408 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:53.235178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:43.047540Z digest=sha256:09e47da9fe660bc71995d607554b4489a5c442b726228420046acebc438cbaac

Observation 38c7cbb9-2171-4374-95fe-3453f32c5ea8 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:53.008973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:43.249635Z digest=sha256:303fb601e01116a258c92e3b5aa01380d85665beb26cee9d51051cc913606eed

Observation 1086a62e-8c4b-4569-b9c5-8c0f98f29315 · outbound

This paper cites GPT-4 Technical Report.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought GPT-4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:43.409280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:43.409280Z digest=sha256:844d1eddde6762cb1d757b7b4e29781d7998fc3a6e9977c6c78f6ca136ef40a2

Observation 07bd95a2-0b61-41ef-8f6f-ef55e92d43b4 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:43.637508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:43.637508Z digest=sha256:2187ced4e45661bc2423346971cc1806e96dc54f6363c86006be8bb783597126

Observation 59c95392-210b-4abb-a5dc-0851f2c3d032 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:52.759769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:43.844069Z digest=sha256:024551bc2ace2d0272c84f081c62945453e8c2b2f25879d607cf0a1a402c28b1

Observation c8cc211c-6e79-45ad-921a-179e08be9277 · outbound

This paper cites TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.046424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.046424Z digest=sha256:7b7b882d9ec16819a6bbd1ce083f7068b17ea36cf260426fbd497a0bdebf70e0

Observation d361bb9e-de9f-4f66-8159-23c9defc968d · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:52.450294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.100873Z digest=sha256:42cce06956104bd113a321e3812c5c610967537e823c8653a4f93db8392f78b7

Observation 666dbb4a-c7d0-4f11-875b-1b49aa22f766 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.190797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.190797Z digest=sha256:a0eab24a020f0a9cfbb8389eca09aec0520582b1b25b47ad0ea4cf495e708719

Observation bc1cc798-6871-4981-ba66-7bedd1d77ff8 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:52.265401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.257827Z digest=sha256:4c02889a486b1b550565cc6b0da3e27b27cd04de6ec61062e0ce708dedd992ba

Observation 28d2689c-4b13-477e-b09d-a028170a9f29 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.329476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.329476Z digest=sha256:0836a44e0f3e0c8f930387baa707ed5eb1866711639586e674836d1cd3522b8a

Observation f553a1ed-58f5-4e64-ade8-c85ee0a09b8d · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:52.024786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.414812Z digest=sha256:7d1223e7702f35be77dd9b262752e40449ff8b5989d9493c1e3558de837f7940

Observation 37da8636-f985-4409-96fb-2f399c848057 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.516389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.516389Z digest=sha256:701ed2c8dda444d187fe1378359ed71c67949b96a5334aee7847ed3df1de1e59

Observation 52cb0a5b-b851-4a36-955c-2efaf3e4ed51 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Gemini: A Family of Highly Capable Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.644702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.644702Z digest=sha256:b3ba0aa8a3004d8decd50e6ee71ed95521c133abb6dc2b80c0f8869cc3516a0b

Observation 1892add0-d767-490e-97af-c7f2addd3e02 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:51.782181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.764286Z digest=sha256:dab9593b64beb55896d475b1d30029be04ac7a9ff96f6f0b332148224c0366ca

Observation e36f32f9-6776-45a0-ae8e-a42c8fc583af · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.852833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.852833Z digest=sha256:f542db532025c51c7f4269207af32b40a8b30bf0845dfad075e419cc15235c17

Observation 65624622-d2fc-4af3-a3c5-ff048ec6b058 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:51.565953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.994367Z digest=sha256:e4b2aea71a4e3fd68c9747a22c02d593558681f6a321d76df42f1781bfe3f1bf

Observation 07a56731-bf64-4a31-adb2-3fa73e7968cd · outbound

This paper cites Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.066926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.066926Z digest=sha256:15cb826a23340b2c7aefe5c104fca1e273f7ab9f0bf9d00f7a4804eb028b37a5

Observation 6d5f953e-df5d-464b-891d-3f3984db2244 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.197404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.197404Z digest=sha256:fef7e8f6b06965f8f59200158269efea40b633235df45ffa665f95d8a367ba44

Observation 9aa33499-ec22-44ca-8371-05a812bd4fe2 · outbound

This paper cites The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.273322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.273322Z digest=sha256:c39f03b4eb02ecf382ff0a24f58f97dfb863b2f912fc3e37c4fa32c29126a9b5

Observation b54b7879-cd9e-4278-a8fd-3b4d1c1f13f1 · outbound

This paper cites DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.339504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.339504Z digest=sha256:136b7c0826f631df583e4187faa2874793aabd5cba632bdfadfa08f38659e4e0

Observation e62bb179-dca9-4451-af74-35c82c258d88 · outbound

This paper cites Number it: Temporal Grounding Videos like Flipping Manga.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Number it: Temporal Grounding Videos like Flipping Manga

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.397760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.397760Z digest=sha256:a515631cf00afbed6c893202c59678c48f6d5a8240affa8f61cfeadd78f3dba2

Observation 0cdca2a2-1491-42a1-84a2-f137d4b60ea8 · outbound

This paper cites Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:08:47.792006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.459689Z digest=sha256:c5f532949191d5789e6f85817268a487a1007f77813b827aeec89f572515cb6f

Observation 51cb02f5-cad9-44dd-bbc9-01845a6243e2 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:51.301911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.523531Z digest=sha256:dd0eea468a3476c8e0c47f48ffa54101c3da330f63d9dca210eb8d1e9610b229

Observation a2e3dec0-6585-4e5c-8a44-a33bd8395fcc · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:51.086971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.602039Z digest=sha256:c6664b9a526a277f70ccf8c3747a24469b1c41530f4f8acc3632060762ff1f36

Observation 7d3edb90-745d-4e70-ad13-949b3b486d8d · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.712931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.712931Z digest=sha256:90a57f1e453532904faa8a3fd130eca1902dbdffb97ae9a53e719d7b537cab8c

Observation a5ccdc74-a977-4c00-bfe6-279859639c39 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:50.838207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.811259Z digest=sha256:9fb2563304a987d6e5fb3a99a6cb81c22a45afe8e22c286f1841b1e6a436b13c

Observation 6e8370fa-e4b9-496a-83dd-5abf89a01f1a · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:50.591962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.874839Z digest=sha256:f8a04aec9abd422714017eaa9363f85e62f81aa1520f1b3071ad3dac2cf9e742

Observation 74d27c5d-3e66-4841-823a-d5a35c83a16f · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:50.368405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.955480Z digest=sha256:a998ccf9c3bab4a4d3410753865c2a517db10b3910d9a1948694754ed1bc1878

Observation f6460ecc-2516-46b6-8951-e023a26d0931 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:46.024434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:46.024434Z digest=sha256:7cc38e4132605c07fec50fb0adfd418b414a4639560bce845a0eecb786983cef

Observation ce95bdd1-bec3-4e60-a458-dfa9cc0aaec8 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:50.131375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.080861Z digest=sha256:54b3c8724f6add8b4af78bc011066c5fd2707747878eff4dd315f8905ebe235b

Observation 1af5bb6a-a3d7-4a4d-b806-8c0e4ecfae47 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:49.859512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.216434Z digest=sha256:52d85527726df00e45c8e4cfe48c0a2aa60b8fefce0ba368053784dbda033b00

Observation 7f8f49d4-b716-4bba-9456-e474910fbc0c · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:49.617597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.357132Z digest=sha256:cb944ee9ccdc17e1c6e2221fad74f71015f66154dd989337256bc08149dba851

Observation 30762b84-bfb6-40ed-9b56-44f745255e11 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:49.407536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.504809Z digest=sha256:d7996227871ed8a92dfc664bc95507d7554dc96e480ffe616c525aca364d5e15

Observation b0b74032-a9cf-4e96-b770-10ae30cdb709 · outbound

This paper cites Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:46.629692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:46.629692Z digest=sha256:f279e1beeef18b387268cd8e3780dcb99a7699c7c1a5f28cd05bc9996ed0cc68

Observation ad0ceb4c-2c7f-4b19-9b02-a574aa1b35a1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:46.785107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:46.785107Z digest=sha256:660e942db509ba190d3eed782224c777f04b212629f0f215545cab423bec4aa6

Observation 30acb137-6da1-4085-9efb-e2691c641a7f · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:49.214546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.880493Z digest=sha256:7c8c3efe916960f03dc6a50643f032bc714514403296929c9941b7f4497e9b40

Observation 3f162f72-5613-42d6-a8d9-85c338a50d2f · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:48.965251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:47.023203Z digest=sha256:5d53c417fe4907ffd08955861cb6e3304b6641a4481c2c1b81d04457fd4971fd

Observation 584af511-db2a-4278-891a-eaf7fb72764c · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:48.633512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:08:47.113216Z digest=sha256:8881d9428f87895b90c5fd79d8ba1ab6a78df5a367497b0fd93313b228e33476

Observation 45e7b5e3-81bc-4183-babd-a1581b8e9909 · outbound

This paper cites online" 'onlinestring :=.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought online" 'onlinestring :=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:47.214335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:47.214335Z digest=sha256:caff8e8b4aaa87d7f0458df812bdeadd6c4c4abf75c499c2f2f490f1a9d715d5

Observation b80bce5c-b602-4db9-862e-9dbc50ccd111 · outbound

This paper cites write newline.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought write newline

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:47.365596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:47.365596Z digest=sha256:add21ffbb92d49dd7e5928de3a2bc47d1446c24f47d7e15c1e8070b1ff0b2c07

Pith citing papers

Observation 0ffd6ff2-f54f-4d8a-9d37-5055ffd9f570 · inbound

SAM 3: Segment Anything with Concepts cites this paper.

SAM 3: Segment Anything with Concepts RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.519580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:14d9bfd4551a9517f9e72fbee53aff503a46c2ab54f3c885d20c0ca0f4d96f5c

Observation 49641d1b-4775-40c6-a75b-a7368f48152b · inbound

A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding cites this paper.

A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T14:10:17.325958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:10:17.325958Z digest=sha256:dc0a1efcc1b3c6dd444b310fab4585e03fb9d7c38882f806bc15ded74aacc36c

Observation f7e37821-8e46-41e3-ba15-3eaf6e4791dc · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:16:34.528679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:12:46.385646Z digest=sha256:43eaabfa521b6b9ab90fad61bbb71e3494ad7b0b52ee9a44473cce7ec3373b85

Observation f33bb049-8095-4067-98f4-d6b7bc5feac5 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.098104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T10:30:06.829915Z digest=sha256:a4d59b8baf7bfd5d201e2777316fb6306e9898b8171e6fe519edd57405a84537

Observation f5f8ffce-3e0a-4b13-a6cb-a408093d7064 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T22:00:07.204123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:00:07.204123Z digest=sha256:780294ed65e6b9da2a96d49957d60e538cd86bd9020b4bf9075be3c768020bba

Observation 97b80294-61cb-429e-b5e1-b3d601a20faa · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.431402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:2482476157e41fa8a8d315bddac1723f6a65ff5304ab638a79d41fe3fcb693db

Observation 5e176c91-e057-457a-99bc-1fa61be974c2 · inbound

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning cites this paper.

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:47:30.633469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T17:01:13.745646Z digest=sha256:83b9600a341b0e476d28b9447008f1c07b2826322bf0679d9fca4b1411f3e17e