Pith. sign in

Paper Citation Record · LEDGER

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

As of 9 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 7 inbound Pith citation observations for arXiv:2506.04277.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04277 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:08:47.365596Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:10:17.325958Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:47:30.631926Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e5f23e84-ba59-44de-bd41-5be8ded3fc81 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Hallucination of Multimodal Large Language Models: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:40.402070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:40.402070Z digest=sha256:ade2f03b9d992a61944911e10b8ec7b56b667a9b499f7b9ec21f5d1b761c36ff

Observation ba55ffd2-d882-44ca-85b4-7107ee658769 · outbound

This paper cites Reframe Anything: LLM Agent for Open World Video Reframing.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Reframe Anything: LLM Agent for Open World Video Reframing

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:08:48.258624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:40.474413Z digest=sha256:b2524d8adcd74f0a7fe90a967f447219b27dc158c139493eb1b9d51d4cd52f07

Observation af604303-f121-4e89-bf79-d0a64c592fa4 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:40.556277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:40.556277Z digest=sha256:53e2f876016c8fe3752c557302916625ce960d9b7e25fc7279fb6bd61b2608f8

Observation c4bbb67e-1a05-4f59-a731-b4446a15aba3 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:54.729787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:40.681671Z digest=sha256:f975c0b8812e10e8a47733bbf25bceb172ca445ef0c3b38e3cc06a8fc4412688

Observation e46b9ce0-3507-42ed-9fa0-5445fb7a13ac · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:54.514876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:40.787454Z digest=sha256:ec2b493f6feec7c726ccf825055c6f6efd6b475e5884bdf1c9a13a4c0e7fafaf

Observation 2ec218c6-b22e-4362-881e-4b7d6aff0071 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:54.264144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:40.984131Z digest=sha256:c4285aa00359125f9ffd25ede97c550d0293ab523eaed5083a6b373108a47e73

Observation 69a6d90d-2fa3-43ed-9d7c-6d693f9152a7 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought LoRA: Low-Rank Adaptation of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.110594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.110594Z digest=sha256:b73345aa371cc6ef4a2e17aa958808fe78dabb8f456af4253305ccda3ccd060e

Observation f75fec81-34f7-44c8-b523-8cb03ab78cd9 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:54.046756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:41.266042Z digest=sha256:7ea5804ef1995412e3ddc59845923afabc202525ccb239a70ba4495435417480

Observation 332913ec-3213-4238-b852-00678b9d6229 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Doll \'a r, and Ross B.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Berg, Wan-Yen Lo, Piotr Doll \'a r, and Ross B

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:08:53.818075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:41.417931Z digest=sha256:e01bba0a7b46fe9f64b2b1049c3dfe99fc1b8ba170458395834a49d641984d04

Observation 6ca3f72d-c73e-40b3-a3bc-0b126bcfcb13 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Gonzalez, Hao Zhang, and Ion Stoica

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.567535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.567535Z digest=sha256:4e52f05b633744e9ba78df6f88848ab1848dbcabb9491be6f2695b4d0cad2336

Observation b7ba99d1-bc73-4c31-8fa0-5837d3cb84a8 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.796732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.796732Z digest=sha256:d9702cb8ed7bfa4164f850bcf6fa8f0f7b4cc1adcd0fea164b89c08330806874

Observation 86bf2cb7-94da-403f-8ed4-05db4472013b · outbound

This paper cites VEU-Bench: Towards Comprehensive Understanding of Video Editing.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought VEU-Bench: Towards Comprehensive Understanding of Video Editing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:41.965687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:41.965687Z digest=sha256:24d744864b531f85480b1b3f3d7344b54742d6c1185c166d8b18d854e6558b6c

Observation 8aa6eeb3-2245-4d99-8f1c-4cef689c819a · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.177990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.177990Z digest=sha256:567d42b28256ddc7ad32ce982bcf15aa7141115b27a450a04ba8edfd387621e3

Observation e32e10fc-095e-4f50-a471-72f556a82b9a · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:53.515400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:42.377897Z digest=sha256:8820afa8a978b3287efc871918ac783bbba0d19c7e4d7cbd3849150877e4ff36

Observation a1aa0c70-a29b-4b96-b5e4-a09febec9bff · outbound

This paper cites Visual Instruction Tuning.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Visual Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.538130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.538130Z digest=sha256:0a4af5751cf870ac7a05ae82c3234a297ae89f379958031ca4274e36a8ef5ce0

Observation 0b696a43-dd1a-4874-b45b-28cd24fca905 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.665200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.665200Z digest=sha256:244cd66caac05c789536ac33eec19e44bb92cfe9c5eaf3a411d5450b6515ce92

Observation 02ce7b08-8488-445d-88ce-899b1600cbcb · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.819240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.819240Z digest=sha256:2579a9494b23dffabf4e6c2f662d244974726daea7f08448826cb93f23166563

Observation d78528e8-4679-4c96-9299-0e685a999408 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:53.235178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:43.047540Z digest=sha256:957eef573f4d6f4b861b7a0fcafd9308c1d1ac770841eeb8718f1d57c957be29

Observation 38c7cbb9-2171-4374-95fe-3453f32c5ea8 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:53.008973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:43.249635Z digest=sha256:35ba08fc46d672208236f8282c8407e7770b12b5bbba5e247b3273821bcf69fe

Observation 1086a62e-8c4b-4569-b9c5-8c0f98f29315 · outbound

This paper cites GPT-4 Technical Report.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought GPT-4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:43.409280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:43.409280Z digest=sha256:fde708a5aeed8033aa16dc24ef70260e6934a1f98ff0eab84d88071f933781a2

Observation 07bd95a2-0b61-41ef-8f6f-ef55e92d43b4 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:43.637508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:43.637508Z digest=sha256:95ae3e5914c202dbb2556f587ac8c91f5ed2d7280941f84b82aa8169902f1fcc

Observation 59c95392-210b-4abb-a5dc-0851f2c3d032 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:52.759769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:43.844069Z digest=sha256:4f1c78da385478eac8a94a2fa1eb5941fab36cd7496a4673337ee14925a67045

Observation c8cc211c-6e79-45ad-921a-179e08be9277 · outbound

This paper cites TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.046424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.046424Z digest=sha256:d073df9944c837c4c96bd847011065f52d5941074ad1b8ba55cb318cf841e3ee

Observation d361bb9e-de9f-4f66-8159-23c9defc968d · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:52.450294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.100873Z digest=sha256:e5c09360b1593a974281536aaa14036d26bf963cec8ec4e36850353b668c0d55

Observation 666dbb4a-c7d0-4f11-875b-1b49aa22f766 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.190797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.190797Z digest=sha256:bfc5d45b36c276d472aa57f6906ba3363d71a515999ff4b30d0becc0e6037ccd

Observation bc1cc798-6871-4981-ba66-7bedd1d77ff8 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:52.265401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.257827Z digest=sha256:81ac81328b89ff5dba9387f46e00a3b4c9521f365066cc0b0a5d25f51139d7f7

Observation 28d2689c-4b13-477e-b09d-a028170a9f29 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.329476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.329476Z digest=sha256:9fade6d5f9e4e7f8276f80f4d92556cb79d28a626dd7a8158e6548a66c1fa8e6

Observation f553a1ed-58f5-4e64-ade8-c85ee0a09b8d · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:52.024786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.414812Z digest=sha256:6d2131038c6ccb82578b97bda4e4a52fd97f5b30e1f7ad05d2d5c98494feadec

Observation 37da8636-f985-4409-96fb-2f399c848057 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.516389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.516389Z digest=sha256:95fbd65659ac08405efa8a4eff9d86ece73db2a5de28b71ccb16cdbfb0a26af0

Observation 52cb0a5b-b851-4a36-955c-2efaf3e4ed51 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Gemini: A Family of Highly Capable Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.644702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.644702Z digest=sha256:bde50734eb953f74d7504a5a7ea6b06ef0de37eeff58e1048ad73de0f9992593

Observation 1892add0-d767-490e-97af-c7f2addd3e02 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:51.782181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.764286Z digest=sha256:0615a3da01441e4f56a9b9695611611603e85887b3b89367d78e72be1849c5a7

Observation e36f32f9-6776-45a0-ae8e-a42c8fc583af · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:44.852833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:44.852833Z digest=sha256:8f3049014b0eff3e6e4e615a7e18a246baea18e1e26541ea37f6db8a79131292

Observation 65624622-d2fc-4af3-a3c5-ff048ec6b058 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:51.565953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:44.994367Z digest=sha256:62e1f338712a13c377049e7ac7823c10ec76f3db9f1a4c7351bbb3c0027fdabc

Observation 07a56731-bf64-4a31-adb2-3fa73e7968cd · outbound

This paper cites Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.066926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.066926Z digest=sha256:f4e9061acc3d5b646061bae7b132419931e28d9a3d8d555e2e2be8abf15cfe24

Observation 6d5f953e-df5d-464b-891d-3f3984db2244 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.197404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.197404Z digest=sha256:581442c2bf8f009e6ecee97cfcf28d1721c9528699381b7d48426e80f6c8b7c6

Observation 9aa33499-ec22-44ca-8371-05a812bd4fe2 · outbound

This paper cites The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.273322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.273322Z digest=sha256:288526bfde9c5700b727471c6e6e42b89e224cadb7d27ac60feae72c279c793a

Observation b54b7879-cd9e-4278-a8fd-3b4d1c1f13f1 · outbound

This paper cites DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.339504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.339504Z digest=sha256:8267af2f9661205df20dc36616374d66dccfc85fb4d20f0fc4b1a80579dad7d7

Observation e62bb179-dca9-4451-af74-35c82c258d88 · outbound

This paper cites Number it: Temporal Grounding Videos like Flipping Manga.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Number it: Temporal Grounding Videos like Flipping Manga

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.397760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.397760Z digest=sha256:fe406bb9bb8d4d790d356254a33cb3fd5392ff2d2d4fa902c0a963df4df45a84

Observation 0cdca2a2-1491-42a1-84a2-f137d4b60ea8 · outbound

This paper cites Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:08:47.792006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.459689Z digest=sha256:842c42c2e5481b3764af5b6786f0239a9939aa7e32fdb897ca3cba819447e602

Observation 51cb02f5-cad9-44dd-bbc9-01845a6243e2 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:51.301911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.523531Z digest=sha256:37a468c687dae2cdcc1688ab9e74eab1422fd54a1ee3f4f9ee1b7c7840c14743

Observation a2e3dec0-6585-4e5c-8a44-a33bd8395fcc · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:51.086971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.602039Z digest=sha256:cf14644e83d43bb58564a45538ed6ebf479705812dfcddd3f8bd2e95ad261b52

Observation 7d3edb90-745d-4e70-ad13-949b3b486d8d · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.712931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.712931Z digest=sha256:544b3761b52e48a9140669c70730beed50171cedf8b0b44daf2c9591b63c20cc

Observation a5ccdc74-a977-4c00-bfe6-279859639c39 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:50.838207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.811259Z digest=sha256:d39a7d8e76528c4df9e5b68598e616cf6e37eccbc23e798e92370667bf4884f9

Observation 6e8370fa-e4b9-496a-83dd-5abf89a01f1a · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:50.591962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.874839Z digest=sha256:14040dbef07f8b9ca85ed71413eb8609da3601423bf4fbca8370ebde9e4561fc

Observation 74d27c5d-3e66-4841-823a-d5a35c83a16f · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:50.368405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:45.955480Z digest=sha256:e89a4c159458343d219029fef9fdbf83cb9270d8d56fd02fa01bc735cc2b147c

Observation f6460ecc-2516-46b6-8951-e023a26d0931 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:46.024434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:46.024434Z digest=sha256:75fa2c85988a38130d5f51dd683934781d7b47479b2d10ac57943f691f650545

Observation ce95bdd1-bec3-4e60-a458-dfa9cc0aaec8 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:50.131375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.080861Z digest=sha256:82d0d3e97456f6ac9a351baad0957a4e8883daad4b2216cdc5834898914ad068

Observation 1af5bb6a-a3d7-4a4d-b806-8c0e4ecfae47 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:49.859512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.216434Z digest=sha256:ef855070b5c4fc0451b4ec4cea30b2702b00d74c00419c64cdefe27ead43e975

Observation 7f8f49d4-b716-4bba-9456-e474910fbc0c · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:49.617597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.357132Z digest=sha256:48623b580dd7495c08d8b7998921186a97d007e8f101fd62688859377bf66e4e

Observation 30762b84-bfb6-40ed-9b56-44f745255e11 · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:49.407536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.504809Z digest=sha256:4622c979c258151f2194924b941ee2af7e7c36515ba6d6c131a2b264a0b6ebe8

Observation b0b74032-a9cf-4e96-b770-10ae30cdb709 · outbound

This paper cites Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:46.629692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:46.629692Z digest=sha256:b1c117fa18ca11e045948ec9c551687d0ef375b2a8146980ce036c4b935b6c5b

Observation ad0ceb4c-2c7f-4b19-9b02-a574aa1b35a1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:46.785107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:46.785107Z digest=sha256:8ee841083ce040e08129061408a1b4114ed3cd128476a1fcb5670923706d0369

Observation 30acb137-6da1-4085-9efb-e2691c641a7f · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:49.214546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:46.880493Z digest=sha256:f704e9c840f763bbda7b141686a1e42e9fd2f4272af8a7d993e0bca3e0f198e2

Observation 3f162f72-5613-42d6-a8d9-85c338a50d2f · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:48.965251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:47.023203Z digest=sha256:f7f3f97121aeca1d6c55d0e3a768e6e6ca72a6692956a9c7cb5b401d43a5bf59

Observation 584af511-db2a-4278-891a-eaf7fb72764c · outbound

This paper cites an unresolved cited work.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:08:48.633512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:08:47.113216Z digest=sha256:a8c9951a8182c24bd1874763bbc9e61c18948f6532681373acfc836d93982616

Observation 45e7b5e3-81bc-4183-babd-a1581b8e9909 · outbound

This paper cites online" 'onlinestring :=.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought online" 'onlinestring :=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:47.214335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:47.214335Z digest=sha256:66101bc8dc7e25248a967bd3c0ac078f66dadb6cd0fb45c59be709b549c70a4c

Observation b80bce5c-b602-4db9-862e-9dbc50ccd111 · outbound

This paper cites write newline.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought write newline

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:47.365596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:47.365596Z digest=sha256:bf0c6dc5129ee56b31bf0c7a575ff4469e80feb0b86c91e67040012c7db87fa6

Pith citing papers

Observation 0ffd6ff2-f54f-4d8a-9d37-5055ffd9f570 · inbound

SAM 3: Segment Anything with Concepts cites this paper.

SAM 3: Segment Anything with Concepts RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.519580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:174278318a7056713127f1a09892d460cc95c9139a4694ba05d4f8f41baaf687

Observation 49641d1b-4775-40c6-a75b-a7368f48152b · inbound

A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding cites this paper.

A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T14:10:17.325958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:10:17.325958Z digest=sha256:a1e037a41deaf6435af89dba7837155efc410f21875303f46fb9569be0c19f77

Observation f7e37821-8e46-41e3-ba15-3eaf6e4791dc · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:16:34.528679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:12:46.385646Z digest=sha256:f47e2c11a3a8b1b23fbe8f4efa2967340ae3a501bfc7dd5e1c9a9741c5e9f55d

Observation f33bb049-8095-4067-98f4-d6b7bc5feac5 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.098104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T10:30:06.829915Z digest=sha256:c3bf82414d556d940ffdcf1efd9c1b2a36d0890ac0c7e6fcacecf127ab9eca18

Observation f5f8ffce-3e0a-4b13-a6cb-a408093d7064 · inbound

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs cites this paper.

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T22:00:07.204123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:00:07.204123Z digest=sha256:c2fe88c9e51fdc3e54033c7b8128413d8f67594f034a1d676d2f75594adee6d6

Observation 97b80294-61cb-429e-b5e1-b3d601a20faa · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.431402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:7ef872a60bb03efc30e894281ae336630d618870b75cfdc98ce63068cf3f845a

Observation 5e176c91-e057-457a-99bc-1fa61be974c2 · inbound

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning cites this paper.

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:47:30.633469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T17:01:13.745646Z digest=sha256:527cf4af6c37459c95159559e18be24cc5681ab7a851eb9343995efc9cc02342