Pith. sign in

Paper Citation Record · LEDGER

Reasoning Segmentation for Images and Videos: A Survey

As of 8 August 2026, this Paper Citation Record lists 100 of 114 outbound references and 5 inbound Pith citation observations for arXiv:2505.18816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18816 v1

Coverage vector

measured 100 of 114 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:18.789405Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:06:22.967237Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.361578Z

Reference resolution

100 of 114 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2535e581-f34c-49aa-a4cd-90ff023a2e52 · outbound

This paper cites UI-Net: Interactive Artificial Neural Networks for Iterative Image Segmentation Based on a User Model.

Reasoning Segmentation for Images and Videos: A Survey UI-Net: Interactive Artificial Neural Networks for Iterative Image Segmentation Based on a User Model

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:27:20.276725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:11.018222Z digest=sha256:39bf7b2160024dde46293ec685ea0c58f419ed5eab598f027c91964562a85f35

Observation 5104d3b1-8a63-48c0-bd05-b019af82bf66 · outbound

This paper cites ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation.

Reasoning Segmentation for Images and Videos: A Survey ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.072412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.072412Z digest=sha256:eaeb23922a32c665b4ead73409455654e0baa503f07576c46835adfda55cd298

Observation 5628c86b-4642-451e-80ca-aa9b60782db4 · outbound

This paper cites Burst: A benchmark for unifying object recognition, 39 segmentation and tracking in video, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Burst: A benchmark for unifying object recognition, 39 segmentation and tracking in video, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.134478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.134478Z digest=sha256:9767ebcdfd0e420565a3dbba96718cc831763e856ed4c1eff4c40e5a84ad3ab4

Observation 1596bf16-837f-4dc4-bc2f-653d81fd4952 · outbound

This paper cites One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos.

Reasoning Segmentation for Images and Videos: A Survey One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.218146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.218146Z digest=sha256:9d3de62eb2e7a54b893200e550fadb9381944089f88cb9dfaeb9856d2f6df671

Observation 98bc1725-92aa-476a-a66b-6da70b393a0b · outbound

This paper cites Cores: Orchestrating the dance of reasoning and segmentation, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Cores: Orchestrating the dance of reasoning and segmentation, in: European Conference on Computer Vision, Springer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.273826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.273826Z digest=sha256:3ff0e077247b12ccefd5cf3bba193733c345713c5f35dd2ad3665fc41219c293

Observation f2799ce2-0ec9-4498-aadb-2f0dd8d18212 · outbound

This paper cites Xmem++: Production-level video segmentation from few annotated frames, in: Pro- ceedings of the IEEE/CVF International Conference on Computer Vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Xmem++: Production-level video segmentation from few annotated frames, in: Pro- ceedings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.329047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.329047Z digest=sha256:62970f562fa168ea3ed0c61e1cd751824d20a4ff1442c74cc66964d3646e1306

Observation 3e7199cf-1e6b-407e-b00d-f1c96ac9424c · outbound

This paper cites Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.431803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.431803Z digest=sha256:278990024c08d486803534381735749eeb1dabbcd0851ec7a2553d9912b18063

Observation 54b9c109-2a0d-4591-b8e9-1cb2d4b7aada · outbound

This paper cites Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.489319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.489319Z digest=sha256:18709513a5719b48d29c50b7d25d6c0d693732f93f68ffefc969f704124bc9bc

Observation f732ecc8-2281-48dc-9fa4-6ee936b81096 · outbound

This paper cites Pixel-Level Reasoning Segmentation via Multi-turn Conversations.

Reasoning Segmentation for Images and Videos: A Survey Pixel-Level Reasoning Segmentation via Multi-turn Conversations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.587733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.587733Z digest=sha256:a8c493207df54c57f513620e6c510dfb66b4e491e1a9cd6e38a9ec5001c4f165

Observation a2672107-2874-4dd1-9b28-4eb09fc36747 · outbound

This paper cites End-to-end object detection with transformers, in: European conference on computer vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey End-to-end object detection with transformers, in: European conference on computer vision, Springer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.653276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.653276Z digest=sha256:6803182d0d2a3bd281bd71cf81f3900de2ec08dd2c1643521db3dd821d0b9a0d

Observation 9c17b8da-d672-4c2d-8dcf-003f75f2aa62 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.754067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.754067Z digest=sha256:9c9a2504cc13897c3e585b6b8812fc5521d38ed57a459dd2a89e015c557bd73f

Observation 0cf067c6-6e1d-49e2-b69c-d469484d8642 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.868491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.868491Z digest=sha256:df6d4340df7e1b393e53f3f875696fb462786527f0fa4a6ad3339099cdc0d510

Observation 0023a925-9097-4578-b971-51f60d688532 · outbound

This paper cites Sam4mllm: Enhance multi-modal large language model for referring ex- pression segmentation, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Sam4mllm: Enhance multi-modal large language model for referring ex- pression segmentation, in: European Conference on Computer Vision, Springer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.942872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.942872Z digest=sha256:a037aaf677e6968443a952b0e6493b01405152c75f3d1b8b22a70d8ff9470017

Observation 4250dca0-38ea-4da1-bb92-80a94ba49762 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Masked-attention mask transformer for universal image segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:11.987613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:11.987613Z digest=sha256:9fedc7cdca8361a6972eab4afa6176578d7fd9319973618fc0d49068cf9cfdb0

Observation 1ed77509-9ee5-4d18-adb0-fd733b49b0dc · outbound

This paper cites Xmem: Long-term video object seg- mentation with an atkinson-shiffrin memory model, in: European Confer- ence on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Xmem: Long-term video object seg- mentation with an atkinson-shiffrin memory model, in: European Confer- ence on Computer Vision, Springer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.062712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.062712Z digest=sha256:7f646ad1ead1552de1eca6f6d9fe5e0322613881496bb79ec7e0009f2afcd2a0

Observation 47fd545a-b396-4222-bb23-690101e7ba9a · outbound

This paper cites Vocabulary-free image classification.

Reasoning Segmentation for Images and Videos: A Survey Vocabulary-free image classification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.140729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.140729Z digest=sha256:819a89652b71f30b28140753d02145f060fb1052d41c5df40d79b58a261749eb

Observation cf1d2c9b-d29f-4510-bfab-118019625441 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Scannet: Richly-annotated 3d reconstructions of indoor scenes, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.225657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.225657Z digest=sha256:8452c6f4925e753a3011807184da1013905a50859f928e4ea970132d96c31d40

Observation e90944a5-7290-4d20-87f3-5469fca853cc · outbound

This paper cites Tao: A large-scale benchmark for tracking any object, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part V 16, Springer.

Reasoning Segmentation for Images and Videos: A Survey Tao: A large-scale benchmark for tracking any object, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part V 16, Springer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.305165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.305165Z digest=sha256:0efcaea0c1d172b7ca7171897f9cbccb235e9ee6f579589cfdc0ea73a5ffb889

Observation 82302091-de9e-46e9-bae2-48fc6dbb0626 · outbound

This paper cites Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level.

Reasoning Segmentation for Images and Videos: A Survey Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.401764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.401764Z digest=sha256:3bed1602d7587d507479b9d9cba032d2e29e65d9c74948986c30714c13ac5b44

Observation a4b13401-579b-4aad-ac1a-9744ffa96064 · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Mevis: A large-scale benchmark for video segmentation with motion expressions, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.500258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.500258Z digest=sha256:abd8213544860c4ab2b2b5ec8a75ea914b157b6b5da329215545b90087b73baf

Observation 783e111d-b744-4e99-b275-4ac4768809cb · outbound

This paper cites Mose: A new dataset for video object segmentation in complex scenes, in: Pro- ceedings of the IEEE/CVF international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Mose: A new dataset for video object segmentation in complex scenes, in: Pro- ceedings of the IEEE/CVF international conference on computer vision, pp

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.592565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.592565Z digest=sha256:c683ca1f5f0c014ad7e27e7cde7496da5b4bfd014330b4062e82e50643678ce6

Observation 579de55c-4d1f-4f10-b152-c97f38c742c9 · outbound

This paper cites Panoptic Segmentation: A Review.

Reasoning Segmentation for Images and Videos: A Survey Panoptic Segmentation: A Review

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:27:20.231845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:12.659787Z digest=sha256:984bf42f43beb57443a8b6234aeb240356f356377ddfa4cd3ca69c35df957b80

Observation bcf025c8-490b-4b28-b496-4f577043241f · outbound

This paper cites Oops! predicting unintentional action in video, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Oops! predicting unintentional action in video, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.718042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.718042Z digest=sha256:0eb9bca000d5f425ce2b136a2f481206707326bc2b81b94fda6cae312553aaa3

Observation 56a0ce34-9da4-4754-8128-2335acf07537 · outbound

This paper cites A Survey for Foundation Models in Autonomous Driving.

Reasoning Segmentation for Images and Videos: A Survey A Survey for Foundation Models in Autonomous Driving

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.818555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.818555Z digest=sha256:e86934a3f1caabce3a492f3227b71366961505fec1dc24a0469f0624fc4dc33a

Observation abe7de4f-a2cf-4111-ac9e-f8fe5fd37ed8 · outbound

This paper cites Part- aware panoptic segmentation, in: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Part- aware panoptic segmentation, in: Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, pp

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:12.927968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:12.927968Z digest=sha256:25be33ab7f13234283a93ed1d9633d0e5cfa04461178e9f0bbb2d7ee389d8afb

Observation 770e0673-4881-4b20-abde-e9f044696652 · outbound

This paper cites The Devil is in Temporal Token: High Quality Video Reasoning Segmentation.

Reasoning Segmentation for Images and Videos: A Survey The Devil is in Temporal Token: High Quality Video Reasoning Segmentation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.039735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.039735Z digest=sha256:6297dd7145ad28533be7ef0f8c934ebb419e0ff7d52049605ba5b22951260d57

Observation d3443753-da5e-48db-b968-d39eaebd1a60 · outbound

This paper cites The iapr tc-12 benchmark: A new evaluation resource for visual information systems, in: International workshop ontoImage, pp.

Reasoning Segmentation for Images and Videos: A Survey The iapr tc-12 benchmark: A new evaluation resource for visual information systems, in: International workshop ontoImage, pp

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.120490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.120490Z digest=sha256:55429efb3dad44fde6430fcb434e84911d394f2ded44bc84e76a335009a971c1

Observation b0496995-8774-403f-a4dd-cde1f9e7a271 · outbound

This paper cites Lvis: A dataset for large vocabu- lary instance segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Lvis: A dataset for large vocabu- lary instance segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.222976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.222976Z digest=sha256:e790278eedd46093716ef67c31523989a7d98606f4495eb89cb6109f400dbf80

Observation 3b31db6e-4004-49a0-9ed4-73dddcf0523d · outbound

This paper cites A survey on instance segmentation: state of the art.

Reasoning Segmentation for Images and Videos: A Survey A survey on instance segmentation: state of the art

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.329467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.329467Z digest=sha256:e0ac920a1bbe9bc317abfcf931d6bb52915f93dcf93e0b77817e485c1ec0c8a2

Observation 99f2b6af-40a8-4263-8878-38917cb669dd · outbound

This paper cites A brief survey on semantic segmentation with deep learning.

Reasoning Segmentation for Images and Videos: A Survey A brief survey on semantic segmentation with deep learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.453671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.453671Z digest=sha256:5db4c9fb4ddb3d138fafbd55fa67f4178d26395fd39e0cb1c6eb4a61f9a83389

Observation c2a94a13-f9d0-4335-aa48-30e7e1af3941 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Reasoning Segmentation for Images and Videos: A Survey LoRA: Low-Rank Adaptation of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.560190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.560190Z digest=sha256:e7e30254402c1d098c1b96847c50cacf14e1c6df2bdaf18eb5698a4236883082

Observation 8ff5aa0d-221e-4702-92d0-32ff4c5ba8a4 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.682259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.682259Z digest=sha256:4afabd5466f054e8bf1cd33536f406ffccc4d4290d43808f7d3546eccdb79ee9

Observation d3d0aa40-7f13-4279-b0aa-637e14b53d26 · outbound

This paper cites MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation.

Reasoning Segmentation for Images and Videos: A Survey MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.729538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.729538Z digest=sha256:6ed16ef44aecf414552e9540dcdcb192ef258fe39f2226adcc91d2321ba497e5

Observation 253e3da3-74e8-4901-9f12-5f9800cc4d08 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes, in: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp.

Reasoning Segmentation for Images and Videos: A Survey Referitgame: Referring to objects in photographs of natural scenes, in: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.814778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.814778Z digest=sha256:586fd25b8556c9b1ca6ff608736a264406e7d1e942997d1299a220c3ed24e8dc

Observation e889dbfc-6c15-40ee-b6c1-ad95c4e89f14 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:13.880321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:13.880321Z digest=sha256:337db890e37a0fad2d47f50b040e827cad8d73009a159452a8c66ec807bd8426

Observation 21a0058f-f297-4cf8-8f18-d9b02ba6c83c · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.033054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.033054Z digest=sha256:73ede327c12a5d14f16da2deb99a6e09b18c9130d250c6a18e8f7e834f4b86a8

Observation 9bdf61cd-b914-4ae6-9b78-ede9ac9bb40a · outbound

This paper cites Segment Anything.

Reasoning Segmentation for Images and Videos: A Survey Segment Anything

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.085972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.085972Z digest=sha256:e012121eed721c62ca6cbde4f6cb845112cf03ce86419fa014b991ad387900c0

Observation 35d9f209-ff49-4b67-aded-df2dbded454b · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annota- tions.

Reasoning Segmentation for Images and Videos: A Survey Visual genome: Connecting language and vision using crowdsourced dense image annota- tions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.115774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.115774Z digest=sha256:6343d7be093eb9e2fc83aec8886bd06594a2435be0d7cb5e2c3e25ff43c6525c

Observation 07e0de33-8137-44ee-ba4c-eb01420f4378 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Reasoning Segmentation for Images and Videos: A Survey The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.277333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.277333Z digest=sha256:44e0a9b2e71ddcc42e7590cecc4f402eb354fc2ef46d271f0fc29e7e201b3e61

Observation c4177e72-78d3-4531-aaf8-79acb0d3d0d0 · outbound

This paper cites Lisa: Reasoning segmentation via large language model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Lisa: Reasoning segmentation via large language model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.380393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.380393Z digest=sha256:039fc13e305b4ad2074d33335fa6da79dc325e7908b635b031e1c5bd08e12c52

Observation 31641866-bf98-47f1-9203-f901a609f30f · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language mod- els, in: International conference on machine learning, PMLR.

Reasoning Segmentation for Images and Videos: A Survey Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language mod- els, in: International conference on machine learning, PMLR

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.526975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.526975Z digest=sha256:1b341b29c359296015c27bc4a8dfdebd1ac29ad79e60375ef795d193dbb21158

Observation a8b427f1-981a-4b88-9931-54c2b2ae607d · outbound

This paper cites SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model.

Reasoning Segmentation for Images and Videos: A Survey SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.624387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.624387Z digest=sha256:16a07f4e0bca1586bef974c8961b7debe30726b0aa27a4d079e0875be3ab2a37

Observation b28ed6f8-e70a-43ca-aa56-40198daea904 · outbound

This paper cites Robust referring video object segmentation with cyclic structural consensus, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Robust referring video object segmentation with cyclic structural consensus, in: Proceed- ings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.796318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.796318Z digest=sha256:f4e22cc4d6608574a8f43981973a17724a8e7ef44dca44150ee4e0dfca8a4b99

Observation ebc5c99f-5a2a-488f-b500-da0ab41c0c1b · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Llama-vid: An image is worth 2 tokens in large language models, in: European Conference on Computer Vision, Springer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:14.946198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:14.946198Z digest=sha256:fb82edf5ca4ba879d98a527d857ae97f1847beaedcd1832a10a11268f11e2de0

Observation 3b393a59-01d6-491b-9e1f-1e25d2744064 · outbound

This paper cites Microsoft coco: Common objects in context, in: Computer Vision–ECCV 2014: 13th European Confer- ence, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, Springer.

Reasoning Segmentation for Images and Videos: A Survey Microsoft coco: Common objects in context, in: Computer Vision–ECCV 2014: 13th European Confer- ence, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, Springer

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.661879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:15.084904Z digest=sha256:735ebdad4fa9181ece5594f45633a83f23c034b46436c548bfee845ca1c3d4ac

Observation 9b0163e8-9b0b-48eb-ba6a-16065bb2e529 · outbound

This paper cites Gres: Generalized referring expression segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Gres: Generalized referring expression segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.653225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:15.241954Z digest=sha256:063aa1e558e5b769a4846483e18650381db93d48b929290d35b5801077e23514

Observation 09a281c4-b093-437f-84cf-ce514a7497c9 · outbound

This paper cites Visual instruction tuning.

Reasoning Segmentation for Images and Videos: A Survey Visual instruction tuning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.643861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:15.352155Z digest=sha256:b12241c22eff8599b815668869dfb403cb0592c1a2aef204bd8905fca34bfe6d

Observation 17fd4db2-02f2-476e-a215-7c8d40dae819 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Reasoning Segmentation for Images and Videos: A Survey Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:15.482347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:15.482347Z digest=sha256:dffbdf764c825bca3df807eaa590ac7ddea84796e4b50e8dc1c95d1c57fc430f

Observation af282039-22b4-42fe-8e33-9d88f65ff8df · outbound

This paper cites Ground abstract structure concepts of scaffolding systems for automatic compliance check- ing based on reasoning segmentation.

Reasoning Segmentation for Images and Videos: A Survey Ground abstract structure concepts of scaffolding systems for automatic compliance check- ing based on reasoning segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.635044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:15.625060Z digest=sha256:517c8371ab1c8e1bf12d989636bc048dcf69cd149f2980dd47858ada8bf5e652

Observation e6c3dddf-c8cf-49f0-8e4d-304427d957a7 · outbound

This paper cites Video anomaly detection and explanation via large language models.

Reasoning Segmentation for Images and Videos: A Survey Video anomaly detection and explanation via large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.627054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:15.763488Z digest=sha256:a3d80fade57590629bc305f5856db216fb3e7b0b40744a1a5f9cac5b44f64e36

Observation dc0c66bc-dbca-49f9-9a4f-6239e505fd84 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Reasoning Segmentation for Images and Videos: A Survey Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:15.874035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:15.874035Z digest=sha256:d8704616451ea40dff81979cb823e90929c65e4a03cdb09dc86492b6ff7659d6

Observation 41bdf037-fc82-43e5-b1cd-8f4d9b98bd02 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:27:20.618602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.015561Z digest=sha256:955878fb473b0cf637bb82702d9e94bd46daf014812b13e0203ad50471e5711a

Observation 4552696c-2c40-4fcf-8ca9-645ee71e25db · outbound

This paper cites Large-scale video panoptic segmentation in the wild: A benchmark, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Large-scale video panoptic segmentation in the wild: A benchmark, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pp

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.603088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.115719Z digest=sha256:677e6007d1356270af55d0f63e4f26629a7fab23b957691ca22f1d9ad282c08a

Observation d7785866-13af-4b67-a815-1f83c89251f6 · outbound

This paper cites Image segmentation using deep learning: A survey.

Reasoning Segmentation for Images and Videos: A Survey Image segmentation using deep learning: A survey

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.594634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.159160Z digest=sha256:8e4cc6c7721c4ab1276d7f821cad69dd4206d868501328dc1ad6996bd2e2aac0

Observation e0d0e917-465d-46e7-9e8d-6def35c33a95 · outbound

This paper cites GPT-4 Technical Report.

Reasoning Segmentation for Images and Videos: A Survey GPT-4 Technical Report

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.586466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.209495Z digest=sha256:24a42ce0d1543da68e87dcec07e50b12c7cf58d4cfe96dfae7277d3b14c88933

Observation 1d4d6a45-bf66-432c-8b22-6c4f1b2ab48f · outbound

This paper cites GPT-4V(ision) System Card.

Reasoning Segmentation for Images and Videos: A Survey GPT-4V(ision) System Card

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.578848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.250907Z digest=sha256:065bb5a9c4afd959a66507ea15e9186dbbb35d3356e2662c17f99130cc5a1e3f

Observation 53191d63-744b-4e11-9acd-8cc39efae377 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Reasoning Segmentation for Images and Videos: A Survey DINOv2: Learning Robust Visual Features without Supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.276213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.276213Z digest=sha256:a175d33cf04b0032cd6c095053f636af7f61a97cb051aabb12a630ca29abc591

Observation 9e9a4395-0340-4ad8-92db-1bdbc82aa975 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models, in: Proceedings of the IEEE international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models, in: Proceedings of the IEEE international conference on computer vision, pp

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.570111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.318324Z digest=sha256:a8ca6ab0f418fc05d1d1cf5e7f8b404a3c4cfc42c9101d0c5426c253ab0eef7e

Observation 264110e8-d488-4fd7-ba05-e7207fa041de · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

Reasoning Segmentation for Images and Videos: A Survey The 2017 DAVIS Challenge on Video Object Segmentation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.379401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.379401Z digest=sha256:3bccdd3ce7be80e89610d26b29c9250081eda4c4b920c28f563df850d4c63539

Observation 184032c2-86a5-4be1-8391-b84572ca2e99 · outbound

This paper cites Occluded video instance segmentation: A benchmark.

Reasoning Segmentation for Images and Videos: A Survey Occluded video instance segmentation: A benchmark

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.560685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.420028Z digest=sha256:7b586689cba0e4ee6dfa893258fb28f0ec572e80e3afdbfcd2e2f91dc4d4c5e1

Observation 316d3aa7-825b-4253-8944-3f4469407fdb · outbound

This paper cites Reasoning to attend: Try to understand how< seg> token works.

Reasoning Segmentation for Images and Videos: A Survey Reasoning to attend: Try to understand how< seg> token works

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.459345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.459345Z digest=sha256:1f724b54ffdab1fcaff63890c82c4d1d19604f85276a25e28bc595d9154efaf3

Observation 1a67f527-dab8-4bf5-bee0-62f351497b25 · outbound

This paper cites Learning trans- ferable visual models from natural language supervision, in: International conference on machine learning, PMLR.

Reasoning Segmentation for Images and Videos: A Survey Learning trans- ferable visual models from natural language supervision, in: International conference on machine learning, PMLR

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.552768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.502970Z digest=sha256:2071a56154de5f67de217c4d4d169604ce710984606b4b0acd4b9979d28e92ba

Observation e994d17d-0097-423f-9f4c-bbda7b8edf20 · outbound

This paper cites Paco: Parts and attributes of common objects, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Paco: Parts and attributes of common objects, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.544550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.551416Z digest=sha256:e43aea0100d11e0a95a706f3335ecf699b0bf2a7fb418b00cdc76bde0d168b56

Observation f5f67165-0da6-4ff5-9b25-611289af9cd7 · outbound

This paper cites Glamm: Pixelground- inglargemultimodalmodel, in: ProceedingsoftheIEEE/CVFConference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Glamm: Pixelground- inglargemultimodalmodel, in: ProceedingsoftheIEEE/CVFConference on Computer Vision and Pattern Recognition, pp

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.536167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.605622Z digest=sha256:73356033eefe614eb0f8990941e88fe465364ad796829eef2c8a60c085bb3d97

Observation d5e00767-6e8e-4f2e-bc96-94c6afd01c47 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Reasoning Segmentation for Images and Videos: A Survey SAM 2: Segment Anything in Images and Videos

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.663543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.663543Z digest=sha256:021cfda0450c04674f99dd7f343d84019b9d38c5ebb315426088502eb9a389cb

Observation 13e78c09-8f72-4f36-a0a7-99b403efb154 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Reasoning Segmentation for Images and Videos: A Survey Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.717849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.717849Z digest=sha256:ad0c07408eb15d8ff0c3bf681a6a0fb605b6aeb6c56b88c88542029c9b1a3f3b

Observation 018d04b0-a9d8-4d8e-9aa5-5213379dccff · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Pixellm: Pixel reasoning with large multimodal model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.527449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.755628Z digest=sha256:28b91849008a871d3cec8ccffb5bdfbde26d5b26fa27b02c3169368dceb0162c

Observation b7865683-1aa7-4cf3-8385-19d6aa57d54a · outbound

This paper cites Object Hallucination in Image Captioning.

Reasoning Segmentation for Images and Videos: A Survey Object Hallucination in Image Captioning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.801705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.801705Z digest=sha256:8ede7901da228724dfdc8f8ec2ab097a4c77f4536f971322f00a2c1db431e96c

Observation f488d60a-fcf2-408d-908e-3ea4f8cb5ed5 · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:27:20.518963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.840520Z digest=sha256:2cc05bfee0256ddc24df7deb1afa4935a0eeca7895270dd23cd925f0a5a9d9b3

Observation 8f6a796f-b9c3-43eb-afa1-2f3b46f58aae · outbound

This paper cites Position: Foundation Models Need Digital Twin Representations.

Reasoning Segmentation for Images and Videos: A Survey Position: Foundation Models Need Digital Twin Representations

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.888010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.888010Z digest=sha256:6e0058c02ef161ee586eb7a88de8cf3f9af1553abb981a6f44d161fb28970fe8

Observation 6154626e-f598-42ac-b4cf-248e91f05f2b · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:27:20.510072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:16.937489Z digest=sha256:9a6572455a94e94f7c4ec7f3509e33f2dc5ac03805effe4fe9b5a84cb3aec78c

Observation ab7aff1f-a7d5-44b9-82d4-6c5c1e4bc53a · outbound

This paper cites RVTBench: A Benchmark for Visual Reasoning Tasks.

Reasoning Segmentation for Images and Videos: A Survey RVTBench: A Benchmark for Visual Reasoning Tasks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:16.976460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:16.976460Z digest=sha256:4a35505fd8d192dad06cc4cc6e25d778bb94acb72893bb56197ddb2bcdf0c1b3

Observation 6260a082-1e7a-49cf-8066-b91dd9288ce4 · outbound

This paper cites Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins.

Reasoning Segmentation for Images and Videos: A Survey Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.030764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.030764Z digest=sha256:a6146f8605218924eca95d65aa941a025f113988465ffb3e924edad73ecbd7d4

Observation 92574bd9-3902-4306-b588-39bf8b41ed14 · outbound

This paper cites Online Reasoning Video Segmentation with Just-in-Time Digital Twins.

Reasoning Segmentation for Images and Videos: A Survey Online Reasoning Video Segmentation with Just-in-Time Digital Twins

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.086609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.086609Z digest=sha256:2facb32e5600253b47deea489bf63e7a315c57fa5f8213821f3b6a54174c1b53

Observation d201cc74-11a5-4ff9-b09c-545f0a9bf9ca · outbound

This paper cites an unresolved cited work.

Reasoning Segmentation for Images and Videos: A Survey Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:27:20.501488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:17.146232Z digest=sha256:c7295c3fdf372dca825efb21071ec2f5a213844fb4ee48a2baf6515601f2e3d8

Observation e75943ca-b73f-4740-b246-86923e47b367 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Reasoning Segmentation for Images and Videos: A Survey EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.198173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.198173Z digest=sha256:7b2598155e96e43251c36315cb0f62d9e61ea4736a4911fd494687b6640249eb

Observation ed7bddb8-2d12-4292-9d58-85e2c9f5c624 · outbound

This paper cites Growcut: Interactive multi-label nd image segmentation by cellular automata, in: proc.

Reasoning Segmentation for Images and Videos: A Survey Growcut: Interactive multi-label nd image segmentation by cellular automata, in: proc

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.493778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:17.223479Z digest=sha256:d307400c0ed670de2292eccffdeea0174f0cf5627c2ba6990521dd98e2d2eff0

Observation 02c0e07d-8a62-4754-b277-3d52456ef1a1 · outbound

This paper cites Reducing the annotation effort for video object segmentation datasets, in: Proceed- ings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Reducing the annotation effort for video object segmentation datasets, in: Proceed- ings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.484452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:17.274833Z digest=sha256:09d5989100f855d01bf80697066989a039c87636ef2ad9c8786e4cd0eb48cc67

Observation 7b617589-b580-4d20-a999-c99295fe8c3a · outbound

This paper cites Prima: Multi-image vision-language models for reasoning segmentation.

Reasoning Segmentation for Images and Videos: A Survey Prima: Multi-image vision-language models for reasoning segmentation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.319384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.319384Z digest=sha256:2fe2a924af87c9d70374d19d14f849c154af9fcceafdcdd995b84bd0b582cce0

Observation a5194841-dd28-44a5-bad6-18c571d4eda9 · outbound

This paper cites Ov-vis: Open-vocabulary video instance segmentation.

Reasoning Segmentation for Images and Videos: A Survey Ov-vis: Open-vocabulary video instance segmentation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.475048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:17.370790Z digest=sha256:ec36967ce4b4cef1209928376cc1ca2aff745e98e49eb61b8aa44ab44597a469

Observation 4f1441ef-eeb7-4181-9c0a-80a193c00a37 · outbound

This paper cites Towards open-vocabulary video instance segmentation, in: proceedings of the IEEE/CVF international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Towards open-vocabulary video instance segmentation, in: proceedings of the IEEE/CVF international conference on computer vision, pp

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.465799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:17.422485Z digest=sha256:1a84f4805b88bd06e9ce4c252d93c2d841e29ea82c1f18844fee78b9469c6249

Observation 77313945-d7f1-4d7f-8ddb-150254cd2b55 · outbound

This paper cites Llm-seg: Bridging image segmentation and large language model reasoning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Llm-seg: Bridging image segmentation and large language model reasoning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.457907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:17.475549Z digest=sha256:2c7cfb85e7d768a06710d4a508b23c839f0b2936e94edb861596291a54c34f97

Observation 135e0f7e-ac33-4b4b-822b-7803a27d02ac · outbound

This paper cites Unidentified video objects: A benchmark for dense, open-world segmentation, in: Proceed- ings of the IEEE/CVF international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Unidentified video objects: A benchmark for dense, open-world segmentation, in: Proceed- ings of the IEEE/CVF international conference on computer vision, pp

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.450167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:17.515490Z digest=sha256:3943744c2b9f6f360109602eb2e8a8581e8c5865b75f4de607b3736b65575a60

Observation 69049e09-7e92-4305-a151-604abd22817f · outbound

This paper cites SegLLM: Multi-round Reasoning Segmentation.

Reasoning Segmentation for Images and Videos: A Survey SegLLM: Multi-round Reasoning Segmentation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.599286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.599286Z digest=sha256:7e484d8d0f2bfbdd419207a9c14275a8beabd1973285c19714444eb299792ee1

Observation c58eabf6-7f01-4e6e-92c4-d9cbc0d0ad04 · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

Reasoning Segmentation for Images and Videos: A Survey LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.662473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.662473Z digest=sha256:86abd58a421baef3b82d4360aaeec0e1e34e2271fe82c8613540008acf61d8de

Observation 84d1e245-ea60-4124-8039-1caec776250d · outbound

This paper cites InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models.

Reasoning Segmentation for Images and Videos: A Survey InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.717014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.717014Z digest=sha256:b88b05295ca2179aa46ba9b9256da2e3cae78bedf750398d0904840b614b1c0d

Observation 87f917ea-32ca-4d10-a3e9-72157c73ce39 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Reasoning Segmentation for Images and Videos: A Survey Chain-of-thought prompting elicits reasoning in large language models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:17.757478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:17.757478Z digest=sha256:253a0a2c51880337050db529935ab199a2c6a9783aba4c63e21280c4b59084c1

Observation f4eeaab8-288b-41c1-a4a6-1bde2a30ff2b · outbound

This paper cites Ov- parts: Towards open-vocabulary part segmentation.

Reasoning Segmentation for Images and Videos: A Survey Ov- parts: Towards open-vocabulary part segmentation

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.437125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:17.799215Z digest=sha256:bb929a04295713e135e1239af6eae6e4d000105f830f988a77ea15b36482239b

Observation d94a26f9-d9b2-495b-a66b-3aec75517bdd · outbound

This paper cites Phrasecut: Language- based image segmentation in the wild, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Phrasecut: Language- based image segmentation in the wild, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.428859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:17.895761Z digest=sha256:05ad0382ab5589711c91e4c05e6682a8dec76a31a9e6431d445c3e32bdc307d4

Observation 482a4945-58ae-4189-81d3-31cdb93c7b47 · outbound

This paper cites See say and segment: Teaching lmms to overcome false premises, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey See say and segment: Teaching lmms to overcome false premises, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.420604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:18.002307Z digest=sha256:08de1b30b36b274e092758f184474b51a10a7a38dd83bb85ed0f900838378693

Observation 1682cf05-b69a-4cd6-a225-5c8f435a83c4 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Gsva: Generalized segmentation via multimodal large language models, in: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.411628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:18.077397Z digest=sha256:5f4e30de4050c02a122868ab505a01fc8369241456ca947b277d018a4ab05e82

Observation 2aeeff76-74e6-4d7e-a14a-eace992817c2 · outbound

This paper cites Youtube-vos: Sequence-to-sequence video object segmentation, in: Proceedings of the European conference on computer vision (ECCV), pp.

Reasoning Segmentation for Images and Videos: A Survey Youtube-vos: Sequence-to-sequence video object segmentation, in: Proceedings of the European conference on computer vision (ECCV), pp

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.402971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:18.152778Z digest=sha256:26e7cbdfe6a8a410b11d63ddd793e6d9499e1de9724efd09cd738d28a7a99f55

Observation a8374649-c868-43fb-a8d0-65a273149fd8 · outbound

This paper cites Visa: Reasoning video object segmentation via large language models, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Visa: Reasoning video object segmentation via large language models, in: European Conference on Computer Vision, Springer

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.395243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:18.244911Z digest=sha256:e92691fb570b31dc7072b48fb02f34c36452710f46029b3efd6c99d05daa495e

Observation a40796a9-b8ff-4ef7-8d6d-11208ac46186 · outbound

This paper cites Panop- tic scene graph generation, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Panop- tic scene graph generation, in: European Conference on Computer Vision, Springer

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.386965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:18.337726Z digest=sha256:1d16fd5b8d0f848fd638ad76c11e08d5eb92761c745d0445ea2c71df54e49dd7

Observation 0f8c2825-a571-46e7-ab1b-7d7dde838826 · outbound

This paper cites Video instance segmentation, in: Pro- ceedings of the IEEE/CVF international conference on computer vision, pp.

Reasoning Segmentation for Images and Videos: A Survey Video instance segmentation, in: Pro- ceedings of the IEEE/CVF international conference on computer vision, pp

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.379641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:18.419486Z digest=sha256:27e7ab34d4ef064839d39a8f8ab6d00a1ebc2a517c457c524740efbc53c700b2

Observation 5eadc24a-072c-40ed-a1b0-bfaa947003b7 · outbound

This paper cites Depth Anything V2.

Reasoning Segmentation for Images and Videos: A Survey Depth Anything V2

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:18.497202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:18.497202Z digest=sha256:6ca01693f2ceede9a85071d4f5f564e9f9fe3acf4f082635741aa9da80abf97e

Observation 2d846f45-b74d-4917-bef0-fe69edc5c59d · outbound

This paper cites An improved baseline for reasoning segmentation with large language model.

Reasoning Segmentation for Images and Videos: A Survey An improved baseline for reasoning segmentation with large language model

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.369615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:18.567677Z digest=sha256:5447cbf2252e77ef651dd6a6683c106cca6406ce78b8677b82c4ea36986e04a5

Observation 61a7ef6b-27f3-4b50-8a02-cc88f72dbcf5 · outbound

This paper cites Empowering Segmentation Ability to Multi-modal Large Language Models.

Reasoning Segmentation for Images and Videos: A Survey Empowering Segmentation Ability to Multi-modal Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:18.655378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:18.655378Z digest=sha256:66c317cd38ce51c8487e4391a14afaca297c8ef4ec26c0e2448442681461c516

Observation add5e5f7-086d-41d7-be11-5af856289345 · outbound

This paper cites Follow the rules: reasoning for video anomaly detection with large language models, in: European Conference on Computer Vision, Springer.

Reasoning Segmentation for Images and Videos: A Survey Follow the rules: reasoning for video anomaly detection with large language models, in: European Conference on Computer Vision, Springer

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.360394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:18.716411Z digest=sha256:a30ab4fbb3cc3639063ed670d4eb76232c31b7813977d371946c335ad08c5d52

Observation a8ad9b85-c90d-4b59-b68f-c088e2d7806d · outbound

This paper cites Lavt: Language-aware vision transformer for referring image segmenta- tion, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Reasoning Segmentation for Images and Videos: A Survey Lavt: Language-aware vision transformer for referring image segmenta- tion, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:20.350859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:27:18.789405Z digest=sha256:de6f2349a9f57932e2ea2915bb2463a65b224b9e265cccbcecb798c70b1a1702

Pith citing papers

Observation 6507a37d-5882-44da-844d-3acb78c39a8e · inbound

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction cites this paper.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Reasoning Segmentation for Images and Videos: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.967237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.967237Z digest=sha256:10ecc596b133c5f602e9f620154c71d778c158dbd3810e4c16f0d9c96adcafec

Observation 8dba12a3-d13f-4f62-8e1e-4ff91d86e7a5 · inbound

GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality cites this paper.

GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality Reasoning Segmentation for Images and Videos: A Survey

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:00.225117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:54:29.027805Z digest=sha256:04e7d2abd7e504ace75f3f5d8043f8ea5d3b49d5a4b20eb598db9c80b78363e1

Observation 99dbdaa5-f54f-43cb-989b-8cf35a603dd1 · inbound

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation cites this paper.

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation Reasoning Segmentation for Images and Videos: A Survey

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:14.156308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:40:16.265708Z digest=sha256:d43acf90b9ebc44ff2111d3623510c2e21cdf17e6bfff5c2923750f93317a651

Observation faaf51be-153a-4a61-9ba3-1c7b12e6a4a4 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Reasoning Segmentation for Images and Videos: A Survey

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.363471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:e09a789fe99ac1f3b5375b045d9250a44eefa422a1978d37395d4d1c4bff4576

Observation b5da47c9-260f-41ab-9dba-72e58d18e44a · inbound

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation cites this paper.

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation Reasoning Segmentation for Images and Videos: A Survey

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T13:38:03.546834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T13:38:03.546834Z digest=sha256:9fe6c78df42890e93649771f079873b33fc65129706f16f69742a878ce8aa82f