Pith. sign in

Paper Citation Record · LEDGER

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction

As of 16 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 2 inbound Pith citation observations for arXiv:2605.20110.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.20110 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T05:23:32.202297Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:26:29.380459Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T00:19:13.707397Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact9
  • verified fuzzy68
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3eb87b22-ca74-4fd6-a80d-0625405867dc · outbound

This paper cites Qwen3-VL Technical Report.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:28:05.178224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:423bd8663f438c5fdbba06e1ed9ba00fe161fd525c984f78e15fff860c037a0d

Observation 76a43aa4-fa76-4894-a94b-cc6ef11e4f70 · outbound

This paper cites One token to seg them all: Language instructed reasoning segmentation in videos.Advances in Neural Information Processing Systems, pages 6833–6859.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction One token to seg them all: Language instructed reasoning segmentation in videos.Advances in Neural Information Processing Systems, pages 6833–6859

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.343595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:e87de416acfcd44885defc14bd91690f7fee74444351f092d2e5d72225aaed07

Observation 537249cd-a9b6-4950-bb87-f14a52f8a6b8 · outbound

This paper cites Sam 3: Segment anything with concepts.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Sam 3: Segment anything with concepts

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.485453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:e7a7b2d2879a6bd596ded8df7998885d06129128edb926fde2db619367daa368

Observation b84410cd-2886-4581-b02f-e5dd1de15705 · outbound

This paper cites End-to-end object detection with transformers.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction End-to-end object detection with transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.420852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:44b15205e985481693e1c2eee22dd179870a6dce83218ac98b7e6acd1d1d0857

Observation b79a0389-ccdb-473c-8194-0d3e54b72a31 · outbound

This paper cites Sam4mllm: Enhance multi-modal large language model for referring expression segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Sam4mllm: Enhance multi-modal large language model for referring expression segmentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.505752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:fd821550fb038c76bd589a0a63b636a51507e85cd0d24b9de3ea4bf2a9a0f5df

Observation adb172a5-1d0f-476b-bde4-aa1fc1c2ed7e · outbound

This paper cites Samwise: Infusing wisdom in sam2 for text-driven video segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Samwise: Infusing wisdom in sam2 for text-driven video segmentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.584627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:941a43508db097ccfc5318160bf52c3ded75553c7c1ab09ae1a42d6f2e5452a1

Observation 4dd02138-6991-4721-a4f2-7517e19b8fc8 · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.547045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:f5db5482509890888892de3169da19fc664e2b007e0d4c963e0da9d1d6d63a68

Observation f0f2a0f5-8269-4eda-a32c-4a73f2323b20 · outbound

This paper cites Mevis: A multi-modal dataset for referring motion expression video segmentation.IEEE Transactions on Pattern Analysis and Machine Intelligence.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Mevis: A multi-modal dataset for referring motion expression video segmentation.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.345611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:ce3a3ebb1a4f74dc30c9100624cbb48b069239172969efe86e00697e34293f7e

Observation 738c484c-458b-4d6f-8100-226a69cd7b6e · outbound

This paper cites Vlt: Vision-language transformer and query generation for referring segmentation.IEEE Transactions on Pattern Analysis and Machine Intelligence, pages 7900–7916.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Vlt: Vision-language transformer and query generation for referring segmentation.IEEE Transactions on Pattern Analysis and Machine Intelligence, pages 7900–7916

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.445528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:af7d7f5bd48803cff49effef5e59b78fd7231c5bb31ea41d898d05329dcc33c2

Observation eba51d53-d74b-4f53-876f-594b053c060d · outbound

This paper cites Sam2long: Enhancing sam 2 for long video segmentation with a training-free memory tree.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Sam2long: Enhancing sam 2 for long video segmentation with a training-free memory tree

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.448119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:31b2e8cbd69ed286b7cebff523565e550b9d70a8334778bd364b49eff2dd4425

Observation f8c2b1c0-d4c3-4e68-b209-7911ffe5e812 · outbound

This paper cites Language-bridged spatial-temporal interaction for referring video object segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Language-bridged spatial-temporal interaction for referring video object segmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.350038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:441e71d49bbfc55e5baa5b81080f88b9030129e718dfc8933149bb45b972b433

Observation 0030514d-a3ea-4d51-a3e7-bdf183d20c61 · outbound

This paper cites Palm-e: an embodied multimodal language model.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Palm-e: an embodied multimodal language model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.346036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:fdd172984b2dba52fe07fae8efce16996bd266a506316b2441138c259e5a2df3

Observation 65221503-1609-4bc9-af69-db1be65ff5b8 · outbound

This paper cites The devil is in temporal token: High quality video reasoning segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction The devil is in temporal token: High quality video reasoning segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.338496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:14b3cb74214edc29693bfba35c4b1a61f2489e8728d8f9c67cd390982a26a973

Observation d5632eb5-18fe-4e92-b789-a9b13b3c8bfd · outbound

This paper cites Anomalygpt: Detecting industrial anomalies using large vision-language models.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Anomalygpt: Detecting industrial anomalies using large vision-language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.580828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:253069b5a1393985a4b1a5e75f303deb34d247dbe9d9447fc306986560936e12

Observation b932d839-ba8d-44d2-99dc-3c6007306863 · outbound

This paper cites Html: Hybrid temporal-scale multimodal learning framework for referring video object segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Html: Hybrid temporal-scale multimodal learning framework for referring video object segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.471242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:b7f40ffea398d347167121c2ea2a46843f551faec8dc78af8897ee3a81a4cd4d

Observation 70349b3b-6e2d-485f-9644-c7d2ba1eebfc · outbound

This paper cites Decoupling static and hierarchical motion perception for referring video segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Decoupling static and hierarchical motion perception for referring video segmentation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.332193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:e462381dc795d1bfecefe19229b486224597244a89c1cae6b3854b3ebf69819f

Observation 8b2e50b9-1027-4c6c-a371-caf6020b12df · outbound

This paper cites Segmentation from natural language expressions.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Segmentation from natural language expressions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.483796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:4eab66388870d8d470f2d7932eb0b5c2530fe03de1c86d75624a5e5a7ceabf96

Observation c817f34c-adf7-4e4d-9535-e79b01cffcd4 · outbound

This paper cites Beyond one-to-one: Rethinking the referring image segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Beyond one-to-one: Rethinking the referring image segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.311963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:682b1f04c69f36a6c9cd2eb4b5a3689f30054ec7735b19e713d31cfaae63bbf9

Observation 4f5d4a88-bb00-459d-90bc-c0eaaac131f5 · outbound

This paper cites Densely connected parameter-efficient tuning for referring image segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Densely connected parameter-efficient tuning for referring image segmentation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.500977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:ee45c4c5e712a121776bf8a102a22b757b6db1b5ced4606627b7bb90ca821209

Observation f5788793-933b-49b6-ba1f-163b149463a7 · outbound

This paper cites Mmr: A large-scale benchmark dataset for multi-target and multi-granularity reasoning segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Mmr: A large-scale benchmark dataset for multi-target and multi-granularity reasoning segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.490898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:d8f8b0b724bb37c3b854374c96c5a5e8dd8273f6b282ca67c9f94baa81489a92

Observation f3999a03-7346-464d-966d-97fe73e50016 · outbound

This paper cites Mdetr-modulated detection for end-to-end multi-modal understanding.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Mdetr-modulated detection for end-to-end multi-modal understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.429307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:eaf9fa4c236ba79798e867e84fbe491d11655278f0142c5ed16dc0fffc1a7e24

Observation 9efdeed2-871b-44f1-845b-42d3e1208133 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Referitgame: Referring to objects in photographs of natural scenes

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.542934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:a8cdb502d9562999c9fe986e3e6ab8bd5a97b3de3bd2bcb4e84a0d03d5cd447d

Observation 19624c22-4726-42c1-830d-1c948fc3fbb2 · outbound

This paper cites Video object segmentation with referring expressions.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Video object segmentation with referring expressions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.487777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:45d64772aa89b9ee67dc8acfdecc575078861aa1bc19c8500dcbf1e096d5736e

Observation 3936b7c1-8714-46ea-b373-06499cc60387 · outbound

This paper cites Segment anything.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Segment anything

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.476625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:29a2bb785b10d8fc62fd06b8ea7453c75f120255caea3991115b22d8fb8ea5a3

Observation b9ad5cbe-4a8d-470e-8645-70df876210ea · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Lisa: Reasoning segmentation via large language model

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.467243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:7ac377a295b48fe77a14da2b2243044b95aa6d833de5b6420327fc54b35057b6

Observation fdc7504d-0d66-48c7-b07f-19182bec6635 · outbound

This paper cites Text4seg: Reimagining image segmentation as text generation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Text4seg: Reimagining image segmentation as text generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.475718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:358af469a7a59eb01d501f78181ba2af543a91bb3b57dca7c80b728d7e6f63e0

Observation 19bc26ee-9c9f-4c10-b491-539122e0fced · outbound

This paper cites Grounded language-image pre-training.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Grounded language-image pre-training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.527409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:000270a841547c2ba7aab08c17a0cea77f3af6f2c72d1a656641120239d2a954

Observation 81cef537-b0dd-4be5-840e-8d5eda63d0b2 · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Gligen: Open-set grounded text-to-image generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.523279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:f80f8ced388ccd52d8c2ed4a075a609ca41fad9db7d60d8c4eb9fc9f01c80ff4

Observation 60a05966-5391-4ded-bcda-7b5eb35ac078 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Open-vocabulary semantic segmentation with mask-adapted clip

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.576332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:cec39811ad189b711fb116f2fa2d05a272ad9ffc17ce85d8002686f1b13694c4

Observation c6d23df5-ee9e-4003-b5e3-cc0a71c9982e · outbound

This paper cites Glus: Global-local reasoning unified into a single large language model for video segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Glus: Global-local reasoning unified into a single large language model for video segmentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.538266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:a56a03c95ac665540cb72f40f9e770e2a3d6526760a6fcc2e841b4eedb0ea6f2

Observation b6194e08-f90d-4b0a-abb2-a6ffd5d65145 · outbound

This paper cites Gres: Generalized referring expression segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Gres: Generalized referring expression segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.327410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:c963adc05aa1ae00cb72fcbc811cc929789aef39c8d225694a8bcbf4b69a5ddd

Observation c83c8857-48a1-4776-8a01-57631030330c · outbound

This paper cites Recurrent multimodal interaction for referring image segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Recurrent multimodal interaction for referring image segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.466234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:549073fb3995fa30922fb73acb99a4b45610b6c5f8c421c0db6f7cfcce4c2285

Observation dbca0f50-7a81-4537-b173-a6b8d49d175a · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.331518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:f2eefe851423a166e853bb6d3d4b1c36c746e5cd99ee1b2bf66b9c278d5dadd0

Observation fe49c8e2-58e5-42b8-872d-80c810ca4bcc · outbound

This paper cites Universal segmen- tation at arbitrary granularity with language instruction.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Universal segmen- tation at arbitrary granularity with language instruction

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.374780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:9e93750dafde7efd6a215f2dad1ef589629a7f6a6965551bf635a31c274d6fdb

Observation 80303571-30e8-4450-9717-77803b3819d2 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:28:05.247212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:93fda08fc0f49677f738e8b2e3066dafd47d220d656679b988c13f7a84cf6a06

Observation 2dd39f48-4e75-447e-8bfd-b9b880995e5a · outbound

This paper cites Visionreasoner: Unified reasoning-integrated visual perception via reinforcement learning.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Visionreasoner: Unified reasoning-integrated visual perception via reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.429006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:9173f5a394ef017719a5324d64e86cb76e53e0e5064476cd716086b0ee12ded4

Observation 834d53f6-83ad-4dba-a2d5-c104decc2f56 · outbound

This paper cites Rsvp: Reasoning segmentation via visual prompting and multi-modal chain-of-thought.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Rsvp: Reasoning segmentation via visual prompting and multi-modal chain-of-thought

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.515127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:b4501f0efd2bfb701ed0ea89c205dc8acedd8a48cbb23d9ba1bacde53fc7be4c

Observation a695338d-6c50-43dc-9a4c-ba3fb332c8df · outbound

This paper cites Image segmentation using text and image prompts.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Image segmentation using text and image prompts

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.531600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:db663aa218ae4d9ddd25b480787d6334fa2280beec1d62b4a6e730e0ebe8853a

Observation 2c35af56-c239-48df-b971-b5e6e1a632ba · outbound

This paper cites Soc: Semantic-assisted object cluster for referring video object segmentation.Advances in Neural Information Processing Systems, pages 26425–26437.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Soc: Semantic-assisted object cluster for referring video object segmentation.Advances in Neural Information Processing Systems, pages 26425–26437

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.463494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:993df66a082f26d8ce129e947e55dd8f295fbb951812791aef994e19d9743090

Observation b1a6f9c8-4336-4365-b003-6975ea7f4923 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Generation and comprehension of unambiguous object descriptions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.370850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:f8bcf592f0f7d67386d03d81df044f90462bbcc828cec47f74e66e3859de87dc

Observation 672f137d-8739-46da-bae2-9a535781bb13 · outbound

This paper cites Spectrum-guided multi-granularity referring video object segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Spectrum-guided multi-granularity referring video object segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.333841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:990c7461502ef1366f13ecd511cea02bfcde120bc14952c4944a38152e530c42

Observation 1f66f7e8-060b-4b29-9689-8f660dfbf472 · outbound

This paper cites Simple open- vocabulary object detection.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Simple open- vocabulary object detection

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.497379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:d75480f82d860a476dd3c386a42fbcb4dc7f6ef40e4b2e696d6feb8776ecfb43

Observation a67dd0cc-66af-40fc-b756-9ff5700963ab · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual grounding in videos.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Videoglamm: A large multimodal model for pixel-level visual grounding in videos

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.354382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:898cd5fc1730c530e56976349c163803d37cfee712a9a0fa222bd0f18a9f31f4

Observation 3bc5c6b4-7b0c-4ab4-b1bc-54d25f6b324b · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Glamm: Pixel grounding large multimodal model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.323072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:9e7e1511e367581a57bd66c58677118a65c62e92951c6a8608bdf615ba96564c

Observation e236a7b2-2272-4002-940f-2fe5098c548b · outbound

This paper cites Sam 2: Segment anything in images and videos.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Sam 2: Segment anything in images and videos

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.314400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:91c1d5464d752c147ba5925b88c3193a3332ed0b4c14cd40260f1bedce8d567a

Observation 0f35b2f6-cb35-4e32-83e4-828eb8f77faf · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Pixellm: Pixel reasoning with large multimodal model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.309741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:7badd3a3f0f67914b126377b88ec2d46c22de9731e55f2af6f495a038c7a19b8

Observation bcd25cdd-85c4-4e87-b2c4-8fc166944435 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.557239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:cf7964d3c2d986fad8e64b4ad826d66a3c3777410cb866e7f023e6bdeea01531

Observation 234bfc0a-998d-41e6-87bc-17a797e80f6c · outbound

This paper cites OpenAI GPT-5 System Card.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction OpenAI GPT-5 System Card

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:28:05.225879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:1af6219cab4a583ee58569df19fd7880630e7d8d10eedf0e7bce04228933e6ca

Observation 20161db5-25b2-4471-800d-308692c2420a · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:28:05.232614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:c714c2f895a9a7239709610820f992be69483560187e153bf19c9d23e7363bc0

Observation 03943dfa-0224-4878-be9e-c5d863b2ab67 · outbound

This paper cites Videoanydoor: High-fidelity video object insertion with precise motion control.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Videoanydoor: High-fidelity video object insertion with precise motion control

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.380574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:dbb64e8b8487fb7a02f453c4602c49b8a9db194e89ec819a469019fa3d297fa8

Observation da2f5a62-fd98-47ae-b2f3-7f058730f80e · outbound

This paper cites X-sam: From segment anything to any segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction X-sam: From segment anything to any segmentation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.519240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:128a3688ee5d4a585060197c7b9f91f06fb5aeb67dde18d7f7c61d4e87e74fd6

Observation 4d8528cd-c447-4f7c-ac86-a99839fa929f · outbound

This paper cites Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:05.220398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:5be27b2a3d7634d3a077c8b220e83b2c1a2a1821d539277d0c8e7157925e1554

Observation 7db59666-b6f0-46df-847c-40edae6a3cbf · outbound

This paper cites Un- veiling parts beyond objects: Towards finer-granularity referring expression segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Un- veiling parts beyond objects: Towards finer-granularity referring expression segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.423872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:c420c0ad4c88e86fdc67590a70d1ba839b65ea62d961deaf823103cefadd0ff5

Observation d7bbf721-287f-4ca6-a965-1586a5e4ff77 · outbound

This paper cites Deforming videos to masks: Flow matching for referring video segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Deforming videos to masks: Flow matching for referring video segmentation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:05.254034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:28b0af77e62623f737074c2a74737846f5ae87b6d0687eb8e78918332a9a02d6

Observation 7be87b18-f613-439b-83b8-ec5a536e582d · outbound

This paper cites Hyperseg: Hybrid segmentation assistant with fine-grained visual perceiver.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Hyperseg: Hybrid segmentation assistant with fine-grained visual perceiver

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.327341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:e30e57569d689b45614a8b2e8ea1fda17574677ed2e0165ed5d02eb14cf03631

Observation a8279055-4d80-4130-8b67-e99c39476047 · outbound

This paper cites Instructseg: Unifying instructed visual segmentation with multi-modal large language models.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Instructseg: Unifying instructed visual segmentation with multi-modal large language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.340829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:b10851165265a27eb3093c7e92a6463fe6a5e8cfba664ec0d33aecf625f7c839

Observation acfc1fc8-4642-44fb-afd1-a42dacd1d863 · outbound

This paper cites Onlinerefer: A simple online baseline for referring video object segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Onlinerefer: A simple online baseline for referring video object segmentation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.336138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:20aeaaf9590918296545760a4c08fa85ac8f8c8a0f1343d0a9832ecf668eb2a1

Observation f4e5db51-586b-4f44-b5b5-598e4646e75b · outbound

This paper cites Language as queries for referring video object segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Language as queries for referring video object segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.393762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:44f44f3ac2f08fc4f6a9b1f681390889b1947bd4339606a6b3e57e1d7581f4c5

Observation 116540d0-c105-4183-be42-9e954e2d362e · outbound

This paper cites General object foundation model for images and videos at scale.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction General object foundation model for images and videos at scale

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.461913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:8c551aaade1ef59aec64d91d55f871a131d4965baf809b6daf897b797a220106

Observation df6a6a94-570c-478d-8ca7-9a71c2f48965 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Gsva: Generalized segmentation via multimodal large language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.457593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:056bf037981e19f0c724c9820e7bf9473144e7a3271bc2cb3a70ad78a09f2e3c

Observation 76775ba7-cdcd-4bd6-bb84-777c3a63fb9a · outbound

This paper cites Region-based cluster discrimination for visual representation learning.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Region-based cluster discrimination for visual representation learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.319424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:387ac4b027bf89134d635cde5d90627ba4f5ac0de8ba5b4fc4f34b2952f12674

Observation 62034e00-7601-49b4-bf7a-e811216db534 · outbound

This paper cites Viddar: Vision language model-based task-detrimental content detection for augmented reality.IEEE transactions on visualization and computer graphics.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Viddar: Vision language model-based task-detrimental content detection for augmented reality.IEEE transactions on visualization and computer graphics

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.570131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:fc61a9f0c9023f79b62a940b5b71b5ca6f09cb67187c97972e37a3e6beda67d9

Observation 9546c80c-1e60-47b5-8233-748e50e06bb4 · outbound

This paper cites Visa: Reasoning video object segmentation via large language models.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Visa: Reasoning video object segmentation via large language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.510040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:6b4343daaafa042eb29a31b08c2a6e77b85a2ebc834a899dc4bccf0e529dd05b

Observation 81062f97-ad17-4ed4-8ca4-92d2b11d0419 · outbound

This paper cites Lavt: Language- aware vision transformer for referring image segmentation.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Lavt: Language- aware vision transformer for referring image segmentation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.323410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:99702b25bb5c8a75d46ce800678a1d1cfc85839c8a097567fb2568effd1b384d

Observation 16ae8d62-a0a5-4cce-80f9-7d80c41808cc · outbound

This paper cites Mattnet: Modular attention network for referring expression comprehension.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Mattnet: Modular attention network for referring expression comprehension

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.320645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:9d39df6b9ceec74fef027e6ef0c4b883f9897029137537dfe12b94361967bb8f

Observation 52b6c867-b539-4fce-9dd7-63f49708b868 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:28:05.189889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:d9877b3a87a7260b26ef69b5ab697fd540d08be26f43d07769f34b22747b4810

Observation fcc88c27-156d-444c-ae29-fa6c9f709434 · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.439311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:47ca32a09c01484bd8cd1de4d68fa6c14b0b90a3d9134503d1317c01bb2e18dd

Observation c36a9fb7-f75a-4e05-bf97-86dff3564869 · outbound

This paper cites EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:05.240709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:fd4508553a61899bcae6462b77d01db2a2153a3820ac68c9ab434752049b3cc4

Observation 13afa7f0-fe46-441f-a54b-7d2a3ca7572f · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Psalm: Pixelwise segmentation with large multi-modal model

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.318635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:9d10b8346e1bb31f355af37626cfbed580f26e38cdc0f4512a6549fcbf24ca00

Observation 349c8681-0bb2-4139-93a2-95bbafd5d4c2 · outbound

This paper cites Sec: Advancing complex video object segmentation via progressive concept construction.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Sec: Advancing complex video object segmentation via progressive concept construction

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.325321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:70d7582a0e2a90d8393c5a25fdd3b864472e0a9440a97372076fba0dc1833cbd

Observation 4e0a5877-bd1d-43d9-bcf8-1291b17b8e2a · outbound

This paper cites Villa: Video reasoning segmentation with large language model.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Villa: Video reasoning segmentation with large language model

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.316618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:5f0b2284c77703137f8448d8b18b2c9265b7d9eba5c9b16d61688699c31484b3

Observation e058cb1e-2a8c-41ea-9554-3320c1480a85 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Regionclip: Region-based language-image pretraining

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.384599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:286d172b5d0034fe44388360f91bc74f95c1a0804e11c414d36e0b57e2582d12

Observation d47a6581-3349-4eb0-85c8-2920f6ea51b3 · outbound

This paper cites Tracking with Human-Intent Reasoning.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Tracking with Human-Intent Reasoning

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:05.184646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:8b61e2b39a73aff0b6189ce7fe01f8a40ea35f847412532b8ef03d0e5d4c70e1

Observation 55fdb31c-232f-46e5-88c3-99b24d8d29c3 · outbound

This paper cites Training-free spatio-temporal decoupled reasoning video segmentation with adaptive object memory.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Training-free spatio-temporal decoupled reasoning video segmentation with adaptive object memory

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:48:05.329461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:4cb485ca002b6494a75a4cd3bb5c7cf35aa62ca6a32b1a18d4722a46eb9f8b55

Observation f98126f4-41db-4573-9b9c-a300445a68e5 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.456728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:d487611080fd8902854951c5686bb08fd32785806168d84e172236317ace011e

Observation 3c562250-8c72-4648-85d5-acefb3f00ea9 · outbound

This paper cites Generalized decoding for pixel, image, and language.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Generalized decoding for pixel, image, and language

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.480873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:38ba3deee364fc81772ba74c9533930486634999f6c2c00350b0da817d7b08a5

Observation e6957b40-145e-4170-95a5-3e24172a1587 · outbound

This paper cites Segment everything everywhere all at once.Advances in neural information processing systems, 36:19769–19782.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Segment everything everywhere all at once.Advances in neural information processing systems, 36:19769–19782

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T05:43:25.434250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:c69ff9230e5d9d4f37e806d96915402590135ff05d8109adb5618c9eaa7b439a

Pith citing papers

Observation 0db5790d-6502-408a-aaf6-ed4cafb42c71 · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:19:13.708634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:ee4cbc814fac404b01e2e382146df8ac3330fd493518f779cc9e0ca6519711b9

Observation 0d7abf58-df95-434f-8ac0-5523cb524756 · inbound

Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds cites this paper.

Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T04:26:29.380459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:26:29.380459Z digest=sha256:a8e7420ceaa13da78db7c33eeff84e65ed4b1576cfff1180deb1157bca7a7059