Pith. sign in

Paper Citation Record · LEDGER

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model

As of 12 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 2 inbound Pith citation observations for arXiv:2412.00153.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00153 v3

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:10:41.196244Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:43:32.954256Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.913171Z

Reference resolution

94 of 94 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc9054c9-867b-423a-b426-2fa2669f5baf · outbound

This paper cites Deep Learning using Rectified Linear Units (ReLU).

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Deep Learning using Rectified Linear Units (ReLU)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.602727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.602727Z digest=sha256:ffd933370c270a75a1a3beec4c9c7ff5b9d417b19bd594c03c79adf783e8079d

Observation 33724053-7262-4073-b3cf-99c545c9044c · outbound

This paper cites Barron, Fer- ran Marques, and Jitendra Malik.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Barron, Fer- ran Marques, and Jitendra Malik

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.608496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.608496Z digest=sha256:0d897bcf313cb15dc116559f2f39a4ac0086b7cc9b950ce5f4cdd7097801a374

Observation 406c57ea-401b-4377-82dd-566577064de0 · outbound

This paper cites Qwen Technical Report.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.614159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.614159Z digest=sha256:92cf82308aebaf3eb2863f5169e6bac9c43097ef38db3070664f8e872879bde2

Observation 8e606117-7d9f-443f-bb03-db732569e074 · outbound

This paper cites CoReS: Orchestrating the Dance of Reasoning and Segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model CoReS: Orchestrating the Dance of Reasoning and Segmentation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.619538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.619538Z digest=sha256:49a55f3fdade15401d4fca1225124cdd255c37095b894fe72dca4fe1b4b1c705

Observation 8670c5ab-c675-4e1c-a7c2-8de69265d678 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.625056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.625056Z digest=sha256:9a5f1ab164be3d4f481b46021a0a25aa9b2084b5154d6ceeaa18f9504877195e

Observation baea3694-043d-48ac-8e92-fe3d7526c3b3 · outbound

This paper cites Coco- stuff: Thing and stuff classes in context.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Coco- stuff: Thing and stuff classes in context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.630566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.630566Z digest=sha256:d33ac10e182c8574e1aafe44b09aa09a7c587e2b5e6df3614d20bf8b25081eaf

Observation a4dae4b5-f4a4-496e-acb3-bec70a780927 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Emerg- ing properties in self-supervised vision transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.636951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.636951Z digest=sha256:40cb80dba8bc714814853d01582c6867aee5f3a65c02075e6e8f0aa7f1559e42

Observation 01ddf6a1-4087-4e6b-96da-84fe45560491 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.646884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.646884Z digest=sha256:6d34d6c8114ec2d968b307519fb663c4c59416df1a52326af9ec29daa2acb963

Observation ee8b1da5-7b65-49ca-b49a-e87c59f6ab99 · outbound

This paper cites an unresolved cited work.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.655072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.655072Z digest=sha256:ab44fdb08c47c35ff3237b357dee11409bb711cd951c2ed65d2532d35db076a4

Observation 7c3a0701-eb07-4aa1-a8e6-ac61d243cd8a · outbound

This paper cites Rethinking Atrous Convolution for Semantic Image Segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Rethinking Atrous Convolution for Semantic Image Segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.663763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.663763Z digest=sha256:7ca0f434072af913c732271ccd35bb9007b3c6d4f29c79f696d7c21d5f597111

Observation 63341173-fefd-42f3-bdce-6bfaa1798857 · outbound

This paper cites A Unified Sequence Interface for Vision Tasks.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model A Unified Sequence Interface for Vision Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.671530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.671530Z digest=sha256:988dcb437ef57bc4ccb0375e67d8abdd49d09835d4da1eae3ff3641a503c90b0

Observation b7eb7c55-f6dd-4d02-bb83-2848e14e3f8e · outbound

This paper cites Schwing, and Alexander Kir- illov.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Schwing, and Alexander Kir- illov

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.678258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.678258Z digest=sha256:db18903c6a26070e7245135fb03bbacefd5d3a260dfb561edd5541c9a2b48510

Observation a79df8a7-3099-4d6b-b159-806618062944 · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.684396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.684396Z digest=sha256:f1eb6e7e663505825b70368924be4e4c79945adee0f9928c4118bcffc74d1e7a

Observation f8a1f219-05f4-46f2-8f6b-6705ee8b53ce · outbound

This paper cites CascadePSP: Toward class-agnostic and very high- resolution segmentation via global and local refinement.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model CascadePSP: Toward class-agnostic and very high- resolution segmentation via global and local refinement

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.690426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.690426Z digest=sha256:39fed60650ee5307928bab387765219875a4b9a7c047c5080c2501e33fbfb3b2

Observation 472714cf-0a71-48e2-ad16-e20305fe19bd · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Gonzalez, Ion Stoica, and Eric P

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.694974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.694974Z digest=sha256:c622eec699774e4eb7566985b52747a49d30a3670671b478be3a4e9c9b5335bb

Observation 0cf0912c-d82e-4403-96eb-81220347d1e6 · outbound

This paper cites Instance-aware se- mantic segmentation via multi-task network cascades.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Instance-aware se- mantic segmentation via multi-task network cascades

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.700692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.700692Z digest=sha256:4f6177f57d563c422cfa51077f126407c5b35826031d9e5b5118daad02f9845a

Observation f9fde1d8-d406-4036-b6e1-d06598d11d26 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.708748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.708748Z digest=sha256:9aa096826fb3f9a403e12cd2cbe2cc2a16cfbecf3a2ee7526b6b7243c3982926

Observation 9b38e67c-72ca-45f5-9857-f083d57431a7 · outbound

This paper cites Vision-language transformer and query generation for refer- ring segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Vision-language transformer and query generation for refer- ring segmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.720146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.720146Z digest=sha256:742ade1e361d7680ffa171ff35bd3779368a03df7ebcc02ba51555280e1b720b

Observation 86d48a15-4367-4cad-b139-7a0fd8be5078 · outbound

This paper cites De- coupling zero-shot semantic segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model De- coupling zero-shot semantic segmentation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.725466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.725466Z digest=sha256:fac98ff1681810d1474f6cd08daae22730b9a363274cd9e6093cf2f3729acc5e

Observation ce5e425e-f964-4565-afbd-cfd6c5d71098 · outbound

This paper cites A discriminatively trained, multiscale, deformable part model.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model A discriminatively trained, multiscale, deformable part model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.732244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.732244Z digest=sha256:700b01c4dae99d360c9c15f3870b048024bceb846cbafd156a6835ebfdbc0f78

Observation 712ad1f6-051a-4afe-b9ae-0cf631fd18fa · outbound

This paper cites Dual attention network for scene segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Dual attention network for scene segmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.738344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.738344Z digest=sha256:c765721e80165e0ebf4dbfa45596f644c3a9a49edebbf3476dfce0ce717a3387

Observation 468bd23d-d2f5-4d5f-a882-d82b4d1655ea · outbound

This paper cites Scal- ing open-vocabulary image segmentation with image-level labels.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Scal- ing open-vocabulary image segmentation with image-level labels

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.744388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.744388Z digest=sha256:d84a77acc79ab33a5faa3706c3338067822a54e3699c9b9be6a894d57b591c82

Observation c3f0ae87-5a05-4d50-a933-63407e84642a · outbound

This paper cites Fast r-cnn.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Fast r-cnn

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.829966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.750583Z digest=sha256:ee25551a947302908cebe81c1e07e7c18ffda474e2be5acadc3622426fd6694a

Observation 4e1b00fc-59f0-437c-8b0b-7adb9f285eb7 · outbound

This paper cites Efficient hierarchical graph-based video segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Efficient hierarchical graph-based video segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.813092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.757926Z digest=sha256:ac5eceee21f8d305a64e308eac2f7ddecbe8267ed4c2a13e1e8959990668ff61

Observation feadfcc1-fa40-435e-9f00-49e411c620cf · outbound

This paper cites SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.764202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.764202Z digest=sha256:dbdff0358fb63dc356c7c53e68180d9877f4a87ed1833ce795d30dc6f4b1b077

Observation 3d406e68-77a8-4547-8ecd-a6be74544aa0 · outbound

This paper cites Global knowledge calibration for fast open-vocabulary segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Global knowledge calibration for fast open-vocabulary segmentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.794206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.770557Z digest=sha256:de9ffe0a1a5596460dec766e9267f929ea970d8c7e48577769541837b730d9e8

Observation 695588bf-8ac2-4c5f-b90a-4ce4999a5c09 · outbound

This paper cites Multi-modal instruction tuned llms with fine-grained visual perception.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Multi-modal instruction tuned llms with fine-grained visual perception

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.772644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.779725Z digest=sha256:fe034d54b4fa597e40a12158862459b6707100494f4b27180087135058d877e7

Observation 8328d994-4c35-4f3c-855e-87f1083541ca · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model LoRA: Low-Rank Adaptation of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.787127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.787127Z digest=sha256:8d10cb8dbf460de77904cf1cc3c59b9418232d40bdfd396513c8d6f52eebad7b

Observation 667a8c61-474c-4eb2-bdc1-a78bf072d809 · outbound

This paper cites Bi-directional relationship inferring network for referring image segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Bi-directional relationship inferring network for referring image segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.751208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.794615Z digest=sha256:cc51b2cf6dd001f8b10ce1e3cba728878ff25e5c46c3f8e940f4ea151c7483a4

Observation f72b2860-a780-43f2-94a4-44a5abbbbb04 · outbound

This paper cites Referring im- age segmentation via cross-modal progressive comprehen- sion.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Referring im- age segmentation via cross-modal progressive comprehen- sion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.727435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.801817Z digest=sha256:476ff61c7f3e27db6ae03cc986a9b0a90ca5a51a7257968ca68ad0e5307c41d0

Observation 8592435d-0e8b-42ed-a9fc-ac42b5963117 · outbound

This paper cites CCNet: Criss-cross attention for semantic segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model CCNet: Criss-cross attention for semantic segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.699887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.807249Z digest=sha256:a3d6750021fe0b7cf8fea619a18fddd1a9245d9dc853e86e161606cd682b0774

Observation 9b3796b3-aed2-4e82-aa22-ac393def43ef · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Scaling up visual and vision-language representation learning with noisy text supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.817090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.817090Z digest=sha256:086bfbcf5052b9ccf10c38e7e116b52c4fee93ca8ce6bce6e5b62cea5a01e218

Observation f72a46c9-8ce4-4987-be4d-80882b00fa95 · outbound

This paper cites Collaborative vision-text rep- resentation optimizing for open-vocabulary segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Collaborative vision-text rep- resentation optimizing for open-vocabulary segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.655968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.825753Z digest=sha256:7d396e6b2eedc0634ff9ae18149beac69035a57e4de34194ebcd6643af7f5d5c

Observation 7fdc2374-2927-4d20-89ef-ce062cd36847 · outbound

This paper cites Locate then segment: A strong pipeline for referring image segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Locate then segment: A strong pipeline for referring image segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.634113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.830882Z digest=sha256:ff95168cd868c7eb1287bc086017bba381e5daeb93d4b1327f2e398415277128

Observation 06feee68-e2b9-4bf9-a4f7-05bc6e827007 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.613695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.837314Z digest=sha256:dfa3aa75ac92d1a52746c98c3c137d4ee5f1b1991bf9ad7e80fcd3c11ee33d96

Observation 6433f450-c112-4599-a060-8f5da956b3be · outbound

This paper cites Segment any- thing.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Segment any- thing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.593290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.845266Z digest=sha256:ea08469fc14634ff8adac42102401cb874bf4ec45e7de99a50b2ef67831c51f2

Observation b2c2b052-bab4-4579-a33b-62edd58b93a9 · outbound

This paper cites Large language models are zero-shot reasoners.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Large language models are zero-shot reasoners

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.569494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.852552Z digest=sha256:1fb64ef8bce1a545a9d790b3587492553360c1d6facf9091e47f20ec31d8191c

Observation 83edde41-c518-440b-80c9-ed98b8846423 · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Lisa: Reasoning segmenta- tion via large language model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.859899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.859899Z digest=sha256:6e171d5bf34c055b7277c7936a83dfd9b7a8ca73f378f11d1f0a7b4bbabd939e

Observation c36fba00-d6af-4aeb-86a7-9886842fa00a · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.867634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.867634Z digest=sha256:e386e0fe635320a6077e9603d02a2949d58ecaf973fb6ae3703641608d98b7eb

Observation 877046cf-543e-44f7-9091-4932d560821a · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.529184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.873795Z digest=sha256:b61e85add4725bb75cd68beb12fdb214e15aa086e1c08b91a212ab00d371eec2

Observation 690077fb-17d7-47a2-b76e-8a6bd67a5223 · outbound

This paper cites Referring transformer: A one- step approach to multi-task visual grounding.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Referring transformer: A one- step approach to multi-task visual grounding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.509319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.879409Z digest=sha256:c700d5d7d214efa125e8d7b11ae9e0bc611d62ff1257f47f81cc882512a6d8a8

Observation 72a2b959-46dd-4dea-a489-e37fced6d16b · outbound

This paper cites Textbooks are all you need ii: phi-1.5 technical report, 2023.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Textbooks are all you need ii: phi-1.5 technical report, 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.486608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.885283Z digest=sha256:93d1b4dad68f201004b0046142d90962a87556a1e954e0d09063ab6792b92d38

Observation dec07de0-38ab-4eaa-ac37-913a67839e81 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Open-vocabulary semantic segmentation with mask-adapted clip

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.452323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.892528Z digest=sha256:08a47499bea0580fdcc526449947a7276adce22500d92e60f0433a5a3de2f5f7

Observation 07ea0b25-15aa-4478-8a64-45321df358fc · outbound

This paper cites Microsoft COCO: Common Objects in Context.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Microsoft COCO: Common Objects in Context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.897561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.897561Z digest=sha256:db73ff220449d1139b86ac1334446cd97a268a11dde6a2b16a3f4a9aae9d90e8

Observation 32b269f7-7bff-4b8b-9bff-f80d8a639b46 · outbound

This paper cites Visual instruction tuning, 2023.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Visual instruction tuning, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.903182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.903182Z digest=sha256:cc608685558fa736cbf86d251c56fb48ce9b10bf0f01628005f9044d5719f47a

Observation 12d2280a-d8f1-466d-9559-30e2cbbcb6aa · outbound

This paper cites Fully convolutional networks for semantic segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Fully convolutional networks for semantic segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.410159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.908369Z digest=sha256:52e461a61e6cae74d785f99521488230fb8159555317e21152961329d168673a

Observation 6e189871-22a0-43f1-b894-35761da8fbf6 · outbound

This paper cites Decoupled Weight Decay Regularization.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Decoupled Weight Decay Regularization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.913998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.913998Z digest=sha256:f507125a6b0eda89a8d7dba23a9741d1b2ceb195a35d9abf197d926b8ff4c5fd

Observation c2ab0f93-7bd6-42cc-91df-66e21610c6b1 · outbound

This paper cites Cascade grouped attention network for referring expression segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Cascade grouped attention network for referring expression segmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.388448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.919613Z digest=sha256:9a80342f2fa3a89686ee1415d4042dcf356407114cb56ffef05565df41aae3f9

Observation 104f859b-3744-4b76-8055-2c3b3d9115d3 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Generation and comprehension of unambiguous object descriptions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.366413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.924413Z digest=sha256:0e36ad92d9f15502aac3d8b94597a3d694c9b571f8a182bc79f9a1e8d2e2f0d9

Observation 872b78b8-0097-48b5-9244-de8b596374c7 · outbound

This paper cites The mapillary vistas dataset for semantic understanding of street scenes.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model The mapillary vistas dataset for semantic understanding of street scenes

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.930956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.930956Z digest=sha256:9304a084414137f33bf1396ab82a01ca98ccc9bf57c6869db78d5b59e18e30f7

Observation 019f3a6a-ded3-4691-9e91-19bfd3467ebf · outbound

This paper cites Chatgpt: A language model for conversational ai.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Chatgpt: A language model for conversational ai

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.319108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.935743Z digest=sha256:57103afdb349c445b835a8039a7bdebf1a848ecfbe313bae32eb9960c5e5a673

Observation dac3c2a1-6ded-4d4f-8369-3e8dc3e16b12 · outbound

This paper cites Gpt-4 technical report, 2023.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Gpt-4 technical report, 2023

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.944616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.944616Z digest=sha256:35f52825627e452797644173335c4188627b8b2d4bd388cf47f433fc10d1eba5

Observation 34c97eb1-eb05-4fbc-b68e-a2a15804a35b · outbound

This paper cites Kosmos-2: Ground- ing multimodal large language models to the world.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Kosmos-2: Ground- ing multimodal large language models to the world

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.278398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.951145Z digest=sha256:35dc401869e69493000ae744fb0447525917083592727c609be6a881b9112a7a

Observation 39d0a7dd-ec4e-4fa7-82c8-81ee1552a315 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.957551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.957551Z digest=sha256:a995df4114cdc6ed948426187126a4b3da41c91872f2bb1d52a4ca412368a67d

Observation 4484be1f-82de-4194-8d0e-62c05722f88a · outbound

This paper cites an unresolved cited work.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:10:42.255454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.963755Z digest=sha256:aec1021199c40ca0d33fa147e4515b5406b06cb178c38eb9dfcd60624ae5f88a

Observation 770cf0eb-9593-48f3-9639-6196fc696c0a · outbound

This paper cites Pinheiro, Tsung-Yi Lin, Ronan Collobert, and Piotr Doll´ar.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Pinheiro, Tsung-Yi Lin, Ronan Collobert, and Piotr Doll´ar

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.235786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.968878Z digest=sha256:dd9647bd468ea799f8cedbe8b61d476a611e6370a0cdd7281ea310356aea709a

Observation c36e3f83-f330-46be-a92a-f67190285945 · outbound

This paper cites Multiscale combinatorial grouping for image segmentation and object proposal gener- ation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Multiscale combinatorial grouping for image segmentation and object proposal gener- ation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.217544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.974233Z digest=sha256:2a65efb4ddf67ff82c4858827f73406f03998dc1b174b19c19b7363b8b97fb9b

Observation 57339680-1a03-4d27-85a4-2380a7780b00 · outbound

This paper cites Learning to segment every referring object point by point.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Learning to segment every referring object point by point

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.197893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:40.980144Z digest=sha256:b8e868cb11197ce2a7923da1054191e3903a6b88447c50867faf33708ab83633

Observation 59a4f07e-f360-429a-b3d5-51a7e8aa0aeb · outbound

This paper cites Learning transferable visual models from natural language supervision.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Learning transferable visual models from natural language supervision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.986159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.986159Z digest=sha256:b5ff92d6da6c31708530535e2df4886790d63c1bb838a95d1fe76e443e76e0f2

Observation d5f58f64-088c-4ebc-88cd-82efb08cd80c · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Learn- ing transferable visual models from natural language super- vision

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:40.993676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:40.993676Z digest=sha256:6d6c69cc8626a84f429cb838ad730e51726b700fbc0c434da1be3ffa419e43df

Observation 95f21376-6334-48a6-8ff1-4d766b914a40 · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.129623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.001120Z digest=sha256:92d822274b689a8444b68261ad28ff06aa590f4eb26b5f7e2a7010550c9b50bc

Observation 3e3e2a8f-2770-47de-9ed6-c23a7c8d921d · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.104654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.006137Z digest=sha256:b8cfea882c87698fd89d71cb6fc5c187a9eec6c5cfc105bed9fe42c11f9979be

Observation 6084e481-980e-45df-863d-ead855e42446 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Pixellm: Pixel reasoning with large multimodal model

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.078881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.011579Z digest=sha256:c2df9216b29f15e1fe5dbba087ee587585f8bb49920ad59a7769410c27e4dde0

Observation 091d006e-4246-4292-9933-ae0e13200fe1 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model U-net: Convolutional networks for biomedical image segmentation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.050276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.017480Z digest=sha256:c5e3b7915c51a5b1441e5f07b3b4a2b3166fe60a9c0859da345337aa2e1922a0

Observation b01dd00a-2855-4aad-ac96-e22bd1db2263 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model LLaMA: Open and Efficient Foundation Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.022299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.022299Z digest=sha256:a652af7992d6f7796e0c5539366cd38fc68d7e5e7d08eebf42f3aef8718b761d

Observation efaf51df-9525-4fd6-ae6d-369540f4f34d · outbound

This paper cites Selective search for object recognition.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Selective search for object recognition

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.024766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.027703Z digest=sha256:d09054f4d07856bd2a21c1e574891296046e74643cf13de421251e37abd7f5e2

Observation 4c2effc2-46d7-4be0-96d6-1d5b96e2e933 · outbound

This paper cites Llm-seg: Bridging image segmen- tation and large language model reasoning.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Llm-seg: Bridging image segmen- tation and large language model reasoning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:42.002608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.034013Z digest=sha256:74f6dde79b4ae4aeda1cecfaaa27204df5e806cf89d627a2725b33739daf7b54

Observation 2a539345-3b75-4d8e-b1c7-e22198fa5d20 · outbound

This paper cites SegRefiner: Towards model- agnostic segmentation refinement with discrete diffusion process.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model SegRefiner: Towards model- agnostic segmentation refinement with discrete diffusion process

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.985143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.040649Z digest=sha256:2056acdfcfa4f401d22ec5bcabd2e17d9f34f2ec4b0606de0226debfcec396d1

Observation 9096e147-9792-4ea6-8cbb-84a8b06d9cec · outbound

This paper cites VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.049912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.049912Z digest=sha256:c89592d6b53b756dd93f9d0078f58dd303956a3950416bfd821ba104de821831

Observation b7b7176e-bea8-4b06-9999-8cea01f96982 · outbound

This paper cites The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.057624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.057624Z digest=sha256:a41ee16a28315cca8f09fdf1eaecedc1dfec26aeb1375b15b6f58ae4385bf40e

Observation 7f637c6f-b367-4b98-bea7-b68c8f2122a3 · outbound

This paper cites Non-local neural networks.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Non-local neural networks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.063230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.063230Z digest=sha256:5690d55f5fbfa6192fe1ffb64a987adbd2f6205a47fa1284d017b1340b95b106

Observation 0374e256-6842-4722-ae9b-0d97a4800603 · outbound

This paper cites SOLO: Segmenting objects by locations.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model SOLO: Segmenting objects by locations

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.956507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.068431Z digest=sha256:3882840684e5652ed24ad8222be24d5fb5dcfa5436b9418e5c2d935404fb7580

Observation 97186277-4ff4-47fe-a2ae-70c769f44316 · outbound

This paper cites Solov2: Dynamic and fast instance segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Solov2: Dynamic and fast instance segmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.939830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.073957Z digest=sha256:ad40f512d775f9afb4640b7630689c6686b0c52252a2d5b10f03faeafbd945a1

Observation df85ce4d-d2c9-4611-9ebd-2dc1a549d400 · outbound

This paper cites Images Speak in Images: A Generalist Painter for In-Context Visual Learning.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Images Speak in Images: A Generalist Painter for In-Context Visual Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.080050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.080050Z digest=sha256:ebfb88396cae012bfa778ac410f3d4cd5d832320d9b5679c92444ef1e24d29ca

Observation 125b212e-419c-4753-8ec3-e38a42756858 · outbound

This paper cites SegGPT: Segmenting Everything In Context.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model SegGPT: Segmenting Everything In Context

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.085777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.085777Z digest=sha256:b031c22fa42edb1a83754a503e19f3536dea373b66678005fe670b556bb75ad0

Observation 59d13369-d98d-4855-b1fd-daaf17007b9c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Chain-of-thought prompting elicits reasoning in large language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.922973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.091613Z digest=sha256:a9b8d5592b624f7e1435cf60fdb512bf6c2783133efc8b90010aa3894ef45b1e

Observation 9ec67287-7a75-47a5-86a9-b6c218785426 · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Gsva: Generalized segmentation via multimodal large language models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.904733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.097842Z digest=sha256:f73fd8d50c8aebc4c899557d0b296b6d49f48211fec8b1cb3b8d45f8ee16c05f

Observation 6104d1f6-685f-4683-bf9a-d92dec4c4a89 · outbound

This paper cites Alvarez, and Ping Luo.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Alvarez, and Ping Luo

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.104345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.104345Z digest=sha256:34ce0c232562532f60b82d010f6216928642f5048dad0f777c8e195f3ee43a14

Observation 4704578f-c2ad-4907-9140-7b5966b79f5d · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.109789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.109789Z digest=sha256:a74ffa372b3b58ba419794045d6bb2bfe65bd223d3020411cb18a4ecd042e6c2

Observation d1d07908-f8ea-431a-8854-4ab39d1af1b0 · outbound

This paper cites A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.859266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.115102Z digest=sha256:1551b8dd812b8bb98b2d88f48fc7933626a0d7b30e798540dbd014860758f00f

Observation 730d6885-8b49-43ac-81d3-67320b97a0be · outbound

This paper cites Fine-grained visual prompting, 2023.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Fine-grained visual prompting, 2023

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.840574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.121101Z digest=sha256:26295f40731f8a461be64e97414bb85ab059ca3f125dfd4e4911ba3b74ef7e60

Observation 68db3912-cd72-4051-b9a4-11fd34e18873 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.126078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.126078Z digest=sha256:7dc2f93c7f8b49bdab4b7921a2133070ae076630c4b8ee2612e6949b1f5a9820

Observation 0a056f1c-adde-4864-8440-8f745078ba01 · outbound

This paper cites Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.822743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.132233Z digest=sha256:c2576d1e1279959aae520c550b81d77257c05c6a399e744d46d4cecad68a0344

Observation 1cc26dc8-4858-4a84-acd7-69b225e49105 · outbound

This paper cites Osprey: Pixel un- derstanding with visual instruction tuning.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Osprey: Pixel un- derstanding with visual instruction tuning

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.805366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.138463Z digest=sha256:35d2442eef7f47eacd96d5a15d71cdbda937517796ef8fc323c797c32b9be15e

Observation 548302c0-d01d-4a26-9eac-6ec52d771371 · outbound

This paper cites Sigmoid loss for language image pre-training,.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Sigmoid loss for language image pre-training,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.144523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.144523Z digest=sha256:f0a3d9602ef840956ef3ceec94a4e6377f939d16d3e6d836dd533554d8bd2456

Observation 7556df97-c4ee-4b7d-b8ab-f513b2f3bb51 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.151302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.151302Z digest=sha256:a213dee4425aac9891eb9ef9fc7b2c80dd17eb8b87cf8b31b06472bba546f8a2

Observation 5923443f-74a8-41ea-ae6e-c52518152f71 · outbound

This paper cites ControlVideo: Training-free Controllable Text-to-Video Generation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model ControlVideo: Training-free Controllable Text-to-Video Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.157888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.157888Z digest=sha256:db99c8d773db71a7d76d91932a957799dae99c03fc0d559310f23b28264c36f7

Observation 891d1b46-5fc9-4f41-a6f8-4fb9f8447cdb · outbound

This paper cites Groundhog: Grounding large language models to holistic segmentation.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Groundhog: Grounding large language models to holistic segmentation

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.775048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.163653Z digest=sha256:3725b0788d8e2379240c2579e0149b1abd3b9301c0b09b30efd2f8b1deb07f86

Observation a43a1024-e009-476b-aede-3d65b8b1cb3a · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Psalm: Pixelwise segmentation with large multi-modal model

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.756026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.169375Z digest=sha256:dcabfe60df0279640be788f742ec7bc0e6006680bc872ff502c0f8defb4c36a8

Observation 46b773f9-a729-4c19-962d-315d1de1023b · outbound

This paper cites Rethinking semantic segmen- tation from a sequence-to-sequence perspective with trans- formers.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Rethinking semantic segmen- tation from a sequence-to-sequence perspective with trans- formers

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.739593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.174302Z digest=sha256:45048b98d6777e0733eeffd0d731f5e858b76f0bc1c72c31ca17f2b2cd6835ca

Observation 6e715acd-3bbc-4c16-88b8-ecfbf0d317ba · outbound

This paper cites Scene parsing through ade20k dataset.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Scene parsing through ade20k dataset

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.722980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.179846Z digest=sha256:8dd259ae7661609758d6fcc0854bab53276769dc15ef3e880f0b46cce483ece9

Observation cf837357-61af-430d-ac08-29f3e48b1a29 · outbound

This paper cites Seqtr: A simple yet universal network for visual grounding.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model Seqtr: A simple yet universal network for visual grounding

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.704525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.185531Z digest=sha256:00fab0decaa361ec7d884494644d226f1094feaf26fde626b30997eecba4061d

Observation 25354d67-8d8e-4442-96ce-34acb6d4f0c4 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:41.190681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:41.190681Z digest=sha256:d938a51a63366384cd1d652ab3781826049196e91c4a1325a1089feb9174f786

Observation 9380f26d-e32a-4e8c-b96d-46447ed59b62 · outbound

This paper cites User: <IMAGE,MASK> Please segment target region with mask and corre- sponding category.

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model User: <IMAGE,MASK> Please segment target region with mask and corre- sponding category

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:10:41.686647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:10:41.196244Z digest=sha256:bf8a16feb57801e45ee9855ddda73dcc666c83139a64b55dd6039106aac6b81a

Pith citing papers

Observation 5a541e7d-6bdb-4f7b-8188-4377cd6f44ad · inbound

Bias from small-scale leakage in Pulsar Timing Array maps cites this paper.

Bias from small-scale leakage in Pulsar Timing Array maps ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:32.954256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:32.954256Z digest=sha256:baa079a442a375992f8e547ea8b4526bb29615219dc7eb479d6f366f5c0be17b

Observation 115d2920-c1bf-478d-af38-9de4301a7d7e · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.914750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:382fe7fdaa8f4433e229da1fe4d497b9414ed26431291e831b0b50e41311ebe7