Pith. sign in

Paper Citation Record · LEDGER

What Holds Back Open-Vocabulary Segmentation?

As of 16 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2508.04211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.04211 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:52:34.595212Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy51
  • unresolved13
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08bccb0c-6ff9-4ea0-b57a-86b2bb22ab30 · outbound

This paper cites Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs.

What Holds Back Open-Vocabulary Segmentation? Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:45.262064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:29.063864Z digest=sha256:d587d194143ff154bd31d9268739e8fc1e1b375909c4c393f9816af8ceacecea

Observation 836284c2-f191-4e0e-a75b-72a72fa9aa63 · outbound

This paper cites Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolu- tion, and fully connected crfs.

What Holds Back Open-Vocabulary Segmentation? Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolu- tion, and fully connected crfs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:45.027960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:29.196234Z digest=sha256:23a6c257eb5282a5e4dc5c16a5639e796fcab52c607d92f397a6114e4fe6c320

Observation 9490c158-dae6-4c24-be8a-979aa108f2ec · outbound

This paper cites Encoder-decoder with atrous separable convolution for semantic image segmentation.

What Holds Back Open-Vocabulary Segmentation? Encoder-decoder with atrous separable convolution for semantic image segmentation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:29.372193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:29.372193Z digest=sha256:a28154842fe650f75c90dd9e3665b839f1ba35a8138a00bec8aa4c3ea539cb55

Observation 806af3ec-7367-49d1-b12a-6b8585006ea6 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

What Holds Back Open-Vocabulary Segmentation? PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:29.522196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:29.522196Z digest=sha256:715671b5349e049feaca606cb8c83a89afce5d42fc3e05bcc352f9f2a7ebbfc5

Observation 38cc5346-8fc7-4c0c-9949-1213fcfecf2c · outbound

This paper cites Per- pixel classification is not all you need for semantic segmen- tation.

What Holds Back Open-Vocabulary Segmentation? Per- pixel classification is not all you need for semantic segmen- tation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:44.798621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:29.645567Z digest=sha256:81d209e93c1b3145f34f062d890b8868505711d031bd2a4509998f52a58b2d22

Observation 191b2a45-989c-4ac5-a9d7-ecd897f398b1 · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

What Holds Back Open-Vocabulary Segmentation? Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:44.496946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:29.769588Z digest=sha256:aa5e22940ed22724de3565ee065419abcda27857acb8f4a585ee2f1abe6366ff

Observation b86c4654-8a19-4e16-8be3-ba7c5dda47d2 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

What Holds Back Open-Vocabulary Segmentation? Reproducible scal- ing laws for contrastive language-image learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:44.259586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:29.857281Z digest=sha256:0b29ff7118271529fbd3128d12b3de5d26a905bd44395b3d1ecba584007f214c

Observation 0e8ead4f-bce7-4f73-9fad-538cb6374bad · outbound

This paper cites Cat- seg: Cost aggregation for open-vocabulary semantic seg- mentation.

What Holds Back Open-Vocabulary Segmentation? Cat- seg: Cost aggregation for open-vocabulary semantic seg- mentation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:43.982668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:29.943467Z digest=sha256:edb18231799c0ab9c8f1d22a40169049ce2b4b1f0501b075f9988d109cd760f3

Observation ba68cbba-3e4f-4171-8486-81cfbd7e7c2b · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

What Holds Back Open-Vocabulary Segmentation? The cityscapes dataset for semantic urban scene understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:43.786629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:30.038805Z digest=sha256:f39e5874471d5cff95bc88725f1eb9cd8516b368d39526fb87f6cea3070cf10b

Observation 3428a1a7-d1b6-4b94-980a-ac9be6c4f41e · outbound

This paper cites Deepglobe 2018: A challenge to parse the earth through satellite images.

What Holds Back Open-Vocabulary Segmentation? Deepglobe 2018: A challenge to parse the earth through satellite images

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:43.649919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:30.141309Z digest=sha256:1f27a6a5a5891d64bd380ae0fdd0d4a1e6f08ecf7d79be7d0d28f3ef759d02d1

Observation d23e21c5-cdd9-4ef5-bc62-9d80e5a84c3f · outbound

This paper cites De- coupling zero-shot semantic segmentation.

What Holds Back Open-Vocabulary Segmentation? De- coupling zero-shot semantic segmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:43.434392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:30.240344Z digest=sha256:01879ad9e26d5e767a6b8735fdd7aa4c9abea99b66e2aa8a487b9b2d3c28d94b

Observation b6a04b08-0d49-4c0f-bdee-f522c436d2c6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

What Holds Back Open-Vocabulary Segmentation? An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:30.331238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:30.331238Z digest=sha256:c4a2ccc022f0ad4365de52d4417bcad567947bf342fd399b41b3b03d00ad7a5c

Observation b671e2d3-bbda-4482-b01f-b35146de6900 · outbound

This paper cites The pascal visual object classes (voc) challenge.

What Holds Back Open-Vocabulary Segmentation? The pascal visual object classes (voc) challenge

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:43.192557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:30.451894Z digest=sha256:9c2cafdb3c3de14e6a2c2a637c4c88ca406fb1b31b79b873ccd0565196cc1758

Observation 803d38f5-55ea-470f-97a1-6d2f039e898a · outbound

This paper cites Data filtering networks.

What Holds Back Open-Vocabulary Segmentation? Data filtering networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:42.949925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:30.565010Z digest=sha256:b28f4162aa6b53dae49b45c1df0b3e40b053c0dd54319a7bdccd0dfec312f71b

Observation b169b8a2-47d1-4296-b7cb-2e6e936684e2 · outbound

This paper cites Scal- ing open-vocabulary image segmentation with image-level labels.

What Holds Back Open-Vocabulary Segmentation? Scal- ing open-vocabulary image segmentation with image-level labels

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:42.773828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:30.657678Z digest=sha256:48fb88029f6ad500449df727b804d6168d98270fc6565ace5022871b4d7000fb

Observation ed2a5457-1b21-40d3-98ef-d176f69ad4a8 · outbound

This paper cites Fast r-cnn.

What Holds Back Open-Vocabulary Segmentation? Fast r-cnn

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:30.726218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:30.726218Z digest=sha256:0b16fbfc8db2853a814bc39861132fac06dcdbaaf742765830de5f4b6dee0600

Observation 28f0228b-916a-437b-8bcb-46114c6a224e · outbound

This paper cites Rich feature hierarchies for accurate object detection and semantic segmentation.

What Holds Back Open-Vocabulary Segmentation? Rich feature hierarchies for accurate object detection and semantic segmentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:30.798509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:30.798509Z digest=sha256:041491a888ff74adb9f54a9c4fba879a5beadab4b6a393094d67af72c0dc82fb

Observation 1e996761-a417-4339-b74e-a6c084f6b76f · outbound

This paper cites Open-vocabulary semantic segmentation with decou- pled one-pass network.

What Holds Back Open-Vocabulary Segmentation? Open-vocabulary semantic segmentation with decou- pled one-pass network

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:42.602852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:30.889096Z digest=sha256:3e7ec9b463cd3a001dd0b798a711629b4ba335b1a6c5c39d5807fac362795b17

Observation 8928180e-a749-424b-94c3-48294938e91a · outbound

This paper cites Simultaneous detection and segmentation.

What Holds Back Open-Vocabulary Segmentation? Simultaneous detection and segmentation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:42.424121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:30.953642Z digest=sha256:72a97fd7e7ee10c631ba08e9d50c5969f3518ccb5e431bc97b90f141eeeffae2

Observation 7b531e68-4a63-4063-91ee-4e5b4de7c130 · outbound

This paper cites Mask r-cnn.

What Holds Back Open-Vocabulary Segmentation? Mask r-cnn

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:31.050425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:31.050425Z digest=sha256:c31de5b5782e9ca2abe36c194a7ed67384d5552bba4f531bd4641976d5911dc1

Observation eb57fcd8-7492-4ae9-a712-4f9bb45c9809 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

What Holds Back Open-Vocabulary Segmentation? Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:42.257739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:31.139595Z digest=sha256:48d4d39873420ee493565020e934faebd74049dd302e6deb67290b913dac149d

Observation bbfe1c44-789c-4386-9490-e9c8499326e8 · outbound

This paper cites Collaborative vision-text rep- resentation optimizing for open-vocabulary segmentation.

What Holds Back Open-Vocabulary Segmentation? Collaborative vision-text rep- resentation optimizing for open-vocabulary segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:42.097447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:31.230619Z digest=sha256:f47753a959c8c03afb271044cbbaae9be20c860916a62620c2c90537f88d96e6

Observation 582cd08c-efe5-4ffb-a139-d357dc1c09ad · outbound

This paper cites Diffusion models for open-vocabulary segmen- tation.

What Holds Back Open-Vocabulary Segmentation? Diffusion models for open-vocabulary segmen- tation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:41.938234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:31.333947Z digest=sha256:82ba58c2de4da1f323fb3587548170a50b0e85614f8389c0f45721a8b9c918ac

Observation b0e45d05-49fb-4239-967d-437a7bdff061 · outbound

This paper cites Panoptic segmentation.

What Holds Back Open-Vocabulary Segmentation? Panoptic segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:41.770170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:31.400137Z digest=sha256:6ffa158f6ee2ee9d28d03e98e01c6587413045f2383f68d3499d6cfd869b771d

Observation f9097f64-90bd-45d4-a0c6-1c5b4c294a00 · outbound

This paper cites TIPS: Text-image pretraining with spatial awareness.

What Holds Back Open-Vocabulary Segmentation? TIPS: Text-image pretraining with spatial awareness

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:41.605623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:31.476294Z digest=sha256:308c11d77d94dd643d32120ba2ed73df34a4bbec0075d4141a47a9b43c66afdb

Observation 477dca5c-e68f-415e-9416-57a2f288c93c · outbound

This paper cites Ladder-style densenets for semantic segmentation of large natural im- ages.

What Holds Back Open-Vocabulary Segmentation? Ladder-style densenets for semantic segmentation of large natural im- ages

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:41.435465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:31.576470Z digest=sha256:1972f7c1e9e853078f6dccc10fe3f03c1148aadb57d52dc2363a4b90e7371850

Observation 4491d09f-d931-4c9a-8077-51e366e8a527 · outbound

This paper cites Clearclip: Decom- posing clip representations for dense vision-language infer- ence.

What Holds Back Open-Vocabulary Segmentation? Clearclip: Decom- posing clip representations for dense vision-language infer- ence

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:41.352324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:31.644242Z digest=sha256:ba60c87bb5811342e8647ba833d9409080d280ef8f17d77b5cb3b404feabc9e4

Observation 0e0f052c-91ce-45d1-86ae-46aed6951a9b · outbound

This paper cites Proxyclip: Proxy attention improves clip for open-vocabulary segmentation.

What Holds Back Open-Vocabulary Segmentation? Proxyclip: Proxy attention improves clip for open-vocabulary segmentation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:41.177292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:31.721672Z digest=sha256:d9270734e57579801d10b894cf463d07e8d9e3ad5578858871074145b5f9d0e8

Observation 7febfaa7-c6ec-411e-ac58-35f10061e988 · outbound

This paper cites Mask dino: Towards a unified transformer-based framework for object detection and segmentation.

What Holds Back Open-Vocabulary Segmentation? Mask dino: Towards a unified transformer-based framework for object detection and segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:40.999444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:31.813349Z digest=sha256:ac3a4516a05fb3e75e527e8329a8d90023787702d1bd27ee3141ff77fab92d7f

Observation b63f0c58-d510-47d8-8436-0afb64394c38 · outbound

This paper cites An inverse scal- ing law for clip training.

What Holds Back Open-Vocabulary Segmentation? An inverse scal- ing law for clip training

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:40.854883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:31.962056Z digest=sha256:e58b85e4dd7df1b1c7deb04ec1ef40b8f982b7939a354e1910fb35dea1cc7805

Observation 8e31c34b-255f-46d7-829e-7ef98889aa10 · outbound

This paper cites Scaling language-image pre-training via masking.

What Holds Back Open-Vocabulary Segmentation? Scaling language-image pre-training via masking

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:40.701891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:32.057501Z digest=sha256:2af702685d354f66b5a7b762cc8f1590cd0853cd2d2d7198697ee8195e362483

Observation a6293139-0be3-4d4b-8f0f-b2935d7c2a7f · outbound

This paper cites Clip surgery for better explainability with enhancement in open- vocabulary tasks.

What Holds Back Open-Vocabulary Segmentation? Clip surgery for better explainability with enhancement in open- vocabulary tasks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:40.572494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:32.190338Z digest=sha256:e48a75e6b9f63941bbbe2db2afa3de6932ec180dafd3d1bc45c08a16a2070a9c

Observation 528ef065-1a13-44b8-aa2c-30253bcbda81 · outbound

This paper cites Mask-adapter: The devil is in the masks for open-vocabulary segmentation.

What Holds Back Open-Vocabulary Segmentation? Mask-adapter: The devil is in the masks for open-vocabulary segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:40.447287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:32.236708Z digest=sha256:a81a283aee6e9fd2ddf82fa99c845743f81eb18de4448f7bf06e4f40c4e32518

Observation bdbb79db-9480-45c4-b2e8-d5de898b0779 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

What Holds Back Open-Vocabulary Segmentation? Open-vocabulary semantic segmentation with mask-adapted clip

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:40.274097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:32.294092Z digest=sha256:9ac946469ccd6501273a5da1a07c68c81ed6efda9072f562205f1c58ab8cabae

Observation a475f8be-bc5c-48a6-8619-975b54c639b7 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

What Holds Back Open-Vocabulary Segmentation? Open-vocabulary semantic segmentation with mask-adapted clip

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:40.125545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:32.381810Z digest=sha256:e13a6be8275c48010c31f42378dc6f39c37d7ba1d6fbd5815bfe4f4f4f8b16c2

Observation 39e0a733-9c49-49b9-b587-4c69162f4bc3 · outbound

This paper cites Microsoft coco: Common objects in context.

What Holds Back Open-Vocabulary Segmentation? Microsoft coco: Common objects in context

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:39.964322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:32.473145Z digest=sha256:4e669b9636a13e8a0d4fb7b9102e2fe199d0f91f99120002514af2cc755082c6

Observation c6eb9401-e27e-4d16-bac8-804b39a030fb · outbound

This paper cites A convnet for the 2020s.

What Holds Back Open-Vocabulary Segmentation? A convnet for the 2020s

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:32.581521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:32.581521Z digest=sha256:37c141ae3d30c9c76fe06911f960fd811fdd40ce70f6f0006315ee2e5d8c5626

Observation a4d70e7d-22b5-410f-9134-7ac259c53b70 · outbound

This paper cites Fully convolutional networks for semantic segmentation.

What Holds Back Open-Vocabulary Segmentation? Fully convolutional networks for semantic segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:39.547599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:32.652442Z digest=sha256:0711ce28d6020c22ebc0674923f4ce2e32ac0baf8ca027b242a0b224e2eb128f

Observation a3a73361-f290-4831-869e-faf0fe434571 · outbound

This paper cites DEARLi: Decoupled Enhancement of Recognition and Localization for Semi-supervised Panoptic Segmentation.

What Holds Back Open-Vocabulary Segmentation? DEARLi: Decoupled Enhancement of Recognition and Localization for Semi-supervised Panoptic Segmentation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:52:34.783235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:32.734021Z digest=sha256:359e0e3995cdd2b932c225009e2c143cd9dc7a4cf4538f40f069a2bda9b6be4d

Observation 9e9685cb-ec8b-4ecc-a3f0-fac3c5f2ba6e · outbound

This paper cites Open vocabulary semantic segmentation with patch aligned con- trastive learning.

What Holds Back Open-Vocabulary Segmentation? Open vocabulary semantic segmentation with patch aligned con- trastive learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:39.375088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:32.831892Z digest=sha256:433635ff31d4b28969ef59315bcfe7ba52157e71e93cf44c203f809592c1c588

Observation 5d0c50ea-f768-48e0-a300-6db953751d91 · outbound

This paper cites Silc: Improving vision language pretraining with self-distillation.

What Holds Back Open-Vocabulary Segmentation? Silc: Improving vision language pretraining with self-distillation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:39.067966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:32.878193Z digest=sha256:4f727b51b322b415f4b71ac737331548910f342564a69c872b7b6d4064ea1e30

Observation 3aa5fea6-f1d4-4906-960b-c919d4f88d33 · outbound

This paper cites The mapillary vistas dataset for semantic understanding of street scenes.

What Holds Back Open-Vocabulary Segmentation? The mapillary vistas dataset for semantic understanding of street scenes

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:38.809930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:32.937921Z digest=sha256:e3b58db8aac233dcb750fa07a812776994b807c713af72dc23be424ee463586f

Observation 2018411c-89ea-4c54-8328-24697be00b76 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

What Holds Back Open-Vocabulary Segmentation? Learning transferable visual models from natural language supervi- sion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:38.630423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:33.074332Z digest=sha256:4978733facf9e222a844a392052d837b009a1dcf507c53390f4fc9a6bd876daf

Observation 11137e0e-28dc-4b4a-b836-327fee72a551 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region 10 proposal networks.

What Holds Back Open-Vocabulary Segmentation? Faster r-cnn: Towards real-time object detection with region 10 proposal networks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:38.451570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:33.169790Z digest=sha256:d612eba89d31cab932f08c3fe4845c65b1311d35c03ad68f55969d2b5923338d

Observation c5604ea4-ce5d-458b-b4c1-ef69dec4152f · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

What Holds Back Open-Vocabulary Segmentation? U- net: Convolutional networks for biomedical image segmen- tation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:38.221021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:33.237726Z digest=sha256:e6b081c65b28bfca37ad0f3b95aa8de774e92d90973e06a8502d9779b9950047

Observation a8068a9a-a40b-40db-b9a6-cc7bc9746531 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

What Holds Back Open-Vocabulary Segmentation? Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:37.974048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:33.297753Z digest=sha256:59893947c3e42897141af420a8f422feb2025d17e760d56d484102672deb39ec

Observation 45d2c4fc-cb69-406c-95b0-d4f156f6fb3d · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

What Holds Back Open-Vocabulary Segmentation? EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:33.344662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:33.344662Z digest=sha256:0c7017745801e2d2dfa32ebc8322f08dd77960217447367a780487f6db6abb42

Observation 70e6db1c-7dcc-490a-a551-7aa169efb490 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

What Holds Back Open-Vocabulary Segmentation? SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:33.410715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:33.410715Z digest=sha256:a352d32605148dbc3a2260d4288fc06b83b2134ce2742e4e36b80a143ff5af7e

Observation c88a9c5d-c065-4fe1-b737-8ad4fde952dd · outbound

This paper cites Sclip: Rethink- ing self-attention for dense vision-language inference.

What Holds Back Open-Vocabulary Segmentation? Sclip: Rethink- ing self-attention for dense vision-language inference

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:37.752484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:33.492924Z digest=sha256:392e7bfa2b6e6448fc784a847b36d32c994d7f7bc3375364d3fb6aba35e8a730

Observation c2bd8ac5-750f-44d1-bd2f-0c67b30f4a03 · outbound

This paper cites Max-deeplab: End-to-end panoptic segmentation with mask transformers.

What Holds Back Open-Vocabulary Segmentation? Max-deeplab: End-to-end panoptic segmentation with mask transformers

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:37.540092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:33.552118Z digest=sha256:1918011f09b6bdcff23559c01fe13a7afa12e05b4f5fda66891c98db45e10b03

Observation c76e95e2-6bc2-474d-b1d4-d9f15dea236b · outbound

This paper cites Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions.

What Holds Back Open-Vocabulary Segmentation? Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:37.290839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:33.615800Z digest=sha256:7c1b4482d3e2feb0782c36aa815c7de4f8ea487bb3311461ca210c9d2e73f2cb

Observation 05ab4961-86c0-4bb0-962d-9764b056adc4 · outbound

This paper cites CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction.

What Holds Back Open-Vocabulary Segmentation? CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:33.694344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:33.694344Z digest=sha256:1c91fccd433fc268642dab94c89e9d08ab41f3151b54e779ecb47c61391ce45b

Observation 77c76741-f118-462a-a791-8a604c287ea7 · outbound

This paper cites Sed: A simple encoder-decoder for open- vocabulary semantic segmentation.

What Holds Back Open-Vocabulary Segmentation? Sed: A simple encoder-decoder for open- vocabulary semantic segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:37.081625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:33.778792Z digest=sha256:356e7e5dfddb633434def0f18c4f29d3c40a5f8e00f8db2e643eaabdabacca54

Observation 20bc680c-a86b-4c4c-95a3-433049a44c8f · outbound

This paper cites Demystifying CLIP Data.

What Holds Back Open-Vocabulary Segmentation? Demystifying CLIP Data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:33.858924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:33.858924Z digest=sha256:e39d8f8f63a3fcda7b2dc2abe017059b8fd81fa387adbb73162c5ee67c96e990

Observation 93334985-0d14-4e54-b9ef-d32c1ac5be10 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

What Holds Back Open-Vocabulary Segmentation? Groupvit: Semantic segmentation emerges from text supervision

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:33.926973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:33.926973Z digest=sha256:f9fa7469249a8df7e89d36ae1d864cc1b6d0f67e73bb2a348c17c1771fe79ccc

Observation ff5eb810-0a4c-40a1-bfb7-9b5b3008d332 · outbound

This paper cites Learning open-vocabulary seman- tic segmentation models from natural language supervision.

What Holds Back Open-Vocabulary Segmentation? Learning open-vocabulary seman- tic segmentation models from natural language supervision

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:36.887396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:34.013450Z digest=sha256:101d027c61e198a88491902d4a014cbcd09674377c1c4d7d2449b5bf4a0aa0d1

Observation d83e1658-cb16-4259-839f-99cdcec37ce1 · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

What Holds Back Open-Vocabulary Segmentation? Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:36.682637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:34.090426Z digest=sha256:7ac512085f05f337728fd7f6bfa57787110c1675a2f8aeb92991fb3b392473a7

Observation 4171ffa6-65a6-44fd-b784-e765653b5266 · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation.

What Holds Back Open-Vocabulary Segmentation? Side adapter network for open-vocabulary semantic segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:36.445159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:34.162129Z digest=sha256:111446f489c2a3a6e0d3de4d11ddc8050b12908db1b98448fb3b5ecfd7b490e6

Observation 61698048-ca34-41ef-b0c6-5d71330e0cee · outbound

This paper cites k-means mask transformer.

What Holds Back Open-Vocabulary Segmentation? k-means mask transformer

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:36.230408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:34.249365Z digest=sha256:54e7302f4b9ff978293e2892855f2f067bbf7439881e18e6124f4fae3533e33b

Observation 81525f0d-9bc9-46d3-95d8-9e82f4efd066 · outbound

This paper cites Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip.

What Holds Back Open-Vocabulary Segmentation? Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:35.965193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:34.303491Z digest=sha256:cf8b7733a80e974ea923755f705903aa24fc3d2d0996eaf6127e589815ed6182

Observation 3e68a02c-c5da-417c-bb19-73f39cd6b51f · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

What Holds Back Open-Vocabulary Segmentation? Lit: Zero-shot transfer with locked-image text tuning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:35.717899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:34.356877Z digest=sha256:38b39f80a54776633e90ebc8c0f9d650a893a5c37524347b4c8b24857819aa01

Observation 4bdd62b8-597d-4771-a03c-e0ee645b8ff3 · outbound

This paper cites Sigmoid loss for language image pre-training.

What Holds Back Open-Vocabulary Segmentation? Sigmoid loss for language image pre-training

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:35.475649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:34.413901Z digest=sha256:d689098c2f3fe29a0d428391efe817c522c0587f567c2d438a87290eb4e9a959

Observation 16226b19-de43-419b-be29-22bd2c83b221 · outbound

This paper cites Pyramid scene parsing network.

What Holds Back Open-Vocabulary Segmentation? Pyramid scene parsing network

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:34.493510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:34.493510Z digest=sha256:ee917b5bbf02edd85fda8f4bda456950d7cae8e964fe21dcde7fd7a16f537d4f

Observation 65388892-487d-40ee-b5a5-11a02ad08bbf · outbound

This paper cites Scene parsing through ade20k dataset.

What Holds Back Open-Vocabulary Segmentation? Scene parsing through ade20k dataset

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:35.190577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:34.528413Z digest=sha256:17da5aeb1592338025ec56f7055e96495fee17c9c7f8db95b2417334fe1ee40e

Observation 0afb6c39-3ebe-45bb-804c-ce44ccc79abd · outbound

This paper cites Extract free dense labels from clip.

What Holds Back Open-Vocabulary Segmentation? Extract free dense labels from clip

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:52:34.939613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:34.595212Z digest=sha256:45864e25d851bcd85fd58312c317a24d6e6933f0f0d3083401abd11e88cea4a9

Observation 650fa148-71b0-4ef2-a8f0-2b8fddfef6f2 · outbound

This paper cites an unresolved cited work.

What Holds Back Open-Vocabulary Segmentation? Unresolved cited work

Reference 755

Resolution
parse uncertain
raw_fallback, observed 2026-08-06T00:52:39.793598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T00:52:32.516086Z digest=sha256:fdcbde1e7a544424a20588934739ce4b183b30d852d3bdfc8cfa8f0748621b87

Pith citing papers

No inbound Pith citation observations are available.