Pith. sign in

Paper Citation Record · LEDGER

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

As of 13 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2411.13836.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13836 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:53:31.889301Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:04:44.106636Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T22:21:15.797283Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy42
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8d4222c-c7a9-42dd-8f17-90c0bb26a74d · outbound

This paper cites Weakly su- pervised learning of instance segmentation with inter-pixel relations.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Weakly su- pervised learning of instance segmentation with inter-pixel relations

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:35.104096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:29.953476Z digest=sha256:8c4425fbbca204807942c3bc5fa950554c2ffa6ceff8777d59fd2943924c5841

Observation 6c6c5a6a-84a4-4045-aa37-4967a6e6d61d · outbound

This paper cites Zero-shot semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Zero-shot semantic segmentation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:35.009572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:29.960642Z digest=sha256:22f467159c4601207ebec976d4818526f076734fb194ff15f55010a7f68e4dc1

Observation d18a48e3-a586-4cd5-b360-d17f0e54e0a9 · outbound

This paper cites Coco- stuff: Thing and stuff classes in context.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Coco- stuff: Thing and stuff classes in context

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.916255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:29.967487Z digest=sha256:32526fed6f259417da3c6d52f31526cc3c87cb8d527d2b0b66c8de69efb3af52

Observation 197c4d1a-34e0-4973-abff-865a2fdd9132 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Emerg- ing properties in self-supervised vision transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.852705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:29.973143Z digest=sha256:121dae4eef88f7ff0433931dc3171df4aabca9797f05b38339ba55d9bfcccda5

Observation 2fb29648-8d04-4c9b-b931-1e7a56e978bd · outbound

This paper cites Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.806179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:29.980180Z digest=sha256:9d440b14d75e652803433f6a5acdf129f835a1cceec707d238de2152617cf0db

Observation dfd3e05b-4488-4f30-9402-a96375d05028 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Masked-attention mask transformer for universal image segmentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.742847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:29.992618Z digest=sha256:6b9c66730bd95363309c204c3d4381bff1292e29684a697d88d81ef9ef5b226c

Observation bfa636b9-7eb7-4292-983e-64c112e2cc83 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Reproducible scaling laws for contrastive language-image learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.704693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.005895Z digest=sha256:10f6b6ea6582a6f0d982bc8f78201d4ce24d25c626ffdb57ee65802a95c72374

Observation 5272cc8f-b55b-45cd-92b1-ba3b62e3fe37 · outbound

This paper cites CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:30.079121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:30.079121Z digest=sha256:ea085537ac6a295fcb312ea299141349a3adf386a6ff4b1ce1b5ff1bb8c92b97

Observation 5c3e7584-fa90-43a4-8185-0a4fe387b6ed · outbound

This paper cites Open-Vocabulary Universal Image Segmentation with MaskCLIP.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Open-Vocabulary Universal Image Segmentation with MaskCLIP

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:30.129614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:30.129614Z digest=sha256:bc9251903fafc587fb5bfecef89b553e584bca1d29c1a73b2aa06782583f1b9a

Observation d277024b-19b1-4fc2-b1f8-077edea43f61 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:30.187878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:30.187878Z digest=sha256:c9b2df64225b233115e993357a20622795a06df7bb069e37d20276264020db49

Observation 78f240fe-ec8b-4796-aadf-3b4b65160abe · outbound

This paper cites The pascal visual object classes (voc) challenge.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation The pascal visual object classes (voc) challenge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.669246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.241754Z digest=sha256:0ee87a032d7b332ff678ffe361f5a83f29448122f757eafc2a7abdc976b3c840

Observation 0881ea7d-4ff2-4390-8134-9d8081342a30 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.636723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.253622Z digest=sha256:ec4c43a39c4d861f9945703ee45394a1acd9cfe6c46ac0997da90b419820e61d

Observation 2274c32d-a5ad-4847-8272-79a291cfca52 · outbound

This paper cites Diffusion models for zero-shot open-vocabulary segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Diffusion models for zero-shot open-vocabulary segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.594676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.262474Z digest=sha256:8026a7f9ffae8262121853b02a033065c61de2b7b6c70c8c7b649ad66e29226a

Observation 6b827282-485e-49ef-88eb-515e84b63026 · outbound

This paper cites Segment anything.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Segment anything

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.561979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.323897Z digest=sha256:9aa24f379ce6b75e9338c34a6dc1c8391faf3612af27b3db8698a416251374a6

Observation 423ec143-ef68-46a0-8dd1-e6d917960360 · outbound

This paper cites ProxyCLIP: Proxy attention improves clip for open-vocabulary segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation ProxyCLIP: Proxy attention improves clip for open-vocabulary segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.527327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.434077Z digest=sha256:d6034b4af9499b25b97834601d0f39d8d42739c15de360293a4088af98bbbc29

Observation 6dc9b584-b0b6-4919-87f8-cb622f46ef15 · outbound

This paper cites ClearCLIP: Decom- posing clip representations for dense vision-language infer- ence.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation ClearCLIP: Decom- posing clip representations for dense vision-language infer- ence

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.469805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.511695Z digest=sha256:621b6b11e48aae20d2f2756e05a283264e6995a6fdaee904e28de3eeef898730

Observation 0e733d73-93f9-4fc4-b743-8420134f2b27 · outbound

This paper cites Anti- adversarially manipulated attributions for weakly and semi- supervised semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Anti- adversarially manipulated attributions for weakly and semi- supervised semantic segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.352659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.524335Z digest=sha256:0532f1b736cab2a6d8055c934292717eda8aa230eb784e4d58d97e7ca31704e0

Observation 9fa5cd2c-22c7-4835-9522-776468f09def · outbound

This paper cites A Closer Look at the Explainability of Contrastive Language-Image Pre-training.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation A Closer Look at the Explainability of Contrastive Language-Image Pre-training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:30.578203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:30.578203Z digest=sha256:bdfc13bb68dbfcd06282c25703f94375e771162744bcffe231b0f27cfb3776e2

Observation 29a748cd-ff89-4e14-8ada-6c5fa9bea170 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Open-vocabulary semantic segmentation with mask-adapted clip

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.308320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.672403Z digest=sha256:427ac14ee31bc40a6a275cee61cca8c1f213e0133fa97428c173f135a4e5fafc

Observation cafa38ba-9d28-4b8a-a449-f375d9a8c235 · outbound

This paper cites CLIP is also an efficient segmenter: A text-driven approach for weakly su- pervised semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation CLIP is also an efficient segmenter: A text-driven approach for weakly su- pervised semantic segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.274869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.744405Z digest=sha256:a1c5de8b36ce65fe38225902eaf57eceb6cdb2b9c148c89eb9bfc03546937605

Observation 449e98ac-1963-48a5-babc-1bd3c2baeffb · outbound

This paper cites TagCLIP: A local-to-global framework to enhance open-vocabulary multi-label classification of clip without training.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation TagCLIP: A local-to-global framework to enhance open-vocabulary multi-label classification of clip without training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.237480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.756293Z digest=sha256:b2b31a448cf054181e639bfcc9a31e87f0b539d4039f980d230ac3960ec38838

Observation 9744087a-2167-4d00-b9a8-cbc741bbe050 · outbound

This paper cites Fully convolutional networks for semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Fully convolutional networks for semantic segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.201120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.764526Z digest=sha256:d1c24c72d1c2d817cc478ae635d2ed4e5cdf4a8ebb07a7e38760dad9080352cc

Observation 66449b3c-c91c-4a61-8ca8-18a167b6a519 · outbound

This paper cites SegCLIP: Patch aggregation with learnable centers for open-vocabulary semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation SegCLIP: Patch aggregation with learnable centers for open-vocabulary semantic segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.169300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.854185Z digest=sha256:b384151c42118a833c8343f517a3534665e170f8ae51f74dc92203916e9215ec

Observation 4cbf94d9-3b2a-4eb4-9e34-5ee0d946aa24 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation The role of context for object detection and semantic segmentation in the wild

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:34.127614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.956732Z digest=sha256:3c342c6c83fe261f0df43238ffe77013961299c1d83faafcc48689a54dc72c8c

Observation 82a288b3-01c7-41c6-9ee1-579ecf5ed70b · outbound

This paper cites Learning transferable visual models from natural language supervision.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Learning transferable visual models from natural language supervision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.983972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:30.980568Z digest=sha256:3439f9d249c25163abe126e00449da45722b9e11259fb8ad5859098c9679a901

Observation d3175b13-b7f6-4075-a543-9417b6f3b717 · outbound

This paper cites Per- ceptual grouping in contrastive vision-language models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Per- ceptual grouping in contrastive vision-language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.861006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.004171Z digest=sha256:5e4403eb3d278917d1c7d95751ace89bc2ab488a5d19512dc625273dd0961717

Observation e7ea2c76-8fe9-4204-97ac-b02c22479857 · outbound

This paper cites ViewCo: Discovering text-supervised segmentation masks via multi-view semantic consistency.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation ViewCo: Discovering text-supervised segmentation masks via multi-view semantic consistency

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.818391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.016261Z digest=sha256:d5eed8229c874b67d4746649e76a2546e32c72251afbcfdd33f8071c2b9e81e3

Observation dc4bb987-e0bd-48da-8858-9c13f054bca6 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation High-resolution image syn- thesis with latent diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.785884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.091640Z digest=sha256:3d22ca5c6d4b36055934a43a893a94c11ad8a2560b0254f68b28d3c2ef0b9491

Observation 8db690b3-cad4-42db-9946-f1bf9512de3a · outbound

This paper cites To- ken contrast for weakly-supervised semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation To- ken contrast for weakly-supervised semantic segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.749028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.205206Z digest=sha256:df76c3982a4e4c3f31ddcc63d5a9e3291a7fa286e5be5d2eeb971f5ae298d8ab

Observation ff6fa4f0-1142-43e1-ac19-a7faab042712 · outbound

This paper cites Laion-5B: An open large-scale dataset for training next generation image-text models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Laion-5B: An open large-scale dataset for training next generation image-text models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.713015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.270985Z digest=sha256:a74b7d5e2cfc5c8ea9ede6783d0dad64ceab21c2433dc4efd8c58e6b3f0fd7a1

Observation 3ef15bd6-675a-46e3-ae5b-d66e8b41b6db · outbound

This paper cites ReCo: Re- trieve and co-segment for zero-shot transfer.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation ReCo: Re- trieve and co-segment for zero-shot transfer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.672270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.288916Z digest=sha256:57134e62e6c2cbcd090c2a7bbc9fc88a1890cf7523611f07d82c0aaab5f5925f

Observation 76af3956-dcb7-43e3-82b4-dade04414c75 · outbound

This paper cites iSeg: An Iterative Refinement-based Framework for Training-free Segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation iSeg: An Iterative Refinement-based Framework for Training-free Segmentation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:53:32.128129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.299753Z digest=sha256:078b2522517ddc562ad51526a45eb8c95a37b4b9d0315f812f4d8511387bdf8f

Observation ee87535e-4d4b-41ad-8387-f8c20bdb7c73 · outbound

This paper cites CLIP as RNN: Segment countless visual concepts with- out training endeavor.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation CLIP as RNN: Segment countless visual concepts with- out training endeavor

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.634456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.324636Z digest=sha256:484616661bdc1ece8000e6d1203b2206eb2064d6ee95bfd46eccc1ebedbfd458

Observation 0e62b1ef-f7ee-410a-9bee-51decc852a1f · outbound

This paper cites Sclip: Rethinking self-attention for dense vision-language inference.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Sclip: Rethinking self-attention for dense vision-language inference

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.574412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.421744Z digest=sha256:b2ba1a27b866f2fa36a34cb38cd3b129cb46bdfb8bc1e16c8943a9b879f89079

Observation 537c9fe7-9180-47fc-beb5-44f643700f30 · outbound

This paper cites Sam-clip: Merging vision foundation models towards semantic and spatial understanding.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Sam-clip: Merging vision foundation models towards semantic and spatial understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.398863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.466108Z digest=sha256:f05d95891fbed5de85f67a5e076835afaa7543d6ce5fa55cb69a23b13159bc0f

Observation f1e187ce-2dfd-4150-8125-3768a669df13 · outbound

This paper cites Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:31.546345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:31.546345Z digest=sha256:fb763bd8bb96e87998c835f4bc5d60b70e6553965d218a78aef5154ce2c270ca

Observation e021f2c3-832d-4647-886f-e9733f0fac57 · outbound

This paper cites Clip-dinoiser: Teaching clip a few dino tricks for open- vocabulary semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Clip-dinoiser: Teaching clip a few dino tricks for open- vocabulary semantic segmentation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.361545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.678102Z digest=sha256:447b3a617d9d0005740bab565d6c71714c4dd514a722117d46717f6ce637eebf

Observation 5268430a-a04e-4686-bde8-0acc42f79d99 · outbound

This paper cites From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:53:32.041705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.718299Z digest=sha256:3320e601209925bd66a3a3fdf435df5989288510aafaa655e31e667dbb32a467

Observation 0fd812d0-79ab-49e5-8126-9f64443fd613 · outbound

This paper cites Sed: A simple encoder-decoder for open- vocabulary semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Sed: A simple encoder-decoder for open- vocabulary semantic segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.203098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.738935Z digest=sha256:4753b951510ad4b47c7f519991ea5bf1c6a7893a7b17c65ff5e573de2cbab36e

Observation 6818b0d7-7fbb-4083-b73b-4807879449f2 · outbound

This paper cites Alvarez, and Ping Luo.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Alvarez, and Ping Luo

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:33.169597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.752197Z digest=sha256:8e2b03b30af37944b63b9f1442d85abc040ce5e29ba058aecf7016572ea20d64

Observation 82e486ea-48b1-475c-bb02-2c9994ba1305 · outbound

This paper cites Clims: Cross language image matching for weakly supervised se- mantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Clims: Cross language image matching for weakly supervised se- mantic segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.990468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.765289Z digest=sha256:3712abc53ca7cc057856ffea4bcf150bf5f64c75985ecc742a26e85e191edeb0

Observation 179b678d-4304-41da-b23a-b38123aa2586 · outbound

This paper cites Rewrite caption semantics: Bridging seman- tic gaps for language-supervised semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Rewrite caption semantics: Bridging seman- tic gaps for language-supervised semantic segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.915948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.782187Z digest=sha256:d5dc419b848bf41421d3cc6accf26f38d5ca3ed01570517c1b9705db2e3bb0b7

Observation aa91ddfa-5545-4106-83cb-1dbe4aaf9647 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Groupvit: Semantic segmentation emerges from text supervision

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.850323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.798853Z digest=sha256:d67baec0f95eb99472960dc8eff010132725229decc4b6e356ae1ffb5b2a0724

Observation 34810fe9-a51d-4920-83db-c970a2641733 · outbound

This paper cites Learning open-vocabulary seman- tic segmentation models from natural language supervision.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Learning open-vocabulary seman- tic segmentation models from natural language supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.809151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.808098Z digest=sha256:5c7f0b308323e7bbb22ba0fa74acd66fb4e2fe2c6c3f49b035b0147e53c9f8eb

Observation 2f00c189-f6a3-4241-a794-bda1d061773d · outbound

This paper cites Open-vocabulary panoptic segmentation with text-to-image diffusion models.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Open-vocabulary panoptic segmentation with text-to-image diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.773430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.817712Z digest=sha256:6a58ec6541b373701399d9b4bcb098b7339dee712b2bb42ae499b5a568a396d7

Observation 16223171-313d-4031-a611-a09d431217d6 · outbound

This paper cites Multi-class token transformer for weakly supervised se- mantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Multi-class token transformer for weakly supervised se- mantic segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.726561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.824494Z digest=sha256:0cdbb2c75300f941c3a252eb5cbd878f897725b13ea962f2488462aca30be24b

Observation 0732b187-27e6-4bd9-9ecd-79a5528df144 · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Side adapter network for open-vocabulary semantic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.660574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.838226Z digest=sha256:d5b9318d2681fe31f0799b195dc829447a8a1ea3f5a7f2cb8f76677acaff0f64

Observation dedbb98a-65d2-424e-b811-32470ae6b714 · outbound

This paper cites Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:31.853023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:31.853023Z digest=sha256:2b0d5b2b5c689310d927582c09ca7d2c1019fa292e48042ffff5a130436b69ec

Observation 91ee6cf4-fc84-47f8-bc5f-272acd54846b · outbound

This paper cites Open vocabulary scene parsing.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Open vocabulary scene parsing

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.616043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.865201Z digest=sha256:fcc60cce963bd4dbd91ea6d1b55859b44e6c7e5fcc003b34240445c827ee7641

Observation c5655f20-8b1e-4db0-8adb-12423fbf6252 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Semantic under- standing of scenes through the ade20k dataset

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:31.878658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:31.878658Z digest=sha256:525afd5d6c995e14947cf71fec0c1a23f6299fc4cb08cd1d34651a05769ee24e

Observation 3241dc5e-106a-4fec-82df-4def85422460 · outbound

This paper cites Extract free dense labels from clip.

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation Extract free dense labels from clip

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:53:32.461065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:53:31.889301Z digest=sha256:1c4ade0bee2bb0b884feed17bee2249cc05817218ed63526cd8831e952efe7f9

Pith citing papers

Observation a7d5046b-d83a-4772-9b7f-346ca2037fe4 · inbound

CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation cites this paper.

CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T20:04:44.106636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:04:44.106636Z digest=sha256:2681dea9191ac7a853a438f4d5aa71c68a997bde7121ebb35026ecacd33dbd7d

Observation 1fee426a-3576-4ef6-b5f8-473170d116dc · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:15.798787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:d95ad605cdb5b6891d1c4cf202ffa8b162a917c9b73a0fae41cd83837642d5cc