Pith. sign in

Paper Citation Record · LEDGER

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements

As of 18 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 2 inbound Pith citation observations for arXiv:2411.12044.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12044 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:02:15.217453Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T07:43:28.627414Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T07:44:02.855304Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy41
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 211d8a3e-c8e6-4690-8246-333147dd6420 · outbound

This paper cites Cdul: Clip-driven unsupervised learning for multi-label image classification.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Cdul: Clip-driven unsupervised learning for multi-label image classification

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.266548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.881814Z digest=sha256:265b6a30413bb8f6c4812e86fe4660892ec2ef57a3fbda1036b5b80834a814de

Observation b6ea9513-2029-4441-8569-f4895551b810 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:14.887139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:14.887139Z digest=sha256:aefad909199435e0df2a79c89a483f7041411aefd63b4412c880957ba474317a

Observation 72f9f8ae-5d9f-4dc8-8eb9-a694084cd171 · outbound

This paper cites Self- supervised multimodal versatile networks.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Self- supervised multimodal versatile networks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.240261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.891219Z digest=sha256:174936edc65327a53368c562d8856017f32833b43eb242508fd97b6a146afc3d

Observation da910b16-5db1-4bf5-a5b7-051b5d23c11a · outbound

This paper cites Single-stage semantic segmentation from image labels.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Single-stage semantic segmentation from image labels

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:14.896060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:14.896060Z digest=sha256:ce3ced8bdf84ca0387a6548bd2e706a034a3e3916f43be1f5a61517e1d43653c

Observation 481c2f86-fb96-4dae-b2d3-1901bb82c67d · outbound

This paper cites Fossil: Free open-vocabulary semantic seg- mentation through synthetic references retrieval.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Fossil: Free open-vocabulary semantic seg- mentation through synthetic references retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.210189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.900159Z digest=sha256:a4c76c0f366a84cfaef22bdbb4f43f1072378e3f2530a895ce9fdb077b2367d3

Observation dff6d89b-1641-4696-843d-b213366b72ab · outbound

This paper cites Grounding everything: Emerging localiza- tion properties in vision-language transformers.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Grounding everything: Emerging localiza- tion properties in vision-language transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.195655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.904567Z digest=sha256:0d465e9121d4f690da365b4f049a89374851a6b0e58fb384beb6c37c60d2447b

Observation 6fb569d3-48eb-4022-a872-9aeb30cb984f · outbound

This paper cites Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.180194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.908491Z digest=sha256:74354ddc36935d67f31c702fb713e1c40c9bba2c39af5762ec46b364991cdb7c

Observation 4ab646cf-a78e-4811-8387-bf3a92a7999e · outbound

This paper cites Coco- stuff: Thing and stuff classes in context.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Coco- stuff: Thing and stuff classes in context

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.166805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.912479Z digest=sha256:eab19123b5979e907d6d59e891916a683041ec8e6d91a131272c64eec9904242

Observation 3e997ae8-6fdf-449d-81fa-949b1d6e4a11 · outbound

This paper cites Cambridge dictionary.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Cambridge dictionary

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.153216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.916571Z digest=sha256:d9bad7d2ac351f830b57ff0fe33335a0e2b16e3d44dabb9ad3929df42d710cd2

Observation 6f44682f-14fd-47f9-b54a-a974ef599626 · outbound

This paper cites Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.140467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.920494Z digest=sha256:ec10aa32f528b738b135bee24ec89f3d8c024ee81c11d38d840550ee90d1669a

Observation 8c13f81f-b621-4f41-8757-12e61577acf3 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Reproducible scal- ing laws for contrastive language-image learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:14.924613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:14.924613Z digest=sha256:a81d6d37fffed66be19370fca026836ddf95dbb25e27c90f56ad937dadc2f989

Observation c0887d83-5411-474e-9f86-d183190c6a44 · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements The cityscapes dataset for semantic urban scene understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.119894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.928945Z digest=sha256:080b6ece18f2f41a09037762fe985b0bef7e73d97b4f9d98f0e85998f6d77000

Observation 0f1cf99a-31f4-4c16-8f8c-dc718bbd7c6d · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:14.934430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:14.934430Z digest=sha256:03744733a96939caedff17c8fdc58e92a0fced8778022ccb9058a6ad9e43c57d

Observation c21f8ba0-67ec-4ace-a42f-393c5e3a02af · outbound

This paper cites Segment Anything Model (SAM) for Digital Pathology: Assess Zero-shot Segmentation on Whole Slide Imaging.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Segment Anything Model (SAM) for Digital Pathology: Assess Zero-shot Segmentation on Whole Slide Imaging

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:14.939125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:14.939125Z digest=sha256:873d29c0b9563a563a4efa6ee21b34825350fd2b8a7454e169fb5095975ade95

Observation 0d0a5f28-718a-46ba-b984-183d0417ccee · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:14.944316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:14.944316Z digest=sha256:3ad19821a6d462ac80fb0b5551da297d1b9ab3189e3a72e3c2a56cbc3b846d88

Observation 3409ab2c-b2b2-478a-9d24-8000c59fcfa1 · outbound

This paper cites De- coupling zero-shot semantic segmentation.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements De- coupling zero-shot semantic segmentation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:14.950514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:14.950514Z digest=sha256:3749d1d0bfc891b86caeaeeda26b6a02e6df3904463ae90e49f6c09092abd66c

Observation 29ad03de-4071-454d-8fe6-17bcb00fcf8c · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.091190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.954949Z digest=sha256:e1b58b44ee01c9b829e0928c5193e1d8339fba3dff72703c2436be78b8919d6a

Observation b466df33-5e4c-481e-923b-266465469f36 · outbound

This paper cites Learning to prompt for open-vocabulary ob- ject detection with vision-language model.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Learning to prompt for open-vocabulary ob- ject detection with vision-language model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.078230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.960267Z digest=sha256:bed4b15a94c7df81070f006bc856d8c4909a1a657fa21ca9e935e46aed6e48b9

Observation 47293814-2e61-4607-9305-b2430122f3db · outbound

This paper cites The Llama 3 Herd of Models.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:14.965778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:14.965778Z digest=sha256:df02f8df19665eb3ae0b63e8b1f17155f9651b713d646b539bd3911e9c0683f6

Observation 5c5c00ca-f4a6-45cf-96db-e7116c81cc76 · outbound

This paper cites The pascal visual object classes (voc) challenge.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements The pascal visual object classes (voc) challenge

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.064978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.970745Z digest=sha256:c550fce74374d6cedbb81b654ae94c4d7c14af33ffd27a6844a407140721cfd0

Observation 0c34932b-d34d-43fd-922d-914cb46083a6 · outbound

This paper cites Improving clip training with language rewrites.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Improving clip training with language rewrites

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.052158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.975042Z digest=sha256:0a89e888227ad269876db6c284b17db96547a929741dc92ad98363c1888dc195

Observation b510a98d-7f1c-42b4-aa19-6017674f0854 · outbound

This paper cites In- terpreting clip’s image representation via text-based decom- position.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements In- terpreting clip’s image representation via text-based decom- position

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.039663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.980372Z digest=sha256:e939343025fe2986df94348c59c710fe53463f93290cf8e47bd47735343798d1

Observation 5dd071c2-6165-4e5d-ba25-5eca120f0394 · outbound

This paper cites Scal- ing open-vocabulary image segmentation with image-level labels.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Scal- ing open-vocabulary image segmentation with image-level labels

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.025991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.984804Z digest=sha256:871a641e8132da484ae25360d063cd063c943f37cc351b25ddeec6c7acf7d839

Observation 9511a246-35da-4c60-9430-440bd34b54b8 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Open-vocabulary object detection via vision and language knowledge distillation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:16.013414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:14.992137Z digest=sha256:d2f996758541aca0b588359f92aba0ee50f21ca40dff463a9dffe3e4c87a290f

Observation 1582c321-d54b-4cc8-bee0-18cbb66c0b0b · outbound

This paper cites Pay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic Segmentation.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Pay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:14.998915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:14.998915Z digest=sha256:55fbdc84b0e456911d07d5acadda2b991087e6184fe27a32f7154fc08b829c88

Observation dd8ae3d4-9d8f-4ce7-b917-34f80779b8b5 · outbound

This paper cites Open-vocabulary semantic segmentation with decou- pled one-pass network.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Open-vocabulary semantic segmentation with decou- pled one-pass network

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.003687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.003687Z digest=sha256:2ce3fdfd66f675940bbe1879c5bd442bb05be70134172b28940e23273aa398e4

Observation 56431c2d-b044-4d6f-809e-53cdbf89ed86 · outbound

This paper cites Deep residual learning for image recognition.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Deep residual learning for image recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.007900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.007900Z digest=sha256:8f56eca83f5e662e95c99f628ae43565207aa4c45b42be1dc64c7fc49248786c

Observation 5e21e778-6fda-49d8-8017-7709597aa0b9 · outbound

This paper cites Computer-Vision Benchmark Segment-Anything Model (SAM) in Medical Images: Accuracy in 12 Datasets.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Computer-Vision Benchmark Segment-Anything Model (SAM) in Medical Images: Accuracy in 12 Datasets

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.013156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.013156Z digest=sha256:ee424a2d236b61bc02bf2c9f81674634a6c1c34a3c9101147983f98c2b37da38

Observation b13db7ec-998b-47b3-abe3-85d987c34028 · outbound

This paper cites Open- clip, July 2021.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Open- clip, July 2021

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.985978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.019669Z digest=sha256:1ca4c6e996a6e2f1ae4c045172290311188aedc01f5ad1e9ba29f41cb11af334

Observation 02084594-29a3-4982-adc4-6e61b7efee96 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Scaling up visual and vision-language representation learning with noisy text supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.024874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.024874Z digest=sha256:3dfef2a5d5b34402ff68da4fc69eb6683c301e24953fe3d0a9cbaaa5f0791f0e

Observation 9e94885b-94b4-46f2-86df-bc7dcf5525de · outbound

This paper cites Learning mask-aware clip representations for zero-shot segmentation.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Learning mask-aware clip representations for zero-shot segmentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.029460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.029460Z digest=sha256:390d4ab053306b3ee1b73682cf18281358793ab0ecb8140334c109c06548af5b

Observation f219d208-1b85-40fc-9e8c-4d9c017d0510 · outbound

This paper cites Segment any- thing.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Segment any- thing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.034305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.034305Z digest=sha256:593b0d45ecd698b5cde754299477bf85f14ad79f52140621671c1b7ffbcd3754

Observation 56afe800-ac58-4364-94b7-1eaecdf3a1e8 · outbound

This paper cites ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.038113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.038113Z digest=sha256:e0b9f5b08c121353f8e86f79d67604ec28a115354d1e217c2bea4d1d826a03ca

Observation 8bf95457-bb15-4e24-9c72-1bb2e9de176d · outbound

This paper cites Language-driven semantic seg- mentation.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Language-driven semantic seg- mentation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.042655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.042655Z digest=sha256:0443587088f86c842365289ee13e285e1e97513302d3e1226673ff5970418664

Observation 01364582-e267-4d1b-8113-0e67a3c6847f · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Align before fuse: Vision and language representation learn- ing with momentum distillation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.046526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.046526Z digest=sha256:cdf9d8f50e48aa6fa9898be8d215de32163f8cf2f8289857ac4444593a2816c6

Observation 0b122750-5ae1-4ae5-a390-22d928c5f2cd · outbound

This paper cites Cascade-clip: Cascaded vision-language embeddings alignment for zero-shot semantic segmentation.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Cascade-clip: Cascaded vision-language embeddings alignment for zero-shot semantic segmentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.933345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.051512Z digest=sha256:bae7197d561ac8137a5c26ba4eb665310641a7bb76189a5365df664fd672f564

Observation 50d91a68-bd76-4a2a-b76a-0a60eca8e541 · outbound

This paper cites A Closer Look at the Explainability of Contrastive Language-Image Pre-training.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements A Closer Look at the Explainability of Contrastive Language-Image Pre-training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.055788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.055788Z digest=sha256:883caaaaec27ea1e6ff4dd2bcfb97317153ded65ac9fbbee12b6f29ad6cc2c79

Observation 46bf1e19-e2f0-4b83-8ed1-b30ae0c03551 · outbound

This paper cites Open-vocabulary object segmentation with diffusion models.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Open-vocabulary object segmentation with diffusion models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.060339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.060339Z digest=sha256:345792b4abc51935dcee54865009e10bc27dbe72f25c06116b05d8622e81f394

Observation f59aa609-d099-4e85-ae70-8469dfecd596 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Open-vocabulary semantic segmentation with mask-adapted clip

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.065044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.065044Z digest=sha256:dad073afde71a8d86dfdac2f040fd123458fd49a5ca99c45a6ab1de72360a413

Observation 75bcf5e0-0dff-4ecf-976b-fb57d54dcc44 · outbound

This paper cites Microsoft coco: Common objects in context.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Microsoft coco: Common objects in context

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.896916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.070556Z digest=sha256:bb196f60ab29d3466f882d13690566515ff577afe0aa1474f45e77e9b5ed962a

Observation 190940b6-a7b6-4443-9446-4fbbdb569e31 · outbound

This paper cites Clip is also an ef- ficient segmenter: A text-driven approach for weakly super- vised semantic segmentation.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Clip is also an ef- ficient segmenter: A text-driven approach for weakly super- vised semantic segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.882322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.075174Z digest=sha256:0b3c02bd16c70d6d1c31c9c4fdb8fbbb48db9089c1fe61e841b63030f6ff5183

Observation b42af303-fa76-4e58-84bd-564ea7e94149 · outbound

This paper cites Tagclip: A local-to-global framework to enhance open-vocabulary multi-label classification of clip without training.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Tagclip: A local-to-global framework to enhance open-vocabulary multi-label classification of clip without training

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.864645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.080381Z digest=sha256:ae80b1f922d3bdd44fdf0019fbbf7e13faa0da74052ac8968d88ceb837a01255

Observation 2d462619-d481-4064-ba83-a4e310900b11 · outbound

This paper cites Object- centric learning with slot attention.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Object- centric learning with slot attention

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.848421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.085665Z digest=sha256:75c50cbe251a6a35814df9e98310ae83f0215f82a63fcae916896c71ead4ddf7

Observation b0e5cd08-77fd-4df4-9e12-04d6a2833d01 · outbound

This paper cites Segclip: Patch aggregation with learn- able centers for open-vocabulary semantic segmentation.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Segclip: Patch aggregation with learn- able centers for open-vocabulary semantic segmentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.824852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.090023Z digest=sha256:6b59b852968bfdced58bb5863afb4f41e3cd2821cd6ac7a9b3adbd17414b3a0f

Observation 964ee126-98fa-4a52-8cad-db981c1d328d · outbound

This paper cites Segment anything model for medical image analysis: an experimental study.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Segment anything model for medical image analysis: an experimental study

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.095233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.095233Z digest=sha256:385eaa82d84154d5ae1010e55551651ad147e77f5313442e96eb5ec1a966bc1a

Observation 6fd68484-45b2-43c2-a4ac-1d7f1d998269 · outbound

This paper cites MetaICL: Learning to learn in context.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements MetaICL: Learning to learn in context

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.793596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.099576Z digest=sha256:a2bcef8a78a8aaa669a121043fa2b8d7183200eabc4c207092c9f216e507c783

Observation db2a48f3-b942-4f89-96ad-20c99eda0f74 · outbound

This paper cites The role of context for object detection and se- mantic segmentation in the wild.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements The role of context for object detection and se- mantic segmentation in the wild

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.104212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.104212Z digest=sha256:cd7c3ad025b5d04f73b62f3a6d623f319b4f9b9f2340e48ef5c68cf12e5509c9

Observation 28814d13-7c08-4ade-90a9-27330caac1a0 · outbound

This paper cites Peters, Mark Neumann, Mohit Iyyer, Matt Gard- ner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Peters, Mark Neumann, Mohit Iyyer, Matt Gard- ner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.767280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.108896Z digest=sha256:f0a3fc1a6394442414b92ad5378f235a7caeab35137a214d952084b149d5d4d7

Observation b772ee8e-83cb-4769-aeb0-9dc5fd81e4de · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Learn- ing transferable visual models from natural language super- vision

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.752229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.113156Z digest=sha256:f201cd67076786181459f3954fa361b1a0ca6ff3253fc050effa16335897c26c

Observation 13723cda-a859-4258-b68d-13e5289e6123 · outbound

This paper cites Improving language understanding by gen- erative pre-training.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Improving language understanding by gen- erative pre-training

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.739310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.117918Z digest=sha256:e0313a5d8a1d4f53b5dd86fb3092fe019179da424b42452826270dca38fb5b9e

Observation 698491c3-2181-491b-b6fc-b0275cf89c4e · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.123028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.123028Z digest=sha256:433b6ad983ff0e952b2c26e906a097a5d5bc911c6ae9758361fe8a3546912871

Observation 5808c737-4193-4742-80be-7f5637b200c0 · outbound

This paper cites Per- ceptual grouping in contrastive vision-language models.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Per- ceptual grouping in contrastive vision-language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.715493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.127248Z digest=sha256:271373c589be914eb1c01f771e2d9b16f41da16c308e6d32b630024db9e9b043

Observation a5c0e07a-e3ef-4355-aa60-abe705c1ca44 · outbound

This paper cites Laion-5b: An open large-scale dataset for train- ing next generation image-text models.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Laion-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.697954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.131814Z digest=sha256:95bc23073a9911958ea3906cbe9bf46b5f48b23fd960708b23173fbfa30d0847

Observation 478430dc-a978-417d-b010-d8b83f948698 · outbound

This paper cites Ex- plore the potential of clip for training-free open vocabulary semantic segmentation.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Ex- plore the potential of clip for training-free open vocabulary semantic segmentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.680188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.136424Z digest=sha256:bd285a31f7fcc5e3bdc1c53396e9f91717e5d55cfe407940bf9f61d00db0b0e1

Observation 64bff08c-1939-4e36-82e3-246c1dc24698 · outbound

This paper cites Edadet: Open-vocabulary ob- ject detection using early dense alignment.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Edadet: Open-vocabulary ob- ject detection using early dense alignment

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.664087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.141043Z digest=sha256:fe4f003d5998180e3635966ee1e39cfd23443487ddfa37b61e7ce95f4465f9aa

Observation f1331559-3d22-4fff-8da3-1fab9c5ad7a4 · outbound

This paper cites Reco: Retrieve and co-segment for zero-shot transfer.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Reco: Retrieve and co-segment for zero-shot transfer

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.647733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.145523Z digest=sha256:9bcfcb6d0091ab84e8b924ceb9279ba15dcffaca8e483e5b03425913fa0f6bc2

Observation 275e2ec4-c1d3-415d-9948-5d5671d3821c · outbound

This paper cites Going denser with open-vocabulary part segmentation.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Going denser with open-vocabulary part segmentation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.632789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.149822Z digest=sha256:5c9206cfc6747ad1786d698bf6c84584058a7798f2d5306d05be4b55870c4df9

Observation b5a22c4c-0561-46b1-80bb-9ae040ee9baf · outbound

This paper cites Clip as rnn: Segment countless visual concepts without training endeavor.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Clip as rnn: Segment countless visual concepts without training endeavor

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.614395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.153522Z digest=sha256:67f8b56bf57b29900ab2ab9dbfe7bf652218dd23e9e9774eb9aa048c20d69941

Observation bcc374ff-4d67-4c54-96a5-7b86198e0e83 · outbound

This paper cites Galip: Generative adversarial clips for text-to-image synthe- sis.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Galip: Generative adversarial clips for text-to-image synthe- sis

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.587655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.157867Z digest=sha256:051b2dcb6fca5a1798bd00e5d59d90b3a57a534b10164c34ef56a58e10bcaee1

Observation f8581781-ecde-49f0-966c-776d4b62e2f6 · outbound

This paper cites Attention is all you need.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Attention is all you need

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.162625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.162625Z digest=sha256:c55310b133998f9214b082b6ed3f3b4681d06af1939c85d185f9bedd4bbf86a6

Observation d9606a73-0f19-4eea-9f67-e31315b3625e · outbound

This paper cites SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.166483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.166483Z digest=sha256:22fc723278ec981bde16e928395f6432c185e12505fa2d81e8077657be7280e4

Observation feb0ec46-cbde-4eec-955d-825dc8e0ba31 · outbound

This paper cites CLIP-GEN: Language-Free Training of a Text-to-Image Generator with CLIP.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements CLIP-GEN: Language-Free Training of a Text-to-Image Generator with CLIP

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.171099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.171099Z digest=sha256:d41e3976f9e7021f2281c90ff19be0a7f40d2bd6f99818b0b144904e904c86ff

Observation 02361cb0-5e06-4ee1-ad18-e1e7941c7d63 · outbound

This paper cites Clip-diy: Clip dense infer- ence yields open-vocabulary semantic segmentation for-free.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Clip-diy: Clip dense infer- ence yields open-vocabulary semantic segmentation for-free

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.555703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.179084Z digest=sha256:b41d295e67f08a908624180d512420973375e8d49d4b3fc31053901748934f70

Observation a58c2522-69dc-41e7-a956-422f5e3d408a · outbound

This paper cites Demystifying CLIP Data.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Demystifying CLIP Data

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.183230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.183230Z digest=sha256:5e2c9620ce82e3f506b48de09fdd5cc7730c9b0b96726ec5538eecfb3f304445

Observation 2b0f967f-b902-4c88-9c79-c0cd8f061570 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Groupvit: Semantic segmentation emerges from text supervision

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.540316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.187351Z digest=sha256:a5d6775fe97ab97896fbf255a8e69be283064c38d68a7577734bc07fae200f5f

Observation 308f03c9-f66e-4035-9c47-0035234cf647 · outbound

This paper cites Learning open-vocabulary semantic segmentation models from natural language supervision.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Learning open-vocabulary semantic segmentation models from natural language supervision

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.525991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.191041Z digest=sha256:6e6f2458ae9de7649b8c85d677a77e7db93067ce72d2aec333aeb5df9cc04230

Observation 3f61d6fc-b9ee-49aa-a0e4-e03d8475c2ab · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.194983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.194983Z digest=sha256:25f77dcdb053f9b7ffe35a823168ac9594dc258a942a89becdbf2be6a00bff02

Observation 5a54b700-5e61-4465-ab40-2ce190c15b47 · outbound

This paper cites A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- language model

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.492430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.200520Z digest=sha256:0a890ffdaa4485dd898eb0aac405ec5fecfc0e701310a79114d609b573b2bfe0

Observation 64285afd-c765-4523-b4fd-10a1301617e5 · outbound

This paper cites Learning deep features for discrimi- native localization.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Learning deep features for discrimi- native localization

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.471607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.207065Z digest=sha256:3c669a7e07d1c846210fc73124b3e564c2a9bf79ece971689970128ce4a4cead

Observation f0908021-ee5c-4ca8-b8e1-68d43e7cfc9b · outbound

This paper cites Extract free dense labels from clip.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements Extract free dense labels from clip

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T18:02:15.212841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:02:15.212841Z digest=sha256:7a4aec9bb723cda03e8e594f051d66dcc98d6cd8770265176b5820facf68abd6

Observation 6d1bb6a9-5310-4e4f-95b2-6d6dfa814bcd · outbound

This paper cites thing” classes and one explicit background class, our method fails to distinguish foreground classes from the background when the word “background.

ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements thing” classes and one explicit background class, our method fails to distinguish foreground classes from the background when the word “background

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:02:15.445785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:02:15.217453Z digest=sha256:4c2db8351501154cd73ca99404bcf55efcd3ed1ca1154d65e6085e3ba40c786b

Pith citing papers

Observation 5b4f4557-10d4-48b9-b2a5-b444dc26a802 · inbound

SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation cites this paper.

SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:53:19.909663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T13:51:35.769341Z digest=sha256:d94c8a4493bde6c4169cad2bcc8e80ebfb8258bfe43f2eda53c0a8d23a2d62bd

Observation a17b1e41-9266-4613-a162-a90dc371c1a4 · inbound

SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation cites this paper.

SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:44:02.856656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T07:43:28.627414Z digest=sha256:45d8ea4859b140a79d76ad28f04213f0faed36a410429894338ceb75bdf1baa2