Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T04:34:24.068813Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.13661.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T04:34:24.068813Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8bd193dc-b30d-4078-8177-d7083e37a48e · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Learning transferable visual models from natural language supervision
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63d26381-afdc-4909-bda2-548169a26959 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddb9cc4e-543c-423d-991f-478e2ee00029 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Domain generalization by mutual-information regularization with pre-trained models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c0642e0-1d44-4817-adae-3bde31c03eab · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neurocomputing, 508:293–304, 2022
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32ba5634-332f-4950-a23b-8ee151898524 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment A simple baseline for zeroshot semantic segmentation with pre-trained vision-language model.ECCV, 2022
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6de4a91c-378c-4c3d-9f61-982631c05f86 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Open-vocabulary semantic segmentation with mask-adapted clip
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11bf04f8-b573-4897-8467-4c55297b54c3 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Open-vocabulary object detection via vision and language knowledge distillation.ICLR, 2022
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b1ce2ea-0b29-4882-b758-c8d3ec22c2e1 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Learning to prompt for open-vocabulary object detection with vision-language model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 565de053-86d3-4a2e-8d37-626adad563dc · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment F-vlm: Open-vocabulary object detection upon frozen vision and language models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8dcd1f4-4855-4adf-963d-405c14e8163e · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6a36991-b888-41d2-8104-2e835b78156e · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip.NeurIPS, 36, 2024
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb9e5c88-b18c-48dd-b5fb-419445fe7bb6 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d417ee-a6ef-4c79-9ebc-15a5b96c56dd · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Fg-clip: Fine-grained visual and textual alignment
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4025cb0e-223d-4a04-9a14-b3b58f99fea0 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Improving fine-grained understanding in image-text pre-training
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77529e6f-7493-4b61-9f75-161c2dccf433 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Position-guided text prompt for vision-language pre-training
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a4a53db-8026-4009-8207-31fb8c29c6fe · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Regionclip: Region-based language-image pretraining
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f74a3a0a-d803-456f-874a-7f819ad08ca3 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Densevlm: A retrieval and decoupled alignment framework for open-vocabulary dense prediction
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fc4abb7-f457-4d88-90c9-7a70dce8e874 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Faster r-cnn: Towards real-time object detection with region proposal networks.NeurIPS, 28, 2015
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32a894cb-2843-4da8-a076-3556dcf10c95 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Fineclip: Self-distilled region-based clip for better fine-grained understanding.Advances in Neural Information Processing Systems, 37:27896–27918, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1be3bf9-5335-4ab8-b12e-af7abd41c793 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f86e41b0-c080-4a67-b931-7a6d6092232a · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fea05e2-1c49-4392-903e-2e2ec146ea25 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Atas: Any-to-any self-distillation for enhanced open-vocabulary dense prediction
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2969ff4-6d48-4357-b08b-10d160aea840 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Clim: Contrastive language-image mosaic for region representation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ec8a351-5b2b-41af-a3d7-29ee243695a9 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Learning to prompt for vision-language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f34aa329-6d56-427e-be60-15064d712336 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0818cb07-8563-4be6-83b6-66b9ae21c98c · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea92005a-c474-4d63-8c91-8cf61cc49883 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Region-aware pretraining for open-vocabulary object detection with vision transformers
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b348c575-dc87-487b-b4c1-837e80beb3cb · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Open-vocabulary object detection upon frozen vision and language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee5b9da2-0c64-4109-bc6c-6afea2ff4193 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Simple open-vocabulary object detection
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 461b57f4-e6e2-4136-8ece-e49634b213ad · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Scaling open-vocabulary image segmentation with image-level labels
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 597862d6-c43b-4b4e-8c4d-21297c1af748 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Cascade-clip: Cascaded vision-language embeddings alignment for zero-shot semantic segmentation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dddf0e2-4cd3-4e4c-8838-73e619e389cc · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Visualizing and understanding convolutional networks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe6d806c-2905-435d-8059-6219fb50845b · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Grad-cam: Visual explanations from deep networks via gradient-based localization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30586115-75d7-4000-bae3-65275a1d9758 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment why should i trust you?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 515bce79-e3d9-414a-8667-04882cf09d0a · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment RISE: Randomized Input Sampling for Explanation of Black-box Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a1c5731-fc8d-4aec-9dff-2fcb024282b4 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Odam: Gradient-based instance-specific visual explanations for object detection
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 600f2382-940b-4e1c-99e6-583c91c474da · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Attcat: Explaining transformers via attentive class activation tokens.NeurIPS, 2022
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fcd5a22-9bb6-4959-906a-6f43fa45edc7 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Vit-cx: Causal explanation of vision transformers
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426aea57-7706-4963-a6ad-75770ba38d26 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment X-pruner: explainable pruning for vision transformers
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 444c21dc-5c62-448f-9e75-d2e059dc324b · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Gradient-based visual explanation for transformer-based clip
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9b6e25d-0459-4d2d-a378-0d1fce8d4c8d · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Extract free dense labels from clip
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94584bb7-1049-4471-ae94-be1cd70f3a1b · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment O’Reilly Media, Inc., 2009
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c64763-8baf-433a-992d-4803d6f90527 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff33e22f-c38a-4388-a02e-9cb2916b1545 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Microsoft coco: Common objects in context
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cda5196-ea45-4be9-b5f0-450626689087 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Scene parsing through ade20k dataset
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ec207f7-1166-4fec-b74b-41a90a1cb363 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60a84edb-86d0-41fd-ba53-cab7dbf6fd26 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Lvis: A dataset for large vocabulary instance segmentation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aedc7db-11ac-459e-878e-0cb683dd086c · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Open-vocabulary object detection using captions
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37a13a99-87f3-4a96-8828-89617d4ca713 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Detecting twenty-thousand classes using image-level supervision
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70f1f7f9-eba6-42e5-9a90-bbe7c2f29837 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Learning Object-Language Alignments for Open-Vocabulary Object Detection
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4015e22e-3554-412d-a9ef-d73385683845 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Everingham, L
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5291320a-5457-4b97-80fa-605b2501cc42 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment The role of context for object detection and semantic segmentation in the wild
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2d19235-5285-41e9-9ae4-545f4796b220 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Cat-seg: Cost aggregation for open-vocabulary semantic segmentation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deaab706-55ba-4d15-928f-54ad94fe9b78 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Coco-stuff: Thing and stuff classes in context
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbedda91-2484-4bca-b016-c1f364cb5091 · outbound
Fine-grained CLIP fine-tuning with self-annotated region alignment Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.