Pith. sign in

Paper Citation Record · LEDGER

Visual Textualization for Image Prompted Object Detection

As of 8 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2506.23785.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23785 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:37:08.897852Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved13
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf10745e-c8e4-4e84-81dc-bf97e68189d9 · outbound

This paper cites Lung image database consor- tium: developing a resource for the medical imaging research community.

Visual Textualization for Image Prompted Object Detection Lung image database consor- tium: developing a resource for the medical imaging research community

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.347738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:02.947288Z digest=sha256:82923364e1e5b8f6906556b45f436a23ac6afa4ccbb84ff734aabcf3317c278f

Observation a8a31ae4-53dd-4930-936c-bc8d7851bbba · outbound

This paper cites Exploring Visual Prompts for Adapting Large-Scale Models.

Visual Textualization for Image Prompted Object Detection Exploring Visual Prompts for Adapting Large-Scale Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.071096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.071096Z digest=sha256:1106a81374ec2c346f1e2d3e761bd9efe9e2ed6debb3cb78b607c0f34da3b838

Observation 494e3f34-7810-4fca-94fa-730bd867c9cc · outbound

This paper cites Fs-detr: Few-shot detection transformer with prompting and without re-training.

Visual Textualization for Image Prompted Object Detection Fs-detr: Few-shot detection transformer with prompting and without re-training

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.332668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.162281Z digest=sha256:09f15f997f7c9a46982b981fef0475c0ab2a15ad51f07ca4745c8ce1d8917acc

Observation c10b563c-4060-489f-b9b8-d1a0a9c4ea95 · outbound

This paper cites Apollo: Unified adapter and prompt learning for vision language models.

Visual Textualization for Image Prompted Object Detection Apollo: Unified adapter and prompt learning for vision language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.316861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.251458Z digest=sha256:4c394bb7322518d302b08f92001fa99b92937f3665014137e366a7bda3758064

Observation f4129d0f-d7d2-41ab-ab60-dc6c1fc394c2 · outbound

This paper cites Coarse-to-fine vision-language pre-training with fusion in the backbone.NeurIPS, 35:32942–32956, 2022.

Visual Textualization for Image Prompted Object Detection Coarse-to-fine vision-language pre-training with fusion in the backbone.NeurIPS, 35:32942–32956, 2022

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.299990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.345641Z digest=sha256:0e929b0806bb2fcdfa1110d979b226e40ed3ed8fdbd8108465670dd668087c50

Observation b47ad70b-726a-4417-8822-148ba1e39462 · outbound

This paper cites s- adaptive decoupled prototype for few-shot object detection.

Visual Textualization for Image Prompted Object Detection s- adaptive decoupled prototype for few-shot object detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.276015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.421737Z digest=sha256:7acfbc8e4a3da47f066569965228e536adeadc57f43ec3fee428ea7fae333029

Observation 02332cf4-9657-440c-b52a-6f5b0fca1064 · outbound

This paper cites Learning to prompt for open-vocabulary object detection with vision-language model.

Visual Textualization for Image Prompted Object Detection Learning to prompt for open-vocabulary object detection with vision-language model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.258469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.502861Z digest=sha256:ccd5287f74ddb89c3b6dd11abf799eaa9a3ef096e72eb9f3ca635ba28da6b38b

Observation fc7a5f35-5371-4a1f-8b5f-125bf38d7737 · outbound

This paper cites The Turking Test: Can Language Models Understand Instructions?.

Visual Textualization for Image Prompted Object Detection The Turking Test: Can Language Models Understand Instructions?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.644571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.644571Z digest=sha256:3e67bf12b226fa9b5e49db5797acd405c46acf4d517a7a77edf11fe5ea0ee218

Observation 5d4286b5-c630-43ea-91e4-2648e1275036 · outbound

This paper cites The pascal visual object classes (voc) challenge.

Visual Textualization for Image Prompted Object Detection The pascal visual object classes (voc) challenge

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.235878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.743469Z digest=sha256:ce076b72d71051b38f3b82ba766d45aea52cf966cb1d25199b81d0a9b4adea62

Observation da991085-3e35-48f9-92af-c7c2e6ba7bef · outbound

This paper cites Few- shot object detection with attention-rpn and multi-relation detector.

Visual Textualization for Image Prompted Object Detection Few- shot object detection with attention-rpn and multi-relation detector

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.829699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.829699Z digest=sha256:4705de6eb4d5c354e2aaabb2868fc96ec147285aec89d1b5973b0f7887e8e2f7

Observation d5dc14cd-cb45-4eec-8789-e2f26e00a6e2 · outbound

This paper cites Nuclei grading of clear cell renal cell carcinoma in histopatho- logical image by composite high-resolution network.

Visual Textualization for Image Prompted Object Detection Nuclei grading of clear cell renal cell carcinoma in histopatho- logical image by composite high-resolution network

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.205060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:03.904842Z digest=sha256:6e69ed8b2768aa6cedbcb19a44dccc66d0b8c9527cca10bf0bae8f87b14cf556

Observation 0c33a867-fc55-40ee-9428-acedea4ca5b9 · outbound

This paper cites Hover-net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images.

Visual Textualization for Image Prompted Object Detection Hover-net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.185000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.004247Z digest=sha256:1e53e356ca0ea8325090e0c7a8fb5e68fb304e27d5a572564ff91beb937a85f9

Observation 482ae711-fa47-4911-ab40-5dd7c88d34a5 · outbound

This paper cites A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models.

Visual Textualization for Image Prompted Object Detection A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.064646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.064646Z digest=sha256:a85866771c69e5bc2503680916d891c3895ca4fdbe41c6b44108dc98f7d56dcf

Observation 966a79fa-1baa-48e3-bde5-e59e1b4c9e28 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Visual Textualization for Image Prompted Object Detection Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.138434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.138434Z digest=sha256:05f1d72dffbfb6119c10d2baed090c7da12bb7c49197b4db87e75537536bddb5

Observation d248f254-2407-4685-a5a7-ce97d22d99b7 · outbound

This paper cites Dp-ddcl: A discriminative prototype with dual decou- pled contrast learning method for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Dp-ddcl: A discriminative prototype with dual decou- pled contrast learning method for few-shot object detection

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.166743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.250057Z digest=sha256:b9768becfad2bed0e09cc5c5c61031c4831fdfe534aeb1ff1fabf64256ed7639

Observation 2c1e0f48-be20-4d20-b198-254041fd36b0 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Visual Textualization for Image Prompted Object Detection Lvis: A dataset for large vocabulary instance segmentation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.149088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.369673Z digest=sha256:6e04879987d666f1dcf378184112a7bb1bbc48b5b2e4a4bc076f6358e680323a

Observation 75988dd5-13ac-45a7-8487-7f4ad5410f54 · outbound

This paper cites Few-shot object detection with foundation models.

Visual Textualization for Image Prompted Object Detection Few-shot object detection with foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.135836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.489927Z digest=sha256:458d97474a1cc09e59e9e90134baf51f0492a08b89a5fb957d4a1bf37a9ce4cc

Observation e848ea3a-8a08-4488-8694-157cb8550f99 · outbound

This paper cites Query adaptive few-shot object detec- tion with heterogeneous graph convolutional networks.

Visual Textualization for Image Prompted Object Detection Query adaptive few-shot object detec- tion with heterogeneous graph convolutional networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.119455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.567751Z digest=sha256:03103b08844db32b4b64a9162e6f126e9fd5ac6f71228710d974f08ec91bacf6

Observation 0eb56578-ed90-408c-b085-e6eee5a18312 · outbound

This paper cites Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting.

Visual Textualization for Image Prompted Object Detection Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.660592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.660592Z digest=sha256:a4305ab441454208ff1c06ea022422e02ee3374d56530461e948e4b7310fa57a

Observation 0b083cbe-5052-4226-b5da-d752dfbf10b3 · outbound

This paper cites Meta faster r-cnn: Towards accurate few-shot object detection with attentive feature alignment.

Visual Textualization for Image Prompted Object Detection Meta faster r-cnn: Towards accurate few-shot object detection with attentive feature alignment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.100727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.727426Z digest=sha256:b7305620af5257827a1a18905e7d9db8b49f12282869669601687770009e07db

Observation 68627033-c1b9-4d89-9ce7-a487d2873efb · outbound

This paper cites Few-shot object detection with fully cross- transformer.

Visual Textualization for Image Prompted Object Detection Few-shot object detection with fully cross- transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.084388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:04.978030Z digest=sha256:b0168d5b305dfef43f08255aa01e0c8063ff2d4b2987072312542803f7f4316d

Observation a3f2ac37-d0a0-4ebb-8c91-099a6c3c8f29 · outbound

This paper cites Few-shot object detection via variational feature aggregation.

Visual Textualization for Image Prompted Object Detection Few-shot object detection via variational feature aggregation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.947853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.120058Z digest=sha256:6ce2dc7dc5f2235382622d498d54cdb3814bf87a41d6e75876993610472be999

Observation 02f4ff15-dfd7-4c74-ba03-bcff5b1ea601 · outbound

This paper cites Visual prompt tuning.

Visual Textualization for Image Prompted Object Detection Visual prompt tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.931188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.202072Z digest=sha256:7c8170c53507d5c898d5649a2af1eced1599df83789a487543e3f7f4da6a32e2

Observation d8e8e50f-2e4b-4c46-af8e-5097595a48d5 · outbound

This paper cites Bert: Pre-training of deep bidirectional transform- ers for language understanding.

Visual Textualization for Image Prompted Object Detection Bert: Pre-training of deep bidirectional transform- ers for language understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.913726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.307636Z digest=sha256:3347548917dc762937d77848f6a015901ec6682d409204e8ce41260e5652a959

Observation 8fb616bd-3570-45dc-8811-9f2bf38cb355 · outbound

This paper cites Maple: Multi- modal prompt learning.

Visual Textualization for Image Prompted Object Detection Maple: Multi- modal prompt learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.891918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.404038Z digest=sha256:627b4b59152e7bc38209aeb75b90f5b3643eae615d73912ca60527ee139a5540

Observation 5b792927-2702-4896-8eec-86b34a410cca · outbound

This paper cites A dataset and a technique for generalized nuclear segmentation for computa- tional pathology.

Visual Textualization for Image Prompted Object Detection A dataset and a technique for generalized nuclear segmentation for computa- tional pathology

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.871760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.530990Z digest=sha256:a88684066092be02c99453462e9e505ee4fef605afaba81214194568f1a4f25d

Observation d737176b-afa7-4a34-9af7-dd51e16057d7 · outbound

This paper cites F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models.

Visual Textualization for Image Prompted Object Detection F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.649074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.649074Z digest=sha256:6cca76b70984ad668947562e2e4ebc34115aef451e2ef38abaea6beab65bed8a

Observation 7be32b2c-81c1-47a8-b5ac-15dbca57af3b · outbound

This paper cites Elevater: A benchmark and toolkit for evaluating language-augmented visual models.

Visual Textualization for Image Prompted Object Detection Elevater: A benchmark and toolkit for evaluating language-augmented visual models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.855850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.790125Z digest=sha256:705f62847d39bb173d03a249a60c14d82d5aae31e5281f6f669d40387bbb1b97

Observation 002258a5-4867-4e66-936b-65487eb0445c · outbound

This paper cites Disentangle and remerge: interventional knowledge distillation for few-shot object detection from a conditional causal perspective.

Visual Textualization for Image Prompted Object Detection Disentangle and remerge: interventional knowledge distillation for few-shot object detection from a conditional causal perspective

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.841297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.875694Z digest=sha256:c2583462b7f13a062a7719ae28439388efd6e43740713e457ba1d03d468b48d7

Observation e4c5e956-6447-4a4a-9cc8-20fcf753f8c8 · outbound

This paper cites Grounded language- image pre-training.

Visual Textualization for Image Prompted Object Detection Grounded language- image pre-training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.820928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:05.969037Z digest=sha256:bc3dffe93bd2626c5877642828812d5055d2338bcf36958177d2ac8a835f9f5c

Observation 9edc8e46-7db7-4c90-9e95-dd493c8605c8 · outbound

This paper cites Microsoft coco: Common objects in context.

Visual Textualization for Image Prompted Object Detection Microsoft coco: Common objects in context

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.788346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.045211Z digest=sha256:461e665540861b59c6f906963881a5a4fb53253b07254d8f56f7915ceb512563

Observation b2801c4e-e958-4df3-9fca-ba4b3f05eab5 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Visual Textualization for Image Prompted Object Detection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.107791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.107791Z digest=sha256:7f5100c6cde6ffb9971095201418ae86b96044269078d5fee58979b4e0207733

Observation 01277d58-5a56-4099-b94e-764c9deffc45 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Visual Textualization for Image Prompted Object Detection Swin transformer: Hierarchical vision transformer using shifted windows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.250780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.250780Z digest=sha256:cc9f3895d35b444422a187283039b21a2a470ab64c4d0f5e4c67a71e7aed320b

Observation 739b0451-22e2-4f15-a6aa-48fd1e79a463 · outbound

This paper cites Breaking immutable: Information-coupled prototype elaboration for few-shot ob- ject detection.

Visual Textualization for Image Prompted Object Detection Breaking immutable: Information-coupled prototype elaboration for few-shot ob- ject detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.756412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.324343Z digest=sha256:628b6336865327a9c6542b09349d3ce99c6f0c437bc2777280a68a164531fb3e

Observation c39a64d9-a64f-48d1-99da-e3a2469eb890 · outbound

This paper cites Image segmentation us- ing text and image prompts.

Visual Textualization for Image Prompted Object Detection Image segmentation us- ing text and image prompts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.737378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.388853Z digest=sha256:4971819857e0b27fd5407c6510bd766936a0984a5872b44344ce6b1e91096997

Observation f0c2cc35-7f82-4ae8-af0e-e8d608ccaf31 · outbound

This paper cites Digeo: Discriminative geometry-aware learning for generalized few-shot object de- tection.

Visual Textualization for Image Prompted Object Detection Digeo: Discriminative geometry-aware learning for generalized few-shot object de- tection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.722011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.440503Z digest=sha256:4360a1ac349de0efbe00f4d1875a1435c24cb4edb52cc6be047e2844929f8aa4

Observation 5f89670f-611c-4991-91f4-bc0e8c45a54a · outbound

This paper cites Simple open-vocabulary object detection.

Visual Textualization for Image Prompted Object Detection Simple open-vocabulary object detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.701594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.497243Z digest=sha256:174540a7e3203a22f3c758dfd0ea9243cc88806159bc17e7ae1d4f0ffb490d79

Observation e0af7dcd-d85c-45f1-b10d-0fdc4ff3770f · outbound

This paper cites Scal- ing open-vocabulary object detection.

Visual Textualization for Image Prompted Object Detection Scal- ing open-vocabulary object detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.678121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.557963Z digest=sha256:d364c49c91043b4634661cfb7a2a7e99a29eac1d6144a89ab1c6a55bd59543b0

Observation a907e535-54c4-4178-8f1b-7ead7169f35e · outbound

This paper cites Defrcn: Decoupled faster r-cnn for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Defrcn: Decoupled faster r-cnn for few-shot object detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.659393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.624223Z digest=sha256:c8ac7e8d0fb1f1398db7b0d67c88cd30c7c45284d676efae7ec8307feae65d2a

Observation 60171d71-81d6-4b39-a81b-16cc0df5ae2a · outbound

This paper cites Language models are unsuper- vised multitask learners.

Visual Textualization for Image Prompted Object Detection Language models are unsuper- vised multitask learners

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.638175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.665207Z digest=sha256:4f885b6d523cae8f6a826adad0b82fa08d2a8495164464d96f17c0fb0bb4b4d3

Observation 3de2fc5b-5b78-44e4-a450-e5c16d79a758 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Visual Textualization for Image Prompted Object Detection Learning transferable visual models from natural language supervi- sion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.611623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.744231Z digest=sha256:26670efac946bcc48612e52c990e0de29161911910b2051b33072d818e9866a5

Observation 6157ef5e-b882-49eb-a9d0-f5d86407f505 · outbound

This paper cites Adaptive multi-task learning for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Adaptive multi-task learning for few-shot object detection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.588564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:06.869279Z digest=sha256:39dafa48d8ec952d0d93ef86b81e816d42e3f8b9a7a348c53113c2b3238f4227

Observation 9042d0fc-ffe6-4c22-afa7-39b41f8cacbf · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Visual Textualization for Image Prompted Object Detection Objects365: A large-scale, high-quality dataset for object detection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.980275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.980275Z digest=sha256:9f7796f9d687306fbd8c001d1d60948992a9524d85712375f0be4760e84b4890

Observation 9b62f9e8-eb53-4952-bd87-25f58b822722 · outbound

This paper cites Few- shot adaptive faster r-cnn.

Visual Textualization for Image Prompted Object Detection Few- shot adaptive faster r-cnn

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.551444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:07.045297Z digest=sha256:09d86c1d89e771ac2a383135b9e7ab32189cad7eef0f5ad11004a8382d396a64

Observation 86f89577-7cb6-453c-810b-285054f4d9b3 · outbound

This paper cites Frustratingly Simple Few-Shot Object Detection.

Visual Textualization for Image Prompted Object Detection Frustratingly Simple Few-Shot Object Detection

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.148413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.148413Z digest=sha256:acee2d69376ac249a5a582d3522005397876561ec900cc9089b2d2278721c485

Observation 3e24ebcb-cddf-4047-90bd-acd227207b64 · outbound

This paper cites Snida: Unlocking few-shot object detection with non- linear semantic decoupling augmentation.

Visual Textualization for Image Prompted Object Detection Snida: Unlocking few-shot object detection with non- linear semantic decoupling augmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.530773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:07.207653Z digest=sha256:dbf73cfb855d7df5877832412fa90848eb1fead9e96140d97c5db8599540ba6f

Observation dfa01c67-5c5f-4333-b46e-b41aa717ed20 · outbound

This paper cites Multi- scale positive sample refinement for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Multi- scale positive sample refinement for few-shot object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.516661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:07.267887Z digest=sha256:10e53db085f220ec075cb5676c7d9a4bd2be5c6bbfdce052558a3f06d29258db

Observation c7e19578-69fe-445c-abf4-20e2d9dc6474 · outbound

This paper cites Multi-faceted distillation of base-novel commonality for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Multi-faceted distillation of base-novel commonality for few-shot object detection

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.501311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:07.331054Z digest=sha256:344eb62a5ea648b74be162e61a5cff44c5aa72432a5c9e1430f8b491c0c0689b

Observation 1767531d-f860-461e-b608-09d2122bdc64 · outbound

This paper cites Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching.

Visual Textualization for Image Prompted Object Detection Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.485161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:07.456827Z digest=sha256:8bcad5a4cb6be756e08d1fe35e09f7062271328c1cd6e10b3f8c5c212c30ccd6

Observation 2feb1fc3-d790-4a12-aaf4-a92cecd2fceb · outbound

This paper cites Generating fea- tures with increased crop-related diversity for few-shot ob- ject detection.

Visual Textualization for Image Prompted Object Detection Generating fea- tures with increased crop-related diversity for few-shot ob- ject detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.469861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:07.524136Z digest=sha256:9b421d22c7132cfc4706eef0f78d64a715d45568021d18c5a90f839a8cea32fc

Observation a057a633-5693-40a9-82c8-13bf30038f4f · outbound

This paper cites Multi-modal queried object detection in the wild.

Visual Textualization for Image Prompted Object Detection Multi-modal queried object detection in the wild

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.454659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:07.577668Z digest=sha256:029dd392afda85248720409d1b14de36b5fed8610c37f1242fd966b0b919c3ce

Observation 0c0f5449-7811-4090-ba43-158692c4fef9 · outbound

This paper cites DeepLesion: Automated Deep Mining, Categorization and Detection of Significant Radiology Image Findings using Large-Scale Clinical Lesion Annotations.

Visual Textualization for Image Prompted Object Detection DeepLesion: Automated Deep Mining, Categorization and Detection of Significant Radiology Image Findings using Large-Scale Clinical Lesion Annotations

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.600258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.600258Z digest=sha256:e6ca89a14bb28fd7b4b9e7d37f09304559be432920063e4efaaf85ff2d99c6ac

Observation 73132bc0-6d32-4dd3-b3f0-56030f5c197d · outbound

This paper cites Meta r-cnn: Towards general solver for instance-level low-shot learning.

Visual Textualization for Image Prompted Object Detection Meta r-cnn: Towards general solver for instance-level low-shot learning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.438524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:07.694527Z digest=sha256:510ebe4bb5283541334cdad2ef5f9be3d30ba3537a2161717f7972a237b52222

Observation e0e62d53-abc9-42a2-90d8-3c4e5d8d3c68 · outbound

This paper cites Meta-detr: Image-level few-shot detection with inter-class correlation exploitation.

Visual Textualization for Image Prompted Object Detection Meta-detr: Image-level few-shot detection with inter-class correlation exploitation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.421162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:07.849351Z digest=sha256:3774b626fba11a2e8c0bd31681388c61d17a766b1a5af58758f21db10074cfc7

Observation e6ad7f42-c487-4313-b2c1-295f968ec37e · outbound

This paper cites Detect Everything with Few Examples.

Visual Textualization for Image Prompted Object Detection Detect Everything with Few Examples

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.017364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.017364Z digest=sha256:54f973dfaff371cda52add9f5db82b80a4ac098e80209aa6ad3605da76cc4a71

Observation 42ad3728-b973-43cf-af2f-2cb443f79a27 · outbound

This paper cites Vlm-guided explicit-implicit complementary novel class semantic learning for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Vlm-guided explicit-implicit complementary novel class semantic learning for few-shot object detection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.399983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.184430Z digest=sha256:2f39dd8bdf23bef3cdf2b45fe4a6250479bf9f30bec2a2add30bcafdb7330f4d

Observation a5d45fcf-aa5b-47a6-8f2a-2ad0baa3c275 · outbound

This paper cites Scene-adaptive and region-aware multi-modal prompt for open vocabulary object detection.

Visual Textualization for Image Prompted Object Detection Scene-adaptive and region-aware multi-modal prompt for open vocabulary object detection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.377603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.351578Z digest=sha256:8318c205aff58524c47ad5abe03cd2e7176f53b968d9646caf7bfd5e1e349563

Observation 08905e76-b956-41e3-930d-b50160e36131 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

Visual Textualization for Image Prompted Object Detection Regionclip: Region-based language-image pretraining

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.355023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.516525Z digest=sha256:837b0630879d241ff14bede6cfe4286f2032dc7868ddf00142dee2023eaec4a5

Observation 5cd787c3-1d4a-4e53-82d9-7dfe83fd4d3c · outbound

This paper cites Conditional prompt learning for vision-language models.

Visual Textualization for Image Prompted Object Detection Conditional prompt learning for vision-language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.336727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.683479Z digest=sha256:845740c49c2d6f5815d073fb4d5fc2c9d4ddc12c9f8b6601ff5d614d9f59529a

Observation f5b47bc4-f299-46af-a8da-1e0ec522b2e6 · outbound

This paper cites Learning to prompt for vision-language models.

Visual Textualization for Image Prompted Object Detection Learning to prompt for vision-language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.322179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.801070Z digest=sha256:4ccc307073731b47574fc6b51a782684c397015079e91c0fac192e40ba78e189

Observation 0e5d6e64-5eae-4609-bc1c-f0d391d14a4d · outbound

This paper cites fully connected (fc) + ReLU.

Visual Textualization for Image Prompted Object Detection fully connected (fc) + ReLU

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.304934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.863444Z digest=sha256:c21fbaba4904d4802a2a54900ad78a8547938aca51de0473c1cbacf498b6df94

Observation 08b86e82-9e0c-4100-8457-b110e286e90a · outbound

This paper cites 4.3 of the main text, we provide detailed transfer results on the ODinW13 subsets [ 31] in Tab.

Visual Textualization for Image Prompted Object Detection 4.3 of the main text, we provide detailed transfer results on the ODinW13 subsets [ 31] in Tab

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.283734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.869019Z digest=sha256:643405286a27b28d54ffe1730dfd6e35973cb8f9e8746f1abb3843f5dedc4dbe

Observation d24f22c8-c160-4757-98b8-9f389092278b · outbound

This paper cites an unresolved cited work.

Visual Textualization for Image Prompted Object Detection Unresolved cited work

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:37:09.267762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.873838Z digest=sha256:8c24e1ca098353d249150addddbd8302240192962945faa0e754bcf0a762cb23

Observation 089e77eb-2e4a-4a55-8010-4c8199a976db · outbound

This paper cites 8, we report the computational overhead for process- ing one image using GLIP-L on RTX3090 with one support image, comparing it to MQ-Det and GLIP-FF.

Visual Textualization for Image Prompted Object Detection 8, we report the computational overhead for process- ing one image using GLIP-L on RTX3090 with one support image, comparing it to MQ-Det and GLIP-FF

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.247972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.880919Z digest=sha256:986ce1fc1f473f8b9462a628613a05885b109c67e451021e6da8202b1cb3185e

Observation 254e521a-4c25-470c-9c16-6511e3952a71 · outbound

This paper cites BG blur" technique performs best. It high- lights the target object while preserving some background, unlike.

Visual Textualization for Image Prompted Object Detection BG blur" technique performs best. It high- lights the target object while preserving some background, unlike

Reference 67

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:37:09.227510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.889886Z digest=sha256:300a0d819eaffa776c1273b6a98d33f7869b65a762d2b45615493fcf92866170

Observation a9ddc82e-4ae4-4035-b9c4-0517e3f24bc1 · outbound

This paper cites Base- line.

Visual Textualization for Image Prompted Object Detection Base- line

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.203949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:37:08.897852Z digest=sha256:8a8a3927605d794f519603ac9a1293ac70f964386910d0372c044f81463d6f40

Pith citing papers

No inbound Pith citation observations are available.