Pith. sign in

Paper Citation Record · LEDGER

Visual Textualization for Image Prompted Object Detection

As of 8 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2506.23785.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23785 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:37:08.897852Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved13
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf10745e-c8e4-4e84-81dc-bf97e68189d9 · outbound

This paper cites Lung image database consor- tium: developing a resource for the medical imaging research community.

Visual Textualization for Image Prompted Object Detection Lung image database consor- tium: developing a resource for the medical imaging research community

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.347738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:02.947288Z digest=sha256:65fe585e6bb31b42864f10f9522687ddd14557c77234006352a11ed7cd4fe9d4

Observation a8a31ae4-53dd-4930-936c-bc8d7851bbba · outbound

This paper cites Exploring Visual Prompts for Adapting Large-Scale Models.

Visual Textualization for Image Prompted Object Detection Exploring Visual Prompts for Adapting Large-Scale Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.071096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.071096Z digest=sha256:1106a81374ec2c346f1e2d3e761bd9efe9e2ed6debb3cb78b607c0f34da3b838

Observation 494e3f34-7810-4fca-94fa-730bd867c9cc · outbound

This paper cites Fs-detr: Few-shot detection transformer with prompting and without re-training.

Visual Textualization for Image Prompted Object Detection Fs-detr: Few-shot detection transformer with prompting and without re-training

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.332668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:03.162281Z digest=sha256:4ba9a0d6581fc816f9af160473dd8786baaa40f1940000d0e80d4e9251218732

Observation c10b563c-4060-489f-b9b8-d1a0a9c4ea95 · outbound

This paper cites Apollo: Unified adapter and prompt learning for vision language models.

Visual Textualization for Image Prompted Object Detection Apollo: Unified adapter and prompt learning for vision language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.316861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:03.251458Z digest=sha256:3419e4a2984a871277f7aea7a9d319b9a9330af0a6ace1837d48301f179b9749

Observation f4129d0f-d7d2-41ab-ab60-dc6c1fc394c2 · outbound

This paper cites Coarse-to-fine vision-language pre-training with fusion in the backbone.NeurIPS, 35:32942–32956, 2022.

Visual Textualization for Image Prompted Object Detection Coarse-to-fine vision-language pre-training with fusion in the backbone.NeurIPS, 35:32942–32956, 2022

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.299990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:03.345641Z digest=sha256:4800ef4a1f144e523d22a4c94c78f208f2a8fd99ef88ecfae11fee2d5fa2c6f7

Observation b47ad70b-726a-4417-8822-148ba1e39462 · outbound

This paper cites s- adaptive decoupled prototype for few-shot object detection.

Visual Textualization for Image Prompted Object Detection s- adaptive decoupled prototype for few-shot object detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.276015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:03.421737Z digest=sha256:65ef82af3e574ca9d9b5a2bf73842da29526ea243085b41ef57a12661479db26

Observation 02332cf4-9657-440c-b52a-6f5b0fca1064 · outbound

This paper cites Learning to prompt for open-vocabulary object detection with vision-language model.

Visual Textualization for Image Prompted Object Detection Learning to prompt for open-vocabulary object detection with vision-language model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.258469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:03.502861Z digest=sha256:e2aa0df491fa7f57f888f184f578db9de89709a3f0c38d8dede5a6397864062e

Observation fc7a5f35-5371-4a1f-8b5f-125bf38d7737 · outbound

This paper cites The Turking Test: Can Language Models Understand Instructions?.

Visual Textualization for Image Prompted Object Detection The Turking Test: Can Language Models Understand Instructions?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.644571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.644571Z digest=sha256:3e67bf12b226fa9b5e49db5797acd405c46acf4d517a7a77edf11fe5ea0ee218

Observation 5d4286b5-c630-43ea-91e4-2648e1275036 · outbound

This paper cites The pascal visual object classes (voc) challenge.

Visual Textualization for Image Prompted Object Detection The pascal visual object classes (voc) challenge

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.235878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:03.743469Z digest=sha256:248d43e3782a0461507cc3e3a6e68434621ce69a9597023ec2bc198d2562d22d

Observation da991085-3e35-48f9-92af-c7c2e6ba7bef · outbound

This paper cites Few- shot object detection with attention-rpn and multi-relation detector.

Visual Textualization for Image Prompted Object Detection Few- shot object detection with attention-rpn and multi-relation detector

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:03.829699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:03.829699Z digest=sha256:4705de6eb4d5c354e2aaabb2868fc96ec147285aec89d1b5973b0f7887e8e2f7

Observation d5dc14cd-cb45-4eec-8789-e2f26e00a6e2 · outbound

This paper cites Nuclei grading of clear cell renal cell carcinoma in histopatho- logical image by composite high-resolution network.

Visual Textualization for Image Prompted Object Detection Nuclei grading of clear cell renal cell carcinoma in histopatho- logical image by composite high-resolution network

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.205060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:03.904842Z digest=sha256:4252109435f9db4300c327312fd0d8841d5f542dcc792e62538182e8edf04f33

Observation 0c33a867-fc55-40ee-9428-acedea4ca5b9 · outbound

This paper cites Hover-net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images.

Visual Textualization for Image Prompted Object Detection Hover-net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.185000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:04.004247Z digest=sha256:789056fe8e2858de2d6ad8c50c7b2068f0c0bfa0278463e606ac971cb1a4c94c

Observation 482ae711-fa47-4911-ab40-5dd7c88d34a5 · outbound

This paper cites A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models.

Visual Textualization for Image Prompted Object Detection A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.064646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.064646Z digest=sha256:a85866771c69e5bc2503680916d891c3895ca4fdbe41c6b44108dc98f7d56dcf

Observation 966a79fa-1baa-48e3-bde5-e59e1b4c9e28 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Visual Textualization for Image Prompted Object Detection Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.138434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.138434Z digest=sha256:e92d3d06fe9338ac88d19eae39cd10642453abe3f3dac762fbfad1ec7a3fb753

Observation d248f254-2407-4685-a5a7-ce97d22d99b7 · outbound

This paper cites Dp-ddcl: A discriminative prototype with dual decou- pled contrast learning method for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Dp-ddcl: A discriminative prototype with dual decou- pled contrast learning method for few-shot object detection

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.166743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:04.250057Z digest=sha256:fbe4ac299a087c73bf9c9e7c48585791b06dd4c5214419d4279205b512540d50

Observation 2c1e0f48-be20-4d20-b198-254041fd36b0 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Visual Textualization for Image Prompted Object Detection Lvis: A dataset for large vocabulary instance segmentation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.149088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:04.369673Z digest=sha256:53723508e314b77345de0d5e8fdb59d97bbb44bd006dc24fada949c7417d62a2

Observation 75988dd5-13ac-45a7-8487-7f4ad5410f54 · outbound

This paper cites Few-shot object detection with foundation models.

Visual Textualization for Image Prompted Object Detection Few-shot object detection with foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.135836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:04.489927Z digest=sha256:361d0b7d76d2d99d380333227cdb03559e0724f622545fbb41d40e2814d3139c

Observation e848ea3a-8a08-4488-8694-157cb8550f99 · outbound

This paper cites Query adaptive few-shot object detec- tion with heterogeneous graph convolutional networks.

Visual Textualization for Image Prompted Object Detection Query adaptive few-shot object detec- tion with heterogeneous graph convolutional networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.119455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:04.567751Z digest=sha256:ec3c2f148e790c75ffa371f80f219f2023a9ba5a1d3a6301e7848306cb61674f

Observation 0eb56578-ed90-408c-b085-e6eee5a18312 · outbound

This paper cites Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting.

Visual Textualization for Image Prompted Object Detection Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:04.660592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:04.660592Z digest=sha256:a4305ab441454208ff1c06ea022422e02ee3374d56530461e948e4b7310fa57a

Observation 0b083cbe-5052-4226-b5da-d752dfbf10b3 · outbound

This paper cites Meta faster r-cnn: Towards accurate few-shot object detection with attentive feature alignment.

Visual Textualization for Image Prompted Object Detection Meta faster r-cnn: Towards accurate few-shot object detection with attentive feature alignment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.100727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:04.727426Z digest=sha256:2a9be9202bc33b48fe53482f0a509b8a501d3e45f07cc5131d39c3978b704c30

Observation 68627033-c1b9-4d89-9ce7-a487d2873efb · outbound

This paper cites Few-shot object detection with fully cross- transformer.

Visual Textualization for Image Prompted Object Detection Few-shot object detection with fully cross- transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:10.084388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:04.978030Z digest=sha256:45848139920e913bf65d71607fd11ca12ef8e8f627e77cef5e265cdc113545f3

Observation a3f2ac37-d0a0-4ebb-8c91-099a6c3c8f29 · outbound

This paper cites Few-shot object detection via variational feature aggregation.

Visual Textualization for Image Prompted Object Detection Few-shot object detection via variational feature aggregation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.947853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:05.120058Z digest=sha256:f43b7706cf0e0efc32685dd1cdca928cac068f3f0effa5c67d004ef2ad12006b

Observation 02f4ff15-dfd7-4c74-ba03-bcff5b1ea601 · outbound

This paper cites Visual prompt tuning.

Visual Textualization for Image Prompted Object Detection Visual prompt tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.931188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:05.202072Z digest=sha256:ffcd7d9edd1b9a3178cd9455fb499d9eddef7df2801275a0d54c0c2c5064d5b0

Observation d8e8e50f-2e4b-4c46-af8e-5097595a48d5 · outbound

This paper cites Bert: Pre-training of deep bidirectional transform- ers for language understanding.

Visual Textualization for Image Prompted Object Detection Bert: Pre-training of deep bidirectional transform- ers for language understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.913726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:05.307636Z digest=sha256:b351a5c51cad0198f04a0ff17384f17ce943ad2f77292d92773667f92b8362ed

Observation 8fb616bd-3570-45dc-8811-9f2bf38cb355 · outbound

This paper cites Maple: Multi- modal prompt learning.

Visual Textualization for Image Prompted Object Detection Maple: Multi- modal prompt learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.891918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:05.404038Z digest=sha256:8ec99e494a5380ae6ff1368297fa1d8888e130d2dbb123dc723099292c3870a2

Observation 5b792927-2702-4896-8eec-86b34a410cca · outbound

This paper cites A dataset and a technique for generalized nuclear segmentation for computa- tional pathology.

Visual Textualization for Image Prompted Object Detection A dataset and a technique for generalized nuclear segmentation for computa- tional pathology

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.871760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:05.530990Z digest=sha256:ab1f9ced8156d550dc20ce8998ff4fd94bb3552f49cde3d7722073d454d5a2ef

Observation d737176b-afa7-4a34-9af7-dd51e16057d7 · outbound

This paper cites F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models.

Visual Textualization for Image Prompted Object Detection F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.649074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.649074Z digest=sha256:6cca76b70984ad668947562e2e4ebc34115aef451e2ef38abaea6beab65bed8a

Observation 7be32b2c-81c1-47a8-b5ac-15dbca57af3b · outbound

This paper cites Elevater: A benchmark and toolkit for evaluating language-augmented visual models.

Visual Textualization for Image Prompted Object Detection Elevater: A benchmark and toolkit for evaluating language-augmented visual models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.855850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:05.790125Z digest=sha256:0566c3d3e52fca66921cf88385108390fd45fc84d47681ec5da3c08289966302

Observation 002258a5-4867-4e66-936b-65487eb0445c · outbound

This paper cites Disentangle and remerge: interventional knowledge distillation for few-shot object detection from a conditional causal perspective.

Visual Textualization for Image Prompted Object Detection Disentangle and remerge: interventional knowledge distillation for few-shot object detection from a conditional causal perspective

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.841297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:05.875694Z digest=sha256:191d196463ec29ec4ef4b153acae8691a81f1c17afe39261a31cfbe1de539962

Observation e4c5e956-6447-4a4a-9cc8-20fcf753f8c8 · outbound

This paper cites Grounded language- image pre-training.

Visual Textualization for Image Prompted Object Detection Grounded language- image pre-training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.820928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:05.969037Z digest=sha256:c755f0ac72fbc1e4ecb8733c98f7a49b26beedbd78036d75bc5b736c1b4961f8

Observation 9edc8e46-7db7-4c90-9e95-dd493c8605c8 · outbound

This paper cites Microsoft coco: Common objects in context.

Visual Textualization for Image Prompted Object Detection Microsoft coco: Common objects in context

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.788346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:06.045211Z digest=sha256:f747c747e0da63347670fab7dbfc1d3e210ce73fc520fa05ac56a74c3b0f848d

Observation b2801c4e-e958-4df3-9fca-ba4b3f05eab5 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Visual Textualization for Image Prompted Object Detection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.107791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.107791Z digest=sha256:7f5100c6cde6ffb9971095201418ae86b96044269078d5fee58979b4e0207733

Observation 01277d58-5a56-4099-b94e-764c9deffc45 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Visual Textualization for Image Prompted Object Detection Swin transformer: Hierarchical vision transformer using shifted windows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.250780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.250780Z digest=sha256:cc9f3895d35b444422a187283039b21a2a470ab64c4d0f5e4c67a71e7aed320b

Observation 739b0451-22e2-4f15-a6aa-48fd1e79a463 · outbound

This paper cites Breaking immutable: Information-coupled prototype elaboration for few-shot ob- ject detection.

Visual Textualization for Image Prompted Object Detection Breaking immutable: Information-coupled prototype elaboration for few-shot ob- ject detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.756412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:06.324343Z digest=sha256:3fd7e3797372f701442284844099a0639f375b696e3e73ebbf0a1469d7413cf6

Observation c39a64d9-a64f-48d1-99da-e3a2469eb890 · outbound

This paper cites Image segmentation us- ing text and image prompts.

Visual Textualization for Image Prompted Object Detection Image segmentation us- ing text and image prompts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.737378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:06.388853Z digest=sha256:806042bd378934de6d6d8b87bdd326e1144b53f930da0e6b793ba7b3f2732115

Observation f0c2cc35-7f82-4ae8-af0e-e8d608ccaf31 · outbound

This paper cites Digeo: Discriminative geometry-aware learning for generalized few-shot object de- tection.

Visual Textualization for Image Prompted Object Detection Digeo: Discriminative geometry-aware learning for generalized few-shot object de- tection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.722011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:06.440503Z digest=sha256:a41a65666cfe2fc88b28eca5f7605ddb1fc88f16762aa1689c86cabb6991ef00

Observation 5f89670f-611c-4991-91f4-bc0e8c45a54a · outbound

This paper cites Simple open-vocabulary object detection.

Visual Textualization for Image Prompted Object Detection Simple open-vocabulary object detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.701594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:06.497243Z digest=sha256:8ccdf1929cee6618c660f68fd46a614f5a85fc4cc4029f62cd5a15ea50e0ab0c

Observation e0af7dcd-d85c-45f1-b10d-0fdc4ff3770f · outbound

This paper cites Scal- ing open-vocabulary object detection.

Visual Textualization for Image Prompted Object Detection Scal- ing open-vocabulary object detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.678121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:06.557963Z digest=sha256:6520e1540ac1b421bbc26a3d380e00242c316318f74090ba9dba31912f0653d0

Observation a907e535-54c4-4178-8f1b-7ead7169f35e · outbound

This paper cites Defrcn: Decoupled faster r-cnn for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Defrcn: Decoupled faster r-cnn for few-shot object detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.659393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:06.624223Z digest=sha256:e6b94fd46ae4b0717f630218ff95b2d13c1e23d20563d6b2613c0f6e871e27af

Observation 60171d71-81d6-4b39-a81b-16cc0df5ae2a · outbound

This paper cites Language models are unsuper- vised multitask learners.

Visual Textualization for Image Prompted Object Detection Language models are unsuper- vised multitask learners

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.638175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:06.665207Z digest=sha256:e6b6a9e21be7dde38cac52014360478072790df77a1720496c2f5c01149a4d4b

Observation 3de2fc5b-5b78-44e4-a450-e5c16d79a758 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Visual Textualization for Image Prompted Object Detection Learning transferable visual models from natural language supervi- sion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.611623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:06.744231Z digest=sha256:7e0f61c2259664c46ed3ab0bf25e20cc986989e983965629b8d02a349cbfa689

Observation 6157ef5e-b882-49eb-a9d0-f5d86407f505 · outbound

This paper cites Adaptive multi-task learning for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Adaptive multi-task learning for few-shot object detection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.588564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:06.869279Z digest=sha256:b1ef4954e7db56f468529b7c51b6047ebab40a84a39120ff409d96a47e73aeaf

Observation 9042d0fc-ffe6-4c22-afa7-39b41f8cacbf · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Visual Textualization for Image Prompted Object Detection Objects365: A large-scale, high-quality dataset for object detection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:06.980275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:06.980275Z digest=sha256:9f7796f9d687306fbd8c001d1d60948992a9524d85712375f0be4760e84b4890

Observation 9b62f9e8-eb53-4952-bd87-25f58b822722 · outbound

This paper cites Few- shot adaptive faster r-cnn.

Visual Textualization for Image Prompted Object Detection Few- shot adaptive faster r-cnn

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.551444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:07.045297Z digest=sha256:e1923f207975d59361e818593e981b4bc613ab0a2fc4fe38c39ca037cc8787fc

Observation 86f89577-7cb6-453c-810b-285054f4d9b3 · outbound

This paper cites Frustratingly Simple Few-Shot Object Detection.

Visual Textualization for Image Prompted Object Detection Frustratingly Simple Few-Shot Object Detection

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.148413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.148413Z digest=sha256:acee2d69376ac249a5a582d3522005397876561ec900cc9089b2d2278721c485

Observation 3e24ebcb-cddf-4047-90bd-acd227207b64 · outbound

This paper cites Snida: Unlocking few-shot object detection with non- linear semantic decoupling augmentation.

Visual Textualization for Image Prompted Object Detection Snida: Unlocking few-shot object detection with non- linear semantic decoupling augmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.530773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:07.207653Z digest=sha256:f60756e8f9c10099dd50aef317904c87e2213f28d08a66904ab2bca746934167

Observation dfa01c67-5c5f-4333-b46e-b41aa717ed20 · outbound

This paper cites Multi- scale positive sample refinement for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Multi- scale positive sample refinement for few-shot object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.516661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:07.267887Z digest=sha256:70113c3cf357180053816053fa3fdc63f80a0cd8c1259c54ba4472ef435c4431

Observation c7e19578-69fe-445c-abf4-20e2d9dc6474 · outbound

This paper cites Multi-faceted distillation of base-novel commonality for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Multi-faceted distillation of base-novel commonality for few-shot object detection

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.501311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:07.331054Z digest=sha256:9d522771253d267eebb462b43a08e913f8db7c6db3c51180492fe4a3f12d8733

Observation 1767531d-f860-461e-b608-09d2122bdc64 · outbound

This paper cites Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching.

Visual Textualization for Image Prompted Object Detection Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.485161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:07.456827Z digest=sha256:d106bee0a1f6e0a711f320e8df905057cd16c939f7dec8c83012bddfd6bec5d9

Observation 2feb1fc3-d790-4a12-aaf4-a92cecd2fceb · outbound

This paper cites Generating fea- tures with increased crop-related diversity for few-shot ob- ject detection.

Visual Textualization for Image Prompted Object Detection Generating fea- tures with increased crop-related diversity for few-shot ob- ject detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.469861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:07.524136Z digest=sha256:c802b501748fa4943cc1c5fe7757ce953accba80dd8beb4434d4fb81fe8c2670

Observation a057a633-5693-40a9-82c8-13bf30038f4f · outbound

This paper cites Multi-modal queried object detection in the wild.

Visual Textualization for Image Prompted Object Detection Multi-modal queried object detection in the wild

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.454659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:07.577668Z digest=sha256:8c641f33bd7eecb966f83f59098af0ae313bd757a40138e7585580c0bcab0864

Observation 0c0f5449-7811-4090-ba43-158692c4fef9 · outbound

This paper cites DeepLesion: Automated Deep Mining, Categorization and Detection of Significant Radiology Image Findings using Large-Scale Clinical Lesion Annotations.

Visual Textualization for Image Prompted Object Detection DeepLesion: Automated Deep Mining, Categorization and Detection of Significant Radiology Image Findings using Large-Scale Clinical Lesion Annotations

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:07.600258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:07.600258Z digest=sha256:e6ca89a14bb28fd7b4b9e7d37f09304559be432920063e4efaaf85ff2d99c6ac

Observation 73132bc0-6d32-4dd3-b3f0-56030f5c197d · outbound

This paper cites Meta r-cnn: Towards general solver for instance-level low-shot learning.

Visual Textualization for Image Prompted Object Detection Meta r-cnn: Towards general solver for instance-level low-shot learning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.438524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:07.694527Z digest=sha256:513fc4c878c8280a016c557e8deb56b2c86485649f74eccbffc350e2f660dad3

Observation e0e62d53-abc9-42a2-90d8-3c4e5d8d3c68 · outbound

This paper cites Meta-detr: Image-level few-shot detection with inter-class correlation exploitation.

Visual Textualization for Image Prompted Object Detection Meta-detr: Image-level few-shot detection with inter-class correlation exploitation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.421162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:07.849351Z digest=sha256:f4583fe800d501c9da12492d8fd4d8bc26de381df7fd9094bd3fe4f936604626

Observation e6ad7f42-c487-4313-b2c1-295f968ec37e · outbound

This paper cites Detect Everything with Few Examples.

Visual Textualization for Image Prompted Object Detection Detect Everything with Few Examples

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:08.017364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:08.017364Z digest=sha256:54f973dfaff371cda52add9f5db82b80a4ac098e80209aa6ad3605da76cc4a71

Observation 42ad3728-b973-43cf-af2f-2cb443f79a27 · outbound

This paper cites Vlm-guided explicit-implicit complementary novel class semantic learning for few-shot object detection.

Visual Textualization for Image Prompted Object Detection Vlm-guided explicit-implicit complementary novel class semantic learning for few-shot object detection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.399983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:08.184430Z digest=sha256:91be119bcc1419a43b8b53ea59fdd14f7b0d447528732a50be8633d0d99521e8

Observation a5d45fcf-aa5b-47a6-8f2a-2ad0baa3c275 · outbound

This paper cites Scene-adaptive and region-aware multi-modal prompt for open vocabulary object detection.

Visual Textualization for Image Prompted Object Detection Scene-adaptive and region-aware multi-modal prompt for open vocabulary object detection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.377603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:08.351578Z digest=sha256:1a0bd6672ef131ba81bfbd2292cde48ecf3339018e68c64b7ea44176e878cad9

Observation 08905e76-b956-41e3-930d-b50160e36131 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

Visual Textualization for Image Prompted Object Detection Regionclip: Region-based language-image pretraining

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.355023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:08.516525Z digest=sha256:19ea80f5da167a2c32cc23bdc6a1672c228c8820031bdc64aeede19a8ef2fd04

Observation 5cd787c3-1d4a-4e53-82d9-7dfe83fd4d3c · outbound

This paper cites Conditional prompt learning for vision-language models.

Visual Textualization for Image Prompted Object Detection Conditional prompt learning for vision-language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.336727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:08.683479Z digest=sha256:158474427d841c69e8cfee5fa336ad99d850d6d31ffe86aae533a4c6498c8340

Observation f5b47bc4-f299-46af-a8da-1e0ec522b2e6 · outbound

This paper cites Learning to prompt for vision-language models.

Visual Textualization for Image Prompted Object Detection Learning to prompt for vision-language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.322179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:08.801070Z digest=sha256:0f433e8a9992913d0021189f80bd8bfe7b936585dcdd2ae93b71fbd7bef23fd2

Observation 0e5d6e64-5eae-4609-bc1c-f0d391d14a4d · outbound

This paper cites fully connected (fc) + ReLU.

Visual Textualization for Image Prompted Object Detection fully connected (fc) + ReLU

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.304934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:08.863444Z digest=sha256:ec6b953a1c62c5d465f64dd3240d15b5c042993cbfe3f9591c6c6a2498e8e6ce

Observation 08b86e82-9e0c-4100-8457-b110e286e90a · outbound

This paper cites 4.3 of the main text, we provide detailed transfer results on the ODinW13 subsets [ 31] in Tab.

Visual Textualization for Image Prompted Object Detection 4.3 of the main text, we provide detailed transfer results on the ODinW13 subsets [ 31] in Tab

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.283734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:08.869019Z digest=sha256:884c3ae3db4849f4cd9124729499fd36d1364911ab889978dd63f2b2809b59d1

Observation d24f22c8-c160-4757-98b8-9f389092278b · outbound

This paper cites an unresolved cited work.

Visual Textualization for Image Prompted Object Detection Unresolved cited work

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:37:09.267762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:08.873838Z digest=sha256:38e0bdbe24c023e9b7b57697e01952ec12f31e5dd220f6a31317317618303a4e

Observation 089e77eb-2e4a-4a55-8010-4c8199a976db · outbound

This paper cites 8, we report the computational overhead for process- ing one image using GLIP-L on RTX3090 with one support image, comparing it to MQ-Det and GLIP-FF.

Visual Textualization for Image Prompted Object Detection 8, we report the computational overhead for process- ing one image using GLIP-L on RTX3090 with one support image, comparing it to MQ-Det and GLIP-FF

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.247972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:08.880919Z digest=sha256:281b6038e0deada7f08ecb55da8b3fd08351bcb814ec77d9135a85de006cbf96

Observation 254e521a-4c25-470c-9c16-6511e3952a71 · outbound

This paper cites BG blur" technique performs best. It high- lights the target object while preserving some background, unlike.

Visual Textualization for Image Prompted Object Detection BG blur" technique performs best. It high- lights the target object while preserving some background, unlike

Reference 67

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:37:09.227510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:08.889886Z digest=sha256:73ca13b731a976fed2fc4c6cd270f01c7ded69008c014596e4c1aa0d2f145440

Observation a9ddc82e-4ae4-4035-b9c4-0517e3f24bc1 · outbound

This paper cites Base- line.

Visual Textualization for Image Prompted Object Detection Base- line

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:37:09.203949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:37:08.897852Z digest=sha256:3c8a090a716a84a7aad8af74ded8f17ce0beff505877cf3cf4ccc85c7828c2d1

Pith citing papers

No inbound Pith citation observations are available.