Pith. sign in

Paper Citation Record · LEDGER

Visual Modality Prompt for Adapting Vision-Language Object Detectors

As of 13 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2412.00622.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00622 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:14:07.816202Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:27:41.333932Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact2
  • verified fuzzy30
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a167833-5801-41e8-8b56-083c57f186d2 · outbound

This paper cites Exploring Visual Prompts for Adapting Large-Scale Models.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Exploring Visual Prompts for Adapting Large-Scale Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.584130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.584130Z digest=sha256:86fbf406f854f4d28c324edb790080b911103c2849974f7b2273c25750c015db

Observation 6fca56ee-a54d-4c6a-8b8e-d89335892ef6 · outbound

This paper cites Zero-shot object detection.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Zero-shot object detection

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.462880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.591559Z digest=sha256:53a226e4f4f6debe47a3f9b7094e491484edf7665559e5ecf0a53d69b47a6e76

Observation 840ae565-9265-4906-8987-7ccfecfc9ccd · outbound

This paper cites A systematic literature review on object detection using near infrared and thermal images.

Visual Modality Prompt for Adapting Vision-Language Object Detectors A systematic literature review on object detection using near infrared and thermal images

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.449862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.596017Z digest=sha256:a7912f9aaca544197a3cba1a88c18b5b0d073f809b1a96e6884e93a5d56a6348

Observation 095d2fb8-baf9-40a2-967a-64c0c81c31a0 · outbound

This paper cites Multimodal object detection by channel switching and spatial attention.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Multimodal object detection by channel switching and spatial attention

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.437720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.599778Z digest=sha256:a38f1e875c0e9fd18cb7fe68bd13e6764c9157edcb90c81fc4b7f7c85b7f2c4e

Observation d4fd734f-ce10-4f19-8d11-55a4c90b169b · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

Visual Modality Prompt for Adapting Vision-Language Object Detectors MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.603671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.603671Z digest=sha256:45ed6807c9d26f1309ea45ea0d7794affddc3bccf33858ceb12ea9348636f4a2

Observation 1c495f39-6dbc-450c-a721-74e228b975d5 · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Yolo-world: Real-time open-vocabulary object detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.426259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.607890Z digest=sha256:90aab43b447f1629e91673cbd0947927ec336fd6aa9995674a456f6ac45feb73

Observation e3fac7cd-91ee-4457-9875-6bd37c11ecbb · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Visual Modality Prompt for Adapting Vision-Language Object Detectors BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.613076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.613076Z digest=sha256:c6902f5d7c5e2f5451bbd3f614e3c0866aab168815d62b6a02af4bfe236d0e6a

Observation 5e86fe4f-a111-4e82-9606-52ea2f79e351 · outbound

This paper cites Privacy-preserving person detection using low-resolution infrared cameras.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Privacy-preserving person detection using low-resolution infrared cameras

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.413352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.617461Z digest=sha256:82a51a400111cd281b5b2c70fac277c748a8213c8445883c281c5f1f0de9d338

Observation 4a41182c-f8d0-46e9-ba77-10b1502ced32 · outbound

This paper cites Multimodal deep learning for robust rgb-d object recognition.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Multimodal deep learning for robust rgb-d object recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.400719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.622230Z digest=sha256:db78d4f33a7da8a66d5986f1ff8e818809664f389de755d46182a7b82983d7bd

Observation 0660f560-94c3-4928-aeae-883dd16cdc53 · outbound

This paper cites Tood: Task-aligned one-stage object detec- tion.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Tood: Task-aligned one-stage object detec- tion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.385766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.627099Z digest=sha256:653fbbb467a91a2fb33585f1bde836b530c9f5ba82065dac70a7d252d5637158

Observation a1f1076c-6eda-4b51-b22f-786bc0cbcacc · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Clip-adapter: Better vision-language models with feature adapters

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.631293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.631293Z digest=sha256:69c37d9ab5164620afadb1fd5bb74e97fafe7fe79066ff03c0eb05391b797f79

Observation da194059-8a97-4f64-a5d1-24a3ad54dddd · outbound

This paper cites Generative adversarial nets.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Generative adversarial nets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.635483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.635483Z digest=sha256:3e7129fde0c4105c3202fc8f44cfed3733272b3e4ca314ab2cfef5acd3703b7f

Observation 225828cf-b3e3-44f4-9d96-34c7faa66649 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.639541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.639541Z digest=sha256:aa5f10fa7cba825b11e93807311f7fee24d103f1d3de211ec03b59f7e072b729

Observation 4d90fa25-4ccf-48a0-b996-2ebae8562ed7 · outbound

This paper cites Deep residual learning for image recognition.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Deep residual learning for image recognition

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.351235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.643921Z digest=sha256:14fc15c806551bfb7f31a29ddf24053d5ddb089caef23f2b433fa56c566226a5

Observation 1c167d08-4ca1-4ade-9490-174fbb91f9a3 · outbound

This paper cites Cnn- based thermal infrared person detection by domain adapta- tion.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Cnn- based thermal infrared person detection by domain adapta- tion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.338030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.648697Z digest=sha256:7b88fbc1b22f7dc7b9385b7d4ecf2c7bf2557b4f5c6f93b70696fac03bb66845

Observation 6d486159-7bd9-4c47-a84c-5c2b691f2c04 · outbound

This paper cites MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications.

Visual Modality Prompt for Adapting Vision-Language Object Detectors MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.652300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.652300Z digest=sha256:8072b008a468033eb13826bbd56ba69c2f6b3e64ddcd0c513bd7d94bebba14ce

Observation b1d9d204-f1cf-49f8-bcb8-33b4860a11ff · outbound

This paper cites Image-to-image translation with conditional adver- sarial networks.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Image-to-image translation with conditional adver- sarial networks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.327234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.656848Z digest=sha256:029d9794c2b64b43778c03aedca7276b7a079194a742125370b22b7d4ac0b284

Observation 0417fcb2-0601-4b50-bfa0-bf93f472e162 · outbound

This paper cites Multimodal computer vision framework for human assistive robotics.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Multimodal computer vision framework for human assistive robotics

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.314408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.660904Z digest=sha256:1f0c8eb5159d6877a84672862fdfc46e2dea33ac5d2259436dc17324614870c1

Observation cc36ed3b-85fb-4de4-9a56-96cad4b77ca7 · outbound

This paper cites Vi- sual prompt tuning.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Vi- sual prompt tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.666025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.666025Z digest=sha256:d7e3e521cc2ec337a6e3de0b743587e044e647b0ad84bf3f0a7da648a9d880f6

Observation a583c228-cd9d-4fd7-9fe8-ef01cbc51566 · outbound

This paper cites Ultralyt- ics yolov8.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Ultralyt- ics yolov8

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.297182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.670492Z digest=sha256:b6f2efceef7020751ada1ee8e6c86e81e8fa2af8aa4322da4ea3e61d62e7bb32

Observation 3d884374-38fc-47c8-b2c5-777e0a6d729f · outbound

This paper cites Mdetr – mod- ulated detection for end-to-end multi-modal understanding,.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Mdetr – mod- ulated detection for end-to-end multi-modal understanding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.675822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.675822Z digest=sha256:a8bbaaff9d2f443305eec898b07929717a74e0fc7336f7a18681b83371daee11

Observation 775e8f50-b2e9-45d3-bc9a-62181c9a2d15 · outbound

This paper cites Auto-encoding varia- tional bayes, 2022.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Auto-encoding varia- tional bayes, 2022

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.680049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.680049Z digest=sha256:2b12408d6b0f590fba2bd0ec5c6f6427423de9638aaaffea014d1c1626052ee7

Observation 21101b2f-d61c-4005-b3f4-ec6bc0aa0138 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning, 2021.

Visual Modality Prompt for Adapting Vision-Language Object Detectors The power of scale for parameter-efficient prompt tuning, 2021

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.684638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.684638Z digest=sha256:be18f8dec2912b63f889381d180b61d0e8d74d0c311cb64db69576d0bcdf4267

Observation 435e7c0b-06a7-4da0-b69d-9fde29cf566c · outbound

This paper cites Grounded language-image pre-training.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Grounded language-image pre-training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.689759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.689759Z digest=sha256:13c1dc7f4bb7bd130305eebb20928397f93515fc8b196666fb6fa44234f7dca1

Observation b9b25b53-87a1-4735-99d4-a794d1b0e708 · outbound

This paper cites Yolo-firi: Improved yolov5 for infrared image object detection.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Yolo-firi: Improved yolov5 for infrared image object detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.250096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.694269Z digest=sha256:fa108462a5119c194d8037e1ac962cb562b2994848750246e5bbd96f9819924a

Observation f19f62c2-cac1-4fa2-8012-83fc0bfc458d · outbound

This paper cites Bevdepth: Acquisition of reliable depth for multi-view 3d object detec- tion.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Bevdepth: Acquisition of reliable depth for multi-view 3d object detec- tion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.235199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.698181Z digest=sha256:fa926a987bce17d8e65cdca428f07b2e60caf1a0ea2dad596e0bf022af60b9f6

Observation b6de8078-fbeb-4780-a43e-f89e47cb20e2 · outbound

This paper cites Learning Object-Language Alignments for Open-Vocabulary Object Detection.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Learning Object-Language Alignments for Open-Vocabulary Object Detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.702032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.702032Z digest=sha256:f0514f1dab078861c9cbd8b6e34aa8a7b706c1d3e2ff5f25e08244157f1f31c1

Observation ccd5248c-08cc-4eb4-a6e6-2f4ad8c3b1f5 · outbound

This paper cites Microsoft coco: Common objects in context.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Microsoft coco: Common objects in context

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.706765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.706765Z digest=sha256:9274b2c9ebcea3fd19868103cb49e82bd06d334440467cbf51b94cb579fd57db

Observation daf16f2c-3223-4d34-8718-85a5d3381314 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.710518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.710518Z digest=sha256:79e1b968a42eef66150d3eeba4925daf943990707ed5e26eeba13d7f1015a357

Observation 25b7252f-2d4d-4dd7-837c-f90d7b0f4559 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Swin transformer: Hierarchical vision transformer using shifted windows

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.714992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.714992Z digest=sha256:6c77d7274be09d486ed2a693eb2b1636dc40bcb16b69be9e9a1edf074eca383a

Observation 1c7c6a05-b942-41e1-a9a6-6598a8cc9b88 · outbound

This paper cites Modality Translation for Object Detection Adaptation Without Forgetting Prior Knowledge.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Modality Translation for Object Detection Adaptation Without Forgetting Prior Knowledge

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:14:07.905453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.718415Z digest=sha256:3437bf1d6eb079ba44174d24fcc600bd4c7ae6ff74ad537e28b331cd1776b327

Observation 367dd8f2-dac4-4329-9015-32375c942547 · outbound

This paper cites MiPa: Mixed Patch Infrared-Visible Modality Agnostic Object Detection.

Visual Modality Prompt for Adapting Vision-Language Object Detectors MiPa: Mixed Patch Infrared-Visible Modality Agnostic Object Detection

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:14:07.889750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.723074Z digest=sha256:4fafd6ddef77c6a2f3819276c341191555aacaf36f0fcc8081108a17a7889d03

Observation bb4f3a5d-53ab-4410-9a56-58a7f29a66c5 · outbound

This paper cites Hallucidet: Hallucinating rgb modality for person de- tection through privileged information.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Hallucidet: Hallucinating rgb modality for person de- tection through privileged information

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.211656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.727176Z digest=sha256:9d9a0cd45da3b04c6fcebf9f3930996ddc0d5cb7404a2bdd3ffca8180a4ee457

Observation a9375eb4-5963-42e8-804e-db5ebc2dc2f3 · outbound

This paper cites Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.730657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.730657Z digest=sha256:2e8bb6082a0741dbde46ac8dee85f32a00ff7a604ff68611092121681d35b3f8

Observation 49209bc5-376c-4f02-b71d-2e8b6095df20 · outbound

This paper cites Infragan: A gan ar- chitecture to transfer visible images to infrared domain.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Infragan: A gan ar- chitecture to transfer visible images to infrared domain

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.200277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.734344Z digest=sha256:8698d6247ee913ae03d8020c3ec7689bce36cdbe4b83d12633a92ad6007912b0

Observation e29cc390-dc07-4947-bc2a-01c12150291d · outbound

This paper cites Image-to-image translation: Methods and applications,.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Image-to-image translation: Methods and applications,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.187048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.737760Z digest=sha256:f2c4852da65324ab35c4269dcd290dccf1865d1e3a04c07c2b64d274338a92af

Observation 958999c3-8af4-4266-8f6b-e8567ff2033f · outbound

This paper cites Deep learning in robotics: a review of recent research.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Deep learning in robotics: a review of recent research

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.175358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.741096Z digest=sha256:44aba53a82b8de1efc97a14a929d6059c52383847534d8ae37bd6759205747db

Observation 534ce3c7-87cb-47ea-9288-7cadbcfeeae3 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.744433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.744433Z digest=sha256:877efb424638d5d26aa1ec3c5b3eca0875b8b6dfdc53335c80b909fc3d3f3b9f

Observation 93d88d10-915e-4702-9451-16f2c5263a17 · outbound

This paper cites A review on object detection in unmanned aerial vehicle surveillance.

Visual Modality Prompt for Adapting Vision-Language Object Detectors A review on object detection in unmanned aerial vehicle surveillance

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.158487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.748144Z digest=sha256:e1347c1010eb5f22978f83eeec0faf4e5f93cd3f0a4942ba92333a47a38be7da

Observation 8b679228-15af-4c9a-86fc-dd29164108a4 · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

Visual Modality Prompt for Adapting Vision-Language Object Detectors U- net: Convolutional networks for biomedical image segmen- tation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.752113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.752113Z digest=sha256:0fc536d453a9b86abd5721c20910e4711c63bd79e2a5304e8607c144e036e73f

Observation 101440e3-f1ef-45dc-8d3c-57f8eec05a7d · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Objects365: A large-scale, high-quality dataset for object detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.143049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.755827Z digest=sha256:e20bb6a46c50b6902179d59f09c078b4f60ed569c909c4f4b22c58f566d7c498

Observation eff25ee5-0fba-4e28-8ad7-bc7a706a5d94 · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Indoor segmentation and support inference from rgbd images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.759913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.759913Z digest=sha256:0084beceb05c874e7ba10cea214906af81a7ee567f63d50a88e06db11a3cf4bf

Observation 0d66c541-bc8a-4e3e-9e70-41c7fb6a9274 · outbound

This paper cites Machine learning, social learning and the gov- ernance of self-driving cars.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Machine learning, social learning and the gov- ernance of self-driving cars

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.125101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.765096Z digest=sha256:f12acb6208aca559f79d8039882bb09e485bf47172cc1ca776826d8a99bda66a

Observation e50ff276-f9ad-41c4-986e-95ebf182fdd8 · outbound

This paper cites Improving rgb-infrared object detec- tion by reducing cross-modality redundancy.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Improving rgb-infrared object detec- tion by reducing cross-modality redundancy

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.112045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.768988Z digest=sha256:8eedc67bb4d4132f46d58015255d74d2393d58d9a4674d88686356fb41af0c51

Observation 9e488f3b-8f29-40c2-be7b-39de08d93926 · outbound

This paper cites Multi-modal deep feature learning for rgb-d object detection.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Multi-modal deep feature learning for rgb-d object detection

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.099080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.773358Z digest=sha256:d4a8972d25a4dc0a8221701b4f6fa9c66f9bb9a59a082e4e7fbf0af3bf2e1476

Observation c8148bdb-46ac-44ad-b77c-cf406f83d721 · outbound

This paper cites Deepinteraction: 3d object detection via modality interaction.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Deepinteraction: 3d object detection via modality interaction

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.083543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.777499Z digest=sha256:dfeea0faadc119f2640cf8c713478eba69e0abf5c1028abeed7ad8ae53d79d87

Observation bdfe7a95-fe66-4a55-9f07-260e4b64e4ae · outbound

This paper cites Task residual for tuning vision-language models.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Task residual for tuning vision-language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.069362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.781750Z digest=sha256:0912d54f70ad0c6c4dbab2009d5fae8eb01a52747e96596c3b23ce53456c5bf4

Observation ace1f26f-f669-4de8-906c-f0861248fe94 · outbound

This paper cites Open-vocabulary detr with conditional matching.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Open-vocabulary detr with conditional matching

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.785634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.785634Z digest=sha256:7b484487cae2ff2ad0e713b78c8d606c43b7d330af1a1caa75f69acb0790747d

Observation 1bfbe532-a403-42f8-9857-4e3187f6915f · outbound

This paper cites Multispectral fusion for object detection with cyclic fuse-and-refine blocks.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Multispectral fusion for object detection with cyclic fuse-and-refine blocks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.048694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.789780Z digest=sha256:a7530f48ec3536aa65cc9f35caa215e8263bc574d3b8706f5adc6159478d23ab

Observation debd8d31-5710-40de-84a3-4d4ee334a89f · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

Visual Modality Prompt for Adapting Vision-Language Object Detectors DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.794382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.794382Z digest=sha256:308d0e50029c3c7be239027bf8b5fc812b50bece560f532e9ed67964b30e4359

Observation d175c05d-ccb5-4eba-ba8e-8d60cb88dce8 · outbound

This paper cites Glipv2: Unifying localiza- tion and vision-language understanding.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Glipv2: Unifying localiza- tion and vision-language understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.038118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.798554Z digest=sha256:8b1289ca7580c62f531798fd4ecf0ddd724e92846add190d14287c50b126dc96

Observation c0e00a36-24c9-4984-8f6b-d8316995c4b3 · outbound

This paper cites Regionclip: Region- based language-image pretraining.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Regionclip: Region- based language-image pretraining

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.026724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.802083Z digest=sha256:26c98fd02011abf213b9530fb8e7603ece0ed94bee636c20ea2779210fd510d0

Observation dbeaa7aa-83ee-4adc-854b-7f2e90212277 · outbound

This paper cites Conditional prompt learning for vision-language mod- els.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Conditional prompt learning for vision-language mod- els

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.805426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.805426Z digest=sha256:b0e055e4f3b95f0b707865135504c95e2f46fe846d67e6c089f7c7f45e3e8b64

Observation cb4cfcbc-858a-4d1b-b4ea-6f8742ee50e4 · outbound

This paper cites Learning to prompt for vision-language models.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Learning to prompt for vision-language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T05:14:07.809125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:14:07.809125Z digest=sha256:f5f5319c072cf14e9ca77003118ca9519dd0fe17c10d3a2922ac13124f9fb4cb

Observation 2e05eda2-8cc2-4b98-a09a-27ce680fde51 · outbound

This paper cites Unpaired image-to-image translation using cycle- consistent adversarial networks.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Unpaired image-to-image translation using cycle- consistent adversarial networks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:08.004209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.812687Z digest=sha256:8627cc87b40b5f09bf3bdb68e886460586252b8f46c662e674d9c47bc8310c0f

Observation 90174e29-0e3f-4a48-bad1-43c2cc7cff63 · outbound

This paper cites Object detection in 20 years: A survey.Proceed- ings of the IEEE, 111(3):257–276, 2023.

Visual Modality Prompt for Adapting Vision-Language Object Detectors Object detection in 20 years: A survey.Proceed- ings of the IEEE, 111(3):257–276, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:14:07.991671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T05:14:07.816202Z digest=sha256:f4eb8baebbf547f79ff8d65f47c667108ba9956c57a7b84e98083f488f76ccff

Pith citing papers

Observation c06edf00-497d-4209-bd28-1461fec91371 · inbound

VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors cites this paper.

VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors Visual Modality Prompt for Adapting Vision-Language Object Detectors

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:41.333932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:27:41.333932Z digest=sha256:5d45271447fe9808f0fb74c0d008eb6a0b74b8ca901b5de0a5275e68204af183