Pith. sign in

Paper Citation Record · LEDGER

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?

As of 7 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2506.23751.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23751 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:36:18.716315Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact3
  • verified fuzzy32
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5687623-310d-422f-a3ae-193d0cacf9d4 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.524118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.524118Z digest=sha256:1fb78489a630ba50f3ca470b476a4561a82b31cc81aa3f58376baeb1e1f042c4

Observation 778040ca-9e55-4ba1-906c-18f08e9b5900 · outbound

This paper cites YOLO-World: Real-Time Open-Vocabulary Object Detection.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? YOLO-World: Real-Time Open-Vocabulary Object Detection

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.528719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.528719Z digest=sha256:1c720d34ae102a126fd4d0d0debf51a42319969886b6014e228e44bd29a90ed1

Observation 54f7e7c1-d14f-4c68-978e-ffe9e38d0606 · outbound

This paper cites InProceedings of the IEEE/CVF International Conference on Computer Vision, 1780–1790 (2021).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? InProceedings of the IEEE/CVF International Conference on Computer Vision, 1780–1790 (2021)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.420862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.532495Z digest=sha256:b4f6e74171315b5b5a5999cc760e28d6afc112a61ffaace9fadc9dbdeee42aeb

Observation 884ade83-e5fe-4d97-b814-0a68e06cf6cb · outbound

This paper cites Glipv2: Unifying localization and vision-language understanding.Advances in Neural Information Processing Systems35, 36067–36080 (2022).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Glipv2: Unifying localization and vision-language understanding.Advances in Neural Information Processing Systems35, 36067–36080 (2022)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.411326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.536062Z digest=sha256:a0afb6932fa80ff513cfca7c8d0cbae5b3a752779a7c3173b1db2681354bc13a

Observation 2f495057-dc90-4445-bde4-174ada3abcbc · outbound

This paper cites InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 16793–16803 (2022).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 16793–16803 (2022)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.401428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.539533Z digest=sha256:d3344d868ac1d42509853b22b769b5df54034c7ef47f903f70bb3d82c5ea1b86

Observation 2b74f78c-c0a6-446a-bf76-4cf0149304f6 · outbound

This paper cites In Proceedings of the IEEE/CVF International Conference on Computer Vision, 6339–6350 (2023).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? In Proceedings of the IEEE/CVF International Conference on Computer Vision, 6339–6350 (2023)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.390251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.542883Z digest=sha256:674cb691470bc19cda7024123c324e3293b3aa386a9be9d4e46ea37ec209b7e5

Observation 0a5d0124-4576-468f-9bfe-074f9e12ae86 · outbound

This paper cites A simple framework for open-vocabulary segmentation and detection.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? A simple framework for open-vocabulary segmentation and detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.379908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.546520Z digest=sha256:7a20ce7d78b0d3c281a501c24d11ac9ed01d3f1d46a0c14a47dfd80a7841bf80

Observation 989687e5-18d6-43b5-bb1c-f1c222a726bd · outbound

This paper cites On the Potential of Open-Vocabulary Models for Object Detection in Unusual Street Scenes.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? On the Potential of Open-Vocabulary Models for Object Detection in Unusual Street Scenes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.549547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.549547Z digest=sha256:d7419e16bc2d42e193194d92941bda5353677e3365ac5e62b237c81e614d0d08

Observation 6f525e28-8210-4bcc-99c1-ee2782f80ad3 · outbound

This paper cites & Sünderhauf, N.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Sünderhauf, N

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.369634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.552934Z digest=sha256:c2a0875ba39f54f623330045b45859098df22e625b24b7aff9880ad065875f23

Observation b42c637e-997c-46b8-b5c0-aab2e4c8c257 · outbound

This paper cites & Sünderhauf, N.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Sünderhauf, N

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.360002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.556189Z digest=sha256:1ffa74ba17cae00f90d0fb9f6372ed67b1e5bdbaf915ef6e79bc6c02f0bb5e8f

Observation 1f083ea2-83ff-47b8-ae5f-c397044fcc8c · outbound

This paper cites & Rottmann, M.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Rottmann, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.349452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.559322Z digest=sha256:99c3ec70d8eac1ffe14633cccc92d4265f853209d587a092175765ab706a1a48

Observation 1fd289e5-8a96-4cf3-a955-833d92d6e5dc · outbound

This paper cites & Gottschalk, H.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Gottschalk, H

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.339661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.562442Z digest=sha256:efae54b79a60f045ac81c14402074e9932a7416d65594b420f66d421b71f7ff6

Observation 44f5711f-afd6-49ed-945c-50c29de8f574 · outbound

This paper cites VOS: Learning What You Don't Know by Virtual Outlier Synthesis.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? VOS: Learning What You Don't Know by Virtual Outlier Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.565567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.565567Z digest=sha256:98ea4d38d296acdd9e1264e0455b741ba8865f82ffd6ee12dd19eb77dab0d508

Observation 288ac21e-10a7-447f-b53f-7c4ee685b7a4 · outbound

This paper cites & Sünderhauf, N.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Sünderhauf, N

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.328919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.569187Z digest=sha256:1e4b7c44dee5f19e26bc16d5b575a7ca9bb1a117520a90878402fa3fd604184b

Observation 94e6f4d1-c9c6-4a2b-8b5a-139da077874d · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.317540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.572179Z digest=sha256:ac10983e32a217db7fe523793adfb3ab8c5c057291cce1ab70def64e4069679b

Observation 8efabe6d-8a82-4b7d-a99e-1b943be4158f · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.575274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.575274Z digest=sha256:1f5b1333e8570fe140dfd30e31310c35e83f4dca22132c7139d650b7d65c89ec

Observation 0b821079-ab36-454a-8f8d-2f95aa026081 · outbound

This paper cites & Nguyen, K.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Nguyen, K

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.306690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.578530Z digest=sha256:65b562b025ade50622a00836238bd938d878c8f2ccdd49d37ffbb38dae219beb

Observation 86b10c76-9deb-4891-9170-411d5abc4c92 · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.296733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.581481Z digest=sha256:2d8a636c4e4759cfc6ac42e2588ab8b013c1fe8f7ef8b27d592c1443f7efdf7e

Observation ba9b86d1-36fc-4666-a3d0-e19663fa8683 · outbound

This paper cites Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.584521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.584521Z digest=sha256:412ee58c9b36c3f32171d0ab8101604d23a0e18f9807ef029153ba76aff2c55a

Observation 683c3c39-2a82-47d4-8efe-b980f5937ef1 · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.587917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.587917Z digest=sha256:3f182c12d5172f7df08b87c8c4a9cbad93191e4effdc4efad1034f5223c59966

Observation 302e95f0-0bb7-4dfd-b5f9-c9a3767853fa · outbound

This paper cites Lost and found: detecting small road hazards for self-driving vehicles.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Lost and found: detecting small road hazards for self-driving vehicles

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.286769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.591490Z digest=sha256:c1ee2528630b902c870ba983dffca4bd79f80015993f5ebfa5ca84f3ce80ead7

Observation 6e6faa80-e0fc-4c6c-89a8-811a575d977a · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.276487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.594534Z digest=sha256:83b75ae4574050c8576d3efb795407d7d20ab0a81e9605af3df06dc3c3346664

Observation 27abc63b-c56c-4aef-a4a0-61e6f7532879 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.597583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.597583Z digest=sha256:1886c63e8e18e226e324f2c9f856c606c970ef2d3c13ec1355ae9256130f9a93

Observation bbc74057-4bbf-485e-a6b1-a7c87506fec1 · outbound

This paper cites & Ommer, B.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Ommer, B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.600664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.600664Z digest=sha256:d33d940fecaa2affb0317b73fd5a4022d812045edc9b15880829bc735a310231

Observation d234a649-8d1c-4b27-8615-25ab634781f9 · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.260347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.603648Z digest=sha256:2c41f413ff2a37c5f58125d30ba5ed8b8034172efba9b72990bb82615053f104

Observation 5deb9100-bb4b-413a-a3e2-4cd695794ef0 · outbound

This paper cites & Zhang, K.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Zhang, K

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.250263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.607169Z digest=sha256:d217e827dec98c01a1ece9b3051e3c1e4f775839c882ec5ac64d8a0574dbf6dd

Observation 781c0f48-99f5-4455-93c0-d0c8d1e07d93 · outbound

This paper cites Repaint: Inpainting using denoising diffusion probabilistic models.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Repaint: Inpainting using denoising diffusion probabilistic models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.240457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.610361Z digest=sha256:9d8df5e1a5813f41416cc173d67e50c479240e410a4e8070eb817b61615bddab

Observation 4c45d631-d538-4be2-ae14-bf4968b30108 · outbound

This paper cites & Martinez, A.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Martinez, A

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.229403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.613417Z digest=sha256:fa04ecf4209970ba08d76db22a0378969d40dc56345751adf27ffcd995f2ed60

Observation 01f3f2e8-63b8-4b2a-a23f-33f439620e0a · outbound

This paper cites & Yang, W.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Yang, W

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.218969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.617018Z digest=sha256:0fa227e3c819e99d0c6f87a657447727608c47f0349feb162b75c12e59df8282

Observation 04c200de-e7f8-41cb-beef-396b2739987d · outbound

This paper cites Stable Diffusion Web UI (2022).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Stable Diffusion Web UI (2022)

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.208258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.620060Z digest=sha256:f1b2935ea5d9e88a1e94d8f5a03c8059750d7a20d5341ebf9d37a1d911dd50ec

Observation 185ed529-6866-42b3-bdda-bc85c0e22016 · outbound

This paper cites ImagenHub: Standardizing the evaluation of conditional image generation models.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? ImagenHub: Standardizing the evaluation of conditional image generation models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.623140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.623140Z digest=sha256:08b59f8608c79e794f8ab6ccb8049046f9bcc024ac84fc071490844b69018f66

Observation 92b6bef2-ac06-4d06-b511-b67f2d2c17d7 · outbound

This paper cites Utilizing Synthetic Data for Medical Vision-Language Pre-training: Bypassing the Need for Real Images.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Utilizing Synthetic Data for Medical Vision-Language Pre-training: Bypassing the Need for Real Images

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:36:18.905441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.626438Z digest=sha256:4a183d8a49af730b149aa3ee61d784c4a688cc71f43407ec639387e8ab026c15

Observation bcc19d0e-d167-4da7-aef3-73a3f6a5e2c1 · outbound

This paper cites & Tang, Y.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Tang, Y

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.196892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.630496Z digest=sha256:c191384abc00dbd6c58d9258f57f9afb1342ce9d59c2271a4e07a874450c966b

Observation 0dc6ffe2-90aa-4566-b9ef-b858c00170ce · outbound

This paper cites How much real data do we actually need: Analyzing object detection performance using synthetic and real data.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? How much real data do we actually need: Analyzing object detection performance using synthetic and real data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.634464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.634464Z digest=sha256:5c3b936ff12aba32e51347aa295480b4247c81384844a57d6a54100847871a92

Observation 982a6467-bfaa-4cf9-b4d7-23fa677f6da7 · outbound

This paper cites & Zhao, R.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Zhao, R

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.186832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.637951Z digest=sha256:0802910f1cf8777980eb5a89da4f159aadce3569594ef4780dee328afeca1b16

Observation d300ba11-ef78-48f0-b3dd-daa3c648b551 · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.176812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.641348Z digest=sha256:c2d4e270171606e16523aef95e5f2599647c068929d35f52fd29c9e16b7d441a

Observation 734124be-46d0-48d3-94f5-86ec2c85fead · outbound

This paper cites Can OOD Object Detectors Learn from Foundation Models?.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Can OOD Object Detectors Learn from Foundation Models?

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:36:18.881224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.645069Z digest=sha256:286b79608b32dcadac65fab2566831127946023a6cb4adff90eac7136dfcce89

Observation 07ea6929-638c-41b5-b492-83b281a620f3 · outbound

This paper cites & Cadena, C.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Cadena, C

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.166956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.648354Z digest=sha256:166b3cb0fe4f934d7f9188f6eecdd57ad111d18b37d955606d2aee33b51e6f61

Observation c672abfc-85dd-43e7-9327-3e0d40a4ed3e · outbound

This paper cites & Cord, M.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Cord, M

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.157509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.651534Z digest=sha256:27c63577b0661a07cf9bff29b11917c55c85c2d2ed6085ae12d1ca8f600b11fe

Observation 5bcb5145-926b-44e3-8f65-cca68fb82b87 · outbound

This paper cites In 2021 IEEE Intelligent Vehicles Symposium Workshops (IV Workshops), 182–189 (IEEE, 2021).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? In 2021 IEEE Intelligent Vehicles Symposium Workshops (IV Workshops), 182–189 (IEEE, 2021)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.146685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.654656Z digest=sha256:6f681feecb2585fb8dc6dfca0cf37a6133e88b82c21993d3a3a5f6b4f14e62f4

Observation 724ef5f0-1d63-4f6e-aacb-aa4e86596c1b · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.136254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.657892Z digest=sha256:77032951c3ff42a02655f7c1ae34e12d22028c839146ea0d1eadf928ad08020f

Observation 46945fdc-6f9d-451a-8324-e0fc534edda8 · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.126061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.661072Z digest=sha256:d9a779b2762e62e4dabd250b9289f951ccc18cc8ad8b6eb17b7a36f0c86b9eb6

Observation 8f8f51d4-909b-4ca6-abc0-bf65973a72f6 · outbound

This paper cites InComputer vision–ECCV 2014: 13th Eu- ropean conference, zurich, Switzerland, September 6-12, 2014, proceedings, part v 13, 740–755 (Springer, 2014).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? InComputer vision–ECCV 2014: 13th Eu- ropean conference, zurich, Switzerland, September 6-12, 2014, proceedings, part v 13, 740–755 (Springer, 2014)

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.116100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.664264Z digest=sha256:9a156027ff455f9c439e4e051c9a911dc8b86e34ef8a7790e4d74c3311eb1646

Observation e9b914be-f8ad-443a-ac52-44ab612d347c · outbound

This paper cites & Girshick, R.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Girshick, R

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.105504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.667338Z digest=sha256:b79f898b0da7d859c2adf004964c1b8098bf006b6a90bf483e72bde553d8b503

Observation a59087b7-e1e5-4527-b185-ffecf43fea79 · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.094384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.670612Z digest=sha256:fe5cd4127cb99fb1ff40bf8dc6f27d0bc84f9ee2c425362b10114fceda3b04d7

Observation 97847394-6117-478f-9244-6da8e53561e0 · outbound

This paper cites & Falchi, F.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Falchi, F

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.084262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.673842Z digest=sha256:19eec2baf79540a7b4dc422a6a2f236aad14a191f7ed6878b6d114a23844ea59

Observation e848b716-1c66-485d-a552-8158fc837812 · outbound

This paper cites Dreamshaper (revision 8c1bfc6) (2023).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Dreamshaper (revision 8c1bfc6) (2023)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.073585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.676876Z digest=sha256:d89ef495a81d67839f31ba34fd5658abad50b3ad6daff2403c728cc67a654971

Observation 7fecbe25-c803-496b-8412-92c0695ceaf0 · outbound

This paper cites Stable Diffusion Inpainting.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Stable Diffusion Inpainting

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.062164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.680204Z digest=sha256:1b9ccf1e69f540a92c43eaf1225018f50507a0b16c31ecb4792a8d726ab28458

Observation 40958ad3-8aa9-49f8-8685-e8c00513fa22 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.683344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.683344Z digest=sha256:ffe2815c04e908a0860f42abcf12090f718020b9cd854325f680b1da47f9ac87

Observation e125fe07-9d49-4589-983b-bee787913275 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Swin transformer: Hierarchical vision transformer using shifted windows

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.050108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.687166Z digest=sha256:04d384b5a0f3a80ffdd652c326b38fe4401434712ab3a907b1fe9c7dbfe31154

Observation 75881d79-0e3c-424a-a6f8-ec2c9ed44463 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.690306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.690306Z digest=sha256:3d0287e5731bf36e67f8bef173788ba35479dd541375f560f186077084919b27

Observation fac2a4bc-5f3a-4c66-a516-0089e2fbcca2 · outbound

This paper cites You only look once: Unified, real-time object detection.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? You only look once: Unified, real-time object detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.038564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.693530Z digest=sha256:298e41d0bafd9552b7c0227f3656507314633daa50c158fef95c819283d74f54

Observation d8239150-8093-4779-9ff2-a591f208e989 · outbound

This paper cites & Qiu, J.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Qiu, J

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.027452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.697037Z digest=sha256:2d852f6d7706ade812f4cb3484582a4b544d0b7efe39c5ef8b1c10739e356d6b

Observation 29ae174e-5e6d-4862-9af4-e1807b3583f6 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Learning transferable visual models from natural language supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.017053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.700069Z digest=sha256:92963cf65cc833a4ebfd88f8ee5377789550abebb78de07064288da1bd0282ab

Observation 4367e64b-8a1a-4b77-83c1-efbc75fabb28 · outbound

This paper cites End-to-end object detection with transformers.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? End-to-end object detection with transformers

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.005713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.703213Z digest=sha256:907cf3835d6fb0d09eef8a506832f708eb40700ce534b6db379ef0db698cad5b

Observation d57a1505-2b3e-41cd-bb85-8d3eedc3d333 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.706367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.706367Z digest=sha256:dec810bfb6f5fa1ab6f04f4e2cd7c4a28eacfb8ece578833ec6f327ba56d54b4

Observation 26a51f06-ae8e-4e81-8948-023bd30e66ef · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.709505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.709505Z digest=sha256:7bcd695b1131769e8c8611efec5713ed10a56b27faded39394963f76c0ea23aa

Observation ab22c0ad-749c-46da-a690-68b553b9401a · outbound

This paper cites Denoising Diffusion Implicit Models.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Denoising Diffusion Implicit Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.712859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.712859Z digest=sha256:f594a146c197ad4474b490017450161209955c65cc41f1e14f915ed9752317b5

Observation c620061b-803e-4e43-8dc3-5c3f49360751 · outbound

This paper cites object in the street.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? object in the street

Reference 59

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:36:18.817206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:36:18.716315Z digest=sha256:eac9a64182ed16067aae1fe1d1643f6e9804324ce7adf7bcadfde74c603493bc

Pith citing papers

No inbound Pith citation observations are available.