Pith. sign in

Paper Citation Record · LEDGER

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?

As of 7 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2506.23751.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23751 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:36:18.716315Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact3
  • verified fuzzy32
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5687623-310d-422f-a3ae-193d0cacf9d4 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.524118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.524118Z digest=sha256:1fb78489a630ba50f3ca470b476a4561a82b31cc81aa3f58376baeb1e1f042c4

Observation 778040ca-9e55-4ba1-906c-18f08e9b5900 · outbound

This paper cites YOLO-World: Real-Time Open-Vocabulary Object Detection.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? YOLO-World: Real-Time Open-Vocabulary Object Detection

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.528719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.528719Z digest=sha256:1c720d34ae102a126fd4d0d0debf51a42319969886b6014e228e44bd29a90ed1

Observation 54f7e7c1-d14f-4c68-978e-ffe9e38d0606 · outbound

This paper cites InProceedings of the IEEE/CVF International Conference on Computer Vision, 1780–1790 (2021).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? InProceedings of the IEEE/CVF International Conference on Computer Vision, 1780–1790 (2021)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.420862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.532495Z digest=sha256:13135c7528e1f7661459779e086aaaf0a02fb45e4b27d4dde11e5c2d19c751ce

Observation 884ade83-e5fe-4d97-b814-0a68e06cf6cb · outbound

This paper cites Glipv2: Unifying localization and vision-language understanding.Advances in Neural Information Processing Systems35, 36067–36080 (2022).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Glipv2: Unifying localization and vision-language understanding.Advances in Neural Information Processing Systems35, 36067–36080 (2022)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.411326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.536062Z digest=sha256:664622c9e23edca1d62755a59e9711765dbced77af100ea2cb6bcda95c88dfaa

Observation 2f495057-dc90-4445-bde4-174ada3abcbc · outbound

This paper cites InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 16793–16803 (2022).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 16793–16803 (2022)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.401428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.539533Z digest=sha256:ce830f9f3115c1e968024b28446ccd4940c6c6e8b7fe0862a4413681d12ccd31

Observation 2b74f78c-c0a6-446a-bf76-4cf0149304f6 · outbound

This paper cites In Proceedings of the IEEE/CVF International Conference on Computer Vision, 6339–6350 (2023).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? In Proceedings of the IEEE/CVF International Conference on Computer Vision, 6339–6350 (2023)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.390251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.542883Z digest=sha256:30d9c8f480d2eaab521f5765524bef83305fd78ef103eb8944560b1362328bb1

Observation 0a5d0124-4576-468f-9bfe-074f9e12ae86 · outbound

This paper cites A simple framework for open-vocabulary segmentation and detection.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? A simple framework for open-vocabulary segmentation and detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.379908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.546520Z digest=sha256:a0688f2ade8334563ca497d222fa0c9604c736b5c6f05e6d89ace33a585b7dbf

Observation 989687e5-18d6-43b5-bb1c-f1c222a726bd · outbound

This paper cites On the Potential of Open-Vocabulary Models for Object Detection in Unusual Street Scenes.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? On the Potential of Open-Vocabulary Models for Object Detection in Unusual Street Scenes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.549547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.549547Z digest=sha256:d7419e16bc2d42e193194d92941bda5353677e3365ac5e62b237c81e614d0d08

Observation 6f525e28-8210-4bcc-99c1-ee2782f80ad3 · outbound

This paper cites & Sünderhauf, N.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Sünderhauf, N

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.369634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.552934Z digest=sha256:f19a380e25bcdaff923f61d6399db2565bf9007ccc54c060329078ddf0854458

Observation b42c637e-997c-46b8-b5c0-aab2e4c8c257 · outbound

This paper cites & Sünderhauf, N.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Sünderhauf, N

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.360002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.556189Z digest=sha256:87b853389fc9baff69a77dfd07fbcab5ca2229ce64a0d032e9de26097079d71c

Observation 1f083ea2-83ff-47b8-ae5f-c397044fcc8c · outbound

This paper cites & Rottmann, M.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Rottmann, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.349452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.559322Z digest=sha256:bc253fbdccb3369f55081b9724726280378a2dd536110c36c58f834ed18ffc46

Observation 1fd289e5-8a96-4cf3-a955-833d92d6e5dc · outbound

This paper cites & Gottschalk, H.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Gottschalk, H

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.339661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.562442Z digest=sha256:d2ee0b5dbab43e146991f4a55bd2af2d3794025d6ffdd0ad6eb825eca8799a99

Observation 44f5711f-afd6-49ed-945c-50c29de8f574 · outbound

This paper cites VOS: Learning What You Don't Know by Virtual Outlier Synthesis.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? VOS: Learning What You Don't Know by Virtual Outlier Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.565567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.565567Z digest=sha256:98ea4d38d296acdd9e1264e0455b741ba8865f82ffd6ee12dd19eb77dab0d508

Observation 288ac21e-10a7-447f-b53f-7c4ee685b7a4 · outbound

This paper cites & Sünderhauf, N.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Sünderhauf, N

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.328919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.569187Z digest=sha256:2205a384773d6482d485958469fa1289d2a55e88815bb597fbee736bbf5b5c4d

Observation 94e6f4d1-c9c6-4a2b-8b5a-139da077874d · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.317540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.572179Z digest=sha256:905dc38278f566c8fad47be47adf4f7d2f0587e6ee3359942cda3afa767d4f62

Observation 8efabe6d-8a82-4b7d-a99e-1b943be4158f · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.575274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.575274Z digest=sha256:1f5b1333e8570fe140dfd30e31310c35e83f4dca22132c7139d650b7d65c89ec

Observation 0b821079-ab36-454a-8f8d-2f95aa026081 · outbound

This paper cites & Nguyen, K.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Nguyen, K

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.306690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.578530Z digest=sha256:ce8603c9fddada1b633602b5990a45fdda2ef5a77f728e22dbe50e67c6e9cc43

Observation 86b10c76-9deb-4891-9170-411d5abc4c92 · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.296733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.581481Z digest=sha256:31b3f26c54c014590d4873169b3f827e97f553033fd8e42c8d893ffbdbef112e

Observation ba9b86d1-36fc-4666-a3d0-e19663fa8683 · outbound

This paper cites Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.584521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.584521Z digest=sha256:412ee58c9b36c3f32171d0ab8101604d23a0e18f9807ef029153ba76aff2c55a

Observation 683c3c39-2a82-47d4-8efe-b980f5937ef1 · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.587917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.587917Z digest=sha256:3f182c12d5172f7df08b87c8c4a9cbad93191e4effdc4efad1034f5223c59966

Observation 302e95f0-0bb7-4dfd-b5f9-c9a3767853fa · outbound

This paper cites Lost and found: detecting small road hazards for self-driving vehicles.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Lost and found: detecting small road hazards for self-driving vehicles

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.286769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.591490Z digest=sha256:b4ae5ec4528f2a58f6407aed24b78dc30bb227e86c7cf9c069dce6eb6852ff25

Observation 6e6faa80-e0fc-4c6c-89a8-811a575d977a · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.276487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.594534Z digest=sha256:8e510eefee53698f4333a4f4eb626e2f7178e0389dcfcff77a9d644117703b97

Observation 27abc63b-c56c-4aef-a4a0-61e6f7532879 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.597583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.597583Z digest=sha256:1886c63e8e18e226e324f2c9f856c606c970ef2d3c13ec1355ae9256130f9a93

Observation bbc74057-4bbf-485e-a6b1-a7c87506fec1 · outbound

This paper cites & Ommer, B.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Ommer, B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.600664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.600664Z digest=sha256:d33d940fecaa2affb0317b73fd5a4022d812045edc9b15880829bc735a310231

Observation d234a649-8d1c-4b27-8615-25ab634781f9 · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.260347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.603648Z digest=sha256:6d62fcb4071b0a4acc45fc92794358bdf9abc3ffa65997de50be99663ac92019

Observation 5deb9100-bb4b-413a-a3e2-4cd695794ef0 · outbound

This paper cites & Zhang, K.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Zhang, K

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.250263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.607169Z digest=sha256:7c1d845be8a47c837310b38b2260d524634f704f4c2743755cdf1305a3f32d4d

Observation 781c0f48-99f5-4455-93c0-d0c8d1e07d93 · outbound

This paper cites Repaint: Inpainting using denoising diffusion probabilistic models.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Repaint: Inpainting using denoising diffusion probabilistic models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.240457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.610361Z digest=sha256:8121bed4ed3a0bd4ef3e062a5990623b3a004cd1a3d1c4a5fbfbe4a35f87ff43

Observation 4c45d631-d538-4be2-ae14-bf4968b30108 · outbound

This paper cites & Martinez, A.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Martinez, A

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.229403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.613417Z digest=sha256:8979199127d4fbd878bcff4d49b2d3295d06c9f197fc3967884144abcbdf1e29

Observation 01f3f2e8-63b8-4b2a-a23f-33f439620e0a · outbound

This paper cites & Yang, W.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Yang, W

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.218969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.617018Z digest=sha256:f5588b897068cdc8d3583b6c51ddf17f52a78158df2842cdbe8f8c42d3256c32

Observation 04c200de-e7f8-41cb-beef-396b2739987d · outbound

This paper cites Stable Diffusion Web UI (2022).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Stable Diffusion Web UI (2022)

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.208258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.620060Z digest=sha256:b7a84ddd44ba60d7228ffbf5ec605f551ba1a004b35e7b6e4ad4afcddb3e16d2

Observation 185ed529-6866-42b3-bdda-bc85c0e22016 · outbound

This paper cites ImagenHub: Standardizing the evaluation of conditional image generation models.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? ImagenHub: Standardizing the evaluation of conditional image generation models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.623140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.623140Z digest=sha256:08b59f8608c79e794f8ab6ccb8049046f9bcc024ac84fc071490844b69018f66

Observation 92b6bef2-ac06-4d06-b511-b67f2d2c17d7 · outbound

This paper cites Utilizing Synthetic Data for Medical Vision-Language Pre-training: Bypassing the Need for Real Images.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Utilizing Synthetic Data for Medical Vision-Language Pre-training: Bypassing the Need for Real Images

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:36:18.905441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.626438Z digest=sha256:a3da2aae9c9012019f220fbef84e825397aa247f6ca7a86552b708f4c13d28e8

Observation bcc19d0e-d167-4da7-aef3-73a3f6a5e2c1 · outbound

This paper cites & Tang, Y.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Tang, Y

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.196892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.630496Z digest=sha256:8dcc0ecb5b6bee6b6ac6ff92961f6cf449eb00cf757fa3f1a27374bed30c0a74

Observation 0dc6ffe2-90aa-4566-b9ef-b858c00170ce · outbound

This paper cites How much real data do we actually need: Analyzing object detection performance using synthetic and real data.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? How much real data do we actually need: Analyzing object detection performance using synthetic and real data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.634464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.634464Z digest=sha256:5c3b936ff12aba32e51347aa295480b4247c81384844a57d6a54100847871a92

Observation 982a6467-bfaa-4cf9-b4d7-23fa677f6da7 · outbound

This paper cites & Zhao, R.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Zhao, R

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.186832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.637951Z digest=sha256:938b522b57e69700211a4390bfb12d12d5ba88216799802d79ee8b910718b612

Observation d300ba11-ef78-48f0-b3dd-daa3c648b551 · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.176812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.641348Z digest=sha256:696f18489d9b179a7485dac760a7843b6f2b03f64d4e53396743d07515182660

Observation 734124be-46d0-48d3-94f5-86ec2c85fead · outbound

This paper cites Can OOD Object Detectors Learn from Foundation Models?.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Can OOD Object Detectors Learn from Foundation Models?

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:36:18.881224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.645069Z digest=sha256:80a3f6528f8c2a34f8d81cd0576e4fd4c4ddc186808899aad875f8aef605f6dc

Observation 07ea6929-638c-41b5-b492-83b281a620f3 · outbound

This paper cites & Cadena, C.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Cadena, C

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.166956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.648354Z digest=sha256:d728cc013bc3559fdc1a2618aa50257a98a7b98529903e7b923c49742f9b0b5e

Observation c672abfc-85dd-43e7-9327-3e0d40a4ed3e · outbound

This paper cites & Cord, M.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Cord, M

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.157509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.651534Z digest=sha256:fcd9be7fa75e7562d9f584558b98ab63a2db57cc3568abb5675fc165df248aaf

Observation 5bcb5145-926b-44e3-8f65-cca68fb82b87 · outbound

This paper cites In 2021 IEEE Intelligent Vehicles Symposium Workshops (IV Workshops), 182–189 (IEEE, 2021).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? In 2021 IEEE Intelligent Vehicles Symposium Workshops (IV Workshops), 182–189 (IEEE, 2021)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.146685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.654656Z digest=sha256:f57636bbefc0441d00ba7a249899464c1ef19043315fab01cd1d5c1fc40696eb

Observation 724ef5f0-1d63-4f6e-aacb-aa4e86596c1b · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.136254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.657892Z digest=sha256:d04e0ad7b5dd555f51136c96e6f5528186b2af844ecd0f45bb8d1e9acff36320

Observation 46945fdc-6f9d-451a-8324-e0fc534edda8 · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.126061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.661072Z digest=sha256:86836a5bd28566ebff9c19107a5b231cc8bbf3d6c585b7c45c15ecf47fc95c57

Observation 8f8f51d4-909b-4ca6-abc0-bf65973a72f6 · outbound

This paper cites InComputer vision–ECCV 2014: 13th Eu- ropean conference, zurich, Switzerland, September 6-12, 2014, proceedings, part v 13, 740–755 (Springer, 2014).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? InComputer vision–ECCV 2014: 13th Eu- ropean conference, zurich, Switzerland, September 6-12, 2014, proceedings, part v 13, 740–755 (Springer, 2014)

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.116100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.664264Z digest=sha256:584f834abfba7ba233cb4e2fac56d9004af20b47e31de608d2064b98b6dd1809

Observation e9b914be-f8ad-443a-ac52-44ab612d347c · outbound

This paper cites & Girshick, R.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Girshick, R

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.105504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.667338Z digest=sha256:843ad9b125bcb66231b9df2af183123f878eef6bac5e10d6e374d54474f39403

Observation a59087b7-e1e5-4527-b185-ffecf43fea79 · outbound

This paper cites an unresolved cited work.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:36:19.094384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.670612Z digest=sha256:b4f7a46ebd855784230056cc54bf7a85afbc6eaea9c2b7e7734c9d49b1b684bb

Observation 97847394-6117-478f-9244-6da8e53561e0 · outbound

This paper cites & Falchi, F.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Falchi, F

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.084262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.673842Z digest=sha256:63fd0ef1a381da21949ac01880a8889a0e65790a1968dd269a50c80b6043f122

Observation e848b716-1c66-485d-a552-8158fc837812 · outbound

This paper cites Dreamshaper (revision 8c1bfc6) (2023).

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Dreamshaper (revision 8c1bfc6) (2023)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.073585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.676876Z digest=sha256:a66da21dc46ca44c277967411bcd7a388e3c2b4c9408d21944f8194bed104789

Observation 7fecbe25-c803-496b-8412-92c0695ceaf0 · outbound

This paper cites Stable Diffusion Inpainting.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Stable Diffusion Inpainting

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.062164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.680204Z digest=sha256:513d283e2d243d9f9e4d4614ad97a211196a719785afe16d87636de7892abbf2

Observation 40958ad3-8aa9-49f8-8685-e8c00513fa22 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.683344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.683344Z digest=sha256:ffe2815c04e908a0860f42abcf12090f718020b9cd854325f680b1da47f9ac87

Observation e125fe07-9d49-4589-983b-bee787913275 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Swin transformer: Hierarchical vision transformer using shifted windows

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.050108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.687166Z digest=sha256:4c277827b34eba5870e3d97e5ec8855aace9f7d9b12d4b44e953c799b8545999

Observation 75881d79-0e3c-424a-a6f8-ec2c9ed44463 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.690306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.690306Z digest=sha256:3d0287e5731bf36e67f8bef173788ba35479dd541375f560f186077084919b27

Observation fac2a4bc-5f3a-4c66-a516-0089e2fbcca2 · outbound

This paper cites You only look once: Unified, real-time object detection.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? You only look once: Unified, real-time object detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.038564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.693530Z digest=sha256:8a37ed4749940cd3bab0b330b7d3d57a66e3fa82f36372d9fd1d772a078cc382

Observation d8239150-8093-4779-9ff2-a591f208e989 · outbound

This paper cites & Qiu, J.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? & Qiu, J

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.027452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.697037Z digest=sha256:1555fe1659c538c5915fc7509a29f5809020a77a61edce1f8353a738d0189e06

Observation 29ae174e-5e6d-4862-9af4-e1807b3583f6 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Learning transferable visual models from natural language supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.017053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.700069Z digest=sha256:578a3847a8e2aa93f2a4da888ed3b0f36b0634f83c8aeaf3345701886b352b62

Observation 4367e64b-8a1a-4b77-83c1-efbc75fabb28 · outbound

This paper cites End-to-end object detection with transformers.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? End-to-end object detection with transformers

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:36:19.005713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.703213Z digest=sha256:342726153830d8235acd1b51912ec645dc660ad3d9fffb9055cde8a6df07d49c

Observation d57a1505-2b3e-41cd-bb85-8d3eedc3d333 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.706367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.706367Z digest=sha256:dec810bfb6f5fa1ab6f04f4e2cd7c4a28eacfb8ece578833ec6f327ba56d54b4

Observation 26a51f06-ae8e-4e81-8948-023bd30e66ef · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.709505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.709505Z digest=sha256:7bcd695b1131769e8c8611efec5713ed10a56b27faded39394963f76c0ea23aa

Observation ab22c0ad-749c-46da-a690-68b553b9401a · outbound

This paper cites Denoising Diffusion Implicit Models.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? Denoising Diffusion Implicit Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:18.712859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:18.712859Z digest=sha256:f594a146c197ad4474b490017450161209955c65cc41f1e14f915ed9752317b5

Observation c620061b-803e-4e43-8dc3-5c3f49360751 · outbound

This paper cites object in the street.

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes? object in the street

Reference 59

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:36:18.817206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T21:36:18.716315Z digest=sha256:e0bbb5cb86f5315f163580e3cb514dc09b7b2db667f25183723b33374d388eb6

Pith citing papers

No inbound Pith citation observations are available.