Pith. sign in

Paper Citation Record · LEDGER

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

As of 22 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 7 inbound Pith citation observations for arXiv:2501.18954.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18954 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T21:55:23.361176Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:17:36.580979Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:46:33.282729Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1573677-fb1b-4b81-b5fe-61fdc1448642 · outbound

This paper cites Nltk: the natural language toolkit.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Nltk: the natural language toolkit

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.933368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.151880Z digest=sha256:d87e1a41cb3d8805ccf8e47511749b5ed854352f4e7e62c2980a7d86f0ffb302

Observation 47ec76e3-c514-4b0b-b9ee-a76b070087c2 · outbound

This paper cites End-to- end object detection with transformers.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models End-to- end object detection with transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.922207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.158157Z digest=sha256:4ec123c6c6fd46535cac90ce2ec5a5b9e7fb6f7a9ad8fb52b2d12de6788b2008

Observation ffaa8454-edde-48f0-8aeb-cbebda13f315 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.161762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.161762Z digest=sha256:e953177a0faf103b3d12a6701528ba04d885ee0ba83e93d5868dd982bb31c681

Observation f7a62b63-cfb7-4405-af9c-40b0302fd257 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Sharegpt4v: Improving large multi-modal models with better captions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.912272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.168129Z digest=sha256:d7a7aca5f0bee159157bfa006a6682cb55d307f8055cfd382fc91d30d00857b4

Observation 84399f18-64fb-43a7-ab53-df1d0b9172fb · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.171583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.171583Z digest=sha256:adec414fb06859be32bd12e64e19dd711b7bf0363f7d5eaf58560fa63d7d8786

Observation 68873fa9-8add-4bbe-86b5-b98e519d44b3 · outbound

This paper cites Yolo-world: Real-time open- vocabulary object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Yolo-world: Real-time open- vocabulary object detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.898006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.175076Z digest=sha256:9c438a6108685d3b191fa5952a0b62e97faeb858358608281af542f2706d872c

Observation 6e36856d-eaae-4521-94f2-71467bd0f696 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Gonzalez, Ion Stoica, and Eric P

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.178553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.178553Z digest=sha256:eb914402c43797b0e4e4b2db04fd2161b2cf129c8586c118b6ea39b94ecfb566

Observation 901397cd-38b0-4712-8301-fbbd736bb35e · outbound

This paper cites Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.181974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.181974Z digest=sha256:499a706c34921ba290e392b43738fc7b20558537f2fda864d5a75a8ea917a334

Observation 44edde5d-1910-4bd9-818a-bb26ab22c5b4 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Imagenet: A large-scale hierarchical image database

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.185433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.185433Z digest=sha256:4639d4702813fa7ceed5fe3eb89617474a42453d59b2edc90e680ce8efe157d8

Observation e407b773-545a-425d-9a2a-bfe6decd759b · outbound

This paper cites Coarse-to-fine vision-language pre-training with fusion in the backbone.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Coarse-to-fine vision-language pre-training with fusion in the backbone

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.877896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.188567Z digest=sha256:d7588719a3b52014306801ab6698fc269e086903afd816bc891ee743be11e374

Observation 175448ee-958f-4fe1-a951-929a89ed7b21 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.191659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.191659Z digest=sha256:5ae74310276e18cdec463ba65d310e04ebc6002b061df543732a63d9fd4f0107

Observation da140a42-570e-4172-b2cc-e20ac0e05163 · outbound

This paper cites Asag: Building strong one-decoder-layer sparse detectors via adaptive sparse anchor generation.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Asag: Building strong one-decoder-layer sparse detectors via adaptive sparse anchor generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.869046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.195329Z digest=sha256:d9ac765a8c583bc64e2266439d0a704ac221bcca0608a71cff7875e6bec9eb5d

Observation d8d29c7b-f8bd-4807-af7b-1f1079d04e04 · outbound

This paper cites Frozen-detr: Enhancing detr with image understanding from frozen foundation models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Frozen-detr: Enhancing detr with image understanding from frozen foundation models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.198600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.198600Z digest=sha256:7401e3394f51424b8c96fce4f93fdd3d10bbf9875cc66faf207a89356835a70f

Observation e2b73a3e-4a5c-4cb3-9251-8fc272f08d66 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Open-vocabulary object detection via vision and language knowledge distillation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.855042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.201596Z digest=sha256:ec7c6357cae02f2f6321e91ace8eab49e969dff3413f64fd64717315d36bc4b9

Observation 71a56cef-51ce-4b25-8fbd-b500139973c8 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Lvis: A dataset for large vocabulary instance segmentation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.204637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.204637Z digest=sha256:6871a6dec4f5498d7e8f94d78db013b859ce6f7c3ab691936e56e01a6e34c59e

Observation d42e2777-51d9-42b8-b4ad-eb92e7620146 · outbound

This paper cites Mask r-cnn.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Mask r-cnn

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.840719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.207869Z digest=sha256:4bf1bd2f2e94a18c160aabda7de645d51d76caebe501e2f8333aba105542b4f5

Observation 0b63dcd3-1a56-409b-b130-18fe2204f11c · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.830403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.210819Z digest=sha256:7baaaf2bba4c9b3c2cce0663185286673b0cc488b9cf89abb435abd4156e4e11

Observation d75afc4b-f51c-4786-af1e-61c751d28bff · outbound

This paper cites GPT-4o System Card.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.213818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.213818Z digest=sha256:e7f2dc7f4e7221080dfedaa57a1471795677e06de43651451d868beeda24b4f9

Observation 928aa28c-d9e7-4b7e-9bd6-83bac3d42c6b · outbound

This paper cites T-rex2: Towards generic object detec- tion via text-visual prompt synergy.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models T-rex2: Towards generic object detec- tion via text-visual prompt synergy

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.820218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.217121Z digest=sha256:2b01c28d807f88a6ee3fa7a9bbdabeb05733c6d10b4d74a169816b4df5becd44

Observation f7fff0fe-110f-4403-980e-52c9497ababb · outbound

This paper cites Mdetr- modulated detection for end-to-end multi-modal understand- ing.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Mdetr- modulated detection for end-to-end multi-modal understand- ing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.809160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.220120Z digest=sha256:87e5b2e0347e8118a8f78f540378f0d72d5fa9bf547b1343ca26fda655e8077d

Observation 22f02348-d893-4d0b-9c76-0394a80040a3 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.798966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.222588Z digest=sha256:1ac5a750ac182eaca3424d76915bc2ff7cef18cb8fb620730940f2f9a611f8a7

Observation 36ab6bc2-80a5-4269-82c7-8b948a420e03 · outbound

This paper cites F-vlm: Open-vocabulary object detection upon frozen vision and language models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models F-vlm: Open-vocabulary object detection upon frozen vision and language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.225150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.225150Z digest=sha256:cd73f4239e73f9ca572fd723e6641a526095252c0caed94fea5cc739c322218d

Observation 76bde895-cbad-43ba-996f-a18531bfe81c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.227783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.227783Z digest=sha256:2d0ef69b549f2e515c5f8460de0478daf2467fc541fd1e31dd0a267ae9f1c13c

Observation 1178452f-d5c1-46a8-af4b-0188eb0c6efb · outbound

This paper cites Elevater: A benchmark and toolkit for evaluating language-augmented visual models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Elevater: A benchmark and toolkit for evaluating language-augmented visual models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.783598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.230754Z digest=sha256:9a6ea990495bdb2bf7e113dde1bde9900e083bb14efb4e85859515164cd48a5b

Observation b405fdf2-d4e0-4aac-adfd-102410a18cda · outbound

This paper cites Desco: Learning object recognition with rich language de- scriptions.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Desco: Learning object recognition with rich language de- scriptions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.774357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.233206Z digest=sha256:8213a2c7786b0e8a945b6699d140858c87c8011c5e12df5c98e06d6a437a1f07

Observation 0778c048-4d9f-4952-b70b-342ad58ac51b · outbound

This paper cites Distilling detr with visual-linguistic knowledge for open-vocabulary object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Distilling detr with visual-linguistic knowledge for open-vocabulary object detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.235667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.235667Z digest=sha256:934ecb1e46d1e143700d2a041ff6ef164c35e07597e37ad94234ff3a50fd27d6

Observation 8d840ada-62b2-4d70-bf8e-129906e71278 · outbound

This paper cites Grounded language-image pre-training.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Grounded language-image pre-training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.760243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.238489Z digest=sha256:59b15fc84a0430bf787409da1babf8d834715734964ff72362f2f2484366a699

Observation e0d13269-2714-4591-9ff8-b2fe067bdcd5 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Evaluating object hallucination in large vision-language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.751470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.241190Z digest=sha256:98c44fc719b0cf4865fa8b368b64f5293aab800633cd0767f0d74a91aa58e6f3

Observation 9323eaa6-52d3-4fb5-a424-9c20843a5806 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.742548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.244124Z digest=sha256:ee3744c0191c187ad51a534cec8fd6783c4c9417be3f62e6b45b0cac0fc81315

Observation b73b9807-e4e7-4d0d-afe8-9b52010fdc72 · outbound

This paper cites Generative region-language pretraining for open-ended object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Generative region-language pretraining for open-ended object detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.733418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.247039Z digest=sha256:53f5513c09d524aa0ef9c2151e5c7594acf4b529b9cc653261d3194c0de7f1d5

Observation 6fe0c555-14ec-4412-900b-f6aebbac2119 · outbound

This paper cites Microsoft coco: Common objects in context.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Microsoft coco: Common objects in context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.724447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.250070Z digest=sha256:a7442feae9763b9e4288329219e8e8628d7d344ace439f3800cfb1f6e93b9076

Observation 525b4851-844e-4af1-ab32-5fd0ae659897 · outbound

This paper cites Focal loss for dense object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Focal loss for dense object detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.715499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.253093Z digest=sha256:7cc51feacc8802ae4eb335f73ae7355e6e0e9c81b79b13dd32f027faffacd653

Observation 40bb858d-7e45-4f0d-9c33-322bdeddbc64 · outbound

This paper cites Gres: Gener- alized referring expression segmentation.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Gres: Gener- alized referring expression segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.706794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.256227Z digest=sha256:3a1ef602a8ffdd23defa2811f0ce0caf7807e54f726d9535ec4f6e7536e585ad

Observation 263e4df8-4a1a-49c4-8417-fdfb524b940f · outbound

This paper cites Visual instruction tuning.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Visual instruction tuning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.698025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.259327Z digest=sha256:437c1e89bdcfc8afce5f6bfc547cc44fa089f9601cd90ee857e42ad0dc2365f7

Observation c0044068-c4d7-4333-92a1-f099cce853d3 · outbound

This paper cites Improved baselines with visual instruction tuning.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Improved baselines with visual instruction tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.262296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.262296Z digest=sha256:d1ff07f9c34893230ffd39a2dcf453ac05c0134d81aedeaeff2dacc66a477b4a

Observation 2aab2b4b-8dc7-4437-93e1-880f1eea506a · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.684444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.265451Z digest=sha256:a0805f48850b7e0af1ba448571bf41fa9600619399775984b29600b06741c1fe

Observation d17e59ce-34b0-4672-953d-50571de4e5a7 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Swin transformer: Hierarchical vision transformer using shifted windows

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.675415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.268477Z digest=sha256:4d16284bb9c17963a71279c2d7f8d0eb72b41b5b825db99e794f65abf7fb2dbf

Observation 8682bee4-9018-49de-9b74-103dff20e1c1 · outbound

This paper cites Capdet: Unifying dense captioning and open-world detec- tion pretraining.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Capdet: Unifying dense captioning and open-world detec- tion pretraining

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.666285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.271366Z digest=sha256:7d3c06850dbf0bb185b52483da44398ef1bcb028adbf4b8dbaeaf2432bce8a61

Observation da99e344-8dd1-4402-94a4-0fa5a80f38b3 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Generation and comprehension of unambiguous object descriptions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.657526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.274310Z digest=sha256:ae8ebbd3c5b80ce6e64b3379a7ce2ab4d114ceaaf575119b824db4d867a92bf6

Observation 1fd409d3-7413-4b59-8c7d-bf22e5ec03c1 · outbound

This paper cites Coco-o: A benchmark for object detectors under natural distribution shifts.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Coco-o: A benchmark for object detectors under natural distribution shifts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.648586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.277500Z digest=sha256:0922356e6bc0be0a85d41baeb3835ca44fb6cb8a31e046aa0b6a39d5d6223d30

Observation 3bae787a-db2c-4633-98d6-3d71437859ce · outbound

This paper cites Scaling open-vocabulary object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Scaling open-vocabulary object detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.639857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.280494Z digest=sha256:951e36b3e2d3906d0756e2a9c771fffd838d320b177e06af40825e193257446f

Observation 5b43bc4b-c505-4adf-913a-2cc5e42c6bd7 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.630943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.283447Z digest=sha256:0aa6781067ab5ca7e500eb1b85f4985f906e5a032302defb5f1ed99a2d9d0d2d

Observation 9b1f626b-f0e1-49f3-8c6b-5cdc6950d1e0 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Learn- ing transferable visual models from natural language super- vision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.286704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.286704Z digest=sha256:a3aa017a838cced0d129eb2ed0c36d5789f0b8f5363088a88bff10f26b611987

Observation c459f17a-db0b-4374-b117-cae7cdcc54d4 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Objects365: A large-scale, high-quality dataset for object detection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.289705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.289705Z digest=sha256:1d30eaef8a80245bc4009c7cdbb046bb9c87e8cf35ceeacd0d0be45a20971d33

Observation a837e111-9987-4453-b261-63cc72631c9b · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.292515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.292515Z digest=sha256:161f0ed1477464c72e44a76f0cc94b4705cd19a00a5038dd33e3125c5298adfa

Observation 685e67aa-bf45-4915-bfc3-5aa41164105c · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.295553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.295553Z digest=sha256:195db5703b138eddd695dde04c7936c3e3173bb72d58aa74ae1b6e6193699504

Observation fab3f358-ac94-47e6-90ac-0b5a9222a54b · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.298981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.298981Z digest=sha256:83f2b412524c4b5ca1327083e0993dc43307b27ce0ce1404a0fb38352bc81a55

Observation 9c2b995e-ef0a-4846-ad97-705907b0062e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.302192Z digest=sha256:11026143edc143cebbafa1a26c34a755f516c982e4abc73c997018c133c5d170

Observation 6ceadd55-2371-4dba-803b-7276702c7e60 · outbound

This paper cites OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.305127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.305127Z digest=sha256:b6ffc780ca446304eb30bf59d0b1a939a44da13067b25e2f3129aa2b10a1440d

Observation 3fe787cf-ccf6-4697-aa3f-ba1ea02fbf10 · outbound

This paper cites V3det: Vast vocabulary visual detection dataset.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models V3det: Vast vocabulary visual detection dataset

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.607473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.308354Z digest=sha256:1708ebde9fba73bb685853cf0d549042010494428e1ce9a0fa00ffa0c37bdf92

Observation 58a6c0de-f752-4992-b1d8-a769f3b7ce75 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.311214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.311214Z digest=sha256:ab89e4b3943f64f3090c588eaf7f4e2f448406c681c01d3a90d97eaa2f76bf1a

Observation 8635b23c-e02a-49cd-8d03-bcc3efcfcb22 · outbound

This paper cites The all-seeing project v2: To- wards general relation comprehension of the open world.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models The all-seeing project v2: To- wards general relation comprehension of the open world

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.598662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.314542Z digest=sha256:0c3324ae516578f60b7be66a138b20c2ae9b6d76c0bf4e4ef4b0ddc689fe9cfe

Observation 30fe3dfc-7868-4563-b9ae-62493519089c · outbound

This paper cites Grit: A gen- erative region-to-text transformer for object understanding.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Grit: A gen- erative region-to-text transformer for object understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.589693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.317329Z digest=sha256:daa454ba7c4f8707ac572f70c0fca4d48dca1a1db4a154145de809917244422e

Observation 379c936c-c3ec-4e99-a626-0661e3828ddd · outbound

This paper cites Aligning bag of regions for open- vocabulary object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Aligning bag of regions for open- vocabulary object detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.321030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.321030Z digest=sha256:75db1ac9dfee4c22367f53f7de07ed3572e90050858e2358dcc7401f01107c79

Observation 49e41f53-ee25-4ceb-af31-9f6821eaa9fe · outbound

This paper cites Qwen2 Technical Report.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Qwen2 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.323886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.323886Z digest=sha256:08291c19881c85888d04f642e02774eaf6d214173901c31f2c02d01b1ba204f2

Observation e5bba39b-7a2a-4bf9-b6f4-4fd26499815e · outbound

This paper cites Detclip: Dictionary-enriched visual-concept paralleled pre- training for open-world detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Detclip: Dictionary-enriched visual-concept paralleled pre- training for open-world detection

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.576360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.326997Z digest=sha256:082cc394321689a8127cabc9f28b65efa4870da70d16db697bd518db7cff3212

Observation a6401795-8ca0-42d3-97ec-24b156b955eb · outbound

This paper cites Detclipv2: Scal- able open-vocabulary object detection pre-training via word- region alignment.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Detclipv2: Scal- able open-vocabulary object detection pre-training via word- region alignment

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.567971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.329849Z digest=sha256:68e8cbab8622ec89c46ba4d5f647427ab77cbf8e385bd353b037650490e7436b

Observation c740e57d-d35b-4a47-b774-49fe0eb0225c · outbound

This paper cites Detclipv3: To- wards versatile generative open-vocabulary object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Detclipv3: To- wards versatile generative open-vocabulary object detection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.559626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.332659Z digest=sha256:e02cd76674caf00dff7568a5b163769409494b7722a21af3e0e29c756b490331

Observation c8c4055f-d33d-4e3c-abf8-528817b67173 · outbound

This paper cites Modeling context in referring expres- sions.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Modeling context in referring expres- sions

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.551103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.335463Z digest=sha256:bb10cf00706e6ef8eb6cf67add67d5a229abae627cb69145c61961170abe3cd4

Observation 186e798d-e220-45a8-9f39-441239d814b0 · outbound

This paper cites Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.542460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.338352Z digest=sha256:7664d182478e5c95a013fe8340c5fb72a254c3177cf4de44bd4f71cbe1a72d3a

Observation 11374560-06eb-44b7-92b9-335ce53c0f5b · outbound

This paper cites Sigmoid loss for language image pre-training.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Sigmoid loss for language image pre-training

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.533868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.341115Z digest=sha256:27f6f2dcf071326f30b0756cd05cdeffc03c845b4106dd83171a6d6dd7030d95

Observation e2f4d276-e4cc-4a3c-a279-99e31ae5b901 · outbound

This paper cites Glipv2: Unifying localiza- tion and vision-language understanding.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Glipv2: Unifying localiza- tion and vision-language understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.524614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.343845Z digest=sha256:007432c8b1c4abf95e486301ff367ec51f91ef8c1e615b30797c0b345cead775

Observation 446fe260-9a89-40b8-8fed-ec76fc9d0af3 · outbound

This paper cites Dino: Detr with improved denoising anchor boxes for end-to-end object detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Dino: Detr with improved denoising anchor boxes for end-to-end object detection

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.515459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.346548Z digest=sha256:7b350c4c5360fba15dde70965876496540d8c7abf1cde18a80c54f8216856bbc

Observation f8cf1f89-d93d-4b5e-a524-0f87991f34d2 · outbound

This paper cites Generating enhanced negatives for training language-based object detec- tors.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Generating enhanced negatives for training language-based object detec- tors

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.506604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.349188Z digest=sha256:871eeb961bea00b0daafd2b43e020e34a6865a559dfd7a5c1e21cc67f566fefb

Observation 2a559df2-72fb-4529-82c1-3936457f60a1 · outbound

This paper cites An Open and Comprehensive Pipeline for Unified Object Grounding and Detection.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T21:55:23.351984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:55:23.351984Z digest=sha256:72106e3eb7cf9773679e9ccfe5f26ff336dcfa6aad2d6957bdd4bdfe3e7e0430

Observation b9aa56ca-f211-4658-a001-dc00bb9d5df4 · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Detecting twenty-thousand classes using image-level supervision

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.497229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.355152Z digest=sha256:692f92090f5e5c4eee80645338234bafdc3c868135bd903aec6de974faf94847

Observation 7bade701-8eca-4ae2-9985-41543c2c0371 · outbound

This paper cites Mova: Adapting mixture of vision experts to multimodal context.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Mova: Adapting mixture of vision experts to multimodal context

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.488169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.357983Z digest=sha256:2914593344428ec827aaa1fb45a151006bae1b460271acba83ee30888d27893d

Observation e1d54f5f-1378-46b5-98c1-321d736e2443 · outbound

This paper cites the man” and “umbrella.

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models the man” and “umbrella

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:55:23.478023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T21:55:23.361176Z digest=sha256:7818664343b4cd6f8980cd2fce55fa82709005d19c5b3885253b1e589b371154

Pith citing papers

Observation ae3df2a4-fd49-4207-9b4b-a7fc500f5417 · inbound

RF-DETR Object Detection vs YOLOv12 : A Study of Transformer-based and CNN-based Architectures for Single-Class and Multi-Class Greenfruit Detection in Complex Orchard Environments Under Label Ambiguity cites this paper.

RF-DETR Object Detection vs YOLOv12 : A Study of Transformer-based and CNN-based Architectures for Single-Class and Multi-Class Greenfruit Detection in Complex Orchard Environments Under Label Ambiguity LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:36.580979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:36.580979Z digest=sha256:4f077829bcc8983ef1f2e51c19cb7a2e09fd50ade550fd3393c53d44459d446a

Observation f63d6b4c-fd85-4bc0-ad99-27228d92063a · inbound

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion cites this paper.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:42.436018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:42.436018Z digest=sha256:ea067b598404805b353318400f2ba6931f18591c187b0522a5b2061fddbd2eaf

Observation b9aa25c1-7b01-4448-b0df-06e09001edc1 · inbound

RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution cites this paper.

RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:35.846198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:33:35.846198Z digest=sha256:f54569eae6250885388ecab1ff45348f7d28f811de8819ce24e7180b6d445f5a

Observation 03db883a-530e-44f5-af38-8974a28cb22c · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:04.920344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:04.920344Z digest=sha256:53264a2b7b1661214f5be5da7117770ee27599e5d99ef31ee6e15eedb27f3f2d

Observation 925712cc-f512-4634-8d6b-39ea59245400 · inbound

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation cites this paper.

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:46:33.284362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T09:40:04.685274Z digest=sha256:cb39a7a6f7a2b376665a9df42f7fda25cbf9c97307ad5d019c7bfa1b3c8337df

Observation 28f99f0d-5609-453e-9ea9-da3eb702dda7 · inbound

MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts cites this paper.

MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T14:44:05.247859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:44:05.247859Z digest=sha256:d027ba5b497aff19545876883ac846b1886f7c35b10357864f737311e03905e7

Observation 2170e321-15aa-4d70-867c-ce1a46fcb057 · inbound

Kitchen Robotic Manipulation utilizing Foundation Models cites this paper.

Kitchen Robotic Manipulation utilizing Foundation Models LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:49:13.593500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:49:13.593500Z digest=sha256:1bae1c25dd215761cab4ca6983219ae41d9e29962676de833d9f23090d00be3c