Pith. sign in

Paper Citation Record · LEDGER

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

As of 13 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 7 inbound Pith citation observations for arXiv:2505.18986.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18986 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:43.959038Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T07:07:59.925795Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:17:58.075590Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8031d5a1-ae22-493d-b571-74ad3918e70f · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.647029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:41.306134Z digest=sha256:66de826b1db272c36362bcfbc5929ebfaa818de965a444a5bc6d48f25e07727f

Observation 80bd6b94-22d0-4e78-98cc-3191d61c86c3 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.407593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.407593Z digest=sha256:a27a4e45d2692332d5264f057ffed6313204e73e7b9935855753083956455ba7

Observation 6815a255-5f92-403c-a406-0daa57900b10 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion A simple framework for contrastive learning of visual representations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.524550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.524550Z digest=sha256:e8f4b1c97050b512b8bf2d5cf1ecb7ef9ce1f7c6dea340aeb86cc50b1fce50e5

Observation 7695381b-dbd6-445b-acb7-586bd352201e · outbound

This paper cites Pix2seq: A Language Modeling Framework for Object Detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Pix2seq: A Language Modeling Framework for Object Detection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.630304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.630304Z digest=sha256:3906a479f4059b198064723e3c0f389800ae46c3de7d8cb0895d6b39b3622ae2

Observation 33c63cf9-3e2e-4215-ada8-85d55d388162 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.747798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.747798Z digest=sha256:b72a6db9db83cabcbb6cd480b1a0f1e5c9ff675f610401f227a8a9970934af5a

Observation 48a071c6-3d4b-4e2f-bffb-25d64aea5eb1 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.836297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.836297Z digest=sha256:58d5f05a7bcee3f40184bc06c9791e16225886d96ac7d41df8167d9134f07b36

Observation 288d0ccd-d231-4159-be59-242400cd12f7 · outbound

This paper cites Yolo-world: Real- time open-vocabulary object detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Yolo-world: Real- time open-vocabulary object detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.612379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:41.984356Z digest=sha256:4bfa8e56dd1e9d79251c41a5b6a99b2b06d7f16234c750c6f379ce58c91e5235

Observation 83de4954-8d7b-44f0-9aa2-636c8924faf8 · outbound

This paper cites Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:42.104267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:42.104267Z digest=sha256:0d677187db488a6c9bd1e709a841e110c03a263125f3373557ca6e7eef22bb7a

Observation b01f4f06-47e3-4fb3-ae57-d385e3b4b817 · outbound

This paper cites Reducing network agnostophobia.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Reducing network agnostophobia

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.596905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:42.210795Z digest=sha256:93c0ea81b61ae5a62bb5225f41866ef1a92284ef22f9864a8ed1cb5cfc8350da

Observation 9db2f7de-1b34-4bd2-9a0c-f139f406c28a · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion An image is worth 16x16 words: Transformers for image recognition at scale

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.582137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:42.302873Z digest=sha256:f5a8f2edbef4055a2899f33da88e64d8b627d984a56103843f5dfa7a99e9e531

Observation f63d6b4c-fd85-4bc0-ad99-27228d92063a · outbound

This paper cites LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:42.436018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:42.436018Z digest=sha256:fab528f98a7b1242a2bbda78889301c5613675b4f64fce8d0ac30c8d59c8894b

Observation 19dd24c6-c3e8-4896-9945-e7cf4b572175 · outbound

This paper cites Recent advances in open set recognition: A survey.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Recent advances in open set recognition: A survey

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.566449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:42.562637Z digest=sha256:91791aa54ac05dff8d45ce6689f964694fe10aa2bab8929e882d442c5f41c1a0

Observation d1c5f081-6694-4ff2-8139-7c0bcc87d935 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Lvis: A dataset for large vocabulary instance segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.550654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:42.643665Z digest=sha256:7c38aae0d1fdee324d42bbbb9b42fd7c72f9cedac22b0cfa97d5184a16e221fd

Observation 129904e7-f0fc-417d-a952-164a0f83f016 · outbound

This paper cites Ow- detr: Open-world detection transformer.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Ow- detr: Open-world detection transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.535337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:42.746308Z digest=sha256:3d00f976a7f3e3a812621aa2ff716bc791a768e4ef3254cf2ecf5b64c44a775d

Observation b5ca2b6e-0f50-4553-a3d8-05fb99cb394f · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Momentum contrast for unsupervised visual representation learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.520086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:42.853037Z digest=sha256:626d9d66505f42bea81637b7340dd4080091e145bd340589f804651a5e330a58

Observation 1c8338aa-67bb-4a63-9bc2-2445730ab42b · outbound

This paper cites Mask r-cnn.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Mask r-cnn

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.505146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:42.964370Z digest=sha256:6ee568bb636b6def4979f224fb0146a0bb7ba775b0966fa3131d7b892c895f0e

Observation e2ea59b2-12b6-42dc-9e04-128fea24ba2f · outbound

This paper cites T-rex2: Towards generic object detection via text-visual prompt synergy.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion T-rex2: Towards generic object detection via text-visual prompt synergy

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.490204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.103990Z digest=sha256:460c30b604ed2000dd0004031628f56dc5b869ab230068e1010209d2ca72a6a4

Observation d3a8e7ce-0b56-4ead-b1f4-6a1d7596a753 · outbound

This paper cites Segment anything.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Segment anything

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.474379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.228584Z digest=sha256:2e04dc45faa2311427cf58c942eecaebe1bf9c5a3d1419f76012ca4d430cdff9

Observation c0219f15-f889-4c94-a70b-64a2b3560a8c · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Lisa: Reasoning segmentation via large language model

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.458987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.350210Z digest=sha256:ed0c742346250fef863aaca9795cf5fc4e20d07f2232c2cab51bd4339501a0b4

Observation 95ca5c2d-c1c2-4d79-807c-bc76acb402a1 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion LLaVA-OneVision: Easy Visual Task Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:43.465321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:43.465321Z digest=sha256:fd9a7853f20f601b3bb009fa34092efe6159608178ecaefd3f620f52a0d64f1f

Observation 0a9ca5e0-6b27-41a1-bd80-8c35644b54c7 · outbound

This paper cites Dn-detr: Accelerate detr training by introducing query denoising.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Dn-detr: Accelerate detr training by introducing query denoising

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.443299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.514620Z digest=sha256:8aeca74d663520183cdf71f7c257fedea24a575e62e48642a91ac80ddd8212c0

Observation a9ccb8e4-1767-4134-806f-97bc6c3f423b · outbound

This paper cites Coda: A real-world road corner case dataset for object detection in autonomous driving.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Coda: A real-world road corner case dataset for object detection in autonomous driving

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.426559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.613826Z digest=sha256:932c4eaf4afda865783a7fc2c0e1a2834ed272165808cd5f52ad81a853dcf132

Observation 1767c512-b56c-4a29-8389-15275d85d357 · outbound

This paper cites Desco: Learning object recognition with rich language descriptions.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Desco: Learning object recognition with rich language descriptions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.409036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.758188Z digest=sha256:dfe5f3539143ff1807a34bc1d135ddc147d120b5e0cc2cea133d36582f53294c

Observation 6188faa5-3e31-4a7e-a0f4-045ff99fa574 · outbound

This paper cites Grounded language-image pre-training.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Grounded language-image pre-training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.393433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.847663Z digest=sha256:02861723cd880e389eac77666b998a93b3caf1df66584a3ccf84e44de7d7c0e4

Observation 1ea8ff82-22a0-4319-a530-fb6425cc897f · outbound

This paper cites Generative region-language pretraining for open-ended object detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Generative region-language pretraining for open-ended object detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.377953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.855977Z digest=sha256:3c32e34725642daacdd0869113801dd79ea33681e860826c700c42239a4ec53b

Observation b1cd2f1d-0e0a-4292-b2cf-c793afca37fb · outbound

This paper cites Training-free open-ended object detection and segmentation via attention as prompts.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Training-free open-ended object detection and segmentation via attention as prompts

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.361788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.862354Z digest=sha256:1eec2e74690a3a670139cd49e12d44f6a3b546b7b9997803b44394e9bc6c6390

Observation e459fbe4-2f7e-433e-943f-ce29b7572115 · outbound

This paper cites Visual instruction tuning.Neural Information Processing Systems (NeurIPS), 2023.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Visual instruction tuning.Neural Information Processing Systems (NeurIPS), 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.344472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.867202Z digest=sha256:67493ee1f3fcf6629f56045232635e6b68f52b50325870b840a12c1e07cda72a

Observation 65e03801-1900-45f7-a0fb-91aff4a1bc80 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.328450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.873335Z digest=sha256:2641b46b4463f04f89f427f199f46928c34594fd61df5717da3d1a3ab628c58b

Observation 66878a2e-4dc8-44c5-a84a-3a94ab7d4a43 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Swin transformer: Hierarchical vision transformer using shifted windows

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.312770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.879076Z digest=sha256:47be183923388ca834f14bb79b5ae7baea5a19fec2c7ff3699c9d54e6e902b5c

Observation 53e798e5-d546-4307-b90d-0fb5d8dc97ad · outbound

This paper cites Capdet: Unifying dense captioning and open-world detection pretraining.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Capdet: Unifying dense captioning and open-world detection pretraining

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.296731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.884458Z digest=sha256:6d45797a702e577648df1841489ebb10f229514082964359586b6970d9ec134e

Observation 86bf87f0-ecd7-4e52-9836-8ac65fd7ce75 · outbound

This paper cites Scaling open-vocabulary object detection.Neural Information Processing Systems (NeurIPS), 2023.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Scaling open-vocabulary object detection.Neural Information Processing Systems (NeurIPS), 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.281439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.889852Z digest=sha256:e44e77870986fcf959a7dcd92b852e90f3b14f6c49205d3f69b4146ad1831ac7

Observation d0e3482c-010f-4af8-8e77-014c190f77be · outbound

This paper cites Learning transferable visual models from natural language supervision.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Learning transferable visual models from natural language supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:43.895759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:43.895759Z digest=sha256:15f1c0d3b2e6bf696f601962ecba6c7694ad9df9340b6f1ced38d002cc1dbc82

Observation 1cf861bd-aa92-4732-b410-788eb9ba0f3d · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.256381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.901395Z digest=sha256:c06fe17b8f18e686f3a257d099adc5e36a8f127e610ab58c4e4011901a9d26bf

Observation 952f7a13-222b-4092-b791-fa62baf447d8 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:43.907816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:43.907816Z digest=sha256:dfb6e9f01f6358e25b15bdbd05713988db77b0fc6deefea8bcb668071d70b67f

Observation bc5ded50-7edf-402d-a22f-6193f73cc3b1 · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Scalability in perception for autonomous driving: Waymo open dataset

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.239758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.914161Z digest=sha256:9df14b288cf3099d43b462231a3491d049707b1d990e35decd58a1e8480d00e5

Observation 46e99aa6-694f-4db8-ab1f-aac1b4d565c3 · outbound

This paper cites OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:43.919389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:43.919389Z digest=sha256:c1ad90edfad2e2c998876842edd87feebbc6e0cca525051a8c3d9c8d88c57edd

Observation c90909b2-a67e-4f66-8a7d-cde1ca2d1edb · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Cogvlm: Visual expert for pretrained language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.223244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.926211Z digest=sha256:b4843a782dce500bd2cacb0a7ee5b053bb5f01be94e09faec35493a09aeb959a

Observation f802fc6a-8f2c-467f-998a-c86d33c79d40 · outbound

This paper cites Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.207117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.931973Z digest=sha256:d0f0235ecd3d48a0311a3d2c8203bc0143a9102d149cb8009f77c1e3d3e7b701

Observation a048e2b5-a136-4b19-ae26-7e841210f4b8 · outbound

This paper cites Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.188398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.937160Z digest=sha256:5469f98e4abf2adc6e4b344bb5734a923c8dfa529ef1373847c79ff3d07aa451

Observation 0cfd0a67-5e40-40cb-ad5a-0454b65051c3 · outbound

This paper cites Detclipv3: Towards versatile generative open-vocabulary object detection.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Detclipv3: Towards versatile generative open-vocabulary object detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.171851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.942085Z digest=sha256:5ec5be71a8daf91ce352a3c3dcd5297ed5998cc412a87988cabf9171efbdf80b

Observation e33a2214-f781-4e90-92e1-5313eccf1022 · outbound

This paper cites Ni, and Heung-Yeung Shum.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Ni, and Heung-Yeung Shum

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.155459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.947653Z digest=sha256:740cdcc7e2c5f914bf060e9ccf6b39b7b58d8ec69aa92b7bee03d8b28ec52d7b

Observation bf19aa99-7d41-4b33-830f-9a967b0586c9 · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Llava-grounding: Grounded visual chat with large multimodal models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.140375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.953447Z digest=sha256:8fa49a6105d75027f2ebd202dcff19e58618a3592490919c65281993d7067bf2

Observation a910885d-f952-4053-a270-93fbdc12f077 · outbound

This paper cites Glipv2: Unifying localization and vision-language understanding.

VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion Glipv2: Unifying localization and vision-language understanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:25:44.123477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:25:43.959038Z digest=sha256:c7d25f7fcad9423ac092bf429473be339ad00df01589108b43514aac928d598b

Pith citing papers

Observation f945f4ea-7141-4be0-a3e6-3f2553c5d7f9 · inbound

Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs cites this paper.

Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:25:51.983415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T19:51:59.102471Z digest=sha256:5e44c469b9effd663d8ab7b8031cd7e34fba879b43dd5e180df9b3fdd13d5374

Observation fdd5e044-6917-4c84-b422-71aaa53799e8 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:30.268838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T01:24:56.236650Z digest=sha256:e68bbc1c7f8f7e03fff10124aec04f29126a7bfd77cb46a224b48f9a802ced51

Observation d0f78be8-c593-4474-959b-3a09d910a640 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:35.844779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T19:14:12.092478Z digest=sha256:002fc230123e6c0f68decf55f50e46c0784c0bdfb603e2927f67b2f8a5416b4d

Observation 36545036-448a-4b16-a2fa-2703d6456869 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:48.393215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T01:33:55.317107Z digest=sha256:59e7730c3dee34d28ab279dc89197b8d89296180ac757d1c985d5c869ed664dc

Observation 0e043359-7b35-4c85-b811-1538a58c95d3 · inbound

FoodCHA: Multi-Modal LLM Agent for Fine-Grained Food Analysis cites this paper.

FoodCHA: Multi-Modal LLM Agent for Fine-Grained Food Analysis VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:21:09.469192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T16:11:26.104498Z digest=sha256:3eed66c7f04f0edd8d915287ad65370ba21e781a7221fb237838dcb0a75336e3

Observation 3bae7f73-e0ef-4b89-8e37-602e64d41309 · inbound

3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding cites this paper.

3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:13:16.671658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T07:07:59.925795Z digest=sha256:3ee6e1d144c19c69ed473db44de77b1be0dccc5a0119645da22ea2117d3abfdb

Observation 0aeb88e6-1235-4112-9af7-b80b89eb9aba · inbound

DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems cites this paper.

DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:17:58.077080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T10:05:38.910257Z digest=sha256:c8c90c4e0470fc09401a40e7cb2168f00638148ce5e9993a9bb2ec73d4f0d33c