Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Grounding as Bidirectional Concept Correspondence

As of 14 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2608.07886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07886 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:50:38.904354Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved32
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a42b76bf-1cfe-4317-8a5c-a46d102f6d8e · outbound

This paper cites Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi.

Vision-Language Grounding as Bidirectional Concept Correspondence Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:40.224535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.596093Z digest=sha256:4ce17d812a84e9c3b65e66b5f38b9ae560e7a57c36be1394cd40ad03a32a8a13

Observation 0da74421-f914-4eba-b191-0ecc1b6324fb · outbound

This paper cites Molmo2: Open weights and data for vision-language models with video understanding and grounding, 2026.

Vision-Language Grounding as Bidirectional Concept Correspondence Molmo2: Open weights and data for vision-language models with video understanding and grounding, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.602120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.602120Z digest=sha256:943b201de2a3b3ba67bca84159c831d83a6eb8c3f996a51820713340c4ac9f9d

Observation 0d8618f7-da09-4da1-a378-680acfce865c · outbound

This paper cites Molmopoint: Better pointing for vlms with grounding tokens.arXiv preprint arXiv:2603.28069, 2026.

Vision-Language Grounding as Bidirectional Concept Correspondence Molmopoint: Better pointing for vlms with grounding tokens.arXiv preprint arXiv:2603.28069, 2026

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.608089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.608089Z digest=sha256:8a4cb4ebd00ea1c59c422b71c8bbdf03534a7d057e184e6a38d618bd5bc6f72a

Observation 66de1c08-f1cb-40b4-b1cb-e060113ae781 · outbound

This paper cites Modeling context in referring expressions.

Vision-Language Grounding as Bidirectional Concept Correspondence Modeling context in referring expressions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.613911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.613911Z digest=sha256:9f3819098c680d2adb295ff3fcdf3279f9e5c9836c0bc5264c4f665d30640ed3

Observation e0b4e406-1737-4ae7-a2d3-9020be4852e4 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

Vision-Language Grounding as Bidirectional Concept Correspondence Generation and comprehension of unambiguous object descriptions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.618665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.618665Z digest=sha256:04af44852f3c92b9d66e246f09b4718665bd4840e534a7dfc68bf60308113f99

Observation 9a4312cc-44bf-460c-a6c8-e2453b8ba94d · outbound

This paper cites Modeling context between objects for referring expression understanding.

Vision-Language Grounding as Bidirectional Concept Correspondence Modeling context between objects for referring expression understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.624092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.624092Z digest=sha256:a1cfadc5da724cb88523d8abed02a2e544d3aef4770dc979fc4157edc704b2ae

Observation adabe624-e8d7-4177-bf97-0a44d01ee053 · outbound

This paper cites Referring relationships.

Vision-Language Grounding as Bidirectional Concept Correspondence Referring relationships

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:40.168590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.629547Z digest=sha256:00530118676ecb62efeee06d1760795511ea598763817c09ace937d8f6bb5900

Observation 184ca318-d921-438b-b73f-68dc0a3b0876 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Vision-Language Grounding as Bidirectional Concept Correspondence Referitgame: Referring to objects in photographs of natural scenes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.634193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.634193Z digest=sha256:a83e50e6dd69e6c0b9d87195cfcb9e0b1a88c625f2943b259b21bc6ed05f9114

Observation 43ee5f8e-a942-4e15-b378-19a5675e0c8a · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

Vision-Language Grounding as Bidirectional Concept Correspondence Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.639144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.639144Z digest=sha256:27647d5f88a9f7a0aad55315d3883994f0900ae4ca8221ea40902e178b3fec7a

Observation fe7e47d9-79fc-495d-b427-85e69978b255 · outbound

This paper cites Phrasecut: Language-based image segmentation in the wild.

Vision-Language Grounding as Bidirectional Concept Correspondence Phrasecut: Language-based image segmentation in the wild

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:40.129902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.644558Z digest=sha256:55ee88c11dbdbc392c794b7aeeb60472da3a1181de6f8582b826fea759f7f861

Observation d7db2878-97fa-4012-ae06-1a882c876e72 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Vision-Language Grounding as Bidirectional Concept Correspondence Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.650685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.650685Z digest=sha256:fdedb0ab1ef0632e82c57294ea27a11e949fd8cf8854e930e7ff99f95287f6bc

Observation da33c1f9-b611-415c-8544-6be218767bee · outbound

This paper cites Sam 3: Segment anything with concepts, 2025.

Vision-Language Grounding as Bidirectional Concept Correspondence Sam 3: Segment anything with concepts, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.655891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.655891Z digest=sha256:4624db97dde83949bf8ffec62601bed5426843d82ffe2395ab541cc7e0d241ff

Observation a60519be-7c22-484e-adc7-3cdc25c8af31 · outbound

This paper cites Clark.Using Language.

Vision-Language Grounding as Bidirectional Concept Correspondence Clark.Using Language

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:40.100677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.661129Z digest=sha256:1e78382e49f9d8a7c6e9e23cf503529a48319aa8cd0f647ac661578ecd8aa609

Observation 5caf5415-ac91-4220-974f-016cfe5ffdd7 · outbound

This paper cites Clark and Susan E.

Vision-Language Grounding as Bidirectional Concept Correspondence Clark and Susan E

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:40.081751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.665612Z digest=sha256:ef6572383c447579c0abf43c89e58e90a12eb323d7e139048105041d800dfd69

Observation 091329e9-49ff-4360-86a8-fe8c0a80beb3 · outbound

This paper cites Spatial mental models.

Vision-Language Grounding as Bidirectional Concept Correspondence Spatial mental models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:40.062604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.670644Z digest=sha256:955b1f3882c9dd18f6ff0bdd5c5330b3588865ac46ac6bfc1cf8b14764ff1439

Observation 16eda70f-eb8d-4b3b-9c5e-1d806f96b16a · outbound

This paper cites Taylor and Barbara Tversky.

Vision-Language Grounding as Bidirectional Concept Correspondence Taylor and Barbara Tversky

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:40.042794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.675891Z digest=sha256:62fdd1d8a8ea0a2faed0db37da87dba13c6dfb5e77bf78446f41a8e62d4cb2a7

Observation 20626ba8-a4d7-44bb-890a-8fdc547ece8b · outbound

This paper cites Selective Visual Representations Improve Convergence and Generalization for Embodied AI.

Vision-Language Grounding as Bidirectional Concept Correspondence Selective Visual Representations Improve Convergence and Generalization for Embodied AI

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.681678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.681678Z digest=sha256:1522e982e997cf0d7b2e54ecf73ef0e57dbfb3bdff881d0e3e336ad28c3bb7b1

Observation e1d81fc1-253e-48a6-8da9-85276ba61ee0 · outbound

This paper cites Treisman and Garry Gelade.

Vision-Language Grounding as Bidirectional Concept Correspondence Treisman and Garry Gelade

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.687795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.687795Z digest=sha256:4c78bf3b1e7ff21eae3021adee83c84e5950bb092667bc5265720857cb7816b5

Observation 1ad7fc1f-0fed-4657-9285-00786092948a · outbound

This paper cites Structure-mapping: A theoretical framework for analogy.Cognitive Science, 7(2):155–170, 1983.

Vision-Language Grounding as Bidirectional Concept Correspondence Structure-mapping: A theoretical framework for analogy.Cognitive Science, 7(2):155–170, 1983

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:40.004128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.692912Z digest=sha256:312742d8482bb5e419745719520bc2094171cce83bf6eff96f4cfc42a7c25105

Observation ece0bc17-9665-4074-b590-4da422e5336c · outbound

This paper cites Who are you referring to? coreference resolution in image narrations.

Vision-Language Grounding as Bidirectional Concept Correspondence Who are you referring to? coreference resolution in image narrations

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.981271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.698415Z digest=sha256:bc6e20de8471d023ae18f3992635a327cba3b3b3c420678a291ef69d840f18d7

Observation 58fd4ad6-37ad-4f50-a2a1-1e1cec8bc804 · outbound

This paper cites Understanding natural language.Cognitive Psychology, 3(1):1–191, 1972.

Vision-Language Grounding as Bidirectional Concept Correspondence Understanding natural language.Cognitive Psychology, 3(1):1–191, 1972

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.962247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.703883Z digest=sha256:46499e54866ddbfd543276723401eba2f7c61989e60f97f252b30a951236b2b5

Observation 62a3f04f-6685-4013-9c57-60541eaf82a1 · outbound

This paper cites Levesque, Ernest Davis, and Leora Morgenstern.

Vision-Language Grounding as Bidirectional Concept Correspondence Levesque, Ernest Davis, and Leora Morgenstern

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.944755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.709378Z digest=sha256:213f23ade3df787fd02f61ebde20bd21b18f1763c9ad141ad98394a3baca38ef

Observation 35e35cf0-2588-47d3-af40-f71304ec8583 · outbound

This paper cites Picturing ambiguity: A visual twist on the winograd schema challenge.

Vision-Language Grounding as Bidirectional Concept Correspondence Picturing ambiguity: A visual twist on the winograd schema challenge

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.926171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.716012Z digest=sha256:66f867f198300bdabc40ca744d89780598a45a9fb12875df43ef1e4570bdd2ce

Observation 37941aa8-3b7e-45aa-9aa0-9a10af435fc8 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

Vision-Language Grounding as Bidirectional Concept Correspondence Qwen3.5: Towards native multimodal agents, February 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.721658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.721658Z digest=sha256:f7054f812eedfd0977887b43527a58fc9b0b9f7770bf1b357d247589465fc755

Observation bd57747f-0275-479e-a477-9c9d0f86943a · outbound

This paper cites Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding.

Vision-Language Grounding as Bidirectional Concept Correspondence Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.728318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.728318Z digest=sha256:312775b3141d7f1c43a1a0dc3f30e8783155f73c8eb35c579f81f1917071973c

Observation be17a4c4-c340-4bbc-9141-a727a01ad8a9 · outbound

This paper cites One trajectory, one token: Grounded video tokenization via panoptic sub- object trajectory.

Vision-Language Grounding as Bidirectional Concept Correspondence One trajectory, one token: Grounded video tokenization via panoptic sub- object trajectory

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.894655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.734939Z digest=sha256:928650efa734538fa825aa88876879e544a6a55d866e3f12bc8a420b57f5ac18

Observation 172d8766-b817-41d9-8805-e24e05b7de94 · outbound

This paper cites TrajTok: Learning Trajectory Tokens enables better Video Understanding.

Vision-Language Grounding as Bidirectional Concept Correspondence TrajTok: Learning Trajectory Tokens enables better Video Understanding

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:39.331928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.739988Z digest=sha256:bf0cb07c009ce11854946f0cc46de1ec71273c785205a73defbe0e39f1ecbc48

Observation 85ea8995-84c7-438a-8382-94d1a09a2c77 · outbound

This paper cites Youtu-vl: Unleashing visual potential via unified vision-language supervision.

Vision-Language Grounding as Bidirectional Concept Correspondence Youtu-vl: Unleashing visual potential via unified vision-language supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.744971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.744971Z digest=sha256:628f44b6c4bed9802113076ac4248608b4f5505476a30d989effc71d3430250b

Observation 7a913f9e-ad34-49cb-a90b-375542c7f88f · outbound

This paper cites Grounded language-image pre-training.

Vision-Language Grounding as Bidirectional Concept Correspondence Grounded language-image pre-training

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.878393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.749934Z digest=sha256:73255600777e98c0c175c4de7cf3277280b6d26aff0952c4aab266b043603300

Observation fddb5876-cb24-4df8-a84f-4831cc1bafc9 · outbound

This paper cites Glipv2: Unifying localization and vision-language understanding.Advances in Neural Information Processing Systems, 35:36067–36080, 2022.

Vision-Language Grounding as Bidirectional Concept Correspondence Glipv2: Unifying localization and vision-language understanding.Advances in Neural Information Processing Systems, 35:36067–36080, 2022

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.861293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.754251Z digest=sha256:5a7c1405ba263b29c60b64032cd9e108ec33ce5af726deb6e1874cfeaa5e9018

Observation 77cd3d39-85d9-4e5d-bfdd-fba5d0e41a17 · outbound

This paper cites Synthetic visual genome.

Vision-Language Grounding as Bidirectional Concept Correspondence Synthetic visual genome

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.840478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.759067Z digest=sha256:bf2e1e4139ac32025c6fae8f73304a8ee8ec8c3431c8070db3e9ae3c12e2d818

Observation b7ea31fb-e3d0-4c8b-b2b6-cece6cbc090d · outbound

This paper cites You, Daniel Ogbu, Chenhao Zheng, Weikai Huang, Yinuo Yang, Winson Han, Quan Kong, Rajat Saini, and Ranjay Krishna.

Vision-Language Grounding as Bidirectional Concept Correspondence You, Daniel Ogbu, Chenhao Zheng, Weikai Huang, Yinuo Yang, Winson Han, Quan Kong, Rajat Saini, and Ranjay Krishna

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.820974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.764173Z digest=sha256:b13577c28a57e7b9c6cf942d38515e122f6c1ddb5a495c147388e716afc868a7

Observation 63783363-2965-4003-b570-7891af021b91 · outbound

This paper cites Mdetr-modulated detection for end-to-end multi-modal understanding.

Vision-Language Grounding as Bidirectional Concept Correspondence Mdetr-modulated detection for end-to-end multi-modal understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.768528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.768528Z digest=sha256:09f81bc9520ad52cfc4c9419afebb0fbf5f2c2a4185026aa1f8d84d3a9477aa8

Observation fe798667-ef84-48ed-a666-839e5fe43cdd · outbound

This paper cites Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection.Advances in Neural Information Processing Systems, 35:9125–9138, 2022.

Vision-Language Grounding as Bidirectional Concept Correspondence Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection.Advances in Neural Information Processing Systems, 35:9125–9138, 2022

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.773695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.773695Z digest=sha256:1a1e37d0abcbb5374ce616262e1dfdc11b00e14f6ec462b300ef0d5c25252642

Observation 40b42858-ba0f-4a7b-9fa0-e971e9b9fbf7 · outbound

This paper cites Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment.

Vision-Language Grounding as Bidirectional Concept Correspondence Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.775293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.778660Z digest=sha256:f0ae0e22806f1f3eb716a48fb181b86bfc8415f9adad641369fecaa3c461d1b8

Observation 33173060-5025-41e1-881b-8a7c7d4b9041 · outbound

This paper cites An Open and Comprehensive Pipeline for Unified Object Grounding and Detection.

Vision-Language Grounding as Bidirectional Concept Correspondence An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.784349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.784349Z digest=sha256:cdda1a4c73697e93d498f13951b825f95a8d5ddfe5553e6aa2365f3ec66a8a25

Observation 31b7c7f1-5b07-4da6-a806-cea59bd7fcfc · outbound

This paper cites Llmdet: Learning strong open-vocabulary object detectors under the supervision of large language models.

Vision-Language Grounding as Bidirectional Concept Correspondence Llmdet: Learning strong open-vocabulary object detectors under the supervision of large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.757128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.789636Z digest=sha256:f48948f5532dfbf7bc37ed71a128681f27b88daf447f45f4f7da9e5f6d313066

Observation a3b87be8-b678-47c9-b383-978b091bd266 · outbound

This paper cites General object foundation model for images and videos at scale.

Vision-Language Grounding as Bidirectional Concept Correspondence General object foundation model for images and videos at scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.795494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.795494Z digest=sha256:c5d0dbbefca6b1d490020b1f6c5a30189b7c9ff08ff2f9116986fbd2f4e4bbeb

Observation de76a102-254c-4923-bf96-75629d0069b5 · outbound

This paper cites Generalized decoding for pixel, image, and language.

Vision-Language Grounding as Bidirectional Concept Correspondence Generalized decoding for pixel, image, and language

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.800990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.800990Z digest=sha256:59b9f0ebff109fa85b0afaeaabb435a78ee7266293c038836f17722a25815285

Observation 7746fd4f-a1ec-43c3-a507-0e0c2ba0844a · outbound

This paper cites A simple framework for open-vocabulary segmentation and detection.

Vision-Language Grounding as Bidirectional Concept Correspondence A simple framework for open-vocabulary segmentation and detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.805958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.805958Z digest=sha256:33dadd3c14c47f3ebe7679937f4a2fad82f96edbcbc795cdec555495e6e15639

Observation 18d6d323-6993-4d3a-91ce-66b7c7f404d9 · outbound

This paper cites Segment everything everywhere all at once.

Vision-Language Grounding as Bidirectional Concept Correspondence Segment everything everywhere all at once

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.689614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.811376Z digest=sha256:ba2709993ad6b469aa64619a3c098211b43ce1d18d00fcdb39c2c313ad676a48

Observation 61cbbf9c-bd7e-40c8-97a6-5cf21afd95ff · outbound

This paper cites Open- worldsam: Extending sam2 for universal image segmentation with language prompts.arXiv preprint arXiv:2507.05427, 2025.

Vision-Language Grounding as Bidirectional Concept Correspondence Open- worldsam: Extending sam2 for universal image segmentation with language prompts.arXiv preprint arXiv:2507.05427, 2025

Reference 42

Resolution
verified exact
raw_fallback, observed 2026-08-12T00:50:39.199874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.816000Z digest=sha256:c3954a960c820c3fd8f649b7d6b5e6fde81f9377c4a9bebda68757058e1f9029

Observation 8d45cf05-1f07-4640-89a8-b4f3cd67e6fc · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

Vision-Language Grounding as Bidirectional Concept Correspondence Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.820534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.820534Z digest=sha256:4175bac2e2a0fe468549364c889d5d0d14998fc1e1ca0ab1ac6a6e7c98dad235

Observation 179f56a2-0934-43a0-b6ba-5795bb5577b5 · outbound

This paper cites Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm Selection.

Vision-Language Grounding as Bidirectional Concept Correspondence Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm Selection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.824993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.824993Z digest=sha256:84251e109815ce9b1038b6ca34563a4edfe7fc9054577f26b05fb69dd3e19c4c

Observation dcdd51bf-6c3a-4e16-a828-df52d40381e8 · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

Vision-Language Grounding as Bidirectional Concept Correspondence LISA: Reasoning Segmentation via Large Language Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.830162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.830162Z digest=sha256:f6199f20c1fe5376e4d267b4b0c33ef63ed394aa0ce8c5776147648e5de230b7

Observation 18dc40db-1ea0-41e8-8a1c-5957d02ea833 · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

Vision-Language Grounding as Bidirectional Concept Correspondence LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.835625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.835625Z digest=sha256:178c88609d3416bca3f8b5b11588c515d549a7375b625e49a807a55d21e3b8ec

Observation ed570fb5-ef30-44b7-a58f-822870c20593 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Vision-Language Grounding as Bidirectional Concept Correspondence Glamm: Pixel grounding large multimodal model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.840703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.840703Z digest=sha256:58b297da312e15ae8d5d96f2421cb4b7f3c033d15a8b6005896249836cbc4c36

Observation aed20469-30a3-4f04-8de3-94dae3870694 · outbound

This paper cites Pointrend: Image segmentation as rendering.

Vision-Language Grounding as Bidirectional Concept Correspondence Pointrend: Image segmentation as rendering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.846502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.846502Z digest=sha256:7bb875d3786b6c0e0e6b64dffdbe70092098e94e3c6f01be3c86f4e961af34dc

Observation c85a48fb-934a-4c15-b4c3-4c7c9819ac59 · outbound

This paper cites V-net: Fully convolutional neural networks for volumetric medical image segmentation.

Vision-Language Grounding as Bidirectional Concept Correspondence V-net: Fully convolutional neural networks for volumetric medical image segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.625488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.851800Z digest=sha256:b6f2ec1194a63916682d51a42378146a4af3834beb10e6e95929c441ca99a538

Observation 6638bb39-258b-4120-8535-7e1d5ea2f75b · outbound

This paper cites COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation.

Vision-Language Grounding as Bidirectional Concept Correspondence COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.856702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.856702Z digest=sha256:b7da5c632642dcbee4abf2c32a1caa91dc1db16d1bfe644e86fcbb9868961499

Observation 36661101-d3c1-4bae-80ee-34586aece948 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Vision-Language Grounding as Bidirectional Concept Correspondence Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.862099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.862099Z digest=sha256:220ab8689bef7fb88465a35d237bebb16f2d69aab6fad59c5437e59547970c12

Observation 10385fd9-fdb0-4493-a342-4b2ea053479b · outbound

This paper cites Lawrence Zitnick, and Piotr Dollár.

Vision-Language Grounding as Bidirectional Concept Correspondence Lawrence Zitnick, and Piotr Dollár

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.867594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.867594Z digest=sha256:e1f1bdc7a1d1c28f6cff91306ef5a4dde95ee3c5fb5a21571eb284df6b0fb40e

Observation 66deab6a-78d7-4abc-8dcb-e8ae0b93c640 · outbound

This paper cites Coconut: Modernizing coco segmentation.

Vision-Language Grounding as Bidirectional Concept Correspondence Coconut: Modernizing coco segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.583042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.873411Z digest=sha256:5ec7e6de6438d9811015493563e7dadda80cb2fed861c6ced5c4d74a1d028d2f

Observation 1ede06cf-8bff-4a53-a624-56b8f28a7dca · outbound

This paper cites High- quality entity segmentation.

Vision-Language Grounding as Bidirectional Concept Correspondence High- quality entity segmentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.566011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.879081Z digest=sha256:15592cb1d938c47ef9844797526576cca80d2d38ebc8b7ec27a0a760668b4a72

Observation 88e66ab2-807d-44f6-90c5-0ff28ddb4407 · outbound

This paper cites Semantic understanding of scenes through the ade20k dataset.International journal of computer vision, 127(3):302–321, 2019.

Vision-Language Grounding as Bidirectional Concept Correspondence Semantic understanding of scenes through the ade20k dataset.International journal of computer vision, 127(3):302–321, 2019

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.884016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.884016Z digest=sha256:94a621a4a808c0e5b92c7f38188aa61f42771d731de3392ede5b64e6d9cc574f

Observation 173dff5c-6072-4a51-aab6-80fcf0136bba · outbound

This paper cites GREC: Generalized Referring Expression Comprehension.

Vision-Language Grounding as Bidirectional Concept Correspondence GREC: Generalized Referring Expression Comprehension

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.888934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.888934Z digest=sha256:a9a03f25370737b1b91f17831e4e1b60da278c2a15003c6c3e79559886fd36ea

Observation 23ee512f-4627-4f97-b5d8-1b16ba8376c4 · outbound

This paper cites Introducing GPT-5.4.

Vision-Language Grounding as Bidirectional Concept Correspondence Introducing GPT-5.4

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:39.535906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:50:38.894120Z digest=sha256:1350869828460994fa5943b7761d4206b73742c743f9a84d0fc88f23779cfc85

Observation dc15b24f-9320-4aac-810b-8490493cc188 · outbound

This paper cites Segment anything.

Vision-Language Grounding as Bidirectional Concept Correspondence Segment anything

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:38.899194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.899194Z digest=sha256:e9ca9a119d9b7025b4e3d749eb59473d46d449670d93b6602d351cac0a660785

Observation 6932e29d-e9b6-4979-baf5-59e71b848e38 · outbound

This paper cites Roboflow100-vl: A multi-domain object detection benchmark for vision-language models.

Vision-Language Grounding as Bidirectional Concept Correspondence Roboflow100-vl: A multi-domain object detection benchmark for vision-language models

Reference 59

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:50:38.904354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:38.904354Z digest=sha256:f91f2f16e0cba61a868becd96fb5095991897f800975533ca5be532ac910152b

Pith citing papers

No inbound Pith citation observations are available.