Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:50:38.904354Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2608.07886.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:50:38.904354Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a42b76bf-1cfe-4317-8a5c-a46d102f6d8e · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0da74421-f914-4eba-b191-0ecc1b6324fb · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Molmo2: Open weights and data for vision-language models with video understanding and grounding, 2026
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d8618f7-da09-4da1-a378-680acfce865c · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Molmopoint: Better pointing for vlms with grounding tokens.arXiv preprint arXiv:2603.28069, 2026
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66de1c08-f1cb-40b4-b1cb-e060113ae781 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Modeling context in referring expressions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0b4e406-1737-4ae7-a2d3-9020be4852e4 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Generation and comprehension of unambiguous object descriptions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4312cc-44bf-460c-a6c8-e2453b8ba94d · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Modeling context between objects for referring expression understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adabe624-e8d7-4177-bf97-0a44d01ee053 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Referring relationships
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 184ca318-d921-438b-b73f-68dc0a3b0876 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Referitgame: Referring to objects in photographs of natural scenes
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43ee5f8e-a942-4e15-b378-19a5675e0c8a · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe7e47d9-79fc-495d-b427-85e69978b255 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Phrasecut: Language-based image segmentation in the wild
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d7db2878-97fa-4012-ae06-1a882c876e72 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da33c1f9-b611-415c-8544-6be218767bee · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Sam 3: Segment anything with concepts, 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a60519be-7c22-484e-adc7-3cdc25c8af31 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Clark.Using Language
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5caf5415-ac91-4220-974f-016cfe5ffdd7 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Clark and Susan E
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 091329e9-49ff-4360-86a8-fe8c0a80beb3 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Spatial mental models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 16eda70f-eb8d-4b3b-9c5e-1d806f96b16a · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Taylor and Barbara Tversky
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 20626ba8-a4d7-44bb-890a-8fdc547ece8b · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Selective Visual Representations Improve Convergence and Generalization for Embodied AI
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1d81fc1-253e-48a6-8da9-85276ba61ee0 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Treisman and Garry Gelade
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ad7fc1f-0fed-4657-9285-00786092948a · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Structure-mapping: A theoretical framework for analogy.Cognitive Science, 7(2):155–170, 1983
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ece0bc17-9665-4074-b590-4da422e5336c · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Who are you referring to? coreference resolution in image narrations
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 58fd4ad6-37ad-4f50-a2a1-1e1cec8bc804 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Understanding natural language.Cognitive Psychology, 3(1):1–191, 1972
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 62a3f04f-6685-4013-9c57-60541eaf82a1 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Levesque, Ernest Davis, and Leora Morgenstern
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 35e35cf0-2588-47d3-af40-f71304ec8583 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Picturing ambiguity: A visual twist on the winograd schema challenge
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 37941aa8-3b7e-45aa-9aa0-9a10af435fc8 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Qwen3.5: Towards native multimodal agents, February 2026
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd57747f-0275-479e-a477-9c9d0f86943a · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be17a4c4-c340-4bbc-9141-a727a01ad8a9 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence One trajectory, one token: Grounded video tokenization via panoptic sub- object trajectory
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 172d8766-b817-41d9-8805-e24e05b7de94 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence TrajTok: Learning Trajectory Tokens enables better Video Understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 85ea8995-84c7-438a-8382-94d1a09a2c77 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Youtu-vl: Unleashing visual potential via unified vision-language supervision
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a913f9e-ad34-49cb-a90b-375542c7f88f · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Grounded language-image pre-training
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fddb5876-cb24-4df8-a84f-4831cc1bafc9 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Glipv2: Unifying localization and vision-language understanding.Advances in Neural Information Processing Systems, 35:36067–36080, 2022
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 77cd3d39-85d9-4e5d-bfdd-fba5d0e41a17 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Synthetic visual genome
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b7ea31fb-e3d0-4c8b-b2b6-cece6cbc090d · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence You, Daniel Ogbu, Chenhao Zheng, Weikai Huang, Yinuo Yang, Winson Han, Quan Kong, Rajat Saini, and Ranjay Krishna
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 63783363-2965-4003-b570-7891af021b91 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Mdetr-modulated detection for end-to-end multi-modal understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe798667-ef84-48ed-a666-839e5fe43cdd · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection.Advances in Neural Information Processing Systems, 35:9125–9138, 2022
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40b42858-ba0f-4a7b-9fa0-e971e9b9fbf7 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 33173060-5025-41e1-881b-8a7c7d4b9041 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence An Open and Comprehensive Pipeline for Unified Object Grounding and Detection
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b7c7f1-5b07-4da6-a806-cea59bd7fcfc · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Llmdet: Learning strong open-vocabulary object detectors under the supervision of large language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a3b87be8-b678-47c9-b383-978b091bd266 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence General object foundation model for images and videos at scale
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de76a102-254c-4923-bf96-75629d0069b5 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Generalized decoding for pixel, image, and language
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7746fd4f-a1ec-43c3-a507-0e0c2ba0844a · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence A simple framework for open-vocabulary segmentation and detection
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d6d323-6993-4d3a-91ce-66b7c7f404d9 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Segment everything everywhere all at once
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 61cbbf9c-bd7e-40c8-97a6-5cf21afd95ff · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Open- worldsam: Extending sam2 for universal image segmentation with language prompts.arXiv preprint arXiv:2507.05427, 2025
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8d45cf05-1f07-4640-89a8-b4f3cd67e6fc · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Florence-2: Advancing a unified representation for a variety of vision tasks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 179f56a2-0934-43a0-b6ba-5795bb5577b5 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm Selection
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcdd51bf-6c3a-4e16-a828-df52d40381e8 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence LISA: Reasoning Segmentation via Large Language Model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18dc40db-1ea0-41e8-8a1c-5957d02ea833 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed570fb5-ef30-44b7-a58f-822870c20593 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Glamm: Pixel grounding large multimodal model
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aed20469-30a3-4f04-8de3-94dae3870694 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Pointrend: Image segmentation as rendering
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c85a48fb-934a-4c15-b4c3-4c7c9819ac59 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence V-net: Fully convolutional neural networks for volumetric medical image segmentation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6638bb39-258b-4120-8535-7e1d5ea2f75b · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36661101-d3c1-4bae-80ee-34586aece948 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10385fd9-fdb0-4493-a342-4b2ea053479b · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Lawrence Zitnick, and Piotr Dollár
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66deab6a-78d7-4abc-8dcb-e8ae0b93c640 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Coconut: Modernizing coco segmentation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1ede06cf-8bff-4a53-a624-56b8f28a7dca · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence High- quality entity segmentation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 88e66ab2-807d-44f6-90c5-0ff28ddb4407 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Semantic understanding of scenes through the ade20k dataset.International journal of computer vision, 127(3):302–321, 2019
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 173dff5c-6072-4a51-aab6-80fcf0136bba · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence GREC: Generalized Referring Expression Comprehension
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ee512f-4627-4f97-b5d8-1b16ba8376c4 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Introducing GPT-5.4
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dc15b24f-9320-4aac-810b-8490493cc188 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Segment anything
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6932e29d-e9b6-4979-baf5-59e71b848e38 · outbound
Vision-Language Grounding as Bidirectional Concept Correspondence Roboflow100-vl: A multi-domain object detection benchmark for vision-language models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.