Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:42:27.272900Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 100 of 115 outbound references and 2 inbound Pith citation observations for arXiv:2505.23766.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:42:27.272900Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T16:13:17.389147Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T05:53:04.514305Z
100 of 115 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a72b6a48-f8a9-4501-a2b2-7024d85d6d1e · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75820004-30b7-44bf-9221-da447c4e1d07 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d3d0c0-a5ce-48c0-ae50-c739bfd28ae4 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Graph of thoughts: Solving elab- orate problems with large language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7799f23-c9d2-4a05-b975-56c038072f1f · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought COYO-700M: Image-text pair dataset
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6c99e94-eb56-490b-a4b2-d9495ed1b9ac · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Image- level or object-level? A tale of two resampling strategies for long-tailed detection
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10317e8e-dfcd-4da6-b919-f54e55f45b5d · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Contrastive Localized Language-Image Pre-Training
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8037f4c1-041c-4ee4-9d2b-f3a5b0f11ff6 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a95082-1f06-46d3-864b-a9c4f7776493 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c50ad31-5af2-49ce-873c-eb9d295a5e35 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought ShareGPT4v: Improving large multi-modal models with better captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca842615-03ed-4778-92f8-614556338954 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59c73fe4-db18-43d4-abc6-fa74c5e38fe5 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought UNITER: Universal image-text representation learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f6b802d-337e-4a18-94f6-af9cec44b46a · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 956ba49b-7030-4de3-b51f-e8d1562315f1 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1e539e5-bd04-4def-9e5e-f38af3bcbc95 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Gonzalez, Ion Stoica, and Eric P
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52568861-4eb7-452f-8468-df21945151b1 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Control of goal- directed and stimulus-driven attention in the brain
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c272db46-e74b-4272-b31b-8a95f9dc2881 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought TransVG: End-to-end visual grounding with transformers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d8760e3-a772-4ce9-8264-2ec9ae37c54f · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought An im- age is worth 16x16 words: Transformers for image recog- nition at scale
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6163e6c-4dc8-4608-8eb2-4c3cd09f1bca · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought EV A: Exploring the limits of masked visual represen- tation learning at scale
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c022133-ab69-429a-8f09-0e77fde0f459 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought EV A-02: A visual representa- tion for neon genesis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 615865d7-9138-4cfe-9a17-53b48207f9f1 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Large-scale adversarial training for vision-and-language representation learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2667aa1f-2260-41a4-987a-8af23b0a4900 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a31c67fc-e199-477a-b0d8-47605887f98b · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Mini-InternVL: a flexible-transfer pocket multi-modal model with 5% parameters and 90% perfor- mance
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eacfc6a9-284e-4cb0-985e-91ae207c9166 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Chain of Thought Prompt Tuning in Vision Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cbf3642-e9bd-4ae4-acc9-f3748b916f3f · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought ICDAR2019 competition on scanned receipt ocr and information extrac- tion
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 349adb94-1f6a-424f-9076-136c1a00cd06 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought GQA: A new dataset for real-world visual reasoning and compositional question answering
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e254aacd-fa5a-4c7b-b740-0849780bae51 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Psychology, briefer course
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59bd9f19-fe59-477f-98c5-76be3a9913a1 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought DVQA: Understanding data visualizations via ques- tion answering
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a4c098-ce32-427b-ba92-0ce4f6a87e66 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought MDETR- modulated detection for end-to-end multi-modal under- standing
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e7df115-ee58-40b9-a260-a522e72ad03e · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Decoupling Representation and Classifier for Long-Tailed Recognition
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7114c00-4945-4a02-96f2-1a02a098d026 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Directed attention as a common resource for executive functioning and self- regulation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2317f3f8-8bfb-4905-9e93-18f907f9a5eb · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Referitgame: Referring to objects in pho- tographs of natural scenes
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1c4ba91-1f06-4c86-8008-525781db0295 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought A diagram is worth a dozen images
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cbfbb20-ff0e-437a-9603-3f7049fb3901 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Ocr- free document understanding transformer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ee7489-05f8-4daf-85be-fc395c2396bc · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Segment Anything
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be2f570d-589b-47ab-87a7-20bf277d305e · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Berg, Wan-Yen Lo, Piotr Doll´ar, and Ross Girshick
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac1fb692-0ba8-46cd-a09f-f0c988b4a917 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Large language models are zero-shot reasoners
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20799561-2b25-4297-b774-f23b9046f6e8 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Shamma, Michael S
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8b71fcb-f634-4abe-9d32-6391b8908497 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beb28aa1-9a29-4ccd-bda1-e105f543785c · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought LLaVA-OneVision: Easy Visual Task Transfer
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bccfa64-bd9d-4c4d-b3c9-a6d40fb91001 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efb08d6d-60ba-44cf-83e0-3faa755e423b · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da18befe-f9d2-4dc4-933e-62124fd12f8a · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23829ac-dcf1-4951-aabf-44575069f9ae · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought VILA: On pre-training for visual language models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c2a42b3-93f0-4925-a7fc-4999f7603b01 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Microsoft COCO: Common objects in context
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df271657-5cac-4b63-a74f-8e561cc63d9b · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6418badd-317f-4b71-8068-579b8eb9ca8f · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Visual spa- tial reasoning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1105ceec-6b3c-40ec-9ba3-f1d9f73e1d1c · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a30631d2-320c-440b-b3d4-c8afdfdd5783 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Visual instruction tuning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b34ff84-c214-423f-99b7-1a56fe1ad065 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge, 2024
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d56ce10-92a8-4789-9fc0-94d03b860da1 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5881f90f-c286-468e-8949-1e900d4945f1 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Grounding DINO: Mar- rying dino with grounded pre-training for open-set object detection
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e27013-fe4f-4292-a61f-e1f0ecc6adf3 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought A convnet for the 2020s
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f78a3c10-f95f-446a-869e-285295b71a53 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Decoupled Weight Decay Regularization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfbf9480-8d1b-454e-8066-530f202a7d25 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Learn to explain: Multimodal reason- ing via thought chains for science question answering
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4c01eca9-412f-40f5-8081-685f8a7bf18e · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4fb2db53-8f1c-44a6-9601-86d9b8792018 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Docvqa: A dataset for vqa on document images
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1a5fcb9f-4b0d-46a9-a983-818e9656c63e · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Infographicvqa
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d246c863-0a36-4ae5-a52f-ff926dd5d1d1 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 714df9be-92dc-4ca0-ad9b-584038d7a9eb · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Architecture of connectivity within a cingulo- fronto-parietal neurocognitive network for directed atten- tion
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2f6a914f-6e71-47db-b797-ef96405b1561 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Brain mechanisms for directed at- tention
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 319f97d0-691f-45aa-bcb7-0271ddc1840b · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b35e0369-055f-4f67-8cfb-a770b12e8cf8 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 927cd690-da00-4c1e-b525-7f40f5ab09fb · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb0ab443-ea3c-4c00-9353-97ead8d5b112 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6646bc2c-a9cc-4453-be73-5dbe9baa0669 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Learning transferable visual models from natural language supervision
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 01b6529b-ebe6-4211-ab78-517f32154a43 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought GLaMM: Pixel grounding large multimodal model
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8a1d9953-bb8d-4e1c-8889-22809e5d3b91 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a36ff0b-81f0-473e-9ef7-dac2d0e3d4f0 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Grounded sam: Assembling open-world models for diverse visual tasks,
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 991f54e3-8231-4670-bdee-c8beff1d6ab8 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32bec0cb-17a0-4933-a838-d9f1587817d4 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Vi- sual CoT: advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8ec44b3f-8926-450d-9f08-f198168a8025 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Objects365: A large-scale, high-quality dataset for object detection
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 554fe99f-b054-4dac-98e2-9f52f98d8e3d · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Eagle: Explor- ing the design space for multimodal llms with mixture of encoders
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ed165910-6713-4ec8-b29f-a8ee59c38191 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Textcaps: a dataset for image caption- ing with reading comprehension
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fb5507e2-77f1-465b-8f74-419354a6b629 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Towards VQA models that can read
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb1b8ea-d531-4f39-b438-6939dc318e59 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Introducing the next generation of claude,
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c5aea925-1146-484b-92a4-af1f85ab80a0 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e59f55fe-a678-4f02-839e-bc218451dfd3 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Emu3: Next-Token Prediction is All You Need
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1285999-0075-4185-9bd3-00990715af43 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Gemini: A Family of Highly Capable Multimodal Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91279993-fbe2-4d3b-9699-8735a99f17dc · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 485ae442-cff4-4a08-9f52-2f58ad1cbb9e · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Laion-gpt4v dataset, 2023
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e16d7b99-ff67-4eea-b2f3-cdf8f51790c2 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought The Llama 3 Herd of Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ce2099-2d64-44d9-9eb9-34ae8cc575df · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought GPT-4 Technical Report
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0259e96-31a0-4aa4-98b2-17eb04358c28 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought InternVL2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4ffdfbc1-cc0d-449e-9556-2dfa81042bec · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants., 2023
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b54b13ef-7543-4d9b-bc37-a5b6a7637745 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb15478-74b7-4f6c-9893-7cbd6890e848 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Eyes wide shut? exploring the visual shortcomings of multimodal LLMs
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d7487d99-fe6f-4567-ab66-d45c5826f27a · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Document understanding dataset and evalu- ation (dude)
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d43d6722-028f-4a63-ac3c-fb3f92504af4 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Unresolved cited work
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1df6994e-e983-4736-9c96-d1dc5913f244 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d422d969-7941-489f-8f72-e2db29a34e9f · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 15d4a4c7-fcf2-4843-b34e-87dac03b76f4 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f4a3485-343b-4250-bec0-8da2edfa7963 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought CogVLM: Visual Expert for Pretrained Language Models
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d80b74-8151-4e78-ba1d-196f8aabaee6 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 81064563-fd71-4858-882d-0404f0a59923 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Chain-of-thought prompting elicits reasoning in large language models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d4241a76-04df-4c98-be96-d1a6782c228f · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a433373-0949-424f-8f61-013d8088efda · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought V*: Guided visual search as a core mechanism in multimodal LLMs
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 14542e57-6b1b-4716-baf2-fca9bf3c314e · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Grok, 2024
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8cfba54f-0f0b-4c89-bef2-f1b4974ecdf7 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Florence-2: Advancing a unified representation for a va- riety of vision tasks
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9a22ca37-65b7-40f7-a567-2579d68079d0 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Denoising vision trans- formers
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a1a23623-bcb0-4da3-874e-c340119c1c88 · outbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought UniTAB: Unifying text and box outputs for grounded vision-language modeling
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation eddf2992-4d4f-48c2-b205-5490f9d40a25 · inbound
Vision Harnessing Agent for Open Ad-hoc Segmentation Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 05d5a947-1c4f-410d-a019-dc3e9c4702cf · inbound
OPLD: On-Policy Latent Distillation for Multimodal Reasoning Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.