Pith. sign in

Paper Citation Record · LEDGER

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

As of 20 August 2026, this Paper Citation Record lists 100 of 115 outbound references and 2 inbound Pith citation observations for arXiv:2505.23766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23766 v1

Coverage vector

measured 100 of 115 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:42:27.272900Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T16:13:17.389147Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T05:53:04.514305Z

Reference resolution

100 of 115 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved73
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a72b6a48-f8a9-4501-a2b2-7024d85d6d1e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:19.402627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:19.402627Z digest=sha256:45282dd3958ba2142cef067d91d6117c907fa2da4bb1df97847dd7666f1fffce

Observation 75820004-30b7-44bf-9221-da447c4e1d07 · outbound

This paper cites Qwen2.5-VL Technical Report.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:19.444526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:19.444526Z digest=sha256:37279fd702c0c56505ba20a3ae61d52359ca5ea61721abf81ce7005e79b83bce

Observation 88d3d0c0-a5ce-48c0-ae50-c739bfd28ae4 · outbound

This paper cites Graph of thoughts: Solving elab- orate problems with large language models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Graph of thoughts: Solving elab- orate problems with large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:19.520327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:19.520327Z digest=sha256:069b280fee99219d35200e7cceac7b03ce6807f79cf3914e971cd28d8ce945fd

Observation f7799f23-c9d2-4a05-b975-56c038072f1f · outbound

This paper cites COYO-700M: Image-text pair dataset.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought COYO-700M: Image-text pair dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:19.601324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:19.601324Z digest=sha256:ee23dcadad40224939cbe9f3f65e8b24afeaed6dc978d9680b2c8d3effb7ac92

Observation f6c99e94-eb56-490b-a4b2-d9495ed1b9ac · outbound

This paper cites Image- level or object-level? A tale of two resampling strategies for long-tailed detection.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Image- level or object-level? A tale of two resampling strategies for long-tailed detection

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:19.671434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:19.671434Z digest=sha256:48b3383f4f4c93affa7df68f8d517dd0ad313d37f01d6e3807f6d43b06091826

Observation 10317e8e-dfcd-4da6-b919-f54e55f45b5d · outbound

This paper cites Contrastive Localized Language-Image Pre-Training.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Contrastive Localized Language-Image Pre-Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:19.724749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:19.724749Z digest=sha256:d3c190ec1cd6c2049b4c9a69929324b4fb73a52566c41f79f99d6e5ae46647fe

Observation 8037f4c1-041c-4ee4-9d2b-f3a5b0f11ff6 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:19.818075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:19.818075Z digest=sha256:ab0dbb44b83f4da47a0ac6d51bc66861c735bbc7e20cfc7afa09c6ba018a6b9f

Observation e2a95082-1f06-46d3-864b-a9c4f7776493 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:19.926696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:19.926696Z digest=sha256:29e9fb7e8c16542c014b29562b27165da7c86354c9599f138680c6dc2faaf280

Observation 3c50ad31-5af2-49ce-873c-eb9d295a5e35 · outbound

This paper cites ShareGPT4v: Improving large multi-modal models with better captions.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought ShareGPT4v: Improving large multi-modal models with better captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:20.016722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:20.016722Z digest=sha256:027fbbaafaddbe96c0e48def63b4637db8379f789b7c6145a68142cc93cb67a0

Observation ca842615-03ed-4778-92f8-614556338954 · outbound

This paper cites an unresolved cited work.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:20.104403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:20.104403Z digest=sha256:d3ab89dd4439e90628cd45158753e1f709a2da0b65977ae569b3f3659a15b149

Observation 59c73fe4-db18-43d4-abc6-fa74c5e38fe5 · outbound

This paper cites UNITER: Universal image-text representation learning.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought UNITER: Universal image-text representation learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:20.109699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:20.109699Z digest=sha256:0609cfe6b82ac9c39067d488445333365de1b620a522be24a24035e7122e9ba2

Observation 1f6b802d-337e-4a18-94f6-af9cec44b46a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:20.115332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:20.115332Z digest=sha256:e5193869ab5aa829a320280ba53758730b62e6cff5ffdeabeac633084f9ca572

Observation 956ba49b-7030-4de3-b51f-e8d1562315f1 · outbound

This paper cites InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:20.167111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:20.167111Z digest=sha256:18bfb78744e9e67a45d042e433c72d102a25f057df95c59c4a6cc840914132e1

Observation e1e539e5-bd04-4def-9e5e-f38af3bcbc95 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Gonzalez, Ion Stoica, and Eric P

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:20.293659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:20.293659Z digest=sha256:3852bc8427bb76a8fbf568be232d3a988da5e70f74bf37ca714016af237a6717

Observation 52568861-4eb7-452f-8468-df21945151b1 · outbound

This paper cites Control of goal- directed and stimulus-driven attention in the brain.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Control of goal- directed and stimulus-driven attention in the brain

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:20.465076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:20.465076Z digest=sha256:33e4c73efe926f9010dbf524d20de575ee2db7dcbd5f708b69d482f98d3eeb05

Observation c272db46-e74b-4272-b31b-8a95f9dc2881 · outbound

This paper cites TransVG: End-to-end visual grounding with transformers.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought TransVG: End-to-end visual grounding with transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:20.584929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:20.584929Z digest=sha256:c612bc0f1595dcde26ce6025c549309960b9900bb4c498382efadcf89859fb94

Observation 0d8760e3-a772-4ce9-8264-2ec9ae37c54f · outbound

This paper cites An im- age is worth 16x16 words: Transformers for image recog- nition at scale.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought An im- age is worth 16x16 words: Transformers for image recog- nition at scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:20.705927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:20.705927Z digest=sha256:f701b7499168e7417529b85e1e60c990179fdbf0be32ebf3a83125cf22c1523b

Observation f6163e6c-4dc8-4608-8eb2-4c3cd09f1bca · outbound

This paper cites EV A: Exploring the limits of masked visual represen- tation learning at scale.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought EV A: Exploring the limits of masked visual represen- tation learning at scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:20.838022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:20.838022Z digest=sha256:8c39b1ab5b586484c1d8d7d6d2100add616d4e8df23ac4e6912d5c56b3bd4ce0

Observation 7c022133-ab69-429a-8f09-0e77fde0f459 · outbound

This paper cites EV A-02: A visual representa- tion for neon genesis.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought EV A-02: A visual representa- tion for neon genesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:21.037629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:21.037629Z digest=sha256:3b13683313f8c5af82a6ac75b4379ecd9940bf5ebcbefc1e6d10227b7469a20d

Observation 615865d7-9138-4cfe-9a17-53b48207f9f1 · outbound

This paper cites Large-scale adversarial training for vision-and-language representation learning.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Large-scale adversarial training for vision-and-language representation learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:21.222589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:21.222589Z digest=sha256:8637977e3702a3682ef5a11d6febb952b823989870fdf4141d5c7e4f73b14ffb

Observation 2667aa1f-2260-41a4-987a-8af23b0a4900 · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:21.316826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:21.316826Z digest=sha256:5e96e2a5fe1fd99e5b0b56c5c863bdf17961ebe82a4290a53df37605bca91e0b

Observation a31c67fc-e199-477a-b0d8-47605887f98b · outbound

This paper cites Mini-InternVL: a flexible-transfer pocket multi-modal model with 5% parameters and 90% perfor- mance.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Mini-InternVL: a flexible-transfer pocket multi-modal model with 5% parameters and 90% perfor- mance

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:21.428454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:21.428454Z digest=sha256:8c02ccf5e71b049e842fe9a7e95067a412af0cac2b91d51ad6f9e2c3640ad776

Observation eacfc6a9-284e-4cb0-985e-91ae207c9166 · outbound

This paper cites Chain of Thought Prompt Tuning in Vision Language Models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Chain of Thought Prompt Tuning in Vision Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:21.531787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:21.531787Z digest=sha256:1cb9dff14cc704ecfc5edf1227f13545e328ddb1ae97fa4d55dd57f870f61a76

Observation 7cbf3642-e9bd-4ae4-acc9-f3748b916f3f · outbound

This paper cites ICDAR2019 competition on scanned receipt ocr and information extrac- tion.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought ICDAR2019 competition on scanned receipt ocr and information extrac- tion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:21.668887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:21.668887Z digest=sha256:427109f30ab531643cfdc29adc04ab88da4b65900e0155b64ef27f433274ecc5

Observation 349adb94-1f6a-424f-9076-136c1a00cd06 · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought GQA: A new dataset for real-world visual reasoning and compositional question answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:21.749739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:21.749739Z digest=sha256:e3912fa04f6fdd334a6d4cae6f80dcc92eb69134695ea547bf6452caed8e8d5e

Observation e254aacd-fa5a-4c7b-b740-0849780bae51 · outbound

This paper cites Psychology, briefer course.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Psychology, briefer course

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:21.843798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:21.843798Z digest=sha256:7c66e8cf3bda98c15db615f22ed375e6a5c0e8fc0ada17f4d313628947286675

Observation 59bd9f19-fe59-477f-98c5-76be3a9913a1 · outbound

This paper cites DVQA: Understanding data visualizations via ques- tion answering.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought DVQA: Understanding data visualizations via ques- tion answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:21.910492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:21.910492Z digest=sha256:8a294793d1f590fc70b6b5854fbf387995baab1cd003cd3488c0d80c646dc31d

Observation d1a4c098-ce32-427b-ba92-0ce4f6a87e66 · outbound

This paper cites MDETR- modulated detection for end-to-end multi-modal under- standing.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought MDETR- modulated detection for end-to-end multi-modal under- standing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:21.952695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:21.952695Z digest=sha256:ae91cda67c358c3f91a45969b0d72f5cc5fe8d4c98fdba04b40d23f0ba07daa3

Observation 2e7df115-ee58-40b9-a260-a522e72ad03e · outbound

This paper cites Decoupling Representation and Classifier for Long-Tailed Recognition.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Decoupling Representation and Classifier for Long-Tailed Recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:21.977055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:21.977055Z digest=sha256:15d5e237979f4b8c7c9b19ed025231dcc2a2b9b2a39128aa28ee7cd68e7fa859

Observation d7114c00-4945-4a02-96f2-1a02a098d026 · outbound

This paper cites Directed attention as a common resource for executive functioning and self- regulation.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Directed attention as a common resource for executive functioning and self- regulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.022331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.022331Z digest=sha256:7d052d2ca6f0e2855d2d00a9101cd20610c089ec4da44eb44042d733c264f3bf

Observation 2317f3f8-8bfb-4905-9e93-18f907f9a5eb · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.115768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.115768Z digest=sha256:bb143455a0aef66cd7fd38de9b02f90d8fba3c7e8b2dcc8bf6786332255177dd

Observation f1c4ba91-1f06-4c86-8008-525781db0295 · outbound

This paper cites A diagram is worth a dozen images.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought A diagram is worth a dozen images

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.148962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.148962Z digest=sha256:4e5c3ba6d4704a4b7aaebfbfb21cc8646953d58ccad3f9be1b1d6735448214ba

Observation 4cbfbb20-ff0e-437a-9603-3f7049fb3901 · outbound

This paper cites Ocr- free document understanding transformer.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Ocr- free document understanding transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.220749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.220749Z digest=sha256:bdb1dd72169c7781f4c23b5e474a2257edd331d308e7cc9b66ca92a88f701bb6

Observation 88ee7489-05f8-4daf-85be-fc395c2396bc · outbound

This paper cites Segment Anything.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Segment Anything

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.248371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.248371Z digest=sha256:b9cf057e6f325f3bfdbacb88f0fbc284dd4f306e959e624b07c68114e1b8fad3

Observation be2f570d-589b-47ab-87a7-20bf277d305e · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Doll´ar, and Ross Girshick.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Berg, Wan-Yen Lo, Piotr Doll´ar, and Ross Girshick

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.283486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.283486Z digest=sha256:5d551ce4d014f8203273e46d81a05ea01f7a044f7e175437c9ee3f33ba0a14cd

Observation ac1fb692-0ba8-46cd-a09f-f0c988b4a917 · outbound

This paper cites Large language models are zero-shot reasoners.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Large language models are zero-shot reasoners

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.340938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.340938Z digest=sha256:9421374d4abe58fd313dd8be09b46cd6ea60eef4c9c07d7c86b43320da39a403

Observation 20799561-2b25-4297-b774-f23b9046f6e8 · outbound

This paper cites Shamma, Michael S.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Shamma, Michael S

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.399259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.399259Z digest=sha256:b9c4ad806372035fce1553f9e27a909599a362909bfec93324193b9668418f8a

Observation e8b71fcb-f634-4abe-9d32-6391b8908497 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.482501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.482501Z digest=sha256:2cf494ce35946a998e51c20b7e0b7d88145a7c5e5393b0ffae90495ca70bbe54

Observation beb28aa1-9a29-4ccd-bda1-e105f543785c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought LLaVA-OneVision: Easy Visual Task Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.562053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.562053Z digest=sha256:007aa8725628e17f043c7aac3af63919dc87d6c45075ad4da1686a5fc0f42f59

Observation 0bccfa64-bd9d-4c4d-b3c9-a6d40fb91001 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.626506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.626506Z digest=sha256:94ef2258ef00ed822da62bf06874b07e4b24bf89a2e33747063f70077d886dfe

Observation efb08d6d-60ba-44cf-83e0-3faa755e423b · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.713377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.713377Z digest=sha256:325afb0a9fc8eb529f9077d5c411ce1e628e987fe8014d0abba40a5c5d740bba

Observation da18befe-f9d2-4dc4-933e-62124fd12f8a · outbound

This paper cites VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.824876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.824876Z digest=sha256:e167a55df644f4defec15f151f3aabb2692e038e32e6c677107d61ac8dc3322f

Observation c23829ac-dcf1-4951-aabf-44575069f9ae · outbound

This paper cites VILA: On pre-training for visual language models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought VILA: On pre-training for visual language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:22.868037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:22.868037Z digest=sha256:a2276880f6f913300af1bc07ab2873191d3d4d328e6f847d66704d2ec355cc1d

Observation 4c2a42b3-93f0-4925-a7fc-4999f7603b01 · outbound

This paper cites Microsoft COCO: Common objects in context.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Microsoft COCO: Common objects in context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.031778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.031778Z digest=sha256:ba208d981c818ad094026dbd41b9d105da85d6aae877f9bfbb0783c46d043078

Observation df271657-5cac-4b63-a74f-8e561cc63d9b · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.171837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.171837Z digest=sha256:50066eea8eb6f090d39cbe4a8aaba6bd75abd169320824f6c144c9919a19f05f

Observation 6418badd-317f-4b71-8068-579b8eb9ca8f · outbound

This paper cites Visual spa- tial reasoning.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Visual spa- tial reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.295233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.295233Z digest=sha256:deeb073bb5ad9501ea5f79ec62604c694d238e8be05f15a68ed8f9346fe3ceb6

Observation 1105ceec-6b3c-40ec-9ba3-f1d9f73e1d1c · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.389391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.389391Z digest=sha256:ceec3132e938c225daa43ef721a7f5c76aa6bbcadc4536326910483ae39c9bbf

Observation a30631d2-320c-440b-b3d4-c8afdfdd5783 · outbound

This paper cites Visual instruction tuning.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Visual instruction tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.469711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.469711Z digest=sha256:e22c0f5e7d3012afabffcc7f4007d2ffb374971d1c5aa5364f671a22f7731cd0

Observation 1b34ff84-c214-423f-99b7-1a56fe1ad065 · outbound

This paper cites LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge, 2024.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.492051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.492051Z digest=sha256:4615707358846e5954864e1620ff9fc8f67697274fe39076c059ac174e463f68

Observation 5d56ce10-92a8-4789-9fc0-94d03b860da1 · outbound

This paper cites LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.566141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.566141Z digest=sha256:93f759342b4b8ad21d1e5f1ebb1b550282f3812bcef60eca715dec057995114f

Observation 5881f90f-c286-468e-8949-1e900d4945f1 · outbound

This paper cites Grounding DINO: Mar- rying dino with grounded pre-training for open-set object detection.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Grounding DINO: Mar- rying dino with grounded pre-training for open-set object detection

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.622303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.622303Z digest=sha256:24ec3168d4b2be04a939c9abcb5c5cf3b42339d57b45afd7a9450368b691411a

Observation 68e27013-fe4f-4292-a61f-e1f0ecc6adf3 · outbound

This paper cites A convnet for the 2020s.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought A convnet for the 2020s

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:34.816112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:23.712415Z digest=sha256:a5d702b250547b3488f1f59e7ca656262a036feda2dd0c0f8dd0fab69849aa4e

Observation f78a3c10-f95f-446a-869e-285295b71a53 · outbound

This paper cites Decoupled Weight Decay Regularization.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Decoupled Weight Decay Regularization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.786924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.786924Z digest=sha256:580e4dc020e48aebb7fa16699952f10f08929ae0295ccb2a9c14f6b97b67de56

Observation cfbf9480-8d1b-454e-8066-530f202a7d25 · outbound

This paper cites Learn to explain: Multimodal reason- ing via thought chains for science question answering.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Learn to explain: Multimodal reason- ing via thought chains for science question answering

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:34.629037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:23.819892Z digest=sha256:aa9f71cda741937ce2b5bc1b0093595885804c8ced6547bc89e87000391bd61e

Observation 4c01eca9-412f-40f5-8081-685f8a7bf18e · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:34.541033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:23.845864Z digest=sha256:7acde175c5009ae16c03de1fe4f5adc1970b667d035a8f1fa4db631c74d7453e

Observation 4fb2db53-8f1c-44a6-9601-86d9b8792018 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Docvqa: A dataset for vqa on document images

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:34.411493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:23.869082Z digest=sha256:9b3010a8e165000800ecd9634d9beac50a6ee58a521ef5807be190503a4af478

Observation 1a5fcb9f-4b0d-46a9-a983-818e9656c63e · outbound

This paper cites Infographicvqa.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Infographicvqa

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:34.312367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:23.903360Z digest=sha256:95cbf7d76be12dd1386688cc2ddde838681c5b88853a45dd0abe1ee10a8b0623

Observation d246c863-0a36-4ae5-a52f-ff926dd5d1d1 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:23.938819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:23.938819Z digest=sha256:0857408fe054e76ed77a10cde02e33fa22f27ee6378fd31f61e49623458bbdff

Observation 714df9be-92dc-4ca0-ad9b-584038d7a9eb · outbound

This paper cites Architecture of connectivity within a cingulo- fronto-parietal neurocognitive network for directed atten- tion.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Architecture of connectivity within a cingulo- fronto-parietal neurocognitive network for directed atten- tion

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:34.232482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:23.992739Z digest=sha256:6f0ad6b7b08f4013c7edbacc3acf4a79a38fa998cdfffeea269370002c2468a3

Observation 2f6a914f-6e71-47db-b797-ef96405b1561 · outbound

This paper cites Brain mechanisms for directed at- tention.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Brain mechanisms for directed at- tention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:34.164846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:24.170796Z digest=sha256:ff7f68e9f462fd04cb392d3f088dd93be18fee4c89eebfadf47c037a4715e535

Observation 319f97d0-691f-45aa-bcb7-0271ddc1840b · outbound

This paper cites Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:24.327499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:24.327499Z digest=sha256:94bb61d468b5541cf46690df461636b3db8413952fcb52e7bbd1a0f639d5075f

Observation b35e0369-055f-4f67-8cfb-a770b12e8cf8 · outbound

This paper cites an unresolved cited work.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:34.102582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:24.467862Z digest=sha256:f515eff529768b5483061f8a5f3e89a29a83aa36356b6cdf4671a69b9092fbf4

Observation 927cd690-da00-4c1e-b525-7f40f5ab09fb · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:24.586170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:24.586170Z digest=sha256:4d1a0cff8a7d7fa58778d6d7ad0a6746508df3b1afbbdde325a12e65242e4bf9

Observation bb0ab443-ea3c-4c00-9353-97ead8d5b112 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:24.625803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:24.625803Z digest=sha256:5baaade37d6221e0e7049ee61523eb6f0a87bf44dd16ffdffee464494e90610d

Observation 6646bc2c-a9cc-4453-be73-5dbe9baa0669 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Learning transferable visual models from natural language supervision

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:34.016050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:24.723342Z digest=sha256:959386e711141b24bf6e386c9d11befee6f6d2a38dcf0be048bfa02783e3b458

Observation 01b6529b-ebe6-4211-ab78-517f32154a43 · outbound

This paper cites GLaMM: Pixel grounding large multimodal model.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought GLaMM: Pixel grounding large multimodal model

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:33.902621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:24.748882Z digest=sha256:f311d271e981a5ce26b01f4ee5fe4e54dff52d3bdb383f4aa7c16a86dc1a6429

Observation 8a1d9953-bb8d-4e1c-8889-22809e5d3b91 · outbound

This paper cites Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:24.889360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:24.889360Z digest=sha256:efe067fb1f289e42f0754143470a8f4b4ec3a6f38ec5755f8ed3972863683fa3

Observation 9a36ff0b-81f0-473e-9ef7-dac2d0e3d4f0 · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks,.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Grounded sam: Assembling open-world models for diverse visual tasks,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:24.956553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:24.956553Z digest=sha256:598f7778f3dcebf8847d4d461de1b14da5ab5f1a8f7a60ed8cb87933e0528793

Observation 991f54e3-8231-4670-bdee-c8beff1d6ab8 · outbound

This paper cites Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:25.021502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:25.021502Z digest=sha256:bf2f51d7a4b2fe646b0c54cdb2f8aff8167098e39108e03ee94b24b8594193b1

Observation 32bec0cb-17a0-4933-a838-d9f1587817d4 · outbound

This paper cites Vi- sual CoT: advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Vi- sual CoT: advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:33.773354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:25.111107Z digest=sha256:f77fdd9a5d425e05562942ef69ddd6b435ed5bd1a1b033d9597096c3a337ea20

Observation 8ec44b3f-8926-450d-9f08-f198168a8025 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Objects365: A large-scale, high-quality dataset for object detection

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:33.613986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:25.224907Z digest=sha256:10bd118d182733d1631e03aa7f1153f1dcb3aa5a550f21a0b7eb3491651e5187

Observation 554fe99f-b054-4dac-98e2-9f52f98d8e3d · outbound

This paper cites Eagle: Explor- ing the design space for multimodal llms with mixture of encoders.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Eagle: Explor- ing the design space for multimodal llms with mixture of encoders

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:33.385784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:25.314795Z digest=sha256:bded1dae8788b07801886d3378e4f1cf21d6b66b344bc7ae45d45ec0ce0f9132

Observation ed165910-6713-4ec8-b29f-a8ee59c38191 · outbound

This paper cites Textcaps: a dataset for image caption- ing with reading comprehension.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Textcaps: a dataset for image caption- ing with reading comprehension

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:33.205047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:25.373716Z digest=sha256:fa2c597327eae26d4e317bd4ac809c4aa925a8269791e5c3999f8a9f7a0a2b7d

Observation fb5507e2-77f1-465b-8f74-419354a6b629 · outbound

This paper cites Towards VQA models that can read.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Towards VQA models that can read

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:25.446368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:25.446368Z digest=sha256:adee38cea92d5efca4e0fffca0ddd7419f83a1161fda867f6ffb58c244ccaf16

Observation fcb1b8ea-d531-4f39-b438-6939dc318e59 · outbound

This paper cites Introducing the next generation of claude,.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Introducing the next generation of claude,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:33.000409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:25.485107Z digest=sha256:2f96e1984c4c7facf4b0046703385e8118bd9f197fc2f3c2b9e882fccf07c416

Observation c5aea925-1146-484b-92a4-af1f85ab80a0 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:25.628926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:25.628926Z digest=sha256:7345225ed32a994a289ca18cdae36f69e39df485a884dbea074a54ad9ce2ce66

Observation e59f55fe-a678-4f02-839e-bc218451dfd3 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Emu3: Next-Token Prediction is All You Need

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:25.735873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:25.735873Z digest=sha256:3b93a8a44d06ad30428ea3f09f31a9b967ae960e0ffbc64a90953af6110dd858

Observation c1285999-0075-4185-9bd3-00990715af43 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Gemini: A Family of Highly Capable Multimodal Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:25.785902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:25.785902Z digest=sha256:6804d93c93f0966d5f0364835bde1b37220254a744b8e6d1ce57747ff5018b6a

Observation 91279993-fbe2-4d3b-9699-8735a99f17dc · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:25.842990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:25.842990Z digest=sha256:2ebee1927d4b7892c29ba292bd8d6fb634000fa68b1f9dec2c61f70e78a40d62

Observation 485ae442-cff4-4a08-9f52-2f58ad1cbb9e · outbound

This paper cites Laion-gpt4v dataset, 2023.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Laion-gpt4v dataset, 2023

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:32.814620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:25.885829Z digest=sha256:8b773ba4ac74d829247a982d09ebd87023473887c6f003769413991d917d18de

Observation e16d7b99-ff67-4eea-b2f3-cdf8f51790c2 · outbound

This paper cites The Llama 3 Herd of Models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought The Llama 3 Herd of Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:25.906540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:25.906540Z digest=sha256:87447a02e53bec9bf870b81587514cb6388a67f78967f00d6601efdba4ffdfea

Observation a0ce2099-2d64-44d9-9eb9-34ae8cc575df · outbound

This paper cites GPT-4 Technical Report.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought GPT-4 Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:25.951435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:25.951435Z digest=sha256:a64267e73e71dc5f5a29bd06543acb248fe9216277f0c2d551fe8dcc81c2345b

Observation a0259e96-31a0-4aa4-98b2-17eb04358c28 · outbound

This paper cites InternVL2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought InternVL2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:32.700118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:26.014272Z digest=sha256:1e25cc6a500f261a3d46ab7992d2a31382b91edba24caf63f3f7d33b80901417

Observation 4ffdfbc1-cc0d-449e-9556-2dfa81042bec · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants., 2023.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants., 2023

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:32.638127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:26.107864Z digest=sha256:d7d9c5e6404b166e983a35ba6c32b0b8aa472a98b5b55939599b9de5ed4a0a13

Observation b54b13ef-7543-4d9b-bc37-a5b6a7637745 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:26.229952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:26.229952Z digest=sha256:f63b29b93119e1a2245b3f45da148e87629d29c748b128dd5b6b83ff6774e851

Observation 2eb15478-74b7-4f6c-9893-7cbd6890e848 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal LLMs.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Eyes wide shut? exploring the visual shortcomings of multimodal LLMs

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:32.551567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:26.323211Z digest=sha256:ad28e55b542f9197720e449b984170f102a1a5ec2ad8bd49159d89cb94ca1310

Observation d7487d99-fe6f-4567-ab66-d45c5826f27a · outbound

This paper cites Document understanding dataset and evalu- ation (dude).

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Document understanding dataset and evalu- ation (dude)

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:32.424203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:26.400879Z digest=sha256:ed6011983004eb37f3cc8d92cc5f8965e8390e4bb8ee84a82e56f371ce198883

Observation d43d6722-028f-4a63-ac3c-fb3f92504af4 · outbound

This paper cites an unresolved cited work.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:42:32.278977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:26.469873Z digest=sha256:f13578fbff8344a73fc3a072307df6773b8616f0f362ff8894be29b9286ac728

Observation 1df6994e-e983-4736-9c96-d1dc5913f244 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:26.521528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:26.521528Z digest=sha256:12355f047f3e9963c4c3a7348631c4015cc6635869eefc03d5c8a8e0c193d6f0

Observation d422d969-7941-489f-8f72-e2db29a34e9f · outbound

This paper cites OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:32.123230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:26.602116Z digest=sha256:2aa445326755e08e0cae2fe241d076f069023cac08d99c143a6999c0acbb966d

Observation 15d4a4c7-fcf2-4843-b34e-87dac03b76f4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:26.691696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:26.691696Z digest=sha256:afaee309b76b3aae7bef7968965a4456dc4a8e24693cadbe1a89f7dec0747894

Observation 0f4a3485-343b-4250-bec0-8da2edfa7963 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought CogVLM: Visual Expert for Pretrained Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:26.766875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:26.766875Z digest=sha256:91e5a246acf37aa4252de33fe382b4da18ea15a8e70eb9d80601ed9581f09841

Observation 64d80b74-8151-4e78-ba1d-196f8aabaee6 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:31.898779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:26.832054Z digest=sha256:54cea03476b1e519a8cd34f09d71b5cc61d786000e41ca86358d95c37550434f

Observation 81064563-fd71-4858-882d-0404f0a59923 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Chain-of-thought prompting elicits reasoning in large language models

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:31.628673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:26.934159Z digest=sha256:cdbf52799bca75a10605a4c6278380086222d851c46cc9d511a09619be0a924d

Observation d4241a76-04df-4c98-be96-d1a6782c228f · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:27.031748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:27.031748Z digest=sha256:7d7c2c4032844ab6a4a754ed36998fc9f41d6e86934cc610f1413d8fa7cc2c6d

Observation 3a433373-0949-424f-8f61-013d8088efda · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal LLMs.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought V*: Guided visual search as a core mechanism in multimodal LLMs

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:31.492476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:27.074230Z digest=sha256:8e5d65db931277a46aaaa46c0bf11e024e7007044f82dcb6b953c1ec2d7ae61a

Observation 14542e57-6b1b-4716-baf2-fca9bf3c314e · outbound

This paper cites Grok, 2024.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Grok, 2024

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:31.326043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:27.113583Z digest=sha256:680b19cdcf8064815c04b8f2bb8770959f476c8a03e107c051737e5140ae07c4

Observation 8cfba54f-0f0b-4c89-bef2-f1b4974ecdf7 · outbound

This paper cites Florence-2: Advancing a unified representation for a va- riety of vision tasks.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Florence-2: Advancing a unified representation for a va- riety of vision tasks

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:31.210669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:27.156407Z digest=sha256:06a77ee464878ecf2726f686dc3b21b9fe183d2593f3a692a1533a1f4acc46fa

Observation 9a22ca37-65b7-40f7-a567-2579d68079d0 · outbound

This paper cites Denoising vision trans- formers.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Denoising vision trans- formers

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:31.042066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:27.195113Z digest=sha256:7452309be185f786bf8a6807e58b5a80dbe21b6d87256b5abd40f3a717d71bba

Observation a1a23623-bcb0-4da3-874e-c340119c1c88 · outbound

This paper cites UniTAB: Unifying text and box outputs for grounded vision-language modeling.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought UniTAB: Unifying text and box outputs for grounded vision-language modeling

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:42:30.851692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:42:27.272900Z digest=sha256:f2fcc84ef2cab9f4305d0b2e38d8a0ad484fe03fd4ff1b81f5b501f200524e5d

Pith citing papers

Observation eddf2992-4d4f-48c2-b205-5490f9d40a25 · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.515869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:1379d63a94cbb9e5d582ddc21503c44b57026a99f004a9f868f64f125555c9d0

Observation 05d5a947-1c4f-410d-a019-dc3e9c4702cf · inbound

OPLD: On-Policy Latent Distillation for Multimodal Reasoning cites this paper.

OPLD: On-Policy Latent Distillation for Multimodal Reasoning Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T16:13:17.389147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T16:13:17.389147Z digest=sha256:36bcc287de3a82088aa87adbced2aa08d5646a3754b6718264e6909b8e489b47