Pith. sign in

Paper Citation Record · LEDGER

Visual Instruction Tuning with Chain of Region-of-Interest

As of 19 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2505.06840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.06840 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:37:17.470659Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy41
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a287be2-67bd-4035-b723-c1121d674098 · outbound

This paper cites https://cdn.openai.com/papers/GPTV_System_Card.pdf, 2023.

Visual Instruction Tuning with Chain of Region-of-Interest https://cdn.openai.com/papers/GPTV_System_Card.pdf, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.602767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.155881Z digest=sha256:36ab89c57ceddfe61aa372b9ebcda6b06ffc7f6817e7db2e767f6972c073b174

Observation 4ac889ba-2a7b-4497-969e-7d02ebbdda3e · outbound

This paper cites https://sharegpt.com, 2023.

Visual Instruction Tuning with Chain of Region-of-Interest https://sharegpt.com, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.587026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.161282Z digest=sha256:76b1773f08af3a45ab9259495e6a63e192e75cc09e88345707e2d0ce5f9a6578

Observation 9d324040-7351-40d1-8e58-4d8d6bbb40b2 · outbound

This paper cites https://www-cdn.anthropic.com/ de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf, 2024.

Visual Instruction Tuning with Chain of Region-of-Interest https://www-cdn.anthropic.com/ de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.569607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.166532Z digest=sha256:848da700176befb0b2be8043cf7389bb7c2cbd40ffff302eec4bbac58a48dd4f

Observation 4fc095c5-bcce-4362-ac18-254801f0d5d3 · outbound

This paper cites https://huggingface.co/datasets/laion/gpt4v-dataset, 2024.

Visual Instruction Tuning with Chain of Region-of-Interest https://huggingface.co/datasets/laion/gpt4v-dataset, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.552767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.171579Z digest=sha256:24b5559c95eaa4894865acc19f49a2fc9a66b031105b9ca6eab12339cebd384d

Observation a5e378aa-f1df-44ce-bbd6-1d0f8620ec8b · outbound

This paper cites https://huggingface.co/NousResearch/Nous-Hermes-2-Yi-34B , 2024.

Visual Instruction Tuning with Chain of Region-of-Interest https://huggingface.co/NousResearch/Nous-Hermes-2-Yi-34B , 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.537366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.176151Z digest=sha256:135e3dfcfc6dd74e4ebe9cb96cf323327502195f58ff45e38ed49a7d87406a79

Observation b571747b-821f-4970-ab66-a7873f24cd97 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Visual Instruction Tuning with Chain of Region-of-Interest Flamingo: a visual language model for few-shot learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.521160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.180913Z digest=sha256:2140824839d77e11ed3b40ff67703956b315f240aae4f00191a1f5e19e295485

Observation 989e12f3-2a98-4604-98f0-572d2c14524f · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Visual Instruction Tuning with Chain of Region-of-Interest Lawrence Zitnick, and Devi Parikh

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.502876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.185939Z digest=sha256:e2373ef1977eedc5391c63db4fa69161335c42171b1f1376dcafe2f777d8e5a9

Observation 3ac63604-86a3-4018-8483-697a7a9175e6 · outbound

This paper cites Openflamingo, Mar.

Visual Instruction Tuning with Chain of Region-of-Interest Openflamingo, Mar

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.484249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.190532Z digest=sha256:7b9e408d4f53ac03262f11568338dcd8e6cf20841df4e937fba9ce1da7a02a0b

Observation e3ffb5c8-188d-411c-9bda-28c431d5f270 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Visual Instruction Tuning with Chain of Region-of-Interest Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.194841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.194841Z digest=sha256:109aadcbd0715cfd94ef2e899121be6c132aceebe2db4fc26ab170438d10f6c2

Observation 5de9a258-b435-4f85-b6b1-ddae3a41fcf0 · outbound

This paper cites Language models are few-shot learners.

Visual Instruction Tuning with Chain of Region-of-Interest Language models are few-shot learners

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.463902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.199706Z digest=sha256:c7be2820e5ac98700ea569b919d4becdd3867f3585fbed375529017d2296cf9f

Observation 0cbeef1e-5d27-4630-a553-675167a3ac28 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Visual Instruction Tuning with Chain of Region-of-Interest ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.204925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.204925Z digest=sha256:e4195324815d9f2c4ab59d00ed974778862fd239d8b402ad32e82030909b9e79

Observation 9d92596d-9fc0-4e35-bfde-20de194ac905 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Visual Instruction Tuning with Chain of Region-of-Interest Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.210203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.210203Z digest=sha256:c275099b956276cd22b5e4dd7ba25c28a6f1deaefab1efc2c6f2a71e98de2e8a

Observation 25120bb8-baa4-4bee-a1da-a4a6a8a00c24 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Visual Instruction Tuning with Chain of Region-of-Interest ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.215019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.215019Z digest=sha256:0b6e2e436b8dbc820164de64f8743d997b289845f2dcc5be1b0437362807ee68

Observation 84cf5b1f-9838-4111-99fa-1eebf1bff84b · outbound

This paper cites CaMML: Context-Aware Multimodal Learner for Large Models.

Visual Instruction Tuning with Chain of Region-of-Interest CaMML: Context-Aware Multimodal Learner for Large Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.220422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.220422Z digest=sha256:341e6c1adbb405a112f1a91720929d6d210e4d4c4d85e18e03bb484d98385a59

Observation e435d66c-1fdb-453d-83dd-ea23f37ff8fd · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites, 2024.

Visual Instruction Tuning with Chain of Region-of-Interest How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.447148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.225726Z digest=sha256:ff5f4ac93ed9a3a97df15760cb268faf9b1785a2ec1468fb84c4f57eca71cf6e

Observation 330339b1-54f0-4fab-b225-c83bfe411f54 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Visual Instruction Tuning with Chain of Region-of-Interest MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.230104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.230104Z digest=sha256:cd40c9291e6a4b1e4fa850f8dca5619159ab206fdc1cd7ce55e58b61b03ec298

Observation da9f4376-e813-40b7-9245-661f3807079b · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

Visual Instruction Tuning with Chain of Region-of-Interest Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.430480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.235144Z digest=sha256:f293327b95887258dd0a2edd0204157ddb24c666169a3db395c27d309b0ab88d

Observation 7cd13423-a335-4eb1-b13c-3259b29aedfe · outbound

This paper cites Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd, 2024.

Visual Instruction Tuning with Chain of Region-of-Interest Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.414109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.239664Z digest=sha256:d2a2369f3aa3b4b66ffd4ac42989f2d68144eaa44cbcd8c5ec4a0eb2bf57fa53

Observation 4640db6b-000f-415c-8d9b-7c2690237bf0 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Visual Instruction Tuning with Chain of Region-of-Interest An image is worth 16x16 words: Transformers for image recognition at scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.244378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.244378Z digest=sha256:16e78f438b009fed31d9c6081b23e0714badda0ced6fe024a67ee11c886c894f

Observation 1ef0dc13-1c7a-46c4-806e-0624221e56bf · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Visual Instruction Tuning with Chain of Region-of-Interest MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.248980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.248980Z digest=sha256:ef7b2dd76bb9476a79b65d1cf0ebf118818977a843a295a9ad3b2b238d4cd1e0

Observation 3a21c9b7-cfd3-4fec-80d0-d08d58f7c302 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Visual Instruction Tuning with Chain of Region-of-Interest Vizwiz grand challenge: Answering visual questions from blind people

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.385984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.253593Z digest=sha256:1ef7fe701b78f53729be4a1e1bef869d10d3ed2eb86b97d39656d6b83b084299

Observation c556b550-d410-486b-afd0-d835d1f22636 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

Visual Instruction Tuning with Chain of Region-of-Interest CogAgent: A Visual Language Model for GUI Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.258489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.258489Z digest=sha256:71c6980d628a66d04b74d44534c05995f7cd6caf38cadeba9d15b176660ed5f3

Observation 079cc4b9-4d25-4ec3-a882-19c9707dce5f · outbound

This paper cites Hudson and Christopher D.

Visual Instruction Tuning with Chain of Region-of-Interest Hudson and Christopher D

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.369113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.263432Z digest=sha256:720d792f6cd7644eb7ef15197edafaf30b5bbd2e84c07f5c232ecaec060f6196

Observation c4697619-5555-48a5-bf31-fd9f83e3f969 · outbound

This paper cites Hénaff, Matthew M.

Visual Instruction Tuning with Chain of Region-of-Interest Hénaff, Matthew M

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.352433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.268222Z digest=sha256:8e79a42f386bf00b076027b6d5ec110427597df79ccd4feb93d2da9ce825ec4f

Observation b4bbc52a-509c-4bfe-8137-2d59fc846468 · outbound

This paper cites Perceiver: General perception with iterative attention.

Visual Instruction Tuning with Chain of Region-of-Interest Perceiver: General perception with iterative attention

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.336682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.272925Z digest=sha256:011bb52c5dfc6a1dc0c89d2bda0eda81d6c34514055cc132f63a4e0b01475e47

Observation 3cf6af1a-dea7-4487-9cc0-4173795c0d7d · outbound

This paper cites an unresolved cited work.

Visual Instruction Tuning with Chain of Region-of-Interest Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:37:18.320239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.277455Z digest=sha256:4b91813f15516d84bbcce67f40496d7c4a311d48e8a57c3ca05ccb915eecbf08

Observation 591e322d-d301-4f49-86fa-f314e6369bdf · outbound

This paper cites an unresolved cited work.

Visual Instruction Tuning with Chain of Region-of-Interest Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:37:18.302359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.282371Z digest=sha256:5afc645c0d4f32247eff2a74f1a06fc894ec27a13208aa1d9af25022a9145c66

Observation d2f2cfe3-d13a-4d97-b3d3-996d296473a4 · outbound

This paper cites Dvqa: Understanding data visualizations via question answering.

Visual Instruction Tuning with Chain of Region-of-Interest Dvqa: Understanding data visualizations via question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.285130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.287735Z digest=sha256:70cee29fd0300cac0795d42f6ee6ff7858be6c56a707891dbf62f295b959e257

Observation 7843d8ba-0381-449d-8b5c-3cf1f1b5cd68 · outbound

This paper cites A Diagram Is Worth A Dozen Images.

Visual Instruction Tuning with Chain of Region-of-Interest A Diagram Is Worth A Dozen Images

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.292309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.292309Z digest=sha256:6f9dabe3a297adc459768aa99652d5437f14b4897d9629b46a02770dede660a6

Observation 8d53e7c6-25de-4c16-a365-66306cfef9fa · outbound

This paper cites Shamma, Michael S.

Visual Instruction Tuning with Chain of Region-of-Interest Shamma, Michael S

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.265417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.297123Z digest=sha256:2c15edd4d85ee831e3d6c9e2af48ffcc6d8d299a3fc2ef7337703d7595fde885

Observation e1afaade-87ce-4d1e-bad3-bacf0af06bb2 · outbound

This paper cites Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh.

Visual Instruction Tuning with Chain of Region-of-Interest Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.247940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.301607Z digest=sha256:dc99d04b580dd2c731b301c48bc59e1b4f95affefbe33f6fa1dc4f46b577f1a9

Observation 12794497-390f-4dd1-a131-303dcc767079 · outbound

This paper cites Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023.

Visual Instruction Tuning with Chain of Region-of-Interest Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.231473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.306160Z digest=sha256:7ef577d7a281eef5dfad803532f20a3cc5b1c15d0756ff385027a046ec226da6

Observation 5ded6dda-4adb-48a2-b999-d05ffc3dd301 · outbound

This paper cites OtterHD: A High-Resolution Multi-modality Model.

Visual Instruction Tuning with Chain of Region-of-Interest OtterHD: A High-Resolution Multi-modality Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.310665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.310665Z digest=sha256:29ee2a4bb34924a83af23f49a8e429b1e6f8e0389db335116a8a8507833cbfac

Observation bb35ca61-05b3-4167-9e45-400fbba3f493 · outbound

This paper cites an unresolved cited work.

Visual Instruction Tuning with Chain of Region-of-Interest Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:37:18.214675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.315645Z digest=sha256:a05758e6a61c63466c948dff581ac3adf4f2b5c32b246a1728f2214d56540711

Observation 3e4290bf-ee0d-46d3-bfc7-554ede6c0307 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Visual Instruction Tuning with Chain of Region-of-Interest Improved Baselines with Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.320912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.320912Z digest=sha256:49c82aed507654991b59d23e7b7f9fbbcdf68c241c03d8b2dcb972472b7c8b1d

Observation 2443bf41-9797-4e95-826d-98b34ed56bf3 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

Visual Instruction Tuning with Chain of Region-of-Interest Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.198080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.325850Z digest=sha256:8e906f6f05a7ff0354d7c1721518cb35680b8a37d6d854f1f2fc09bdef6bf2c3

Observation 1abdd7bc-d975-4a3d-b3ff-94a8c2699068 · outbound

This paper cites Visual instruction tuning.

Visual Instruction Tuning with Chain of Region-of-Interest Visual instruction tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.180787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.330537Z digest=sha256:046aeed6f1961542ccea66bad9a183a5d4a31ee901a2d1ec79eb0d4e446dea6f

Observation accf8c68-4581-4dfe-a57e-3834d3bb63d9 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Visual Instruction Tuning with Chain of Region-of-Interest MMBench: Is Your Multi-modal Model an All-around Player?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.335514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.335514Z digest=sha256:1032d8fe0975ea460fcf1c79a4172733d4e48bb873efda0ff5ab85828f9284d0

Observation 3de3b107-023e-49c7-9f93-5cc5b2c9b169 · outbound

This paper cites UNIFIED-IO: A unified model for vision, language, and multi-modal tasks.

Visual Instruction Tuning with Chain of Region-of-Interest UNIFIED-IO: A unified model for vision, language, and multi-modal tasks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.162203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.340154Z digest=sha256:18993bd76d0f8c6f7e3bf04f65b634cb71d005c65f77f64cd302d9cd445af443

Observation a4311996-20f9-404d-ac32-e78646047c19 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Visual Instruction Tuning with Chain of Region-of-Interest Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.145856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.344615Z digest=sha256:fc75e57b913e323baaec968792b1a9f2ee2a0922ad2bd6340440fa176d4cc4e2

Observation 5d36567e-9a57-4797-a4db-8fc03cdcd8f2 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Visual Instruction Tuning with Chain of Region-of-Interest Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.128235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.349090Z digest=sha256:d0a08b85c55f6cbff567d7d6121dcf995e13d0d114739c88b56d939dd86ad862

Observation ffff0f23-5f26-451e-8e93-3658624d185b · outbound

This paper cites A survivor in the era of large-scale pretraining: An empirical study of one-stage referring expression comprehension.

Visual Instruction Tuning with Chain of Region-of-Interest A survivor in the era of large-scale pretraining: An empirical study of one-stage referring expression comprehension

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.110682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.353541Z digest=sha256:c4b6c3cc6c4ab637c17db2e91e9236912ddd2ce5f893f534e54bf8041e5fcbb2

Observation d394467a-a626-49c3-9447-95529818fc5a · outbound

This paper cites Feast your eyes: Mixture-of-resolution adaptation for multimodal large language models, 2024.

Visual Instruction Tuning with Chain of Region-of-Interest Feast your eyes: Mixture-of-resolution adaptation for multimodal large language models, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.095311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.358115Z digest=sha256:a2b3ae743ca5549df6ab390f256d2616d11364ec341d8d39b183d16ce270b5ee

Observation f5093c0c-ac0d-46d3-85b2-0fc5a29b8acd · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Visual Instruction Tuning with Chain of Region-of-Interest Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.362591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.362591Z digest=sha256:4fef967306a12254623274b0b9afbda6b6deca83f1a141bbbbfcd64896a2b604

Observation 643815db-9c99-4828-ad3b-482dac649944 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

Visual Instruction Tuning with Chain of Region-of-Interest ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.068053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.367414Z digest=sha256:6aadf2a4c4bb954a0f2fc304ec2c13b4e6f56a12b6aa494a137cd00ddb46cf64

Observation b4ce2632-ae02-4819-9afc-37b92cd0e315 · outbound

This paper cites an unresolved cited work.

Visual Instruction Tuning with Chain of Region-of-Interest Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:37:18.051687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.372229Z digest=sha256:d345a120d86a187f933bcb1e0ceef9b4f80cfd2a6f547350d8c665dca6e12d60

Observation 6a537e57-98fe-481b-93f2-3d14dccbf1ab · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Visual Instruction Tuning with Chain of Region-of-Interest Ocr-vqa: Visual question answering by reading text in images

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.376842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.376842Z digest=sha256:6d230130944a7046abd2ca338dd144fed2a958e75c3c9eb001d5c3ba8751a671

Observation 57f11aad-dc40-4278-9c8c-0a7153dafbc6 · outbound

This paper cites Gpt-4 technical report, 2023.

Visual Instruction Tuning with Chain of Region-of-Interest Gpt-4 technical report, 2023

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.381421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.381421Z digest=sha256:ade5b201e9208f1aaf804c005a9bf3d2204fde7b3a4b99fa941bd0d009892d3e

Observation 38811b5a-1a36-4c63-a18e-b5d8698edeed · outbound

This paper cites Learning transferable visual models from natural language supervision.

Visual Instruction Tuning with Chain of Region-of-Interest Learning transferable visual models from natural language supervision

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:18.013848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.385729Z digest=sha256:368b2cc43000754fa04ea02a8739b21579d7e5125c0d49f701f0a7088c488e4f

Observation 644f9a0f-3f3d-4334-af87-ca6ce011f14a · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Visual Instruction Tuning with Chain of Region-of-Interest Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:17.997294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.390340Z digest=sha256:e917e8d01a4e63ea62f20963c70eeba023f17c5361fdf2d9f5f91dfcb21d3152

Observation fbc1b3c0-c457-481f-b7ca-e37b6a1f6c36 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

Visual Instruction Tuning with Chain of Region-of-Interest A-okvqa: A benchmark for visual question answering using world knowledge

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:17.979168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.394934Z digest=sha256:d16b21d45f5b3cd5caaa23e85940642f54414be4f299a3eea702b009a2ee512e

Observation dd55ed0c-9ba9-42fc-b859-2a79c31590df · outbound

This paper cites Unified model for image, video, audio and language tasks, 2023.

Visual Instruction Tuning with Chain of Region-of-Interest Unified model for image, video, audio and language tasks, 2023

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:17.963174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.399666Z digest=sha256:fa59c931a79514dc01b6a41a490a8c6acb2963eeb4c3f9d2227b7c9f654c096c

Observation 2528b9fb-0b48-43cc-b69d-6c13e0ac91c3 · outbound

This paper cites Textcaps: a dataset for image captioningwith reading comprehension.

Visual Instruction Tuning with Chain of Region-of-Interest Textcaps: a dataset for image captioningwith reading comprehension

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:17.946025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.403889Z digest=sha256:9b881472c049efc67438523a71c75985db38ab20d3b327d25de126aa163cb6c4

Observation 01a5cfba-bf17-49c6-9feb-8a83590156c9 · outbound

This paper cites Towards vqa models that can read.

Visual Instruction Tuning with Chain of Region-of-Interest Towards vqa models that can read

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:17.929242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.408537Z digest=sha256:dbb482055a3d8ac6921c445e0176fa3832f531527e943717c5302971cdaee1aa

Observation b1f1336e-84bb-485a-b837-13f6c6bf7ba7 · outbound

This paper cites Chameleon: Mixed-modal early-fusion foundation models.

Visual Instruction Tuning with Chain of Region-of-Interest Chameleon: Mixed-modal early-fusion foundation models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:17.912465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.412857Z digest=sha256:3705bcd77d80cff9fc8d6c3ab297f4cef110b5c3789e955b1f25e942337000ba

Observation 81263de6-0bd7-480e-9e89-d89e75387bbf · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Visual Instruction Tuning with Chain of Region-of-Interest Gemini: A Family of Highly Capable Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.417414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.417414Z digest=sha256:5ca26b7760856d3e2e12ab4e82681140755aa6c4de9f4351fcaf257117c99003

Observation 53f83287-8d9f-4f53-a903-57551cc69c45 · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

Visual Instruction Tuning with Chain of Region-of-Interest Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.421900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.421900Z digest=sha256:40a99fbb1fe24fd98c45858938987ba7bdb84d93aed57d0309696fbef2fbf038

Observation ec62cae5-d286-45dc-827e-d32018c112ae · outbound

This paper cites Llama: Open and efficient foundation language models, 2023.

Visual Instruction Tuning with Chain of Region-of-Interest Llama: Open and efficient foundation language models, 2023

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:17.896807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.426731Z digest=sha256:67fdebefbccaaf2e5dee6a018b7251814c5e3f7ce213002c5486f90691011bb3

Observation d5c8882a-c4d9-4b1a-a521-c035ec3b7c72 · outbound

This paper cites an unresolved cited work.

Visual Instruction Tuning with Chain of Region-of-Interest Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:37:17.880212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.431671Z digest=sha256:7e7069e57a0beb035d1f7e3509093d35b07b0f330d207ad43bec30b4caa5833a

Observation bd9ee375-7009-48ba-825e-079455b080cf · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to- sequence learning framework.

Visual Instruction Tuning with Chain of Region-of-Interest Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to- sequence learning framework

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:17.862966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.436174Z digest=sha256:d3c8a65ce4825f2504b98995ba385b21acd4a5387112607d3fc2db3ccb56d4c1

Observation f6c21bdc-d0a1-4ced-be6a-1080b311c758 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Visual Instruction Tuning with Chain of Region-of-Interest CogVLM: Visual Expert for Pretrained Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.440571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.440571Z digest=sha256:733a4b4fa9de5e460746a27a012c17c3e269591ca3bf7db48215251ca79dc478

Observation e53a4edd-b4df-4dc6-b17a-9165e5a7d895 · outbound

This paper cites Berg, and Tamara L.

Visual Instruction Tuning with Chain of Region-of-Interest Berg, and Tamara L

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:17.846393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.445203Z digest=sha256:f2193866617e725e99230106da13e669a08d9a72b61e744236072631aa5d2990

Observation 13e047fc-b32f-47a1-81ff-3a562b728ff1 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023.

Visual Instruction Tuning with Chain of Region-of-Interest Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:17.830646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.450074Z digest=sha256:57ab6b8bcb4714567a1ef88d628aeeba905a430d911acaf6493b5ea3e462cc78

Observation 71e12d3c-17b0-4a0e-98be-a4dff0d7ba6e · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Visual Instruction Tuning with Chain of Region-of-Interest Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:17.814495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.455372Z digest=sha256:6647f91fc4c5a3e02a1969b664c11b427463a32aa54e94e05b7103f49cd70e6a

Observation dc659125-cc7c-43f0-a077-5dd32e9c7da3 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Visual Instruction Tuning with Chain of Region-of-Interest LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.460410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.460410Z digest=sha256:4e1ee1be29d4af9100df94a8b2923d45d8800dcb8f4f284821ef615166bfcb86

Observation 664cb106-88ba-4595-a124-35a26f2fd489 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Visual Instruction Tuning with Chain of Region-of-Interest MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T22:37:17.465095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:37:17.465095Z digest=sha256:1640b7b52eab4ed54ea30911a4cddb5c158bd83e8d6148a8a1562f13525c7954

Observation 625d38c0-d572-49d9-9c86-30596b0c8a6b · outbound

This paper cites Uni- perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks.

Visual Instruction Tuning with Chain of Region-of-Interest Uni- perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:37:17.797604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:37:17.470659Z digest=sha256:bdad12c7d17f236ade56eed727154768984ec02032a7b78debce462952d23b07

Pith citing papers

No inbound Pith citation observations are available.