Pith. sign in

Paper Citation Record · LEDGER

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 3 inbound Pith citation observations for arXiv:2506.12776.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12776 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:48:35.027792Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T21:58:53.702009Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:37:14.448914Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0e82b7c-81b1-41c4-a0b5-6c052a6f472b · outbound

This paper cites GPT-4 Technical Report.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.130945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.130945Z digest=sha256:944593beadc2436f332c60eb9ebf33043eb7350ea50f70e343bd52392135a143

Observation 6369f684-748a-4e43-9414-9bc5a166ee8b · outbound

This paper cites Qwen2.5-VL Technical Report.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.212811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.212811Z digest=sha256:b6d809b5e533716e3fc8154f0a0e969374b94a274b4ff7fde2e7b4f8c683ba11

Observation 80ce4dc1-81d8-4543-8384-df02b0d133c3 · outbound

This paper cites an unresolved cited work.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:48:37.930240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:48:28.348014Z digest=sha256:9621165d33f540866a5bb4984e8131a6897b2a66d2482d31561b7c4a50747a3a

Observation 3a605ce2-499d-4e75-a0b5-7159b53b774b · outbound

This paper cites Ocean-OCR: Towards General OCR Application via a Vision-Language Model.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.448857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.448857Z digest=sha256:0071571ec9f01f8f7d5200d2a0bea8226d5a6867ccda44a3ed1fac589b30977d

Observation 92ffee03-72d9-4c73-9c81-4513847c6dcf · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.559091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.559091Z digest=sha256:e69845a88c864bf19acaa9c61ce840b0535b1cb2bf8f2820242e179b5aee9273

Observation f7d0afd1-97d5-4d80-80f4-e03ef8c6561c · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.686334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.686334Z digest=sha256:74458f26d7338378f1633adf013e1b7323b0c827d181a380d115815c14a7e703

Observation b260ec6d-0bbb-4899-8e6c-51b9aff3a854 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.785387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.785387Z digest=sha256:56a39b4d9201cb162da12169a80afc8870e260b15641a5e61d4715ceb941ffef

Observation 72c52b55-10e8-4431-bf04-c0040d6601aa · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.885481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.885481Z digest=sha256:cd92a4ced5a61415d24e614964f99b669cde57d871fb67d94e674f1317e88358

Observation 15e0b855-cade-44c4-8722-8e0b0d8f28aa · outbound

This paper cites Bert: Pre-training of deep bidi- rectional transformers for language understanding.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Bert: Pre-training of deep bidi- rectional transformers for language understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.989104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.989104Z digest=sha256:1a70a9bc45b8fff57f06dbb6034b37364329d8e3787c76741b0be6372846c6b7

Observation afc9fdcc-c093-4040-8fc8-8c54c2b820d5 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.154407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.154407Z digest=sha256:63dc692fa80de3f9fba6d8b594c88fad312019d220278d800f1a6a7f935106f5

Observation 34d6d148-e214-4a23-b785-302d79a40750 · outbound

This paper cites Gpt-3: Its nature, scope, limits, and consequences.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Gpt-3: Its nature, scope, limits, and consequences

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:37.749092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:48:29.255022Z digest=sha256:1f831baa50e7cf3a23ccdde00bc485655fbe12148c5ea172c7bbad7507666db9

Observation e32e37fb-45e4-4006-ada8-2c8766dc3bcc · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.381519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.381519Z digest=sha256:2761d63e666bab577b79ed96f239a74bdf264a7f5871ed7109c07fb5a1d26d71

Observation 79bd0a28-93a8-4458-9fc0-5186a0340428 · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.512536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.512536Z digest=sha256:8e7d81737220f84837cb42d1fb9e6d94c81f8fff25f080ceae00d9715f1209db

Observation 1d2629bd-4eee-4440-90d2-4227f3240ef6 · outbound

This paper cites Seed1.5-vl technical report, 2025.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Seed1.5-vl technical report, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:37.524920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:48:29.634978Z digest=sha256:f265ef4e665d2873c950e55f30f7fa230b9116f795cad7f1a2e4628d2b9bced3

Observation 383e8978-26b3-4a4a-8418-e1043bfa123e · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:37.308102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:48:29.741955Z digest=sha256:fbdc2ef7a329d47511afb4d6f01c326bb4a7cc3b3b1547caf6d5d2ddf6b0dd3f

Observation e7ff2c4b-6a7e-4a70-a0bc-1de1c4fcf400 · outbound

This paper cites A diagram is worth a dozen images.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models A diagram is worth a dozen images

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.843867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.843867Z digest=sha256:70cc9ed960c561aebbdd77fe880cf03c580d1381862a471c0d1161ebe533c4ef

Observation 757c8926-79bc-4b83-8eef-656cf9bd9112 · outbound

This paper cites BERT: A Review of Applications in Natural Language Processing and Understanding.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models BERT: A Review of Applications in Natural Language Processing and Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.931868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.931868Z digest=sha256:2c2675454620b47cb5b4f970ecd3f458a9950f11197af219019f5c2d6dfbf039

Observation f45a04d0-49c7-4083-9a7b-458e621671fc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.085311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.085311Z digest=sha256:5850361ffa717e334c01326d944421500f0f6ef9a2e451a22ca645ae03f47333

Observation d0599d4f-1320-4427-957d-93ceb82ad4d3 · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.185030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.185030Z digest=sha256:5969b3dc98801eaae5ff725d536279bb1dbb9cf4e1b3c141915383c9e28d7a98

Observation ecc99e65-9cfe-4be9-8062-c03588f67105 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.316110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.316110Z digest=sha256:f4521bb08c3e08762e7b8e18c45a6ec7023c4e0d987691b96a9e626e457fe36e

Observation 4294a87b-37ca-4dd7-8684-3793d217ceef · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.416053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.416053Z digest=sha256:ce2d387b58ca1f8bbdd3543cdaacb9ba95f53211c75f4418070b8b4529cac7be

Observation c8e8a796-7b0c-4d36-bc25-67623a38812a · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.535552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.535552Z digest=sha256:8a4b4a4a92dd1e52bf1d28705b2149f5ec2a2bef34beab5745179ab5a15234e1

Observation 4544fb5c-1eb9-4760-9932-a995bd17dc91 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.624519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.624519Z digest=sha256:6c3f7eb95e4a51e500f308893ae48d195524524f1c9c29dbbebd417b29bc988b

Observation d39cb784-ac20-49f6-858e-3926719fd61f · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Monkey: Image resolution and text label are important things for large multi-modal models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.741494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.741494Z digest=sha256:51761220b9b977bf9f09d0601b9d1b6b7eed09483d41199ca4f20feed8464b82

Observation bf36f4bd-1fae-47c5-85ba-3a755db47a5a · outbound

This paper cites Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.849435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.849435Z digest=sha256:57eb88b54c539a41e18725bf0b319b93c6547eb56e1ccdb619a908ec55d09d8b

Observation 316f67e6-7fba-4b5e-84ee-da9e18e278e2 · outbound

This paper cites Improved baselines with visual instruction tuning.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Improved baselines with visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.977433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.977433Z digest=sha256:34dfe7b7ec20414002174db9e04771cfb14147b35260109444956e22b1660d05

Observation c7c5f576-1235-4f76-ae9b-4cc6d7c4d51b · outbound

This paper cites LLaV A-NeXT: Improved reasoning, ocr, and world knowledge.https://llava-vl.github.io/blog/ 2024-01-30-llava-next/.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LLaV A-NeXT: Improved reasoning, ocr, and world knowledge.https://llava-vl.github.io/blog/ 2024-01-30-llava-next/

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:37.142439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:48:31.082464Z digest=sha256:72c1d7a4b59d6f77282649175f0a36a1673a2ba123cc0867e0bd391f5e59010f

Observation 50640e39-879f-48dd-9f2f-013a9676da49 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.182343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.182343Z digest=sha256:cc10bdfd860848492e5c30bb0ee307d6f03262363d60200b1d8602d79e2615ea

Observation ef186e24-61e5-41cf-bf81-fda06bac4f5c · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.252154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.252154Z digest=sha256:05a1e7e2a613da8950e1321748b8617ceb7c66159cd3d3e690282f7e8f4ed507

Observation 7f8843cc-3055-4ec6-9c48-159ad1fb97a6 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.366111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.366111Z digest=sha256:ea97d6d772249bb627d53b121c1950788562334890b293fbb57e34d48ae65efd

Observation e39b5d73-5718-41e6-9886-f2394b594b86 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.424586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.424586Z digest=sha256:53be77fb967c2cfb78600ebafe2e1e08f1333713c617891f7b4ae6b3fedba769

Observation a607c0fa-11a5-46ca-8516-3e96c1384d3a · outbound

This paper cites Decoupled Weight Decay Regularization.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.521261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.521261Z digest=sha256:c15f19df14117a61734f441458d4940b2b324b576ff30e89f1cb8f948288174d

Observation 5fb20c4f-02ed-43fd-a417-9700eb54ee93 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.608767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.608767Z digest=sha256:112dc824c6ed95090ee2d91a584b27ce4fc4f15582a957e065709d2d22002355

Observation 56555408-5b6c-498d-a63f-ba3f12f10b05 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.688789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.688789Z digest=sha256:02010a4dea94661def7dac757acf17317e72ca05a0f67d21bb37ba2bd3c44aa8

Observation 888b9194-5eee-423a-a6c0-3abbdc2e8561 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.784648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.784648Z digest=sha256:2d9721b3bd0a108fdfe1e535a992b00a17c92ed2501467b23905c2ce7bd1c679

Observation 3622c02b-b745-4a0a-84e8-765ad3c35cd1 · outbound

This paper cites Infographicvqa.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Infographicvqa

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:31.859231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:31.859231Z digest=sha256:6fa4d97c07d93d37d8fcca853cbfde57087eabb70668b423cb93373c806fff92

Observation f2035eb1-a169-47b4-a1c2-e9e7a65c513c · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Docvqa: A dataset for vqa on document images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:36.930932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:48:31.955507Z digest=sha256:dcabf855fdc2b4992dfecfe15b557045a3d85ee026cc2fab3b5e2178c931bba3

Observation 6a20a55d-df00-4e28-a867-2b132bc3d205 · outbound

This paper cites Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution, 2023.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:36.759217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:48:32.047540Z digest=sha256:066b264defb9bc6dc91ad049c86fab7000aa1eeefab9d13356ae9bc70a53e45a

Observation 9b0c52fe-66d5-49d3-8796-b11c09c98a05 · outbound

This paper cites Ovo-bench: How far is your video-llms from real-world online video understanding? In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Ovo-bench: How far is your video-llms from real-world online video understanding? In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:36.557679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:48:32.169757Z digest=sha256:ee53d71d382f9f67ef6aadf49fcc84f780d30ebc25371f6bf907e4721f13b6c6

Observation fd369666-9d8e-44d8-a40e-dd555445bd52 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Learning transferable visual models from natural language supervision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.297562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.297562Z digest=sha256:6f595c924571de15cad41be508f0d6787c425635a6aed6174eb7687a5581b136

Observation 396cba8e-2143-49c2-a0a5-5cb1159e68b6 · outbound

This paper cites When do we not need larger vision models? In European Conference on Computer Vision, pages 444–462.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models When do we not need larger vision models? In European Conference on Computer Vision, pages 444–462

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:36.296805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:48:32.404618Z digest=sha256:525b36f90387ef72c5785d25e95fc171f155ff935d472db9089ff1253f8b9193

Observation 7ebac53f-50e3-4040-a501-c1ec785f2b0b · outbound

This paper cites Towards vqa models that can read.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Towards vqa models that can read

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.488516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.488516Z digest=sha256:b55dba33b4c46f6386093c238e167e4a24ac57d2c87f85ae2ce090d6eae22548

Observation 7f22c7c6-e324-451c-bd75-660a10d00b58 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Roformer: Enhanced transformer with rotary position embedding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.571426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.571426Z digest=sha256:094cd72c026ae9481fe961b8dde529fee512a88594d742e44e3c0e8b06b8b972

Observation 26a6e65a-bf4a-4826-8a9c-fafbb8a22589 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities, 2023.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Internlm: A multilingual language model with progressively enhanced capabilities, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.667962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.667962Z digest=sha256:81dde5106b09497961989b6ea39f426901f2b0d4276afd569e652a667b5e5ba5

Observation 7d899ee3-e8d1-4120-8aae-aebb65305b14 · outbound

This paper cites Kimi-VL Technical Report.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Kimi-VL Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.740707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.740707Z digest=sha256:1e2718dc537feaab748b1f67cd8d9749b74db07bc4749d5654dc702d91c28a6d

Observation d2e9f43c-e693-4959-852e-547fcdf21f5a · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.855555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.855555Z digest=sha256:7d695eeff369bb95dca0bd7cc3562e0c5a41ea7d2a9d846e7d9c772f24eabdeb

Observation f4b0045e-7682-452e-b341-2d0e1956b838 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:32.992231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:32.992231Z digest=sha256:0a24058bbfa71405313497956da902f2fbb66277f481f38fdced210bbe07da7f

Observation 2cc87ade-6243-4ef1-8afa-2e443bdfbefc · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.114735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.114735Z digest=sha256:bf742926335b7cc45bdcdb55f1fe1adfa870bd6abffe38f27ca24078d9a3aee1

Observation 32f6e633-21d2-455d-a48f-02b01855bc70 · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:36.091872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:48:33.213010Z digest=sha256:0eb2df07103ea9de589db8cfa58f65344518e416e93752adb97696ef8827f4aa

Observation 3b412e13-b0c9-482f-9f09-a2dae75db265 · outbound

This paper cites Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.338434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.338434Z digest=sha256:5ebc30fe64900cdaf1807d9ce64f452189671bc1ee935d465bc585277641dd5d

Observation e80137b3-24e5-4535-8511-8cd6d336417a · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.446924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.446924Z digest=sha256:d2da88c5cb57b684d585d9505af87b777bae836a9745e1c3e5dbc1b4dbac2667

Observation 1b8db740-51b7-4499-9a91-3e129af10f55 · outbound

This paper cites Qwen2.5 Technical Report.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Qwen2.5 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.535196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.535196Z digest=sha256:950872667cc7f8f921c228736ccaeb61077b8de8f5e15fcf5f7ee395f6238186

Observation 5a2db8e0-340b-4ab7-b439-f2c2c86854f8 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.654963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.654963Z digest=sha256:35d24aa64420403e97ef332645cf7842b0b87b17d9bb1567105e5b6e10a75687

Observation 94d73980-c679-4485-84c2-0af603b7dc22 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.765571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.765571Z digest=sha256:23c4e3afb2d3790aefd7c867445ed8d8cffc96b8a2605071009956baf38fc7ac

Observation bd675666-0726-40a7-ada7-dcd14b2e0ccb · outbound

This paper cites TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.913826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.913826Z digest=sha256:78d374458729d8d972d1f5183c6555e31408baee97d106e8ef02ac1e943e2073

Observation 727ce860-a44f-441f-8fed-e0382e7af1ba · outbound

This paper cites Sigmoid loss for language image pre-training, 2023.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Sigmoid loss for language image pre-training, 2023

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.049608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.049608Z digest=sha256:0c195c84a3c6d416a34434aa081ebe0a0002b966b4a479542fecd1263121252e

Observation 4cef6c13-6340-4263-8359-99ce2afbf1c6 · outbound

This paper cites InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.167998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.167998Z digest=sha256:f72243473f1fe1786c7cea0e112a977d073a9b82b4d17cf7ed8b4fb3add654c6

Observation 309da78b-8d1e-4153-836e-aa64ee4ed65e · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.329669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.329669Z digest=sha256:5c0259cd8d8ab018a710d884ba76d4a6322d5f137e1c381e3c5fc8e3e337a87b

Observation 3d866024-da5b-4073-b695-8bf8b0b3a9ae · outbound

This paper cites LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.447271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.447271Z digest=sha256:472b8c1cba0e92fab0c637a0627c34dbbc1a72f4ca2fb010514ca34b564610ad

Observation 80f41d3b-5c2a-4624-901b-1e90b1c58f28 · outbound

This paper cites MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.538261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.538261Z digest=sha256:7ee3e530e7d8b8bfe7cd2372adbc8a980530d2961dc8d9fa097957d56a3e0424

Observation ebd9d7f2-d8c0-4b09-95a8-8012cb60fa00 · outbound

This paper cites Swift: a scalable lightweight infrastructure for fine-tuning.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Swift: a scalable lightweight infrastructure for fine-tuning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:35.879362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:48:34.646887Z digest=sha256:d7537485f47e85096429b121ce5ea80e79812379a1e79358c7bd19c284c88057

Observation d7e63517-4296-4df2-99a0-37102c2a6997 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.833805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.833805Z digest=sha256:9f9ce33e9fb4c1c6b524e1b18f9eb4fa6d7117eb176b7e69799e675528901d59

Observation fce9c156-8fc7-47df-8f60-e76731845313 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:34.911797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:34.911797Z digest=sha256:1acc1c90c332fbbd7b58737990de9f649f4b73e4435354c837c24870ee532c01

Observation 7a882fd8-dd3a-4dfc-ba9a-baa3e1f05b5f · outbound

This paper cites $” or measurement units such as “cm.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models $” or measurement units such as “cm

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:35.630405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:48:35.027792Z digest=sha256:3caeaea195365c49c49c1a327e61cf0c5ef929677b48be322d8904fdf0ea0100

Pith citing papers

Observation e7ab0eda-b5a6-4a65-9a70-0eda9239920e · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.034155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:24e44203282dbc12dd78157b30c11494b01ec7fe6723208727fce2c5e84995d8

Observation 112900c6-95c3-433a-b8c0-8835c23ea18c · inbound

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale cites this paper.

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:35:52.439397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:58:41.377996Z digest=sha256:e6f0532d192edc47de6886008a2f296e70e3206513e810877290896b2b3a04b7

Observation 0fe9215b-8e3c-4eae-9b83-3f1e8e898893 · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.450399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:b665c0e62f6c16df8b174d849f1727a4db32d52d231f046a12ea9c72c73f495e