Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T15:18:49.974575Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2501.14276.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T15:18:49.974575Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:38:36.041216Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T11:38:38.471219Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b953dbc5-4c2c-4380-ab0a-4568c43acca0 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Visual instruction tuning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 793e77b5-ae36-4a4e-8708-d4d4d1ad1c04 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 06c17ce1-689b-4fa8-a059-5500e98a098d · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Sharegpt4v: Improving large multi- modal models with better captions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a2eaa09e-434f-4f68-b289-9f04a195ccdd · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Monkey: Image resolution and text label are important things for large multi-modal models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4c8928aa-2cb5-4c60-94ad-ace2a76da0d9 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models InternLM-XComposer2-4KHD: A pioneering large vision-language model handling resolutions from 336 pixels to 4k HD
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c450716a-35f9-47b7-9d62-b9aa6d260320 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Salgan: Visual saliency prediction with adversarial networks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 83e6ff61-f14e-45ea-8bab-98341f2d92cb · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models PaLI: A jointly-scaled multilingual language-image model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 341bd568-5bef-4ff7-9209-2e8df930f74b · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ac08dd5-2ae7-486a-b65d-a0f9052462d9 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Cambrian-1: A fully open, vision-centric exploration of multimodal LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0b98434d-d5c4-473f-ab85-df6b0825bfaf · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Llava-next: Improved reasoning, ocr, and world knowledge, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ac4578-9d36-44c6-bc74-9aac4380cf8c · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1365c077-116c-4169-a8d4-ffc36c18a8c9 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Instruct- blip: Towards general-purpose vision-language models with instruction tuning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 695af14c-d233-47bc-b4bd-f3520a045059 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Learning transferable visual models from natural language supervision
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1602eccf-23fe-4268-b4c2-4910c6fce142 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Dual modality prompt tuning for vision-language pre-trained model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c149faca-a31e-419f-903f-fbdc5d9bfb1c · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Sgva-clip: Semantic-guided visual adapting of vision-language mod- els for few-shot image classification
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f3181c5a-fa3f-4cb3-afc8-eff3fb1e2313 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Gpt4ego: Unleashing the potential of pre-trained models for zero- shot egocentric action recognition
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c5de4ea7-eb2f-442a-bde9-424ba3531185 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ca60c17-f493-44a1-8a96-f9d0d7887b4b · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models InternLM2 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f95fd1fc-71b6-444c-bea8-964308349d3e · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a553936-ea40-47d6-8922-850dd0ea4399 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 90c802b5-471a-4462-b9a4-4c44ce254abf · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Towards vqa models that can read
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fc766ca9-d6d0-4117-a0b4-122ad133a92f · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Mmbench: Is your multi-modal model an all-around player? In Proceed- ings of the European Conference on Computer Vision (ECCV) , pages 216–233, 2025
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 22cb8890-990b-464b-9528-904eec590064 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models A diagram is worth a dozen images
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 45db2590-02e4-4fa6-ad30-f86564e89be3 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9fb2e153-f3fc-4ab9-a3de-3e38e67e7fb4 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Improved baselines with visual instruction tuning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 92118df8-568f-47dc-bc9d-2b271f751bbf · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Decoupled weight decay regular- ization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 54c31879-1f12-40ce-818f-3f1fbf90adc1 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Dvqa: Understanding data visualizations via question answering
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 73927c6d-1325-4fb3-9fa0-b15755cc4b35 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation df0f00d9-5505-4a04-9084-806373e0ec1c · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Docvqa: A dataset for vqa on document images
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 423dafa5-36b5-4243-b489-c6c85c98862a · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3fea2abc-bf7f-49fa-877d-8fc5f3113e0d · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Ocr-free document understanding trans- former
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 24fae8a2-ce73-4e34-b13b-0bee0ad822f5 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Are we on the right way for evaluating large vision-language models? In Proceedings of the International Conference on Neural Information Processing Systems (NIPS) , 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5a44332a-664a-4c49-8cee-a2ec6c247fdc · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Mm-vet: evaluating large multimodal models for integrated capabilities
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9906a31c-bea3-48b3-9336-8f37e7158ecd · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Seed-bench: Benchmarking multimodal large language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c6c03810-9ee8-4a14-8e6e-a27d589e770a · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Grok-1.5 vision preview
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6aef7a1f-6d98-4fb2-91b5-1776dc5c62f7 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85204ccc-18ad-4c77-bee0-92926ec24000 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d55b6464-93a5-4e79-8433-028d34009484 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Infographicvqa
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 05ce2f96-c2e0-4b66-a5b8-5075adda0bfc · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Evaluating object hallucination in large vision-language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dabf7dd9-6e3c-4415-ae52-f3c04b9c7b12 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b290609f-f6fa-4c5e-b999-c5d4635ae25f · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models What matters when building vision-language models? In Proceedings of the International Conference on Neural Information Processing Systems (NIPS), 2024
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cc3526a8-d06a-4df2-b2c6-bbb387b08b1c · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Vila: On pre-training for visual language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5bf02c69-1611-4eb0-888a-7a11ed03c12f · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a48a38b5-7dce-4027-9b86-642a409825f6 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26ee26dc-7886-4915-ab75-859a9ae4212c · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models PaliGemma: A versatile 3B VLM for transfer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 140b4b66-190e-45ad-9a71-68564329d3f7 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de49709a-2c31-437b-9bed-827cca1ca6f6 · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc3821e-0ff2-44bb-84c7-380649c8191e · outbound
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377603f8-1932-42d4-ac27-1cb2d7df6fed · inbound
HRVVS: A High-resolution Video Vasculature Segmentation Network via Hierarchical Autoregressive Residual Priors Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.