Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:51:42.939404Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2507.20156.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:51:42.939404Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0e6b01c7-ae0a-4c47-aef8-eb42e524c7a7 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e00b4115-514b-4eee-95e4-1bb26842c2bc · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Wang et al., Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution ,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 183d1dc8-1e3a-4d5c-be39-1e6d9576aa0c · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality GPT-4 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c8b985-9a1b-4bb4-b3dc-5917fd79b9e6 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1318520-c311-4073-a0d8-7985c669838e · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality LLaV A-onevision: Easy visual task transfer,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 73b8b5eb-d6ea-4da7-9035-8b5ca9246bfb · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 336b6437-7ad8-4a96-aab0-9fafdb0b7ece · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Masry et al
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1d86c0e1-6bd8-4667-8769-8edaf95e3ad4 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b92f115-f401-4d14-ad10-ef1761793569 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Mplug-owi2: Revolutionizing multi- modal large language model with modality collab- oration,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 78af3003-449a-47cc-91f8-2db9ee342ba2 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Conceptual 12m: Pushing web- scale image-text pre-training to recognize long-tail visual concepts,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 473a62d3-b1fd-4d9b-b6e9-383c663df6d5 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality What If We Recaption Billions of Web Images with LLaMA-3?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c429a6-d137-4f20-96d0-0f6be4310713 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Obelics: An open web-scale fil- tered dataset of interleaved image-text documents,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9e104414-5404-4c31-9018-ac5f49acf053 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Bai et al
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0e493f98-bf3c-4af0-b44d-12a222dfb98d · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Towards efficient visual-language align- ment of the q-former for visual reasoning tasks,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6e260517-0b60-4f18-a560-8cb159f743b1 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Training language models to fol- low instructions with human feedback,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 54f10874-2138-4c03-9c82-14b38eb9cb4b · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Let's Go Shopping (LGS) -- Web-Scale Image-Text Dataset for Visual Concept Understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2ec930d8-ce85-4785-bf34-a4aa211b036f · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Lima: Less is more for alignment,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 00ccbdd9-15cd-48ed-8bc3-9d6e4dc55e88 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Agarwal and D
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ed77cb0e-c99b-4a29-b816-52dd411f3a29 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Visual instruction tuning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9d9c0583-18e0-48b0-a1a4-381266f45c07 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality A Survey on Hallucination in Large Vision-Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5d58d02-0272-421e-b5ab-3d5a4059181f · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality On the origin of hallucinations in con- versational models: Is it the datasets or the models?
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4b5aff2c-9fb3-422d-887f-aa0e71838b34 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality 48550 / ARXIV
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a5b05a76-f4e4-4ee9-8d8f-eba98670ca9a · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9137c9a2-2555-463d-9124-98d294a67a3f · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality MiniGPT-4: Enhancing vision-language understanding with advanced large language mod- els,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b65c4bde-955a-4669-bced-69f5fb7df324 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f842be4f-5590-4ba2-a228-f4a8114b7da0 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86302a80-cbc6-4715-ac71-616b76a72c2a · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Alpagasus: Training a better alpaca with fewer data,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b7cf5628-da59-48db-85ee-ea0ce844e3bc · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Language models are unsupervised multitask learners,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4cd92eee-129b-4164-9e4e-b0d6cd0874c2 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality A frustratingly simple approach for end-to-end image captioning,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9b75eceb-4c74-4217-bd8a-9a4b8eece9c6 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Learning transferable visual mod- els from natural language supervision,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5a46c7e5-1d23-4691-88a8-cc08dfe4f5cf · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec5e046e-f3f5-49b7-8e2c-91e7c739fae6 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Available: https://cdn.openai
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a81406a3-565c-48ea-904b-866e719a8904 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality A survey on enhancing image caption- ing with advanced strategies and techniques,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2a3bc4dd-9bb9-48eb-ad34-b5cb2aa1b190 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 57702283-bbd4-4e3a-b763-cefa40cf6a70 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality 32604 / cmes
Reference 1506
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a99d609a-3c26-45d0-afd2-016f4732510b · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5a050c-5f47-489a-928f-c215d8208db3 · outbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality 48550 / ARXIV
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.