Pith. sign in

Paper Citation Record · LEDGER

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality

As of 17 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2507.20156.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20156 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:51:42.939404Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact6
  • verified fuzzy19
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e6b01c7-ae0a-4c47-aef8-eb42e524c7a7 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.792055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.792055Z digest=sha256:60851af615729c6b86f0f69f7e0527c39cf9246f286cbdf3ef32d5cacac9399f

Observation e00b4115-514b-4eee-95e4-1bb26842c2bc · outbound

This paper cites Wang et al., Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution ,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Wang et al., Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution ,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.777817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.797227Z digest=sha256:a788d9087eaebc8dc9895288b33aaadd49abc29bce35e8a759c2d050ac65065b

Observation 183d1dc8-1e3a-4d5c-be39-1e6d9576aa0c · outbound

This paper cites GPT-4 Technical Report.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.805659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.805659Z digest=sha256:f5d4aee15549146a2cf720ed88fc8db434ab20c71fd75c8b3eed1462f5ba9cf7

Observation c2c8b985-9a1b-4bb4-b3dc-5917fd79b9e6 · outbound

This paper cites Qwen2.5-VL Technical Report.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.809672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.809672Z digest=sha256:7b15a4a0d58b6e09beae545054f92a560e98d1aaa9ce69115642e2cc7ca6fc23

Observation b1318520-c311-4073-a0d8-7985c669838e · outbound

This paper cites LLaV A-onevision: Easy visual task transfer,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality LLaV A-onevision: Easy visual task transfer,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.765837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.814331Z digest=sha256:74b7649c373eb3e34ab2b7a8c433223edd512916533e248222efc960c43f6476

Observation 73b8b5eb-d6ea-4da7-9035-8b5ca9246bfb · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.818265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.818265Z digest=sha256:e3c3e007966b60fa62d22c51de5879e4ba86da5d4682f0d83e0ccf2e77b8f6ea

Observation 336b6437-7ad8-4a96-aab0-9fafdb0b7ece · outbound

This paper cites Masry et al.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Masry et al

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.753778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.822523Z digest=sha256:d7050b90157c1d000d7d244e13b39dfcc94915f2b7957d24a3df04aab54731cd

Observation 1d86c0e1-6bd8-4667-8769-8edaf95e3ad4 · outbound

This paper cites Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.830659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.830659Z digest=sha256:c4cc7017ab62fed50f9fcfa9e10e3f4a66b5d4da79018101d27f45d2e7c056df

Observation 6b92f115-f401-4d14-ad10-ef1761793569 · outbound

This paper cites Mplug-owi2: Revolutionizing multi- modal large language model with modality collab- oration,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Mplug-owi2: Revolutionizing multi- modal large language model with modality collab- oration,

Reference 9

Resolution
verified exact
raw_fallback, observed 2026-08-15T17:51:43.543200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.835619Z digest=sha256:56efed9435c8b8d12c29b055b6b28fcbad04c92900ac75fc3ab0b025c0835978

Observation 78af3003-449a-47cc-91f8-2db9ee342ba2 · outbound

This paper cites Conceptual 12m: Pushing web- scale image-text pre-training to recognize long-tail visual concepts,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Conceptual 12m: Pushing web- scale image-text pre-training to recognize long-tail visual concepts,

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-15T17:51:43.480475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.839683Z digest=sha256:2cad9fbe8bd43060e4a8ebc397f2669743e924604277664ee1e7257619ff6f50

Observation 473a62d3-b1fd-4d9b-b6e9-383c663df6d5 · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality What If We Recaption Billions of Web Images with LLaMA-3?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.843903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.843903Z digest=sha256:c2f2384d4a0774d9ff3531c8a99507df37a9f6debb0696e22401cb0db66bf9f0

Observation 34c429a6-d137-4f20-96d0-0f6be4310713 · outbound

This paper cites Obelics: An open web-scale fil- tered dataset of interleaved image-text documents,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Obelics: An open web-scale fil- tered dataset of interleaved image-text documents,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.741786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.848435Z digest=sha256:01ec1071ec8af006f29e6d07649c09bdb976fe81b3b9be60b3113f80513631c2

Observation 9e104414-5404-4c31-9018-ac5f49acf053 · outbound

This paper cites Bai et al.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Bai et al

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.729804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.852105Z digest=sha256:54c382b6452912d0efad01bdad2591768d8e4c342929707dd630d037205723b4

Observation 0e493f98-bf3c-4af0-b44d-12a222dfb98d · outbound

This paper cites Towards efficient visual-language align- ment of the q-former for visual reasoning tasks,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Towards efficient visual-language align- ment of the q-former for visual reasoning tasks,

Reference 14

Resolution
verified exact
doi, observed 2026-08-15T17:51:43.158871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.861283Z digest=sha256:bf6b7f32e669c79e9bdd3e81cd38999651d7dd2bb02f0ed1544276b9263c4bda

Observation 6e260517-0b60-4f18-a560-8cb159f743b1 · outbound

This paper cites Training language models to fol- low instructions with human feedback,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Training language models to fol- low instructions with human feedback,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.717929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.865159Z digest=sha256:334803bba04422d1ba5c9a259734bb4a4ed59d5450f697e38f1bc3ed4d62de6e

Observation 54f10874-2138-4c03-9c82-14b38eb9cb4b · outbound

This paper cites Let's Go Shopping (LGS) -- Web-Scale Image-Text Dataset for Visual Concept Understanding.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Let's Go Shopping (LGS) -- Web-Scale Image-Text Dataset for Visual Concept Understanding

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:51:43.176356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.856037Z digest=sha256:b5b401cc8135248694d78e876e0ae7704278b830e3a41889a8d2ef006231634d

Observation 2ec930d8-ce85-4785-bf34-a4aa211b036f · outbound

This paper cites Lima: Less is more for alignment,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Lima: Less is more for alignment,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.693170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.872904Z digest=sha256:37030b8483542fb4a0f927b81f8056f02325395bb58b6acd01bed9a879a91526

Observation 00ccbdd9-15cd-48ed-8bc3-9d6e4dc55e88 · outbound

This paper cites Agarwal and D.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Agarwal and D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.680958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.876778Z digest=sha256:2645ee5a26378bddaad5c14c57259e9ca132301e29a05835c23bd96c360f9432

Observation ed77cb0e-c99b-4a29-b816-52dd411f3a29 · outbound

This paper cites Visual instruction tuning,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Visual instruction tuning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.704893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.868943Z digest=sha256:af65f040d684a5f9ddf9df930c4d4ada4474486ec1575dd32fa9db1a7245b968

Observation 9d9c0583-18e0-48b0-a1a4-381266f45c07 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality A Survey on Hallucination in Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.888839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.888839Z digest=sha256:1b2b8f7637126536e15f607a268f4996603bcb35257faf94f37f2ce9c145175f

Observation c5d58d02-0272-421e-b5ab-3d5a4059181f · outbound

This paper cites On the origin of hallucinations in con- versational models: Is it the datasets or the models?.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality On the origin of hallucinations in con- versational models: Is it the datasets or the models?

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.669088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.893005Z digest=sha256:87d24e341e1b148b694eb070a3e280368310f13d55cf515a39677d8583ee5553

Observation 4b5aff2c-9fb3-422d-887f-aa0e71838b34 · outbound

This paper cites 48550 / ARXIV.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality 48550 / ARXIV

Reference 22

Resolution
verified exact
doi, observed 2026-08-15T17:51:43.145209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.880779Z digest=sha256:ca91286066252e61a7c821924c3cfa7beea613692882fa2b141868c65d9b688e

Observation a5b05a76-f4e4-4ee9-8d8f-eba98670ca9a · outbound

This paper cites Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.884682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.884682Z digest=sha256:4c1d62674f04f3ad0aa7e50ac399544a2e098807c8a46b0bb76e0c05eeda848d

Observation 9137c9a2-2555-463d-9124-98d294a67a3f · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language mod- els,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality MiniGPT-4: Enhancing vision-language understanding with advanced large language mod- els,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.644770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.905045Z digest=sha256:6e07e4f610fcf60dd44c288715b5e459f0b6118193daa90af75ee57376dded64

Observation b65c4bde-955a-4669-bced-69f5fb7df324 · outbound

This paper cites Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.632535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.909037Z digest=sha256:64487f39c41176f32b754fb3fe764ff5b73331e9fe7eb55a8471bba5cc03a5b8

Observation f842be4f-5590-4ba2-a228-f4a8114b7da0 · outbound

This paper cites MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.896673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.896673Z digest=sha256:16d48018508cc60cd6e630faa842bc45dff1706b0b424dce5fb41b7c9502888f

Observation 86302a80-cbc6-4715-ac71-616b76a72c2a · outbound

This paper cites Alpagasus: Training a better alpaca with fewer data,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Alpagasus: Training a better alpaca with fewer data,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.656968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.900720Z digest=sha256:cba05d337f462be42db1235242dce31b70e767bd4b706e87cf555653d17723e2

Observation b7cf5628-da59-48db-85ee-ea0ce844e3bc · outbound

This paper cites Language models are unsupervised multitask learners,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Language models are unsupervised multitask learners,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.607479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.920881Z digest=sha256:cddb53ff87c1e4bcc0560550df417f31d25b4e114646a4d0f8903ac812c5bcf0

Observation 4cd92eee-129b-4164-9e4e-b0d6cd0874c2 · outbound

This paper cites A frustratingly simple approach for end-to-end image captioning,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality A frustratingly simple approach for end-to-end image captioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.582723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.928389Z digest=sha256:33aa33daa21c102e9e8dfff229324a405bfb2bb15b622bddb6139aa1384df3ed

Observation 9b75eceb-4c74-4217-bd8a-9a4b8eece9c6 · outbound

This paper cites Learning transferable visual mod- els from natural language supervision,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Learning transferable visual mod- els from natural language supervision,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.620012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.912892Z digest=sha256:3eeafcbc9022fa80b0b7e55df152e8f882b9d94e47cce61491cb105af59f0ceb

Observation 5a46c7e5-1d23-4691-88a8-cc08dfe4f5cf · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.916590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.916590Z digest=sha256:11f7364ad6b2024bf7cdf0bcc81bb64978767abba529b7725f6c87cd794dee45

Observation ec5e046e-f3f5-49b7-8e2c-91e7c739fae6 · outbound

This paper cites Available: https://cdn.openai.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Available: https://cdn.openai

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.595346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.924572Z digest=sha256:ca9f59ab249b54d1643a876877712848bd5047597043b95fa245f39a74634e23

Observation a81406a3-565c-48ea-904b-866e719a8904 · outbound

This paper cites A survey on enhancing image caption- ing with advanced strategies and techniques,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality A survey on enhancing image caption- ing with advanced strategies and techniques,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.569802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.932021Z digest=sha256:778e2cdaf8af07101d2be6077eeb4f94ecd27040011cbcf6a87a9cb7a2e07d8a

Observation 2a3bc4dd-9bb9-48eb-ad34-b5cb2aa1b190 · outbound

This paper cites SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python,.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:51:43.555538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.939404Z digest=sha256:96570bdfbe8142c349cc2879a2ad4eac3edd4e4eef137d3997cd064bfb36a30b

Observation 57702283-bbd4-4e3a-b763-cefa40cf6a70 · outbound

This paper cites 32604 / cmes.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality 32604 / cmes

Reference 1506

Resolution
metadata mismatch
raw_fallback, observed 2026-08-15T17:51:43.399739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.935633Z digest=sha256:c39762c0e0bd8921bb1326230e412bafde06b49fd74cc2add0eb473792ef8bd2

Observation a99d609a-3c26-45d0-afd2-016f4732510b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.801311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.801311Z digest=sha256:f201f3561f8ee0a22b8cffc8657e1c95498f1fcfc67a3ba2b44c7662199c2ca2

Observation bb5a050c-5f47-489a-928f-c215d8208db3 · outbound

This paper cites 48550 / ARXIV.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality 48550 / ARXIV

Reference 2025

Resolution
verified exact
doi, observed 2026-08-15T17:51:43.260921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:51:42.826469Z digest=sha256:3647cd640c7885d67d9c0b31cf340ac6576652505f4c166bbcbb258a518b4b0e

Pith citing papers

No inbound Pith citation observations are available.