Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:51:25.887796Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2501.09672.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:51:25.887796Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8c55e834-d644-4453-8609-4029da6d3356 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e8a8d6-36d6-4fbd-932d-ce30f36737e1 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Scaling laws for generative mixed-modal language models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 965ccf86-cf63-490a-897b-12180214a7c0 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Lawrence Zitnick, Dhruv Batra, and Devi Parikh
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6d0c078c-27ad-407d-a6ed-4e6ece8b1d9c · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Pythia: A suite for analyzing large language models across training and scaling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c1e84e37-393c-4d22-8f32-ba5994008391 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models PaLI: A Jointly-Scaled Multilingual Language-Image Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a69c8f5-4d07-41d0-a6df-3c0be1f202f7 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Reproducible scaling laws for contrastive language-image learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 071960fb-a29d-496e-8816-80874bbf38f9 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models A coefficient of agreement for nominal scales
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988c5c71-494e-4961-830f-8d3f9587a9b7 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c6eb6c26-6232-4cd0-9d38-a5c1ee17599a · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6c988c71-d243-4fc3-9eef-0e37e0e2f402 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 774765d2-7178-4461-83e5-b0083ccf2e89 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Hudson and Christopher D
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8db6cee1-d9a2-4de2-87ea-42797afbda03 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Scaling laws for downstream task performance of large language models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b5fe3cc-cb98-4dd6-aad3-c45194caa9b6 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Large Language Models as Automated Aligners for benchmarking Vision-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef415a1-fb34-4b4f-abe4-43da84becfff · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b7b52ec-d510-4b2a-ae3e-41f8d6522a8a · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e27249bc-83bf-4033-be89-10cf8fb087ee · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Shamma, Michael S
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 467e27ca-f6fd-446c-abb4-d96ef405a73f · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Richard Landis and Gary G
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 295d2df2-2f49-4ca0-9b70-0c7fd212d330 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832ef200-8aa7-41d0-b0c2-4e95a6ee8bcc · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Hashimoto
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 42bc1c55-1abe-433d-827b-4849628fe22f · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f93b1673-c777-4a6a-804c-9f1346b0fbe6 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Lawrence Zitnick, and Piotr Dollár
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f734faa6-9cdb-458f-896e-ed141fc023fb · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Improved Baselines with Visual Instruction Tuning , 2023 a
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2334d8b9-6bf9-41b2-a240-d44b5099203c · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Visual Instruction Tuning , 2023 b
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 02fdbc1d-1582-4b13-afc6-ae23afd06637 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ea1380d3-e1ef-4445-b172-78babf8c9908 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models OCR-VQA: Visual Question Answering by Reading Text in Images
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee15cd68-f4a2-422f-953d-733d0c764825 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Gpt-4v(ision) system card
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4899d78b-0e84-4e0c-bf73-8b28a0560d7b · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Openai model documentation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 85cee05d-d53c-48d8-804c-823981f7e093 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5520d006-a753-4346-baf9-c42044b301bb · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6f9c3a3-9ba9-439d-9ba8-d1b2ac43344c · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Revisiting the train loss: an efficient performance estimator for neural architecture search, 06 2020
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2671ea9d-312b-46c0-9bf5-8dc13fd2d1c5 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Towards vqa models that can read
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20c47a55-3e47-4ebe-9ece-ddbf32f870df · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Generative Multimodal Models are In-Context Learners
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85d87149-19c4-45c4-bece-1d2d06fd2cbb · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Székely, Maria L
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454a078a-a3e9-4704-9f36-cc8df588f1eb · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Style Over Substance: Evaluation Biases for Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 223c606b-c254-4773-86a6-2f50624c24d7 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities , 2023
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7f77c188-126e-456f-9203-a7c136f166a3 · outbound
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Xing, Hao Zhang, Joseph E
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
No inbound Pith citation observations are available.